TL;DR of Deep Dive into LLMs Like ChatGPT by Andrej Karpathy
Andrej Karpathy's deep dive into LLMs like ChatGPT reveals that effective training involves extensive data filtering and tokenization, with models like GPT-4 utilizing 100,277 tokens for optimal performance.
LIMO: Less Is More for Reasoning
LIMO challenges the belief that extensive training data is essential for complex reasoning, achieving 57.1% accuracy on AIME and 94.8% on MATH with only 817 samples, significantly outperforming previous models.
Three Observations
AGI's potential is framed as a transformative tool that could enable everyone to achieve more than today's most impactful individuals, fundamentally reshaping productivity and creativity across society.
Undergraduate Upends a 40-Year-Old Data Science Conjecture
Undergraduate Andrew Krapivin and colleagues have developed a new type of hash table that significantly speeds up data searches, contradicting a 40-year-old conjecture about their efficiency.
Scaling Up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
The study introduces a novel language model architecture that enhances test-time computation by leveraging latent reasoning through a recurrent depth approach, allowing for arbitrary depth unrolling without the need for specialized training data. Link to article
AI-designed proteins utilizing AlphaFold 2 and RFdiffusion effectively bind to and neutralize cytotoxins in cobra venom, presenting a novel approach to antivenom development.
[R] LIMO: Less is More for Reasoning
LIMO reveals that complex reasoning in large language models can be achieved with as few as 817 training samples, defying the belief that extensive data is essential for sophisticated tasks.
CAPTCHAs: 'a tracking cookie farm for profit masquerading as a security service'
A 2023 study from UC Irvine reveals that CAPTCHAs serve as a tracking cookie farm, generating nearly $1 trillion for Google while wasting 819 million hours of user time on tasks like identifying traffic lights.
Explainable Linear Programs
Explainable Linear Programs enhance understanding of complex models by allowing interactive modifications, which aids in debugging and refining supply chain optimization processes.
What Happens to SaaS in a World with Computer Using Agents?
AI agents are transforming SaaS by automating interactions, making traditional user interfaces obsolete as agents handle tasks directly, thus shifting the value proposition from user experience to backend functionality.
[D] KL divergence as a primary reward in LLM post-training RL?
KL divergence as a reward model in RL for LLMs could yield sequences with significantly lower KL divergence than those generated by the pretrained model, potentially enhancing coherence in outputs.
[R] 3D Point Regularization for Physics-Aware Video Generation
This work presents a 3D point cloud regularization method that enhances physical realism in video generation by constraining outputs with learned 3D point trajectories, akin to motion capture techniques.
[R] Multi-View Scene Completion Using Latent Diffusion Transformers for Uncalibrated Image Sets
This work introduces a transformer-based method for completing missing regions in multi-view scenes, achieving a 30% improvement in visual quality compared to previous techniques, particularly with uncalibrated casual photos.
Flexible and Efficient Grammar-Constrained Decoding
New GCD algorithm significantly enhances efficiency by achieving 17.71x faster offline preprocessing compared to existing methods, ensuring structured outputs from LLMs adhere to specified syntactic rules.
ChallengeMe: An Adversarial Learning-enabled Text Summarization Framework
ChallengeMe introduces an adversarial learning-based prompt framework that enhances text summarization by addressing issues like hallucination and lack of specificity in large language models (LLMs).