ML Times
Things we learned about LLMs in 2024
In 2024, the GPT-4 barrier was decisively surpassed, with 18 organizations now boasting models that outperform it, including Google's Gemini 1.5 Pro, which introduced unprecedented capabilities like a 2 million token input context and video input.
Deepseek's R1 model has outperformed OpenAI’s o1 on reasoning benchmarks, showcasing its potential as a formidable player in the AI landscape, driven by a commitment to foundational technology and open-source principles.
Coconut introduces a novel reasoning method for LLMs, allowing them to operate in a continuous latent space, enhancing their problem-solving capabilities beyond traditional word-based reasoning.
Nvidia is investing in robotics, launching its Jetson Thor computers in 2025, aiming to lead the anticipated robotics revolution with a comprehensive "full stack" solution for AI-powered robots.
Putnam-AXIOM introduces a new benchmark for evaluating higher-level mathematical reasoning in large language models (LLMs), highlighting significant performance gaps and the effects of data contamination.
The Large Concept Models (LCM) utilize a higher-level semantic representation called "concepts," enabling language- and modality-agnostic sentence modeling across 200 languages in text and 57 languages in speech, leveraging the SONAR embedding space.
The RT-2 model demonstrates advanced capabilities in vision-language-action tasks, successfully executing commands like "pick up the extinct animal" through integrated perception and action.
Submerging electrical circuits in nonconductive liquid coolant can capture 100% of generated heat, potentially revolutionizing data-center design by eliminating the need for traditional cooling systems.
A new state-of-the-art (SOTA) Text to Audio model utilizes rectified flow matching and FLUX architecture, achieving rapid inference times of about 3 seconds on a GPU.
The current SOTA for biomedical encoder models includes PubMedBERT and BioBERT, which serve as strong baselines for tasks like sentence similarity and document retrieval.
Multihead latent attention compresses the entire sequence into a latent vector, allowing the model to access information from the whole sequence, which disrupts traditional autoregressive constraints.
CNN with regression heads is proposed for predicting absolute clearance distances without reference objects, addressing the challenge of varying object dimensions and camera intrinsics.
Multimodal models, particularly Vision-Language Models (VLMs), are increasingly utilized for content moderation, addressing challenges like detecting hateful memes and violence in images.