The Large Memory Model (LM2) enhances standard Transformers by integrating an auxiliary memory module, enabling improved performance in multi-step reasoning and long-context synthesis. Link to article
Gemini beats everyone on new OCR benchmark
This paper presents an open-source benchmark for evaluating Vision-Language Models (VLMs) on Optical Character Recognition (OCR) tasks in dynamic video environments, featuring a dataset of 1,477 annotated frames across various domains.
We built GenAI at Google and Apple, then left to build an open source AI lab
Oumi is an open-source AI lab founded by ex-Google and Apple engineers, aiming to foster collaboration in AI development by making code, weights, and training data accessible to the community.
SWE-agent is the new open-source SOTA on SWE-bench Lite
SWE-agent is an open-source software engineering agent that supports massively parallel runs and cloud-based deployment, enhancing its usability across various models, including local LMs like Qwen and Llama.
Text-to-SQL in Enterprises: Comparing approaches and what worked for us
Fine-tuning open-weight LLMs on business-specific query-SQL pairs achieved 95% accuracy, significantly outperforming previous methods that plateaued at 85% accuracy.
Evaluating RAG for large scale codebases
RAG systems are essential for enhancing generative AI coding assistants, necessitating a robust evaluation framework to ensure accuracy and comprehensiveness in large-scale codebases.
AlignRec Outperforms SOTA Models in Multimodal Recommendations
AlignRec significantly enhances multimodal recommendation systems by optimizing three alignment tasks: inter-content (ICA), content-category (CCA), and user-item (UIA), effectively bridging semantic gaps between diverse content types.
Diffusion Without Tears
Notion serves as a versatile workspace, integrating notes, tasks, wikis, and databases into a single platform, enhancing productivity and organization for users. For a deeper understanding of its capabilities, refer to the article.
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
The study introduces a novel language model architecture that enhances test-time computation by leveraging latent reasoning through a recurrent block, allowing for arbitrary depth unrolling without the need for specialized training data.
CoT-Valve introduces a novel method for dynamically controlling reasoning chain lengths, allowing models to adapt their inference costs based on task difficulty, thus enhancing efficiency without sacrificing performance.
Fixing Open LLM Leaderboard with Math-Verify
The introduction of Math-Verify has enabled a comprehensive re-evaluation of 3,751 models on the Open LLM Leaderboard, correcting previous inaccuracies in math performance assessments.
Advancements in Embedding-Based Retrieval at Pinterest Homefeed
Pinterest's embedding-based retrieval has evolved with advanced feature crossing and ID embeddings, enhancing user engagement metrics by up to 1.2% through innovative model architectures like MaskNet and DHEN.
LP-LM: No Hallucinations in Question Answering with Logic Programming
LP-LM eliminates hallucinations in question answering by grounding responses in a knowledge base (KB), ensuring reliability through semantic parsing in Prolog.
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
CopySpec is a novel technique that enhances LLM efficiency by identifying repeated sequences in chat history, allowing for speculative copying without sacrificing output quality or increasing GPU memory usage.
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Self-Consistent Internal Rewards (SCIR) enhances the reliability of internal reward models in Self-Rewarding Language Models, addressing inconsistencies that hinder alignment with human preferences.