# Feb 14, 2025

## Daily

### LM2: Large Memory Models
- The **Large Memory Model (LM2)** enhances standard Transformers by integrating an **auxiliary memory module**, enabling improved performance in **multi-step reasoning** and **long-context synthesis**. [Link to article](https://arxiv.org/abs/2502.06049)

### Gemini beats everyone on new OCR benchmark
- This paper presents an **open-source benchmark** for evaluating **Vision-Language Models (VLMs)** on **Optical Character Recognition (OCR)** tasks in dynamic video environments, featuring a dataset of **1,477 annotated frames** across various domains.

### We built GenAI at Google and Apple, then left to build an open source AI lab
- **Oumi** is an open-source AI lab founded by ex-Google and Apple engineers, aiming to foster **collaboration** in AI development by making **code, weights, and training data** accessible to the community.

### SWE-agent is the new open-source SOTA on SWE-bench Lite
- **SWE-agent** is an open-source software engineering agent that supports **massively parallel runs** and **cloud-based deployment**, enhancing its usability across various models, including local LMs like Qwen and Llama.

### Text-to-SQL in Enterprises: Comparing approaches and what worked for us
- **Fine-tuning open-weight LLMs** on business-specific query-SQL pairs achieved **95% accuracy**, significantly outperforming previous methods that plateaued at **85%** accuracy.

### Evaluating RAG for large scale codebases
- **RAG systems** are essential for enhancing generative AI coding assistants, necessitating a **robust evaluation framework** to ensure accuracy and comprehensiveness in large-scale codebases.

### AlignRec Outperforms SOTA Models in Multimodal Recommendations
- **AlignRec** significantly enhances multimodal recommendation systems by optimizing three alignment tasks: **inter-content (ICA)**, **content-category (CCA)**, and **user-item (UIA)**, effectively bridging semantic gaps between diverse content types.

### Diffusion Without Tears
- **Notion** serves as a versatile workspace, integrating **notes**, **tasks**, **wikis**, and **databases** into a single platform, enhancing productivity and organization for users. For a deeper understanding of its capabilities, refer to the [article](https://baincapitalventures.notion.site/Diffusion-Without-Tears-14e1469584c180deb0a9ed9aa6ff7a4c).

### Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- The study introduces a **novel language model architecture** that enhances test-time computation by leveraging **latent reasoning** through a recurrent block, allowing for **arbitrary depth unrolling** without the need for specialized training data.

### CoT-Valve: Length-Compressible Chain-of-Thought Tuning
- **CoT-Valve** introduces a novel method for **dynamically controlling reasoning chain lengths**, allowing models to adapt their inference costs based on task difficulty, thus enhancing efficiency without sacrificing performance.

### Fixing Open LLM Leaderboard with Math-Verify
- The introduction of **Math-Verify** has enabled a comprehensive re-evaluation of **3,751 models** on the Open LLM Leaderboard, correcting previous inaccuracies in math performance assessments.

### Advancements in Embedding-Based Retrieval at Pinterest Homefeed
- **Pinterest's embedding-based retrieval** has evolved with advanced feature crossing and ID embeddings, enhancing user engagement metrics by up to **1.2%** through innovative model architectures like MaskNet and DHEN.

### LP-LM: No Hallucinations in Question Answering with Logic Programming
- **LP-LM** eliminates hallucinations in question answering by grounding responses in a **knowledge base (KB)**, ensuring reliability through **semantic parsing in Prolog**.

### CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
- **CopySpec** is a novel technique that enhances **LLM efficiency** by identifying repeated sequences in chat history, allowing for **speculative copying** without sacrificing output quality or increasing GPU memory usage.

### Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
- **Self-Consistent Internal Rewards (SCIR)** enhances the reliability of internal reward models in **Self-Rewarding Language Models**, addressing inconsistencies that hinder alignment with human preferences.
