ML Times

Feb 14, 2025

Daily

LM2: Large Memory Models

Gemini beats everyone on new OCR benchmark

We built GenAI at Google and Apple, then left to build an open source AI lab

SWE-agent is the new open-source SOTA on SWE-bench Lite

Text-to-SQL in Enterprises: Comparing approaches and what worked for us

Evaluating RAG for large scale codebases

AlignRec Outperforms SOTA Models in Multimodal Recommendations

Diffusion Without Tears

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

CoT-Valve: Length-Compressible Chain-of-Thought Tuning

Fixing Open LLM Leaderboard with Math-Verify

Advancements in Embedding-Based Retrieval at Pinterest Homefeed

LP-LM: No Hallucinations in Question Answering with Logic Programming

CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models