# Jan 24, 2025

## Daily

- **Rapid progress in AI**: Anthropic's latest model, Sonnet 3.5, achieved **50%** on the SWE-bench, a significant leap from **3%** at the beginning of 2024, indicating a swift approach to human-level capabilities in AI.

- **Humanity’s Last Exam**, a new benchmark by Scale AI and CAIS, reveals that current AI models answered fewer than **10%** of expert-level questions correctly, indicating significant room for improvement in AI reasoning capabilities.

- **ColBERT** enhances vector search by utilizing **token-level multi-vectors**, which preserves fine-grained details and improves search accuracy over traditional sentence embeddings.

- **Energy-based diffusion models** enhance **text generation** by effectively merging diffusion techniques with energy-based modeling, tackling the complexities of discrete generative tasks.

- **Monte Carlo Tree Search (MCTS)** is innovatively applied to enable language models to **self-reflect**, enhancing their ability to evaluate and refine responses through structured self-criticism.

- This paper presents a **novel orderless compression technique** for vector IDs in approximate nearest neighbor (ANN) search systems, achieving a **70% compression ratio** while maintaining search accuracy.

- **Vision support has been integrated into smolagents**, enabling them to utilize vision language models (VLMs) for enhanced web browsing capabilities, allowing agents to interpret visual content alongside text.

- **AI has revolutionized the mapping of Titan's methane clouds**, enabling researchers to analyze years of Cassini data in mere seconds using NVIDIA GPUs, thus enhancing productivity in planetary science.

- **K-COMP** enhances retrieval-augmented question answering by integrating **prior knowledge** into the compression of retrieved passages, improving accuracy in medical domain QA.

- The **self-referencing causal cycle (RECALL)** enhances large language models (LLMs) by allowing them to overcome the **reversal curse**, which typically hinders their ability to recall prior context in sequential data.

- The **Softplus Attention** mechanism, enhanced with a **re-weighting** strategy, significantly improves **length extrapolation** in large language models, outperforming traditional Softmax attention in various inference scenarios.

- **ReLLaX** (Retrieval-enhanced Large Language models Plus) addresses the **lifelong sequential behavior incomprehension** problem in LLMs for recommendations by optimizing data, prompts, and parameters, enhancing their ability to extract key information from long user behavior sequences.

- **Intel's AI Playground**, powered by **PyTorch**, showcases advanced **Generative AI** workloads, integrating features like image generation and chatbots, all optimized for **Intel Arc GPUs**.
