# Feb 28, 2025

### Fire-Flyer File System from DeepSeek
- The **Fire-Flyer File System (3FS)** is engineered for **high-performance distributed file storage**, optimizing AI workloads through **disaggregated architecture** and **strong consistency** via Chain Replication with Apportioned Queries (CRAQ).

### Putting Andrew Ng's OCR models to the test
- **Andrew Ng's new OCR models** exhibit significant flaws, including **over 50% hallucinated values** and **30+ second processing times**, raising concerns for industries reliant on accurate data extraction.

### [R] Beyond Dot Products: Retrieval with Learned Similarities
- **New approach**: The paper introduces **Mixture of Logits (MoL)**, a method that enables **learned similarity functions**, surpassing traditional dot product methods in efficiency and effectiveness for **recommendation systems** and **question answering**.

### [R] Training-free Chroma Key Content Generation Diffusion Model
- The **TKG-DM** model enables **training-free** generation of foreground objects on chroma key backgrounds, utilizing any pre-trained diffusion model without the need for fine-tuning.

### Merlion: A Machine Learning Framework for Time Series Intelligence
- **Merlion** is a comprehensive **Python library** designed for **time series intelligence**, offering features like anomaly detection, forecasting, and change point detection, all unified under a single interface.

### [R] Belief State Transformers
- The **Belief State Transformer** introduces a dual-input mechanism that predicts both the next token for a prefix and the previous token for a suffix, enhancing performance in complex tasks where traditional transformers falter. [Link to article](https://arxiv.org/abs/2410.23506)

### [R] Dynamic Vocabulary Curriculum Learning Improves LLM Pre-training Efficiency
- **Dynamic vocabulary curriculum learning** enhances LLM pre-training efficiency by starting with a **smaller vocabulary** (~5k tokens) and expanding to full size (~50k) based on model convergence metrics, leading to a **25% reduction in training time** without sacrificing quality.

### [R] FFTNet: Linear-Time Global Token Mixing via Adaptive Spectral Filtering
- **FFTNet** replaces **quadratic complexity** self-attention with **linear complexity** using Fast Fourier Transforms, enabling efficient global token mixing while preserving performance.

### [R] Dynamic Planning induction in Large Language Models
- **DyPlan** introduces a **dynamic strategy selection** process in Large Language Models (LLMs), enhancing their ability to answer queries by adapting to the specific context of each question.

### Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
- **Multi-Agent Verification (MAV)** enhances large language models (LLMs) by utilizing multiple verifiers to evaluate outputs, leading to improved performance without additional training.

### Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
- **Distilled Mamba models** outperform traditional Transformers in **mathematical reasoning** by leveraging faster inference speeds, achieving better accuracy under fixed computational budgets despite a slight drop in zero-shot performance.

### Collaborative Stance Detection via Small-Large Language Model Consistency Verification
- The **CoVer framework** enhances stance detection by leveraging **Small-Large Language Model consistency verification**, allowing for efficient batch processing and logical checks to improve accuracy in social media monitoring.

### Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
- **Meta-Reasoner** enhances Large Language Models (LLMs) by enabling them to **optimize inference-time reasoning**, reducing computational overhead and mitigating error propagation through strategic guidance.
