# Jun 18, 2025  
## ML Times  
### AI is eating our brains. MIT study: Your brain on ChatGPT  
- The study reveals that using **Large Language Models (LLMs)** like ChatGPT in educational settings may lead to a **cognitive cost**, as participants using LLMs exhibited weaker neural connectivity and lower engagement compared to those relying solely on their brains or search engines.

### MiniMax-M1 open-weight, large-scale hybrid-attention reasoning model  
- **MiniMax-M1** is the **first open-weight hybrid-attention reasoning model**, featuring a **Mixture-of-Experts architecture** and a **lightning attention mechanism**, enabling it to handle **1 million tokens** and outperform competitors in complex tasks.

### Making 2.5 Flash and 2.5 Pro GA, and introducing Gemini 2.5 Flash-Lite  
- **Gemini 2.5** now includes **Flash** and **Pro** models, which are generally available, alongside the introduction of **2.5 Flash-Lite**, the fastest and most cost-efficient model yet.

### AMD's CDNA 4 Architecture Announcement  
- **AMD’s CDNA 4 architecture** enhances matrix multiplication performance for machine learning by optimizing lower precision data types, while maintaining a competitive edge in vector operations through a familiar chiplet design.

### Timescale Is Now TigerData  
- **TigerData**, formerly Timescale, has evolved from a time-series database to a modern PostgreSQL solution, catering to the needs of over **2,000 customers** and achieving **mid 8-digit ARR** with significant year-over-year growth.

### Andrej Karpathy's YC AI SUS talk on the future of the industry  
- **Andrej Karpathy** emphasizes a transformative era in software development, coining **Software 3.0** as a paradigm shift where neural networks, particularly large language models (LLMs), are now programmable in natural language, making programming more accessible to everyone.

### LLMs pose an interesting problem for DSL designers  
- **LLMs challenge traditional DSL design** by raising the question of necessity; if LLMs can generate code efficiently, why invest in creating domain-specific languages that eliminate boilerplate?

### Is There a Half-Life for the Success Rates of AI Agents?  
- **AI agents exhibit an exponentially declining success rate** on longer tasks, characterized by a unique half-life, suggesting that as task duration increases, the likelihood of failure rises due to the accumulation of subtasks, as shown in Kwa et al. (2025) [arXiv](https://doi.org/10.48550/arXiv.2503.14499).

### The Unreasonable Effectiveness of Fuzzing for Porting Programs  
- **Fuzzing with LLMs** has proven effective for automating the porting of programs from **C to Rust**, allowing for significant reductions in manual coding effort and potential errors during the process.

### Time Series Forecasting with Graph Transformers  
- **Graph Transformers enhance time series forecasting** by leveraging interconnected data from relational databases, allowing for more accurate predictions through a comprehensive end-to-end pipeline that integrates various signals from the entire graph structure.

### Real-time action chunking with large models  
- **Real-time chunking (RTC)** enables robots to execute actions without delays or discontinuities, significantly improving performance in dynamic tasks like striking a match or plugging in an Ethernet cable, even under high latency conditions exceeding 300 milliseconds.

### [D] CausalML : Causal Machine Learning  
- **CausalML** is a significant advancement in **Causal Machine Learning**, offering a comprehensive **140-page survey** that explores its methodologies and applications. [Read the survey here](https://arxiv.org/abs/2206.15475).

### Revisiting Minsky's Society of Mind in 2025  
- **Minsky’s vision of intelligence as a modular society of agents** is gaining traction in AI development, as researchers pivot from monolithic models to **multi-agent systems** that enhance robustness and scalability.

### Reasoning by Superposition: A Perspective on Chain of Continuous Thought  
- **Continuous chain-of-thought (CoT) techniques** in Large Language Models (LLMs) outperform discrete methods by enabling **parallel breadth-first search**, solving directed graph reachability with **D steps**, where D is the graph's diameter, compared to **O(n²)** steps for discrete CoTs.

### [R] Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons  
- The proposed **non-attention based architecture** for large language models (LLMs) can efficiently manage **context windows** of **hundreds of thousands to millions of tokens**, overcoming the limitations of traditional Transformer designs that face **quadratic memory and computation overload**.
