# Jan 23, 2025

## Stargate Project: SoftBank, OpenAI, Oracle, MGX to build data centers

- **Trump announces a partnership** involving OpenAI, Oracle, and SoftBank to invest **up to $500 billion** in AI infrastructure, with initial funding of **$100 billion** aimed at building data centers in Texas, reflecting a significant commitment to advancing AI technology in the U.S.

## Lossless Compression of Vector IDs for Approximate Nearest Neighbor Search

- **Lossless compression schemes** for vector IDs can achieve a **7x reduction** in storage size without sacrificing accuracy or search runtime, significantly optimizing approximate nearest neighbor search.

## How to solve computational science problems with AI: PINNs

- **Physics-Informed Neural Networks (PINNs)** effectively integrate **physical laws** into neural network training, enabling the solution of complex **partial differential equations (PDEs)** with improved accuracy and efficiency.

## [R] Learning to Continually Learn with the Bayesian Principle

- **Novel meta-continual learning framework** combines the **representational power of neural networks** with the **robustness of Bayesian models**, effectively preventing catastrophic forgetting during continual learning.

## Scale AI Unveil Results of Humanity's Last Exam, a Groundbreaking New Benchmark

- **Humanity’s Last Exam**, a new benchmark by Scale AI and CAIS, reveals that current AI models answered fewer than **10%** of expert-level questions correctly, indicating significant room for improvement in AI reasoning capabilities.

## DeepSeek and the Effects of GPU Export Controls

- **DeepSeek's V3 model**, trained on **2,048 H800 GPUs**, claims to match or exceed benchmarks set by **GPT-4** and **Claude**, demonstrating that efficiency can rival sheer computational power.

## Kimi k1.5: Scaling Reinforcement Learning with LLMs

- **Kimi k1.5** leverages **reinforcement learning (RL)** to enhance large language models (LLMs), achieving **state-of-the-art reasoning performance** across various benchmarks, including **77.5 on AIME** and **96.2 on MATH 500**.

## Autonomy-of-Experts Models

- **Autonomy-of-Experts (AoE)** enhances **Mixture-of-Experts (MoE)** models by allowing experts to autonomously select themselves for processing, improving both **expert selection** and **learning efficiency**.

## [R] ENERGY-BASED DIFFUSION LANGUAGE MODELS FOR TEXT GENERATION

- **Energy-based diffusion models** enhance **text generation** by effectively merging diffusion techniques with energy-based modeling, tackling the complexities of discrete generative tasks.

## Into the Omniverse: OpenUSD Workflows Advance Physical AI for Robotics, Autonomous Vehicles

- **OpenUSD workflows** are enhancing **physical AI** capabilities, enabling robots and autonomous vehicles to understand and interact with the real world through advanced simulation environments that replicate physical dynamics and spatial relationships.

## NExtLong: Toward Effective Long-Context Training without Long Documents

- **NExtLong** introduces a **novel framework** for synthesizing long-context data by utilizing **Negative document Extension**, which enhances long-range dependency modeling in large language models (LLMs) through the use of **hard negative distractors**.

## Architectural Fusion Through Contextual Partitioning in Large Language Models: A Novel Approach to Parameterized Knowledge Integration

- **Contextual Partitioning** revolutionizes large language models by dynamically segmenting parameters into context-aware regions, enhancing **task-specific specialization** without the need for external fine-tuning.
