# Jun 25, 2025

- **Federal judge rules copyrighted books are fair use for AI training**  
  A **federal judge** ruled that AI developers can train models on copyrighted books without consent, establishing a significant precedent for **fair use** in generative AI contexts.

- **Gemini Robotics On-Device brings AI to local robotic devices**  
  **Gemini Robotics On-Device** introduces a powerful **on-device VLA model** that enhances robotic dexterity and task adaptation, operating independently of data networks for improved latency and robustness in various environments.

- **The bitter lesson is coming for tokenization**  
  **Tokenization's fragility** in LLMs is a significant bottleneck, with the potential to be replaced by more efficient methods like the **Byte Latent Transformer (BLT)**, which aims to leverage byte-level modeling for improved performance and scalability.

- **MCP is eating the world**  
  **MCP (Model Context Protocol) is gaining traction due to its simplicity and timely execution**, enabling the creation of agents and workflows atop LLMs, unlike previous complex attempts that faltered under integration challenges.

- **AlphaGenome: AI for better understanding the genome**  
  **AlphaGenome** is a groundbreaking AI tool that predicts the effects of **genetic variants** on biological processes, utilizing long DNA sequences of up to **1 million base pairs** for high-resolution insights into gene regulation.

- **XBOW, an autonomous penetration tester, has reached the top spot on HackerOne**  
  **XBOW** has achieved a historic milestone as the first autonomous penetration tester to reach the **top spot** on the US HackerOne leaderboard, demonstrating its capability to discover vulnerabilities in real-world environments.

- **Bot or Human? Creating the Invisible Turing Test for the Internet**  
  **AI systems exhibit detectable behavioral signatures** that can enhance bot detection, as demonstrated by Roundtable's [Proof-of-Human API](https://www.roundtable.ai/), which verifies human presence without intrusive methods.

- **PhD (non-US) → Research Scientist jobs in CV/DL at top companies—how much DSA grind is essential?**  
  **Research Scientist roles** at top tech companies increasingly prioritize **DSA skills**, with algorithm interviews becoming a significant hurdle despite strong publication records in CV/DL fields.

- **Extremely low(<0.2) train/val loss after 1.96 billion tokens when pretraining GPT-2 small**  
  **Extremely low train/val loss (<0.2)** was achieved after processing **1.96 billion tokens** during the pretraining of GPT-2 small, utilizing advanced techniques like **RoPE positional embeddings** and **SwiGLU MLP layers**.

- **Advanced Python Function Debugging with MCP Integration**  
  **Gnosis Mystic** enables **AI assistants** to interact with Python functions through **runtime hijacking**, allowing real-time analysis and control with minimal code changes, enhancing development efficiency.

- **OMEGA: Can LLMs Reason Outside the Box in Math?**  
  **OMEGA** explores whether **LLMs** can demonstrate **reasoning** capabilities in **mathematics** beyond conventional methods, revealing significant limitations in compositional generalization.

- **Applying COCONUT continuous reasoning into a learnt linear layer that produces sampling parameters (temp, top-k, top-p, etc.) for the current token?**  
  **Dynamic modulation of sampling parameters** in LLMs could enhance creativity and precision by allowing models to learn when to adjust **temperature**, **top-p**, and **top-k** during token generation, rather than relying on static values.

- **KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality**  
  **KnowRL** introduces a **factuality reward** mechanism in Reinforcement Learning (RL) to combat hallucinations in slow-thinking models, enhancing their ability to recognize knowledge boundaries during reasoning.

- **Presenting Flux Fast: Making Flux go brrr on H100s**  
  **Flux** has achieved a **~2.5x speedup** on H100 GPUs through optimizations primarily using **native PyTorch code**, enhancing its performance as a leading open-weight model in image generation.

- **HPE and NVIDIA Debut AI Factory Stack to Power Next Industrial Shift**  
  **HPE and NVIDIA's new AI Factory Stack** aims to accelerate enterprise AI adoption with innovations like the **NVIDIA Blackwell architecture** and **HPE's RTX PRO Servers**, providing a comprehensive framework for generative and industrial AI applications.
