# May 23, 2024

## Nvidia announces financial results for first quarter fiscal 2025
- **NVIDIA** announced its **financial results for the first quarter of fiscal 2025**, highlighting significant growth and strategic advancements.

## Show HN: Route your prompts to the best LLM
- **Unify's chat interface** is under development for mobile, aiming to **route prompts to the most suitable Large Language Models (LLMs)** for optimized responses.

## [D] AI Agents: too early, too expensive, too unreliable
- **Autonomous agent-based LLM workflows** are underperforming against the hype, with the best models achieving only a **35.8% success rate** on the [WebArena leaderboard](https://docs.google.com/spreadsheets/d/1M801lEpBbKSNwP-vDBkC_pF7LdyGU1f_ufZb_NWNBZQ/edit#gid=0).

## [P] Fish Speech TTS: clone OpenAI TTS in 30 minutes
- **Fish Speech TTS** successfully **cloned OpenAI's TTS** performance using **supervised fine-tuning** on 10 hours of OpenAI TTS data, achieving similar emotion, rhythm, accent, and timbre.

## [P] ReproModel: Open Source ML Research Toolbox.
- **ReproModel** is an **open-source, no-code toolbox** designed to streamline the process of testing and reproducing **ML models**, addressing the significant time investment required to replicate studies from existing research.

## Show HN: Open-Source Real Time Data Framework for LLM Applications
- **Indexify** is an **open-source data framework** designed for building **data-intensive LLM applications**, featuring a **real-time extraction engine** and **pre-built extraction adapters** for **fast, reliable, and precise AI application development**.

## [Research] How Can Understanding Sparse Autoencoders in Claude 3 Sonnet Influence Practical AI Applications?
- The **Anthropic study** on **sparse autoencoders** in the **Claude 3 Sonnet** demonstrates their ability to **extract interpretable, multilingual, and multimodal features** from transformer models, enhancing **AI's understanding** of complex data. [Read the paper](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)

## [R] Introducing SSAMBA: The Self-Supervised Audio Mamba!
- **SSAMBA**, a **self-supervised state-space model**, outperforms or matches transformer-based models in **audio tasks** without relying on attention mechanisms.

## [R] Geometry of data for ML
- **Diffusion Geometry** defines the **geometry of data** as a **powerful method**, surpassing traditional approaches like **persistent homology** in effectiveness. [Read more](http://arxiv.org/abs/2405.10858)

## TwoMinutePapers - 50,000,000 Point Bouncy Jelly Simulation!
- The **50,000,000 point bouncy jelly simulation** showcases **elastic bodies**, like squishy balls and armadillos, interacting in complex ways, demonstrating **advanced computational physics**.

## GigaPath: Whole-Slide Foundation Model for Digital Pathology
- **GigaPath**, developed by Microsoft in collaboration with Providence Health System and the University of Washington, is a **novel vision transformer** designed for whole-slide digital pathology, leveraging **dilated self-attention** to manage the computational challenges posed by gigapixel images.

## YannicKilcher - [ML News] OpenAI is in hot waters (GPT-4o, Ilya Leaving, Scarlett Johansson legal action)
- **OpenAI's GPT-4o**, an advancement over previous models, **natively processes text, images, and voice** in real-time, offering a more unified and human-like interaction experience.

## Into the Omniverse: SoftServe and Continental Drive Digitalization With OpenUSD and Generative AI
- **SoftServe and Continental** have developed the **Industrial Co-Pilot**, a **virtual agent powered by generative AI**, to streamline maintenance workflows in manufacturing, enhancing productivity and reducing downtime.
