# Main Content

## DisTrO – a family of low latency distributed optimizers

- **DisTrO** significantly **reduces inter-GPU communication** requirements, enhancing distributed training efficiency by **three to four orders of magnitude**.

## Cerebras Inference: AI at Instant Speed

- **Cerebras introduces the world's fastest AI inference solution**, delivering up to **1,800 tokens per second** for Llama3.1 8B and **450 tokens per second** for Llama3.1 70B, outpacing NVIDIA GPU-based solutions by **20x**.

## Tinybox of tinygrad by George Hotz is finally entering production

- **Tinyboxes** now features a **"buy it now" button**, 18 months after the company's inception, with **13 units** available for immediate purchase.

## Splatt3R: Zero-Shot Gaussian Splatting from Uncalibrated Image Pairs

- **Splatt3R** introduces a **pose-free, feed-forward method** for 3D reconstruction and novel view synthesis from **uncalibrated stereo pairs**, predicting 3D Gaussian Splats without needing camera parameters or depth information.

## AI predicts earthquakes with unprecedented accuracy

- **The University of Texas developed an AI that predicted 70% of earthquakes** in a trial in China, showcasing its potential to enhance earthquake preparedness and risk management.

## Many FDA-approved AI medical devices are not trained on real patient data

- **Approximately 43% of FDA-approved AI medical devices lack published clinical validation data**, indicating a significant gap in the evidence supporting their effectiveness and safety.

## Cerebras launches inference for Llama 3.1; benchmarked at 1846 tokens/s on 8B

- **Cerebras** has achieved a **new AI inference speed record**, serving **Llama 3.1 8B** at **1,850 output tokens/s** and **70B** at **446 output tokens/s**, marking a significant advancement in AI processing capabilities.

## Launch HN: Bucket Robotics (YC S24) – Defect detection for molded and cast parts

- **Bucket Robotics** introduces a **novel approach to defect detection** in manufacturing by transforming CAD models into custom defect detection models, aiming to **improve upon the traditional 80% human success rate** and **manual or traditional ML-based methods** that rely on real-world sample imaging.

## [R] DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

- **DeepSeek-Prover-V1.5** enhances theorem proving capabilities in Lean 4 by **optimizing training and inference**, leveraging **reinforcement learning from proof assistant feedback** and a **Monte-Carlo tree search variant** named RMaxTS for diverse proof path generation.

## [D] Image segmentation converges to all zeros when masks are too small

- **Image segmentation models**, such as **UNet**, **struggle with small masks**, leading to convergence on all-zero outputs, particularly in **breast cancer image segmentation** tasks.

## [D] Exploring the Potential of Edge Computing/Federated Learning in Continuous Training for GPT/LLMs

- **Exploring the potential of Edge Computing and Federated Learning** could revolutionize the way **GPT and Large Language Models (LLMs)** are continuously trained, by decentralizing the process.

## PyRoboCOP: Python-Based Robotic Control and Optimization Package

- **PyRoboCOP** is a **Python-based package** designed for **robotic control and optimization**, specifically addressing **manipulation tasks and collision avoidance** through handling **complex complementarity constraints**.

## [D] Trackers like SAM2 but faster

- **Combining RT-DETR with SAM2** enhances object tracking capabilities, **outperforming RT-DETR + DeepSort** in tracking tiny vehicles in low-quality surveillance videos.

## TwoMinutePapers - DeepMind’s New AI Looked At 100,000,000 Examples!

## NVIDIA Launches Array of New CUDA Libraries to Expand Accelerated Computing and Deliver Order-of-Magnitude Speedup to Science and Industrial Applications

- **NVIDIA's new CUDA libraries** significantly **speed up** and **reduce energy consumption** in diverse fields such as data processing, AI, and 6G research, by leveraging accelerated computing.
