# Aug 28, 2024

- **Sapiens:** Foundation for Human Vision Models  
  **Sapiens** models, designed for **human-centric vision tasks** like 2D pose estimation and depth estimation, leverage **self-supervised pretraining** on over 300 million human images for enhanced performance.

- **Diffusion Models Are Real-Time Game Engines**  
  **GameNGen**, the **first game engine powered entirely by a neural model**, enables real-time interaction with complex environments, simulating the classic game **DOOM** at over **20 frames per second** on a single TPU.

- **DisTrO – a family of low latency distributed optimizers**  
  **DisTrO** significantly **reduces inter-GPU communication** requirements, enhancing distributed training efficiency by **three to four orders of magnitude**.

- **[R] Playable 20FPS Doom via a finetuned SD1.4 model from Google research team**  
  **GameNGen**, the **first game engine** powered by a neural model, enables **real-time interaction** with complex environments, simulating DOOM at over **20 frames per second** on a single TPU.

- **Tesla's TTPoE at Hot Chips 2024: Replacing TCP for Low Latency Applications**  
  **Tesla's TTPoE protocol** aims to **replace TCP** for low latency applications by **simplifying state machines** and **removing wait states**, targeting microsecond scale latencies.

- **Splatt3R: Zero-Shot Gaussian Splatting from Uncalibrated Image Pairs**  
  **Splatt3R** introduces a **pose-free, feed-forward method** for 3D reconstruction and novel view synthesis from **uncalibrated stereo pairs**, predicting 3D Gaussian Splats without needing camera parameters or depth information.

- **"Writing in the Margins (WiM)" - a better inference pattern for long context LLMs that solves the Lost-in-the-Middle problem**  
  **Writing in the Margins (WiM)** introduces a **new inference pattern** for Large Language Models, optimizing long input sequence handling through **chunked prefill** and segment-wise inference.

- **Launch HN: Bucket Robotics (YC S24) – Defect detection for molded and cast parts**  
  **Bucket Robotics** introduces a **novel approach to defect detection** in manufacturing by transforming CAD models into custom defect detection models, aiming to **improve upon the traditional 80% human success rate** and **manual or traditional ML-based methods** that rely on real-world sample imaging.

- **OpenAI shows 'Strawberry' to feds, races to launch it**  
  **OpenAI's new AI model, Strawberry, is designed to solve complex problems accurately on the first attempt,** leveraging a technique known as [process supervision](https://openai.com/index/improving-mathematical-reasoning-with-process-supervision/).

- **The Mamba in the Llama: Distilling and Accelerating Hybrid Models**  
  **Linear RNN architectures like Mamba** can rival Transformer models in language tasks, benefiting from **efficient deployment characteristics** and the ability to distill large Transformers into more deployable forms using existing linear projection weights.

- **When A.I.'s Output Is a Threat to A.I. Itself**  
  **A.I.-generated data's indistinguishability** poses a risk of creating feedback loops where A.I. trains on its own output, leading to degraded performance and **model collapse**.

- **PyRoboCOP: Python-Based Robotic Control and Optimization Package**  
  **PyRoboCOP** is a **Python-based package** designed for **robotic control and optimization**, specifically addressing **manipulation tasks and collision avoidance** through handling **complex complementarity constraints**.

- **Show HN: Repo2vec – an open-source library for chatting with any codebase**  
  **`repo2vec` is a modular library designed for chatting with codebases**, simplifying the process of understanding and integrating code without manual code review.

- **[D] Image segmentation converges to all zeros when masks are too small**  
  **Image segmentation models**, such as **UNet**, **struggle with small masks**, leading to convergence on all-zero outputs, particularly in **breast cancer image segmentation** tasks.

- **[D]Exploring the Potential of Edge Computing/Federated Learning in Continuous Training for GPT/LLMs**  
  **Exploring the potential of Edge Computing and Federated Learning** could revolutionize the way **GPT and Large Language Models (LLMs)** are continuously trained, by decentralizing the process.
