# Aug 22, 2024

## Daily

### Second human being implanted with Neuralink brain chip

- **Alex, the second participant in Neuralink's PRIME Study, quickly adapted to using the Neuralink implant**, demonstrating enhanced digital device control, including video gaming and CAD software for 3D design, shortly after implantation.

### Launch HN: Outerport (YC S24) – Instant hot-swapping for AI model weights

- **Outerport introduces 'hot-swapping' for AI models**, allowing different models to be served on the same GPU with swap times of approximately **2 seconds**, a process **150x faster** than the current baseline, demonstrated through a [live demo](https://hotswap.outerport.com/).

### Self-Supervised Learning for Videos

- **Self-supervised learning for videos** leverages the **VideoMAE architecture**, adapting masked autoencoders from images to videos, addressing the unique spatio-temporal complexity of video data. [VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training, 2022](https://arxiv.org/abs/2203.12602)

### GPU utilization can be a misleading metric

- **GPU Utilization, often measured by tools like nvidia-smi, can misleadingly indicate full GPU engagement** through memory activities without actual computation, revealing a gap in assessing true performance.

### Trailer Faces HQ Dataset

- The **Trailer Faces HQ Dataset** comprises **186,553 high-resolution face images** from movie trailers, aiming to provide a diverse range of facial expressions for advanced machine learning applications, available for download at [Hugging Face](https://huggingface.co/datasets/justinpinkney/trailer-faces-hq).

### Transformers learn in-context by gradient descent [R]

- The paper demonstrates that **transformers can emulate the effect of a gradient descent step** on a linear model's weights by **transforming input data** through a single-head attention mechanism, aligning the **transformed data's loss** with that of the **post-gradient descent model**. [Read the paper](https://arxiv.org/pdf/2212.07677)

### [D] Beta9: Open Serverless GPU Cloud

- **Beta9**, now open-sourced, **enables developers to deploy serverless functions on cloud GPUs** efficiently, building on the success of its predecessor, Beam, to offer a Python-first experience for rapid ML model deployment. [GitHub Repo](https://github.com/beam-cloud/beta9)

### Lightweight Champ: NVIDIA Releases Small Language Model With State-of-the-Art Accuracy

- **NVIDIA's Mistral-NeMo-Minitron 8B** is a compact, high-accuracy language model, leveraging **pruning and distillation** techniques to reduce its size from 12 billion to 8 billion parameters without sacrificing performance.

### [P] Formatron: a high-performance constrained decoding library

- **[Formatron](https://github.com/Dan-wanna-M/formatron) excels in controlling language model outputs** with features like **fluent formatting**, **regex and CFG support**, and **efficient JSON generation**, setting it apart as a **lightweight and flexible library**.

### FocusLLM: Scaling LLM's Context by Parallel Decoding

- **FocusLLM** enhances decoder-only LLMs by enabling them to **process long text inputs efficiently**, focusing on relevant information from very long sequences through a novel parallel decoding mechanism.

### Predicting Phenotypes from Omics Data [R] [P]

- **Deep learning** is being explored to **predict phenotype variation** from **omics data**, indicating a growing interest in leveraging computational methods for biological insights.

### TwoMinutePapers - New AI Makes The Mona Lisa Come Alive!

### [R] HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models in Resource-Constrained Environments

- **HiRED**, a **token-dropping scheme**, enhances the efficiency of **high-resolution Vision-Language Models (VLMs)** by selectively processing visual tokens within a **fixed token budget**, preserving accuracy in **resource-constrained environments**.

### SLMming Down Latency: How NVIDIA’s First On-Device Small Language Model Makes Digital Humans More Lifelike

- **NVIDIA's first on-device small language model (SLM), Nemotron-4 4B Instruct, enhances digital human interaction** in games by understanding and responding to player instructions more intuitively and accurately.

### Bidirectional Gated Mamba for Sequential Recommendation

- The **Selective Gated Mamba (SIGMA)** framework introduces a **bidirectional architecture** using Partially Flipped Mamba (PF-Mamba) to enhance **contextual modeling** in Sequential Recommender Systems (SRS), addressing the limitations of unidirectional models.
