# May 28, 2024

## Reproducing GPT-2 in llm.c

- **Reproducing GPT-2 (124M) using llm.c** on a single 8X A100 80GB SXM node takes approximately 90 minutes and costs around $20, showcasing an efficient use of model flops utilization at up to ~60%.

## Grokked Transformers Are Implicit Reasoners

- **Transformers can learn implicit reasoning** through grokking, which involves extended training well beyond the point of overfitting, particularly in the domains of **composition and comparison reasoning**.

## Transformers Can Do Arithmetic with the Right Embeddings

- **Transformers improve in arithmetic tasks** by incorporating **position-specific embeddings** for each digit, addressing their previous limitations in tracking digit positions within large numbers.

## Tinygrad 0.9.0

- **tinygrad 0.9.0** introduces significant usability improvements, including **over 1200 commits** since version 0.8.0, and approaches the new line limit with **7958 lines** of code.

## Llama 3-V: Matching GPT4-V with a 100x smaller model and 500 dollars

- **Llama 3-V** is a **multimodal model** that matches GPT4-V's performance with a **100x smaller model size**, achieving a 10-20% performance boost over Llava, the current state-of-the-art (SOTA) in multimodal understanding, by leveraging a novel architecture and training approach.

## [R] Poisson Variational Autoencoder

- The **Poisson Variational Autoencoder** introduces a novel approach to modeling count data, leveraging the Poisson distribution for enhanced performance in tasks involving discrete numerical data.

## [D] Multi-Step Parallel Prediction for Train Delays Using Graph Neural Networks

- The project aims to **predict train delays** using Graph Neural Networks (GNNs) by considering the **complex spatial-temporal interactions** between trains across the railway network, a challenge highlighted by the network's intricate structure depicted in the provided image.

## [D] Question about You Only Cache Once: Decoder-Decoder Architectures for Language Models

- The **You Only Cache Once (YOCO)** architecture introduces a **novel split** in the network into **self-decoder layers** for generating a **global KV-Cache** and **cross-decoder layers** that reuse this cache, aiming to **enhance efficiency** in Large Language Models (LLMs) as illustrated in [Figure 1](https://preview.redd.it/n6iitz36873d1.png?width=804&format=png&auto=webp&s=597102302acd26f27e28e99b366c13b7b135457a).

## NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

- The **NV-Embed model** introduces **latent attention layers** and a **two-stage contrastive instruction-tuning method**, significantly enhancing LLMs' performance in text embedding tasks.

## [P] DARWIN - open-sourced Devin alternative is back with updates

- **DARWIN's latest update** enhances its ability to **understand and map existing projects' structures** and **extract class and function signatures**, addressing the challenge of context length.

## TwoMinutePapers - NVIDIA Just Supercharged Ray Tracing!

- **NVIDIA and the University of Utah** have developed a **new ray tracing technique** that significantly reduces noise, making real-time, beautifully ray-traced images a reality.

## Assessing LLMs Suitability for Knowledge Graph Completion

- **Large Language Models (LLMs)** like **Mixtral-8x7B-Instruct-v0.1** and **gpt-3.5-turbo-0125** show promise in **Knowledge Graph Completion tasks**, even under **Zero- or Few-Shot settings**, despite their tendency to **hallucinate answers**.

## NVIDIA Scoops Up Wins at COMPUTEX Best Choice Awards

- **NVIDIA AI Enterprise** won a **Golden Award** at COMPUTEX for its cloud-native software platform that enhances the development and deployment of generative AI applications, offering a significant reduction in energy costs and data center footprint for businesses.

## 🤗Training and Finetuning Embedding Models with Sentence Transformers v3

- **Sentence Transformers v3 introduces a comprehensive approach to training and finetuning embedding models**, enhancing performance on specific tasks by incorporating datasets, loss functions, training arguments, evaluators, and a new trainer.

## Think Before You Act: A Two-Stage Framework for Mitigating Gender Bias Towards Vision-Language Tasks

- **GAMA**, a **two-stage framework**, targets **mitigating gender bias** in vision-language models (VLMs) by focusing on **object hallucination** as the core issue, producing gender-neutral narratives to avoid bias.
