# Jul 11, 2024

### Daily

- **DoLa Decoding by Contrasting Layers Improves Factuality in Large Language Models**  
  **DoLa**, a novel decoding strategy, **significantly reduces hallucinations** in large language models by contrasting logits from different transformer layers, **enhancing factuality** without external knowledge or fine-tuning.

- **Vision language models are blind**  
  **VLMs struggle with basic visual tasks** such as identifying overlapping circles or counting objects, tasks trivial for humans.

- **Turbopuffer: Fast search on object storage**  
  **turbopuffer** is a **cost-efficient, high-performance search engine** that leverages **object storage and smart caching**, designed to scale to billions of vectors and millions of tenants/namespaces, addressing the high costs and operational challenges of existing search solutions.

- **Inference Performance Optimization for Large Language Models on CPUs**  
  **Optimizing inference performance** for **Large Language Models (LLMs) on CPUs** addresses the challenge of deploying high-performance LLMs in **low-resource environments**, focusing on reducing financial and hardware constraints.

- **RouteLLM: A framework for serving and evaluating LLM routers**  
  **RouteLLM introduces a framework for efficiently routing queries between large language models (LLMs), achieving up to an 85% reduction in costs while maintaining 95% of GPT-4's performance**, as detailed in their [paper](https://arxiv.org/abs/2406.18665).

- **FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-Precision**  
  **FlashAttention-3** introduces techniques exploiting **asynchrony and low-precision (FP8)** on Hopper GPUs, achieving **1.5-2.0x speed improvements** and **75% utilization** of H100 GPU capabilities, significantly enhancing the efficiency of attention mechanisms in large language models (LLMs).

- **Karpathy: Let's reproduce GPT-2 (1.6B): one 8XH100 node 24h $672 in llm.c**  
  Reproducing **GPT-2 (1.6B parameters)** using **llm.c** on a single 8XH100 node took **24 hours and cost $672**, showcasing significant advancements in compute, software, and data availability since its initial release by OpenAI in 2019.

- **Training of Physical Neural Networks**  
  **Physical Neural Networks (PNNs)** leverage the properties of physical systems to **perform computation**, potentially enabling AI models **1000x larger** than current ones to operate **locally and privately on edge devices**.

- **Korvus: Single-Query RAG with Postgres**  
  **Korvus integrates the entire RAG pipeline into a single Postgres database query**, offering a unified search SDK with Python, JavaScript, and Rust bindings for high-performance, customizable search functionalities.

- **GitHub Copilot is not infringing your copyright**  
  **GitHub Copilot**, trained on publicly available source code, **does not infringe copyright** by generating code suggestions, despite using GPL-licensed repositories for training.

- **From Unlabeled Data to Rich Segmentation: The Magic of Self-Supervised Models**  
  **Self-supervised learning** with **DINOv2 ViT weights** from Facebook Research enables **rich image segmentation** by finetuning with Low-Rank Adaptation (LoRA) and simple decoders, achieving **solid validation IoU scores**.

- **Memory^3: Language Modeling with Explicit Memory**  
  **Memory^3** introduces a **novel approach** to language modeling by equipping large language models (LLMs) with **explicit memory**, significantly reducing parameter size, training, and inference costs while outperforming larger models and text retrieval-augmented generation (RAG) models. [Read the paper](https://arxiv.org/pdf/2407.01178)

- **Controllable Navigation Instruction Generation with Chain of Thought Prompting**  
  **C-Instructor** leverages **chain-of-thought-style prompting** to enable **style-controllable and content-controllable navigation instruction generation**, addressing limitations of existing models by incorporating **landmark identification** for clearer guidance.

- **TwoMinutePapers - NVIDIA’s New Tech Runs A Virtual City!**  
  NVIDIA's breakthrough allows for the creation of **virtual worlds** from **250,000 photos**, synthesizing unseen parts between images using **Neural Radiance Fields (NERFs)**, enabling unprecedentedly large and detailed virtual scenes.
