# Jan 14, 2025

## AI agents may soon surpass people as primary application users

- **AI agents are projected to become the primary users of enterprise digital systems by 2030**, marking a significant shift in user interaction as they surpass human engagement in application usage by 2032, according to Accenture's predictions.

## AI Engineer Reading List

- The **2025 AI Engineering Reading List** curates **50 essential papers** across **10 AI fields**, including LLMs, Vision, and CodeGen, aimed at providing practical insights for engineers starting from scratch.

## Cosine Similarity Isn't the Silver Bullet We Thought It Was

- **Cosine similarity** has been found to be unreliable due to **arbitrary scaling** introduced by regularization in linear matrix factorization models, which can lead to **meaningless results** in applications like recommendation systems.

## Lines of code that will beat A/B testing every time (2012)

- A **20-line code modification** can significantly enhance A/B testing by implementing a multi-armed bandit approach, allowing for real-time optimization of user interactions on websites.

## Reversible Computing Escapes the Lab

- **Reversible computing** is transitioning from theory to practice, with Vaire Computing aiming to commercialize chips that recover energy in arithmetic circuits, potentially achieving a **4,000x energy-efficiency gain** over traditional methods.

## voyage-code-3

- **`voyage-code-3`** enhances code retrieval accuracy, outperforming OpenAI-v3-large and CodeSage-large by **13.80%** and **16.81%** respectively, while utilizing lower-dimensional, quantized embeddings to reduce storage costs significantly.

## Training AI models might not need enormous data centres

- **AI model training may evolve** to require significantly less physical infrastructure, potentially eliminating the need for large data centers altogether, as advancements in distributed computing and resource-sharing technologies emerge.

## Cheating Is All You Need

- **LLMs represent a monumental shift in software engineering**, comparable to the advent of the World Wide Web, yet many engineers remain skeptical, viewing them as just another trend like crypto.

## Hallucination Detection Benchmarks

- **LLM-as-a-Judge** framework for hallucination detection outperforms many advanced research methods, with a **CoT (Chain-of-Thought)** prompting achieving an impressive **accuracy of 0.833** on the HaluBench dataset.

## Fast Semantic Text Deduplication

- **SemHash** offers a novel approach to **semantic text deduplication**, addressing the complexities of duplicate samples that can distort model training and lead to unreliable results, with added **explainability features** for transparency in the deduplication process.

## Breaking Memory Limits: Gradient Wavelet Transform Enhances LLMs Training

- The **Gradient Wavelet Transform (GWT)** method significantly reduces memory requirements for training large language models (LLMs) by applying wavelet transforms to gradients, enhancing efficiency without compromising performance.

## NVIDIA and IQVIA Build Domain-Expert Agentic AI for Healthcare and Life Sciences

- **NVIDIA and IQVIA are developing custom AI agents** using the NVIDIA AI Foundry to enhance drug research, clinical development, and commercialization, ultimately aiming to improve patient outcomes in healthcare and life sciences.

## AutoGen v0.4: Reimagining the foundation of agentic AI for scale, extensibility, and robustness

- **AutoGen v0.4** introduces a **redesigned architecture** that enhances **robustness, scalability, and extensibility** for agentic AI applications, addressing previous user feedback on architectural constraints and API inefficiencies.

## Parallel Key-Value Cache Fusion for Position Invariant RAG

- The proposed **Parallel Key-Value Cache Fusion** framework enhances **Retrieval Augmented Generation (RAG)** by ensuring **position invariance**, allowing decoder-only models to produce consistent outputs regardless of input context order.

## GenAI Acceleration for PyTorch 2.5 on Intel® Xeon®Processors

- **GenAI acceleration** for PyTorch 2.5 on Intel® Xeon® Processors enhances performance for models like GPTFast, Segment Anything Fast, and Diffusion Fast, achieving speedups of up to **3.95x** through optimizations like weight-only quantization and BFloat16 support.
