AI agents may soon surpass people as primary application users
AI agents are projected to become the primary users of enterprise digital systems by 2030, marking a significant shift in user interaction as they surpass human engagement in application usage by 2032, according to Accenture's predictions.
AI Engineer Reading List
The 2025 AI Engineering Reading List curates 50 essential papers across 10 AI fields, including LLMs, Vision, and CodeGen, aimed at providing practical insights for engineers starting from scratch.
Cosine Similarity Isn't the Silver Bullet We Thought It Was
Cosine similarity has been found to be unreliable due to arbitrary scaling introduced by regularization in linear matrix factorization models, which can lead to meaningless results in applications like recommendation systems.
Lines of code that will beat A/B testing every time (2012)
A 20-line code modification can significantly enhance A/B testing by implementing a multi-armed bandit approach, allowing for real-time optimization of user interactions on websites.
Reversible Computing Escapes the Lab
Reversible computing is transitioning from theory to practice, with Vaire Computing aiming to commercialize chips that recover energy in arithmetic circuits, potentially achieving a 4,000x energy-efficiency gain over traditional methods.
voyage-code-3
voyage-code-3 enhances code retrieval accuracy, outperforming OpenAI-v3-large and CodeSage-large by 13.80% and 16.81% respectively, while utilizing lower-dimensional, quantized embeddings to reduce storage costs significantly.
Training AI models might not need enormous data centres
AI model training may evolve to require significantly less physical infrastructure, potentially eliminating the need for large data centers altogether, as advancements in distributed computing and resource-sharing technologies emerge.
Cheating Is All You Need
LLMs represent a monumental shift in software engineering, comparable to the advent of the World Wide Web, yet many engineers remain skeptical, viewing them as just another trend like crypto.
Hallucination Detection Benchmarks
LLM-as-a-Judge framework for hallucination detection outperforms many advanced research methods, with a CoT (Chain-of-Thought) prompting achieving an impressive accuracy of 0.833 on the HaluBench dataset.
Fast Semantic Text Deduplication
SemHash offers a novel approach to semantic text deduplication, addressing the complexities of duplicate samples that can distort model training and lead to unreliable results, with added explainability features for transparency in the deduplication process.
Breaking Memory Limits: Gradient Wavelet Transform Enhances LLMs Training
The Gradient Wavelet Transform (GWT) method significantly reduces memory requirements for training large language models (LLMs) by applying wavelet transforms to gradients, enhancing efficiency without compromising performance.
NVIDIA and IQVIA Build Domain-Expert Agentic AI for Healthcare and Life Sciences
NVIDIA and IQVIA are developing custom AI agents using the NVIDIA AI Foundry to enhance drug research, clinical development, and commercialization, ultimately aiming to improve patient outcomes in healthcare and life sciences.
AutoGen v0.4: Reimagining the foundation of agentic AI for scale, extensibility, and robustness
AutoGen v0.4 introduces a redesigned architecture that enhances robustness, scalability, and extensibility for agentic AI applications, addressing previous user feedback on architectural constraints and API inefficiencies.
Parallel Key-Value Cache Fusion for Position Invariant RAG
The proposed Parallel Key-Value Cache Fusion framework enhances Retrieval Augmented Generation (RAG) by ensuring position invariance, allowing decoder-only models to produce consistent outputs regardless of input context order.
GenAI Acceleration for PyTorch 2.5 on Intel® Xeon®Processors
GenAI acceleration for PyTorch 2.5 on Intel® Xeon® Processors enhances performance for models like GPTFast, Segment Anything Fast, and Diffusion Fast, achieving speedups of up to 3.95x through optimizations like weight-only quantization and BFloat16 support.