Google CEO says more than a quarter of the company's new code is created by AI
Over 25% of new code at Google is generated by AI, significantly enhancing productivity and efficiency, as confirmed by CEO Sundar Pichai during the Q3 earnings call.
Chain-of-Thought Can Hurt Performance on Tasks Where Thinking Makes Humans Worse
Chain-of-thought (CoT) prompting can lead to a significant drop in model performance, with reductions up to 36.3% absolute accuracy in certain tasks, highlighting the need for careful application in specific contexts.
AI Flame Graphs
AI Flame Graphs are a new tool from Intel designed to visualize AI accelerator performance, potentially reducing resource costs and contributing to a 10% decrease in US power usage by 2030.
Pushing the frontiers of audio generation
Google DeepMind's latest audio generation technology enables the creation of 2-minute dialogues with improved naturalness and speaker consistency, utilizing advanced models that process dialogue scripts and speaker markers in under 3 seconds on a TPU v5e chip.
DeepSeek v2.5 – open-source LLM comparable to GPT-4, but 95% less expensive
DeepSeek-V2.5 excels in performance, ranking in the top 3 on AlignBench and rivaling GPT-4-Turbo, showcasing its advanced capabilities in math, coding, and reasoning.
Cerebras Trains Llama Models to Leap over GPUs
Cerebras has achieved a remarkable 3.5X increase in AI inference performance with its Llama 3.2 models, significantly outperforming Nvidia's H100 GPUs.
Support for Claude Sonnet 3.5, OpenAI O1 and Gemini 1.5 Pro
Qodo now integrates Claude Sonnet 3.5, OpenAI o1, and Google Gemini 1.5 Pro, enhancing its platform with advanced models that improve coding tasks and problem-solving capabilities, with full access set to launch next week.
The National EUV Accelerator comes to Albany
The NSTC EUV Accelerator will be established at the Albany NanoTech Complex, enhancing North America's semiconductor research and manufacturing capabilities, crucial for maintaining technological leadership.
Physical Intelligence's first generalist robotic model
π0 is a groundbreaking generalist robot foundation model designed to enable robots to perform a wide range of tasks by learning from embodied experiences, akin to how large language models (LLMs) operate with text and images.
[R] Experimental Design for Multi-Channel Imaging via Task-Driven Feature Selection (ICLR)
The paper introduces a novel method for supervised feature selection that enhances task-driven image channel selection, significantly impacting MRI acquisition times and multispectral image reconstruction.
[R] Our results experimenting with different training objectives for an AI evaluator
Preference optimisation techniques like DPO and RPO outperform traditional supervised fine-tuning (SFT) in training LLM-as-a-judge models, suggesting a shift in effective training objectives for AI evaluators.
[D] using LLMs to power novelty-seeking adaptive learning agents ?
The proposed novelty-seeking machine learning algorithm leverages Gödel's completeness/incompleteness theorems, aiming to enhance efficiency by integrating LLMs for data generation and parsing, thus reducing the resource burden of training each component independently.
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
Unified framework for evaluating hallucinations in Large Vision-Language Models (LVLMs) addresses both object and relation hallucination, revealing a critical gap in existing benchmarks that focus solely on object-related issues.
Introducing DRIFT Search: Combining global and local search methods to improve quality and efficiency
DRIFT Search enhances local query responses by integrating community insights, allowing for a broader range of facts and more nuanced answers compared to traditional methods.
BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference
BUZZ introduces a beehive-structured sparse KV cache that optimizes memory usage and enhances inference speed for large language models (LLMs) by dynamically segmenting historical tokens and utilizing a sliding window for recent information capture.