Willow, Google's new quantum chip, achieves exponential error reduction as it scales, marking a pivotal advancement in quantum error correction that has eluded researchers for nearly three decades.
Trellis – 3D Mesh Generative Model
TRELLIS introduces a unified Structured Latent (SLAT) representation that enables the generation of high-quality 3D assets across various formats, including Radiance Fields and meshes, by integrating sparse 3D grids with dense visual features from advanced vision models.
The Google Willow Thing
Google's Willow chip, a 105-qubit superconducting device, marks a significant advancement in quantum computing, demonstrating improved coherence times and gate fidelity, with a 2-qubit gate fidelity of 99.7% for controlled-Z gates and 99.85% for iswap gates, compared to ~99.5% in 2019.
Training LLMs to Reason in a Continuous Latent Space
Coconut introduces a novel reasoning paradigm for large language models (LLMs) by utilizing a continuous latent space instead of traditional language space, enhancing reasoning capabilities through a breadth-first search approach.
AI Model for Near-Instant Image Creation on Consumer-Grade Hardware
NitroFusion is the first AI model enabling near-instant image creation on consumer-grade hardware, allowing users to generate images as they type, thus revolutionizing creative workflows.
[R] Diffusion Models, Image Super-Resolution, and Everything: A Survey
This survey paper explores diffusion models in the context of image super-resolution, highlighting their effectiveness and potential applications in enhancing image quality.
Google Says AI Weather Model Masters 15-Day Forecast
Google's GenCast AI model delivers 15-day weather forecasts with superior accuracy, outperforming the ECMWF in over 97% of tested scenarios, showcasing its potential for life-saving applications amid climate change challenges.
Long Convolutions via Polynomial Multiplication
Long convolutions in GPT-like models leverage polynomial multiplication and Fast Fourier Transforms (FFTs) to efficiently handle sequences, enabling models to process longer contexts than traditional methods allow.
[R] Understanding Transformer Limitations in Graph Search: A Mechanistic Analysis of Learning and Scaling Behavior
Transformers struggle with graph search due to inherent scaling limitations, as they can only learn basic search operations effectively when trained on simpler, smaller graphs.
[R] The Well: A Large-Scale Collection of Diverse Physics Simulations for Machine Learning
The Well is a 15TB dataset comprising 16 diverse datasets of numerical simulations, enabling researchers to evaluate machine learning models across various spatiotemporal physical systems like fluid dynamics and supernova explosions.
[D] Meta's New LLama Model
Meta's new Llama 3.3 70B model significantly enhances efficiency, aiming to reduce compute costs for large AI models, which is crucial for scaling AI applications.
[R] Monet: Mixture of Monosemantic Experts for Transformers
Monet introduces a Sparse Mixture-of-Experts (SMoE) architecture that enhances mechanistic interpretability in large language models by addressing polysemanticity through monosemantic experts.
From Uncertainty to Trust: Enhancing Reliability in Vision-Language Models with Uncertainty-Guided Dropout Decoding
Dropout Decoding enhances reliability in large vision-language models (LVLMs) by quantifying and addressing uncertainty in visual token interpretation, leading to improved output quality.
[R] Improving Robustness to Corruptions with Multiplicative Weight Perturbations - A Simple Yet Effective Approach to Robustify Neural Networks to Corruptions
DAMP (Data augmentation via multiplicative perturbations) enhances neural network robustness by applying multiplicative weight perturbations during training, achieving ResNet50-level performance on ImageNet without complex augmentations.
[R] Distillation-Based Colorization of 3D Neural Radiance Fields for Consistent Novel View Synthesis
This paper presents a knowledge distillation approach that effectively colorizes 3D neural representations from grayscale images, leveraging pre-trained 2D models to ensure view consistency across novel perspectives.