ML Times

Stable Virtual Camera: Multi-View Video Generation with 3D Camera Control

Stable Virtual Camera is a multi-view diffusion model that converts 2D images into immersive 3D videos, allowing for dynamic camera control across various paths without complex preprocessing.


Karatsuba Matrix Multiplication and Its Efficient Hardware Implementations

The Karatsuba algorithm, when extended to matrix multiplication, not only retains its reduced multiplication complexity but also minimizes the overhead of additional operations, making it more efficient for larger datasets.


Preview: Amazon S3 Tables and Lakehouse in DuckDB

DuckDB now supports Apache Iceberg REST Catalogs, allowing seamless connections to Amazon S3 Tables and Amazon SageMaker Lakehouse, enhancing data accessibility for users.


Nvidia Dynamo: A Datacenter Scale Distributed Inference Serving Framework

NVIDIA Dynamo is a high-throughput, low-latency inference framework tailored for generative AI and reasoning models, supporting multiple inference engines like TRT-LLM and vLLM, while optimizing GPU performance through features like dynamic scheduling and KV cache offloading.


Introducing KBLaM: Bringing plug-and-play external knowledge to LLMs

KBLaM introduces a novel integration of structured knowledge bases into LLMs, utilizing a rectangular attention mechanism that allows for efficient, linear scaling without the need for costly retraining or complex retrieval modules.


RWKV-7 "Goose" with Expressive Dynamic State Evolution

RWKV-7 "Goose" achieves state-of-the-art performance in multilingual tasks at the 3 billion parameter scale, outperforming other models despite training on significantly fewer tokens, while maintaining constant memory usage and inference time per token.


NVIDIA Accelerated Quantum Research Center to Bring Quantum Computing Closer

NVIDIA's Accelerated Quantum Research Center (NVAQC) aims to integrate quantum processing units (QPUs) with AI supercomputers, enhancing capabilities to tackle complex quantum computing challenges and advancing quantum error correction techniques.


[R] Jagged Flash Attention Optimization

Jagged Flash Attention optimizes large-scale recommendation systems, achieving up to 9× speedup and 22× memory reduction compared to dense attention methods, marking a significant leap in performance.


The clustering behavior of sliding windows

Clustering timeseries data with a sliding window can lead to unexpected failures based on the relationship between window size and timeseries length, revealing critical insights into data preprocessing.


Notebooks as reusable Python programs

marimo transforms notebooks into reusable Python programs, allowing for version control with Git, testing with pytest, and execution as scripts, thus enhancing maintainability and interoperability in data workflows.


An early look at cryptographic watermarks for AI-generated content

Cryptographic watermarks for AI-generated content aim to ensure robustness, undetectability, and unforgeability, enhancing the identification of AI artifacts while maintaining quality.


[R] Forget Chain-of-Thought reasoning! Introducing Chain-of-Draft: Thinking Faster (and Cheaper) by Writing Less.

Chain-of-Draft (CoD) offers a faster and cheaper alternative to Chain-of-Thought (CoT) reasoning by significantly reducing the length of model prompts, thus minimizing token usage and computational costs.


ByteCraft: Generating video games and animations through bytes

ByteCraft generates executable video games and animations from text prompts by fine-tuning a 7B parameter LLM (Qwen2.5) over 4 months on 4 GPUs, showcasing its potential despite challenges in byte-level accuracy.


[R] RWKV-7 "Goose" with Expressive Dynamic State Evolution

RWKV-7 "Goose" achieves state-of-the-art performance in multilingual tasks at the 3 billion parameter scale, outperforming other models trained on more data while maintaining constant memory and inference time per token.


[R] SmolDocling: A Compact Vision-Language Model for Complete Document Element Recognition and Markup Generation

SmolDocling is an ultra-compact vision-language model that integrates a 2B parameter vision encoder with a 5B parameter language decoder, achieving end-to-end document processing while being significantly smaller than competitors.