ML Times
Stable Virtual Camera: Multi-View Video Generation with 3D Camera Control
Stable Virtual Camera is a multi-view diffusion model that converts 2D images into immersive 3D videos, allowing for dynamic camera control across various paths without complex preprocessing.
Karatsuba Matrix Multiplication and Its Efficient Hardware Implementations
The Karatsuba algorithm, when extended to matrix multiplication, not only retains its reduced multiplication complexity but also minimizes the overhead of additional operations, making it more efficient for larger datasets.
Preview: Amazon S3 Tables and Lakehouse in DuckDB
DuckDB now supports Apache Iceberg REST Catalogs, allowing seamless connections to Amazon S3 Tables and Amazon SageMaker Lakehouse, enhancing data accessibility for users.
Nvidia Dynamo: A Datacenter Scale Distributed Inference Serving Framework
NVIDIA Dynamo is a high-throughput, low-latency inference framework tailored for generative AI and reasoning models, supporting multiple inference engines like TRT-LLM and vLLM, while optimizing GPU performance through features like dynamic scheduling and KV cache offloading.
Introducing KBLaM: Bringing plug-and-play external knowledge to LLMs
KBLaM introduces a novel integration of structured knowledge bases into LLMs, utilizing a rectangular attention mechanism that allows for efficient, linear scaling without the need for costly retraining or complex retrieval modules.
RWKV-7 "Goose" with Expressive Dynamic State Evolution
RWKV-7 "Goose" achieves state-of-the-art performance in multilingual tasks at the 3 billion parameter scale, outperforming other models despite training on significantly fewer tokens, while maintaining constant memory usage and inference time per token.
NVIDIA Accelerated Quantum Research Center to Bring Quantum Computing Closer
NVIDIA's Accelerated Quantum Research Center (NVAQC) aims to integrate quantum processing units (QPUs) with AI supercomputers, enhancing capabilities to tackle complex quantum computing challenges and advancing quantum error correction techniques.
[R] Jagged Flash Attention Optimization
Jagged Flash Attention optimizes large-scale recommendation systems, achieving up to 9× speedup and 22× memory reduction compared to dense attention methods, marking a significant leap in performance.
The clustering behavior of sliding windows
Clustering timeseries data with a sliding window can lead to unexpected failures based on the relationship between window size and timeseries length, revealing critical insights into data preprocessing.
Notebooks as reusable Python programs
marimo transforms notebooks into reusable Python programs, allowing for version control with Git, testing with pytest, and execution as scripts, thus enhancing maintainability and interoperability in data workflows.
An early look at cryptographic watermarks for AI-generated content
Cryptographic watermarks for AI-generated content aim to ensure robustness, undetectability, and unforgeability, enhancing the identification of AI artifacts while maintaining quality.
[R] Forget Chain-of-Thought reasoning! Introducing Chain-of-Draft: Thinking Faster (and Cheaper) by Writing Less.
Chain-of-Draft (CoD) offers a faster and cheaper alternative to Chain-of-Thought (CoT) reasoning by significantly reducing the length of model prompts, thus minimizing token usage and computational costs.
ByteCraft: Generating video games and animations through bytes
ByteCraft generates executable video games and animations from text prompts by fine-tuning a 7B parameter LLM (Qwen2.5) over 4 months on 4 GPUs, showcasing its potential despite challenges in byte-level accuracy.
[R] RWKV-7 "Goose" with Expressive Dynamic State Evolution
RWKV-7 "Goose" achieves state-of-the-art performance in multilingual tasks at the 3 billion parameter scale, outperforming other models trained on more data while maintaining constant memory and inference time per token.
[R] SmolDocling: A Compact Vision-Language Model for Complete Document Element Recognition and Markup Generation
SmolDocling is an ultra-compact vision-language model that integrates a 2B parameter vision encoder with a 5B parameter language decoder, achieving end-to-end document processing while being significantly smaller than competitors.