ML Times
Dec 16, 2024
ML Times
Daily
Weekly
Konwinski Prize
- $1M K Prize will reward the first open source AI to achieve 90% accuracy on a new, contamination-free version of SWE-bench, promoting integrity in AI benchmarking.
Veo 2: Our video generation model
- Veo 2 is a cutting-edge video generation model that produces realistic motion and high-quality output up to 4K resolution, significantly enhancing detail and reducing artifacts compared to previous models.
Fast LLM Inference From Scratch (using CUDA)
- Building an LLM inference engine from scratch using C++ and CUDA allows for deep insights into optimizations that enhance inference speed, particularly for single-batch processing on consumer devices, surpassing existing implementations like llama.cpp.
Best-of-N Jailbreaking
- Best-of-N (BoN) Jailbreaking is a novel black-box algorithm that effectively exploits AI systems by generating variations of prompts, achieving high attack success rates (ASRs) of 89% on GPT-4o and 78% on Claude 3.5 Sonnet with 10,000 samples.
In Search of a Faster SQLite
- Researchers at the University of Helsinki and Cambridge have achieved a 100x reduction in tail latency for SQLite by implementing asynchronous I/O and disaggregated storage in their paper, “Serverless Runtime / Database Co-Design With Asynchronous I/O”.
🤗Introducing the Synthetic Data Generator - build Datasets with Natural Language
- The Synthetic Data Generator enables users to create custom datasets using natural language prompts, streamlining the process with a no-code interface powered by Hugging Face's text-generation API and distilabel.
Researchers discover new third class of magnetism
- Altermagnetism, a newly discovered class of magnetism, could revolutionize digital devices by enabling magnetic memory with speeds up to a thousand times faster than current technologies.
Load is not what you should balance: Introducing Prequal
- PReQuaL (Probing to Reduce Queuing and Latency) is a novel load balancer that prioritizes minimizing real-time request latency over traditional CPU load balancing, utilizing asynchronous and reusable probes to select servers based on estimated latency and active requests-in-flight (RIF).
High-Fidelity 3D Mesh Generation at Scale with Meshtron – Nvidia Technical Blog
- Meshtron enables high-fidelity 3D mesh generation at scale, significantly enhancing the efficiency and quality of 3D modeling in various applications, including gaming and simulations.
[R] Optimizing LLM merging to reduce performance tradeoffs
- Optimizing LLM merging can yield Pareto-optimal models by effectively combining checkpoints trained under varying conditions, enhancing performance across multiple tasks without the need for extensive hyperparameter tuning.
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
- LinGen introduces a linear-complexity framework for text-to-video generation, enabling high-resolution minute-length videos on a single GPU without sacrificing quality, a significant advancement over traditional models limited to 10-20 seconds.
[N] Special session on Privacy-Preserving Machine and Deep Learning at IJCNN 2025
- The special session on Privacy-Preserving Machine and Deep Learning at IJCNN 2025 seeks innovative papers on techniques like Homomorphic Encryption to enhance AI applications, emphasizing the need for new ML and DL models.
State-of-the-art video and image generation with Veo 2 and Imagen 3
- Veo 2 and Imagen 3 are now available, showcasing state-of-the-art capabilities in video and image generation, with Veo 2 offering high-quality video creation at 4K resolution and Imagen 3 enhancing artistic diversity and detail accuracy.