ML Times
Jan 28, 2025
Nvidia’s $589B DeepSeek rout
Nvidia and ASML shares fell sharply as the Chinese AI startup DeepSeek demonstrated competitive performance against Western chatbots at a significantly lower cost, raising concerns about the sustainability of the current AI business model reliant on high-end chips.
DeepSeek releases Janus Pro, a text-to-image generator [pdf]
The Janus Pro Technical Report details advancements in AI technology, specifically focusing on enhanced inference capabilities and model efficiency that significantly improve performance in various applications.
DeepSeek's training efficiency is 45x greater, suggesting they could have led the LLM market, yet they opted for open-sourcing, likely to foster collaboration and innovation within the community.
Nvidia's stock faces significant challenges as competition intensifies from innovative architectures like Cerebras and Groq, which threaten its data center dominance and high margins, while major customers develop custom silicon to reduce reliance on Nvidia's products.
ggml achieves a remarkable x2 speed increase for WASM by optimizing SIMD instructions in the
qX_K_q8_KandqX_0_q8_0dot product functions, significantly enhancing performance for web applications.DeepSeek-R1 is a groundbreaking open weights model that excels in reasoning tasks by utilizing a unique training method that incorporates long chains of reasoning and reinforcement learning, resulting in a model capable of generating detailed thought processes.
SLAP and FLOP are speculative execution attacks targeting Apple CPUs, exploiting the Load Address Predictor (LAP) and Load Value Predictor (LVP), respectively, to access sensitive data through incorrect memory operations.
DeepSeek v3 achieves state-of-the-art performance with only 2.8 million H800 hours of training, significantly less than Llama 3.1, by innovating with multi-head latent attention (MLA) and mixture-of-experts (MoE) techniques.
Cleveland police's reliance on AI facial recognition to justify a search warrant led to the dismissal of key evidence in a murder case, as the technology's results are deemed inadmissible in court, highlighting the critical need for regulatory oversight in law enforcement practices.
Berkeley researchers have successfully replicated DeepSeek R1's core technology for under $30, demonstrating that small models can achieve complex reasoning capabilities, thus democratizing AI research.
Researchers at the University of Toronto have developed nano-architected materials that combine the strength of carbon steel with the lightness of Styrofoam, utilizing machine learning to optimize their design for enhanced performance.
ErisForge is a Python library that enables the modification of Large Language Models (LLMs) by transforming their internal layers, allowing for the creation of both ablated and augmented model versions tailored to specific inputs.
Sam Altman's assertion that startups with only $10 million are "totally hopeless" against giants like OpenAI has been challenged by DeepSeek, which claims to have trained its model for just $5.6 million, showcasing a potential shift in the competitive landscape of AI startups.
Qwen2.5-Max is a large-scale Mixture-of-Expert (MoE) model pretrained on over 20 trillion tokens, utilizing advanced techniques like Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) to enhance its intelligence.
Diffusion models, particularly the Diffusion Transformer (DiT), have transformed content synthesis, yet they struggle with generation diversity, which this work addresses through targeted attention feature injection for consistent image edits.