DeepSeek V4–almost on the frontier, a fraction of the price
DeepSeek V4 introduces two models, DeepSeek-V4-Pro and DeepSeek-V4-Flash, featuring 1 million token context Mixture of Experts, with Pro boasting 1.6T total parameters and Flash at 284B total parameters.
LLMs consistently pick resumes they generate over ones by humans or other models
LLMs exhibit a significant self-preference bias, favoring their own generated resumes over human-written ones, with a bias range of 67% to 82% across various models, even when controlling for content quality.
Eka’s robotic claw feels like we're approaching a ChatGPT moment
Eka's robotic claw demonstrates unprecedented dexterity, capable of tasks like screwing in light bulbs and handling delicate items, suggesting a potential ChatGPT moment for robotics in the physical realm.
Uber wants to turn its drivers into a sensor grid for self-driving companies
Uber aims to transform its millions of drivers into a vast sensor grid, collecting real-world data for autonomous vehicle (AV) companies, thereby addressing the critical data bottleneck in AV development.
I spent years building a 103B-token Usenet corpus (1980–2013) and finally documented it
The Usenet corpus comprises 103.1 billion tokens and 408 million posts, offering a rich dataset for language model training that spans 33 years of online discourse from 1980 to 2013.
LFM2-24B-A2B: Scaling Up the LFM2 Architecture
LFM2-24B-A2B is a sparse Mixture of Experts (MoE) model with 24 billion parameters, demonstrating effective scaling and consistent quality improvements across benchmarks as it expands from 350M to 24B parameters.
Real World Physics-Informed AI Applications
Physics-informed AI integrates physical laws into machine learning models, enhancing their predictive capabilities in real-world scenarios, such as engineering and environmental science.