ML Times
Nov 28, 2025
Got burned by an Apple ICLR paper — it was withdrawn after my Public Comment.
A critical bug in an Apple ICLR paper's code led to unexpected low scores for a model, revealing potential 30% ground truth error rate in the dataset used, which was likely generated with poor quality control.
TPUs vs. GPUs and why Google is positioned to win AI race in the long term
Google's TPU, designed for AI inference, emerged from a need to efficiently handle massive data loads, outperforming traditional CPUs and GPUs by focusing on specific tasks like running TensorFlow neural networks.
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
DeepSeek-Math-V2 is a significant advancement in mathematical modeling, enhancing the capabilities of AI in solving complex mathematical problems, as detailed in the DeepSeekMath_V2.pdf.
Vsora Jotunn-8 5nm European inference chip
Jotunn 8 is the world’s most efficient AI inference chip, designed to optimize speed, cost, and scalability for real-time applications, ensuring maximum impact on AI investments.
A trillion dollars (potentially) wasted on gen-AI
Scaling improvements in AI are stagnating, as Ilya Sutskever warns that reliance on larger models without innovative techniques may lead to diminishing returns, echoing concerns raised in previous critiques of deep learning methodologies.
Open (Apache 2.0) TTS model for streaming conversational audio in realtime
Dia2 is a streaming dialogue TTS model that generates audio from partial text input, enabling real-time conversations and supporting up to 2 minutes of English audio generation.
Inverse hyperbolic sine as an activation function and its anti-derivative as a loss function
The inverse hyperbolic sine activation function demonstrates superior performance in regression tasks, outperforming traditional functions like ReLU and sigmoid in cross-validation metrics such as WMAPE.
Any VLMs that are fully reproducible with clear documentation on how to do so?
Recent VLMs with true reproducibility are sought after, as many papers lack clear instructions, risking extensive GPU time without guaranteed results.