ML Times
Oct 3, 2025
Daily
We bought the whole GPU, so we're damn well going to use the whole GPU
New megakernel for tensor-parallel inference with Llama-70B on H100s achieves >22% higher throughput than SGLang by optimizing resource usage across compute, memory, and communication operations.
Fp8 runs ~100 tflops faster when the kernel name has "cutlass" in it
The persistent attention kernel rewrite enhances performance at low contexts, but fp16 performance suffers at larger contexts due to a ptxas instruction scheduling issue in the softmax partition, while fp8 shows a significant speed increase when using "cutlass" in the kernel name.
Microsoft CTO says he wants to swap most AMD and Nvidia GPUs for homemade chips
Microsoft aims to primarily utilize its own AI chips in data centers, reducing dependence on Nvidia and AMD, as indicated by CTO Kevin Scott during a recent discussion at Italian Tech Week.
Be Worried
AI's influence on human behavior is profound and immediate, as it operates without the need for intelligence, raising concerns about societal control through non-sentient systems.
Researchers develop molecular qubits that communicate at telecom frequencies
Molecular qubits developed by researchers can communicate at telecom frequencies, bridging light and magnetism, which positions them as essential components for future quantum networks and technologies.
D: I’m looking for papers, preprints, datasets, or reports where an LLM is trained to only know what humans knew before a major scientific breakthrough, and is then asked to propose a new theoretical framework without using post-breakthrough knowledge and without requiring experimental validation.
Training an LLM on historical physics texts up to 1904 could yield novel theoretical frameworks, challenging the model to generate ideas without modern scientific knowledge or experimental validation.
Move over Dijkstra: New Algorithm Just Rewrote 70 Years of Computer Science
A new algorithm developed by a team at Tsinghua University has shattered the “sorting barrier”, achieving a deterministic O(m log^(2/3) n)-time complexity for single-source shortest paths, thus challenging the long-standing dominance of Dijkstra’s algorithm.
R: New paper: LLMs don't have privileged self knowledge, which means we can efficiently train a General Correctness Model to predict the correctness of multiple models. Surprising or expected?
LLMs lack privileged self-knowledge, meaning they do not predict their own correctness effectively; instead, they excel when trained as a General Correctness Model (GCM) that evaluates multiple models simultaneously.
Learning to Reason for Hallucination Span Detection
LLMs often produce hallucinations, and this study reveals that Chain-of-Thought (CoT) reasoning can enhance the detection of hallucinated spans, yielding at least one correct identification upon multiple samplings.
P: Building a Music Search Engine + Foundational Model on 100M+ Latent Audio Embeddings
EmergeSound.ai leverages 100M+ audio embeddings to enable users to query by sound, enhancing music discovery beyond traditional text-based searches.
VideoNSA: Native Sparse Attention Scales Video Understanding
VideoNSA enhances video understanding in multimodal language models by integrating Native Sparse Attention (NSA) with a focus on maintaining coherence across long time scales, achieving superior performance on various benchmarks.
D: Will fine-tuning LLaMA 3.2 11B Instruct on text-only data degrade its vision capabilities?
Fine-tuning LLaMA 3.2 11B Instruct on text-only data may risk multimodal forgetting, potentially impairing its image processing capabilities. This concern arises from the hypothesis that training on a single modality could diminish performance in others, particularly in tasks like OCR.
Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language Generation
Hallucinations, particularly confabulations, significantly compromise the reliability of large language models (LLMs) due to their predictive uncertainty, necessitating improved evaluation methods for uncertainty estimation in natural language generation (NLG).
DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
DiFFPO introduces a novel framework for training masked diffusion large language models (dLLMs), enhancing their reasoning speed and accuracy through reinforcement learning (RL), which significantly improves sample efficiency and task performance.
Constrained Adaptive Rejection Sampling
Constrained Adaptive Rejection Sampling (CARS) enhances sample efficiency in constrained language model outputs by adaptively ruling out invalid continuations, ensuring that only valid samples are generated without distorting the model's distribution.