ML Times
Jun 28, 2026
Daily
DSpark: Speculative decoding accelerates LLM inference
The DSpark paper from the DeepSpec project presents innovative methodologies for enhancing deep learning model performance, particularly in the context of data efficiency and scalability.
AI learns the “dark art” of RFIC design
AI is revolutionizing RFIC design by employing reinforcement learning and inverse design techniques, enabling the rapid creation of radio chips that outperform traditional designs in both performance and efficiency.
Asian AI startups launch Mythos-like models
Asian AI startups like 360 and Sakana AI have launched models, Tulongfeng and Fugu, respectively, to compete with Anthropic’s Mythos, amid ongoing U.S. export bans that limit access to advanced AI technologies.
AMD Strix Halo RDMA Cluster Setup Guide
The AMD Strix Halo cluster setup guide outlines the configuration of a two-node system using Intel E810 (RoCE v2) for efficient distributed vLLM inference, emphasizing the importance of Tensor Parallelism for handling large models.
Semgrep: GLM 5.2 beats Claude in our Cyber Benchmarks
GLM 5.2, an open-weight model from Zhipu AI, achieved a 39% F1 score in IDOR detection, outperforming Claude Code (32%) at a cost of $0.17 per vulnerability found, showcasing its competitive edge in security tasks.
A way to exclude sensitive files issue still open for OpenAI Codex
A proposed feature seeks to implement a mechanism for marking files and paths that should not be accessed or sent to the model, enhancing security and usability across repositories with a deterministic configuration.
Wayfinder Router: deterministic routing of queries between local and hosted LLM
Wayfinder enables deterministic prompt-complexity routing, allowing users to send prompts to either local or cloud models without incurring additional model calls, thus optimizing costs and efficiency.
Programmable Probabilistic Computer with 1M p-bits
This research introduces a programmable probabilistic computer with 1,000,000 p-bits, enabling unprecedented Gibbs sampling speeds exceeding one trillion flips per second while maintaining local memory for coupling weights. arXiv
MathFormer: Testing whether symbolic math is pattern matching or reasoning
A 4M parameter seq2seq model achieves ~ 98.6% accuracy on symbolic math tasks, indicating it relies on structural token transformations rather than understanding operators or variables.
TOP500 at ISC'26: We Have a New Number 1 – By George Cozma
LineShine Supercomputer in Shenzhen, China, has claimed the top spot on the TOP500 list, marking the first Chinese entry in nine years with a CPU-only system that boasts 2.198 Exaflops of sustained FP64 performance.
Reflecting to optimise
Optimisation in protein binder design reveals that traditional methods may overlook more effective approaches, as demonstrated through the exploration of projected gradient descent (PGD) and mirror descent techniques.
Built an LLM training framework that actually runs on older GPUs without crashing
Picotron is a newly developed LLM training framework that eliminates GPU-specific dependencies, enabling it to run on older GPUs like T4 and V100 without crashing during import.
Evaluating long-term memory limits in stateless LLM chatbots — feedback needed
The research aims to assess long-term memory limits in stateless LLM chatbots by testing their ability to recall key facts after numerous unrelated messages, providing insights into their conversational retention capabilities.
Benchmarking Self-Hosted Gemma 2 9B vs. Frontier APIs: The FP8 Quantization Prefill Tax and VRAM Realities on an NVIDIA L4
FP8 quantization incurs a significant 58% latency penalty on Time to First Token (TTFT) for complex prompts, revealing a trade-off between speed and memory efficiency on an NVIDIA L4 GPU.
NagaTranslate: Building a translation and voice pipeline for low-resource Nagaland creoles (Whisper, VITS, LLMs)