ML Times
Sep 6, 2025
Articles
Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
Apertus-70B-2509 is a new addition to the Apertus LLM collection, showcasing advancements in large language models with enhanced capabilities.ML needs a new programming language – Interview with Chris Lattner
Mojo, a new programming language by Chris Lattner, aims to enhance GPU programming by combining ease of use with type-safe metaprogramming, allowing developers to harness the full power of modern hardware while maintaining productivity.Novel hollow-core optical fiber transmits data 45% faster with record low loss
A novel hollow-core optical fiber achieves 45% faster data transmission with a record low loss of 0.091 dB/km, surpassing traditional silica fibers which have a minimum loss of 0.14 dB/km.How big are our embeddings now and why?
Embedding sizes have significantly increased, with models like OpenAI's using 1536 dimensions, reflecting a shift from the previous standard of 300 dimensions due to advancements in training data and architecture.Now Live: Europe’s First Exascale Supercomputer, JUPITER, Accelerates Climate Research, Neuroscience, Quantum Simulation
JUPITER, Europe’s first exascale supercomputer, is now operational, enabling unprecedented computational power for scientific research and AI applications.Europe enters the exascale supercomputing league with Jupiter
The European Commission's Press Corner serves as a vital hub for accessing official communications, including press releases and statements, enhancing transparency and public engagement. Link to article[R] The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
Current benchmarks for hallucination detection in LLMs are inadequate, often yielding low-signal results that fail to address real-world high-stakes scenarios, as highlighted in the paper here.[P] I Was Wrong About Complex ML Solutions - Gower Distance Beat My UMAP Approach
Gower distance, a method for calculating distances in mixed categorical and numerical data, often outperforms complex ML solutions like UMAP, especially in small-to-medium datasets, highlighting the value of simplicity in machine learning.[P] Knowledge Distillation for Text-to-SQL — Training GPT-2 with Qwen2-7B as Teacher
Knowledge Distillation (KD) enables GPT-2 to effectively generate SQL queries from natural language by learning from the larger Qwen2-7B model, demonstrating that smaller models can be trained for specific tasks without the need for massive resources.[D] Baseten raises $150M Series D for inference infra. where’s the real bottleneck?
Baseten's $150M Series D emphasizes their commitment to enhancing inference infrastructure, targeting low latency and throughput optimization, with a valuation of $2.1B.[P] An Open-Source Pipeline for Speech-to-Speech Translation with Voice Preservation (RVC) and Lip-Sync
The project presents an open-source pipeline for speech-to-speech translation that preserves the speaker's voice and synchronizes lip movements, initially targeting Telugu for low-resource language dubbing.Fast 2-Simplicial Attention: Hardware-Efficient Kernels in TLX
Fast 2-Simplicial Attention achieves 588 Tensor Core BF16 TFLOPs with 60% tensor core utilization, marking a 1.74x speedup over the original Triton kernel, thanks to a hardware-aligned design and the use of TLX for modern GPU techniques.