Hashed sorting outperforms hash tables in counting unique values in large arrays, achieving up to 4× faster performance in well-tuned implementations, particularly for larger datasets.
Defeating Nondeterminism in LLM Inference
Nondeterminism in LLM inference arises from floating-point non-associativity and batch size variability, which can lead to different outputs even with identical inputs, as demonstrated by the varying completions generated by models like Qwen-3.
Top model scores may be skewed by Git history leaks in SWE-bench
Loopholes in SWE Bench Verified allow agents to access future repository states, revealing solutions and approaches through commands like git log --all, which can expose future commits that directly address issues.
Spiral
Spiral introduces a new data infrastructure designed for the Third Age of data systems, enabling machine-scale outputs that legacy platforms cannot support, with backing from major tech firms like Microsoft and Snowflake.
Intel's E2200 "Mount Morgan" IPU at Hot Chips 2025
Intel’s E2200 “Mount Morgan” IPU enhances cloud infrastructure with 24 Arm Neoverse N2 cores, improved accelerators, and 400 Gbps Ethernet throughput, doubling its predecessor's capabilities.
DeepCodeBench: Real-World Codebase Understanding by Q&A Benchmarking
DeepCodeBench introduces a benchmark dataset of 1,144 real-world questions derived from complex code repositories, enhancing the understanding of codebases through Q&A benchmarking. This dataset aims to improve AI-assisted workflows by addressing the challenges developers face when navigating large codebases.
Hot Chips 2025: Session 1 – CPUs – By George Cozma
Hot Chips 2025 showcased significant advancements in CPU technology, featuring presentations on Condor Computing's Cuzco core, PEZY's SC4s chip, IBM's Power11, and Intel's Clearwater Forest Xeon CPU.
[D]NVIDIA Blackwell Ultra crushes MLPerf
NVIDIA's Blackwell Ultra achieved 5× throughput on DeepSeek-R1 and set records on Llama 3.1 and Whisper, showcasing innovative techniques like FP8 KV-cache and disaggregated serving.
Stability AI Introduces Stable Audio 2.5
Stable Audio 2.5 is the first audio generation model tailored for enterprise-grade sound production, enabling brands to create distinct audio identities across various channels, enhancing memorability by up to eight times.
ApeRAG: Production-ready GraphRAG
ApeRAG is a production-ready RAG platform that integrates Graph RAG, vector search, and full-text search, enabling the development of sophisticated AI applications with hybrid retrieval and intelligent agents.
🤗Tricks from OpenAI gpt-oss
OpenAI's GPT-OSS models introduce advanced techniques like MXFP4 quantization and custom kernels, enhancing the efficiency of loading, running, and fine-tuning within the transformers library.
[P] Semlib: LLM-powered Data Processing
Semlib introduces a novel approach to semantic data processing by utilizing functional programming primitives, effectively separating data pipeline logic from LLM orchestration, which enhances efficiency in handling complex tasks.
RenderFormer
RenderFormer is a groundbreaking neural architecture that enables full 3D rendering without traditional graphics computations, marking a shift towards neural rendering that leverages machine learning for complex light transport simulation.
[D] Universal Deep Research (UDR)
Universal Deep Research (UDR) by Nvidia redefines AI research agents by allowing users to create research strategies in plain English, which are then compiled into executable code, enhancing flexibility and control over the research process.
Delta Flow
Delta Flow's AI platform revolutionizes the architecture, engineering, and construction (AEC) sector by generating buildable digital twins in minutes, drastically reducing the pre-construction timeline from months to mere minutes.