Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
Kimi K2 Thinking explores advanced methodologies in artificial intelligence, emphasizing the integration of cognitive processes to enhance machine learning capabilities.
Mathematical Exploration and Discovery at Scale
AlphaEvolve, developed in collaboration with Google DeepMind, utilizes a large language model (LLM) to evolve computer code for solving mathematical problems, enhancing traditional optimization methods by focusing on code structure rather than raw input data.
Open Source Implementation of Apple's Private Compute Cloud
OpenPCC is an open-source framework for provably private AI inference, enabling users to run AI models without compromising data privacy through encrypted streaming and unlinkable requests.
Reasoning models don't degrade gracefully - they hit a complexity cliff and collapse entirely [Research Analysis] [R]
Reasoning models exhibit a stark performance drop: They maintain 85% accuracy until a complexity threshold, after which they collapse to near-random guessing by step 15, indicating a complexity cliff rather than gradual degradation.
Learning from failure to tackle hard problems
BaNEL (Bayesian Negative Evidence Learning) leverages failed attempts to train generative models, addressing the challenge of extremely sparse rewards in complex problem-solving scenarios, such as drug discovery and theorem proving.
LLMs Encode How Difficult Problems Are
LLMs encode problem difficulty in a manner that aligns with human judgment, revealing a strong linear decodability of human-labeled difficulty (AMC: $\rho \approx 0.88$) across various model sizes, while LLM-derived difficulty shows poor scaling.
The Parallel Search API
The Parallel Search API enables AIs to efficiently navigate the web, enhancing their ability to retrieve and process information in real-time, thus revolutionizing AI interactions with online data.
Show HN: TabPFN-2.5 – SOTA foundation model for tabular data
TabPFN-2.5 significantly enhances tabular AI, scaling to 20× data cells compared to its predecessor, and matches the accuracy of complex models like AutoGluon 1.4 while outperforming tuned tree-based models on industry benchmarks.
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
Brain-IT employs a Brain Interaction Transformer (BIT) to reconstruct images from fMRI data, achieving high fidelity with only 1 hour of recordings, comparable to methods requiring 40 hours.
[R][N] TabPFN-2.5 is now available: Tabular foundation model for datasets up to 50k samples
TabPFN-2.5 is a pretrained transformer that significantly enhances tabular data processing, now accommodating 50,000 samples × 2,000 features, a 5x increase from its predecessor.
Benchmarking the Most Reliable Document Parsing API
Tensorlake's Document Parsing API achieves 91.7% accuracy, surpassing competitors like Azure and AWS Textract by focusing on both structural preservation and usability for downstream applications.
Show HN: The Legal Embedding Benchmark (MLEB)
The Massive Legal Embedding Benchmark (MLEB) is the largest and most diverse benchmark for legal text embedding models, featuring 10 datasets that cover various document types, jurisdictions, and legal tasks, ensuring comprehensive evaluation of legal reasoning and domain knowledge. Learn more about MLEB.
[D] Trajectory Distillation for Foundation Models
Trajectory distillation offers a leaner alternative to traditional reinforcement learning (RL) for post-training foundation models, achieving comparable performance at a 10× lower cost.
KernelFalcon: Autonomous GPU Kernel Generation via Deep Agents
KernelFalcon is a pioneering deep agent architecture that autonomously generates GPU kernels, achieving 100% correctness across all 250 tasks in the KernelBench suite, utilizing a unique combination of hierarchical task decomposition and execution-based verification.
DeepInverse Joins the PyTorch Ecosystem: the library for solving imaging inverse problems with deep learning
DeepInverse is an open-source library that simplifies deep learning for imaging across various domains, including medical imaging and computational photography, by providing tools for image reconstruction and state-of-the-art neural networks.