Circuit Tracing: Revealing Computational Graphs in Language Models (Anthropic)
Circuit tracing reveals the inner workings of language models by constructing attribution graphs that detail how features interact to produce outputs, utilizing a cross-layer transcoder to enhance interpretability and accuracy in understanding model behavior.
Show HN: Qwen-2.5-32B is now the best open source OCR model
The Omni OCR Benchmark evaluates the OCR and data extraction capabilities of various large multimodal models, aiming to provide a comprehensive assessment of OCR accuracy across traditional providers and multimodal language models, with all methodologies being open source.
UCSD: Large Language Models Pass the Turing Test
GPT-4.5 achieved a 73% identification rate as human in Turing tests, outperforming human participants, marking a significant milestone in AI evaluation.
How Google built its Gemini robotics models
Gemini Robotics models enable robots to learn complex tasks such as preparing salads and performing a "slam dunk" on their first attempt, showcasing a significant leap in robotic capabilities.
AI image recognition detects bubble-like structures in the universe
AI image recognition has successfully identified previously unrecorded bubble-like structures in the Milky Way, enhancing our understanding of star formation and galaxy evolution.
[R] Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
Current LLMs struggle with rigorous mathematical reasoning, achieving less than 5% accuracy on the 2025 USAMO problems, indicating a significant gap in their capabilities compared to human competitors.
[D][P] Turning Knowledge Graphs into Memory with Ontologies?
Integrating ontologies into the cognee AI memory tool enhances the grounding of knowledge graphs by aligning them with external system rules through RDF + OWL frameworks.
Practical Prolog Planner Prompting
Large Language Models (LLMs) enhance Prolog's combinatorial search capabilities, effectively bridging the gap between natural language processing and automated planning, as evidenced by their application in generating Prolog planners for complex logistics problems.
Speed Demon: NVIDIA Blackwell Takes Pole Position in Latest MLPerf Inference Results
NVIDIA Blackwell achieved record-breaking performance in the latest MLPerf Inference V5.0, showcasing the GB200 NVL72 system with up to 30x higher throughput on the Llama 3.1 405B benchmark compared to previous submissions.
[D] Relevance of Minimum Description Length to understanding how Deep Learning really works
Minimum Description Length (MDL) offers insights into the underlying mechanisms of deep learning, addressing phenomena like overparameterization and double descent that remain poorly understood.
[Project] AxiomGPT – programming with LLMs by defining Oracles in natural language
AxiomGPT revolutionizes programming by allowing users to define Oracles in natural language, enabling a new paradigm of latent-space programming that treats language as an invocation rather than mere instruction.
[R] Neuron-based explanations of neural networks sacrifice completeness and interpretability (TMLR 2025)
Neuron-based explanations of neural networks often lack completeness and interpretability, as demonstrated by the finding that principal components yield superior insights.
[D] What are the current challenges in deepfake detection (image)?
Current challenges in deepfake detection include the need for improved dataset generalization and the identification of specific problems that remain unsolved, particularly in image detection.
[D] [P] We created a Transcription API with an open-source, multi-step, multi-modal approach instead of custom models. The result? No.1 in an accuracy benchmark (You can recreate the benchmark).
The Salad Transcription API achieved a 95.1% accuracy for English and 96.8% for German, outperforming competitors by utilizing an open-source, multi-modal approach rather than custom models.
Evaluating potential cybersecurity threats of advanced AI
AI's role in cybersecurity is evolving, as advanced models can automate defenses but also pose risks of misuse in cyberattacks, necessitating a comprehensive evaluation framework to address these threats effectively.