Project Glasswing has identified over 10,000 critical vulnerabilities in essential software using Claude Mythos Preview, significantly enhancing the speed of vulnerability detection compared to traditional methods.
Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
The OpenSCAD LLM benchmark evaluated multiple AI coding tools on their ability to generate a detailed Pantheon model, revealing significant differences in output quality and speed among Codex 5.5 High, Claude Sonnet, and Google Antigravity 2.0, with Antigravity achieving the best autonomous result.
Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
Domain camouflaged injection attacks exploit the vocabulary and authority structures of target documents, leading to a dramatic drop in detection rates from 93.8% to 9.7% on Llama 3.1 8B and from 100% to 55.6% on Gemini 2.0 Flash, revealing a critical vulnerability in current detection systems.
NuExtract3 released: open-weight 4B VLM for Markdown, OCR and structured extraction (self-hostable)
NuExtract3 is a 4B model designed for information extraction from complex documents, including PDFs and forms, and is available under the Apache-2.0 license.
Live Human Detector on Outbound Phone Calls
The Live Human Detector aims to enhance call center efficiency by accurately identifying when a call transitions from an automated system to a live agent within a 1-2 second window, utilizing advanced audio classification techniques.
Towards Speed-of-Light Text Generation with Nemotron-Labs Diffusion Language Models
Nemotron-Labs Diffusion Language Models aim to achieve speed-of-light text generation, leveraging advanced diffusion techniques to enhance performance and efficiency in natural language processing tasks.
Tested chunking + embeddings data from 3 production websites.
Chunking and embeddings from three production websites reveal significant variance in content density, with Intercom achieving a 31% yield score while KPMG lags at 8%, indicating the effectiveness of content in retrieval tasks.
Alignment: Higher order prioritizing over constraints
Transformers exhibit a behavior termed "clarity seeking," where their inherent design prioritizes understanding meaning over adhering strictly to imposed constraints, suggesting a potential avenue for alignment and safety research.
LQS v3.1 — an open methodology for rating AI training data (multi-oracle consensus + signed certificates)
LQS v3.1 introduces a comprehensive rating methodology for AI training data, featuring 19 dimensions such as label correctness and oracle agreement, designed to enhance buyer confidence in data quality.