Semble is a code search library that enables agents to retrieve relevant code snippets instantly, utilizing ~98% fewer tokens than traditional methods like grep+read, achieving indexing in under a second and query responses in approximately 1.5 ms.
GenCAD
GenCAD is an innovative image-conditional CAD generation model that produces both 3D CAD and the complete parameterized CAD command history, enhancing the design process significantly.
Project Glasswing: what Mythos showed us
Mythos Preview represents a significant advancement in vulnerability detection, enabling the construction of exploit chains and proof generation, which enhances its ability to identify and validate security flaws effectively.
Anthropic acquires Stainless
Anthropic's acquisition of Stainless enhances its capabilities in SDK generation and agent connectivity, crucial for the evolving landscape of AI where agents must act effectively within their environments.
Reviving PapersWithCode (by Hugging Face)
Hugging Face is revitalizing PapersWithCode, utilizing AI to parse high-impact papers and generate leaderboards, ensuring the platform remains a vital resource for the ML community.
Voice AI Systems Are Vulnerable to Hidden Audio Attacks
Hidden audio signals can manipulate AI voice systems, demonstrating vulnerabilities that allow unheard sounds to alter model behavior, posing significant security risks.
Agora-1: The Multi-Agent World Model
Agora-1 introduces a multi-agent world model that allows up to four participants, human or AI, to interact in a shared, real-time simulation, enhancing experiences across various fields like gaming and robotics.
Cutting inference cold starts by 40x with LP, FUSE, C/R, and CUDA-checkpoint
Modal's engineering innovations have reduced inference cold start times by 40x, enabling rapid scaling of AI applications through techniques like cloud buffers, custom filesystems, and checkpoint/restore methods for both CPU and GPU memory.
Fabricked: Misconfiguring Infinity Fabric to Break AMD SEV-SNP
Fabricked is a novel attack vector targeting machine learning models, exploiting vulnerabilities in their training data to manipulate outcomes.
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
Recent LLM architectures like Gemma 4 and DeepSeek V4 are innovating with techniques such as KV sharing and compressed attention to enhance long-context efficiency, significantly reducing memory and computational costs.
Sub-JEPA: a simple fix to LeCun group's LeWorldModel that consistently improves performance
Sub-JEPA enhances LeCun's LeWorldModel by applying Gaussian regularization within multiple frozen random orthogonal subspaces, improving performance on low-dimensional tasks without introducing new hyperparameters.
Sense Humans with WiFi – Ruview
Cognitum.One offers advanced intelligence solutions tailored for real-world applications, enhancing decision-making through data-driven insights.
A Rust-Python thing I am working on. Apache 2 licence
Nairobi OS is a high-performance, distributed data science infrastructure that processes massive datasets efficiently in constrained environments, utilizing a Rust-based refinery daemon and advanced kernel features for optimal performance.
Alignment pretraining: AI discourse creates self-fulfilling (mis)alignment
Discourse on AI significantly influences LLM behavior, with negative narratives leading to self-fulfilling misalignment; upsampling misalignment discourse increases misaligned behavior, while aligned discourse reduces misalignment scores from 45% to 9%.
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
AIRA-Compose and AIRA-Design enable LLM agents to autonomously create advanced neural architectures, yielding 14 new models that outperform Llama 3.2 and Composer baselines in accuracy and efficiency.