An autonomous agent exploited a SQL injection vulnerability in McKinsey's AI platform, Lilli, gaining full access to sensitive data within two hours without any insider knowledge or credentials.
BitNet: 100B Param 1-Bit model for local CPUs
bitnet.cpp is a cutting-edge inference framework for 1-bit LLMs, achieving 1.37x to 6.17x speedups on various CPU architectures while significantly reducing energy consumption by up to 82.2%.
Google closes deal to acquire Wiz
Wiz officially joins Google, enhancing its mission to secure cloud environments at the speed of AI, ensuring organizations can protect their innovations without sacrificing speed.
Many SWE-bench-Passing PRs would not be merged
Approximately 50% of test-passing SWE-bench Verified PRs generated by AI agents from mid-2024 to late-2025 would not be merged into the main branch, highlighting a significant gap between benchmark performance and real-world applicability.
Kotlin creator's new language: a formal way to talk to LLMs instead of English
CodeSpeak is a revolutionary programming language that leverages LLMs to reduce codebases by 5-10x, enhancing maintainability through specification rather than traditional coding.
Show HN: Open-source browser for AI agents
Agent Browser Protocol (ABP) transforms web browsing into a step machine, allowing agents to interact with a stable, frozen state of the web, enhancing automation efficiency.
[D] ICML paper to review is fully AI generated
A paper submitted to ICML is entirely AI-generated, raising concerns about adherence to guidelines that prohibit LLM assistance in writing or reviewing submissions.
Reliable Software in the LLM Era
LLMs enhance code generation but complicate validation, as they produce text that appears correct, making it challenging to ensure reliability in software development.
Are LLM merge rates not getting better?
LLMs have not improved in programming abilities for over a year, as evidenced by a lack of increase in merge rates since early 2025, contradicting claims of ongoing advancements in AI capabilities.
Show HN: Autoresearch_at_home – SETI_at_home but for LLM training
Ensue facilitates a collaborative research environment where AI agents share GPU resources, enhancing the capabilities of language models through collective effort.
Preliminary data from a longitudinal AI impact study
AI productivity gains are only ~10%, contrary to inflated claims of 2-3x, based on a longitudinal study analyzing 40 companies from November 2024 to February 2026, revealing a 65% increase in AI usage but only a 9.97% rise in pull request throughput.
High fidelity font synthesis for CJK languages
zi2zi-JiT is a conditional variant of JiT that enables Chinese font style transfer by synthesizing characters based on a source and a style reference, utilizing a novel architecture with a Content Encoder, Style Encoder, and Multi-Source In-Context Mixing.
IonRouter delivers high throughput and low-cost inference through its IonAttention engine, achieving a throughput of 7,167 tokens per second on a single GH200 GPU, significantly outperforming traditional inference providers which average around 3,000 tokens per second.
Lf-lean: The frontier of verified software engineering
lf-lean achieves a remarkable 350x speed-up in verified software engineering, translating 1,276 statements from Rocq to Lean with minimal human effort, demonstrating the potential of task-level specification generators to automate correctness verification across complex codebases.
The Biggest Identity Sandpiles and How to Compute Them
The largest identity sandpile computed is 16384 by 16384, achieved in under an hour, significantly faster than the previous record of 10 days for a 10,000 by 10,000 sandpile, showcasing advancements in computational methods.