LiteLLM Python package compromised by supply-chain attack
The litellm==1.82.8 package on PyPI contains a malicious .pth file that executes a credential-stealing script upon Python interpreter startup, compromising sensitive data without requiring an import statement.
Epoch confirms GPT5.4 Pro solved a frontier math open problem
A Ramsey-style problem on hypergraphs seeks to construct large hypergraphs without a specific property, with a recent solution confirmed by AI models, notably GPT-5.4 Pro, which improved the understanding of lower bounds in hypergraph theory.
LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language?
RYS method demonstrates that relayering in modern LLMs like Qwen3.5-27B enhances performance, confirming that this technique is not merely a quirk of earlier models but a robust feature of Transformer architecture.
MSA: Memory Sparse Attention
Memory Sparse Attention (MSA) offers a scalable, end-to-end trainable framework that efficiently handles 100M-token contexts, achieving <9% degradation from 16K to 100M tokens through innovative techniques like sparse attention and document-wise RoPE.
Finding all regex matches has always been O(n²)
Regex engines universally exhibit O(n²) complexity when finding all matches, despite claims of linear performance; this issue has persisted since the 1970s, affecting all languages and implementations.
Show HN: Gemini can now natively embed video, so I built sub-second video search
SentrySearch enables semantic search over dashcam footage, utilizing Google's Gemini Embedding model to convert video into a searchable vector space, allowing users to retrieve specific clips based on text queries.
Arm AGI CPU
Arm AGI CPU marks a pivotal advancement in AI infrastructure, delivering breakthrough performance and efficiency tailored for agentic AI workloads, enabling real-time decision-making across distributed systems.
Run a 1T parameter model on a 32gb Mac by streaming tensors from NVMe
Hypura is a storage-tier-aware LLM inference scheduler designed for Apple Silicon, enabling the execution of models larger than the system's memory by intelligently distributing tensors across GPU, RAM, and NVMe based on usage patterns and hardware capabilities.
ARM AGI CPU: Specs and SKUs
ARM AGI CPU is Arm's inaugural production silicon, engineered for AI infrastructure with up to 136 Neoverse V3 cores and a 3nm process, enhancing performance and density for modern data centers.
Designing AI for Disruptive Science
Scaling AI does not guarantee paradigm shifts; true innovation requires AI to generate new conceptual frameworks rather than merely enhancing existing predictive capabilities.
Welcome to FastMCP
FastMCP is the premier framework for developing MCP applications, enabling seamless integration of LLMs with tools and data while automating schema, validation, and documentation processes.
[R] Causal self-attention as a probabilistic model over embeddings
Support tokens emerge from a new interpretation of causal self-attention transformers, revealing a stability-margin analogous to support vector machines, which enhances the robustness of large language models (LLMs).
[R] VLouvain: Louvain Community Detection Directly on Vectors, No Graph Construction
VLouvain enables community detection directly on embedding matrices, eliminating the need for graph construction and reducing complexity from O(n²) to O(n*d), thus maintaining performance even with large datasets.
PyTorch 2.11 Release Blog
PyTorch 2.11 introduces significant enhancements such as Differentiable Collectives for Distributed Training and FlexAttention with FlashAttention-4, which promise to improve performance and flexibility in deep learning workflows.
[D] Modeling online discourse escalation as a state machine (dataset + labeling approach)
The proposed framework models online discourse escalation as a state machine, identifying seven distinct states from Neutral to Threats of violence, with each comment labeled according to its local state while the thread evolves globally.