Training an LLM in Swift, Part 1: Taking matrix mult from Gflop/s to Tflop/s
Optimizing matrix multiplication in Swift can elevate performance from 2.8 Gflop/s to 1.1 Tflop/s, showcasing the potential of Apple Silicon's architecture for training Large Language Models (LLMs) without relying on external libraries.
Lakebase architecture achieves up to 5x faster write throughput for Postgres by offloading crash-recovery tasks to distributed storage, effectively eliminating traditional bottlenecks associated with Write-Ahead Logging (WAL).
Interfaze: A new model architecture built for high accuracy at scale
Interfaze is a new model architecture that significantly outperforms existing models like Gemini-3-Flash and Claude-Sonnet-4.6 across nine benchmarks in tasks such as OCR, vision, and structured output, achieving accuracy rates as high as 89.9% in specific tests.
Interaction Models
Interaction models enable real-time collaboration with AI by processing audio, video, and text simultaneously, enhancing responsiveness and intelligence without external scaffolding.
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
MISA (Mixture of Indexer Sparse Attention) enhances DeepSeek Sparse Attention by utilizing a mixture-of-experts approach, allowing for a significant reduction in computational cost while maintaining performance on long-context tasks.
Fast Byte Latent Transformer
The Byte Latent Transformer (BLT) introduces BLT Diffusion (BLT-D), enabling parallel byte generation, which significantly accelerates the autoregressive process compared to traditional methods.
A hackable compiler to generate efficient fused GPU kernels for AI models
The hackable LLM compiler generates efficient fused GPU kernels, achieving 1.11× faster performance than PyTorch eager and 1.20× faster than torch.compile on RTX 5090 for specific models like TinyLlama and Qwen2.5-7B.
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
The Memory-Efficient Looped Transformer (MELT) architecture innovatively decouples reasoning depth from memory consumption, allowing for efficient multi-step computation without the linear memory growth seen in traditional models like Ouro.
🤗MachinaCheck: Building a Multi-Agent CNC Manufacturability System on AMD MI300X
MachinaCheck is a multi-agent AI system that automates CNC manufacturability assessments, generating comprehensive reports in under 30 seconds from a STEP file, thus saving significant time and reducing errors in job feasibility analysis.
Looking for arXiv endorsement (cs.CV) to post my ViT positional embeddings paper
The paper titled "Positional Encodings in Vision Transformers" explores how various positional encoding schemes influence the internal representations of Vision Transformers, introducing the Spatial Similarity Distance Correlation (SSDC) metric to assess spatial structure in token representations.
Signals: finding the most informative agent traces without LLM judges
Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation
Trajectory-Shaped Discrete Flow Matching (TS-DFM) enhances text generation by replacing blind stochastic jumps with a guided energy compass, significantly improving coherence in fewer steps.
MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing
MAVEN introduces a blackboard-inspired framework that enhances LLMs by implementing an adversarial Skeptic-Researcher-Judge loop, enabling explicit role-decoupling for improved reasoning and interpretability.
Confidence-Aware Alignment Makes Reasoning LLMs More Reliable
CASPO (Confidence-Aware Step-wise Preference Optimization) enhances reasoning reliability in large language models by aligning token-level confidence with logical correctness, eliminating the need for separate reward models.
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective
The cumulative token IS ratio offers a theoretically sound solution to the bias-variance dilemma in LLM policy optimization, ensuring unbiased prefix corrections with lower variance than traditional full sequence ratios.