Block Diffusion: Interpolating between autoregressive and diffusion models
Block diffusion language models bridge the gap between autoregressive and diffusion models, enhancing flexible-length generation and inference efficiency through techniques like KV caching and parallel token sampling.
Transformers Without Normalization
Dynamic Tanh (DyT) serves as a drop-in replacement for normalization layers in Transformers, enabling models to achieve comparable or superior performance without the need for hyperparameter tuning.
Ask HN: Any insider takes on Yann LeCun's push against current architectures?
Yann LeCun argues that current LLM architectures will perpetually struggle with hallucinations due to runaway errors from token choice methods, suggesting a shift towards an 'energy minimization' architecture that evaluates the overall response energy during training.
Mayo Clinic's secret weapon against AI hallucinations: Reverse RAG in action
Mayo Clinic employs a reverse RAG technique to combat AI hallucinations, effectively linking extracted data back to its original sources, which has nearly eradicated retrieval-based inaccuracies in non-diagnostic scenarios.
[R] Transformers without Normalization (FAIR Meta, New York University, MIT, Princeton University)
Transformers without normalization can achieve equal or superior performance through a novel technique called Dynamic Tanh (DyT), which replaces traditional normalization layers with a simple element-wise operation, DyT(x)=tanh(αx).
Arbitrary-Scale Super-Resolution with Neural Heat Fields
Thera employs a hypernetwork to estimate pixel-wise parameters, ensuring aliasing-free super-resolution by utilizing local neural heat fields that adapt based on frequency and scaling factors.
Show HN: OCR Benchmark Focusing on Automation
Document processing automation is increasingly critical as new OCR solutions emerge, yet enterprises face challenges in discerning genuine capabilities from marketing claims, necessitating robust benchmarks for evaluation.
[R] How Pickle Files Backdoor AI Models—And What You Can Do About It
Pickle files can introduce vulnerabilities in AI models by allowing malicious code execution during the deserialization process, potentially compromising model integrity and security.
Show HN: Open-Source MCP Server for Context and AI Tools
The JigsawStack MCP Server enables AI models to access live search, OCR, and structured data extraction through a standardized interface, enhancing their capabilities beyond fixed context windows and outdated knowledge.
Owl: Optimized Workforce Learning for multi-agent collaboration
OWL is a pioneering framework for multi-agent collaboration in task automation, achieving a 58.18 average score on the GAIA benchmark, ranking it #1 among open-source frameworks.
[R] Block Diffusion: A Hybrid Language Model Combining Autoregressive and Diffusion Approaches for Flexible-Length Generation
Block Diffusion innovatively combines autoregressive and diffusion models by processing text in flexible blocks, achieving a 9.37 perplexity on C4 validation, which sets a new SOTA for diffusion language models.
[R] Multi-View Video Generation via View-Invariant Motion Learning and Cross-View Consistent Translation
Reangle-A-Video innovatively frames 4D video generation as a video-to-video translation problem, enabling the synthesis of arbitrary camera viewpoints from a single input video while ensuring temporal consistency.
AutoHete: An Automatic and Efficient Heterogeneous Training System for LLMs
AutoHete is an innovative training system that enhances the efficiency of heterogeneous training for large language models (LLMs) by dynamically optimizing resource allocation based on hardware and training needs.