US Copyright Office found AI companies breach copyright. Its boss was fired
The US Copyright Office has determined that AI companies often exceed fair use by utilizing copyrighted material without consent, leading to the dismissal of its head shortly after the report's release.
Byte Latent Transformer: Patches Scale Better Than Tokens
The Byte Latent Transformer (BLT) introduces a novel byte-level architecture that achieves tokenization-level performance while enhancing inference efficiency and robustness through the use of dynamically sized patches based on byte entropy.
Continuous Thought Machines
The Continuous Thought Machine (CTM) introduces a novel neural network architecture that incorporates neural timing and synchronization, bridging the gap between biological intelligence and modern AI efficiency, yielding promising results in various tasks.
Intellect-2 Release: The First 32B Model Trained Through Globally Distributed RL
INTELLECT-2 is the first 32B parameter model trained through globally distributed reinforcement learning, utilizing a unique asynchronous framework that allows for decentralized contributions from various compute sources, enhancing training efficiency and scalability.
Absolute Zero Reasoner
The Absolute Zero Reasoner (AZR) implements a novel paradigm that enables models to autonomously generate and solve tasks without relying on human-curated data, fostering continuous self-improvement through self-play.
[R] Continuous Thought Machines: neural dynamics as representation.
Continuous Thought Machines (CTMs) leverage neural dynamics as a core representation, challenging traditional deep learning by reintroducing temporal processing at the neuron level, which enhances the richness of neuron dynamics.
ParaQuery offers a GPU-accelerated Spark/SQL solution that claims to be more cost-efficient and performant than BigQuery, enabling users to save over 60% on their data processing costs while achieving 2x faster performance without data migration.
Why National Labs are investing (heavily) in AI
Jason Pruet, Director of the National Security AI Office at Los Alamos National Laboratory, emphasizes that AI is not merely a tool but a transformative force reshaping scientific inquiry and national security, driven by advancements in large AI models.
Build Your Own Siri. Locally. On-Device. No Cloud
Build your own local voice assistant that respects privacy by executing commands directly on your device without cloud dependency, utilizing models like LLaMA 3.1 and Whisper for seamless interaction.
SDFs and the Fast sweeping algorithm in Jax
The Fast Sweeping Method (FSM) efficiently solves the Eikonal equation in O(n) time, making it suitable for applications like signed distance functions (SDFs) in machine learning, where it represents 3D surfaces implicitly.
In-Memory Ferroelectric Differentiator
The in-memory ferroelectric differentiator utilizes ferroelectric domain dynamics to perform differential calculations directly within memory, significantly enhancing efficiency in data processing and reducing energy consumption to 0.24 fJ per operation.
[R] Zero-shot forecasting of chaotic systems (ICLR 2025)
Foundation models demonstrate zero-shot learning capabilities in forecasting chaotic systems, achieving competitive results against custom-trained models across 135 chaotic systems and 108 timepoints.
Toward a Sparse and Interpretable Audio Codec
This work presents a novel audio encoder that utilizes a sparse representation of audio events, enhancing interpretability compared to traditional codecs like MP3 and Ogg Vorbis.
[D] Simulating Bias with Bayesian Networks - Feedback wanted!
The project simulates bias in machine learning using a Bayesian network to analyze recruitment, revealing that any observed bias stems from the models rather than the training data itself.
[P] Llama 3.2 1B-Based Conversational Assistant Fully On-Device (No Cloud, Works Offline)
The Llama 3.2 1B Instruct model powers a privacy-first mobile assistant that operates entirely offline, ensuring user data remains secure and unclouded.