The Image-based Joint-Embedding Predictive Architecture (I-JEPA) enables the learning of semantic image representations without hand-crafted data-augmentations, focusing on predicting target blocks from a single context block.
We hacked Gemini's Python sandbox and leaked its source code (at least some)
Google's A.I. Gemini source code was partially leaked during a bug bounty event, revealing internal structures and proprietary code, including sensitive Google3 directories and internal proto files that should not have been exposed.
Bolt Graphics Zeus a New GPU Architecture with Up to 2.25TB of Memory and 800GbE
Bolt Graphics' Zeus architecture features up to 2.25TB of memory and 800GbE connectivity, utilizing a RISC-V core for enhanced performance and flexibility in GPU design.
Optimizing Matrix Multiplication on RDNA3
The optimized FP32 matrix multiplication on AMD RDNA3 GPU achieves a 60% performance increase over rocBLAS, demonstrating significant improvements through iterative kernel enhancements and architectural insights.
[R] Anthropic: On the Biology of a Large Language Model
Claude 3.5 Haiku demonstrates advanced capabilities, such as multi-step reasoning and poetic planning, revealing its internal processes through attribution graphs, which enhance our understanding of model behavior.
Low responsiveness of ML models to critical or deteriorating health conditions
Machine learning models exhibit significant deficiencies in recognizing critical health conditions, failing to identify 66% of severe injuries in mortality predictions, highlighting the urgent need for improved model responsiveness in clinical settings.
Plasmonic Modulators Can Break the Wireless Terahertz Barrier
Plasmonic modulators enable the transfer of signal information from electrical to optical waves at unprecedented speeds, potentially revolutionizing 6G networks and AI data centers.
[R] DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
DeltaProduct enhances state-tracking in linear RNNs by utilizing multiple steps of online gradient descent per token, leading to improved expressivity without sacrificing efficiency.
[R] Enhancing GUI Agent Reasoning Through Rule-Based Reinforcement Learning
UI-R1 innovatively combines rule-based reinforcement learning with large language models to enhance GUI agents, enabling them to learn from mistakes and adapt to new interfaces effectively.
[R] Evaluating Multi-Step Spatial Reasoning in MLLMs Through LEGO-Based Visual Tasks
The LEGO-Puzzles benchmark evaluates multimodal language models (MLLMs) on spatial reasoning tasks, revealing a significant performance gap between human (85.8%) and AI (59.8%) capabilities, particularly in complex scenarios.