OpenDevin: An Open Platform for AI Software Developers as Generalist Agents
OpenDevin is introduced as a platform for developing AI agents that can write code, use command lines, and browse the web, aiming to mimic the capabilities of human software developers.
GIL Become Optional in Python 3.13
Python 3.13 introduces an experimental feature allowing the Global Interpreter Lock (GIL) to be disabled, enhancing thread concurrency.
Improving Alignment and Robustness with Circuit Breakers
Circuit breakers in AI interrupt models to prevent harmful outputs, offering a novel approach beyond traditional refusal and adversarial training methods.
1.5-Pints Technical Report: Pretraining in Days, Not Months -- Your Language Model Thrives on Quality Data (2408.03506)
The 1.5-Pints Language Model pre-trains in just 9 days, leveraging a 57 billion token dataset to surpass benchmarks set by tech giants like Apple and Microsoft, focusing on quality data for enhanced instruction-following capabilities.
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
The Tree Attention algorithm leverages a novel scalar energy function for efficient self-attention computation, offering a Bayesian interpretation and linking to energy-based models like Hopfield Networks.
Segment Anything Model and Friends
The Segment Anything Model (SAM) and its variants, introduced in Segment Anything, 2023, aim to enable zero-shot generalization in image segmentation by leveraging a broad dataset and a task formulation that supports flexible prompting.
Gaussian Splatting Slam [CVPR 2024]
Gaussian Splatting SLAM introduces a novel application of 3D Gaussian Splatting for incremental 3D reconstruction using a single moving camera, achieving live performance at 3fps and enabling high-quality rendering and accurate tracking.
New Llama scaling laws?
Llama 3's development revealed that model performance improves log-linearly with training on up to 15T tokens, surpassing the previously optimal 200B tokens for an 8B parameter model.
🤗Tool Use, Unified
The unified tool use API now allows for seamless integration of tools across various LLM platforms like Mistral, Cohere, NousResearch, and Llama, with Transformers library offering enhanced support for tool calling, complete with documentation and examples.
New Apache Airflow Operators for Google Generative AI
Google Cloud introduces new Apache Airflow operators for seamless integration with Vertex AI's generative models, enhancing data analytics pipelines with advanced AI capabilities. Read more
Show HN: Simple Science – The Newest Science Explained Simply
Electric drills cause more brain damage than hand drills in surgical procedures, according to a recent study.
Transformers are Universal In-context Learners
Transformers can approximate continuous in-context mappings with arbitrary precision, handling an infinite number of context tokens without increasing the embedding dimension or the number of heads.
Multi-Agent Imitation Learning: Value is Easy, Regret is Hard
Multi-agent imitation learning (MAIL) faces a unique challenge: while matching an expert's behavior can close the value gap, it fails to address the regret gap caused by strategic deviations of agents.
Workshop: Scalable MatMul-free Language Modeling
MatMul operations, which dominate the computational cost of large language models, can be eliminated while maintaining performance at billion-parameter scales, as demonstrated by experiments up to 2.7B parameters.
🤗Welcome FalconMamba: The first strong attention-free 7B model
Falcon Mamba is the first strong attention-free 7B model that overcomes sequence scaling limitations without sacrificing performance, utilizing the Mamba architecture with added RMS normalization layers for stable scaling.