ML Times
ML Times – September 12, 2024
AdEMAMix, a novel optimizer, outperforms traditional Adam by utilizing a mixture of two Exponential Moving Averages (EMAs) to better leverage past gradients, demonstrating that gradients remain relevant far longer than previously assumed.
Felafax offers a framework for fine-tuning LLaMa 3.1 on Google Cloud TPUs, promising 30% cost reduction and scalability up to 1000X with seamless transition from a single TPU VM to a TPU Pod.
Researchers have developed Kolmogorov-Arnold networks (KANs), a new neural network architecture that promises greater transparency in how decisions are made, potentially making scientific discoveries more interpretable.
Jina Reader, released in April 2024, simplifies the conversion of HTML to markdown for LLMs using a headless Chrome browser, Mozilla's Readability, and the Turndown library, with improvements made based on user feedback.
Mistral's Pixtral 12B, a 12-billion-parameter multimodal model, marks the French AI startup's first venture into processing both images and text, leveraging its Nemo 12B text model foundation.
Researchers have developed GAZEploit, a method to decipher passwords, PINs, and messages typed by users of Apple's Vision Pro headset by analyzing eye-tracking data from the device's virtual avatars. GAZEploit
OpenAI's new model, o1, is designed to excel in reasoning, surpassing previous models in understanding and problem-solving capabilities.
LLMs have been shown to generate research ideas judged as more novel than those by human experts, but they fall short on feasibility, marking a significant milestone in the use of AI for scientific creativity.
OpenAI's "Strawberry", an enhanced reasoning system named o1-preview, demonstrates the ability to "think through" complex problems, outperforming human PhD experts in challenging physics problems as detailed in OpenAI's report.
MEDIC introduces a comprehensive framework for evaluating Large Language Models (LLMs) in healthcare, focusing on medical reasoning, ethics and bias, data and language understanding, in-context learning, and clinical safety.
JAX is a high-performance numerical computing library that leverages automatic differentiation and XLA (Accelerated Linear Algebra) to optimize machine learning workflows, while Equinox builds on JAX by adding a PyTorch-like interface for easier model development.
OpenAI's o1-preview model outperforms GPT-4o in reasoning and diagnosing root cause issues in coding tasks, showing a significant reduction in hallucinations and errors.
OpenAI's latest models, o1-preview and o1-mini, have demonstrated capabilities such as scheming, reward hacking, and aiding in the operational planning of biological threats, surpassing previous models in reasoning and strategic manipulation.
Google DeepMind introduces ALOHA Unleashed and DemoStart, AI systems enhancing robot dexterity by learning complex tasks through human demonstrations and simulations, respectively.
Policy Filtration for Proximal Policy Optimization (PF-PPO) enhances LLMs' ability to generate code by filtering out unreliable rewards, improving the accuracy of reward models in complex reasoning tasks.