AudioFlux: A C/C++ library for audio and music analysis
audioFlux is a deep learning library designed for audio and music analysis, offering extensive feature extraction capabilities across various time-frequency analysis methods.
Grok-2 Beta Release
Grok-2 outperforms Claude 3.5 Sonnet and GPT-4-Turbo on the LMSYS leaderboard, showcasing advancements in chat, coding, and reasoning capabilities.
MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images
MVSplat introduces an efficient model that predicts clean feed-forward 3D Gaussians from sparse multi-view images, leveraging a cost volume representation for accurate localization.
ARPA-H announces awards to develop novel technologies for precise tumor removal
ARPA-H's PSI program awards $150 million to develop technologies for precise tumor removal, aiming to reduce repeat surgeries and unintentional injuries during cancer operations. PSI program page
[R] The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
The AI Scientist framework enables frontier large language models to independently conduct scientific research, generating novel ideas, writing code, executing experiments, and communicating findings through full scientific papers, including a simulated review process. Paper
"Mutual Reasoning" improves GSM8K accuracy from 13% to 64% [R]
rStar, a self-play mutual reasoning approach, significantly enhances the reasoning capabilities of small language models (SLMs) by employing a novel generation-discrimination process without the need for fine-tuning or larger models.
Wet-lab innovations will lead the AI revolution in biology
Abhishaike Mahajan specializes in biology posting on Substack, offering insights into biology and machine learning through essays.
[R] Grokfast: Accelerated Grokking by Amplifying Slow Gradients
Grokfast significantly accelerates the grokking phenomenon in machine learning by focusing on amplifying slow-varying gradient components, which are crucial for generalization. Read the paper
[R] Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Gemma Scope introduces an open suite of JumpReLU Sparse Autoencoders (SAEs) trained across various layers of Gemma 2 models, aiming to democratize access to advanced neural network interpretability tools.
[R] New Paper on Mixture of Experts (MoE) 🚀
The new paper on Mixture of Experts (MoE) delves into advancements in balancing computational efficiency with high performance in AI systems, addressing current challenges and future directions.
[P] New open-source release: SOTA multimodal embedding models for fashion
Marqo-FashionCLIP & Marqo-FashionSigLIP, new open-source multimodal models, have outperformed existing SOTA models like FashionCLIP2.0 and OpenFashionCLIP on 7 fashion evaluation datasets including DeepFashion and Fashion200K by up to 57%.
LLMs as Optimizers - Theory Paper Recommendation [R]
Transformers are theorized to perform first-order optimization, such as gradient descent, expanding their utility beyond traditional applications.
[R] Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness
The proposed multi-resolution input representations and dynamic self-ensembling technique enhances adversarial robustness by leveraging intermediate layer predictions, which are naturally resistant to attacks.
[D] opinions on keras 3 with jax
The interest in JAX and its ecosystem, including Flax (linen and nnx), stems from its mathematical simplicity and functional programming approach, contrasting with TensorFlow's complexity.
Decoding NVIDIA Edify — The Technology That Helps Developers Create Custom Models Trained on Their Data
NVIDIA Edify enables developers to create custom generative AI models using their own data, supporting a wide range of content including images, videos, and 3D assets.