- Reasoning in Large Language Models: A Geometric Perspective
The expressive power of large language models ( LLMs) is intricately linked to the density of their self-attention graphs, which determines the intrinsic dimension of inputs, enhancing their reasoning capabilities.
- LivePortrait: A fast, controllable portrait animation model
LivePortrait introduces an efficient method for animating portraits, leveraging advanced stitching and retargeting techniques to control animations dynamically.
- Show HN: A fast OSS voice assistant
Swift is a fast, open-source voice assistant that leverages the power of Groq, Cartesia, VAD, and Vercel for enhanced speech detection capabilities.
- [R] Learning to (Learn at Test Time): RNNs with Expressive Hidden States
The Test-Time Training (TTT) layers introduce a novel approach by making the hidden state a machine learning model itself, enhancing expressiveness and performance in long contexts.
- Micro-agent: make an AI write code until it passes an unit test
Micro Agent is an AI tool designed to write and iteratively fix code until all test cases pass, streamlining the development process by automating code generation and correction.
- [R] An Empirical Study of Mamba-based Language Models (8B Mamba-2-Hybrid on 3.5T tokens data)
Mamba-based models, particularly the 8B Mamba-2-Hybrid, outperform 8B Transformer models across 12 standard tasks by an average of +2.65 points, showcasing superior efficiency and effectiveness in language modeling. Study details
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
Test-Time Training (TTT) layers introduce a novel approach by making the hidden state a machine learning model itself, enhancing expressiveness and performance in long contexts.
- [R] A Universal way to Jailbreak LLMs' safety inputs and outputs if provided a Finetuning API
A researcher has developed a universal method to bypass safety checks in Large Language Models (LLMs) using a finetuning API, employing a Caesar Cipher with 25 shifts to encode harmful instructions that the LLMs' safety mechanisms fail to detect.
- [R] What is GraphRAG? Explained
GraphRAG represents an advancement over the baseline RAG by utilizing Knowledge Graphs for retrieval, which enhances the quality of its output.
- [P] ReproModel: Open Source ML Research Toolbox Update!
ReproModel, an open-source toolbox, aims to simplify the testing and reproduction of machine learning models, addressing common issues like missing code and unclear experiment parameters.
- [Research] Neural decoding - mapping EEG data of song listening to respective audio files
The research aims to reconstruct songs participants listened to by mapping EEG data to audio files using a regression model, with a CNN-based approach detailed in a study.
- TwoMinutePapers - DeepMind’s New AI Found The Sound Of Pixels!
DeepMind's new AI technique synthesizes sound by analyzing video content, a significant leap towards more immersive AI-generated media.
- In It for the Long Haul: Waabi Pioneers Generative AI to Unleash Fully Driverless Autonomous Trucking
Waabi is leveraging generative AI and NVIDIA's DRIVE Thor technology to revolutionize autonomous trucking, aiming for fully driverless operations by next year.
- YannicKilcher - Scalable MatMul-free Language Modeling (Paper Explained)
Researchers from UC Santa Cruz, Sucha University, UC Davis, and Loxy Tech have developed a scalable, MatMul-free language model that replaces matrix multiplication in large language models with ternary accumulators and a parallelizable form of a ternary recurrent network, aiming for greater hardware efficiency.
- ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
The paper introduces ANAH-v2, an iterative self-training framework designed to scale and improve the accuracy of hallucination annotation in large language models (LLMs) by leveraging the Expectation Maximization algorithm.