ML Times
Jul 9, 2024
MobileLLM: Optimizing sub-billion parameter language models for on-device use, integrating SwiGLU activation, deep and thin architectures, embedding sharing, and grouped-query attention, achieving up to 4.3% accuracy improvement over state-of-the-art models on zero-shot commonsense reasoning tasks. Read the paper
Surprising Gender Biases in GPT: The OSF (Open Science Framework) platform has implemented Google Tag Manager and Sentry for enhanced tracking and error logging, respectively, to improve site functionality and user experience.
turbopuffer: A cost-efficient, high-performance search engine that leverages object storage and smart caching, designed to scale to billions of vectors and millions of tenants/namespaces, addressing the high costs and operational challenges of existing search solutions.
[R] Learning to (Learn at Test Time): RNNs with Expressive Hidden States. The Test-Time Training (TTT) layers introduce a novel approach by making the hidden state a machine learning model itself, enhancing expressiveness and performance in long contexts.
LightRAG: The PyTorch Library for Large Language Model Applications. LightRAG is designed to assist developers in creating and optimizing Retriever-Agent-Generator (RAG) pipelines for a wide range of applications, from chatbots to text classification, emphasizing its lightweight, modular, and robust architecture.
LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages. LLaMAX enhances translation capabilities of Large Language Models (LLMs) to support over 100 languages, focusing on improving performance in low-resource languages through extensive multilingual continual pre-training.
[R] What is GraphRAG? Explained: GraphRAG represents an advancement over the baseline RAG by utilizing Knowledge Graphs for retrieval, which enhances the quality of its output.
[D] Inference-time gradients: Diffusion models excel in inference-time gradients by predicting score/noise/gradient to reverse the diffusion process, notably through rectified flow in Stable Diffusion 3.
[Research] Neural decoding - mapping EEG data of song listening to respective audio files: The research aims to reconstruct songs participants listened to by mapping EEG data to audio files using a regression model, with a CNN-based approach detailed in a study.
Multi-Object Hallucination in Vision-Language Models: Large vision language models (LVLMs) exhibit increased object hallucination when processing multiple objects, revealing a significant challenge in accurately recognizing and reasoning about complex visual scenes.
[R] Literature on Lipsync/Body Gestures: The inquiry seeks research papers on body gesture and lipsync training, aiming to understand the methodologies behind technologies like Heygen and Synthesia.
[R] How GraphRAG works? Explained: GraphRAG enhances RAG (Retrieval-Augmented Generation) by improving the connection of external documents to Large Language Models (LLMs), facilitating more informed and context-rich responses.
TwoMinutePapers - DeepMind’s New AI Found The Sound Of Pixels! DeepMind's new AI technique synthesizes sound by analyzing video content, a significant leap towards more immersive AI-generated media.
In It for the Long Haul: Waabi is leveraging generative AI and NVIDIA's DRIVE Thor technology to revolutionize autonomous trucking, aiming for fully driverless operations by next year.
YannicKilcher - Scalable MatMul-free Language Modeling (Paper Explained): Researchers from UC Santa Cruz, Sucha University, UC Davis, and Loxy Tech have developed a scalable, MatMul-free language model that replaces matrix multiplication in large language models with ternary accumulators and a parallelizable form of a ternary recurrent network, aiming for greater hardware efficiency.