Infrastructure set-up & open-source scripts to train a 70B model from bare metal
Imbue trained a 70B parameter model that surpassed GPT-4o in reasoning tasks, leveraging a custom-built infrastructure with 4,092 H100 GPUs across 511 computers, facilitated by partnerships with Voltage Park, Dell, H5, and NVIDIA.
Are Language Models Actually Useful for Time Series Forecasting?
Removing or replacing the LLM component in time series forecasting methods does not degrade results, often improving them instead, challenging the utility of LLMs in this domain.
Llama-agents: an async-first framework for building production ready agents
llama-agents is an async-first framework designed for building and deploying multi-agent systems, facilitating features like multi-agent communication and human-in-the-loop processes.
Interpretability research in LLMs
Most interpretability research in LLMs has pivoted towards mechanistic interpretability, diverging from traditional methods like counterfactuals and saliency maps.
Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence Segmentation
The Segment any Text (SaT) model introduces a new pretraining scheme to enhance robustness by reducing reliance on punctuation, addressing a common shortfall in existing sentence segmentation methods.
Deep Learning Paper Summaries
The Vision Language Group at IIT Roorkee has crafted detailed summaries of deep learning papers from 2016 to 2024, covering major conferences like NeurIPS, CVPR, ICCV, and ICML.
Efficient World Models with Context-Aware Tokenization
$\Delta$-IRIS, a new agent, leverages context-aware tokenization to encode changes between time steps, significantly reducing the computational load required for simulating environments in model-based RL.
YannicKilcher - Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (Paper Explained)
Researchers from Stanford and Yale have critically evaluated the accuracy of AI legal research tools, focusing on their propensity for "hallucinations"—the tendency to generate incorrect or misleading information.
Into the Omniverse: SyncTwin Helps Democratize Industrial Digital Twins With Generative AI, OpenUSD
SyncTwin GmbH leverages OpenUSD and NVIDIA's technologies to create digital twins that optimize industrial efficiency and sustainability.
Fine-tuning retrieval models (DeBERTa/RoBERTa/e5) for biomedical/STEM: Seeking advice on unsupervised fine tuning, query/instruct formatting and loss functions
The individual is fine-tuning DeBERTa models for medical/STEM knowledge retrieval, exploring configurations and strategies to optimize performance, including unsupervised fine-tuning with TSDAE and supervised fine-tuning with various loss functions.
The Remarkable Robustness of LLMs: Stages of Inference?
Deleting and swapping layers in Large Language Models ( LLMs) retains 72-95% of the original model's prediction accuracy, suggesting remarkable robustness without the need for fine-tuning.
EmPO: Theory-Driven Dataset Construction for Empathetic Response Generation through Preference Optimization
The paper introduces a novel approach for enhancing empathetic response generation in conversational agents by constructing theory-driven preference datasets and aligning large language models (LLMs) with preference optimization algorithms.
T-FREE: Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
T-FREE directly embeds words using sparse activation patterns over character triplets, eliminating the need for traditional tokenizers and their associated limitations.
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
Finetuning Large Language Models (LLMs) on a synthetic dataset significantly enhances their ability to retrieve information and reason over long-context inputs, as demonstrated in experiments with GPT-3.5 Turbo and Mistral 7B.
Jump Starting Bandits with LLM-Generated Prior Knowledge
Integrating Large Language Models (LLMs) with Contextual Multi-Armed Bandit frameworks significantly reduces online learning regret by simulating human behaviors for personalized recommendations.