Prolog enhances LLM reasoning by serving as an intermediate language that simplifies the generation of code for symbolic reasoning tasks, allowing models to leverage its declarative nature for improved logic processing.
Grandmaster-Level Chess Without Search
Grandmaster-Level Chess is achieved through a 270M parameter transformer model trained on 10 million chess games, utilizing 15 billion data points annotated by Stockfish 16, demonstrating that strong performance emerges only at sufficient scale.
D PyTorch 2.5.0 released!
PyTorch 2.5 introduces a new CuDNN backend for SDPA, enhancing performance on H100 GPUs, and features regional compilation to minimize cold start times for repeated modules, significantly benefiting large models like transformers.
Bugs in LLM Training – Gradient Accumulation Fix
Unsloth's recent fix for gradient accumulation addresses a critical bug that inflated loss calculations during LLM training, ensuring accurate training runs and loss metrics.
Microsoft BitNet: inference framework for 1-bit LLMs
bitnet.cpp is a cutting-edge inference framework for 1-bit LLMs, achieving 1.37x to 6.17x speedups on various CPU architectures while significantly reducing energy consumption by up to 82.2%.
P How to build a custom text classifier without days of human labeling
Custom text classifiers can be built efficiently by leveraging LLMs for auto-labeling datasets, significantly reducing the need for extensive human labeling while maintaining high accuracy.
LLMD: A Large Language Model for Interpreting Longitudinal Medical Records
LLMD is a large language model specifically designed to analyze longitudinal medical records, leveraging a vast corpus of data collected over an average of 10 years and 140 care sites per patient, enhancing the accuracy of patient health assessments.
R DART can generate high-quality human motions in real-time, achieving over 300 frames per second on a single RTX 4090 GPU! It combines text inputs with spatial constraints, allowing for tasks like reaching waypoints and interacting with scenes.
DART achieves over 300 frames per second on a single RTX 4090 GPU, enabling real-time generation of high-quality human motions by integrating text inputs with spatial constraints for tasks like waypoint navigation and scene interaction.
PyTorch 2.5 Release Blog
PyTorch 2.5 introduces a new CuDNN backend for SDPA, offering up to 75% speedup on H100 GPUs, alongside enhancements like regional compilation for reduced cold start times and improved TorchInductor performance with FP16 support.
An Evolved Universal Transformer Memory
Neural Attention Memory Models (NAMMs) enhance transformer efficiency by learning to manage memory, allowing for improved performance without the need for hand-designed context rules.
LLMOPT: Learning to Define and Solve General Optimization Problems from Scratch
LLMOPT introduces a unified learning-based framework that automates the formulation and solving of optimization problems from natural language, enhancing generalization across diverse problem types.
SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction
SimLayerKV effectively reduces inter-layer KV cache redundancies by identifying and dropping cache from "lazy" layers, which contribute less to long-range dependencies in large language models (LLMs).
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
The Chain-of-Embedding (CoE) method allows LLMs to conduct output-free self-evaluation by analyzing the differences in their latent thinking paths during inference, enhancing reliability in response correctness estimation.
Cerberus: Efficient Inference with Adaptive Parallel Decoding and Sequential Knowledge Enhancement
Cerberus introduces an adaptive parallel decoding framework that utilizes a gating mechanism, allowing large language models (LLMs) to select optimal decoding strategies dynamically, enhancing both speed and accuracy.
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
Summary-Guided Decoding (SGD) effectively mitigates hallucinations in Large Vision-Language Models (LVLMs) by prioritizing image information over linguistic priors, thus enhancing response accuracy.