ML Times

Sep 2, 2024

Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer

Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs

AI-Implanted False Memories

Population Minimizer of The Categorical Cross Entropy Loss (a blog post)

Fine-tuning coding LLMs on Git histories rather than just final code?

MemLong: Memory-Augmented Retrieval for Long Text Modeling

Investigating Neuron Ablation in Attention Heads: The Case for Peak Activation Centering

A Unified Benchmark for Federated Unsupervised Anomaly Detection in Tabular Data

Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning

EMPOWER: Embodied Multi-role Open-vocabulary Planning with Online Grounding and Execution

Dynamic Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling

Safety Layers of Aligned Large Language Models: The Key to LLM Security

UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches