Diffusion models for images approximate autoregression in the frequency domain, revealing a closer relationship between diffusion and autoregressive models than previously understood, especially when analyzing visual data through spectral analysis.
AI-Implanted False Memories
Conversational AI, particularly generative chatbots using large language models (LLMs), significantly amplifies the formation of false memories in individuals during simulated crime witness interviews, compared to control, survey-based, and pre-scripted chatbot interactions.
Graph Language Models
The Graph Language Model (GLM) innovatively combines the strengths of traditional Language Models (LMs) and Graph Neural Networks (GNNs), enhancing the representation of structured knowledge graphs without sacrificing text feature quality.
Smaller, Weaker, yet Better: Training LLM Reasoners via Compute-Optimal Sampling
Training language models (LMs) on synthetic data from weaker, cheaper models can outperform data from stronger, expensive ones in improving reasoning performance, challenging conventional strategies.
Population Minimizer of The Categorical Cross Entropy Loss (a blog post)
The population minimizer of conditional risk associated with Categorical Cross Entropy is proven to be the true probability label distribution, a significant insight for neural network outputs.
Fine-tuning coding LLMs on Git histories rather than just final code?
Fine-tuning coding LLMs on Git histories could enhance software development by leveraging the evolutionary context of code, beyond its final state.