Improving Recommendation Systems and Search in the Age of LLMs
Recommendation systems are evolving by integrating large language models (LLMs) and multimodal content, enhancing their ability to address cold-start and long-tail item challenges through hybrid architectures that combine content understanding with behavioral modeling.
Qwen2.5-VL-32B: Smarter and Lighter
Qwen2.5-VL-32B-Instruct enhances human-like responses and excels in mathematical reasoning and image understanding, outperforming previous models and competitors in multimodal tasks.
Bitter Lesson is about AI Agents
Raw computing power consistently outperforms intricate human-designed solutions in AI, as demonstrated by Richard Sutton's essay, The Bitter Lesson, which emphasizes that systems improve with increased compute rather than complex rules.
Euclid Opens Data Treasure Trove, Offers Glimpse of Deep Fields
Euclid's first data release on March 19, 2025, showcases 26 million galaxies and includes a detailed catalogue of over 380,000 galaxies, revealing their shapes and structures through advanced AI and citizen science collaboration.
GRPO-Based Reinforcement Learning Improves Math Reasoning in Small LLMs with Limited Resources
Small language models (3B-7B params) can achieve significant reasoning improvements through reinforcement learning, with a combination of PPO and DPO yielding up to 74.2% accuracy on the GSM8K benchmark using a 7B model.
The Prospero Challenge
The Prospero Challenge involves rendering a 1024×1024 image from 7866 math expressions in a plain-text file, with a basic Python implementation taking about 15 seconds and consuming 60+ GB of RAM for intermediate results.
BeeFormer: CF and CBF Hybrid Approach for Recommendation Systems
beeFormer enhances cold-start recommendations by leveraging user interaction patterns, allowing new items to benefit from learned behaviors, thus overcoming limitations of traditional collaborative filtering and content-based methods.
Reviewed Several ACL Papers on Data Resources and Feel that LLMs are Undermining this Field
LLMs are increasingly used to generate benchmark datasets, yet this trend raises concerns about the quality and representativeness of the data, as many researchers opt for convenience over rigorous curation methods.
Aircraft Detection at Planetary Scale
Planet's Aircraft Detection Analytic Feed utilizes machine learning to automate the detection of aircraft globally, identifying those ≥25 meters in length or wingspan with unprecedented frequency and accuracy.
Instella: New Open 3B Language Models
Instella is a new family of 3-billion-parameter language models developed by AMD, trained from scratch on 128 Instinct MI300X GPUs, showcasing superior performance over existing fully open models and competitive results against state-of-the-art open-weight models like Llama-3.2-3B.
Aiter: AI Tensor Engine for ROCm
AITER (AI Tensor Engine for ROCm) is a centralized repository of high-performance AI operators that significantly enhances the efficiency of AI workloads on AMD GPUs, allowing seamless integration into various frameworks.
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
FactSelfCheck introduces a black-box sampling-based method for fine-grained detection of hallucinated content in LLMs, utilizing knowledge graphs to represent facts as triples, thus enhancing accuracy in identifying inaccuracies.