Outcome-driven fine-tuning enhances large language models' (LLMs) forecasting abilities by utilizing self-play to generate diverse reasoning trajectories, improving prediction accuracy by 7-10% over base models.
High-Fidelity Datasets for League of Legends
High-fidelity datasets for League of Legends can be created by reverse engineering the game engine, allowing for precise capture of gameplay events like ability usage and player movements, which are typically inaccessible through public APIs.
o3 Achieves a Gold Medal at the 2024 IOI
o3 achieves a gold medal at the 2024 IOI without relying on hand-crafted strategies, showcasing the power of reinforcement learning in large language models (LLMs) for complex tasks.
Automated Capability Discovery via Foundation Model Self-Exploration
Automated Capability Discovery (ACD) leverages a foundation model as a "scientist" to autonomously propose tasks that probe the abilities of other models, revealing thousands of capabilities and failures that would be difficult to uncover manually.
Text-to-SQL in Enterprises
Fine-tuning open-weight LLMs on business-specific query-SQL pairs achieved 95% accuracy, significantly outperforming previous methods that plateaued at 85% accuracy.
New Paper on Frontier Models Self-Exploration
Automated Capability Discovery (ACD) enables foundation models to autonomously explore and identify their own capabilities, revealing thousands of abilities that human evaluators might miss.
SWE-agent: Open-Source SOTA on SWE-bench Lite
SWE-agent is an open-source software engineering agent that supports massively parallel runs and cloud-based deployment, enhancing its usability across various models, including local LMs like Qwen and Llama.
Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
TAID introduces a novel knowledge distillation method that utilizes a time-dependent intermediate distribution to effectively bridge the gap between student and teacher models, addressing challenges like capacity differences and mode collapse.
Novel Clustering Metric - The Jaccard-Concentration Index
The Jaccard-Concentration Index (JCI) is a novel Python library that evaluates clustering quality by combining the Jaccard index with a custom concentration score, offering a more nuanced view of cluster purity. Learn more about JCI.
AlignRec Outperforms SOTA Models in Multimodal Recommendations
AlignRec significantly enhances multimodal recommendation systems by optimizing three alignment tasks: inter-content (ICA), content-category (CCA), and user-item (UIA), effectively bridging semantic gaps between diverse content types.
LLMs as Few-Shot Data Annotators for Multilingual Text Detoxification
LLMs can generate high-quality parallel datasets for text detoxification through few-shot learning, effectively creating toxic/non-toxic text pairs that retain semantic meaning while significantly reducing toxicity.
Creating a Causal DAG for Irregular Time-Series Data
Dynamic Bayesian networks can effectively model causal relationships in irregular time-series data, offering a structured approach to causal inference.
How Scaling Laws Drive Smarter, More Powerful AI
Scaling laws in AI reveal that performance improves with increased training data, model parameters, and computational resources, leading to the emergence of three distinct laws: pretraining, post-training, and test-time scaling.
Stable Diffusion 3.5 Large Available on Microsoft Azure AI Foundry
Stable Diffusion 3.5 Large (SD3.5 Large) is now accessible on Microsoft Azure AI Foundry, enabling businesses to leverage professional-grade image generation seamlessly within their existing workflows.
Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding
Top-Theta Attention introduces a method that prunes less essential attention elements through calibrated thresholds, enhancing efficiency in transformer models without sacrificing accuracy.