ML Times
Jul 1, 2025
Daily
The New Skill in AI Is Not Prompting, It's Context Engineering
- Context Engineering is the emerging skill in AI, emphasizing the importance of providing comprehensive context to enhance the performance of language models, as highlighted by Tobi Lutke's definition of it as "the art of providing all the context for the task to be plausibly solvable by the LLM.”
Small Language Models Are the Future of Agentic AI
- Small language models (SLMs) are poised to dominate agentic AI due to their efficiency and specialized capabilities, making them more suitable than large language models (LLMs) for repetitive tasks in various applications.
BIG-Bench Extra Hard
- BIG-Bench Extra Hard (BBEH) introduces a new benchmark to evaluate general reasoning capabilities in large language models (LLMs), replacing tasks in the previous BIG-Bench Hard (BBH) with significantly more challenging ones to address saturation in model performance.
The Bitter Lesson is coming for Tokenization
- Tokenization's fragility is under scrutiny as researchers advocate for a general method that optimally utilizes compute and data, potentially revolutionizing how we process information in machine learning.
Researchers Uncover Hidden Ingredients Behind AI Creativity
- Creativity in AI arises from the inherent imperfections in the denoising process of diffusion models, suggesting that their ability to generate novel images is a deterministic outcome of their architecture rather than mere randomness.
Entropy of a Mixture
- The entropy of a mixture ( H(p_\lambda) ) is concave in relation to the interpolation factor ( \lambda ), indicating that as the similarity between distributions ( p_0 ) and ( p_1 ) decreases, the curve bulges upwards, revealing deeper insights into their relationship through metrics like JSD and KL divergence.
Show HN: Arch-Router – 1.5B model for LLM routing by preferences, not benchmarks
- Arch-Router is a 1.5B parameter model that enables preference-based routing for LLMs, allowing users to define routing rules in plain language without the need for retraining classifiers.
Simulations reveal the secret to strengthening carbon fiber
- ORNL researchers have doubled the tensile strength of carbon-fiber composites by incorporating a thin layer of PAN nanofibers, enhancing load distribution and overall strength at the atomic level through advanced simulations.
Training and Finetuning Sparse Embedding Models with Sentence Transformers v5
- Sparse embedding models in Sentence Transformers v5 enable efficient training and finetuning, allowing for high-dimensional vector representations that enhance semantic search and retrieval tasks, particularly in hybrid scenarios.
Inference-Time Scaling and Collective Intelligence for Frontier AI
- AB-MCTS enables multiple frontier models to collaborate at inference time, significantly enhancing performance on the ARC-AGI-2 benchmark compared to individual models.
Building a Personal AI Factory
- Personal AI factories leverage multiple AI agents to autonomously generate, verify, and improve code, enhancing efficiency and reducing manual oversight.
Interpreting Large Language Models' Personality through Critical Event Analysis
- The Supernova Event Dataset introduces a benchmark for analyzing the "personality" of large language models (LLMs) through critical event analysis, utilizing real-world Wikipedia articles to assess behavioral patterns.
AI Testing and Evaluation: Learnings from genome editing
- Generative AI's governance can benefit from insights gained in genome editing, emphasizing the need for tailored testing and evaluation frameworks to ensure responsible technology deployment.