Pico-Banana-400K is a comprehensive dataset featuring ~400K text–image–edit triplets, facilitating advancements in text-guided image editing through diverse transformations across 35 edit operations and 8 semantic categories.
Definition of AGI
This paper defines Artificial General Intelligence (AGI) as the ability to match the cognitive versatility of a well-educated adult, utilizing a framework based on the Cattell-Horn-Carroll theory to assess AI systems across ten cognitive domains.
Nvidia DGX Spark: When Benchmark Numbers Meet Production Reality
DGX Lab benchmarks reveal significant discrepancies between expected and actual performance metrics, highlighting the need for more accurate testing environments.
TOON – Token Oriented Object Notation
Token-Oriented Object Notation (TOON) is a compact format that reduces token usage by 30-60% compared to JSON, making it ideal for structured data input to Large Language Models (LLMs) while maintaining readability through an indentation-based structure.
DeepAgent: A General Reasoning Agent with Scalable Toolsets
DeepAgent is an end-to-end deep reasoning agent that autonomously discovers tools and executes actions, addressing the limitations of traditional agent frameworks that rely on predefined workflows.
We're in the Wrong Moment
Generative AI has transformed the coding landscape, diminishing the joy of programming for many, as it automates tasks that once required creativity and problem-solving skills.
Cutting Inference Costs from $46K to $7.5K by Fine-Tuning Qwen-Image-Edit
Fine-tuning Qwen-Image-Edit reduced inference costs from $46K to $7.5K, enabling the generation of a 1.2 million image catalog with significant efficiency improvements.
PKBoost: Gradient Boosting that Stays Accurate Under Data Drift
PKBoost demonstrates superior performance in handling data drift, achieving only a 2% degradation in PR-AUC compared to XGBoost's 32%, making it a robust choice for imbalanced datasets like credit card fraud detection.
The New Calculus of AI-Based Coding
AI agents like Amazon Q and Kiro are revolutionizing coding practices, enabling teams to achieve 10x throughput by collaborating with human engineers who ensure code quality through rigorous review processes.
Sparser Block-Sparse Attention via Token Permutation
PBS-Attn enhances block-sparse attention by utilizing token permutation, significantly improving computational efficiency in large language models (LLMs) during long-context prefilling.
Streaming Datasets: 100x More Efficient
Streaming datasets have achieved a remarkable 100x efficiency improvement, allowing users to train on multi-TB datasets without the hassle of downloading, thus eliminating common issues like "disk out of space" errors and significantly speeding up data resolution times.
A Geometric Interpretation of the Weight Update in GPTQ Quantization Algorithm and a Novel Solution
GPTQ quantization modifies weights row-wise, utilizing a Lagrangian approach to derive updates, which enhances efficiency in matrix quantization.
huggingface_hub v1.0: Five Years of Building the Foundation of Open Machine Learning
huggingface_hub v1.0 marks a significant milestone, supporting 200,000 dependent libraries and providing access to over 2 million models, 500,000 datasets, and 1 million Spaces, reflecting its evolution into a robust platform for open machine learning.
Co-Sight: Enhancing LLM-Based Agents
Co-Sight enhances LLM-based agents by implementing Conflict-Aware Meta-Verification (CAMV) and Trustworthy Reasoning with Structured Facts (TRSF), transforming reasoning into a falsifiable and auditable process that improves efficiency and reliability.
Boosting Accuracy and Efficiency of Budget Forcing in LLMs via Reinforcement Learning for Mathematical Reasoning
Reinforcement learning (RL) enhances budget forcing in Large Language Models (LLMs), improving mathematical reasoning accuracy while reducing token usage by over 40% compared to supervised fine-tuning (SFT) alone.