NeuralSVG: An Implicit Representation for Text-to-Vector Generation
NeuralSVG introduces an innovative approach to text-to-vector graphics generation, leveraging a small MLP network to encode entire scenes, enhancing the layered structure crucial for vector graphics.
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
rStar-Math demonstrates that small language models (SLMs) can achieve or exceed the math reasoning capabilities of larger models like OpenAI's o1 through innovative techniques such as Monte Carlo Tree Search (MCTS) and a self-evolution process.
Nvidia releases its own brand of world models
Nvidia's Cosmos World Foundation Models (Cosmos WFMs) are now available, enabling developers to create physics-aware videos and synthetic data for applications like robotics and autonomous vehicles, with models ranging from 4 billion to 14 billion parameters.
TabPFN v2: Accurate predictions on small data with a tabular foundation model
TabPFN v2 is a pretrained transformer that excels in small tabular data, achieving superior performance in 2.8 seconds for classification and 4.8 seconds for regression tasks, even outperforming strong baselines tuned for hours.
Show HN: TabPFN v2 – A SOTA foundation model for small tabular data
TabPFN is a novel tabular foundation model that significantly outperforms traditional methods, achieving superior predictions on datasets with up to 10,000 samples in just 2.8 seconds, compared to 4 hours for conventional models like gradient-boosted decision trees.
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
The Video-of-Thought (VoT) framework introduces a novel Multimodal Large Language Model (MLLM), MotionEpic, which enhances video comprehension by integrating spatial-temporal scene graph (STSG) representation for pixel-level grounding.
ObliqueTree: Advanced Decision Tree Implementation
ObliqueTree offers a high-performance decision tree implementation that excels in both classification and regression tasks, utilizing oblique splits for enhanced flexibility and generalization with shallow trees.
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
LongBench v2 is a comprehensive benchmark featuring 503 challenging questions across six task categories, designed to evaluate LLMs' capabilities in handling long-context problems that require deep understanding and reasoning, with contexts ranging from 8k to 2M words.
Unveiling a New Era of Local AI With NVIDIA NIM Microservices and AI Blueprints
NVIDIA's new NIM microservices and AI Blueprints enable generative AI on RTX AI PCs, enhancing capabilities in digital human creation, content generation, and productivity applications, all powered by the advanced GeForce RTX 50 Series GPUs.
Hyundai Motor Group Embraces NVIDIA AI and Omniverse for Next-Gen Mobility
Hyundai Motor Group is leveraging NVIDIA AI and Omniverse technologies to enhance vehicle safety, manufacturing efficiency, and robotics, marking a significant step towards next-gen mobility solutions.
CO₂ Emissions and Models Performance: Insights from the Open LLM Leaderboard
CO₂ emissions from model inference have become a critical concern, with recent evaluations revealing that community fine-tunes often outperform official models in carbon efficiency, highlighting a shift towards more sustainable AI practices.
Supervision-free Vision-Language Alignment
SVP (Supervision-free Visual Projection) enhances vision-language alignment by utilizing self-captioning and a pre-trained grounding model, eliminating the need for curated image-text pairs.
CURing Large Models: Compression via CUR Decomposition
CURing leverages CUR matrix decomposition to compress large models, approximating weight matrices with selected columns (C), rows (R), and a linking matrix (U), achieving significant size reduction with minimal performance loss.
Integrating Ascend Backend with Torchtune through PyTorch Multi-Device Support
Torchtune is a PyTorch-native library that simplifies the fine-tuning of Large Language Models (LLMs) by providing modular building blocks and supporting various training methods across different GPUs, including Ascend NPU integration for enhanced performance.