NeuralSVG: An Implicit Representation for Text-to-Vector Generation
NeuralSVG introduces an innovative approach to text-to-vector graphics generation, leveraging a small MLP network to encode entire scenes, enhancing the layered structure crucial for vector graphics.
Minimum bipartite matching via Riemann optimization
Riemann optimization transforms the minimum weight bipartite matching problem into a differentiable framework, leveraging the Birkhoff-von Neumann theorem to navigate the convex space of doubly stochastic matrices for efficient solutions.
[R][D] White Box Transformers
The White Box Transformers approach redefines data representation learning as a quest for a dictionary of multivariate gaussians, emphasizing sparse coding for effective feature extraction.
NVIDIA DRIVE Partners Showcase Latest Mobility Innovations at CES
NVIDIA DRIVE AGX Thor, built on the NVIDIA Blackwell architecture, is designed to manage demanding data workloads in autonomous vehicles, enhancing capabilities in generative AI and large language models.
[D] Positional Embeddings in Embedding Space
Positional Encodings are crucial for understanding the distribution of embeddings in feature space, influencing how models interpret sequential data.
[R][P] distillKitPlus: High Performance Knowledge Distillation for LLMs
distillKitPlus is an open-source toolkit designed for knowledge distillation in large language models (LLMs), enabling efficient transfer of capabilities from a 70B model to a 7B model without significant cost or performance loss.
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
The Video-of-Thought (VoT) framework introduces a novel Multimodal Large Language Model (MLLM), MotionEpic, which enhances video comprehension by integrating spatial-temporal scene graph (STSG) representation for pixel-level grounding.
[D] Program Of Thought Prompting (PoT) vs Chain Of Thought Prompting (CoT)
Program of Thought Prompting (PoT) enhances the Chain of Thought (CoT) methodology by decoupling reasoning from computation, allowing for more accurate problem-solving through executable Python code.
NVIDIA Makes Cosmos World Foundation Models Openly Available to Physical AI Developer Community
NVIDIA's Cosmos World Foundation Models are now openly available, enabling developers to create physics-aware simulations for robotics and autonomous vehicles, leveraging 9,000 trillion tokens from extensive real-world data.
[R] LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
LongBench v2 is a comprehensive benchmark featuring 503 challenging questions across six task categories, designed to evaluate LLMs' capabilities in handling long-context problems that require deep understanding and reasoning, with contexts ranging from 8k to 2M words.
SWE Bench just got updated – new #1s
SWE-bench evaluates language models' ability to resolve GitHub issues, utilizing a dataset of 2,294 issue-pull request pairs from popular Python repositories, with performance measured through unit test verification.
Why World Foundation Models Will Be Key to Advancing Physical AI
World foundation models (WFM) are pivotal for advancing physical AI, enabling systems like robots and self-driving cars to simulate and predict real-world interactions effectively.
CES 2025: AI Advancing at ‘Incredible Pace,’ NVIDIA CEO Says
NVIDIA's CES 2025 keynote highlighted the launch of NVIDIA Cosmos and the Blackwell RTX 50 Series GPUs, marking a significant leap into physical AI, which enables machines to reason, plan, and act autonomously.
NVIDIA Unveils ‘Mega’ Omniverse Blueprint for Building Industrial Robot Fleet Digital Twins
NVIDIA's Mega Omniverse Blueprint revolutionizes industrial AI by enabling the development and optimization of digital twins for robot fleets, enhancing operational efficiency before real-world deployment.
Building Smarter Autonomous Machines: NVIDIA Announces Early Access for Omniverse Sensor RTX
NVIDIA's Omniverse Sensor RTX APIs enable high-fidelity sensor simulation, allowing developers to generate diverse datasets essential for training autonomous machines, addressing challenges in data collection for edge cases and hazardous scenarios.