Data Exfiltration from Slack AI via indirect prompt injection
Attackers can exfiltrate data from private Slack channels they are not a member of by using indirect prompt injection, exploiting the Slack AI's inability to distinguish between legitimate user queries and malicious instructions.
NVIDIA Announces First Digital Human Technologies On-Device Small Language Model, Improving Conversation for Game Characters
NVIDIA's Nemotron-4 4B Instruct, a small language model, is showcased in Mecha BREAK, enhancing game character conversations for a more immersive experience. NVIDIA ACE
Launch HN: Outerport (YC S24) – Instant hot-swapping for AI models
Outerport introduces 'hot-swapping' for AI models, allowing different models to be served on the same GPU with swap times of approximately 2 seconds, a process 150x faster than the current baseline, demonstrated through a live demo.
Join Our Global Paper Reading Group for a Deep Dive into "Plan Like a Graph (PLaG)" - Enhancing LLMs in Asynchronous Plan Reasoning | ICML 2024 with the author Fangru Lin!
The "Plan Like a Graph (PLaG)" method significantly enhances LLM performance by decomposing tasks into sub-tasks that form an execution graph, enabling parallel and sequential task execution without fine-tuning.
HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models in Resource-Constrained Environments
HiRED, a token-dropping scheme, enhances the efficiency of high-resolution Vision-Language Models (VLMs) by selectively processing visual tokens within a fixed token budget, preserving accuracy in resource-constrained environments.
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Transfusion introduces a novel training recipe that merges language modeling and diffusion techniques, enabling a single transformer to process both text and image data efficiently.
Beta9: Open Serverless GPU Cloud
Beta9, now open-sourced, enables developers to deploy serverless functions on cloud GPUs efficiently, building on the success of its predecessor, Beam, to offer a Python-first experience for rapid ML model deployment. GitHub Repo
Predicting Phenotypes from Omics Data
Deep learning is being explored to predict phenotype variation from omics data, indicating a growing interest in leveraging computational methods for biological insights.
PEDAL: Enhancing Greedy Decoding with Large Language Models using Diverse Exemplars
PEDAL introduces a method to enhance greedy decoding in large language models by incorporating diverse exemplars, aiming to improve output quality and relevance.
Using synthetic data from LLMs to train/finetune other models
Generating synthetic data with LLMs aims to enhance the fine-tuning of sentence transformers, exploring the potential of artificial intelligence in creating valuable datasets for machine learning models.
NVIDIA Showcases New AI Capabilities With ACE, RTX Games and More at Gamescom 2024
NVIDIA announced its first on-device small language model (SLM), NVIDIA Nemotron-4 4B Instruct, aimed at enhancing conversational AI in games, showcased through Mecha BREAK, a mech combat game demonstrating AI-powered characters' potential.
SLMming Down Latency: How NVIDIA’s First On-Device Small Language Model Makes Digital Humans More Lifelike
NVIDIA's first on-device small language model (SLM), Nemotron-4 4B Instruct, enhances digital human interaction in games by understanding and responding to player instructions more intuitively and accurately.
Lightweight Champ: NVIDIA Releases Small Language Model With State-of-the-Art Accuracy
NVIDIA's Mistral-NeMo-Minitron 8B is a compact, high-accuracy language model, leveraging pruning and distillation techniques to reduce its size from 12 billion to 8 billion parameters without sacrificing performance.
Improving Hugging Face Training Efficiency Through Packing with Flash Attention
Hugging Face introduces packing with Flash Attention 2, significantly improving training efficiency by allowing training with packed instruction tuning examples, eliminating the need for padding, as detailed in a recent PR and facilitated by the new DataCollatorWithFlattening.