# Aug 21, 2024

## Data Exfiltration from Slack AI via indirect prompt injection
- **Attackers can exfiltrate data from private Slack channels** they are not a member of by using indirect prompt injection, exploiting the Slack AI's inability to distinguish between legitimate user queries and malicious instructions.

## NVIDIA Announces First Digital Human Technologies On-Device Small Language Model, Improving Conversation for Game Characters
- NVIDIA's **Nemotron-4 4B Instruct**, a **small language model**, is showcased in _Mecha BREAK_, enhancing **game character conversations** for a more immersive experience. [NVIDIA ACE](https://developer.nvidia.com/ace)

## Self-Supervised Learning for Videos
- **Self-supervised learning for videos** leverages the **VideoMAE architecture**, adapting masked autoencoders from images to videos, addressing the unique spatio-temporal complexity of video data. [VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training, 2022](https://arxiv.org/abs/2203.12602)

## Launch HN: Outerport (YC S24) – Instant hot-swapping for AI models
- **Outerport introduces 'hot-swapping' for AI models**, allowing different models to be served on the same GPU with swap times of approximately **2 seconds**, a process **150x faster** than the current baseline, demonstrated through a [live demo](https://hotswap.outerport.com/).

## Join Our Global Paper Reading Group for a Deep Dive into "Plan Like a Graph (PLaG)" - Enhancing LLMs in Asynchronous Plan Reasoning | ICML 2024 with the author Fangru Lin!
- The **"Plan Like a Graph (PLaG)" method** significantly **enhances LLM performance** by decomposing tasks into sub-tasks that form an execution graph, enabling parallel and sequential task execution without fine-tuning.

## HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models in Resource-Constrained Environments
- **HiRED**, a **token-dropping scheme**, enhances the efficiency of **high-resolution Vision-Language Models (VLMs)** by selectively processing visual tokens within a **fixed token budget**, preserving accuracy in **resource-constrained environments**.

## Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
- **Transfusion** introduces a **novel training recipe** that merges language modeling and diffusion techniques, enabling a single transformer to process both text and image data efficiently.

## Beta9: Open Serverless GPU Cloud
- **Beta9**, now open-sourced, **enables developers to deploy serverless functions on cloud GPUs** efficiently, building on the success of its predecessor, Beam, to offer a Python-first experience for rapid ML model deployment. [GitHub Repo](https://github.com/beam-cloud/beta9)

## Predicting Phenotypes from Omics Data
- **Deep learning** is being explored to **predict phenotype variation** from **omics data**, indicating a growing interest in leveraging computational methods for biological insights.

## PEDAL: Enhancing Greedy Decoding with Large Language Models using Diverse Exemplars
- **PEDAL** introduces a method to **enhance greedy decoding** in large language models by incorporating **diverse exemplars**, aiming to improve output quality and relevance.

## Using synthetic data from LLMs to train/finetune other models
- **Generating synthetic data with LLMs** aims to **enhance the fine-tuning of sentence transformers**, exploring the potential of artificial intelligence in creating valuable datasets for machine learning models.

## NVIDIA Showcases New AI Capabilities With ACE, RTX Games and More at Gamescom 2024
- **NVIDIA announced its first on-device small language model (SLM), NVIDIA Nemotron-4 4B Instruct**, aimed at enhancing conversational AI in games, showcased through _Mecha BREAK_, a mech combat game demonstrating AI-powered characters' potential.

## SLMming Down Latency: How NVIDIA’s First On-Device Small Language Model Makes Digital Humans More Lifelike
- **NVIDIA's first on-device small language model (SLM), Nemotron-4 4B Instruct, enhances digital human interaction** in games by understanding and responding to player instructions more intuitively and accurately.

## Lightweight Champ: NVIDIA Releases Small Language Model With State-of-the-Art Accuracy
- **NVIDIA's Mistral-NeMo-Minitron 8B** is a compact, high-accuracy language model, leveraging **pruning and distillation** techniques to reduce its size from 12 billion to 8 billion parameters without sacrificing performance.

## Improving Hugging Face Training Efficiency Through Packing with Flash Attention
- **Hugging Face introduces packing with Flash Attention 2**, significantly **improving training efficiency** by allowing training with packed instruction tuning examples, eliminating the need for padding, as detailed in a [recent PR](https://github.com/huggingface/transformers/pull/31629) and facilitated by the new [`DataCollatorWithFlattening`](https://huggingface.co/docs/transformers/main/en/main_classes/data_collator#transformers.DataCollatorWithFlattening).
