# Aug 8, 2024

**FlexAttention: The Flexibility of PyTorch with the Performance of FlashAttention**  
**FlexAttention** introduces a **flexible PyTorch API** that enables the implementation of various attention variants with performance on par with optimized kernels like FlashAttention, addressing the trade-off between flexibility and efficiency in attention mechanisms.

**Qwen2-Math**  
**Qwen2-Math introduces a series of math-specific large language models**, including Qwen2-Math-Instruct-1.5B/7B/72B, designed to significantly outperform both open-source and proprietary models like GPT-4o in solving complex mathematical problems.

**[Research] The Puzzling Failure of Multimodal AI Chatbots**  
**Multimodal AI chatbots** like **GPT-4o and Gemini**, despite their advanced capabilities in processing images and texts, **fail to match human-level general intelligence and reasoning**, as highlighted by the new **[PuzzleVQA benchmark](https://arxiv.org/abs/2403.13315)**.

**[D] FlexAttention: Flexibility of PyTorch with Performance of FlashAttention**  
**FlexAttention** combines **PyTorch's flexibility** with the **high performance** of FlashAttention, offering a **new tool** for deep learning practitioners.

**🤗XetHub is joining Hugging Face!**  
**Hugging Face acquires XetHub**, a company specializing in scaling Git for AI development, to enhance the storage backend of HF datasets and models.

**GPUDrive: Data-driven, multi-agent driving simulation at 1M FPS**  
**GPUDrive** accelerates multi-agent learning by generating **over a million steps of experience per second**, leveraging the **Madrona Game Engine** for high-scale simulation.

**[P] GroundedAI: Open-Source Framework/Models for Efficient LLM Evaluation**  
**GroundedAI** is an **open-source framework** designed for **evaluating large language model (LLM) outputs**, focusing on **toxicity, RAG relevance, and hallucination** using **fine-tuned small language models** and **specialized adapters**.

**WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models**  
**WalledEval** is a **comprehensive AI safety testing toolkit** designed for evaluating large language models (LLMs), featuring over **35 safety benchmarks** including multilingual and exaggerated safety, and prompt injections.

**Recursion CEO Chris Gibson on Accelerating the Biopharmaceutical Industry With AI**  
**Recursion** utilizes **AI and machine learning** to significantly **enhance drug discovery and development**, aiming to **increase efficiency** and **reduce costs** in the biopharmaceutical industry.

**Figure Unveils Next-Gen Conversational Humanoid Robot With 3x AI Computing for Fully Autonomous Tasks**  
**Figure's next-gen Figure 02 humanoid robot** utilizes **NVIDIA Omniverse and GPUs** for enhanced autonomy, achieving **3x AI computing power** for real-world tasks.

**Large-scale pathology foundation models show promise on a variety of cancer-related tasks**  
**Microsoft Research and Paige** have developed **Virchow2 and Virchow2G**, foundation models for computational pathology, demonstrating **unprecedented accuracy in detecting various cancers** by analyzing over **3.1 million whole slide images** from **225,000 patients** across **45 countries**.

**[D] OpenAI: Structured Outputs in the API**  
**OpenAI introduces structured outputs in their API**, enhancing the way developers interact with AI models by allowing for more complex and formatted responses.

**Generative Language Models with Retrieval Augmented Generation for Automated Short Answer Scoring**  
This study introduces a **novel pipeline** that integrates **vector databases, transformer-based encoders, and Generative Language Models (GLMs)** to **enhance Automated Short Answer Scoring (ASAS)** accuracy, marking a departure from traditional rule-based or complex deep learning approaches.

**Large Language Models for Base Station Siting: Intelligent Deployment based on Prompt or Agent**  
**Large Language Models (LLMs)** are revolutionizing **base station siting (BSS)** by leveraging **prompt and agent engineering** to infuse human expertise into AI, aiming for more efficient and cost-effective network deployments.

**MPC-Minimized Secure LLM Inference**  
**Marill**, a new framework, **adapts LLM fine-tuning** to **minimize MPC usage** during secure inference, addressing privacy concerns without compromising security.
