# Aug 15, 2024

## Artists Score Major Win in Copyright Case Against AI Art Generators

A federal judge has allowed **key copyright and trademark claims** to proceed in a lawsuit by artists against AI art generators, marking a significant development in the legal battle over the use of copyrighted images to train AI.

## We have discovered antibiotics in the global microbiome with AI, ask us anything

Researchers have **discovered nearly 1 million new antibiotic molecules** in the global microbiome, leveraging **machine learning** to mine through **63,410 metagenomes and 87,920 microbial genomes**. [Read the study](https://doi.org/10.1016/j.cell.2024.05.013)

## YouTube Video to Tabs and Lyrics

**Fish** is an AI-powered multimodal project that generates **chords, beats, lyrics, melody, and tabs** for any song, leveraging a transformer-based hybrid model to address music information retrieval challenges.

## Gemlite: Towards Building Custom Low-Bit Fused CUDA Kernels

**[Gemlite](https://github.com/mobiusml/gemlite/)** is introduced as a toolkit for developers to easily create custom low-bit fused General Matrix-Vector Multiplication (GEMV) CUDA kernels, aiming to simplify the process of implementing quantization-aware runtime for AI models on GPUs.

## Launch HN: Hamming (YC S24) – Automated Testing for Voice Agents

**Hamming introduces an automated testing service for LLM voice agents**, aiming to streamline the iterative process of improving voice agent performance by simulating challenging user interactions.

## 
[R] I've devised a potential transformer-like architecture with O(n) time complexity, reducible to O(log n) when parallelized.

The **Equinox architecture** employs a **divide and compute method** to achieve a **time complexity of O(n)**, which can be further reduced to **O(log n)** through parallelization, suggesting a significant efficiency improvement over traditional transformer models.

## 
[R] Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

**Gemma Scope introduces an open suite of JumpReLU Sparse Autoencoders (SAEs)** trained across various layers of Gemma 2 models, aiming to democratize access to advanced neural network interpretability tools.

## Hermes 3: The First Fine-Tuned Llama 3.1 405B Model

**Hermes 3**, developed by Nous Research and hosted on Lambda's Cloud, represents the **first full-parameter fine-tune** of Meta's Llama 3.1 405B model, boasting **exceptional reasoning capabilities** and designed for the open-source community.

## 
[R] New Paper on Mixture of Experts (MoE) 🚀

The new paper on **Mixture of Experts (MoE)** delves into **advancements** in balancing **computational efficiency** with **high performance** in AI systems, addressing **current challenges** and **future directions**.

## 
[P] New open-source release: SOTA multimodal embedding models for fashion

**Marqo-FashionCLIP & Marqo-FashionSigLIP**, new **open-source multimodal models**, have **outperformed existing SOTA models** like FashionCLIP2.0 and OpenFashionCLIP on **7 fashion evaluation datasets** including DeepFashion and Fashion200K by **up to 57%**.

## Detecting Sign Language in News Videos

The **Hand-Signer detection model** developed by Vrroom's team can identify sign language communicators in crowded news videos, marking a significant step towards creating a **large-scale dataset** for bi-directional translation between English and Indian Sign Language (ISL). [Longtail AI Foundation](https://longtailai.org/)

## 
[R] Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

**Scaling test-time computation** in Large Language Models (LLMs) can significantly **enhance performance** on complex prompts, suggesting a pivotal shift towards more **compute-efficient strategies** for improving LLM outputs.

## ALS Stole His Voice. A.I. Retrieved It

**Casey Harrell**, suffering from **A.L.S.**, regained the ability to **communicate** through a **brain-computer interface** that translates his intended speech into an **A.I.-powered voice**, closely mimicking his own.

## 
[R] AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation

**AgentGen** enhances **LLM-based agents' planning abilities** by automating the synthesis of diverse environments and planning tasks, moving beyond the limitations of manually designed tasks.

## 
[R] LayerMerge: Neural Network Depth Compression through Layer Pruning and Merging (ICML 2024)

**LayerMerge** innovatively **reduces CNN and diffusion model depth** by **pruning and merging** both convolution and activation layers, maintaining performance while enhancing efficiency. [Paper](https://arxiv.org/abs/2406.12837)
