# May 29, 2024

## Daily Updates

- **Reproducing GPT-2 (124M) using llm.c** on a single 8X A100 80GB SXM node takes approximately 90 minutes and costs around $20, showcasing an efficient use of model flops utilization at up to ~60%.

- **Codestral**, Mistral AI's inaugural code model, excels in **code generation** across **80+ programming languages**, promising to enhance software development with its advanced AI capabilities.

- The **world's first bioprocessor**, developed by Swiss startup FinalSpark, **utilizes 16 human brain organoids** to achieve **‘a million times less power’ consumption** than traditional digital chips.

- **Llama 3-V** is a **multimodal model** that matches GPT4-V's performance with a **100x smaller model size**, achieving a 10-20% performance boost over Llava, the current state-of-the-art (SOTA) in multimodal understanding, by leveraging a novel architecture and training approach.

- **Over the past year, building with large language models (LLMs) has transitioned from experimental to practical**, with real-world applications benefiting from the technology's advancements and broad accessibility.

- **tinygrad 0.9.0** introduces significant usability improvements, including **over 1200 commits** since version 0.8.0, and approaches the new line limit with **7958 lines** of code.

- **AdFlush**, a machine learning model, was developed for real-world browsers to effectively prevent advertisements and web trackers, selecting **27 key features from an evaluation of 883** for optimal performance.

- **Cohere's release of the Wikipedia dataset**, embedded into vectors using their multilingual-v3 model, **makes it feasible to index Wikipedia on a personal laptop**, sidestepping the previously prohibitive $5000 cost of computing such embeddings. [Cohere's dataset](https://huggingface.co/datasets/Cohere/wikipedia-2023-11-embed-multilingual-v3)

- **ChatTTS** is a text-to-speech model tailored for dialogue scenarios, supporting English and Chinese with a training dataset of over 100,000 hours, while the open-source version is based on 40,000 hours of pre-training without Speaker Feature Transfer (SFT).

- **Hallucination in LLMs** is perceived as less prioritized than **safety**, despite its critical role in ensuring the **accuracy of generated information**.

- **Elixir's machine learning landscape** has evolved with **Nx v0.7** introducing **MLIR support**, enhancing capabilities like **Apple Silicon Metal support** and **cross-compilation** for embedded devices. [MLIR](https://mlir.llvm.org/)

- **Optimised Attention** reduces parameters by **25%** and matrix multiplications per head, maintaining performance akin to standard attention.

- **Era3D** introduces a **novel multiview diffusion method** that overcomes camera prior mismatch and inefficacy, generating high-resolution images from a single image without shape distortions.

- **REX Computing** is innovating with a **new processor architecture**, the Neo, aiming for a **10 to 25x increase in energy efficiency** over current CPUs and GPUs by simplifying design and focusing on software improvements.

- The **You Only Cache Once (YOCO)** architecture introduces a **novel split** in the network into **self-decoder layers** for generating a **global KV-Cache** and **cross-decoder layers** that reuse this cache, aiming to **enhance efficiency** in Large Language Models (LLMs) as illustrated in [Figure 1](https://preview.redd.it/n6iitz36873d1.png?width=804&format=png&auto=webp&s=597102302acd26f27e28e99b366c13b7b135457a).
