# Oct 7, 2024

## Nobel Prize in Physiology or Medicine 2024

- The **2024 Nobel Prize in Physiology or Medicine** was awarded to **Victor Ambros** and **Gary Ruvkun** for their groundbreaking work on **microRNA** and its critical function in **post-transcriptional gene regulation**.

## Longwriter – Increase llama3.1 output to 10k words

- **LongWriter** is a groundbreaking model capable of generating **over 10,000 words in just one minute**, utilizing advanced long-context LLMs, and is now deployable via [vllm](https://github.com/vllm-project/vllm) for enhanced performance.

## ByteDance’s Bytespider is scraping at much higher rates than other platforms

- **ByteDance's Bytespider** is a web scraper that operates **25 times faster than OpenAI's GPTbot**, rapidly accumulating data to enhance its generative AI models, reflecting a strategic push to catch up in the AI race.

## The magic (image resampling) kernel

- **Magic Kernel Sharp** is a superior image resizing algorithm, utilized by major platforms like **Facebook** and **Instagram**, enhancing image quality while improving CPU and storage efficiency since its inception in 2013.

## Sorbet: A neuromorphic hardware-compatible transformer-based spiking model

- **Sorbet** introduces a **neuromorphic hardware-compatible** transformer-based spiking language model that utilizes **PTsoftmax** and **BSPN** to replace energy-intensive operations, enhancing efficiency for edge deployment.

## [P] Model2Vec: Distill a Small Fast Model from any Sentence Transformer

- **Model2Vec** distills Sentence Transformer models into **30mb static embeddings** that are **up to 500x faster** than their original counterparts, enabling efficient CPU usage without requiring extensive hardware.

## Llamafile for Meltemi: The First LLM for Greek

- **Meltemi 7B Instruct v1.5** is the **first Large Language Model (LLM) for Greek**, developed by the **Athena Research & Innovation Center**, and is available in both `llamafile` and `gguf` formats on [HuggingFace](https://huggingface.co/Florents-Tselai/Meltemi-llamafile).

## [R] MaskBit: Embedding-free Image Generation via Bit Tokens

- **MaskBit introduces an innovative, embedding-free image generation model** that operates directly on **bit tokens**, achieving a state-of-the-art FID of **1.52** on the ImageNet 256x256 benchmark with a compact generator of **305M parameters**.

## [Project] Optimizing Neural Networks with Language Models

- **Dux** is a **meta-optimizer** leveraging **GPT-4o-mini** for the **adaptive optimization** of neural networks, aiming to enhance performance through innovative techniques.

## Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models

- **Object hallucinations** in CLIP models are a significant concern, revealing that these issues arise independently of the interaction between vision and language modalities, indicating a deeper flaw within the model itself.

## MELODI: Exploring Memory Compression for Long Contexts

- **MELODI** introduces a **hierarchical compression scheme** that optimally balances short-term and long-term memory, enabling efficient processing of long documents with limited context windows.

## LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy

- **LoRC** introduces a **low-rank approximation** for KV weight matrices, enabling **significant memory reduction** in transformer-based LLMs without the need for model retraining or extensive parameter tuning.
