# Nov 8, 2024

## Daily

### Roaring Bitmap Compression

- The **new Roaring hybrid** technique enhances bitmap compression by integrating uncompressed bitmaps, packed arrays, and RLE compressed segments, resulting in superior performance on both speed and size.

### AI for real-time fusion plasma behavior prediction and manipulation

- **AI-driven multimodal super-resolution** enhances the understanding of fusion plasma behavior by revealing hidden inter-correlations between diagnostics, crucial for stabilizing Edge Localized Modes (ELMs) that threaten reactor integrity.

### Perceptually lossless (talking head) video compression at 22kbit/s

- **LivePortrait achieves perceptually lossless video compression at an impressive 22kbit/s**, leveraging advanced facial keypoint transformations to minimize data transmission while maintaining quality, outperforming traditional codecs like H.264.

### LoRA vs. Full Fine-Tuning: An Illusion of Equivalence

- **LoRA and full fine-tuning yield distinct weight matrix structures**, revealing that LoRA introduces new, high-ranking singular vectors termed **_intruder dimensions_**, which do not emerge in full fine-tuning.

### [R] State-space models can learn in-context by gradient descent

- **Deep state-space models (Deep SSMs)** can achieve in-context learning through **gradient descent**, demonstrating that a single structured layer with local self-attention can replicate outputs of an implicit linear model after just one gradient descent step.

### [D] Discovery: Anthropic somehow injecting/hiding safety warnings in user prompts, telling Claude to keep it secret. [Content Warning: Violence]

- **Claude** is programmed to inject **dynamic safety warnings** into user prompts, indicating a sophisticated mechanism to manage sensitive content, which may involve **surgical tuning** techniques as outlined in Anthropic's research on model interpretation.

### [N] Super fast and SOTA Visual Tokenizers

- **Tokenizers are essential** for advancing image and video generative models, with this work introducing **multiple causal tokenizers** that support both continuous and discrete spaces, enhancing their utility in diffusion and autoregressive contexts.

### [R] Benchmarking Large Language Models with Integer Sequence Generation Tasks

- This benchmark evaluates **large language models (LLMs)** on their ability to generate code for integer sequences from the **Online Encyclopedia of Integer Sequences (OEIS)**, revealing that the **o1 series** models excel in both accuracy and cheating detection compared to competitors like OpenAI and Google.

### [D] Directions on drug-target interaction prediction

- **Drug-target interaction (DTI) prediction** typically involves generating **target embeddings** with **PLMs** like **ESM2** and **drug embeddings** using **CLMs** such as **ChemBERTa**, followed by employing **cross-modal attention mechanisms** for integration.

### [R]: How much is a noisy image worth? 👀

### SuffixDecoding: A Model-Free Approach to Speeding Up Large Language Model Inference

- **SuffixDecoding** is a **model-free** method that accelerates **large language model (LLM)** inference by utilizing **suffix trees** from prior outputs, enabling efficient token sequence predictions without the need for additional models.
