# Feb 1, 2025

## Daily

### Auto-Differentiating Any LLM Workflow: A Farewell to Manual Prompting

- **LLM-AutoDiff** introduces a framework for **Automatic Prompt Engineering (APE)** that enhances multi-component LLM workflows by treating textual inputs as trainable parameters, enabling iterative prompt updates through feedback akin to textual gradients.

### Notes on OpenAI O3-Mini

- **OpenAI's o3-mini model** outperforms GPT-4o and o1 in competitive programming benchmarks, notably achieving a **2130 score on Codeforces ELO**, indicating its potential for specialized tasks.

### \\D Non-deterministic behavior of LLMs when temperature is 0

- **LLMs are expected to be deterministic at temperature 0**, yet practical observations reveal _non-deterministic behavior_ influenced by hardware variations and other factors.

### Large Language Models Think Too Fast to Explore Effectively

- **Large Language Models (LLMs) struggle with effective exploration**, particularly in open-ended tasks, as they often make **premature decisions** due to their rapid processing speed, which contrasts with human strategies that balance uncertainty and empowerment.

### Theoretical limitations of multi-layer Transformer

- This work establishes the **first unconditional lower bound** for multi-layer decoder-only Transformers, demonstrating that any $L$-layer model requires a **polynomial model dimension** ($n^{\Omega(1)}$) to execute sequential compositions of $L$ functions over $n$ tokens.

### \\R Fully open source codebase to train SOTA VLMs

- **Hugging Face** has released a **fully open-source codebase** for training **SmolVLM**, enabling users to train state-of-the-art vision-language models (VLMs) on **256 H100 GPUs**.

### 3D Scene Reconstruction in Adverse Weather Conditions via Gaussian Splatting

- **WeatherGS** enhances **3D scene reconstruction** by effectively addressing artifacts from adverse weather, utilizing a novel **dense-to-sparse preprocess strategy** to improve clarity in reconstructed scenes.

### \\R Molecular Fingerprints Are Strong Models for Peptide Function Prediction

- **Molecular fingerprints** outperform complex models like GNNs and transformers in peptide classification, achieving **state-of-the-art results** on 126 datasets, including LRGB, without hyperparameter tuning.

### \\2412.20302 EXAdam: The Power of Adaptive Cross-Moments

- **EXAdam** is an advanced optimization algorithm that enhances the **Adam optimizer** with new debiasing terms, a gradient-based acceleration mechanism, and a dynamic step size formula, leading to improved convergence and robustness.

### \\News Tulu 3 model performing better than 4o and Deepseek?

- The **Tulu 3 model**, released by the **Allen Institute for AI**, reportedly **outperforms** both **4o** and **DeepSeek** in several benchmarks, showcasing advancements in AI model performance.

### Toward a Sparse Interpretable Audio Codec

- This work presents a **sparse audio codec** that encodes audio as a set of events, enhancing interpretability through a physics-based model that captures the **resonance** of instruments and environments, aiming for a more intuitive representation than traditional codecs.

### \\Discussion Reason for Activation Steering over finetuning?

- **Activation steering** offers a method to control language model behavior by adjusting neuron activations rather than retraining, allowing for _real-time intervention_ and fine-grained control without altering model weights.

### Accelerate DeepSeek Reasoning Models With NVIDIA GeForce RTX 50 Series AI PCs

- The **DeepSeek-R1 model family** leverages **NVIDIA GeForce RTX 50 Series GPUs**, achieving up to **3,352 trillion operations per second**, enabling unprecedented speed for reasoning models that excel in problem-solving and code capabilities.

### Research Focus: Week of January 27, 2025

- **FLAVARS** is a new multimodal foundation model that enhances remote sensing by combining contrastive learning and masked modeling, achieving a **+6% mIOU** improvement over SkyCLIP in vision-only tasks while maintaining zero-shot classification capabilities.
