# Jan 17, 2026

## ML Times

### Daily

### Weekly

- **FLUX.2 [Klein]**: Towards Interactive Visual Intelligence  
  **FLUX.2 [klein]** introduces a **compact architecture** that achieves **sub-second inference** for image generation and editing, making it the fastest model family available, optimized for consumer hardware with just **13GB VRAM**.

- **Mamba-2**: Why Mamba rewrote its core algorithm and Microsoft abandoned RetNet  
  **Mamba-2** improved its algorithm by shifting from **parallel scans** to **block-diagonal GEMMs**, achieving **60-70% Tensor Core utilization**, significantly enhancing performance.

- **GLM-Image**: China just released first SOTA multimodal model trained entirely on domestic chips  
  **GLM-Image** is the first **SOTA multimodal model** trained entirely on **Chinese chips**, utilizing a hybrid architecture of autoregressive and diffusion decoders, showcasing impressive capabilities in **Chinese text rendering**.

- **Counterfactual evaluation** for recommendation systems  
  **Counterfactual evaluation** for recommendation systems addresses the **interventional nature** of recommendations, contrasting with traditional observational methods that fail to account for how recommendations influence user behavior.

- **vLLM-MLX**: Native Apple Silicon LLM inference - 464 tok/s on M4 Max  
  **vLLM-MLX** leverages Apple's **MLX** for **native GPU acceleration**, achieving **464 tok/s** on the M4 Max, making it a powerful tool for LLM inference.

- **Weight decay in RealNVP**: Does weight decay in RealNVP (Normalizing flows) encourage identity transforms?  
  **Weight decay in RealNVP may bias models towards the identity transform**, potentially undermining their capacity to learn meaningful representations, as small weights can push transformations to do nothing.
