ML Times

Mar 17, 2025

Undergraduate Disproves 40-Year-Old Conjecture, Invents New Kind of Hash Table

An undergraduate, Andrew Krapivin, has developed a new type of hash table that significantly speeds up data searches, contradicting a 40-year-old conjecture by Andrew Yao regarding search efficiency limits.

Akira ransomware can be cracked with sixteen RTX 4090 GPUs in around ten hours

Akira ransomware can be cracked using sixteen RTX 4090 GPUs in about ten hours, leveraging a new brute-force method that exploits its outdated encryption techniques, allowing affected companies to recover their data without paying the ransom.

Can a LLM convert C, to ASM to specs and then to a working Z/80 Speccy tape? Yes

LLMs can effectively convert C code to assembly and generate functional specifications, enabling the creation of a working Z/80 Spectrum application, as demonstrated by a sales tax calculator project that transitioned from C to ASM and then to a high-level specification.

I fine-tuned Qwen 2.5 Coder on a single repo and got a 47% improvement in code completion accuracy

Fine-tuning the Qwen 2.5 Coder model on a single repository resulted in a 47% improvement in code completion accuracy, elevating performance from 25% to 36% after just 500 iterations on an RTX 4090 GPU.

Milestone XAI/Interpretability papers?

Key papers in XAI include Axiomatic Attribution for Deep Networks and Sanity Checks for Saliency Maps, which introduce foundational concepts that reshape our understanding of interpretability in AI.

Double Descent in neural networks

Double descent in neural networks refers to a phenomenon where model performance improves after initial overfitting, leading to a second descent in error rates as model complexity increases.

Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models

State Space Models (SSMs) are gaining traction as an efficient alternative to transformer-based models, particularly excelling in tasks involving sequential data and longer contexts, showcasing comparable performance with notable efficiency improvements.

Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

Block diffusion language models merge the strengths of autoregressive and diffusion models, enabling flexible-length generation and enhanced inference efficiency through techniques like KV caching and parallel token sampling.

4D Language Fields for Dynamic Scenes via MLLM-Guided Object-wise Video Captioning

4D LangSplat integrates 4D Gaussian Splatting with multimodal LLMs to create language-aware scene representations, enabling 3D-aware grounding of language in dynamic environments without scene-specific training.

GTC 2025 – Announcements and Live Updates

GTC 2025 will showcase cutting-edge advancements in AI, robotics, and accelerated computing, featuring influential speakers like Yann LeCun and Frances Arnold, who will challenge conventional thinking and inspire innovation.

BriLLM: Brain-inspired Large Language Model

BriLLM is the first brain-inspired large language model that diverges from traditional architectures, utilizing a Signal Fully-connected flowing (SiFu) approach to enhance interpretability across all nodes in the model's directed graph.

AIstorian lets AI be a historian: A KG-powered multi-agent system for accurate biography generation

AIstorian is a KG-powered multi-agent system that significantly enhances biography generation by addressing challenges like factual fidelity and stylistic adherence that traditional LLMs struggle with.

Collaboration is all you need: LLM Assisted Safe Code Translation

UniTranslator redefines code translation by utilizing a collaborative framework of multiple compact LLMs, achieving accuracy and efficiency comparable to larger models while focusing on specialized tasks.

Limits of KV Cache Compression for Tensor Attention based Autoregressive Transformers

KV cache compression in autoregressive transformers is a critical factor limiting the context length of large language models (LLMs), with our research extending previous findings on space complexity to tensor attention mechanisms.