# Oct 5, 2025

### ProofOfThought: LLM-based reasoning using Z3 theorem proving

- **ProofOfThought** leverages **LLM-based reasoning** with **Z3 theorem proving**, enabling users to query complex questions and receive logical answers, as demonstrated in the example where it predicts Nancy Pelosi's stance on abortion as **False**.

### Which Table Format Do LLMs Understand Best?

- **Markdown-KV format achieved the highest accuracy** at **60.7%**, outperforming CSV by approximately **16 points**, indicating its effectiveness for LLM data comprehension.

### NSA and IETF: Can an attacker purchase standardization of weakened cryptography?

- **The NSA and GCHQ are attempting to influence cryptographic standards by promoting weakened post-quantum cryptography**, potentially compromising security. This effort is likened to past instances where the NSA manipulated standards to favor weaker algorithms, raising concerns about the integrity of the standardization process.

### The Demonization of DeepSeek: How NIST Turned Open Science into a Security Scare

- **NIST's report on DeepSeek is a politically charged critique**, lacking evidence of malicious intent, and instead aims to undermine open science and protect corporate interests in AI development.

### Anthropic Release Memory API

- **New capabilities** on the Claude Developer Platform, including **context editing** and the **memory tool**, enhance agent performance by allowing for longer, uninterrupted tasks while preserving critical information across sessions.

### Matrix Core Programming on AMD GPUs

- **Matrix Cores** in AMD's CDNA™3 and CDNA™4 architectures significantly enhance performance for matrix operations, achieving up to **64x speedup** with low-precision types like FP4 compared to FP32, particularly in mixed-precision modes.

### Newton: physics simulation engine built upon NVIDIA Warp

- **Newton** is a **GPU-accelerated physics simulation engine** designed for roboticists, leveraging **NVIDIA Warp** and integrating **MuJoCo Warp** for enhanced performance and flexibility in simulations.

### What GPT-OSS leaks about OpenAI's training data

- **OpenAI's GPT-5 model** has been found to include phrases from **adult websites** in its training data, revealing potential vulnerabilities in the model's training stack.

### How to inject knowledge efficiently? Knowledge infusion scaling law for LLMs

- **Knowledge infusion** during pretraining can significantly enhance large language models' (LLMs) performance on specialized tasks, but it requires careful management to avoid _catastrophic forgetting_ of prior knowledge.

### XiangShan Vector Floating-Point Unit Design

- The **Vector Floating-Point Unit (VFPU)** supports a range of operations including **vector floating-point multiplication, addition, division, and square root** calculations, with capabilities for mixed-precision formats such as **fp16, fp32, and fp64**.

### Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR

- The proposed **PACS** framework reformulates **Reinforcement Learning with Verifiable Rewards (RLVR)** into a supervised learning task, enhancing stability and efficiency in training by coupling actor and critic roles implicitly.

### Machine Learnability as a Measure of Order in Aperiodic Sequences

- **Machine learning models** can effectively measure the **regularity of prime number fields** in the Ulam spiral, revealing that regions around **500m** are more learnable than those below **25m**.

### Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking

- The proposed framework, **$\mathbf{Li	ext{ }2}$**, delineates three stages of grokking in 2-layer nonlinear networks: **lazy learning**, **independent feature learning**, and **interactive feature learning**, revealing how gradient dynamics influence feature emergence.

### [P] How we used token healing to build a better autocomplete model

- **Token healing enhances autocomplete accuracy** by addressing the issue of LLMs mispredicting completions due to tokenization mismatches, allowing models to generate suggestions that align with user input more effectively.
