ProofOfThought: LLM-based reasoning using Z3 theorem proving
ProofOfThought leverages LLM-based reasoning with Z3 theorem proving, enabling users to query complex questions and receive logical answers, as demonstrated in the example where it predicts Nancy Pelosi's stance on abortion as False.
Which Table Format Do LLMs Understand Best?
Markdown-KV format achieved the highest accuracy at 60.7%, outperforming CSV by approximately 16 points, indicating its effectiveness for LLM data comprehension.
NSA and IETF: Can an attacker purchase standardization of weakened cryptography?
The NSA and GCHQ are attempting to influence cryptographic standards by promoting weakened post-quantum cryptography, potentially compromising security. This effort is likened to past instances where the NSA manipulated standards to favor weaker algorithms, raising concerns about the integrity of the standardization process.
The Demonization of DeepSeek: How NIST Turned Open Science into a Security Scare
NIST's report on DeepSeek is a politically charged critique, lacking evidence of malicious intent, and instead aims to undermine open science and protect corporate interests in AI development.
Anthropic Release Memory API
New capabilities on the Claude Developer Platform, including context editing and the memory tool, enhance agent performance by allowing for longer, uninterrupted tasks while preserving critical information across sessions.
Matrix Core Programming on AMD GPUs
Matrix Cores in AMD's CDNA™3 and CDNA™4 architectures significantly enhance performance for matrix operations, achieving up to 64x speedup with low-precision types like FP4 compared to FP32, particularly in mixed-precision modes.
Newton: physics simulation engine built upon NVIDIA Warp
Newton is a GPU-accelerated physics simulation engine designed for roboticists, leveraging NVIDIA Warp and integrating MuJoCo Warp for enhanced performance and flexibility in simulations.
What GPT-OSS leaks about OpenAI's training data
OpenAI's GPT-5 model has been found to include phrases from adult websites in its training data, revealing potential vulnerabilities in the model's training stack.
How to inject knowledge efficiently? Knowledge infusion scaling law for LLMs
Knowledge infusion during pretraining can significantly enhance large language models' (LLMs) performance on specialized tasks, but it requires careful management to avoid catastrophic forgetting of prior knowledge.
XiangShan Vector Floating-Point Unit Design
The Vector Floating-Point Unit (VFPU) supports a range of operations including vector floating-point multiplication, addition, division, and square root calculations, with capabilities for mixed-precision formats such as fp16, fp32, and fp64.
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
The proposed PACS framework reformulates Reinforcement Learning with Verifiable Rewards (RLVR) into a supervised learning task, enhancing stability and efficiency in training by coupling actor and critic roles implicitly.
Machine Learnability as a Measure of Order in Aperiodic Sequences
Machine learning models can effectively measure the regularity of prime number fields in the Ulam spiral, revealing that regions around 500m are more learnable than those below 25m.
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
The proposed framework, $\mathbf{Li ext{ }2}$, delineates three stages of grokking in 2-layer nonlinear networks: lazy learning, independent feature learning, and interactive feature learning, revealing how gradient dynamics influence feature emergence.
[P] How we used token healing to build a better autocomplete model
Token healing enhances autocomplete accuracy by addressing the issue of LLMs mispredicting completions due to tokenization mismatches, allowing models to generate suggestions that align with user input more effectively.