# AI and Technology Updates

## Recent Articles

- **AI accelerators** are gaining traction, with the author designing a **JTAG** for a **systolic array** in a **Tiny Tapeout** experimental shuttle, emphasizing the need for in-silicon debug infrastructure as a critical component of the project.

- **LLM inference** presents unique challenges, primarily due to the **autoregressive Decode phase** of Transformer models, which shifts focus from compute to **memory and interconnect** issues.

- **ANN v3** achieves **200ms p99 latency** for **100 billion vectors** in a single search index, enabling over **1,000 queries per second (QPS)** while maintaining **92% recall**.

- **Vortex** is a new **columnar file format** designed to enhance performance in reading and writing data, addressing limitations of **Parquet** by enabling **late materialization** and compute functions on compressed data.

- **HAT (Heterogeneous Accelerator Toolkit)** enables Java developers to offload workloads to GPUs, achieving performance improvements from **7 GFLOP/s on CPUs to 14 TFLOP/s on NVIDIA A10 GPUs** through advanced programming abstractions and optimizations like ND-Range API and memory management.

- **High-bandwidth flash (HBF) technology** is projected to surpass high-bandwidth memory (HBM) in market size by 2038, as highlighted by Professor Kim Jung-ho from KAIST during a recent forum in Seoul.

- **KAOS** is a **Kubernetes-native framework** designed for deploying and orchestrating AI agents, enabling **multi-agent coordination** and **LLM integration** for enhanced functionality.

- The **motcpp** library, an **open-source C++17** project, achieves **10–100× speedup** over Python implementations for **multi-object tracking** in video frames, enhancing real-time performance and deployment ease.
