# Dec 21, 2025

## Measuring AI Ability to Complete Long Tasks
- **AI's capability to complete long tasks** is assessed through a new framework that evaluates performance across various metrics, revealing significant insights into task management efficiency.

## Structured Outputs Create False Confidence
- **Structured outputs degrade response quality** by forcing models to prioritize format over accuracy, leading to errors in data extraction and reasoning, as demonstrated with receipt parsing examples.

## EGGROLL: Trained a Model Without Backprop and Found It Generalized Better
- **EGGROLL** demonstrates that optimizing **NDCG directly** using evolution strategies can outperform traditional contrastive loss methods, achieving a **22% improvement** in validation scores despite a lower training score.

## Why I Built KnowGraph: Static Knowledge Graphs for LLM-Centric Code Understanding
- **KnowGraph** innovates by creating **static knowledge graphs** from code repository artifacts, enhancing **structural awareness** and **explainability** in LLM-centric systems.

## Benchmarking Semantic vs. Lexical Deduplication on the Banking77 Dataset
- Result: 50.4% redundancy found using Vector Embeddings (all-MiniLM-L6-v2).

## A Memory Efficient TF-IDF Project in Python to Vectorize Datasets Larger Than RAM
