The Feb 3, 2024 release of "Speech and Language Processing" by Dan Jurafsky and James H. Martin introduces updated chapters and slides, with a promise of more content soon.
What Extropic is Building
Extropic is developing a full-stack hardware platform that leverages matter's natural fluctuations for Generative AI, promising to extend hardware scaling beyond digital computing limits and enable AI accelerators that are significantly faster and more energy-efficient.
How to Write a Lisp Interpreter in Python
Peter Norvig demonstrates how to implement a Lisp interpreter in Python, focusing on the Scheme dialect, to illustrate the foundational concepts of software, as highlighted by Alan Kay's "Maxwell's Equations of Software."
Are We Watching the Internet Die?
Reddit's IPO, valued at $6.5bn, is criticized as a major swindle, exploiting millions of unpaid contributors for the financial gain of a few, including CEO Steve Huffman and investor Sam Altman, despite the company's consistent financial losses and controversial management decisions.
Behind the Compute: Benchmarking Compute Solutions
Stability AI's benchmarking reveals Intel Gaudi 2 accelerators outperform Nvidia's A100 and H100 in training Stable Diffusion 3, showcasing 1.5 times faster image processing and superior scalability.
Prompts as WASM Programs
The Artificial Intelligence Controller Interface (AICI) by Microsoft Research enables real-time control and direction of Large Language Model outputs through flexible Controllers, which can implement constrained decoding, dynamic editing, and coordinate execution across multiple generations.
Engineering Feedback at Figma with Engineering Critiques
Figma's engineering critiques (eng crits) foster a culture of early feedback on in-progress work, encouraging diverse perspectives and unblocking teams to explore new ideas without the pressure of approval processes.
Diffusion Models from Scratch
Diffusion models, known for their effectiveness in generative modeling across various domains like text-to-image, audio, video, and protein design, are explored from an optimization perspective, offering insights into their training and implementation for both simple and complex datasets as detailed in a tutorial based on a new paper.
Gemini 1.5 Pro Recall Information in 10M Context
Google's Gemini 1.5 Pro utilizes modern approaches like recurrent memory and ring attention to enhance its long-context capabilities, as detailed in a technical report.
ShortGPT: Layers in Large Language Models
ShortGPT demonstrates that many layers in Large Language Models (LLMs) are redundant, with some contributing minimally to the model's overall functionality, as revealed through a novel metric called Block Influence (BI). Read the paper
OpenAI: JSON Mode vs Functions
JSON mode in OpenAI's API forces GPT models to generate outputs as valid JSON strings, requiring explicit JSON structure specification within the prompt, enhancing structured data generation. JSON mode guide
Gemini 1.5: Unlocking Multimodal Understanding
Gemini 1.5 Pro significantly advances multimodal understanding, demonstrating near-perfect recall in long-context retrieval tasks and setting new benchmarks in long-document and long-video QA, as well as long-context ASR.
Bias-Augmented Consistency Training
Bias-augmented consistency training (BCT) significantly reduces biased reasoning in chain-of-thought prompting by training models to maintain consistent reasoning across biased and unbiased prompts.
GEAR: An Efficient KV Cache Compression Recipe
GEAR introduces an efficient KV cache compression framework for LLMs, significantly reducing memory constraints by applying quantization, low rank, and sparse matrix techniques for near-lossless compression.
Can't Remember Details in Long Documents? You Need Some R&R
R&R, combining reprompting and in-context retrieval (ICR), significantly enhances QA accuracy in long-context large language models by addressing the issue of missing crucial information within extensive documents. Read more