ML Times
Mar 12, 2024
OpenAI – transformer debugger release
Transformer Debugger (TDB), developed by OpenAI's Superalignment team, facilitates the investigation of small language models by combining automated interpretability techniques with sparse autoencoders, enabling users to understand and manipulate model behaviors without coding.
Building Meta's GenAI infrastructure
Meta announces the creation of two 24k GPU clusters designed for high throughput and reliability across various AI workloads, including the training of Llama 3, a next-generation AI model.
Diffusion models from scratch, from a new theoretical perspective
Diffusion models, known for their effectiveness in generative modeling across various domains like text-to-image, audio, video, and protein design, are explored from an optimization perspective, offering insights into their training and implementation for both simple and complex datasets as detailed in a tutorial based on a new paper.
Devin: AI Software Engineer
Devin, the world's first fully autonomous AI software engineer, has been introduced by Cognition Labs, showcasing abilities to plan, execute complex engineering tasks, and learn over time.
Simpson's paradox
Simpson's paradox occurs when a trend observed within multiple groups reverses when the groups are combined, highlighting the complexity of statistical data interpretation.
Stealing Part of a Production Language Model
Researchers have developed a model-stealing attack capable of extracting the embedding projection layer from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2, revealing their hidden dimensions.
Behind the Compute: Benchmarking Compute Solutions
Stability AI's benchmarking reveals Intel Gaudi 2 accelerators outperform Nvidia's A100 and H100 in training Stable Diffusion 3, showcasing 1.5 times faster image processing and superior scalability.
Show HN: Prompts as WASM Programs
The Artificial Intelligence Controller Interface (AICI) by Microsoft Research enables real-time control and direction of Large Language Model outputs through flexible Controllers, which can implement constrained decoding, dynamic editing, and coordinate execution across multiple generations.
Is Cosine-Similarity of Embeddings Really About Similarity?
Cosine-similarity, often used to quantify semantic similarity between high-dimensional objects, can produce arbitrary and potentially meaningless results when applied to embeddings from regularized linear models.
(How to Write a (Lisp) Interpreter (In Python)) (2010)
Peter Norvig demonstrates how to implement a Lisp interpreter in Python, focusing on the Scheme dialect, to illustrate the foundational concepts of software, as highlighted by Alan Kay's "Maxwell's Equations of Software."
[D] How does Gemini 1.5 Pro recall information in 10M context?
Google's Gemini 1.5 Pro utilizes modern approaches like recurrent memory and ring attention to enhance its long-context capabilities, as detailed in a technical report.
[R] ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
ShortGPT demonstrates that many layers in Large Language Models (LLMs) are redundant, with some contributing minimally to the model's overall functionality, as revealed through a novel metric called Block Influence (BI). Read the paper
Among the A.I. doomsayers
A.I. doomsayers and techno-optimists are deeply divided over whether artificial intelligence will lead to humanity's utopia or its destruction, with debates often centered in the Bay Area's "Cerebral Valley."
[D] Can someone please clarify if web search LLMs like Perplexity, You.com, or Coral Search are crawling the entire web themselves? Otherwise, how do they differ from simply combining a search API with any LLM model?
Web search LLMs like Perplexity, You.com, or Coral Search may or may not be crawling the web themselves, raising questions about their operational uniqueness and the source of their search results.
Steve Wozniak and Stuart Brand Discuss Control of IP (1984)
At the first Hackers Conference in 1984, Stewart Brand famously articulated the dual nature of information, stating it wants to be both expensive and free, reflecting on the tension between intellectual property control and the inherent desire for information dissemination.