ML Times

May 18, 2024

OpenAI's superalignment team faces significant departures

OpenAI's superalignment team faces significant departures, including co-founder Ilya Sutskever and co-team leader Jan Leike, amid concerns over the company's direction and safety-focused culture.

Computer scientists invent an efficient new way to count

Computer scientists have developed a new algorithm, named the CVM algorithm, that efficiently estimates the number of distinct elements in a data stream using a minimal memory footprint, leveraging randomness to simplify the process. Read the paper

Ubershaders: A Ridiculous Solution to an Impossible Problem (2017)

Ubershaders are a groundbreaking solution to the shader compilation stuttering issue in Dolphin Emulator, designed to emulate the GameCube/Wii's rendering pipeline directly on the GPU.

Ilya Sutskever: “If you learn all of these, you’ll know 90% of what matters”

Ilya Sutskever provided John Carmack with a reading list of approximately 30 research papers, claiming mastery of these would cover 90% of current essential knowledge in machine learning/AI.

Toon3D: Seeing cartoons from a new perspective

Toon3D innovatively recovers camera poses and dense geometry from hand-drawn scenes, overcoming their inherent lack of 3D consistency by employing piecewise-rigid deformation optimization at hand-labeled keypoints and leveraging monocular depth as a prior.

38% of webpages that existed in 2013 are no longer accessible a decade later

Link rot and digital decay significantly affect government, news, and other webpages, leading to the loss of online content over time.

Multi AI agent systems using OpenAI's assistants API

Experts.js simplifies the creation and deployment of OpenAI's Assistants, enabling them to be linked as Tools within a Panel of Experts system, enhancing memory and attention to detail.

LoRA Learns Less and Forgets Less

LoRA finetuning underperforms compared to full finetuning in programming and mathematics domains, but preserves base model performance on external tasks better.

Exact binary vector search for RAG in 100 lines of Julia

The Julia implementation for exact binary vector search in RAG (Retrieval-Augmented Generation) demonstrates state-of-the-art performance, significantly reducing server costs and memory requirements by converting 32-bit vectors to binary, shrinking a 1TB database to approximately 32GB.

How are subspace embeddings different from basic dimensionality reduction?

Subspace embeddings and basic dimensionality reduction techniques like PCA differ in their approach to uncovering latent structures, with subspace methods often targeting more complex relationships and structures within data.

LLM-generated code must not be committed without prior written approval by core

NetBSD's Commit Guidelines emphasize familiarity with code, legal clarity on code origins, and thorough testing before committing to the source tree, ensuring code integrity and compliance.

HMT: Hierarchical Memory Transformer for Long Context Language Processing

The Hierarchical Memory Transformer (HMT) introduces a novel framework that mimics human memory behavior to enhance long-context processing in language models, addressing the limitations of "flat" memory architectures in current transformer-based models.

Chrome DevTools now uses Gemini to help with JavaScript Errors in the console

Chrome DevTools introduces "Understand console messages with AI" to help developers get detailed explanations for errors and warnings in the Console, enhancing debugging efficiency.

Are PyTorch high-level frameworks worth using?

Exploring high-level frameworks like PyTorch Lightning and Ignite can enhance experiment tracking and hyperparameter management, potentially offering support for custom metrics like MAE.

Seminal papers list since 2018 that will be considered cannon in the future

Seminal papers since 2018 in machine learning are considered foundational, including works on Attention mechanisms, CLIP, and Vision Transformers.