ML Times
Mar 16, 2024
Ollama now supports AMD graphics cards
Ollama now accelerates its features using AMD graphics cards on both Windows and Linux, broadening its hardware compatibility.TextSnatcher: Copy text from images, for the Linux Desktop
TextSnatcher enables users to copy text from images swiftly using OCR technology, specifically leveraging the Tesseract OCR 4.x for character recognition.What are some well-written ML codebases to refer to get inspiration on good ML software design?
Scikit-learn exemplifies good ML software design with its intuitive fit/predict paradigm, serving as a benchmark for newcomers in machine learning.Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
Quiet-STaR generalizes the Self-Taught Reasoner (STaR) approach, enabling language models (LMs) to generate rationales for each token to elucidate future text, thereby enhancing prediction accuracy.Show HN: Matrix Multiplication with Half the Multiplications
The Free-pipeline Fast Inner Product (FFIP) algorithm and its hardware architecture, introduced by T. E. Pogue and N. Nicolici, significantly enhance ML accelerator efficiency by reducing the need for multiplier units by nearly half without sacrificing performance.Show HN: Open-source, browser-local data exploration using DuckDB-WASM and PRQL
Pretzel is an open-source, offline browser-based tool designed for fast and intuitive data exploration and visualization, leveraging WebAssembly-based DuckDB and PRQL for high performance.MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
A careful mix of image-caption, interleaved image-text, and text-only data is essential for achieving state-of-the-art few-shot results in multimodal large language models (MLLMs), outperforming other pre-training methods.ASCII art elicits harmful responses from 5 major AI chatbots
ASCII art tricks AI chatbots, including GPT-4 and Google’s Gemini, into bypassing safety protocols and providing harmful responses, exploiting a gap in their training on text interpretation.Show HN: Flash Attention in ~100 lines of CUDA
tspeterkim/flash-attention-minimal offers a simplified CUDA and PyTorch re-implementation of Flash Attention, aiming to be accessible and educational with a concise codebase. Link to articleAutoDev: Automated AI-driven development by Microsoft
AutoDev introduces a comprehensive AI-driven framework for software development, enabling autonomous planning and execution of complex engineering tasks beyond mere code suggestions.Reproducing the "Self-Rewarding Language Models" Paper by MetaAI
A team successfully reproduced the "Self-Rewarding Language Models" paper by MetaAI, implementing a process that enhances any base model through a self-improvement loop.ArXiv Papers as Audiobooks
The ArXiv Paper Reader algorithm transforms ArXiv papers into engaging video or audio formats, catering to both visual learners and those preferring auditory learning, by leveraging tools likelatex2html, OpenAI's GPT, and Google's Text-to-Speech API.Loongson 3A6000: A Star Among Chinese CPUs
The Loongson 3A6000 CPU marks a significant leap in China's domestic CPU development, featuring the new LA664 core architecture, which offers a 38% performance gain over its predecessor, the 3A5000, and introduces Simultaneous Multithreading (SMT) support for enhanced multithreaded performance.Y Combinator's chief startup whisperer is demoting himself
Michael Seibel steps down as Y Combinator's managing director to focus more on direct mentorship within the startup incubator, amidst its strategic downsizing and external political challenges.80 Level: AI Simulation Platform Where Characters Make Their Own Decisions
SAGA, a generative AI framework developed by Fable studio, enables AI agents to autonomously make decisions based on their skills, goals, and the context of their environment, drawing inspiration from NVIDIA's Voyager and research on simulating human behavior.