ML Times
Jul 21, 2024
Google Distributed Cloud air-gapped appliance
- The Google Distributed Cloud air-gapped appliance is now generally available, offering real-time local data processing for AI use cases in harsh, disconnected, or mobile environments.
txtai: Open-source vector search and RAG for minimalists
- txtai acts as an all-in-one embeddings database for semantic search, LLM orchestration, and language model workflows, leveraging a combination of vector indexes, graph networks, and relational databases.
Tenstorrent Unveils High-End Wormhole AI Processors, Featuring RISC-V
- Tenstorrent's Wormhole AI processors leverage the RISC-V architecture to offer scalable high-performance computing, aiming to deliver cost-effective AI solutions.
[R] Perpetual: a gradient boosting machine which doesn't need hyperparameter tuning
- PerpetualBooster is a gradient boosting machine that eliminates the need for hyperparameter tuning by using a single
budgetparameter to control predictive power, simplifying the model training process.
- PerpetualBooster is a gradient boosting machine that eliminates the need for hyperparameter tuning by using a single
NYT: The Data That Powers AI Is Disappearing Fast
- A study by the Data Provenance Initiative reveals a significant reduction in publicly available data for A.I. training, with 5% of all data and 25% of high-quality data now restricted. Study details
DeepL's LLM Outperforms Google Translate, ChatGPT-4, and Microsoft
- DeepL's next-gen language model significantly outperforms Google Translate, ChatGPT-4, and Microsoft in translation quality, setting a new benchmark in the field.
Artificial consciousness: a perspective from the free energy principle
- The free energy principle (FEP) suggests a distinction between artificial systems that simulate consciousness and those that replicate it, focusing on causal flow as a key factor.
Invalid SMILES beneficial rather than detrimental to chemical language models
- Invalid SMILES strings, previously considered a limitation in chemical language models, actually enhance model performance by acting as a self-corrective mechanism that filters out low-likelihood samples, improving the quality of generated molecules.
[P] ChessGPT, 100,000x smaller than GPT-4, plays chess at 1500 Elo. By finding a skill vector, we can increase its win rate by 2.6x in out-of-distribution games.
- ChessGPT, with 25M and 50M parameters, achieves a 1500 Elo rating in chess, significantly smaller yet capable compared to GPT-4's 1.8T parameters, demonstrating efficient learning and application of complex rules without explicit instruction. GPT-4's parameters
[R] Discussion of ReFT Paper with lead author Zhengxuan Wu
- ReFT, a fine-tuning technique, outperforms LoRA by being 15x-60x more parameter efficient, showcasing its effectiveness in model optimization.
[R] VLLMs OCR capabilities
- Multimodal Large Language Models (VLLMs) are being explored for their OCR capabilities to potentially replace traditional OCR engines in processing industrial documents.
[R] Weak baselines and reporting biases lead to overoptimism in machine learning for fluid-related partial differential equations
- 79% of ML-for-PDE studies reviewed use weak baselines, leading to overoptimistic results about their performance compared to standard numerical methods.
TwoMinutePapers - NVIDIA’s Crazy New AI Paints With Images!
- NVIDIA's new AI technology allows users to paint with images instead of brushes, transforming noise into detailed images through a diffusion-based process, offering a novel approach to digital art creation.