ML Times
Jun 1, 2024
Daily Updates
Arthur Whitney releases an open-source K with MIT license
Shakti introduces a fast, fun, universal database and language aimed at high-demand sectors like hedge funds, banks, and Formula 1, claiming 100 times faster performance than competitors like Polars, DataTable, BigQuery, Redshift, Databricks, and Snowflake.
[R] CoPE: Contextual Position Encoding: Learning to Count What's Important
- Contextual Position Encoding (CoPE) enhances Large Language Models by allowing positioning to be context-dependent, enabling models to attend to more abstract concepts like the i-th sentence.
Superconducting Computer: Imec's plan to shrink datacenters
- Imec's superconducting computer technology promises to reduce data center sizes to that of a shoebox by leveraging superconductors' zero-resistance properties for energy-efficient computing, potentially revolutionizing AI processing and cloud-based training.
Transformers Represent Belief State Geometry in Their Residual Stream
- Transformers learn to represent the geometry of belief state updates from data-generating processes, such as Hidden Markov Models (HMMs), beyond merely predicting the next token.
LLMs Aren't "Trained on the Internet" Anymore
- LLMs are evolving beyond their initial training on internet data, incorporating custom and synthetic data to address their traditional limitations.
[D] KAN == multi-layer GAM ?
- The KAN paper introduces a method for stacking multiple layers of Generalized Additive Models (GAMs), suggesting that KANs can be viewed as multi-layer GAMs with the incorporation of an activation function.
[R] Lipreading with LipNet: End-to-End Sentence-level Lipreading
- The LipNet model has been re-implemented from scratch to predict sentences by analyzing lip movements, now using a 3DConv-LSTM (bi-directional) architecture instead of the original 3DConv-GRU.
[P] Automated LoRA Discovery
- The Automated LoRA Discovery project introduces a suite of methods to optimize the training of LoRAs by constraining the search space and using generative models, aiming to reduce the number of trainable parameters significantly. Full writeup and resources
[D] Can other areas researches such as the recent mapping of a cubic millimeter of a human brain tissue, help the Machine Learning field?
- The recent mapping of a cubic millimeter of human brain tissue offers unprecedented detail, potentially guiding the development of more efficient AI models by mimicking neural structures and functions.
[D] Is sequence packing common for training transformers?
- Sequence packing, as detailed in the EFFICIENT SEQUENCE PACKING WITHOUT CROSS-CONTAMINATION paper, involves combining multiple sequences into a single sample to enhance training efficiency without impacting performance.
[D] Bigram tokenizers better than status quo? Especially for multilingual
- Bigram tokenizers are proposed as a more efficient alternative to current models for multilingual language processing, especially for languages with a high number of word forms like Icelandic, which has over 1 million.
TwoMinutePapers - NVIDIA’s New AI: 5,000x Faster Virtual Worlds!
- NVIDIA's new AI technique transforms text into 3D worlds with unprecedented speed and quality, outperforming previous methods by being 5,000 times faster.