Sophia: Scalable Stochastic 2nd-Order Optimizer for Language Model Pre-Training
Sophia, a scalable second-order optimizer, significantly reduces the time and cost of language model pre-training by using a lightweight estimate of the diagonal Hessian for preconditioning and element-wise clipping for update control.
Language models are Super Mario: Absorbing abilities from homologous models
Language Models (LMs) can now absorb new abilities from similar models through a novel technique called DARE, which simplifies the merging of capabilities without the need for retraining or advanced hardware.
More Agents Is All You Need: LLMs performance scales with the number of agents
Large language models (LLMs) improve in performance through a simple sampling-and-voting method, which scales with the number of agents used, demonstrating a straightforward yet effective enhancement technique.
Chisel: A fast TCP/UDP tunnel over HTTP
Chisel is a fast TCP/UDP tunnel over HTTP, secured with SSH, combining both client and server functionalities into a single Go executable, designed for firewall traversal and secure network entry.
What John von Neumann Did at Los Alamos (2020)
John von Neumann's contributions at Los Alamos extended beyond his minor role in the Manhattan Project, significantly impacting the development of computing and nuclear weapons.
[D] ML researchers who are not in NLP, what are you researching? Please share.
ML researchers outside the NLP domain are invited to share their research areas, aiming for a comprehensive view of the spectrum of ML research.
Mixture-of-Depths: Dynamically allocating compute in transformers
Transformers can learn to dynamically allocate FLOPs to specific sequence positions, optimizing compute across different layers, enhancing efficiency.
Loki: An open-source tool for fact verification
Loki is an open-source tool designed to automate fact verification, providing a pipeline that includes decomposing texts into claims, assessing their significance, generating queries, crawling for evidence, and verifying claims, aimed at journalists and researchers.
Dot – A standalone open source app meant for easy use of local LLMs and RAG
Dot is an open-source application designed for seamless interaction with documents via local Large Language Models (LLMs), specifically Retrieval Augmented Generation (RAG), without requiring programming knowledge, and is bundled with Mistral 7B for out-of-the-box functionality.
Tokens, n-grams, and bag-of-words models (2023)
Tokens and n-grams serve as foundational elements in Natural Language Processing (NLP), enabling the understanding and manipulation of text by breaking it down into manageable units for analysis.
SentenceTransformers: Python framework for sentence, text and image embeddings
Schedule-Free learning employs a novel approach by replacing the momentum of an underlying optimizer with interpolation and averaging, facilitating faster training without the need for predefined schedules.
[D] what do you do with paper with no code published
Many papers introduce models with minor modifications to existing ones, targeting specific problems, yet often lack published code, complicating result reproduction.
Deep Aphantasia: a visual brain with minimal influence from priors?
Deep Aphantasia is characterized by minimal influence from prior expectations or inhibitory feedback on visual experiences, leading to atypical experiences of actual visual inputs, as detailed in a study by Loren N. Bouyer and Derek H. Arnold. Read more
Any statisticians who decided on a PhD in CS rather than a PhD in Stats? [D]
The MS stats student expresses a shift in interest from classical statistics to deep learning applications in time series forecasting, motivated by the desire for more modern research topics.