JetMoE: Reaching Llama2 Performance with 0.1M Dollars
JetMoE-8B, a new Large Language Model (LLM), achieves superior performance to Llama2 models with a budget of less than $0.1 million, utilizing a 1.25T token dataset and 30,000 H100 GPU hours.
Anyone got a contact at OpenAI. They have a spider problem
OpenAI's GPTBot fetched over 3 million pages from a single content farm, indicating a potential oversight in its web crawling algorithm.
Tinygrad: Hacked 4090 driver to enable P2P
A fork of NVIDIA's driver now includes P2P support for 4090 GPUs, enhancing direct GPU-to-GPU communication by utilizing the large BAR support and bypassing traditional limitations.
I Lost Faith in Kagi
Kagi's diversification into multiple unprofitable projects such as AI tools, a Mac-only browser, and a planned email service, alongside a t-shirt factory venture, stretches its resources thin, raising sustainability concerns.
DwarFS – Deduplicating Warp-Speed Advanced Read-Only File System
DwarFS achieves very high compression ratios for redundant data without compromising on speed, outperforming SquashFS in both compression and build speed, with detailed comparisons provided for various use cases including astrophotography and handling bit rot.
Adobe Is Buying Videos for $3 per Minute to Build AI Model
Adobe is actively purchasing videos at $3 per minute to amass content for its AI text-to-video generator, aiming to rival OpenAI's advancements in similar technology.
Quantum Algorithms for Lattice Problems
Yilei Chen introduces a quantum algorithm for the Learning With Errors (LWE) problem, leveraging Gaussian functions with complex variances and windowed quantum Fourier transform techniques, aiming to solve LWE in polynomial time.
Fivefold Slower Compared to Go? Optimizing Rust's Protobuf Decoding Performance
GreptimeDB's optimization journey revealed Rust's Protobuf decoding was initially five times slower than Go's implementation in VictoriaMetrics, prompting a deep dive into performance enhancements.
McCarthy's Ambiguous Operator (2005)
John McCarthy's amb operator, introduced in 1961, elegantly solves computational problems by selecting values that prevent future errors, a concept that predates but anticipates modern backtracking algorithms.
Amazon virtually kills efforts to develop Alexa Skills
Amazon will cease providing monthly AWS credits for hosting Alexa Skills by June 30, effectively ending its Alexa Developer Rewards program and disincentivizing third-party development.
NeurIPS 2024 Adds a New Paper Track for High School Students
NeurIPS 2024 introduces a new paper track specifically for high school students, focusing on machine learning for social impact, a move that underscores the conference's commitment to nurturing young talent and broadening the scope of contributions to the field.
Publication rat race for PhD in ML
PhD programs in ML at top universities now require applicants to have multiple first author papers and strong recommendations from reputable researchers, a standard once reserved for faculty position qualifications in other fields.
Infinite context Transformers
The Infinite Context Transformers paper introduces a novel approach to significantly extend the context length capabilities of transformer models, potentially explaining the Gemini 1.5's reported 10 million token context length.
An Auto-Regression Model for Object Recognition
The auto-regression model introduced at CVPR enables label prediction directly from images, bypassing the need for a predefined query gallery or class concepts, marking a significant shift towards more flexible object recognition systems.
Fine-Tuning Increases LLM Vulnerabilities and Risk
Fine-tuning and quantization significantly reduce jailbreak resistance in Large Language Models (LLMs), making them more susceptible to attacks such as jailbreaking, prompt injection, and privacy leakage.