The GPU rental market has dramatically shifted from a peak of $8/hr for H100s to under $2/hr, driven by oversupply, the rise of compute resellers, and a pivot towards fine-tuning existing models rather than training new ones from scratch.
The First GPU and Its Impact
The NVIDIA GeForce 256, launched 25 years ago, was the first GPU, revolutionizing gaming by offloading tasks from the CPU and enabling more detailed graphics, which laid the groundwork for advancements in AI.
ARIA: An Open Multimodal Model
Aria is an open multimodal native model that excels in integrating diverse information, achieving best-in-class performance across various multimodal, language, and coding tasks, with 3.9B and 3.5B activated parameters for visual and text tokens, respectively.
AMD's Latest GPU Developments
AMD's Instinct MI325X will now feature 256GB of HBM3e memory due to supply issues with 36GB stacks, while the upcoming MI355X will boast 288GB and an increased bandwidth of 8 TB/s.
SCUDA: Virtual GPU Technology
SCUDA enables remote GPU access for CPU-only machines, allowing developers to leverage powerful GPUs over a network, enhancing flexibility in resource allocation and application deployment.
Limitations of LLMs in Mathematical Reasoning
GSM-Symbolic reveals that LLMs struggle with mathematical reasoning, showing a 65% performance drop when a single irrelevant clause is added to questions, indicating a reliance on training data rather than true logical reasoning.
Vulnerabilities in E2EE Cloud Storage
End-to-end encrypted (E2EE) cloud storage is compromised: A cryptographic analysis reveals severe vulnerabilities across major providers like Sync, pCloud, and Seafile, allowing malicious servers to inject files, tamper with data, and access plaintext information.
Conway's Gradient of Life
Conway's Gradient of Life utilizes gradient descent to approximate the reversal of configurations in Conway's Game of Life, transforming a complex discrete problem into a manageable continuous optimization task.
Challenges with Formal Reasoning in LLMs
LLMs lack formal reasoning, as demonstrated by a recent study from Apple, which reveals that their performance is primarily based on sophisticated pattern matching, leading to significant variability in results with minor changes in input. Read the study.
Advances in Normalized Transformers
The normalized Transformer (nGPT) architecture employs unit norm normalization for all vectors, enabling faster learning and reducing training steps by a factor of 4 to 20 compared to traditional methods.
Grokking Phenomenon in Models
Grokking in binary logistic classification reveals a delayed generalization phenomenon, where models may overfit on nearly linearly separable data before achieving perfect generalization asymptotically.
Rule-Based Systems and Intelligent Behavior
Intelligent behavior in artificial systems emerges from the complexity of rule-based systems, with our study on elementary cellular automata (ECA) revealing that higher complexity correlates with enhanced model performance on reasoning tasks and chess move predictions.
The First LLM for Greek
Meltemi 7B Instruct v1.5 is the first Large Language Model (LLM) for Greek, developed by the Athena Research & Innovation Center, and is available in both llamafile and gguf formats on HuggingFace.
Efficient Attention Mechanisms
Rodimus introduces a data-dependent tempered selection (DDTS) mechanism within a linear attention framework, achieving significant accuracy while reducing memory usage, thus addressing the computational costs of traditional softmax attention in LLMs.
Adaptive Learning in LLMs
Composite Learning Units (CLUs) enable Large Language Models (LLMs) to engage in generalized, continuous learning without traditional parameter updates, enhancing their reasoning through iterative feedback and interaction.