Qwen2 introduces models across five sizes, with the largest being Qwen2-72B, and extends support for 27 additional languages, enhancing multilingual capabilities and performance in coding and mathematics.
How Does GPT-4o Encode Images?
GPT-4o encodes a 512x512 image tile as 170 embedding vectors, suggesting a sophisticated internal representation strategy beyond simple tokenization, with a cost implication of approximately 227 words per picture according to OpenAI's pricing.
AI in software engineering at Google: Progress and the path ahead
Google has significantly advanced AI in software engineering, with AI-powered tools now completing as much code as developers write manually, leveraging large language models (LLMs) for inline code completion and other tasks.
σ-GPTs: A new approach to autoregressive models
σ-GPTs revolutionize autoregressive models by introducing positional encoding for output, enabling dynamic order modulation and enhanced sampling flexibility.
Dragonfly: A large vision-language model with multi-resolution zoom
Dragonfly introduces a novel vision-language model architecture that leverages multi-resolution zoom-and-select for enhanced visual understanding, demonstrated through its application in both general and biomedical domains with open-source models Llama-3-8b-Dragonfly-v1 and Llama-3-8b-Dragonfly-Med-v1.
Hybrid Bonding: 3D Chip Tech to Save Moore's Law
Hybrid bonding, a technology enabling millions of connections within a square millimeter of silicon, marks a significant leap in 3D chip fabrication.
The illusion of state in state-space models
State-space models (SSMs), designed to surpass transformers in sequential computation and state tracking, fail to offer any real advantage in expressive power for these tasks.
[R] Scalable MatMul-free Language Modeling
MatMul-free models maintain strong performance at billion-parameter scales, eliminating the need for MatMul operations in large language models (LLMs), as detailed in the Arxiv study.
[R] Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Model
State-of-the-art Large Language Models (LLMs) exhibit a dramatic breakdown in reasoning when faced with simple, common sense problems, contradicting their purported strong function across various tasks.
Meta's Ad Algorithm Directs Black Users to For-Profit Colleges
Meta's ad algorithms have been found to exhibit racial bias, disproportionately directing Black users towards for-profit colleges, which are often more expensive and have predatory marketing practices.
[D] Is it me or does it seem like benchmarks are making language models worse?
Benchmarks may be detracting from the quality of language models, with LLama3 outperforming others like GPT-4o, which often ignores instructions.
[R] Testing LoRA initialisations
Orthogonal initializations for LoRA outperform standard methods, aiming for ΔW = 0 with fewer zero parameters, as explored through various strategies including reversing and purely orthogonal approaches.
[D] How to reconcile double descent with chincilla scaling laws?
The Chinchilla scaling law suggests LLMs should be trained on more data with fixed model size, contrasting with the 'double descent' concept which advocates for increasing model parameters as much as compute allows.
[P] Lightning-Fast Text Classification with LLM Embeddings on CPU
fastc introduces a Python library for efficient text classification on CPUs, leveraging small models and cosine similarity for tasks like sentiment analysis and spam detection without the need for fine-tuning.
[R] Extracting Concepts from GPT-4
OpenAI has published a report along with code and a visualizer for extracting concepts from GPT-4, building on similar efforts like the one by Anthropic. Anthropic report