ML Times
Jul 4, 2024
Beating NumPy matrix multiplication in 150 lines of C
A custom C implementation of matrix multiplication outperforms NumPy's performance by achieving over 1 TFLOPS on an AMD Ryzen 7700 CPU, utilizing a design inspired by the BLIS framework and optimized through parallelization with OpenMP.AI's $600B Question
AI's revenue gap has widened from $200B to $600B, reflecting a significant discrepancy between the revenue expectations from AI infrastructure investments and actual growth in the ecosystem.Diffusion Forcing: Next-Token Prediction Meets Full-Sequence Diffusion
Diffusion Forcing introduces a novel training paradigm where a diffusion model denoises tokens with varied noise levels, merging next-token prediction's variable-length generation with full-sequence diffusion's trajectory guidance.[P] New collection of Llama, Mistral, Phi, Qwen, and Gemma models for function/tool calling
Rubra v0.1 introduces a collection of open-weight, tool-calling large language models (LLMs) including Llama, Mistral, Phi, Qwen, and Gemma, aiming to bridge the gap in function/tool calling capabilities between proprietary and open-source models. Try it out here[R] Taxonomy for Data Transformations in AI Systems
The Data Transformation Taxonomy introduced in the SIGMOD'24 paper categorizes data transformations in AI systems into reusable across models, model-specific, and real-time request-specific. Read the SIGMOD paperLikelihood computation in diffusion models [P]
Diffusion models, specifically the SDE approach detailed in Song et al. and Song et al., require the probability flow ODE for exact likelihood computation of generated data, due to the intractability of SDE likelihoods.TwoMinutePapers - ChatGPT Just Learned To Fix Itself!
OpenAI's new paper introduces a groundbreaking concept where AI systems critique and correct errors in other AI-generated outputs, enhancing code quality and security.MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context
The Medical Visual Hallucination Test (MedVH) introduces a benchmark dataset to assess hallucinations in domain-specific Large Vision Language Models (LVLMs) within the medical context, addressing a gap in evaluating these models' robustness.Improving Retrieval-augmented Text-to-SQL with AST-based Ranking and Schema Pruning
The approach dynamically retrieves database information and utilizes abstract syntax trees for selecting few-shot examples, aiming to improve Text-to-SQL semantic parsing with Large Language Models.Multiple-Resolution Tokenization for Time Series Forecasting with an Application to Pricing
The proposed transformer architecture focuses on multiple-resolution tokenization for time series forecasting, enhancing representation learning across various scales in the pricing domain.Let the Code LLM Edit Itself When You Edit the Code
The Positional Integrity Encoding (PIE) method significantly reduces computational overhead in real-time code editing by maintaining accurate positional relationships between tokens without the need for full re-encoding.SemioLLM: Assessing Large Language Models for Semiological Analysis in Epilepsy Research
SemioLLM evaluates state-of-the-art Large Language Models (LLMs) like GPT-3.5, GPT-4, Mixtral 8x7B, and Qwen-72chat for their efficacy in diagnosing epilepsy by analyzing unstructured text descriptions of seizures.LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
LoRA-Guard introduces a parameter-efficient guardrail adaptation method for content moderation on resource-constrained devices, leveraging knowledge sharing between LLMs and guardrail models.Large Language Models as Evaluators for Scientific Synthesis
Large Language Models (LLMs) like GPT-4 and Mistral are being tested for their ability to evaluate the quality of scientific syntheses, comparing their assessments to human annotators using a dataset of 100 research questions.GPTQT: Quantize Large Language Models Twice to Push the Efficiency
GPTQT introduces a novel post-training quantization method for Large Language Models (LLMs), significantly reducing memory usage and enhancing processing speed by quantizing model weights to 3bit/2bit.