ML Times
May 31, 2024
News Highlights
Imprecise language models are smaller, speedier, and nearly as accurate.
1-bit large language models (LLMs) offer a promising solution to AI's energy consumption by being smaller and faster, yet maintaining near-par accuracy with traditional models.
Better RAG Results with Reciprocal Rank Fusion and Hybrid Search
Reciprocal Rank Fusion (RRF) and Hybrid Search significantly enhance the performance of Retrieval Augmented Generation (RAG) systems by combining keyword and vector search results for more accurate query responses.
Legal models hallucinate in 1 out of 6 (or more) benchmarking queries
AI-driven legal research tools like LexisNexis and Thomson Reuters still produce incorrect information in more than 17% of cases, despite claims of reducing errors compared to general-purpose models like GPT-4.
Data drift detection methods fail to accurately predict model performance degradation.
I ran 580 model-dataset experiments to show that, even if you try very hard, it is almost impossible to know that a model is degrading just by looking at data drift results.
YOLOv5 on FPGA with Hailo-8 and 4 Pi Cameras
The multi-camera YOLOv5 project on Zynq UltraScale+ leverages the Hailo-8 AI accelerator for enhanced performance in machine vision applications, demonstrating a significant improvement over FPGA-based neural network implementations.
Superconducting Computer: Imec's plan to shrink datacenters
Imec's superconducting computer technology promises to reduce data center sizes to that of a shoebox by leveraging superconductors' zero-resistance properties for energy-efficient computing, potentially revolutionizing AI processing and cloud-based training.
NPGA: Neural Parametric Gaussian Avatars – high-fidelity digital faces
NPGA introduces a data-driven approach to create high-fidelity, controllable avatars from multi-view video recordings, leveraging 3D Gaussian Splatting for efficient rendering and topological flexibility.
Incorporating RoPE for positional encoding and Flash Attention could modernize BERT, enhancing its understanding of position and attention efficiency.
KAN == multi-layer GAM?
The KAN paper introduces a method for stacking multiple layers of Generalized Additive Models (GAMs), suggesting that KANs can be viewed as multi-layer GAMs with the incorporation of an activation function.
Lipreading with LipNet: End-to-End Sentence-level Lipreading
The LipNet model has been re-implemented from scratch to predict sentences by analyzing lip movements, now using a 3DConv-LSTM (bi-directional) architecture instead of the original 3DConv-GRU.
Machine learning introspection aims to equip AI with human-like intuition for problem-solving, a concept with roots in Newell and Simon's 1958 "General Problem Solver."
ANAH: Analytical Annotation of Hallucinations in Large Language Models
ANAH is a bilingual dataset designed to measure and annotate hallucinations in Large Language Models (LLMs) within Generative Question Answering, aiming to address the critical issue of hallucination in LLM applications.
NVIDIA’s New AI: 5,000x Faster Virtual Worlds!
NVIDIA's new AI technique transforms text into 3D worlds with unprecedented speed and quality, outperforming previous methods by being 5,000 times faster.
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
DITTO-2 significantly speeds up music generation by distilling a pre-trained diffusion model, enabling faster-than-real-time generation with enhanced control over music inpainting, outpainting, intensity, melody, and structure.
LLMs exhibit sensitivity to the sequence of options in multiple-choice questions
This observation underscored by experiments with the LLAMA3-8B model and supported by research, including studies found on arXiv and arXiv.