Benchmarking Chinese Knowledge Rectification in Large Language Models
Large Language Models (LLMs) struggle with Chinese ancient poetry, proverbs, and idioms, leading to the generation of nonsensical information due to a lack of specific knowledge.
Transfusion: Predict the next token and diffuse images with one multimodal model
Transfusion introduces a novel training recipe that merges language modeling and diffusion techniques, enabling a single transformer to process both text and image data efficiently.
Radiology specific foundation model released by Harrison.ai
Harrison.rad.1 significantly outperforms other AI models, including OpenAI's GPT-4o and Google's Gemini 1.5 Pro, in the FRCR 2B Rapids exam, scoring 51.4 out of 60 (85.67%), where most competitors scored below 30.
Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Deepsilicon is developing software and hardware to train and run ternary transformer models, which compress weight matrices by almost 8x and reduce arithmetic intensity, offering a significant leap in efficiency for large transformer-based models.
Deductive Verification for Chain-of-Thought Reasoning in LLMs
Chain-of-Thought (CoT) prompting enhances Large Language Models' (LLMs) reasoning capabilities but risks introducing hallucinations and errors, necessitating a method for rigorous deductive reasoning and self-verification.
Sail – Faster and cheaper PySpark-compatible computation framework
Sail aims to unify stream processing, batch processing, and compute-intensive AI workloads, positioning itself as a versatile tool in data processing and AI fields.
[R] Revisiting Sparse Convolutional Model for Visual Recognition
Sparse convolutional models, bridging the gap between interpretability and empirical performance, employ differentiable optimization layers as replacements for standard convolutional layers in deep neural networks.
We're in the brute force phase of AI – once it ends, demand for GPUs will too
Gartner's chief of research for AI, Erick Brethenoux, claims that the current "brute force" phase of AI, heavily reliant on GPUs, is temporary, as history shows that specialized hardware becomes obsolete when general-purpose machines catch up.
Satellites Spotting Aircraft
Umbra Space operates a fleet of Synthetic Aperture Radar (SAR) satellites capable of capturing high-resolution images through obstacles like clouds and camouflage, with their first satellite launched by SpaceX in 2021.
[R] Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
LLM-generated ideas are judged more novel than those by human experts in a large-scale study involving over 100 NLP researchers, highlighting LLMs' potential in creative ideation.
Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
Talaria is a model visualization and optimization system designed to address the challenge of fitting machine learning models on devices with limited resources by enabling interactive optimization and visualization.
[P] I built a tool to minimize hallucinations with 1 hyperparameter search - Nomadic
The Nomadic tool significantly reduces hallucinations in Retrieval Augmented Generation pipelines by 4X with a single hyperparameter search, enhancing the reliability of generated content.
Is telling a model to "not hallucinate" absurd?
Instructing a Large Language Model (LLM) to "not hallucinate" is feasible, especially if it has been preference-fine-tuned with such directives, challenging initial skepticism about the effectiveness of this approach.
LowFormer introduces a hardware-efficient design for vision backbones by blending convolutions and transformer blocks, focusing on actual throughput and latency rather than just MACs for a more accurate efficiency metric.
[P] costly: a package for estimating costs & running times of LLM projects in advance
costly is a Python package designed to estimate costs and running times for LLM projects, addressing the need for pre-expenditure planning in complex LLM workflows.