Financial Statement Analysis with Large Language Models
GPT4 outperforms financial analysts in predicting future earnings changes from standardized and anonymous financial statements, even without narrative or industry-specific information.
Mistral Fine-Tune
mistral-finetune employs LoRA for efficient finetuning, allowing for the training of only 1-2% of additional weights, optimizing for memory efficiency and performance on A100 or H100 GPUs.
Thermodynamic Natural Gradient Descent
Natural Gradient Descent (NGD), a second-order method, achieves comparable computational complexity to first-order methods when paired with specific hardware, sidestepping the traditional computational overhead associated with second-order training.
YOLOv10 introduces NMS-free training and a holistic model design strategy, significantly enhancing real-time object detection's performance and efficiency. Read the paper
[P] State-of-the-art, open source, Computer Vision models that are not ultra resource intensive?
Leading-edge Computer Vision (CV) models sought for inference on mid-tier GPUs like the A4000, with a preference for models beyond ResNet or YOLO and not necessarily CNN-based.
Lessons from the trenches on reproducible evaluation of language models
Effective evaluation of language models faces challenges such as sensitivity to setup, difficulty in comparisons, and lack of reproducibility, drawing on three years of experience for insights.
[R] Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
Dataset decomposition introduces a variable sequence length training technique that avoids the inefficiencies of the traditional concat-and-chunk method by preventing cross-document attention, leading to more effective and efficient training of large language models (LLMs). Read the paper