Lossless LLM compression for efficient GPU inference via dynamic-length float
DFloat11 achieves a 30% reduction in LLM size while ensuring outputs are bit-for-bit identical to the original, utilizing entropy coding for optimal compression.
The Policy Puppetry Prompt: Novel bypass for major LLMs
HiddenLayer introduces a novel universal bypass that enhances security for all major LLMs, effectively protecting against inference, bypass, extraction attacks, and model theft without complicating existing models or requiring access to raw data.
Berkeley Humanoid Lite – open-source robot
Berkeley Humanoid Lite is an open-source humanoid robot designed to be accessible and customizable, with a total hardware cost under $5,000, promoting community engagement in robotics.
World Emulation via DNN
The project demonstrates a neural network capable of generating a playable world from real-world video data, showcasing the unique ability of neural worlds to create environments from any video, not just game footage.
Paper2Code: Automating Code Generation from Scientific Papers
PaperCoder automates the transformation of machine learning papers into functional code repositories through a structured three-stage process: planning, analysis, and generation, utilizing specialized agents for collaboration.
We compress any BF16 model to ~70% size during inference, while keeping the output LOSSLESS so that you can fit in more context or run larger models.
DF11 compresses BF16 models to ~70% size during inference while maintaining lossless output, allowing for larger context or model sizes without sacrificing accuracy, as detailed in the arXiv paper.
LLMs can see and hear without any training
MILS enables LLMs to process visual and auditory data without prior training, showcasing a significant advancement in cross-modal learning capabilities.
Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
PaperCoder automates the transformation of machine learning papers into functional code repositories, utilizing a multi-agent LLM framework that enhances reproducibility and accelerates research progress.
Cross-Encoder Rediscovers a Semantic Variant of BM25
BERT-based cross-encoders not only outperform BM25 but may also reimplement it semantically, revealing how MiniLM learns components akin to BM25 through mechanistic interpretability.
CosAE: Learnable Fourier Series for Image Restoration
CosAE (Cosine Autoencoder) innovatively combines Fourier series with a feed-forward neural network, enabling extreme spatial compression while preserving image detail during restoration.
Intuition behind Load-Balancing Loss in the paper OUTRAGEOUSLY LARGE NEURAL NETWORKS: THE SPARSELY-GATED MIXTURE-OF-EXPERTS LAYER
The Load-Balancing Loss in the paper "OUTRAGEOUSLY LARGE NEURAL NETWORKS" aims to ensure that all experts in a sparsely-gated mixture-of-experts model are utilized effectively, preventing any single expert from becoming overloaded while others remain underused.
Accelerate PyTorch 2.7 on Intel® GPUs
PyTorch 2.7 enhances performance on Intel GPUs, achieving up to 3x speedup in inference for models like Stable Diffusion, thanks to optimizations in scaled dot-product attention (SDPA) and the introduction of torch.compile on Windows.