ML Times
Oct 21, 2024
Daily
Janus: Decoupling Visual Encoding for Multimodal Understanding and Generation
Janus is an innovative autoregressive framework that enhances multimodal understanding by decoupling visual encoding, allowing for improved flexibility and performance compared to traditional models.
3D-Printed Active Electronics
MIT researchers have developed semiconductor-free logic gates using 3D printing, enabling the potential for widespread electronics fabrication without traditional semiconductor facilities.
[R] Google Shopping 10M dataset for large scale multimodal product retrieval and ranking
The Marqo Google Shopping 10M dataset is now available on Hugging Face, featuring 10 million rows of data, including queries, product titles, images, and relevance ranks, making it a premier resource for multimodal product retrieval research.
All-optical switch device paves way for faster fiber-optic communication
An ultrafast all-optical switch developed by a University of Michigan team utilizes circularly polarized light to control light signals without electrical conversion, enhancing speed and energy efficiency in fiber-optic communication.
[R] RWKV-7: attention-free and surpassing strong Modded-GPT baseline (the one with Muon optimizer), while only using headsz 64
RWKV-7, an entirely RNN-based model, has demonstrated the ability to surpass the strong Modded-GPT baseline, achieving a loss of 3.26xx with potential for further optimization.
[D] Last Week in Medical AI: Top LLM Research Papers/Models (October 12 - October 19)
MedLFQA introduces a benchmark dataset for evaluating the factuality of long-form answers from medical LLMs, enhancing the reliability of AI-generated medical information.
Enhancing Large Language Models' Situated Faithfulness to External Contexts
Situated faithfulness in Large Language Models (LLMs) is crucial, as it allows them to dynamically assess the reliability of external information against their internal knowledge, addressing the issue of inaccurate or misleading contexts.
🤗Llama 3.2 in Keras
Llama 3.2 is fully operational in Keras, allowing users to load models directly from Hugging Face checkpoints with seamless conversion, enhancing accessibility for developers.
Large Language Models Are Overparameterized Text Encoders
Pruning the last $p%$ layers of large language models (LLMs) before supervised training can significantly reduce memory and inference time while maintaining performance, with up to 30% layers pruned with negligible impact and 80% with modest drop.
LoGU: Long-form Generation with Uncertainty Expressions
The LoGU framework addresses the challenge of long-form generation in Large Language Models (LLMs) by enabling them to express uncertainty, thus reducing the incidence of hallucinations in generated content.