# Dec 7, 2024

## Llama-3.3-70B-Instruct

- **Llama 3.3** is a significant update in the Llama series, featuring **transformers** and original repositories that enhance its capabilities for various applications.

## The Curse of Recursion: Training on Generated Data Makes Models Forget

- **Model Collapse** occurs when training on **model-generated content** leads to the loss of original content distribution, resulting in **irreversible defects** in generative models like Variational Autoencoders and LLMs.

## DSPy – Programming–not prompting–LMs

- **DSPy** is a framework that enables **programming language models** through modular AI systems, allowing for rapid iteration and optimization of prompts and weights, enhancing the quality of outputs without relying on fragile prompts.

## MIT largest open-source car design dataset, incl aerodynamics, to speed design

- MIT engineers have created **DrivAerNet++**, the largest open-source dataset of **over 8,000 car designs**, which includes detailed **aerodynamics** simulations to enhance the design of eco-friendly vehicles.

## How does OpenAI’s O1 outperform others in math despite limitations noted in recent papers?

- **OpenAI’s O1 model** outperforms other LLMs in mathematical reasoning by addressing limitations such as **memorization reliance** and **self-correction failures**, as highlighted in recent benchmarks.

## Google's AI weather prediction model is pretty darn good

- **Google's GenCast AI model outperformed traditional forecasting systems**, achieving accuracy over 97% against the ENS model, showcasing its potential to enhance weather prediction capabilities.

## Nucleotide Transformer: building robust foundation models for human genomics

- The **Nucleotide Transformer (NT)** models, with parameters ranging from **50 million to 2.5 billion**, are pre-trained on extensive genomic datasets, enabling accurate predictions of molecular phenotypes from DNA sequences, even in low-data scenarios.

## Ultralytics AI model hijacked to infect thousands with cryptominer

- The **Ultralytics YOLO11 AI model** was compromised in a **supply chain attack**, deploying a cryptominer on devices using versions **8.3.41 and 8.3.42** from PyPI, affecting thousands of users.

## JAX vs TensorFlow-XLA

- **JAX outperforms TensorFlow** due to its reliance on **XLA's JIT compilation**, which significantly enhances execution speed compared to TensorFlow's implementation.

## Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis

- **Switti** introduces a **scale-wise transformer** that significantly enhances **text-to-image generation** speed, outperforming traditional T2I AR models and rivaling advanced diffusion models.

## Towards Time Series Reasoning with LLMs

- **Novel multi-modal time-series LLM** approach demonstrates **zero-shot performance** in reasoning tasks, leveraging a lightweight encoder to extract time-series information effectively.

## Abstracts: NeurIPS 2024 with Dylan Foster

- **Dylan Foster's research** at NeurIPS 2024 investigates how existing reinforcement learning (RL) algorithms can be adapted to tackle high-dimensional observations and latent dynamics, aiming for faster learning in complex environments.

## Abstracts: NeurIPS 2024 with Pranjal Chitale

- **CVQA** is a new benchmark for **multilingual visual question answering**, encompassing **31 languages** and **30 cultures**, developed to enhance model inclusivity and cultural understanding in AI systems.

## An EPYC Exclusive for Azure: AMD's MI300C – By George Cozma

- **AMD's MI300C** powers Azure's new **HBv5 VMs**, featuring **96 Zen 4 cores** and **128GB of HBM3E** per EPYC 9v64H CPU, delivering unprecedented performance for high-demand applications.

## For a change of topic: some nonLLM focused work of mine: Bias-Free Sentiment Analysis through Semantic Blinding and Graph Neural Networks

- **Bias-Free Sentiment Analysis** through **Semantic Blinding** and **Graph Neural Networks**.
