Inference API Limits? - DeepTalk - Deep Learning Community
Inference API Limits?
post by junruilee on Feb 14, 2025
Does anyone know the specific API limits (e.g. Tokens per minute (TPM), Request per minute (RPM), etc.) applicable to the Inference API, per model? Could not find any documentation on the website. Apologies if I missed anything.
145 views
1 link
post by cody_b on Feb 14, 2025
There are no rate limits.
I just now added to the docs, “No limits are placed on the rate of requests.”
post by junruilee on Feb 14, 2025
Thank you for confirming.
Related topics
| Topic | Replies | Views | Activity |
|---|---|---|---|
| Model and content limit | 2 | 174 | Feb 2025 |
| No way to establish spending limits | 0 | 131 | Mar 2025 |
| Did inference api change recently? | |||
| Technical Help | 0 | 119 | Mar 2025 |
| Lambda <> Openrouter Woes | |||
| Model Debugging | 11 | 370 | Jan 2025 |
| Inference API Timeout | |||
| Technical Help | 2 | 158 | Jun 2025 |