Inference API Limits? - DeepTalk - Deep Learning Community

Inference API Limits?

post by junruilee on Feb 14, 2025

Does anyone know the specific API limits (e.g. Tokens per minute (TPM), Request per minute (RPM), etc.) applicable to the Inference API, per model? Could not find any documentation on the website. Apologies if I missed anything.

145 views
1 link

post by cody_b on Feb 14, 2025

@junruilee

There are no rate limits.

I just now added to the docs, “No limits are placed on the rate of requests.”

post by junruilee on Feb 14, 2025

Thank you for confirming.

Related topics

Topic Replies Views Activity
Model and content limit 2 174 Feb 2025
No way to establish spending limits 0 131 Mar 2025
Did inference api change recently?
Technical Help 0 119 Mar 2025
Lambda <> Openrouter Woes
Model Debugging 11 370 Jan 2025
Inference API Timeout
Technical Help 2 158 Jun 2025