What is DeepInfra

DeepInfra provides hosted inference for more than 100 open and commercial AI models through OpenAI-compatible and native APIs, alongside private model deployments and on-demand GPU instances.

Serverless inference pricing

As of August 5, 2026, serverless inference has no monthly subscription, seat fee or minimum commitment. Each model has its own per-token or output-unit rate. Current examples per million tokens include GLM 5.2 at US$0.93 input and US$3 output, Kimi K2.7 Code at US$0.74 input and US$3.50 output, NVIDIA Nemotron 3 Ultra at US$0.50 input and US$2.20 output, and DeepSeek V4 Flash at US$0.09 input and US$0.18 output.

DeepInfra also lists DeepSeek V4 Pro at US$1.30 input and US$2.60 output, Qwen 3.6 35B at US$0.15 input and US$0.95 output, and Gemma 4 26B at US$0.07 input and US$0.34 output per million tokens. Cached-input prices are model specific, and Priority Inference adds 20% to the selected model's standard rate.

Private models and compute

Private models are billed by GPU time on dedicated or autoscaling A100, H100, H200, B200 or B300 infrastructure. On-demand DGX B300 currently starts at US$4.89 per instance hour. The platform includes compatible APIs, streaming, structured outputs and tool support where the chosen model provides them, plus monitoring and private deployments for production workloads.

Official sources: DeepInfra live model and compute pricing, DeepInfra documentation.