What is Novita AI

Novita AI combines serverless APIs for language, image, audio and video models with on-demand GPU infrastructure and isolated agent sandboxes.

Serverless model pricing

As of August 5, 2026, serverless APIs are pay as you go with no monthly commitment. Current language-model rates per million tokens include DeepSeek V4 Pro at US$1.74 input, US$0.145 cached input and US$3.48 output; MiniMax M2.7 at US$0.30 input, US$0.06 cached input and US$1.20 output; GLM 5.1 at US$1.40, US$0.26 and US$4.40; and Gemma 4 31B at US$0.14 input and US$0.40 output.

Novita exposes more than 200 models through OpenAI- and Anthropic-compatible APIs. Other published examples include GLM text-to-speech at US$0.28 per million characters, speech recognition at US$0.021 per million characters and GLM voice cloning at US$0.83 per million characters.

GPU and agent infrastructure

GPU Instances provide dedicated machines with predictable resources, while Serverless GPU allocates compute for individual jobs and scales to zero when idle. Agent Sandbox is billed per second by requested vCPU and memory. Enterprise buyers can request dedicated clusters, higher-volume arrangements and support.

Official sources: Novita AI pricing, Novita platform and current model rates, Novita GLM 5.1 pricing and API guide.