What is NextBit
NextBit provides serverless language-model inference and dedicated GPU endpoints through an OpenAI-compatible API, with prepaid usage credits and European infrastructure options.
Serverless pricing
As of August 5, 2026, Serverless API has no setup fee or commitment and is billed per million input and output tokens. Current examples include Qwen 3.5 35B at US$0.23 input and US$1.60 output, Qwen 3 30B at US$0.14 and US$0.55, and Qwen 3 14B at US$0.10 and US$0.24.
The serverless plan includes more than 30 ready-to-use models, shared managed infrastructure, automatic scaling, streaming and compatible SDK support. Services draw down a prepaid credit balance, and optional automatic top-ups can replenish the account at a chosen threshold.
Dedicated endpoints
Dedicated Endpoints use custom fixed monthly pricing based on the selected model, GPU configuration and capacity. They support catalogue, fine-tuned, private or customer-supplied models with single-tenant resources, guaranteed performance, predictable latency and privacy controls.
Official sources: NextBit model rates and deployment options, NextBit API documentation, NextBit billing terms.