What is Nebius Token Factory
Nebius Token Factory is a managed inference and post-training platform for open models, with public serverless endpoints, dedicated deployments, batch processing and an OpenAI-compatible API.
Serverless and batch pricing
As of August 5, 2026, new accounts receive US$1 in free credit. Public endpoints use per-token pay-as-you-go pricing with economical base and higher-throughput fast flavours. Current base rates per million tokens include GPT OSS 120B at US$0.15 input and US$0.60 output, Kimi K2 Instruct at US$0.50 and US$2.40, Qwen 3 Coder 480B at US$0.40 and US$1.80, and Llama 3.3 70B at US$0.13 and US$0.40.
Where available, fast endpoints cost more for lower latency; DeepSeek R1 0528, for example, is US$2 input and US$6 output per million tokens in fast mode versus US$0.80 and US$2.40 in base mode. Batch inference is billed at 50% of the base real-time model price, rounded up to the nearest cent.
Dedicated and enterprise inference
Dedicated endpoints are billed per GPU hour with per-minute granularity and provide isolated capacity, selectable regions, custom weights, controlled autoscaling and no standard shared-service rate limits. Enterprise arrangements add volume discounts, governed workspaces, SSO, unified billing, auditability and a 99.9% uptime SLA.
Official sources: Nebius Token Factory pricing, Dedicated endpoint documentation, Token Factory platform overview.