What is Cerebras Inference
Cerebras Inference serves open models on Cerebras systems through compatible APIs, offering free-trial, pay-as-you-go and dedicated-endpoint access.
Public endpoint pricing
As of August 4, 2026, OpenAI GPT OSS 120B costs US$0.35 per million input tokens and US$0.75 per million output tokens. Gemma 4 31B costs US$0.99 input and US$1.49 output, while the preview Z.ai GLM 4.7 costs US$2.25 input and US$2.75 output per million tokens. Cerebras says GLM 4.7 is scheduled for deprecation on August 17, 2026.
Public models are available on free-trial and pay-as-you-go tiers subject to rate limits. GPT OSS is the current production reasoning model; Gemma adds vision input, and the preview catalogue is intended for evaluation rather than production. Supported API features vary by model and include streaming, tools, structured outputs and reasoning controls.
Dedicated endpoints
Dedicated access uses custom flat monthly pricing based on required input and output token throughput, with reserved capacity, additional model families, higher throughput and production SLAs. Cerebras documents flexible contract terms of three, six or twelve months.
Official sources: Cerebras live public model catalogue and prices, Cerebras model catalogue, Cerebras pricing tiers.