What is Chutes

Chutes provides public model inference and self-serve private deployments on confidential GPU infrastructure, with USD or TAO billing.

Pay as you go

As of August 4, 2026, public inference has no subscription or minimum and is billed per million input and output tokens at each model's live rate. Current examples include GLM 5.1 at US$0.98 input and US$3.08 output and GLM 5.2 at US$1.25 input and US$3.95 output per million tokens.

Plus, Pro and Enterprise

Plus costs US$10 per month and includes a bundled daily request quota plus 6% off pay-as-you-go rates after the quota. Pro costs US$20 per month with a larger daily quota and a 10% overage discount. Enterprise uses custom pricing for volume discounts, custom rate limits and dedicated support.

Private Chutes

Private deployments are billed by the second. The current verified confidential-GPU listing starts at US$1.80 per hour for an RTX Pro 6000, with a one-time deployment fee of three times the hourly rate, currently US$5.40 at that tier. Private Chutes support custom code, private weights, fine-tunes and dedicated production instances.

Official source: Chutes pricing.