What is Parasail

Parasail is an inference cloud for serverless, batch and dedicated deployment of open models and customer fine-tunes.

Serverless pricing

As of August 5, 2026, serverless inference is billed per million input, output and cached-input tokens, with no subscription amount. Current examples include DeepSeek V4 Flash at US$0.14 input, US$0.07 cache read and US$0.28 output; Gemma 4 26B at US$0.13, US$0.05 and US$0.40; GLM 5.2 at US$1.40, US$0.26 and US$4.40; and MiniMax M3 at US$0.30, US$0.06 and US$1.20.

Standard self-serve accounts have a published 500-request-per-minute limit and are charged automatically each time accrued spend reaches US$25. Serverless Free is rate-limited to five requests per minute, while higher dedicated and Enterprise tiers raise or remove limits.

Batch, dedicated and enterprise

Batch processing costs 50% of the corresponding serverless rate. Cached batch tokens receive another 30% discount, while FP16 models add 30% relative to FP8. Dedicated deployments are billed by GPU hour using the configuration shown at launch and support private Hugging Face models, LoRA adapters and autoscaling. Enterprise uses negotiated monthly invoicing, capacity and support terms.

Official sources: Parasail live pricing, Parasail pricing documentation, Parasail plan rate limits.