What is FriendliAI

FriendliAI is an inference platform for open-weight models with pay-per-token Model APIs, dedicated GPU endpoints and container deployment for private environments.

Model API pricing

Current serverless examples per million tokens include GLM-5.2 at US$1.40 input, US$0.26 cached input and US$4.40 output; Gemma 4 31B at US$0.14 input and US$0.40 output; DeepSeek V3.2 at US$0.50 input, US$0.25 cached and US$1.50 output; and MiniMax M2.5 at US$0.30 input, US$0.06 cached and US$1.20 output. Whisper Large v3 speech-to-text is listed at US$0.0015 per audio minute.

Dedicated and enterprise pricing

On-demand dedicated endpoints are billed per second with no startup charge: A100 80GB at US$2.90 per GPU-hour, H100 at US$3.90, H200 at US$4.50, B200 at US$8.90 and B300 at US$12. Container deployment and Enterprise are contact-sales routes.

Enterprise is a configurable contract rather than a fixed bundle. Published options include custom API rate limits, reserved GPU capacity, priority hardware access, custom regions, VPC or on-premises deployment, dedicated support, named customer success and custom commercial terms.

Official sources: FriendliAI pricing, FriendliAI documentation.