What is Replicate

Replicate is an AI infrastructure and model-hosting platform built around running, deploying, and serving models through a common API. It is best known for giving developers access to a large catalogue of public models alongside tools for deploying their own private models.

Replicate’s pricing is unlike a seat-based AI subscription because it is fundamentally usage-based. Buyers need to understand whether their workload is public-model inference, private-model hosting, or custom deployment, because the commercial mechanics differ meaningfully.

Core offerings

  • Public-model access through a common API.
  • Private model deployment for custom workloads.
  • Hardware-based billing for hosted workloads and time-based inference.
  • Enterprise options including volume discounts and support.

Pricing

As of July 19, 2026, Replicate’s official pricing page says you only pay for what you use. Public models are generally billed by runtime on the hardware used, while some models are billed by input and output tokens or by generated assets. Replicate gives per-model cost estimates directly on each model page.

The hardware pricing table currently lists examples such as CPU Small at US$0.000025 per second (US$0.09 per hour), CPU at US$0.000100 per second (US$0.36 per hour), Nvidia T4 at US$0.000225 per second (US$0.81 per hour), Nvidia L40S at US$0.000975 per second (US$3.51 per hour), Nvidia A100 80GB at US$0.001400 per second (US$5.04 per hour), and Nvidia H100 at US$0.001525 per second (US$5.49 per hour). Multi-GPU tiers are also published.

Replicate also documents private-model hosting as dedicated hardware billing, where customers pay for online instance time including setup, idle, and active processing time. Enterprise offerings add higher GPU limits, support, onboarding help, performance SLAs, and volume discounts for larger spend.

Model footprint

Replicate’s real strength is model breadth rather than one first-party model family. Its pricing model and product structure are designed for developers who want broad inference access and custom deployment rather than a polished consumer assistant.

Why select Replicate

Replicate is worth shortlisting when you need flexible developer access to many models and want hosting, deployment, and inference under one API-friendly roof. It is especially relevant for product teams that care more about programmatic model access than seat-based SaaS plans.

Official source: Replicate pricing.