What is Inceptron
Inceptron provides serverless language-model inference and dedicated NVIDIA GPU deployments from European infrastructure.
Serverless pricing
As of August 5, 2026, serverless rates are billed per model and token. Current prices per million tokens include GLM 5.2 at US$1.20 input, US$0.26 cached input and US$4.20 output; Kimi K2.7 Code at US$0.75 input, US$0.20 cached input and US$3.50 output; and Kimi K2.6 at US$0.73 input, US$0.25 cached input and US$3.50 output.
Dedicated deployments
Dedicated GPU prices are US$3 per hour for an H100 80GB, US$5 for an H200 141GB and US$8 for a B200 180GB. Commitments receive 5% off for one month, 10% for six months and 15% for twelve months.
Serverless deployments scale to zero and include compatible APIs, automatic failover and analytics. Enterprise capabilities include European data residency, ISO and GDPR controls, SSO, role-based access and audit logs.
Official sources: Inceptron pricing, Inceptron model rates.