What is Together AI?
Together AI provides serverless and dedicated inference, fine-tuning, GPU infrastructure, code execution and model deployment for developers.
Core offerings
- Usage-priced serverless inference across text, image, video and audio.
- Dedicated model endpoints and GPU clusters.
- Fine-tuning and custom-model deployment.
- Embeddings, moderation and code execution services.
Pricing
Together AI uses prepaid usage billing and requires a minimum US$5 credit purchase; there is no general free trial. Serverless models are priced per token, image, video or audio unit, while dedicated deployments are billed for provisioned hardware. Enterprise and high-volume arrangements can include custom capacity and discounts.
API and model availability
The live serverless catalogue currently documents 103 distinct non-retired entries: 25 chat, 30 image, 37 video, nine audio, one embedding and one moderation model. They are reproduced on the API Cost page from Together's official catalogue; rerank models are excluded because Together states none are currently offered serverlessly.
Why select Together AI?
Together AI is best suited to engineering teams that value a broad, frequently updated model catalogue and multiple deployment modes.
Official sources: Together pricing, Together serverless catalogue.