What is Together AI?

Together AI provides serverless and dedicated inference, fine-tuning, GPU infrastructure, code execution and model deployment for developers.

Core offerings

  • Usage-priced serverless inference across text, image, video and audio.
  • Dedicated model endpoints and GPU clusters.
  • Fine-tuning and custom-model deployment.
  • Embeddings, moderation and code execution services.

Pricing

Together AI uses prepaid usage billing and requires a minimum US$5 credit purchase; there is no general free trial. Serverless models are priced per token, image, video or audio unit, while dedicated deployments are billed for provisioned hardware. Enterprise and high-volume arrangements can include custom capacity and discounts.

API and model availability

The live serverless catalogue currently documents 103 distinct non-retired entries: 25 chat, 30 image, 37 video, nine audio, one embedding and one moderation model. They are reproduced on the API Cost page from Together's official catalogue; rerank models are excluded because Together states none are currently offered serverlessly.

Why select Together AI?

Together AI is best suited to engineering teams that value a broad, frequently updated model catalogue and multiple deployment modes.

Official sources: Together pricing, Together serverless catalogue.