Two tiers built around one vector space

Cohere launched Embed 5 on 30 September as a pair of embedding models for search, retrieval-augmented generation and agent workflows. Embed 5 Pro is aimed at maximum retrieval quality, while Embed 5 Fast targets interactive and high-volume query paths. Both accept text, images and fused text-image inputs, cover more than 100 languages and support a 128,000-token context window.

The distinctive deployment feature is a shared embedding space. Cohere says organisations can index documents with Pro and query the resulting vectors with either Pro or Fast, provided both sides use the same output dimension. That lets a team pay for the higher-quality model during offline ingestion and use the cheaper, faster tier on every live request without rebuilding its entire index.

Availability and published pricing

Embed 5 is generally available through the Cohere API and Model Vault, Microsoft Foundry and Amazon SageMaker. Cohere lists text pricing at US$0.12 per million tokens for Pro and US$0.08 for Fast, with image inputs priced at US$0.40 per million tokens for both. The models can also be privately deployed with vLLM, giving regulated organisations an alternative to shared API inference.

Fast retains the same context length, modalities, language coverage and output formats as Pro. Cohere reports that it processes an average of 2.4 times as many documents per second across the tested context sizes. That difference is especially relevant for agents that issue repeated searches, because embedding latency and cost accumulate every time a tool retrieves more context.

Multimodal and enterprise-document retrieval

The model family is designed for material where meaning is not confined to clean paragraphs. Cohere highlights tables, charts, diagrams, scanned pages, financial filings and technical manuals. Embed 5 can represent page images directly or combine an image with metadata in one vector, addressing cases where PDF text extraction destroys layout or removes the relationships between rows, columns and labels.

Cohere reports an average score of 85.8 for Pro on the ViDoRe V3 suite, compared with 77.0 for Embed 4, while Fast scores 84.5. It also publishes results for parsed PDFs, financial documents and multilingual retrieval. These are vendor-reported comparisons and some use Cohere’s new AI-judged metric, so customers should reproduce tests with their own corpora and failure costs before treating rankings as decisive.

Compression can change index economics

Embed 5 supports output dimensions from 256 to 2,048 and returns float, int8 or binary embeddings. It uses Matryoshka representation learning so vectors can be shortened while retaining useful structure. Cohere illustrates the storage impact with 100 million chunks: a 2,048-dimensional float32 index would require roughly 819GB for raw vectors, while 256-dimensional binary outputs would use about 3.2GB.

That reduction is not free. Smaller dimensions and lower precision can trade retrieval quality for memory and search speed. Cohere recommends 1,024-dimensional int8 vectors as a practical balance for many deployments, while suggesting binary output for first-pass retrieval followed by higher-precision reranking. Teams should test this pipeline end to end, because storage savings are valuable only if the retrieved context remains good enough for downstream answers.

Migration requires more than swapping a model name

An Embed 5 evaluation should include the documents that cause real search failures: long reports, duplicated policies, multilingual queries, image-heavy files and questions with several plausible answers. Teams should measure first-stage recall as well as reranking, inspect latency under concurrency and check whether cross-tier querying behaves consistently across departments and languages.

The shared space gives existing retrieval teams an unusually flexible cost lever, but moving from an older embedding family still requires re-indexing because the compatibility promise applies within Embed 5, not across generations. A controlled migration can build a new index in parallel, replay production queries and compare the effect on answer citations and user outcomes. Embed 5 offers strong technical options; evidence from the organisation’s own data should decide how they are used.

Operational compatibility deserves its own checklist. Vector databases must support the chosen dimension and data type, and applications need consistent normalisation and input-type settings for documents and queries. Multimodal ingestion should retain enough provenance to link a vector back to the page image, parsed text and original file. Teams should also budget for the temporary storage and compute needed to maintain two indexes during migration. Once cutover is complete, monitoring should track changes in empty-result rates, click-through, cited passages and downstream answer quality. These measures reveal whether a benchmark improvement translates into a better search experience rather than merely a different vector representation.

Because Pro and Fast can be mixed, configuration should be visible in telemetry rather than buried in deployment code. That lets operators relate quality or latency shifts to the tier actually used for each stage and reproduce unexpected results accurately over time.