By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Cohere Releases Embed 5 With Dual Tiers for Search
Cohere has released its new embedding model family, Embed 5, designed for enterprise search, retrieval-augmented generation (RAG), and agentic retrieval applications. The model family is offered in two distinct tiers: Embed 5 Pro, which prioritizes maximum retrieval quality, and Embed 5 Fast, which is optimized for latency and cost efficiency during live query operations. A key design innovation is that both Pro and Fast tiers share a unified embedding space, allowing users to index data using one tier and query it using the other, providing significant deployment flexibility. Both tiers are capable of processing text, images, and fused text-plus-image inputs, supporting over 100 languages and handling contexts of up to 128,000 tokens. Embed 5 is available today through the Cohere API and on platforms such as Microsoft Foundry and Amazon SageMaker, with options for private VPC or on-premise serving via vLLM. The API model identifiers are embed-v5.0-pro and embed-v5.0-fast, as detailed in Cohere's model documentation. These models can output embeddings in various dimensions, including 2048 (default), 1536, 1024, 768, 512, and 256, with embeddings returned as float, int8, or binary formats. Pricing for Embed 5 Pro is set at $0.12 per 1 million text tokens, while Embed 5 Fast costs $0.08 per 1 million text tokens. Image inputs are priced at $0.40 per 1 million tokens for both tiers. The models can embed page images directly or fuse an image with its associated metadata into a single vector, a feature particularly beneficial for scanned documents, slide decks, schematics, and charts where text extraction alone may lose critical information. Cohere conducted extensive testing across 40 development datasets, evaluating various corpus and query pairings. In these tests, an index created with Embed 5 Pro and queried with Embed 5 Fast achieved a score of 98.4, normalized against a Pro-plus-Pro setup at 100. An all-Fast configuration scored 96.6. Cohere recommends a pattern of indexing with the Pro tier for optimal quality and querying with the Fast tier for speed, provided both sides utilize the same output dimension. This split is particularly advantageous for agentic workloads, where an agent might perform dozens of searches per task, making query latency a critical factor. Cohere's internal benchmarks indicate that the Fast tier processed 377.3 documents per second, compared to 159.7 documents per second for the Pro tier. On the ViDoRe V3 benchmark, Embed 5 Pro achieved an average score of 85.8, representing an 8.8-point improvement over its predecessor, Embed 4. Embed 5 Fast averaged 84.5 on the same benchmark. For comparative context, Voyage 4 Large scored 83.7, Gemini Embedding 2 scored 83.2, and OpenAI's text-embedding-3-large scored 83.0.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.