HeadlinesBriefing favicon HeadlinesBriefing.com

Voyage Multimodal 3.5 Adds Video Retrieval

Hacker News: Front Page •
×

Voyage AI has released voyage-multimodal-3.5, its next-generation embedding model for retrieval across text, images, and now videos. Building on its predecessor, the new model uses a unified transformer architecture to process interleaved content in a shared vector space, avoiding the modality gap found in CLIP-based systems.

The model demonstrates improved performance, outperforming Cohere Embed v4 by 4.56% on visual documents and Google's Multimodal Embeddings by 4.65% on video datasets. It also matches state-of-the-art text embedding models while supporting Matryoshka embeddings for flexible dimensionality and multiple quantization options, reducing costs for production use.

This release targets developers building search pipelines for complex documents and video libraries. By providing a single model for multimodal retrieval, Voyage AI aims to simplify infrastructure and improve accuracy for applications like document analysis and video content search, competing directly with offerings from Google and Cohere.