Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers
Sentence Transformers just made multimodal embeddings trainable. Here's why that matters for your search stack.

Why it matters
Open-source tooling for training multimodal embeddings democratizes reranker capability building. Companies no longer need proprietary models to optimize search and retrieval—Sentence Transformers now enables fine-tuning on custom data, shifting competitive advantage from model access to training methodology.
The key facts
9 to knowSentence Transformers adds native multimodal embedding training
Supports both text and image inputs in single embedding space
Enables fine-tuning on custom datasets without proprietary infrastructure
Published on Hugging Face as open-source resource
Applicable to reranking, semantic search, and retrieval workflows
Sentence Transformers framework now supports multimodal embedding training
Enables fine-tuning of reranker models on custom datasets
Published on Hugging Face blog — indicating broad distribution/accessibility push
Targets enterprise adoption of open-source alternatives to commercial embedding APIs
Go to the source
Hugging Face Bloghuggingface.co