Introducing ConTextual: How well can your Multimodal model jointly reason over text and image in text-rich scenes?
A new benchmark just exposed a blind spot in every major multimodal model — including GPT-4V and Claude.

Why it matters
ConTextual is a new academic benchmark that measures how well multimodal models reason over text and images in real-world, text-rich scenes. This matters because it reveals capability gaps that standard benchmarks miss, forcing model builders to optimize for a new dimension of performance.
The key facts
5 to knowConTextual benchmark measures joint text-image reasoning in text-rich scenes
Published March 5, 2024 on Hugging Face
Tests multimodal models including GPT-4V and Claude
Addresses gap in existing multimodal evaluation standards
New capability dimension for model comparison and ranking
Go to the source
Hugging Face Bloghuggingface.co