PRISM2 model uses clinical dialogue to interpret pathology slides
Paige and Microsoft's PRISM2 reads pathology slides like a diagnostic radiologist—not by classifying pixels, but by reasoning through clinical dialogue.

Why it matters
A new multimodal model architecture (perceiver-based encoder, dialogue-grounded training) demonstrates how foundation models can be adapted for specialized medical imaging workflows. This is a capability advance in vision-language reasoning for high-stakes domains, not just a product feature.
The key facts
7 to knowPRISM2 built by Paige and Microsoft
Perceiver-based encoder architecture
Trained jointly on tissue tiles and clinical dialogue from pathology reports
Aggregates thousands of tile embeddings per slide into single representation
Generates diagnostic text answers rather than pixel classification
Training data spans 2.3 million whole-slide images
Focus on reasoning over classification
Go to the source
AI Newsartificialintelligence-news.com
Publisher excerpt: Built by Paige and Microsoft, PRISM2 reads whole-slide images through a perceiver-based encoder trained jointly on tissue tiles and clinical dialogue drawn from pathology reports. The model aggregates thousands of tile embeddings per slide into one representation, then generates text that answers…