A Multimodal Whole-Slide Foundation Model for Pathology (TITAN)
Pretrained on 335,645 whole-slide images with vision-language alignment, TITAN generates slide-level pathology reports and performs rare cancer retrieval — without fine-tuning — outperforming all prior slide-level foundation models.
TITAN was pretrained on 335,645 whole-slide images using a combination of visual self-supervised learning and vision-language alignment — the latter powered by 423,122 synthetic captions generated from paired pathology reports. The dual training objective gives the model both strong morphological representations and the ability to reason about histological content in natural language.
In zero-shot evaluation, TITAN generates coherent slide-level pathology reports and retrieves visually similar cases for rare cancer types — tasks that prior region-of-interest models could not perform at the slide level. The model outperformed previous slide-level foundation models across classification, survival prediction, and retrieval benchmarks.
The work represents a key step toward slide-level pathology AI that pathologists can interrogate in natural language, rather than only receiving a binary classification output.
Read in Nature Medicine →