Why Indian datasets matter

Western validation doesn't transfer automatically

Prov-GigaPath and comparable pathology foundation models are trained predominantly on Western cohorts. Several factors mean that performance doesn't automatically generalize to Indian clinical settings.

Population-specific biology

Tumor genomics, mutation prevalence and disease presentation in Indian cancer cohorts differ meaningfully from Western datasets.

Lab & scanner variability

Staining protocols, tissue processing and digital scanner hardware vary widely across Indian pathology labs.

Access beyond metros

Genomic testing is concentrated in a few metropolitan labs, making independent validation the prerequisite for wider access.

Current collaborations

Hospital and research partners in Phase 1

Retrospective cohorts and slide digitization are being coordinated with the following categories of partner — full details on the Collaborations page.

Causal ITResearch & AI collaboration
--Clinical collaboration
SRMIST BioNESTIncubation partner
Future multicentre expansion

From one dataset to a national validation network

If Phase 1 results support the hypothesis, the natural next step is a prospective, blinded study spanning multiple hospitals and geographies — covering a broader range of scanners, staining protocols and patient populations.

That expansion is deliberately sequenced after, not alongside, Phase 1 — we want the retrospective feasibility signal before asking partner hospitals for a heavier prospective commitment.

Funnel diagram showing the validation pipeline stages: hospitals, slides, AI validation, and benchmarking