OFICIAL Databricks Newsroom

Biomedical Imaging's Real Bottleneck Is the Data, Not the Model

What happened
Based on Databricks Newsroom · Oct 08, 2026

Medical imaging AI faces a data bottleneck rather than model limitations, as data remains siloed in PACS and vendor archives, complicating sharing and analysis.

Biomedical Imaging's Real Bottleneck Is the Data, Not the Model
Databricks Newsroom — Databricks
Key points
·
Radiology accounts for 76% of FDA-authorized AI/ML medical devices, most of which are narrow and human-supervised tools.
·
The EXAM study demonstrated a 16% AUC improvement and 38% generalizability gain using federated learning across 20 institutions without sharing patient data.
·
DICOM parsing libraries are single-core, but distributed techniques like zipdcm can catalog over 107,000 DICOMs in 3.5 minutes on two 8-core workers.
Key numbers
·
Hospitals generate vast imaging data, yet most AI tools remain narrow and supervised, with models accounting for only about 10% of the challenge.
·
However, reproducibility remains a major credibility problem, with a 2020 BMJ review finding that 93% of deep-learning imaging studies lacked available code and 95% lacked data.
·
Federated learning offers a partial solution, as demonstrated by the EXAM study, where 20 institutions improved model performance by 16% in AUC and 38% in generalizability without sharing patient data.

Hospitals generate vast imaging data, yet most AI tools remain narrow and supervised, with models accounting for only about 10% of the challenge. The primary obstacle lies in accessing and processing data locked in PACS and vendor-neutral archives, which are designed for viewing rather than research. Extracting, de-identifying, and computing with this data is labor-intensive, often requiring manual effort to navigate DICOM headers and pixel-level PHI. This structural issue underscores why collaboration across institutions is essential but difficult to achieve.

Academic medical centers drive open science in imaging AI, contributing to shared datasets like MIMIC-CXR and The Cancer Imaging Archive. However, reproducibility remains a major credibility problem, with a 2020 BMJ review finding that 93% of deep-learning imaging studies lacked available code and 95% lacked data. The difficulty in sharing and re-running studies stems from the complexity of imaging data formats and governance requirements. Federated learning offers a partial solution, as demonstrated by the EXAM study, where 20 institutions improved model performance by 16% in AUC and 38% in generalizability without sharing patient data.

Medtech companies integrate AI directly into scanners to address challenges like dose reduction and triage, but generalization across vendors and protocols remains a hurdle. Pharma and biotech rely on standardized imaging criteria like RECIST and QIBA profiles to ensure consistency in multi-site trials. Without uniform acquisition and analysis standards, variability undermines trial endpoints and biomarker reliability. The EXAM study’s federated approach highlights how collaborative modeling can overcome these barriers while preserving data privacy.

The broader challenge extends beyond imaging to multi-omics and real-world data, where siloed formats and governance issues persist. A Bain analysis found that 41% of companies cite data access and integration as the biggest barrier to AI adoption. The solution lies in consolidating governance, making imaging data queryable, and linking it to other patient data. Platforms like Databricks adapt existing medallion patterns to handle binary imaging files, using distributed parsing techniques to improve throughput and efficiency.

Original source → Deals on Clipraptor.com →