Benchmarking the Impact of Data Quality Dimensions on Predictive Analytics Systems Across Healthcare and Life Science Workflows

Main Article Content

Dharmateja Priyadarshi Uddandarao, Rohit Nagpal, Alekhya Challa, Ankit Pushpam, Samridhi Vats, Naina Gupta

Abstract

From hospital readmission risk to patient stratification based on biomarkers, predictive analytics systems have become integral to the healthcare and life-science value chain. These systems are traditionally judged by their predictive accuracy with the tacit assumption that their input data is clean, complete and stable. In practice, operational data pipelines have measurable defects in several dimensions of data quality - incompleteness, inconsistency, inaccuracy and staleness - but none of these dimensions is standardized with a method of measuring how sensitive a predictive system is to it. Consequently, two models that perform the same accuracy on the benchmark data could fail very differently when applied to imperfect data and this fragility would not be apparent until the model fails to make accurate predictions. This paper introduces a dimension-resolved benchmarking framework which treats data quality as an experimental factor to be controlled by the experimenter instead of fixed background conditions. This analysis (1) combines the existing dimensions of data quality into an operational taxonomy applicable to predictive pipelines; (2) provides a library of controlled perturbation operators to add graded, reproducible degradation to each dimension; (3) introduces a set of sensitivity statistics, including dose–response degradation curves, an Area-Under-the-Degradation-Curve (AUDC) robustness index, calibration drift, and a version-stability coefficient of variation, that summarize the fragility of a system per dimension; and (4) defines a Data-Quality Robustness Scorecard for standardized reporting. This analysis uses it to a public healthcare readmission dataset to empirically show that models that have comparable clean-data accuracy can have significantly different per-dimension fragility profiles, and extends its use to intensive-care and omics workflows. The framework is independent of the data set and independent of the model, allowing for reproducible, comparable robustness benchmarking and data-quality governance for predictive analytics in healthcare and in the life sciences.

Article Details

Section
Articles