Cloud-Based Integration of Vitis Vinifera Phenolic and Cohort Data Questions Beverage-Type Polyphenol Proxies

Main Article Content

Deepak Yadav, Savita Kumari Sheoran

Abstract

Background. Research on plant-derived bioactives routinely substitutes the food or beverage that carries a compound for the compound itself, and the heterogeneous data needed to test that substitution are rarely brought into a single queryable structure.
Methods. Two openly available sources were integrated in a containerised, cloud-executed pipeline: a phenolic analysis of 178 wine samples from three Vitis vinifera cultivars, and the NHANES I Epidemiologic Follow-up Study cohort of 1,629 adults. A six-table star schema was loaded into a row store and a columnar store, and query latency, data-volume scaling and parallel execution were benchmarked. Composition was compared by analysis of variance and principal component analysis; 10-year all-cause mortality was modelled by sequentially adjusted logistic regression, four classifiers, feature ablation and E-value analysis.
Results. The columnar store was 34 times faster than the row store at 1.63 million fact rows, and the single-threaded analytical workload reached 79% parallel efficiency on two cores. Cultivar explained 24–73% of variance in phenolic constituents, with a 3.8-fold difference in flavanoid content. Wine drinkers had lower crude mortality than beer drinkers (9.8% versus 18.3%; odds ratio 0.43, 95% CI 0.24–0.77), but adjustment removed 58% of the association (0.70, 0.35–1.39; E-value 1.68) and removing beverage type changed discrimination by under 0.001 AUC.
Conclusion. Beverage type is a poor surrogate for polyphenol dose. Natural-product epidemiology should quantify the bioactive, and the integration infrastructure to do so is inexpensive.

Article Details

Section
Articles