Do Quality-Assurance Scores Predict Competitive Rankings Fairly? A Cross-Sectional Audit of NAAC and NIRF in Indian Universities
Main Article Content
Abstract
India’s two national instruments for evaluating higher-education quality, the National Assessment and Accreditation Council (NAAC) and the National Institutional Ranking Framework (NIRF), score institutions on largely different logics: NAAC audits governance and process on a seven-criterion, four-point scale, while NIRF ranks institutions on five outcome-oriented parameters scored out of 100. Whether the two systems agree, and whether any statistical model built to forecast one from the other treats different categories of universities equally, has not been examined empirically. Using a matched cross-sectional sample of 132 Indian universities linking their most recent NAAC accreditation cycle (2018–2025) to the nearest subsequent NIRF ranking cycle (2016–2025), we (a) test whether NAAC’s seven criterion-level scores predict NIRF outcomes, and (b) audit whether prediction errors from a gradient-boosted model differ systematically across Central, State, and Private/Deemed universities. NAAC’s overall CGPA explained only 16.1% of the variance in NIRF’s overall score (R² = .161), and the criterion most theoretically aligned with NIRF’s research parameter (Criterion 3: Research and Innovation) correlated with it only weakly (r = .297), well below the a priori benchmark of r > .60. A cross-validated gradient boosting model modestly outperformed linear specifications (R² = .138 on 5-fold cross-validation) and attributed most of its predictive power to teaching-learning and research criteria. Contrary to the hypothesis that state universities would be disadvantaged by such a model, Central universities showed the largest mean absolute error (7.64 points vs. 5.94 for State and 6.08 for Private/Deemed universities), though the difference was not statistically significant (Mann-Whitney U, p = .726) and the bootstrapped 95% confidence interval for the disparate impact ratio was wide and crossed 1 ([0.12, 1.61]), reflecting the small Central-university subsample (n = 19). We conclude that NAAC and NIRF measure substantially non-overlapping constructs, that any predictive substitution of one for the other would be unreliable and potentially unfair to a specific institutional category depending on model choice, and that larger, better-linked administrative data are needed before either instrument is used to forecast the other. All analyses were conducted in Python.
