New W² Metric to Boost AI Model Evaluation in Health

Shenyang Agricultural University Collaborative Journals

Artificial intelligence is increasingly used to estimate pollutant concentrations, human exposure, and disease risks. However, researchers warn that commonly reported performance measures may give an incomplete picture of whether these predictions are truly reliable.

In a new viewpoint published in Artificial Intelligence & Environment, researchers clarify the difference between the coefficient of determination, known as R², and the predictive squared correlation coefficient, Q². They also introduce W², a complementary metric designed to evaluate both predictive accuracy and systematic prediction bias.

"A model can show a strong statistical correlation while still consistently overestimating low values or underestimating high values," said corresponding author Bin Wang of Peking University. "W² provides an additional way to identify these hidden prediction patterns and assess whether a model reproduces the actual scale of environmental exposure."

R² is traditionally used to describe how well a regression model fits the data used to build it. For predictions involving new and unseen data, the researchers argue that Q² is more appropriate. Yet even a high Q² value may not reveal whether predictions systematically deviate from observed values.

To address this limitation, W² combines Q² with a penalty based on the angular difference between the model's prediction line and the ideal 1:1 line. Predictions that closely match observed values receive a higher W² score, while models showing consistent overprediction or underprediction are penalized.

The researchers demonstrated the metric using human blood exposome data and six representative modeling approaches. Although several models achieved relatively high Q² values, W² revealed important differences in their calibration and quantitative reliability.

The authors emphasize that W² should complement, rather than replace, established indicators such as root mean squared error, mean absolute error, cross-validation, and external validation. Together, these measures may support more transparent and rigorous evaluation of AI models used in environmental health research.

===

Journal reference: Wang B; Xia F; Pan B; et al. From R2 to W2: rethinking predictive metrics for AI models in environmental health. AI Environ. 2026, 1(3): xx-xx. DOI: 10.66178/aie-0026-0015

https://www.the-newpress.com/aie/article/doi/10.66178/aie-0026-0015

===

About the Journal:

Artificial Intelligence & Environment is an international multidisciplinary platform for communicating advances in fundamental and applied research on the intersection of environmental science and artificial intelligence (AI). It is dedicated to serving as an innovative, efficient and professional platform for researchers in the cross-discipline fields of earth and environmental sciences, big data science and AI around the world to deliver findings from this rapidly expanding field of science. It is a peer-reviewed, open-access journal that publishes critical review, original research, rapid communication, view-point, commentary and perspective papers.

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.