Lab Variations May Mislead Scientific AI Models

Pennsylvania State University

To turn abundant carbon dioxide into valuable fuel, researchers need a fast and efficient way of determining which catalysts work best over the longest time. Artificial intelligence (AI) models have the potential to help guide catalyst selection, but as with internet chatbots, AI models are only as good as the data put into them.

By convening four laboratories from across the nation to test an experimental carbon monoxide-producing catalyst, a key first step in turning carbon dioxide into fuels, scientists at SLAC National Accelerator Laboratory, alongside two researchers from Penn State, have demonstrated the importance of generating highly reproducible experimental data when building AI models for investigations in science. They published the results in Nature Catalysis.

"An AI model is only as good as the data used to create it, which will undoubtedly come from multiple sources," said Robert Rioux, Friedrich G. Helfferich Professor of Chemical Engineering at Penn State and co-author on the study. "This study represents the first round-robin study focused on heterogeneous catalysis that quantifies the uncertainty from studies done between labs, and the impact this uncertainty will have on future data-driven AI modeling."

Speeding up catalyst development - with AI

With a good AI model, researchers can enter conditions such as temperature, length of time of the reaction, and catalyst formulation, then run the simulation and see a prediction of how well the catalyst performs. They can then confirm the predictions with a few well-designed experiments, ultimately speeding up catalyst discovery and implementation at a global scale.

In addition to saving time and money, such models can also explore conditions that are difficult to achieve in the lab. Most lab catalysis studies can only look at short time periods - days - but catalyst deactivation occurs over the course of months or years due to buildup of impurities and repeated exposure to high temperatures.

AI models need large amounts of high-quality data for training. To generate the data, the four labs performed a set of round-robin experiments, in which multiple laboratories conduct the same tests to evaluate reproducibility using previously agreed upon protocols and the same rhodium-based catalyst.

Squaring the data from round-robin experiments

To the researchers' surprise, achieving the same results from four labs working independently was harder than anticipated. When they got together to share their results, they realized that they had a problem.

Each of the four research teams produced results that varied in the amounts of carbon monoxide and methane, an undesirable side product, produced. The computer would not be able to learn from four sets of data that contain different outcomes.

"It was a bit of an eye-opener," said SLAC staff scientist Adam Hoffman, senior author of the study. "This experience shines light on the practical challenges of including real-world data into machine learning models."

Painstakingly, the teams evaluated their methods. Through rigorous testing, they found a handful of sources of mismatch, with one of the biggest contributors to the variability coming down to how hard the mixture was shaken or stirred.

"Our findings are a reminder to exercise caution about what information we feed into a machine-learning model, and how the consistency of experimental data can influence the reliability of the outcomes," said Selin Bac, a postdoctoral researcher at the University of California, Santa Barbara, and first author on the study.

With further standardization across the four labs - which in addition to SLAC included groups at Penn State, Stanford University and University of California, Santa Barbara - the results began to look more consistent. The team outlined several recommendations to strengthen experimental reproducibility, including enhancing the consistency of reactor design, operating protocols and experimental conditions.

"We contributed to the characterization of the catalytic materials used in the round-robin study," Rioux said. "Our findings reinforce the idea that high-quality experimental data dictates the success of any AI-based model by determining model accuracy, reliability and scope."

Hoffman said he hopes the study will help experimentalists and data scientists who are designing AI models to consider how small variations in experimental design across labs can lead to problems with reproducibility and impact on long-term predictions for AI modes.

"We see this work as a guide for the community as to how to think about designing experiments for inclusion in machine learning models," Hoffman said.

Greg Barber, assistant professor of chemistry at Penn State Altoona and affiliate researcher in the Institute of Energy and the Environment and co-author on the paper, also contributed to this research.

This work was supported in part by the U.S. Department of Energy's Office of Science under award number FWP 101064. Testing equipment was supplied in part by Co-ACCESS, part of the SUNCAT Center for Interface Science and Catalysis, a joint research center supported by SLAC National Accelerator Laboratory and Stanford University. This content is solely the responsibility of the authors and does not necessarily represent the views of the Department of Energy.

At Penn State, researchers are solving real problems that impact the health, safety and quality of life of people across the commonwealth, the nation and around the world.

For decades, federal support for research has fueled innovation that makes our country safer, our industries more competitive and our economy stronger. Recent federal funding cuts threaten this progress.

Learn more about the implications of federal funding cuts to our future at Research or Regress.

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.