Ensuring Scientific Dataset Quality

University of Vienna

A research team at the University of Vienna, led by pharmaceutical chemist Johannes Kirchmair, has developed a new statistical method that can be used to identify hidden anomalies in scientific datasets. The method can be applied automatically to large datasets. This makes it easier for researchers to identify problematic data sets in a targeted manner. The new approach significantly improves the quality of data-driven research and promotes the development of reliable AI applications.

Modern scientific methods increasingly rely on large volumes of data. In modern drug discovery, for example, millions of experimental measurements are often collated from databases and fed into AI models to aid the discovery and development of new medicines. The quality of this data is crucial: errors in the collection, processing or integration of data can not only slow down subsequent analyses, findings and research processes, but also render them futile. However, the systematic verification of large datasets poses considerable challenges for researchers.

Uday Abu-Shehab, Matthias Welsch and Johannes Kirchmair from the Christian Doppler Laboratory for Molecular Informatics in the Biosciences at the Department of Pharmaceutical Sciences, University of Vienna, have therefore developed a method capable of automatically identifying anomalous data sets. At the heart of this is what is known as Benford's Law. This law describes the surprising observation that, in many naturally occurring data sets, certain digits occur more frequently as the first digit than others. If data sets deviate substantially from this pattern, this may indicate irregularities or potential quality issues.

New method as a filter for datasets

"Our method does not provide direct evidence that data is flawed or falsified," explains Johannes Kirchmair. "Rather, it helps to filter out datasets from large collections that should be subjected to closer scrutiny. This enables researchers to focus their attention specifically on potentially problematic datasets."

To enable the prioritisation of datasets, the researchers developed the "Simulation-Based Benfordness Estimation" (SBBE) method. The approach combines extensive computer simulations with Bayesian statistics and provides not only an assessment of data quality but also an estimate of the associated uncertainty. This makes it possible to compare datasets much more effectively than with previous methods.

"With many existing approaches, the results of quality analyses depend heavily on the number of available data points," explains Uday Abu-Shehab. "Our method explicitly takes this uncertainty into account, thereby enabling fair comparisons between datasets of different sizes. This makes the results substantially more robust."

Test run with bioactivity databases

To assess the effectiveness of the approach, the researchers tested it using bioactivity data. Such data describe the strength of interaction between bioactive compounds and their target proteins and are essential in modern drug discovery. The analysis showed that large bioactivity databases generally exhibit high data quality. The new method revealed that certain data sets exhibit unusual statistical characteristics. The researchers were ultimately able to attribute these to particularities in the experimental design, data processing or data curation.

New method ensures the quality of scientific data collections

"Tools for automated quality control are becoming increasingly important, particularly for data-driven research," says Matthias Welsch. "Our aim is to provide researchers with an easy-to-use tool to identify anomalies at an early stage and further improve the quality of scientific data collections."

As SBBE is not limited to any specific field of research, the authors see potential for numerous further applications. Wherever large volumes of numerical data are generated, the method could help to reveal hidden irregularities and strengthen the basis for well-founded scientific decisions as well as applications in the field of AI.

Summary:

  • In science, large datasets are often used, for example, in drug discovery. However, systematically analysing large datasets is challenging.
  • Scientists at the University of Vienna have developed a new method to easily identify hidden anomalies in such data.
  • "Simulation-Based Benfordness Estimation" (SBBE) helps to filter out potential sources of error from datasets and can be applied across various fields of research.
  • The new method significantly improves the quality of data-driven research and promotes the development of reliable AI applications.

About the University of Vienna:

For over 650 years the University of Vienna has stood for education, research and innovation. Today, it is ranked among the top 100 and thus the top four per cent of all universities worldwide and is globally connected. With degree programmes covering 188 disciplines, and approximately 11,000 employees, we are one of the largest academic institutions in Europe. Here, people from a broad spectrum of disciplines come together to carry out research at the highest level and develop solutions for current and future challenges. Its students and graduates develop reflected and sustainable solutions to complex challenges using innovative spirit and curiosity.

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.