Scientists Challenge No Difference Conclusion

Researchers from the Universities of Manchester, Oxford and Arkansas are urging scientists to stop treating "non-significant" results as proof that nothing happened.

The common statistical mistake, they warn in a new paper in PNAS on 10 08 26 - one of the world's leading scientific journals - could be leading researchers to draw the wrong conclusions from their data.

A p-value, which measures how surprising the observed data would be if there were truly no effect, of greater than 0.05 does not, they say, show there is 'no effect'.

However, the interpretation remains widespread in about 50% of research papers and conference presentations, according to sources.

Instead, the authors say, a non-significant result simply means there is insufficient evidence to conclude a difference exists, not proof that a difference does not exist.

Crucially, the same statistical result can arise either because there is genuinely no meaningful effect or because an important effect is hidden by small sample sizes or highly variable data.

The researchers argue that failing to recognise this distinction risks oversimplifying scientific findings and may cause potentially important effects to be overlooked.

With many studies not having a large-enough sample size to detect the real difference, reporting that there is none can miss promising therapies or fail to detect genuine risk, derailing subsequent research in the area.

Equivalence testing, they say, guards against declaring an effect unimportant when the data cannot support that claim. A small or noisy study will usually return an inconclusive result rather than a verdict of no difference, which is the honest answer.

To address the problem, the team is promoting a statistical approach known as equivalence testing.

Rather than asking whether there is evidence for a difference, equivalence testing asks whether any difference that exists is too small to matter in scientific, clinical or practical terms.

The method allows researchers to distinguish between effects that are genuinely negligible and results that remain inconclusive because there is not enough reliable evidence.

They focus on the two one-sided tests procedure, or TOST, which has already gained traction in psychology, medicine and pharmaceutical regulation but remains underused across many areas of the life and natural sciences.

Wider adoption of equivalence testing, they add, could improve the quality of scientific reporting and help prevent non-significant findings from being misrepresented as evidence of no effect.

Co-author, David Eisner, Professor of Cardiac Physiology from The University of Manchester, said: "Researchers are often interested in whether an effect is absent or too small to be important, but traditional statistical testing cannot answer that question.

"A non-significant result is frequently interpreted as proof that nothing happened, when it may simply mean there is not enough evidence to be certain."

Co-author Jakub Tomek from the University of Oxford said: "Equivalence testing helps separate genuinely trivial effects from unresolved questions, giving scientists a much clearer picture of what their data are actually telling them and improving confidence in the conclusions that are reported.

"We hope the method will encourage more careful interpretation of scientific data and improve the way research findings are reported, understood and acted upon."

To make the approach more accessible, the team has also developed a free online calculator that allows researchers to perform common equivalence tests without writing computer code.

The calculator supports a range of widely used statistical comparisons and has been validated against established statistical software.

Co-author Aaron Caldwell from the University of Arkansas for Medical Sciences added: "Equivalence testing forces you to answer a question most studies never ask: how small is small enough to be uninteresting?

""That is a scientific judgement, not a statistical one, and it has to be made before the data are collected. The payoff is that you can finally say something affirmative about an absence of effect rather than just failing to find one."

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.