Today's Clicks Can Reveal Tomorrow's Breakthroughs

The trail of likes, shares and downloads we leave across the internet could help predict successful innovations years in advance, Cornell researchers showed by curating two datasets that allowed them to compare early engagement to future impact.

Researchers from the Cornell Ann S. Bowers College of Computing and Information Science published new datasets from arXiv, an online repository of unreviewed scientific papers, and GitHub, a platform where developers store and share their code. This data can be used to test new methods for "lead-lag forecasting" - the use of early likes, views and downloads to forecast which contributions will be important, even years later.

"In the modern digital world, every platform is gathering this type of data. What are people clicking on? What are they doing?" said senior author Sarah Dean, assistant professor of computer science in Cornell Bowers. "It's extremely powerful for predicting all sorts of things."

Yangfanyu Yang, a doctoral student in the field of information science, will present the new study, "Benchmark Datasets for Lead-Lag Forecasting on Social Platforms," at the Association for Computing Machinery Knowledge Discovery in Data conference, held Aug. 9-13 in Jeju, Korea.

The project began as discussions over how to forecast emerging fields and scientific breakthroughs more quickly, which can yield incredibly valuable advantages.

"During COVID, economists estimated that accelerating solutions by even a single day was worth tens of billions of dollars," said co-author Yian Yin, assistant professor of information science in Cornell Bowers. "Those who can anticipate where the frontier moves next - even just a few days ahead - stand to capture competitive advantages, whether for firms, funders or nations."

Currently, a paper's impact is measured by how many times other researchers cited it in their own work. But years can pass between reading a new paper and citing it in a publication.

"We had this very simple, yet powerful idea," Yin said. "What if instead of looking at what papers people are citing, we simply look at how people are reading papers?"

It works like a prediction market, Yin explained, except instead of money, scientists are betting their time. Reading a new piece of work is a costly commitment, and aggregated across thousands of experts, those commitments offer a view of the future scientific and technological landscape.

Working with arXiv, the researchers accessed anonymized data for the number of downloads for 2.3 million papers uploaded to the site. Then they checked how the number of early downloads related to citations in the scientific literature five years later.

Using established machine learning models, the researchers showed they could predict popularity at five years with as little as a month of data. "Most papers get read one or two days after they get published, but the citation takes three years or even longer, which means our prediction is way more useful," Yin said.

This finding not only demonstrates that early downloads are a good indicator of valuable research, but also validates the dataset as a useful tool for developing more effective forecasting models.

The team also created a similar dataset by compiling GitHub data from about 938,000 repositories of code. This data showed a similar pattern: early pushes (uploading personal code changes) and stars (saves) translated to greater numbers of forks (when programs build on earlier code) five years later.

Previously, people had developed prediction methods for rolling forecasting tasks - such as estimating upcoming energy usage or traffic levels - but these methods predict just hours or days into the future.

In related work, the team is interested in whether early downloads can indicate paper quality on arXiv. With the advent of AI, arXiv has experienced a flood of low-quality papers, and weeding out the AI slop is challenging.

They may also be able to forecast emerging fields by looking at whether scientists read pairs or groups of papers concurrently.

Ultimately, the team hopes these datasets will enable others to advance more effective models for lead-lag forecasting.

"Methods that work well in this setting would impact lots of other domains, like social media, online sales and investing," Dean said.

Kimia Kazemian and Zhenzhen Liu, both Ph.D. students in the field of computer science, are co-first authors on the paper. Katie Luo, Ph.D. '25, Sherman Gu, M.S. '25, Moyun Du '26, Xinyu Yang, a doctoral student in the field of information science, Jack Jansons '25, Kilian Weinberger, professor of computer science, and John Thickstun, assistant professor of computer science, also contributed to the research.

The work received support from the National Science Foundation, NASA, a PCCW Affinito-Stewart Award, the LinkedIn-Cornell Bowers Strategic Partnership, an AI2050 Early Career Fellowship from Schmidt Sciences and NewYork-Presbyterian for the NYP-Cornell Cardiovascular AI Collaboration.

Patricia Waldron is a writer for the Cornell Ann S. Bowers College of Computing and Information Science.

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.