Agentic AI Premiums Don't Ensure Fastest Service

University of Michigan

Study: Pricing Delayed Agentic AI Services: When High-Value Jobs Wait Longer

It's conventional wisdom in the business world: Premium customers get a pass to the front of the line. Think airports or amusement parks.

But new research from the University of Michigan finds the best approach-at least in the realm of data processing services-may be placing those big spenders in the "better-but-later" queue and completing the smaller jobs first.

It's especially relevant in the realm of artificial intelligence, where data center demand driven in part by AI workloads is straining power grids and creating new capacity bottlenecks.

The challenge comes in how the providers of those services can structure their menu of options so those premium customers see the value in getting more even if they wait longer for it.

Mojtaba Abdolmaleki
Mojtaba Abdolmaleki

The study focused on agentic AI, though the research can apply more broadly to services in which providers can spend more processing effort to improve output quality-and that effort consumes scarce shared capacity.

Their focus, they say, speaks to a timely managerial question: How should firms design service menus when better answers take more time to produce?

The researchers from U-M's Ross School of Business found in their numerical analysis that completing smaller jobs first increases profit by 2.5% relative to first-come, first-served and by nearly 5.5% relative to conventional, high-value-first scheduling.

Izak Duenyas
Izak Duenyas

"When premium AI performs substantially more work, putting it first can make many smaller jobs wait," said Mojtaba Abdolmaleki, doctoral candidate at the Ross School and co-author of the study with business professors Izak Duenyas and Roman Kapuscinski. "The premium customer can still receive the more thorough result without automatically being first in line.

"Premium AI may need a quality lane, not only a fast lane."

Roman Kapuscinski
Roman Kapuscinski

Abdolmaleki and his co-authors offer the following example from the world of tax preparation, already a real business workflow:

A graduate student has one wage or fellowship statement, no business income and only a few supporting documents. The student selects basic preparation, and the resulting AI job occupies the compute pipeline for one minute.

A business owner has multiple entities, dozens of tax forms, income in several states, investment transactions and foreign accounts. That customer selects comprehensive preparation, which occupies the pipeline for two hours.

Now, if the longer job runs first, the shorter one waits two hours, and if the shorter one runs first, the longer one is delayed by only one additional minute. Completing the one-minute return first reduces total waiting cost by about 58%, even though the business owner is 50 times more costly to delay.

The premium job isn't delayed because the return is less important; rather, the customer selected that service because the return warrants substantially more analysis. So the service providers must clearly define quality-first service: more analysis, a better expected result and a disclosed completion window.

"That doesn't mean secretly degrading the service received by a premium customer," Abdolmaleki said.

The study is under review for publication.

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.