AI Recommendations and Incentives Under Time Constraints: Evidence from an RCT

Last registered on August 20, 2026

Pre-Trial

Trial Information

General Information

Title
AI Recommendations and Incentives Under Time Constraints: Evidence from an RCT
RCT ID
AEARCTR-0018623
Initial registration date
August 11, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
August 20, 2026, 8:31 AM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information

Primary Investigator

Affiliation
Clemson University

Other Primary Investigator(s)

Additional Trial Information

Status
Completed
Start date
2026-08-13
End date
2026-08-20
Secondary IDs
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
This paper studies how the size of the monetary penalty for inaccurate responses affects subjects' reliance on AI-generated recommendations. Subjects complete a series of CAPTCHA-style object-counting tasks, with and without access to a pre-generated recommendation from an AI system, permitting within-subject comparison of task accuracy and AI reliance across AI-available and AI-unavailable conditions; this contrast is implemented under one of three between-subjects penalty levels for incorrect responses. To help explain over- or under-reliance relative to the AI's true, undisclosed accuracy, we separately elicit subjects' subjective belief distribution over the AI's accuracy using an incentive-compatible quadratic scoring rule instrument (Matheson and Winkler, 1976), and their willingness to pay for AI access using a Becker-DeGroot-Marschak (1964) mechanism. Subjects' subjective beliefs about the AI's accuracy are compared to their own self-assessed and task-measured accuracy to provide further evidence on why subjects may over- or under-rely on AI recommendations (Moore and Healy, 2008). Subjects are recruited via Prolific and complete the experiment on Qualtrics.
External Link(s)

Registration Citation

Citation
Chupak, Maxwell. 2026. "AI Recommendations and Incentives Under Time Constraints: Evidence from an RCT." AEA RCT Registry. August 20. https://doi.org/10.1257/rct.18623-1.0
Sponsors & Partners

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information
Experimental Details

Interventions

Intervention(s)
Subjects complete three modules, each containing 30 unique judgment tasks. Subjects are randomly assigned to one of three financial incentive conditions that differ in the monetary penalty for incorrect responses. AI recommendations are available in exactly one of the first two modules and unavailable in the other; this availability is varied within-subject and counterbalanced across subjects. Access to AI recommendations in the third module is instead determined by a Becker-DeGroot-Marschak (BDM) mechanism eliciting each subject's willingness to pay. In every module, the first task is one for which the associated AI answer is fully accurate, and the second task is one for which the associated AI answer is off by one; the relative order of these two tasks is randomized independently at each module load, regardless of whether AI recommendations are available in that module. All remaining tasks are presented in a fixed, predetermined order. When an AI recommendation is available, subjects may reveal it and, with an additional click, incorporate it into their response, so that viewing and adopting a recommendation are captured as distinct decisions.
Intervention Start Date
2026-08-13
Intervention End Date
2026-08-20

Primary Outcomes

Primary Outcomes (end points)
For each subject, the primary outcomes are

1. The within-subject difference between AI and non-AI modules in number of tasks answered correctly, number of completed tasks, and percent of tasks correctly completed.

2. The subject’s willingness to pay for access to AI recommendations for the final module.

Behavioral measures of reliance (e.g., AI-recommendation reveal and adoption rates) are analyzed as secondary outcomes; in the absence of a structural model dictating a single theoretically-preferred reliance measure, we do not designate any one specification as primary.
Primary Outcomes (explanation)
Number of Tasks Completed is defined as the number of tasks completed within a given module.

Number of Tasks Answered Correctly is defined as the number of tasks answered correctly within a given module.

Percent of Tasks Answered Correctly is defined as the number of tasks answered correctly divided by the number of tasks attempted. The number of tasks answered correctly divided by the fixed number of questions per module is not included as a separate outcome, as it is a rescaling of Number of Tasks Answered Correctly.

Within-Subject AI Effect is computed, separately for each of the outcomes above, as the difference in a subject's outcome between their AI-available and non-AI module.

Willingness To Pay (WTP) is measured using the BDM elicitation and recorded as the maximum amount the subject is willing to pay for AI access in the third module.

Secondary Outcomes

Secondary Outcomes (end points)
Secondary outcomes include,
1. The mean of each subject’s elicited subjective belief distribution over the AI’s accuracy.
2. The deviation between the mean of each subject's elicited belief distribution and the AI’s true accuracy.
3. The deviation between the mean of each subject's expected own accuracy and the subject's true accuracy.
4. Behavioral measures of reliance (e.g., AI-recommendation reveal and adoption rates) are analyzed as secondary outcomes; in the absence of a structural model dictating a single theoretically-preferred reliance measure, we do not designate any one specification as primary.
Secondary Outcomes (explanation)
Elicited subjective belief distribution over the AI’s accuracy is obtained by asking participants to allocate a fixed budget of tokens across bins spanning the possible accuracy range of the provided AI recommendations.

Deviation of belief distribution’s mean from the AI’s true accuracy is calculated by taking the mean of each participant’s belief distribution over the AI accuracy and subtracting it from the true AI accuracy.

Deviation of participant’s beliefs over their own accuracy from revealed participant accuracy, calculated as the participant’s stated expected number of correct answers minus the participant’s realized number of correct answers, divided by the total number of questions
answered.

Behavioral measures include whether participants viewed an AI recommendation, whether they incorporated it into their submitted response, whether they changed their response after viewing the recommendation, whether they ultimately accepted or rejected the recommendation, and whether the submitted response matched the recommendation. Furthermore, the timing between behavioral measures acts as supportive evidence of who these actions relate to each other.

Experimental Design

Experimental Design
Participants complete three modules of incentivized recognition tasks under time constraints. AI recommendations are available in one of the first two modules, with AI availability varied within subjects and module order counterbalanced across participants. Participants are randomly assigned to one of three financial incentive conditions that differ in the penalty for incorrect responses. In the third module, access to AI recommendations is determined through a Becker-DeGroot-Marschak (BDM) mechanism eliciting willingness to pay. Participants also complete an incentivized elicitation of their beliefs regarding the AI’s accuracy. Earnings are determined by one randomly selected module.
Experimental Design Details
Not available
Randomization Method
Randomization of first task accuracy, module presentation order, BDM price, and assignment to penalty level is conducted by the Qualtrics survey software.
Randomization Unit
Penalty level is randomized at the subject level. First task accuracy for modules with AI recommendations and module order across the experiment are randomized within subject.
Was the treatment clustered?
No

Experiment Characteristics

Sample size: planned number of clusters
Not applicable. Randomization occurs at the individual participant level.
Sample size: planned number of observations
400 individual participants. Up to 36,000 task-level observations (400 participants × 90 tasks), conditional on completion of all tasks..
Sample size (or number of clusters) by treatment arms
Approximately 133 participants per penalty treatment arm.
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
Power calculations are based on the primary outcome of the within-subject difference in accuracy between AI and non-AI modules. Using pilot data, the standard deviation of the participant-level difference in accuracy between AI and non-AI conditions is 0.138. With a planned sample of 400 participants, a two-sided significance level of 5\%, and 80\% power, the minimum detectable effect size is approximately 0.019 accuracy points, equivalent to a 1.9 percentage point change in task accuracy.
Supporting Documents and Materials

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information
IRB

Institutional Review Boards (IRBs)

IRB Name
Clemson University Institutional Review Board
IRB Approval Date
2026-07-26
IRB Approval Number
IRB2026-0020