You need to sign in or sign up before continuing.
Back to History

Fields Changed

Registration

Field Before After
Trial Title AI Recommendations and Incentives Under Time Constraints: Evidence from an RCT Human Reliance on AI Recommendations Under Time Constraints and Varying Financial Incentives
Trial Status completed on_going
Abstract This paper studies how the size of the monetary penalty for inaccurate responses affects subjects' reliance on AI-generated recommendations. Subjects complete a series of CAPTCHA-style object-counting tasks, with and without access to a pre-generated recommendation from an AI system, permitting within-subject comparison of task accuracy and AI reliance across AI-available and AI-unavailable conditions; this contrast is implemented under one of three between-subjects penalty levels for incorrect responses. To help explain over- or under-reliance relative to the AI's true, undisclosed accuracy, we separately elicit subjects' subjective belief distribution over the AI's accuracy using an incentive-compatible quadratic scoring rule instrument (Matheson and Winkler, 1976), and their willingness to pay for AI access using a Becker-DeGroot-Marschak (1964) mechanism. Subjects' subjective beliefs about the AI's accuracy are compared to their own self-assessed and task-measured accuracy to provide further evidence on why subjects may over- or under-rely on AI recommendations (Moore and Healy, 2008). Subjects are recruited via Prolific and complete the experiment on Qualtrics. This research studies how monetary penalties for inaccurate responses affect participants’ reliance on AI-generated recommendations. Participants complete a series of CAPTCHA-style object-counting tasks, with and without access to a pre-generated recommendation from an AI agent. The experiment permits within-participant comparisons of task attempts and accuracy across AI-available and AI-unavailable conditions. Participants are paid a piece-rate reward for each correctly completed task and are assigned to one of three treatment levels that differ in the penalty for incorrect responses. To study differences in AI reliance across participants, we elicit participants’ beliefs about the accuracy of the AI’s recommendations using an incentive-compatible quadratic scoring rule, as well as their beliefs about their own accuracy under both AI conditions. After participants have completed tasks under both conditions, we elicit their willingness to pay for access to the AI using a Becker-DeGroot-Marschak (1964) mechanism. Our main hypotheses are that (i) participants with AI access will attempt more tasks, (ii) willingness to pay for AI access will increase with the penalty for incorrect responses, and (iii) willingness to pay will be higher among participants who perceive a larger improvement in their probability of answering correctly when using AI. Participants are recruited via Prolific and complete the experiment on Qualtrics.
Trial End Date August 20, 2026 October 07, 2026
Last Published August 20, 2026 08:31 AM September 15, 2026 01:33 PM
Intervention (Public) Subjects complete three modules, each containing 30 unique judgment tasks. Subjects are randomly assigned to one of three financial incentive conditions that differ in the monetary penalty for incorrect responses. AI recommendations are available in exactly one of the first two modules and unavailable in the other; this availability is varied within-subject and counterbalanced across subjects. Access to AI recommendations in the third module is instead determined by a Becker-DeGroot-Marschak (BDM) mechanism eliciting each subject's willingness to pay. In every module, the first task is one for which the associated AI answer is fully accurate, and the second task is one for which the associated AI answer is off by one; the relative order of these two tasks is randomized independently at each module load, regardless of whether AI recommendations are available in that module. All remaining tasks are presented in a fixed, predetermined order. When an AI recommendation is available, subjects may reveal it and, with an additional click, incorporate it into their response, so that viewing and adopting a recommendation are captured as distinct decisions. 1. Participants complete three modules, each containing 30 CAPTCHA-style object-counting tasks. 2. Participants are randomly assigned to one of three financial incentive conditions that differ in the monetary penalty (relative to a constant piece-rate reward) for incorrect responses. AI recommendations are available in one of the first two modules and unavailable in the other; AI availability is varied within-participant and counterbalanced across participants. Access to AI recommendations in the third module is determined by a Becker-DeGroot- Marschak (BDM) mechanism eliciting each participant’s willingness to pay. 3. Following each module, participants are asked to estimate the number of tasks they believe they answered correctly. After completing their first module with AI access, participants are also asked to estimate the accuracy of the AI recommendations. Interactions with the AI button, including revealing and auto-filling the recommendation, are recorded to provide additional information on AI usage.
Intervention Start Date August 13, 2026 September 30, 2026
Intervention End Date August 20, 2026 October 07, 2026
Primary Outcomes (End Points) For each subject, the primary outcomes are 1. The within-subject difference between AI and non-AI modules in number of tasks answered correctly, number of completed tasks, and percent of tasks correctly completed. 2. The subject’s willingness to pay for access to AI recommendations for the final module. Behavioral measures of reliance (e.g., AI-recommendation reveal and adoption rates) are analyzed as secondary outcomes; in the absence of a structural model dictating a single theoretically-preferred reliance measure, we do not designate any one specification as primary. For each participant, the primary outcomes are, 1. Task productivity and accuracy. The number of tasks attempted within the time limit and the number of tasks answered correctly in the AI-available and AI-unavailable modules, including differences across the three penalty conditions. 2. Willingness to pay for AI access. Participants’ willingness to pay for AI access in Module C.
Primary Outcomes (Explanation) Number of Tasks Completed is defined as the number of tasks completed within a given module. Number of Tasks Answered Correctly is defined as the number of tasks answered correctly within a given module. Percent of Tasks Answered Correctly is defined as the number of tasks answered correctly divided by the number of tasks attempted. The number of tasks answered correctly divided by the fixed number of questions per module is not included as a separate outcome, as it is a rescaling of Number of Tasks Answered Correctly. Within-Subject AI Effect is computed, separately for each of the outcomes above, as the difference in a subject's outcome between their AI-available and non-AI module. Willingness To Pay (WTP) is measured using the BDM elicitation and recorded as the maximum amount the subject is willing to pay for AI access in the third module. 1. Task productivity and accuracy. Task productivity is measured by the number of tasks attempted within the time limit, and accuracy is measured by the number of tasks answered correctly in the AI-available and AI-unavailable modules. 2. Willingness to pay. WTP is the participant’s maximum stated willingness to pay for AI access elicited using the Becker-DeGroot-Marschak mechanism in Module C.
Experimental Design (Public) Participants complete three modules of incentivized recognition tasks under time constraints. AI recommendations are available in one of the first two modules, with AI availability varied within subjects and module order counterbalanced across participants. Participants are randomly assigned to one of three financial incentive conditions that differ in the penalty for incorrect responses. In the third module, access to AI recommendations is determined through a Becker-DeGroot-Marschak (BDM) mechanism eliciting willingness to pay. Participants also complete an incentivized elicitation of their beliefs regarding the AI’s accuracy. Earnings are determined by one randomly selected module. Participants complete three modules of CAPTCHA-style object-counting tasks under time constraints. AI recommendations are available in one of the first two modules, with AI availability varied within participants and module order counterbalanced across partici- pants. Participants are randomly assigned to one of three financial incentive conditions that differ in the penalty for incorrect responses with the reward for correctly completed tasks held constant across all participants. In the third module, access to AI recommendations is determined through a Becker-DeGroot-Marschak (BDM) mechanism eliciting willingness to pay. Participants also complete an incentivized elicitation of their beliefs about the AI’s accuracy once during the experiment and a non-incentivized elicitation of their beliefs about their own accuracy following each module. Earnings are determined by one randomly selected module.
Randomization Method Randomization of first task accuracy, module presentation order, BDM price, and assignment to penalty level is conducted by the Qualtrics survey software. Randomization of the accuracy of the first two AI recommendations, module presentation order (excluding the final module), BDM price, and assignment to penalty level is conducted by the Qualtrics survey software.
Randomization Unit Penalty level is randomized at the subject level. First task accuracy for modules with AI recommendations and module order across the experiment are randomized within subject. Penalty level is randomized at the participant level. First task accuracy for modules with AI recommendations and module order across the experiment are randomized within participant.
Intervention (Hidden) Subjects complete three modules, each containing 30 unique judgment tasks. Participants are randomly assigned to one of three financial incentive conditions, which differ in the monetary penalty for an incorrect response: $0.125, $0.25, or $0.375. The reward for each correct response is held constant at $0.25 across all treatment conditions. AI recommendations are available in one of the first two modules, with module order counterbalanced across participants. In the third module, access to AI recommendations is determined using a Becker-DeGroot-Marschak (BDM) mechanism. Each participant states their maximum willingness to pay for AI recommendations, after which a purchase price is drawn uniformly at random from $0.05 to $1.05 (inclusive). Participants receive access to AI recommendations in the third module if their stated willingness to pay is greater than or equal to the randomly drawn price and pay the realized price; otherwise, they complete the module without AI recommendations and incur no charge. The AI recommendation accuracy is fixed at two-thirds (2/3) throughout the experiment. This accuracy level is held constant across all participants and treatment conditions, implying that following the AI recommendation has a positive expected monetary value under each of the three financial incentive conditions. Participants complete three modules (A, B, and C), each containing 30 unique judgment tasks. Participants are randomly assigned to one of three financial incentive conditions, which differ in the monetary penalty for an incorrect response: $0.125, $0.25, or $0.375. The reward for each correct response is held constant at $0.25 across all treatment conditions. AI recommendations are available in one of the first two modules, with module order counter- balanced across participants. In the third module, access to AI recommendations is determined using a Becker-DeGroot-Marschak (BDM) mechanism. Each participant states their maximum willingness to pay for AI recommendations, after which a purchase price is drawn uniformly at random from $0.05 to $1.05 (inclusive). Participants receive access to AI recommendations in the third module if their stated willingness to pay is greater than or equal to the randomly drawn price and pay the realized price; otherwise, they complete the module without AI recommendations and incur no charge. The AI recommendation accuracy is fixed at two-thirds (2/3) throughout the experiment. This accuracy level is held constant across all participants and treatment conditions, implying that following the AI recommendation has a positive expected monetary value under each of the three financial incentive conditions.
Secondary Outcomes (End Points) Secondary outcomes include, 1. The mean of each subject’s elicited subjective belief distribution over the AI’s accuracy. 2. The deviation between the mean of each subject's elicited belief distribution and the AI’s true accuracy. 3. The deviation between the mean of each subject's expected own accuracy and the subject's true accuracy. 4. Behavioral measures of reliance (e.g., AI-recommendation reveal and adoption rates) are analyzed as secondary outcomes; in the absence of a structural model dictating a single theoretically-preferred reliance measure, we do not designate any one specification as primary. Willingness To Pay (other measures): Measures of participants’ perceived and realized value of access to AI assistance.
Secondary Outcomes (Explanation) Elicited subjective belief distribution over the AI’s accuracy is obtained by asking participants to allocate a fixed budget of tokens across bins spanning the possible accuracy range of the provided AI recommendations. Deviation of belief distribution’s mean from the AI’s true accuracy is calculated by taking the mean of each participant’s belief distribution over the AI accuracy and subtracting it from the true AI accuracy. Deviation of participant’s beliefs over their own accuracy from revealed participant accuracy, calculated as the participant’s stated expected number of correct answers minus the participant’s realized number of correct answers, divided by the total number of questions answered. Behavioral measures include whether participants viewed an AI recommendation, whether they incorporated it into their submitted response, whether they changed their response after viewing the recommendation, whether they ultimately accepted or rejected the recommendation, and whether the submitted response matched the recommendation. Furthermore, the timing between behavioral measures acts as supportive evidence of who these actions relate to each other. 1. Perceived value of AI access. For each participant, perceived incremental value is constructed from the participant’s elicited beliefs about their probability of answering correctly with and without AI. Given the $0.25 reward for a correct response and the participant’s assigned penalty, perceived incremental value is the difference between the perceived expected payoff with AI and without AI. 2. Realized value of AI access. Realized incremental value is measured as the difference between the participant’s realized payoff in the AI-available module and their realized payoff in the AI-unavailable module.
Back to top

Other Primary Investigators

Field Before After
Affiliation Clemson University
Back to top
Field Before After
Affiliation Clemson University
Back to top