Human Oversight in AI-Based Performance Evaluation: A Field Experiment

Last registered on October 07, 2026

Pre-Trial

Trial Information

General Information

Title
Human Oversight in AI-Based Performance Evaluation: A Field Experiment
RCT ID
AEARCTR-0019837
Initial registration date
October 05, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
October 07, 2026, 10:59 AM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information

Primary Investigator

Affiliation
WU Vienna University of Economics and Business

Other Primary Investigator(s)

PI Affiliation
PI Affiliation
PI Affiliation

Additional Trial Information

Status
In development
Start date
2026-11-27
End date
2027-05-30
Secondary IDs
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
This study examines how the design of AI-based evaluation systems affects employee work strategies and task behavior. We conduct a field experiment in collaboration with a Taiwanese company (referred to as "FocalCompany" for privacy reasons) that provides field services for consumer goods producers. Specifically, we compare two conditions: one in which field agents' work outputs are evaluated solely by an AI-based checking system providing immediate feedback, and one in which AI judgments may additionally be reviewed by a human evaluator. We expect that the prospect of human oversight — which introduces delayed, less predictable judgment and social considerations — leads agents to adopt a more cautious and effort-intensive work strategy, relative to AI-only evaluation under which agents can act upon immediate feedback and optimize their approach over time. We examine these predictions using behavioral performance data from the company's internal information system as well as self-reported measures from a multi-wave survey. In addition, we examine whether individual-level or situational characteristics moderate agents' responses to the evaluation system design.
External Link(s)

Registration Citation

Citation
Feichter, Christoph et al. 2026. "Human Oversight in AI-Based Performance Evaluation: A Field Experiment." AEA RCT Registry. October 07. https://doi.org/10.1257/rct.19837-1.0
Experimental Details

Interventions

Intervention(s)
Intervention Start Date
2026-12-01
Intervention End Date
2027-02-28

Primary Outcomes

Primary Outcomes (end points)
Outcome variables are drawn from two sources:
• Behavioral performance data from FocalCompany's internal information system, including: number of stores visited, visit duration, tasks completed, task characteristics, or quality-related indicators. These data are available for the pre-experimental period as well as during the experiment, enabling difference-in-differences analyses.
• Self-reported measures from a three-wave survey (pre, mid, and post intervention), capturing agents' perceived work approach, caution, effort, and perceptions of the evaluation system.
Primary Outcomes (explanation)

Secondary Outcomes

Secondary Outcomes (end points)
Secondary Outcomes (explanation)

Experimental Design

Experimental Design
The study uses a two-condition between-participants field experiment conducted in collaboration with FocalCompany, a Taiwanese company that provides field services for consumer goods producers. FocalCompany employs approximately 100 field agents who regularly visit and serve retail locations and document their completed tasks by uploading pictures via the company app. FocalCompany is implementing an AI-based checking system to evaluate these pictures, and our study accompanies this implementation. Field agents are organized in eight existing work teams; to reduce spillovers among employees who occasionally work together, treatment assignment occurs at the team level.
Condition 1 – AI only: Employees are informed that uploaded pictures of completed tasks are evaluated by an AI-based checking system that provides immediate feedback.
Condition 2 – AI plus human oversight: Employees are informed that pictures are initially checked by AI, but may additionally be reviewed by a human evaluator to detect possible inconsistencies.
Participants will be informed about the AI pilot and the accompanying research, but will not be informed that different teams are assigned to different conditions. This incomplete disclosure is necessary to avoid contamination and demand effects. The manipulation does not affect formal compensation or require participants to perform tasks beyond their regular work responsibilities.
In addition to the field experiment, a small number of voluntary interviews will be conducted with managers and field agents before the experiment (to understand organizational context) and after the experiment (to interpret findings and understand potential null results). A three-wave survey is administered before, during, and after the intervention.

No pilot study was conducted. The survey instrument was developed and refined prior to the pre-registration, but no data collection involving the experimental manipulation has taken place.
Experimental Design Details
Not available
Randomization Method
FocalCompany's approximately 100 field agents are organized into eight geographically separated work teams across different regions of Taiwan. Within each region, teams are assigned to conditions in equal numbers using a computer-generated random draw, ensuring balanced treatment assignment across regions. Within teams, agents occasionally interact, but cross-team contact is rare, limiting the risk of contamination across conditions.
Randomization Unit
The randomization unit is the work team. All members of a given team are assigned to the same condition.
Was the treatment clustered?
Yes

Experiment Characteristics

Sample size: planned number of clusters
8 work teams, with 4 teams assigned to each condition.
Sample size: planned number of observations
Behavioral data from the company's internal information system will be available for all participating field agents (approximately 100). Survey observations may be lower depending on participation rates; to incentivize completion, participants will receive a financial incentive for completing each survey wave, the details of which are to be determined. We expect up to approximately 100 survey responses per wave.
Sample size (or number of clusters) by treatment arms
4 teams per treatment arm. The exact number of individual agents per arm depends on team sizes. We expect approximately 50 agents per condition, though the precise split will be determined by the actual team composition at the time of the intervention.
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
IRB

Institutional Review Boards (IRBs)

IRB Name
WU ETHICS BOARD
IRB Approval Date
2026-05-29
IRB Approval Number
WU-RP-2026-038