Setting Expectations: Randomizing the Structure of Quarterly Reviews

Last registered on August 10, 2026

Pre-Trial

Trial Information

General Information

Title
Setting Expectations: Randomizing the Structure of Quarterly Reviews
RCT ID
AEARCTR-0019320
Initial registration date
August 06, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
August 10, 2026, 3:48 PM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information

Primary Investigator

Affiliation
Michigan State University

Other Primary Investigator(s)

PI Affiliation
University of Arizona, Eller College of Management

Additional Trial Information

Status
In development
Start date
2026-04-03
End date
2027-12-31
Secondary IDs
MSU IRB STUDY00011827 -- FWA00004556
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
The partner organization is a U.S. credit union of roughly 1,250 employees that conducts quarterly manager–employee performance and development reviews through its HR platform in conjunction with one-on-one meetings. This study is a manager-clustered randomized field experiment testing whether restructuring these reviews around substantially more detailed documentation improves employee outcomes. Treated managers and their teams follow an enhanced protocol prompting both parties to document expectations, progress against competencies and goals, and development discussion in greater detail than the business-as-usual process; control teams continue unchanged.

Randomization was conducted at the team level — a manager together with all direct reports — so each manager delivers the same version of the review to every one of their direct reports. Of 152 teams, 76 are assigned to the enhanced protocol and 76 to business-as-usual, balanced within division on age, tenure, gender, function, and job level; matched to the most recent roster these teams comprise 1,207 regular employees (interns excluded) (597 treated, 610 control). A short survey attached to each review supplies the primary and some secondary outcomes. Two pre-treatment waves are complete; two post-treatment waves follow the intervention, giving a difference-in-differences design layered on the randomization.

The primary outcome is employee-reported expectations clarity: i.e., whether the most recent review strengthened the employee's understanding of what is expected of them. Secondary outcomes are role satisfaction, one-year retention intent, recognition adequacy, survey completion, employee–manager performance-rating calibration, and voluntary attrition. Analysis is intention-to-treat by assigned arm, using a linear probability model (LPM) with standard errors clustered at the team level and two-sided tests.
External Link(s)

Registration Citation

Citation
Sandvik, Jason and Richard Saouma. 2026. "Setting Expectations: Randomizing the Structure of Quarterly Reviews." AEA RCT Registry. August 10. https://doi.org/10.1257/rct.19320-1.0
Experimental Details

Interventions

Intervention(s)
Under the business-as-usual process, each quarterly review is organized around a small set of open-text reflections and the manager's response, with light documentation of the underlying performance and development content. The intervention keeps the quarterly cadence but restructures the pre-review prompts so that managers and employees document the components of the review in substantially greater detail: the expectations set for the period, progress against competencies and goals, development and career discussion, and the link between individual and organizational objectives. Manager- and employee-facing prompts differ. Because the enhanced documentation is embedded in the managerial review process rather than offered as an optional survey, take-up among treated units is expected to be high. All active employees on a treated team receive the treated condition.
Intervention Start Date
2026-08-06
Intervention End Date
2027-03-01

Primary Outcomes

Primary Outcomes (end points)
Employee-reported expectations clarity, measured by the review-linked survey item: “Following your most recent quarterly review, which statement best represents your sentiment regarding what is expected of you at [organization]?” (stronger / unchanged / weaker understanding). The registered primary is the binary indicator: stronger versus not stronger.
Primary Outcomes (explanation)
Intention-to-treat effect of assignment to the enhanced protocol, estimated by difference-in-differences (two pre-treatment versus two post-treatment quarters of data, treated versus control teams) using an LPM with standard errors clustered at the team level, two-sided tests, and pre-specified covariate adjustment for baseline team characteristics. The item is re-anchored at each review, so the post-period treated-versus-control comparison is identifying on its own; the pre-treatment waves add precision and a balance check. The estimation sample is a repeated cross-section of valid responses per wave; a balanced-panel analysis is reported as a standing robustness check, as is the ordinal coding (−1/0/+1 linear, and ordered logit).

Secondary Outcomes

Secondary Outcomes (end points)
● Role satisfaction (0–10)
● One-year retention intent (0–10)
● Recognition adequacy (adequately recognized / recognized but not adequately / not recognized)
● Survey completion
● Employee–manager rating calibration — employee's predicted manager rating vs. the manager's survey rating
● Voluntary attrition — voluntary separations, through 1 Mar 2027
● Exploratory: Gallup Q12 group-level engagement, especially Q01 (“I know what is expected of me at work”); and internal coherence between self-rated and predicted manager performance ratings
Secondary Outcomes (explanation)
Role satisfaction, one-year retention intent, and recognition adequacy are analyzed separately, each using the same intention-to-treat difference-in-differences specification as the primary: team-clustered, two-sided, covariate-adjusted; satisfaction and retention intent as bounded integer scores entered linearly, recognition as a binary. Recognition adequacy collapses “not recognized” and “recognized but not adequately” against “adequately recognized,” with the full three-level distribution reported descriptively by arm and wave. Results across the secondary outcomes are reported without adjustment for multiple comparisons. Survey completion is analyzed as a response indicator. Calibration is computed for employee–manager pairs in which both the employee and their manager responded in the same wave. Employees report how they believe their manager would rate their performance against expectations and managers rate them separately on the research survey, with both reports collected after the review has taken place, on the same three-point scale — did not meet, met, or exceeded expectations — coded 1 to 3 so that higher values are more favorable. Two measures are reported: an agreement indicator equal to one when the employee's perception matches the manager's realized evaluation, and the employee's perception net of the manager's realized evaluation, where a positive value indicates the employee over-estimated. Voluntary attrition is measured at the individual level through 1 March 2027. The Gallup Q12 is exploratory only: it is observed at group level rather than individually, and its 2026 wave is fielded 7–22 October, capturing only partial first-cycle exposure; the pre-period wave is September 2025 and the realized exposure share as of 22 October 2026 will be reported alongside any comparison.

Experimental Design

Experimental Design
Two-arm, parallel, cluster-randomized controlled trial using a difference-in-differences model. The unit of randomization is the team: a manager together with all direct reports, which prevents within-manager contamination. 152 teams were randomized 76/76, balanced within the institution's 11 divisions; matched to the most recent roster these teams comprise 1,207 regular employees (interns excluded) (mean team size 8.0, range 1–19); realized arms are balanced on age, tenure, gender, member-facing versus back-office status, job level, and pay type. Each wave is defined by the review cycle it covers rather than by a calendar cut-off.
Experimental Design Details
Not available
Randomization Method
Computer-generated (Stata) pseudo-random assignment executed prior to launch from the organization's roster as finalized at randomization, balanced within division.
Randomization Unit
Team — a manager and all of their direct reports.
Was the treatment clustered?
Yes

Experiment Characteristics

Sample size: planned number of clusters
152 teams: 76 treatment, 76 control
Sample size: planned number of observations
1,207 regular employees (interns excluded), as matched to the most recent roster; the operative headcount at each wave is the roster in force at that wave. The per-wave analysis sample is the subset completing the outcome item; response rate is the binding constraint on power.
Sample size (or number of clusters) by treatment arms
Treatment — enhanced protocol: 76 teams and 597 employees
Control — business as usual: 76 teas and 610 employees
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
Power is governed by the survey response rate rather than the number of clusters. The baseline share reporting strengthened understanding is 0.409. The intra-cluster correlation of the primary outcome, estimated from the first baseline wave across 94 teams by one-way random-effects ANOVA, is 0.049 (cluster bootstrap 95% CI −0.09 to 0.19); 0.05 is used as the planning value. At two-sided α = 0.05 and 80% power, with 76 clusters and 604 employees per arm and two pooled post-treatment waves, the minimum detectable effect on the treated-versus-control post-period comparison is 13.8 percentage points at the observed 18% response rate, 10.5 points at 35%, and 9.2 points at 50%. Treating the pre-period as contributing independent sampling variance of equal magnitude yields conservative upper bounds of 19.5, 14.8, and 13.0 points respectively. Varying the intra-cluster correlation across the bootstrap interval changes the minimum detectable effect by under two percentage points, so response rate rather than clustering is the operative constraint. Voluntary attrition, given a low base rate over a six-month window, is the least well-powered outcome.
IRB

Institutional Review Boards (IRBs)

IRB Name
Michigan State University Social Science / Behavioral / Education Institutional Review Board
IRB Approval Date
2025-05-16
IRB Approval Number
STUDY00011827