Rating robustness tests

Last registered on September 28, 2026

Pre-Trial

Trial Information

General Information

Title
Rating robustness tests
RCT ID
AEARCTR-0019805
Initial registration date
September 24, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
September 28, 2026, 9:43 AM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information

Primary Investigator

Affiliation
University of Pittsburgh

Other Primary Investigator(s)

PI Affiliation
Stockholm School of Economics
PI Affiliation
Nova School of Business and Economics
PI Affiliation
Trinity College Dublin

Additional Trial Information

Status
In development
Start date
2026-09-25
End date
2027-12-31
Secondary IDs
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
We survey economists (have or pursuing a PhD in economics or finance) to elicit ratings of the validity and relevance of robustness checks for a key empirical finding in each of several papers published in the American Economic Review. Each participant is randomly assigned four papers and, for each paper, evaluates six robustness checks drawn without replacement from a larger fixed pool of candidate checks for that paper's key result. Check order and paper order are randomized. For each check, participants rate its validity and relevance on a 0–10 scale and separately predict the average rating other participants will give the same check (their second-order belief), with a financial incentive tied to forecast accuracy. Participants also answer incentivized comprehension questions about each paper's main result, rate the paper's credibility and their own familiarity with its methods, and answer background questions on seniority, sub-field, region, and gender. We test whether ratings and second-order beliefs differ by type of the check, and whether raters agree with each other more for checks from one source than the other — separately for two sets of papers — along with exploratory associations with rater seniority, sub-field match, familiarity, and engagement with the source paper. Recruitment is by invitation to faculty and PhD students at economics departments, participants in the Institute for Replication games, and professional networks, targeting 300 participants who complete the survey (up to 300 × 24 = 7,200 check-level rating and forecast observations). Compensation includes a show-up fee, payment for correct comprehension answers, and a forecast-accuracy bonus, for a maximum of $62 per participant.
External Link(s)

Registration Citation

Citation
Johannesson, Magnus et al. 2026. "Rating robustness tests." AEA RCT Registry. September 28. https://doi.org/10.1257/rct.19805-1.0
Experimental Details

Interventions

Intervention(s)
Participants are shown, for each of four randomly assigned papers published in the American Economic Review, a description of the paper's key empirical result (hypothesis, coefficient, standard error, t/z-value, p-value) and a link to the published article. After answering incentivized comprehension questions and rating the paper's credibility and their own familiarity with its methods, participants are shown six candidate robustness checks for that paper's key result, drawn without replacement from a larger, fixed pool of candidate checks assembled in advance for that paper. For each check, participants (i) rate its validity and relevance on a 0–10 scale and (ii) predict, also on a 0–10 scale, the average rating other participants will give the same check, excluding their own rating; the second task carries a monetary accuracy incentive. Checks are presented in random order. The target population is economists and economics/finance PhD students, recruited by invitation, targeting 300 participants who complete all four assigned papers.
Intervention Start Date
2026-09-25
Intervention End Date
2027-01-31

Primary Outcomes

Primary Outcomes (end points)
(1) Robustness rating: a participant's 0–10 rating of the validity and relevance of a given robustness check. (2) Second-order belief: a participant's 0–10 prediction of the average robustness rating that other participants will give the same check, excluding the participant's own rating.
Primary Outcomes (explanation)
Both outcomes are collected once for every robustness check a participant evaluates (24 per participant: 6 checks × 4 papers). The second-order belief targets the same 0–10 scale as applied by other complete participants who rated the same check; participants receive a monetary bonus based on the squared error of this forecast, averaged across their 24 predictions. Both outcomes are analyzed separately for two sets of papers (11 papers drawn from a prior robustness study, and 12 recently published or accepted papers), each contributing two of the four papers a given participant evaluates. See Experimental Design for how checks are sourced and selected within each paper.

Secondary Outcomes

Secondary Outcomes (end points)
Disagreement: for each rated robustness check, the mean absolute difference in ratings across all distinct pairs of complete participants who rated that check, in rating-scale points.
Secondary Outcomes (explanation)
Measured at the level of the individual robustness check (one observation per check rated in the study), using raw 0–10 ratings without adjusting for individual scale use. Participants also rate each paper's overall credibility (0–10) and their own familiarity with its methods (0–10); these and background measures (seniority, sub-field match, having read the paper before, whether the participant clicked through to the published article) are used as covariates in exploratory analyses, not as trial outcomes.

Experimental Design

Experimental Design
Within-subject design with no between-participant treatment arms. Each participant is randomly assigned four of 23 eligible papers (two from each of two paper sets) and, for each paper, six robustness checks drawn from a larger fixed per-paper pool, split across two comparator sources without disclosing what differentiates them. Paper order and check order are randomized; assignment balances how many times each check and each paper is shown across participants. Recruitment is by invitation to faculty and PhD students at economics departments, Institute for Replication replication-game participants, professional networks, and the authors' personal networks, targeting 300 participants who complete the full survey (four papers each).
Experimental Design Details
Not available
Randomization Method
Computer-generated random assignment with fixed seeds recorded in the pre-analysis plan, determining: (a) which four of 23 papers a participant sees (two from each paper set) and their presentation order; (b) for each paper, which six of a larger fixed per-paper shortlist of checks the participant sees, and their order — subject to balancing presentation counts across checks and across papers within each set.
Randomization Unit
Individual participant, for both paper assignment and check assignment. There is no cluster- or group-level randomization.
Was the treatment clustered?
No

Experiment Characteristics

Sample size: planned number of clusters
N/A
Sample size: planned number of observations
Target of 300 participants who complete the survey. Each rates 24 robustness checks (6 per paper × 4 papers) and provides 24 corresponding second-order-belief forecasts — up to 7,200 check-level observations of each outcome.
Sample size (or number of clusters) by treatment arms
This is a within-subject design with no between-participant treatment arms: all ~300 participants complete an identical task structure (see Experimental Design) and are exposed to checks from both comparator sources described there. The relevant comparison is at the level of the individual check rating, not the participant. The composition of the two check pools is detailed in the embargoed portion of the Experimental Design field and in the confidential Pre-Analysis Plan.
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
Not computed ex ante, since the standard deviation of ratings and the within-participant correlation across observations aren't known before data collection. The pre-analysis plan instead commits to reporting, after data collection, the minimum detectable effect size with 90% power for each primary hypothesis test: 3.24 × the observed standard error for two-sided tests at the 5% ("suggestive evidence") threshold, and 4.09 × the observed standard error at the 0.5% ("statistically significant evidence") threshold.
Supporting Documents and Materials

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information
IRB

Institutional Review Boards (IRBs)

IRB Name
University of Pittsburgh Institutional Review Board
IRB Approval Date
2026-08-18
IRB Approval Number
N/A
Analysis Plan

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information