Abstract
We survey economists (have or pursuing a PhD in economics or finance) to elicit ratings of the validity and relevance of robustness checks for a key empirical finding in each of several papers published in the American Economic Review. Each participant is randomly assigned four papers and, for each paper, evaluates six robustness checks drawn without replacement from a larger fixed pool of candidate checks for that paper's key result. Check order and paper order are randomized. For each check, participants rate its validity and relevance on a 0–10 scale and separately predict the average rating other participants will give the same check (their second-order belief), with a financial incentive tied to forecast accuracy. Participants also answer incentivized comprehension questions about each paper's main result, rate the paper's credibility and their own familiarity with its methods, and answer background questions on seniority, sub-field, region, and gender. We test whether ratings and second-order beliefs differ by type of the check, and whether raters agree with each other more for checks from one source than the other — separately for two sets of papers — along with exploratory associations with rater seniority, sub-field match, familiarity, and engagement with the source paper. Recruitment is by invitation to faculty and PhD students at economics departments, participants in the Institute for Replication games, and professional networks, targeting 300 participants who complete the survey (up to 300 × 24 = 7,200 check-level rating and forecast observations). Compensation includes a show-up fee, payment for correct comprehension answers, and a forecast-accuracy bonus, for a maximum of $62 per participant.