Ambiguous Reliability and the Weight of Advice: Experimental Evidence

Last registered on August 31, 2026

Pre-Trial

Trial Information

General Information

Title
Ambiguous Reliability and the Weight of Advice: Experimental Evidence
RCT ID
AEARCTR-0019547
Initial registration date
August 31, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
August 31, 2026, 9:02 AM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

Primary Investigator

Affiliation
VNU University of Economics and Business

Other Primary Investigator(s)

Additional Trial Information

Status
In development
Start date
2026-08-31
End date
2026-12-31
Secondary IDs
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
Advice is often received from a source whose reliability is not precisely known. This study asks whether reluctance to rely on such advice reflects ambiguity aversion, and whether it can be changed by altering how reliability is communicated while holding its expected value fixed.

Participants estimate the caloric content of photographed meals, receive an estimate from an AI system, and revise. We randomize what participants are told about the system's accuracy: nothing (control), a point value ("accurate in 80 of 100 cases"), or an interval with the same expected value ("accurate in 65 to 95 of 100 cases; no more precise figure is available"). Expected accuracy is identical across the two disclosure arms, so expected-utility accounts predict no difference between them.

Guided by an alpha-maxmin expected utility framework, we predict that interval disclosure lowers reliance for ambiguity-averse participants and raises it for ambiguity-loving participants, with the contrast reversing sign at ambiguity neutrality. We elicit ambiguity attitude and risk attitude with multiple incentivized instruments and correct for measurement error using ORIV.

We additionally elicit participants' own beliefs about the system's accuracy as an interval, which allows us to estimate how much of the disclosed midpoint and the disclosed width they adopt, and to test whether that adoption depends on elicited ambiguity attitude.
External Link(s)

Registration Citation

Citation
Nguyen, Tuan. 2026. "Ambiguous Reliability and the Weight of Advice: Experimental Evidence." AEA RCT Registry. August 31. https://doi.org/10.1257/rct.19547-1.0
Sponsors & Partners

Sponsors

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information
Experimental Details

Interventions

Intervention(s)
Participants complete three estimation tasks in which they judge how many calories are contained in a photographed meal. In each task, they first record their own estimate, then see an estimate produced by an artificial intelligence system, and then record a final estimate. Payment depends on the accuracy of the final estimate, so participants have a financial reason to use the advice when they believe it is informative and to ignore it when they do not.

The intervention is a single page shown before the estimation tasks. This page describes the system that produced the advice. Its content is identical across treatment arms except for one statement about how accurate the system has been in prior testing. Participants in the control arm read no statement about accuracy. Participants in the second arm read that the system was accurate in 80 of 100 test cases. Participants in the third arm read that testing produced varying results, with accuracy falling somewhere between 65 and 95 of 100 cases, and that no more precise figure could be determined.

The second and third statements describe the same average level of accuracy. They differ only in whether that level is presented as a single number or as a range. This allows us to separate the effect of knowing how accurate a source is from the effect of knowing that accuracy precisely.

Accuracy is defined for participants as an estimate falling within 15 per cent of the true value. Before the study begins, we verify the system's accuracy on 100 comparable meal photographs whose true calorie content has been measured, so that both statements shown to participants are factually correct. No misleading information is used at any point in the study.

The advice shown for each of the three experimental photographs is generated once
and held fixed for every participant. The advice itself is therefore identical
across arms, and any difference in behaviour can be attributed to the accuracy
statement rather than to the content of the advice.
Intervention Start Date
2026-09-01
Intervention End Date
2026-12-31

Primary Outcomes

Primary Outcomes (end points)
Weight on advice, measured as the extent to which participants move their estimate toward the advice they receive.
Primary Outcomes (explanation)
For each estimation task we record three numbers: the participant's initial estimate, the advice shown, and the participant's final estimate. Weight on advice is the share of the distance between the initial estimate and the advice that the participant travels when revising. A value of zero means the participant ignored the advice entirely, a value of one means the participant adopted it completely, and intermediate values mean the participant combined the two. The primary outcome is the average of this measure across the three tasks.

This measure is standard in the literature on advice taking and allows our results to be compared with published benchmarks. It also has the useful property of being independent of the size of the numbers involved, so tasks with larger or smaller calorie values contribute on the same scale.

The measure becomes unstable when the advice happens to fall very close to the participant's own estimate, because the denominator then approaches zero. For this reason our main specification does not compute the ratio directly. Instead we regress the final estimate on the initial estimate and on the advice, and take reliance to be the coefficient on the advice divided by the sum of the two coefficients. This produces the same quantity while remaining well behaved.

We report two further versions to show that the finding does not depend on how the outcome is constructed. The first is the ratio measure itself, restricted to values between zero and one and excluding cases where the advice falls within 10 per cent of the initial estimate. The second is a simple indicator of whether the final estimate moved toward the advice at all. If the three versions agree, the result is not an artefact of the measure.

Secondary Outcomes

Secondary Outcomes (end points)
Perceived reliability of the source, measured by the width and midpoint of the belief interval participants report about the system's accuracy; the degree to which participants adopt the disclosed midpoint and width; confidence in one's own estimate; and perceived clarity, certainty, and trust regarding the accuracy statement.
Secondary Outcomes (explanation)
Before the estimation tasks, participants state the smallest number, the largest number, and their best guess for how many correct answers the system would give in 100 comparable tasks. The distance between the smallest and largest figures measures how much ambiguity the participant perceives about the source. This elicitation is incentivized: a participant is paid if the system's verified accuracy falls inside the interval they state, minus an amount proportional to how wide that interval is. The penalty for width means participants gain nothing by stating a very wide interval simply to be safe.

These beliefs allow us to estimate two quantities that describe how disclosed information is absorbed. The first is the weight participants place on the midpoint of the statement they read. The second is the weight they place on its width. The control arm identifies what participants believe when they are told
nothing, and each disclosure arm then reveals how far beliefs move toward what was stated. Because the point arm and the interval arm each provide an estimate of the same width parameter, the two estimates can be compared, and disagreement between them would indicate that the framework does not describe belief formation well. We regard this as a strength of the design, since it allows the framework to be rejected rather than only confirmed.

Three further measures serve as checks on the intervention rather than as outcomes of interest in themselves. We ask how clear the accuracy statement was, how certain the participant feels about the system's accuracy, and how much they trust the system. The design requires that the interval statement lowers certainty without lowering clarity. If both fall together, the range was confusing rather than ambiguous, and the results would need to be interpreted differently. We report these measures before the main results so that readers can judge whether the intervention worked as intended.

After the estimation tasks we also ask how accurate participants believe such systems to be. Because this question comes after treatment, it is affected by the treatment and is used only descriptively. Its purpose is to confirm that participants regarded the system as capable of error. If participants believe the system is rarely wrong, a statement about its accuracy carries little information, and the intervention has nothing to work with.

Experimental Design

Experimental Design
The study is a laboratory experiment conducted on paper with groups of up to 20 participants. Each session lasts approximately 41 minutes. Participants are randomly assigned to one of three arms, and every participant completes three estimation tasks, so that treatment varies between participants while the tasks vary within them.

The session proceeds in a fixed order. Participants first answer background questions and report general attitudes toward risk and uncertainty, their experience with artificial intelligence tools, and whether they experienced particular difficulties before the age of 18. They then complete an incentivized task measuring their attitude toward ambiguity, using two physical urns, one of known composition and one of unknown composition. An incentivized measure of risk attitude follows this.

The intervention page comes next, followed immediately by questions checking attention and comprehension and by ratings of how clear the statement was and how certain participants feel about the system's accuracy. Participants then state what accuracy they expect from the system, and complete the three estimation tasks. Placing these parts together and instructing participants not to return to earlier pages ensures that beliefs are recorded after the statement is read and before any advice is seen.

The final parts of the session measure ambiguity attitude and risk attitude a second time, using different question formats and unrelated sources, and collect measures of cognitive reflection and numeracy along with closing questions. Measuring each preference more than once is a deliberate feature of the design
rather than a repetition, because preferences estimated from a small number of choices contain substantial measurement error, and correcting for that error requires more than one measure.

Every decision in the session is incentivized. One decision is selected at random by dice at the end of the session and carried out with real money, in addition to a fixed participation fee. Because any decision may be the one that counts, participants have a reason to answer each question according to their true preferences.
Experimental Design Details
Not available
Randomization Method
Randomization was carried out in advance of data collection by the principal investigator using statistical software. Assignment was blocked in groups of 30 and stratified by gender and by self reported frequency of artificial intelligence use, which keeps the arms balanced on these characteristics even if data
collection is interrupted before the target sample is reached.

Booklets were printed in three versions, numbered according to the allocation table, and sealed before sessions began. The three versions are identical in layout, length and appearance, and differ only in the wording of one statement on one page. Research assistants distribute booklets in numerical order and therefore do not know which arm any participant has been assigned to. No randomization takes place in the session itself.
Randomization Unit
The individual participant. There is a single level of randomization. Treatment is assigned at the individual level and does not vary within a participant, while the three estimation tasks are completed by every participant regardless of arm.
Was the treatment clustered?
No

Experiment Characteristics

Sample size: planned number of clusters
We intend to collect data from 300 individual participants. The design is not clustered, so the number of clusters equals the number of participants.
Sample size: planned number of observations
900 participant task observations, arising from 300 participants who each complete three estimation tasks. The primary outcome is computed as the average across a participant's three tasks, giving 300 observations at the level of analysis.
Sample size (or number of clusters) by treatment arms
50 participants control, no statement about accuracy
125 participants point disclosure, accurate in 80 of 100 cases
125 participants interval disclosure, accurate in 65 to 95 of 100 cases

The allocation is deliberately unequal. The comparison between the point and interval arms is an interaction between treatment and a measured characteristic, and interactions require considerably more observations than differences in means. The comparison between the point arm and the control arm is a difference in means and reaches acceptable power with a much smaller control group. Dividing the sample equally would have put participants on the easier comparison and reduced power for the harder one.
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
The outcome is weight on advice, which is a proportion. A value of zero means the participant ignored the advice and a value of one means the participant adopted it completely. Effects below are expressed in units of that proportion. Power was calculated by Monte Carlo simulation rather than by formula, because the primary test is an interaction between a randomly assigned treatment and a continuous measured characteristic, and because the outcome is averaged over three tasks per participant. The simulation assumes a standard deviation of 0.12 across participants and 0.22 across tasks within a participant in weight on advice, and a standard deviation of 0.22 in normalized ambiguity attitude. These values will be re-estimated from pilot data and the calculation repeated before the main wave begins. Primary test. The interaction between interval disclosure and ambiguity attitude, comparing 125 participants in the point arm with 125 in the interval arm. At 80 per cent power and a 5 per cent significance level, the smallest detectable slope is 0.30 units of weight on advice per unit of normalized ambiguity attitude. In practical terms, this means the design can detect a situation in which the interval statement raises reliance by about 0.15 for the most ambiguity-loving participants and lowers it by about 0.15 for the most ambiguity-averse. The effect implied by the theoretical framework under our calibration is 0.315, which the design detects with 85 per cent power. If measurement error in the elicited preference attenuates the slope to 0.24, power falls to 62 per cent. Secondary test. The difference in mean weight on advice between the point arm and the control arm, comparing 125 participants with 50. The smallest detectable difference is 0.08 units of weight on advice at 80 per cent power. This comparison is well powered at a small control group because it is a difference in means rather than an interaction, which is the reason for the unequal allocation. We state plainly that the design is powered to detect the predicted pattern but is not powered to rule it out. If the interaction is not statistically significant, that result will not distinguish between a genuine absence of effect and a sample too small to detect one of the size we expect. We note this in advance so that a null result is interpreted correctly rather than treated as evidence against the prediction.
Supporting Documents and Materials

Documents

Document Name
IRB
Document Type
irb_protocol
Document Description
File
IRB

MD5: 4d61073bc748c43c7c57694def541847

SHA1: bd693185e9f1ebeaf1b4d90ce0b96f2339932b1b

Uploaded At: August 31, 2026

IRB

Institutional Review Boards (IRBs)

IRB Name
VNU University of Economics and Business
IRB Approval Date
2026-07-30
IRB Approval Number
2026-REC-UEB-12
Analysis Plan

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information