Back to History

Fields Changed

Registration

Field Before After
Abstract Generative AI can now be used to help citizens move from a mere social concern to a concrete, structured policy proposal. This creates a question for participatory politics: what changes when part of the work of generating a proposal can be supported by AI? AI may matter on the formulation side, by shaping the clarity, structure, feasibility, and appeal of the proposals citizens formulate. It may also matter on the reception side, by shaping how other citizens react to those proposals, especially when they know that AI was involved in producing them. We study these questions in the context of Le Primarie delle Idee, an Italian participatory platform launched in April 2026. The platform allows registered users to develop policy proposals either with the support of a generative AI assistant or on their own. In the AI-assisted route, users interact with a multi-turn conversational facilitator based on a generative AI model, called AIdea and designed to help them turn an initial concern into a more concrete proposal. In the self-authored route, users draft their proposal directly using a template. Once submitted, proposals can be read, discussed, supported, and voted on by the user community. The experiment therefore studies both the formulation and reception of policy proposals in an online political platform. On the formulation side, we ask whether AI-assisted proposals differ from self-authored proposals in quality, structure, and electoral appeal. Among users who receive AI assistance, we also ask whether the design of that assistance matters, by comparing a procedural, scaffolded version of the assistant with a more open-ended version and measuring effects on users' policy-reasoning skills. On the reception side, we ask whether disclosure matters: do users engage differently with a proposal when they are told whether it was AI-assisted or self-authored? We answer these questions through two main randomizations: one over the proposal-development route and one over whether proposals display an authorship tag indicating that they were either AI-assisted or self-authored, plus a nested randomization over the version of the assistant assigned to users. Generative AI can now be used to help citizens move from a mere social concern to a concrete, structured policy proposal. This creates a question for participatory politics: what changes when part of the work of generating a proposal can be supported by AI? AI may matter on the formulation side, by shaping the clarity, structure, feasibility, and appeal of the proposals citizens formulate. It may also matter on the reception side, by shaping how other citizens react to those proposals, especially when they know that AI was involved in producing them. We study these questions in the context of Le Primarie delle Idee, an Italian participatory platform launched in April 2026. The platform allows registered users to develop policy proposals either with the support of a generative AI assistant or on their own. In the AI-assisted route, users interact with a multi-turn conversational facilitator based on a generative AI model, called AIdea and designed to help them turn an initial concern into a more concrete proposal. In the self-authored route, users draft their proposal directly using a template. Once submitted, proposals can be read, discussed, supported, and voted on by the user community. The experiment therefore studies both the formulation and reception of policy proposals in an online political platform. On the formulation side, we ask whether AI-assisted proposals differ from self-authored proposals in quality, structure, and electoral appeal. Among users who receive AI assistance, we also ask whether the design of that assistance matters, by comparing a procedural, scaffolded version of the assistant with a more open-ended version and measuring effects on users' policy-reasoning skills. On the reception side, we ask whether disclosure matters: do users engage differently with a proposal when they are told whether it was AI-assisted or self-authored? We answer these questions through three main randomizations: one over the proposal-development route and one over whether proposals display an authorship tag indicating that they were either AI-assisted or self-authored, plus a nested randomization over the version of the assistant assigned to users and a further randomization in the final voting voting of the platform.
Trial End Date October 05, 2026 January 31, 2027
Last Published October 02, 2026 05:12 AM October 02, 2026 05:25 AM
Primary Outcomes (End Points) The primary outcomes, grouped by randomization, are: (1) For the proposal-development route: proposal quality (a blind-evaluated rubric score), proposal-level engagement on the platform (supports, comments, event participation, and sharing), and individual-level submission behavior (whether a user submits any proposal, and the number of proposals submitted). (2) For the facilitator mode (procedural vs. open-ended): proposal quality, and characteristics of the development conversation, including its length and how it concludes, the breadth of policy-reasoning components covered, and the extent of post-generation editing of the proposal. (3) For the authorship-disclosure tag: engagement with proposals on the platform (supports, comments, event participation, sharing, and final voting), and whether users open or submit a proposal via the AI-assisted or self-authored route. The partition of outcomes into primary and secondary is preliminary and depends on the statistical power available given platform adoption; it will be finalized in an updated plan once the realized sample is known. The primary outcomes, grouped by randomization, are: 1. Proposal-development route: proposal quality; platform engagement, including the ranking score and top-50 status; submission and proposal count; thematic and semantic dispersion; within-user proposal diversity; and linguistic polarization. 2. Facilitator mode: proposal quality; conversation length and how it concludes; coverage of four complexity components; and user initiative. 3. Authorship-disclosure tag: user-level and user–proposal engagement through supports, comments, event participation and sharing. 4. Final-vote advice: submission of any ballot; inclusion of a pillar among submitted ranked choices; and opening a pillar’s description. Primary and secondary outcomes are specified in the updated pre-analysis plan.
Primary Outcomes (Explanation) Several outcomes are constructed. Proposal quality is a 0–12 score from a six-item rubric (clarity/coherence, causal mechanism, feasibility acknowledgment, trade-off recognition, evidence referenced, distributional specificity), each scored 0–2 by two blind independent human evaluators, with a third evaluator triggered by disagreement greater than two points on any item; the full proposal corpus is additionally scored by LLM evaluators using the same rubric as a robustness check, validated against human scores, with human scores as the primary measure. Platform engagement is constructed from logged user interactions (supports, comments, event participations, and shares); a platform "most shared" score weights these three points per support, two per event participation, and one per comment. Conversation complexity is constructed from the coverage of policy-reasoning components and a measure of user initiative (the share of components first introduced by the user rather than the facilitator). Post-generation editing is measured as both a binary indicator of any manual modification and the magnitude of change (semantic and textual distance between the generated and submitted versions). Detailed operational definitions are provided in the uploaded pre-analysis plan. Proposal quality is a 0–12 score covering complexity, language clarity and voter appeal, each comprising four binary criteria. Three independent LLM evaluators code each criterion; majority codes are summed into dimension and total scores. Platform engagement uses logged supports, comments, event participations and shares. The platform ranking score assigns three points per support, two per event participation and one per comment. Dispersion and within-user diversity use CAP topic differences and pairwise semantic distances. Linguistic polarization indicates the presence of any of five prespecified hostile or polarizing language features. Secondary ideological extremism measures distance from the CHES scale midpoint, using the average ideological position of parties coded as supporting the proposal. Conversation coverage scores four complexity components: trade-offs, implementation constraints, second-order effects and uncertainty. User initiative is the share of covered components first raised by the user, undefined when none is covered. Post-generation editing is a descriptive diagnostic measuring any modification and semantic distance between generated and submitted proposals. Final-vote outcomes are binary indicators of ballot submission, pillar selection and information gathering, coded zero when no corresponding action is recorded. Detailed definitions and aggregation procedures appear in the updated pre-analysis plan.
Experimental Design (Public) Registered users are randomized at the individual level along three orthogonal dimensions. The first assigns users to one of two proposal-development routes: an AI-assisted route, in which they develop a proposal through conversation with a generative AI facilitator, or a self-authored route, in which they draft directly using a template. Nested within this, a second assignment sets the AI facilitator to one of two modes: a procedural, scaffolded style or an open-ended style, and applies to all users, becoming operative only for those who use the facilitator. A third, independent assignment determines whether the platform interface displays a tag disclosing each proposal's mode of authorship (AI-assisted or self-authored). The three assignments are orthogonal by design. Analysis is primarily by intent-to-treat. The study runs on a live participatory platform from June to October 2026, concluding with a final platform vote among the most-supported proposals. Because the realized sample depends on platform adoption over the experimental window, the partition of outcomes and some specifications are preliminary and will be finalized in an updated plan. Further design detail is provided in the uploaded pre-analysis plan. During June 25–September 20, 2026, users received three orthogonal assignments at the individual level: AI-assisted versus self-authored proposal development (R1); procedural versus open-ended AI facilitation (R1.1); and visible versus hidden authorship tags (R2). Facilitator mode was assigned to all users and applied whenever they used AI. Existing users were randomized within registration-phase strata; subsequent registrants were assigned through a fixed eight-registration cycle. Route compliance was observed but not enforced. A separate randomization (R3) concerns the October 2–4 final vote, in which registered users can rank up to three of twelve policy pillars. Eligible users receive no advice, comfort advice or outside-comfort-zone advice. Eligibility requires registration and qualifying activity before the September 20 closure, plus nonempty comfort and outside-comfort sets. Assignment occurs within eight prior-activity strata, with rerandomization to balance pillar-set membership. AI-framed recommendations are drawn uniformly from pillars linked to users’ prior activity or from the remaining pillars. Analysis prioritizes intent-to-treat estimates using prespecified samples; proposal outcomes conditional on submission are interpreted with selection caveats. Full design and analysis details appear in the updated pre-analysis plan.
Randomization Method Randomization done by computer. For users registered before the experimental start, assignment is performed by the research team, stratified by registration phase, with balanced allocation across the eight experimental cells within each stratum. For users registering during the experimental period, assignment is made on a rolling basis using a deterministic eight-registration cycle that preserves the same target proportions. All three randomized dimensions are assigned at the user level and remain fixed throughout the experimental period. ssignment is implemented by computer. For the proposal phase (June 25–September 20, 2026), the research team randomized previously registered users within registration-phase strata, balancing the eight R1 × R1.1 × R2 cells. Subsequent registrants were assigned through a deterministic eight-registration cycle preserving the same proportions. These assignments are at the user level and remain fixed throughout the proposal phase. R3 is randomized separately at the user level within eight prior-activity strata, targeting equal allocation to control, comfort advice and outside-comfort-zone advice. Candidate assignments are redrawn until each stratum meets its prespecified pillar-membership balance criterion
Randomization Unit Individual (user). All three randomized dimensions, proposal-development route, AI-facilitator mode, and authorship-disclosure tag, are assigned at the individual user level. For users registered before the experimental start, randomization is stratified by registration phase. Note: the realized sample size is not yet known, as it depends on platform adoption over the experimental window; the figures below are planned targets derived from ex-ante power calculations and will be updated once the realized sample is known. Individual (user). All four randomized dimensions, proposal-development route, AI-facilitator mode, authorship-disclosure tag and voting advice, are assigned at the individual user level. For users registered before the experimental start, randomization is stratified by registration phase. Note: the realized sample size is not yet known, as it depends on platform adoption over the experimental window; the figures below are planned targets derived from ex-ante power calculations and will be updated once data is received.
Planned Number of Clusters The realized number is not yet known, as it depends on platform adoption over the experimental window (June–October 2026); this field will be updated once the realized sample is known. The realized number is not yet known, as it depends on platform adoption over the experimental window and data not yet received.
Planned Number of Observations The realized number of individuals, proposals, and conversations is not yet known, as it depends on platform adoption over the experimental window; these figures will be updated once the realized sample is known. The realized number of individuals, proposals, and conversations is not yet known, as it depends on platform adoption over the experimental window and data not yet received.
Sample size (or number of clusters) by treatment arms The three randomizations are orthogonal, so each user contributes to all three contrasts simultaneously; the arm sizes above are not additive across randomizations. Realized arm sizes are not yet known and depend on platform adoption; these are planned targets and will be updated once the realized sample is known. The three randomizations are orthogonal, so each user contributes to all three contrasts simultaneously; the arm sizes above are not additive across randomizations. Realized arm sizes are not yet known and depend on platform adoption; these are planned targets and will be updated once the realized sample data is received.
Power calculation: Minimum Detectable Effect Size for Main Outcomes Ex-ante calculations assume significance α = 0.05, power = 0.80, minimum detectable effect = 0.20 standard deviations, within-individual correlation across proposals ρ = 0.50, and an average of two proposals per individual. For R1, accounting for two-sided non-compliance, calculations additionally assume a first-stage compliance rate κ = 0.80. Under these assumptions the planned per-arm targets are approximately 614 individuals (R1, participant level) and 393 individuals per arm (R1.1 and R2, participant level); lower per-arm requirements apply at the proposal/conversation level (approximately 460 for R1 and 295 for R1.1 and R2). All minimum detectable effects are expressed in standard-deviation units of the respective outcome. These calculations are ex-ante; realized power depends on platform adoption and will be revisited once the realized sample is known. Ex ante calculations assume significance α = 0.05, power = 0.80 and a minimum detectable effect of 0.20 standard deviations. Repeated-proposal calculations assume two proposals per individual and within-individual correlation ρ = 0.50; R1 additionally assumes a first-stage coefficient κ = 0.80. Per-arm planning requirements are approximately 614 individuals for R1 participant-level analyses and 460 for proposal-level analyses; 393 AI users for R1.1 primary first-conversation and participant-level analyses, and 295 for proposal-level robustness analyses; and 393 individuals for R2 participant-level analyses. For R3, user-level comparisons require approximately 393 users per arm, or 1,179 across three arms. This also provides an equal-variance benchmark for pillar-specific advice-versus-control comparisons. Under common variances and uncorrelated control-set averages, the pillar-specific comfort-versus-outside-comfort contrast requires approximately 785 users per arm. Actual power depends on sample size, outcome variances and, for pillar-specific comparisons, the variances and covariance of weighted control-set averages. These benchmarks concern unadjusted tests; multiple-testing adjustments affect power.
Intervention (Hidden) The AI facilitator is implemented as a two-step generative-AI pipeline: a chat assistant conducts a multi-turn conversation with the user, after which a separate proposal-generation model reads the completed conversation and produces a structured proposal that the user can manually revise before submitting. The two facilitator modes differ only in a middle "conduct" block of the chat prompt: the procedural arm follows a predefined sequence of phases (problem, context, proposal, synthesis, confirmation), while the open-ended arm conducts the conversation without a fixed sequence. The shared prompt components (role, opening, length and format constraints, system tags, and guardrails) are identical across arms. The self-authored route presents the same proposal fields via a template. The authorship-disclosure tag, when active, labels proposals as either AI-assisted or self-authored; when inactive, no label is shown. Crossover between routes is observed but not enforced and is made costly (switching requires abandoning the current draft or conversation). Full prompt text and implementation detail are provided in the uploaded pre-analysis plan. The AI facilitator combines a multi-turn chat with a separate proposal-generation model that produces a structured proposal users can edit before submission. The procedural and open-ended modes differ only in the chat prompt’s conduct block: the former follows problem, context, proposal, synthesis and confirmation phases; the latter has no fixed sequence. Other prompt components and the proposal-generation step are identical across arms. The self-authored route uses the same proposal fields. Route compliance is observed but not enforced; switching requires abandoning the current draft or conversation. Tagged viewers see “Creata con AIdea” or “Scritta in proprio”; untagged viewers see neither. A separate experiment during the October 2–4 final vote on twelve policy pillars assigns eligible users to no advice, comfort advice or outside-comfort-zone advice. Recommendations are presented as coming from AI and drawn uniformly from pillars linked to the user’s prior platform activity or from the remaining pillars, respectively. Assignment is stratified by prior activity, with rerandomization to balance pillar-set membership. Full prompts, eligibility rules and implementation details are provided in the updated pre-analysis plan.
Secondary Outcomes (End Points) Secondary outcomes, by randomization: (1) Proposal-development route: self-assessed expected impact and approval of the proposal; off-platform mobilization (whether a user organizes at least one event); a platform engagement index based on login frequency; and the conversation-level outcomes listed below. (2) Facilitator mode (procedural vs. open-ended): satisfaction with the AI-assisted experience and an open-text report of what the assistant was most useful for (collected from AI-assisted users at submission). (3) Conversation-level (analyzed where applicable): conversation length and how it concludes; breadth of policy-reasoning components covered and user initiative; and the extent of post-generation editing of the generated proposal. We additionally report secondary estimands alongside the primary intent-to-treat analyses, including a complier-average (instrumental-variables) effect for the proposal-development route. The primary/secondary partition is preliminary and will be finalized once the realized sample is known. Secondary Outcomes (Explanation) Secondary outcomes, by randomization: 1. Proposal-development route: self-assessed expected impact and approval; similarity to Italian political and programmatic texts; ideological extremism and related statistics; within-topic semantic dispersion; event organization; and logins per active week. 2. Facilitator mode: satisfaction with the AI-assisted experience and an open-text report of what the assistant was most useful for, collected from AI users at submission. 3. Authorship-disclosure tag: whether users open or submit proposals through the AI-assisted or self-authored route. 4. Final-vote advice: number of ranked choices submitted; submission of second- and third-place choices; placement of individual pillars first, second or third; and opening proposals linked to a pillar. For the proposal-development route, instrumental-variable estimates supplement the primary intent-to-treat analysis for user-level outcomes defined for all assigned participants. IV analyses conditional on submission are exploratory. Post-generation editing is reported as a descriptive diagnostic.
Secondary Outcomes (Explanation) The engagement index is constructed from logins per active week. Off-platform mobilization is an indicator for organizing at least one online or in-person event linked to a proposal. Self-assessed impact and approval are measured via fixed-response survey items at submission. Conversation-level outcomes (length, conclusion type, policy-reasoning breadth, user initiative, and editing magnitude) are constructed as described under the primary-outcome explanation. Satisfaction is a fixed-response item; the usefulness report is free text. Detailed operational definitions are in the uploaded pre-analysis plan. The engagement index is logins per active week. Mobilization indicates organizing at least one online or in-person event linked to a proposal. Expected impact, approval and AI satisfaction use fixed-response survey items at submission; the usefulness report is free text. Political-text similarity is assessed by LLM evaluators against Italian party programs and policy initiatives. Ideological extremism is the distance from the CHES scale midpoint of the average position of parties coded as supporting a proposal. Related statistics include average ideological position and the shares supported by all or no parties. Within-topic dispersion averages semantic distances within treatment-arm and CAP-topic cells containing at least five proposals. Route-choice outcomes indicate opening or submitting through each route. Final-vote outcomes count ranked choices (0–3) and indicate second-/third-place choices, each pillar’s rank-specific placement, and opening linked proposals. Voting and browsing indicators equal zero when no corresponding action occurs. Detailed definitions appear in the updated pre-analysis plan.
Back to top