Back to History

Fields Changed

Registration

Field Before After
Trial Title Complementarities of AI-Enabled and Human Recommendation to Jobseekers: Experimental Evidence from Kenya AI, Human, and Hybrid Career Guidance in Kenya: Does Advice Inform or Activate?
Abstract Career guidance is one of the few scalable tools that targets the information frictions behind occupational mismatch in low- and middle-income countries, yet its welfare value depends on whether it works by informing jobseekers or by persuading them to act. We run a randomized controlled trial with about 4,000 jobseekers in coastal Kenya that compares human-only, AI-only, and hybrid AI-plus-human career guidance, and cross-randomizes whether AI support stops at a recommendation or adds persuasion and action support. Primary outcomes are employment and earnings, match quality, and persistence; we anchor short-run measures to long-run welfare with a surrogate index estimated in external panel data. The design uses AI as a research instrument, fixing the informativeness of advice and the strength of persuasion separately across arms in a way human counselling cannot. We ask whether AI substitutes for or complements human guidance, and whether guidance works by informing or by persuading. Career guidance may reduce occupational mismatch, but personalized human provision is costly to scale. I randomize 4,000 young Kenyan jobseekers to human-only, AI-only, hybrid, and control conditions, and separately randomize whether AI support stops at a recommendation or adds a behavioral-activation conversation. Because AI can be constrained, the activation contrast holds the recommendation fixed and measures what the subsequent conversation adds. This turns AI into a research instrument, making otherwise bundled features of counseling experimentally separable. Primary outcomes, measured six to nine months later, are employment, earnings, and match quality, defined as the utility of the job held under participants’ pre-treatment preferences over job attributes. An externally estimated surrogate index further projects longer-run earnings, job quality, and occupational persistence. The design tests whether AI substitutes for or complements scarce counselors, and whether activation changes valuation, reduces implementation frictions, or both. Substitution is assessed using a budget-equivalence criterion at delivery costs.
Last Published June 18, 2026 09:25 AM August 31, 2026 03:18 PM
Intervention (Public) Participants are offered structured career guidance, delivered through one of three channels depending on random assignment: a trained near-peer human mentor, an AI-enabled career-guidance tool (Compass), or a combination of the two. A control group receives no additional personalized guidance during the study period. Human mentors are drawn from Swahilipot Hub and the National Council of Churches of Kenya, and are recruited from the same coastal communities as participants. They provide one-on-one counselling: eliciting skills and preferences, discussing options, recommending a path, and supporting the jobseeker to act on it. Compass (developed by Tabiya) elicits a jobseeker's skills and preferences through a structured conversation, translates informal experience into a structured skill profile, and recommends jobs and career paths matched both to the jobseeker and to local labour demand. Within the AI-supported arms, the design also varies whether AI support stops once a recommendation has been delivered (information only) or continues with persuasion and action support (planning, encouragement, and follow-through). This separates the informational content of guidance from its persuasive, action-shifting content. Participants are offered structured career guidance, delivered through one of three channels depending on random assignment: a trained near-peer human mentor, an AI-enabled career-guidance tool (Compass), or a combination of the two. A control group receives no additional personalized guidance during the study period. Human mentors are drawn from Swahilipot Hub and the National Council of Churches of Kenya, and are recruited from the same coastal communities as participants. They provide one-on-one counselling: eliciting skills and preferences, discussing options, recommending a path, and supporting the jobseeker to act on it. Compass (developed by Tabiya) elicits a jobseeker's skills and preferences through a structured conversation, translates informal experience into a structured skill profile, and recommends jobs and career paths matched both to the jobseeker and to local labour demand. Within the AI-supported arms, the design also varies whether AI support stops once a recommendation has been delivered (information only) or continues with behavioral activation and action support (planning, encouragement, and follow-through). This separates the informational content of guidance from its behavioral activation content. The activation session is delivered by the AI directly to the jobseeker, in a conversation separate from any human counseling and never routed through the counselor. Counselors receive identical material for both subarms and are blind to which of their mentees receive the activation session.
Intervention End Date August 14, 2026 September 18, 2026
Primary Outcomes (End Points) Primary outcomes are pre-specified in families defined by domain and by measurement wave, across two waves. (Yet to be refined) Short-run primary (midline, ~4 weeks): 1. Information uptake and decision quality: preference clarity and belief accuracy (pre/post preference elicitation; calibration of subjective probabilities; choice uncertainty / second-order beliefs), consideration-set breadth and calibration, dominated-option avoidance, and search targeting (a search-focus / Herfindahl measure over occupations considered and applied to). 2. Job search and advice take-up: job-search effort and application yield, active vs. passive search (on/off-path search-help decision), information-bundle choice, implementation of the recommended path and action-plan completion, and the subjective probability of pursuing the recommended occupation. Longer-run primary (endline, ~6-9 months): 3. Economic opportunity: employment or income-generating activity (incl. a count of IGAs), earnings, hours worked, number of income sources, job applications, interviews, offers, and participation in training. 4. Match quality and persistence: skill-occupation alignment (KeSCO 2- and 3-digit), targeting of search toward a coherent occupation or sector, occupational/sectoral persistence over time, tenure and work-history length, and work satisfaction. 5. Welfare and subjective well-being: consumption, savings, perceived agency, and life satisfaction. Long-run targets: occupation/industry persistence (2-digit ISIC) and total real earnings ~10 years out, via a pre-specified surrogate index (Kenya Life Panel Survey). The two short-run families map to the two channels the design separates: information (uptake, decision quality, search focus) and persuasion (take-up, effort). They are primary both because the intervention targets them most directly (proximal behavioural responses) and because they identify the mechanism (RQ2). The endline families test whether these proximal effects translate into distal labour-market and welfare gains. Primary Outcomes (end points) All primary outcomes are measured at the endline (~6-9 months) to align with what the design can support. They consist of three pre-specified objects: 1. Labor-market index: Equal-weighted average of paid employment (income-generating work in the last seven days) and monthly earnings. 2. Quality-adjusted placement: The utility of the respondent's main occupation evaluated at their own baseline preferences. 3. Recommendation-directed action index: Averages applications in the target career family, fields searched, and training discipline (the primary outcome for the activation contrast). Long-run targets: Four-year job-quality index, log labor earnings, and longest continuous employment spell, projected via a pre-specified surrogate index using the Kenya Life Panel Survey (KLPS). Detailed construction and contrasts can be found in the uploaded PAP.
Primary Outcomes (Explanation) Each family is summarized by a standardized index (Anderson 2008), reducing each family to one primary test; across the family indices within a wave we report sharpened FDR q-values, and within each family disaggregated components use Romano-Wolf corrections. Corrections are applied within family and within measurement wave. The short-run (midline, ~4 weeks) and longer-run (endline, ~6-9 months) primary sets are pre-specified as distinct outcome sets: measured at different times, addressing distinct questions (proximal behavioural mechanism vs. distal labour-market welfare). We therefore correct within each set and do not apply a single joint correction across the midline and endline primary outcomes. This split is registered in advance, not chosen after seeing results. The surrogate index maps short-run proxies to the long-run targets (Athey-Chetty-Imbens-Kang 2024), under unconfoundedness, Prentice surrogacy, and cross-cohort comparability. Each family is summarized by an equal-weighted standardized index (Kling, Liebman, Katz 2007). Inverse-covariance weighting (Anderson 2008) is used as a robustness check. The registered index carries the within-family correction using Romano-Wolf step-down, with remaining disaggregated members reported beneath it unadjusted. Sharpened FDR q-values are reported alongside the Romano-Wolf p-values. Corrections are applied within family and not across families. Midline outcomes (~6 weeks) are secondary and descriptive. The surrogate index maps endline proxies to the long-run four-year targets (Athey-Chetty-Imbens-Kang 2024).
Experimental Design (Public) Individually randomized controlled trial with approximately 4,000 jobseekers. Six operational arms pool into four primary categories forming a 2×2 factorial (AI access × human counselling): Control, Human-only, AI-only, and AI+Human. Within the AI-supported arms, a secondary cross-randomization varies whether AI support stops at a recommendation (information only) or adds persuasion/action support. Primary estimands are intention-to-treat effects of the four pooled categories together with the factorial main effects (of AI and of human counselling) and their interaction (the substitutes-vs-complements test). The timing of AI introduction is separately randomized at the mentor level to diagnose counsellor learning and spillovers. Individually randomized controlled trial with approximately 4,000 jobseekers.Six operational arms pool into four primary categories forming a 2x2 factorial in AI access (Compass on or off) and human counseling (on or off): Control, Human-only, AI-only, and AI-and-Human. Within the two AI-receiving categories, a nested randomization varies whether AI support stops at a recommendation (recommendation only) or adds a behavioral-activation session. Primary estimands are intention-to-treat effects of the pooled active guidance offer vs. control, the factorial main effects of AI and human counseling, and the pooled activation contrast (recommendation-plus-activation vs. recommendation-only). The AI-by-human interaction (the substitutes-vs-complements test) is exploratory. The timing of AI introduction is separately randomized at the mentor level to diagnose counsellor learning and spillovers.
Randomization Method Randomization is carried out by a reproducible, seeded computer script, run after baseline data collection. The procedure combines blocking with rerandomization for covariate balance (Morgan and Rubin 2012, 2015) and is implemented batch-by-batch as recruitment proceeds, conditioning each batch on all previously randomized participants via cumulative-deficit apportionment of arm totals. The script blocks on geography (county cell), gender, and baseline labour-market state, then rerandomizes on pre-specified baseline predictors of outcomes and attrition, accepting an assignment only if it passes a calibrated, tiered Mahalanobis balance criterion. Confirmatory inference replays this exact mechanism (randomization inference). Randomization was carried out by a reproducible, seeded computer script, run after baseline data collection, and is now complete across five recruitment batches. The procedure combines blocking with rerandomization for covariate balance (Morgan and Rubin 2012, 2015) and was implemented batch by batch as recruitment proceeded, conditioning each batch on all previously randomized participants through cumulative-deficit apportionment of arm totals. The randomization is hierarchical. The first level assigns jobseekers in equal proportion (1:1:1:1) across the four primary categories: Control, AI-only, AI-and-Human, and Human-only. The second level applies a nested 1:1 split between recommendation and activation inside the two AI-receiving categories only. Collapsing the two levels yields six atomic cells in the marginal ratio 2:1:1:2:1:1. The script blocks on geography, gender, and baseline labour-market state, then rerandomizes on pre-specified baseline predictors of outcomes and attrition, accepting an assignment only if it passes a calibrated, tiered Mahalanobis balance criterion. Confirmatory inference replays this exact mechanism (randomization inference).
Randomization Unit The individual jobseeker. A separate, secondary randomization is conducted at the mentor level (timing of AI introduction: early vs. late, 1:1), used only to diagnose counsellor learning and spillovers, not to define the primary treatment effects. The individual jobseeker. A separate, secondary randomization is conducted at the mentor level, varying the timing of AI introduction (early versus late, 1:1). It is used only to diagnose counselor learning and implementation spillovers and does not define any primary treatment effect.
Planned Number of Clusters Individual randomization (not clustered): approximately 4,000 individuals. Guidance in the human-contact arms is delivered by approximately 80 near-peer mentors; inference clusters at the counsellor × session level. The secondary mentor-level randomization (AI-introduction timing) involves approximately 80 mentor clusters. Individual randomization, not clustered: approximately 4,000 individuals. Guidance in the human-contact arms is delivered by roughly 80 to 100 near-peer mentors. Because randomization is individual, no design-based clustering enters the power calculations. Model-based standard errors are clustered on the assigned counselor, with wild cluster bootstrap p-values, as a secondary diagnostic alongside the primary randomization-inference statements. The secondary mentor-level randomization of AI-introduction timing involves those same mentor clusters.
Sample size (or number of clusters) by treatment arms Approximately 4,000 jobseekers, assigned across six operational arms in a 2:1:1:2:1:1 ratio: - T0 Control: 1,000 - T1-R (AI-only, recommendation): 500 - T1-P (AI-only, recommendation + persuasion): 500 - T2-R (AI+Human, recommendation): 500 - T2-P (AI+Human, recommendation + persuasion): 500 - T3 Human-only: 1,000 Pooled into four primary categories of ~1,000 each: Control; Human-only (T3); AI-only (T1-R + T1-P); AI+Human (T2-R + T2-P). Approximately 4,000 jobseekers, assigned across six operational arms in a 2:1:1:2:1:1 ratio: - T0 Control: 1,000 - T1-R (AI-only, recommendation): 500 - T1-A (AI-only, recommendation + behavioral activation): 500 - T2-R (AI and Human, recommendation): 500 - T2-A (AI and Human, recommendation + behavioral activation): 500 - T3 Human-only: 1,000 Pooled into four primary categories of approximately 1,000 each: Control; Human-only (T3); AI-only (T1-R and T1-A); AI-and-Human (T2-R and T2-A).
Power calculation: Minimum Detectable Effect Size for Main Outcomes Assuming 10% attrition, covariate adjustment, and an intra-cluster correlation of ~0.05 at the counsellor × session level (α = 0.05, power 0.80): pooled active-vs-control comparisons have a minimum detectable effect of approximately 0.13–0.16 SD; the factorial main effects of AI and of human counselling are more precise, ~0.10–0.12 SD (each compares ~2,000 vs. ~2,000); the AI×human interaction is powered for moderate departures from additivity. Effects are in standard-deviation units of each standardized (Anderson) outcome index. Assuming 10% attrition and covariate adjustment, with individual-level randomization (no design-based clustering) ($\alpha = 0.05$, power 0.80): the pooled guidance offer vs. control has a minimum detectable effect (MDE) of 0.096 SD (one-sided); the factorial main effects of AI and human counseling are more precise at 0.084 SD (one-sided); the pooled activation contrast is detectable at 0.118 SD (two-sided). The AI×human interaction is exploratory (MDE 0.167 SD). Effects are in standard-deviation units of each equal-weighted standardized outcome index (Kling, Liebman, Katz 2007).
Additional Keyword(s) AI in development, artificial intelligence, behavioral interventions, career guidance, digital labor markets, field experiment, human-AI complementarity, information and persuasion, information frictions, job matching, job search, job search assistance, Kenya, labor market frictions, low- and middle-income countries, occupational matching, public employment services, randomized controlled trial, recommendation, skills signaling, youth employment AI in development, artificial intelligence, behavioral interventions, career guidance, digital labor markets, field experiment, human-AI complementarity, behavioral activation, information, persuasion, information frictions, job matching, job search, job search assistance, Kenya, labor market frictions, low- and middle-income countries, occupational matching, public employment services, randomized controlled trial, recommendation, skills signaling, youth employment
Intervention (Hidden) The study evaluates multiple modes of delivering career guidance through a randomized controlled trial with four primary groups: (i) human-only mentoring, (ii) AI-only guidance, (iii) combined human + AI support, and (iv) a control group. The AI intervention (“Compass”) is an interactive, agent-based career guidance system that guides participants through a structured workflow. It elicits information on past experiences, identifies skills using a standardized taxonomy, supports CV generation, and provides personalized recommendations for career paths and jobs opportunities. The system logs user interactions, recommendations, and engagement metrics. Compass operates in three stages: (1) preference elicitation via an adaptive discrete-choice and best-worst-scaling instrument that recovers a random-utility preference profile; (2) recommendation and matching, scoring opportunities by the product of preference fit and a success/feasibility propensity grounded in Kenyan labour-market data; and (3) behavioural activation, which converts a recommendation into action via implementation intentions and light-touch commitment. Recommendation-only arms deliver stages 1–2 and stop; persuasion arms add stage 3. AI is used as a research instrument: it fixes the informativeness of the signal and the strength of persuasion separately across arms, which human counselling cannot. We additionally randomize the timing of AI introduction at the mentor level (early vs. late, 1:1) to diagnose within-mentor learning and spillovers. The human mentoring intervention consists of structured sessions delivered by peer case managers affiliated with the implementing partner. These sessions focus on career exploration, goal setting, job search strategies, and follow-up support. Mentors may provide accountability, encouragement, and context-specific advice, including referrals or networking support where available. In the combined treatment arm(s), participants receive both AI-based and human support. Mentors receive a report on what the mentee discussed with the AI to inform their own support to the mentee. The intervention is delivered over multiple interactions following baseline data collection. AI interactions occur via mobile devices, while human mentoring sessions take place in person through partner facilities. Engagement, take-up, and adherence to the intervention are tracked using both administrative and survey data. The study evaluates multiple modes of delivering career guidance through a randomized controlled trial with four primary groups: (i) human-only mentoring, (ii) AI-only guidance, (iii) combined human + AI support, and (iv) a control group. The AI intervention (“Compass”) is an interactive, agent-based career guidance system that guides participants through a structured workflow. It elicits information on past experiences, identifies skills using a standardized taxonomy, supports CV generation, and provides personalized recommendations for career paths and jobs opportunities. The system logs user interactions, recommendations, and engagement metrics. Compass operates in three stages: (1) skills and preference elicitation to create a detailed profile using state of the art AI-driven elicitation methods; (2) recommendation and matching, scoring opportunities by the product of preference fit and a success/feasibility propensity grounded in Kenyan labour-market data; and (3) behavioural activation, which converts a recommendation into action via implementation intentions and light-touch commitment. Recommendation-only arms deliver stages 1-2 and stop; activation arms add stage 3. AI is used as a research instrument: it fixes the informativeness of the signal and the strength of behavioral activation separately across arms, which human counselling cannot. The activation agent operates under fixed constraints. The formal skills profile, preference profile, ranked recommendation, and all factual claims are identical to the corresponding recommendation-only arm. The agent is instructed to remain truthful, to use probabilistic rather than guaranteeing language, to present trade-offs honestly, to respect the jobseeker's autonomy, and to avoid pressure toward easy-to-place options. The study involves no deception. I additionally randomize the timing of AI introduction at the mentor level (early vs. late, 1:1) to diagnose within-mentor learning and spillovers. The human mentoring intervention consists of up to six one-on-one sessions of roughly thirty minutes, which a mentor may combine into three sessions of roughly an hour, delivered by near-peer case managers affiliated with the implementing partners. Sessions cover career exploration, goal setting, job-search strategies, and follow-up support. Mentors may provide accountability, encouragement, and context-specific advice, including referrals or networking support where available. Mentors file a tracking form after each check-in through KoboToolbox. In the combined arms, participants receive both AI-based and human support. The counselor-facing dossier is restricted to three AI-generated items for each mentee who completes the Compass conversation: the skills profile, the preference profile, and the ranked recommendations. Counselors do not observe conversation logs, and the dossier is identical across the recommendation and activation subarms. The intervention is delivered over multiple interactions following baseline data collection. AI interactions occur via mobile devices, while human mentoring sessions take place in person through partner facilities. Engagement, take-up, and adherence to the intervention are tracked using both administrative and survey data.
Secondary Outcomes (End Points) (Yet to be refined) Additional midline mechanism measures, each its own family: incentivized revealed preferences and WTP (counselling-format WTP; stepping-stone vignette; stated usefulness and a recommendation-regret index) ; beliefs and expectations, extended (top-option subjective probabilities; pre/post preference-shift magnitude); decision-process robustness (stated-chosen consistency; choice stability under re-elicitation); time use and substitution (displaced leisure); agency and discouragement (locus of control, self-efficacy, fatigue, aspirations, controllability slope). Secondary treatment estimands: CACE; recommendation-vs-persuasion contrasts (T1-P vs. T1-R; T2-P vs. T2-R); "any-AI vs. control" and "any-human vs. control" pooled policy contrasts; counsellor-side mechanism outcomes. Secondary Outcomes (end points) Midline mechanism measures (~4-6 weeks): Career-direction index, advice take-up index, beliefs, agency, intentions, and delivery fidelity. Endline secondary outcomes: Allocation quality (misallocation gap, skill-occupation alignment), labor market search and training (job-search index), psychological well-being index, consumption, savings, and life satisfaction. Unsigned measures (hours, work multiplicity, reservation wage, activity type) are descriptive and enter no correction family. I also test a budget-equivalence criterion for cost-effectiveness (AI vs. Human). Secondary treatment estimands: Within-channel activation contrasts (T1-A vs. T1-R, T2-A vs. T2-R). The AI-by-human interaction (factorial superadditivity) is exploratory. Detailed construction and contrasts can be found in the uploaded PAP.
Back to top

Partners

Field Before After
Partner Website (URL) https://swahilipothub.co.ke/ https://ncck.org/
Back to top