How Will AI Agents Produce Research? Literature Context, Estimator Choice, and Specification Search in Autonomous Empirical Analysis”
Artificial intelligence agents have shown the capacity to autonomously conduct empirical research by selecting outcome variables, acquiring data, selecting identification strategies, and running econometric estimators end-to-end. How do AI agents behave when they conduct empirical research on their own, and what happens as they are given increasing discretion over methodological choices? I deploy approximately 600 agent sessions using a single pinned model on the same task, estimating the effect of minimum wage increases on teen employment from the same state-year panel, and randomising only the literature context each agent reads before working. Four pre-registered waves vary the content of the literature summaries and the set of permitted estimators. When discretion is narrow, and agents are restricted to a single heterogeneity-robust difference-in-differences estimator, estimates have a similar distribution across arms while written interpretations diverge: agents describe their results in the direction of whichever literature they were assigned. When a subsequent wave grants discretion over estimator choice, estimates diverge entirely through that choice. Fourteen of fifty negative-context agents report two-way fixed effects estimates; no agent in any other arm does, and agents who retain the robust estimator produce indistinguishable estimates across arms. Session transcripts reveal the mechanism, and it is sharper than selective searching: agents in every arm experiment with two-way fixed effects during their sessions. Agents assigned null-effects literature, or none, treat it as a robustness check and report the robust estimator every time. Negative-context agents, whose robust estimates typically contradict the literature they were assigned, instead promote two-way fixed effects to their reported estimate, and in later waves, they cite the result itself rather than diagnostics as their reason for revising at three times the rate of other arms. None of this is visible in the final code an agent delivers; only the transcripts reveal it. Requiring robust estimators closes the channel: a design lever available to anyone deploying AI agents for empirical work, and one worth understanding before they produce research at scale.
Join at imt.lu/sagrestia
Speakers
- Scott Cunningham, Baylor University
Unità di Ricerca
- AXES