Earning while learning: How to run batched bandit experiments

Refereed Journal // 2026
Refereed Journal // 2026

Earning while learning: How to run batched bandit experiments

Researchers typically collect experimental data sequentially, allowing early outcome observations and adaptive treatment assignment to reduce exposure to inferior treatments. In this article, we review multiarmed bandit adaptive experimental designs that balance exploration and exploitation. Because adaptively collected experimental data through bandit algorithms violate standard asymptotics, inference is challenging. We implement an estimator that yields valid heteroskedasticity-robust confidence intervals in batched bandit designs and compare coverage in Monte Carlo simulations. We introduce bbandits for Stata, a community-contributed package for designing experiments via simulation, running interactive bandit experiments, and implementing and analyzing adaptively collected data. bbandits includes three common assignment algorithms—ε-first, ε-greedy, and Thompson sampling—and supports estimation, inference, and visualization.

Kemper, Jan and Davud Rostam-Afschar (2026), Earning while learning: How to run batched bandit experiments, Stata Journal 26(3)

Authors Jan Kemper // Davud Rostam-Afschar