Papers › Blurring cluster randomized trials and observational studies using Two-Stage TMLE to...
Blurring cluster randomized trials and observational studies using Two-Stage TMLE to address sub-sampling, missingness, and minimal independent units
Joshua R. Nugent, Carina Marquez, Edwin D. Charlebois, Rachel Abbott, Laura B. Balzer
The archive published only this paper's code-link row. Authors, date and abstract are from arXiv's metadata (CC0), read from the Kaggle arXiv metadata snapshot of 2026-09-12 where its title matched the archive's; the title is the archive's.
Cluster randomized trials (CRTs) often enroll large numbers of participants, but due to logistical and fiscal challenges, only a subset of participants may be selected for measurement of certain outcomes, and those sampled may, purposely or not, be unrepresentative of all participants. Missing data also present a challenge: if sampled individuals with measured outcomes are dissimilar from those with missing outcomes, unadjusted estimates of arm-specific outcomes and the intervention effect may be biased. Further, CRTs often enroll and randomize few clusters by necessity, limiting statistical power and raising concerns about finite sample performance. Motivated by a sub-study of the SEARCH community randomized trial on the incidence of TB infection, we demonstrate interlocking methods to handle these challenges. First, we extend Two-Stage targeted minimum loss-based estimation (TMLE) to account for three sources of missingness: (1) sampling for the sub-study; (2) measurement of baseline status among those sampled, and (3) measurement of final status among those in the incidence cohort (i.e., persons known to be at risk at baseline). Second, we critically evaluate the assumptions under which sub-units of the cluster can be considered the conditionally independent unit, improving precision and statistical power but also causing the CRT to behave more like an observational study. Our application to the SEARCH highlights the impact of different assumptions on measurement and dependence as well as the real-life gains of our approach for bias reduction and efficiency improvement.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Results from the paper archive 2025-07-28
No leaderboard rows for this paper in the archive.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections