AI for Academic
Research MethodologyJune 25, 2026

Sample Size Calculation with Claude: When to Trust the Output

A sample-size answer can be arithmetically correct and scientifically wrong. The usual failure occurs before calculation: the wrong design family, unit of analysis, estimand, or variance assumption has been selected.

Claude can help expose the required inputs and explain how they affect recruitment. It should not be the only calculator, and its choice of formula should never be accepted without independent verification.

Start with the design narrative

Do not begin with a list of percentages and a request for n. Supply the primary objective, design, unit of allocation, unit of analysis, endpoint type and time point, effect measure, allocation ratio, repeated or clustered structure, and planned analysis.

Ask the model to return four things before any arithmetic:

  1. the proposed calculation framework and why it matches the estimand;
  2. every required input and its source;
  3. any design feature that requires inflation or adjustment;
  4. the missing decisions that prevent a valid calculation.

If a parameter has no empirical or published basis, label it as an assumption and plan a sensitivity analysis.

Three Design Types and Where It Goes Wrong

Two-arm parallel designs

Even a familiar two-arm design requires choices about endpoint scale, superiority versus non-inferiority, effect measure, allocation, sidedness, multiplicity, attrition, and the analysis model. A generic two-proportion or two-mean formula may not match an adjusted primary analysis.

Cluster trial

Clustering reduces the effective information contributed by each participant. Under a simple equal-cluster approximation, the design effect is 1 + (m − 1) × ICC, where m is mean cluster size. Real calculations may also need the number and variability of cluster sizes, cluster-level attrition, and a method appropriate to a small number of clusters. Applying a single inflation factor is not universally sufficient.

Diagnostic accuracy study

Diagnostic studies may be sized for precision of sensitivity or specificity, comparison of tests, or another primary objective. The required number of participants is not the same as the required number with and without the target condition; prevalence therefore affects recruitment. State the target measure and precision explicitly.

Use the model for explanation and sensitivity work

After the framework has been approved, a model can help create a sensitivity grid across plausible effect sizes, ICC values, attrition rates, or prevalence assumptions. It can also translate the calculation into a methods paragraph.

Those outputs remain drafts. Compare the generated code and numbers with dedicated software, published formulas, or a statistician’s calculation. Preserve the software version, command, and assumptions so the result can be reproduced.

Before approval, confirm the unit of randomization and analysis, primary estimand, parameter sources, adjustments for clustering or repeated measures, multiplicity, attrition, and feasibility. Resolve disagreement between methods rather than averaging two answers.

Report enough to reproduce the calculation

The CONSORT–SPIRIT guidance makes sample-size reporting part of a reproducible trial record. A methods paragraph should name the calculation framework or software, version, primary outcome, target effect, variance or baseline event-rate assumption, alpha, power or precision, allocation ratio, and every inflation factor. Cite the source for externally derived parameters. Report the calculated analyzable sample separately from the final recruitment target after attrition.

For cluster designs, include the assumed ICC and cluster-size information. For diagnostic studies, distinguish the total sample from the required numbers with and without the target condition. These details let a reviewer reconstruct the logic instead of judging a lone number.

The framing problem that precedes calculation — whether your study is powered to answer the right question at all — is covered in When a Study Is Too Small to Matter. The same attention to assumption documentation applies to your analysis plan: see Statistical Assumption Check with Claude for the equivalent audit at the results stage.

AI for Academic’s workspace can help stress-test the methods narrative before submission. Treat any resulting warning as a review prompt, then verify the calculation in the original software or with a statistician.

Sample Size Calculation with Claude: When to Trust the Output | AI for Academic