Protocol

Preregister for the next LLM

LLM experiments are easy to p-hack, that is, tune after results are visible, by changing the model, prompt, etc. until a preferred result appears. This initiative instead exploits new model releases as a holdout set. It asks to preregister the full procedure and then run it on the next eligible model, thereby letting researchers provide a certificate against p-hacking the model that is used.

Reference: “Mitigating LLM-based p-Hacking by Preregistering for the Next LLM”
Maria Thomas, Kristina Gligorić, Nihar B. Shah (2026). arXiv:2606.27687.

Effectiveness of the protocol

73.9%Share of p-hacked analysis configurations blocked by the protocol in one evaluation task.
72.7%Share of p-hacked analysis configurations blocked by the protocol in another evaluation task.
6 of 7Of seven p-hacked configurations, six failed to reproduce the favorable result on the next model.
Researcher guide

Complete the AsPredicted template

Use the initiative’s AsPredicted template at AsPredicted.org to create a timestamped preregistration of the configuration and eligible models.

1. Data collected yet?

Select “No.” The chosen future model must still be unavailable, so confirmatory outputs cannot yet exist.

2. Hypothesis

Record the predicted result, any directional prediction, and the threshold for support. Define eligible families and variants, plus a selection rule that an independent reader can apply to future releases.

3. Dependent variable(s)

Provide a deterministic formula that converts acceptable model responses into the study outcome. Specify the required syntax, label set, parsing procedure, and normalization.

4. Conditions

Record every model-use setting: the full prompt, decoding controls, token limit, and whether execution is batched or item-by-item.

5. Analyses

Name the planned statistical procedures and the decision rules used to interpret them.

6. Exclusions

List every invalid-output case and the deterministic response to it. Set a model-level failure threshold, the switch-to-next-model rule, and any additional validity checks.

7. Sample size

Give the item count and identify the dataset and version precisely enough to reconstruct the analyzed sample.

8. Other

Precommit to any follow-up models and to rules for model failure, newly introduced model families, and discontinued families.

After preregistration is complete, use the next model released within the eligible class you specified for the confirmatory analysis.
Audit tool

Compare a preregistration and paper

The audit reports whether the paper uses the first eligible model released after preregistration, whether the reported configuration matches the preregistration, and whether the preregistration follows the template.

The preregistration PDF, paper PDF, and API key are sent directly to the selected model provider for this request. This page does not store them.

or
or