Human Reliability and Adjudication Completion Checklist
This is the final milestone gate for human inter-annotator reliability and adjudication. It is intentionally stricter than designed, partial, or awaiting-adjudication status. A case, source language, task layer, and cohort may support a publication-ready human reliability claim only when every relevant artifact, authority boundary, publication disclosure, and repository validation check below passes.
The checklist is cohort-scoped. Do not use one complete cohort to certify a different case, language, task layer, sample, packet version, codebook version, or coder population. Incomplete cases and languages remain draft disclosures, not reliability evidence.
Cohort identity and scope gate
Training and calibration gate
Sample and blind packet gate
Two-coder submission gate
Metrics and reference-comparison gate
Disagreement, adjudication, and codebook gate
Protected-path and authority gate
Publication package gate
Repository validation gate
All four commands must pass in the same repository state used for the completion decision:
npm run statusnpm run validatenpm run pipelinequarto render
These commands are required even when the human reliability state is honestly blocked. A blocked human reliability checklist with passing repository validation means the repository reports the limitation correctly; it does not mean the human reliability study is complete.
Completion meaning
The milestone gate is complete only when every checked item above is satisfied for the cohorts being claimed as reliable. It is blocked or incomplete when training, calibration, samples, two qualified primary coders, validated submissions, metrics, reference comparison, disagreements, required adjudication, codebook notes, protected-path evidence, publication updates, or repository validation are missing or invalid.
Current publication materials report no complete human reliability cohort rows. Until that changes, the workbench may disclose human reliability architecture and status, but it must not make a publication-ready human reliability claim for any case, language, layer, or cohort.