Model-Reliability Disclosure
Generated by scripts/generate-traceability-report.py on 2026-06-28T10:31:47.
The workbench keeps four forms of review evidence separate. They answer different questions and must not be combined into one reliability claim.
| Layer | Role | Publication authority |
|---|---|---|
| Accepted annotations | Canonical project references used by analysis after human review | May support claims when traceable to sources and appropriately qualified |
| Prior review artifacts | Existing human sampling, review, and adjudication records | Document prior quality control; they are not model outputs |
| Multi-model stress test | Blind comparison of AI model behavior against the same bounded tasks and project references | Diagnostic only; can identify instability and questions for human review |
| Human reliability study | Separately designed blind human double-coding and adjudication | Must be reported with its own sample, coders, metrics, and limitations |
No model-reliability command changes an accepted annotation. A model disagreement creates, at most, a review candidate. A human must inspect the source and codebook, make any authorized correction through the ordinary governance process, and then regenerate downstream artifacts.
Reproducibility boundary
Rebuildable packets, hashes, schemas, prompts, run metadata, and diagnostic reports make the stress-test procedure auditable and computationally repeatable. They do not make a substantive annotation, interpretation, or historical claim correct. Model agreement may reflect clear instructions, shared bias, common training data, prompt effects, or an easy item. It therefore does not prove scholarly reproducibility and cannot substitute for independent human review, source corroboration, or a human reliability study.
Rights and publication boundary
Packet payloads and raw model responses inherit each source document’s rights status and storage policy. Local-only, restricted, metadata-only, or unresolved source spans must not be uploaded to a hosted model or committed unless the rights record explicitly permits that use. Public reporting should prefer stable identifiers, aggregate diagnostics, provenance metadata, and short compliant spans. Withholding a restricted payload does not prevent publication of a rights-safe aggregate result, but the data-availability statement must say what was withheld and why.
See the stress-test methodology, current execution results, the human reliability methodology, and external review procedure.