Corpus Limitations and Future Expansion

V1 prioritizes public-domain and open-access materials, but may also use rights-reviewed gitignored-local texts for scholarly analysis when the rights policy permits local storage and public artifacts do not reproduce restricted text. Missing, inaccessible, uncertain, or locally restricted corpora should be documented here.

Frozen Corpus Boundary

The expanded pre-v1 corpus boundary is frozen as of 2026-06-28. The verified working corpus now contains six cases and 41 manifest documents:

  • Lincoln: 5 documents.
  • American Revolution: 11 documents.
  • Napoleon: 11 documents.
  • Hitler: 8 documents.
  • French Revolution: 2 documents.
  • British World War I: 4 documents.

The pre-v1 expansion window admitted the two new cases and targeted existing-case repairs selected in the expansion matrix. It is now closed for corpus growth before v1. Wholesale redownload, starter-corpus replacement, and opportunistic source growth remain out of scope. Reopening acquisition is appropriate only for a document-specific defect inside the frozen boundary, such as failed corpus verification, damaged or misleading OCR, missing provenance, a rights-status change, or a source-quality problem that would materially affect annotation.

The pre-v1 expansion candidate matrix and rights/provenance review are recorded in docs/corpus/pre-v1-expansion-candidate-matrix.md and docs/corpus/pre-v1-expansion-rights-provenance-review.md. Deferred candidates from that review, including Imperial Japan, Wilson WWI, Stalin WWII, Mao/CCP, Saint-Just additions, Napoleon’s Farewell to the Old Guard, and Hitler support material, are future-work candidates. They are not part of the frozen pre-v1 corpus and should not enter the v1 reliability or claim-promotion cycle.

Known source limitations remain part of the v1 record. Napoleon is normalized from Gallica OCR and should preserve OCR uncertainty where it affects French MIPVU decisions. Hitler source-language review depends on rights-reviewed local/fair-use German texts with English glosses used only as analytical aids.

Multi-model reliability limits

The multi-model stress test does not widen the corpus’s publication rights. Packet payloads and raw responses inherit each source’s storage and third-party-transmission constraints. Local-only, restricted, metadata-only, or unresolved spans must remain local or be withheld; public summaries should prefer stable IDs, aggregate diagnostics, provenance metadata, and short compliant spans.

Accepted annotations and prior human-review artifacts remain separate from model diagnostics. AI agreement cannot establish human inter-annotator reliability, prove that an interpretation is correct, or demonstrate scholarly reproducibility. A publication-grade human reliability claim would require a separately designed blind human coding study with its own sample, metrics, adjudication, and limitations.