Research Design

This document is the publication-facing research design for The Sacrifice Law Workbench. It defines the project aim, research questions, corpus boundaries, units of analysis, method sequence, validation strategy, AI-use policy, data availability policy, and expected outputs that govern the scholarly pipeline.

The design is intentionally evidence-facing. Koenigsberg’s Law of Sacrifice is treated as a theoretical claim to be assessed, revised, limited, complicated, or rejected by the corpus and historical evidence. The project does not treat the Law as proven in advance.

Project Aim

The workbench assesses the degree to which recurring metaphor systems in leader-centered political corpora support, complicate, or limit Koenigsberg’s Law of Sacrifice, its body-politic corollary, and the construction of enemies as death.

The project combines:

  • rights-aware corpus construction;
  • MIPVU-style metaphor-related lexical-unit identification;
  • Conceptual Metaphor Theory source-target mapping;
  • corpus-assisted discourse analysis;
  • rhetorical and genre analysis;
  • systematic absence and agency analysis;
  • historical enactment or alignment review;
  • four-dimension evidentiary support scoring;
  • cross-case comparison with explicit moral-equivalence guardrails.

The central methodological commitment is staged interpretation. The pipeline must not move directly from striking language to broad claims about sacrifice, violence, national fantasy, redemption, purification, or enemy destruction. Major claims should pass through an inspectable chain from source text to annotation, mapping, cluster, support dimension, historical corroboration where claimed, and synthesis.

Primary Research Question

The working publication-level question is:

To what degree do recurring metaphor systems in leader-centered political corpora, when compared with historically documented practices of mobilization, killing, dying, purification, and enemy destruction, support, complicate, or limit Koenigsberg’s Law of Sacrifice, the body-politic corollary, and the construction of enemies as bringers of death across the initial pilot cases defined in the case-selection protocol?

This question is deliberately framed as an assessment of evidentiary support. It asks what the selected corpora and corroborating historical record can show. It does not ask the corpus to prove broad historical causation, diagnose private mental states, or explain war and genocide as monocausal outcomes of metaphor.

Secondary Research Questions

  1. What metaphor-related lexical units recur across the selected corpora, and which conceptual mappings do they support?
  2. Which collective objects are represented as sacred, ultimate, transcendent, immortal, or worth dying or killing for?
  3. How are bodies represented as offerings, wounds, instruments, extensions of the collective, contaminants, parasites, martyrs, debts, or redeemed matter?
  4. How are enemies represented as death-bearing threats to the sacred object: agents, carriers, embodiments, or signs of mortality, doubt, dissolution, unreality, contamination, enslavement, desecration, or destruction?
  5. What forms of violence are made morally necessary, obligatory, defensive, redemptive, purifying, providential, or historically inevitable?
  6. How do metaphor systems vary by case, phase, genre, register, audience, and rhetorical occasion?
  7. What absences or suppressions matter: civilian suffering, enslaved people, enemy subjectivity, perpetrator agency, coercion, contingency, or failed sacrifice?
  8. Where do textual-symbolic patterns align with historically documented practices, policies, institutions, mobilizations, or outcomes?
  9. Where do the corpora resist, complicate, limit, or fail to support Koenigsbergian expectations?
  10. What differences distinguish war-oriented rhetoric from genocidal rhetoric in enemy construction, body imagery, purification logic, sacrifice logic, and imagined exit conditions?

Corpus Scope

The project uses leader-centered political corpora as an auditable entry point into symbolic systems that organize sacred objects, sacrificial obligations, enemy construction, and political violence. Leader-centered does not mean leader-caused. It means that public and semi-public texts attributed to historically consequential political actors provide a bounded corpus for methodical comparison.

V1 uses two corpus layers:

  • Balanced core: high-priority texts selected for comparability across case phases, genres, registers, and rhetorical settings. Cross-case frequency and distribution claims should identify whether they use the balanced core.
  • Extended corpus: additional texts that support source discovery, close reading, hypothesis generation, and case-local interpretation. Extended material may support case-level claims when its status is disclosed, but it should not silently control cross-case comparison.

The first project phase is a pilot stage. Its purpose is to test whether the method is usable, auditable, and capable of distinguishing strong, moderate, weak, complicated, and unsupported case patterns. Expansion to additional cases should revise the method where the evidence requires it.

Pre-v1 case or document expansion is governed by the expansion rubric in case-selection.qmd. Additions must have a documented selection rationale, rights and provenance review, language or translation policy, expected comparative gain, and an explicit freeze point before reliability and publication promotion resume.

Inclusion Criteria

A source may enter the analytical corpus when it has:

  • a stable document identifier;
  • clear title, date, date precision, genre, register, phase, authorship, and source metadata;
  • an explicit inclusion rationale tied to one or more research questions;
  • documented rights status and storage decision;
  • provenance sufficient for citation or reproducible retrieval;
  • verification expectations where feasible, such as minimum word count and anchor phrases;
  • risk flags for translation, OCR, authorship, excerpting, editorial mediation, or source uncertainty;
  • case relevance for sacred object, sacrificial body, enemy as bringer of death, historical enactment or alignment, rhetorical setting, or negative evidence.

Exclusion Criteria

A source should be excluded from the committed analytical corpus, or preserved only as metadata, when:

  • rights status does not permit the intended storage or reuse;
  • provenance is too unclear for scholarly citation;
  • authorship or attribution risk is too high for the intended claim;
  • the text is too fragmentary, corrupted, or mediated to support reliable annotation without special handling;
  • the source duplicates another text without adding edition, register, audience, or phase value;
  • inclusion would distort the balanced core without an explicit extended-corpus rationale;
  • the source is useful only as background context rather than as analyzable corpus evidence.

Exclusion is not a scholarly failure. Excluded, unavailable, metadata-only, and rights-blocked sources should be documented when they materially affect corpus coverage or interpretation.

Units Of Analysis

The workbench preserves multiple units because different claims require different evidence:

Unit Role
Case Historical-political corpus under study
Corpus layer Balanced-core or extended-corpus scope
Document Stable bibliographic and historical unit
Section or paragraph Rhetorical movement and local context
Sentence Citable annotation context
Lexical unit MIPVU decision unit
Source span Textual evidence for CMT or interpretation
Conceptual mapping Source-target relation and entailments
Cluster Recurring metaphor-system grouping
Historical corroboration note Contextual alignment with documented practice or outcome
Support rating Document-level or case-level dimension score
Claim Publication-facing statement classified by evidentiary status

Every major claim should be traceable through the relevant units. A typical audit chain is:

claim
  -> claim status
  -> support dimension and score
  -> cluster
  -> conceptual mapping
  -> MIPVU lexical-unit IDs
  -> sentence IDs
  -> document metadata
  -> source text
  -> historical corroboration where claimed

Methodological Frameworks

The project uses each method for a specific task:

Framework Role
Rights-aware corpus construction Controls what text can be stored, cited, and reused
MIPVU Identifies metaphor-related lexical units
Historical semantics Controls for period-appropriate meanings and semantic shift
Conceptual Metaphor Theory Interprets source-target mappings and entailments
Corpus-assisted discourse analysis Examines frequency, distribution, concentration, collocation, co-occurrence, register, genre, and chronology
Critical Metaphor Analysis Interprets persuasive and ideological function
Rhetorical criticism Interprets occasion, audience, genre, posture, and action
Systematic absence and agency analysis Records who acts, who suffers, who is displaced, and who is excluded
Historical enactment or alignment review Tests whether textual-symbolic patterns align with documented policies, practices, implementation, social uptake, or outcomes
Koenigsbergian analysis Assesses support for sacred object, sacrificial body, enemy as bringer of death, and body-politic claims
Comparative analysis Compares cases while preserving historical specificity and moral-equivalence cautions

MIPVU identifies metaphor-related words. It does not identify conceptual metaphors, ideological fantasies, or historical causation. CMT, corpus analysis, rhetorical criticism, historical corroboration, and Koenigsbergian synthesis occur downstream.

Annotation Sequence

Annotation proceeds in staged layers:

  1. Corpus and segmentation review: confirm metadata, rights status, source-language policy, stable sentence IDs, and document provenance.
  2. Historical semantics preparation: identify difficult terms and period meaning risks before metaphor decisions harden.
  3. MIPVU lexical-unit decisions: record contextual meaning, basic meaning, contrast, comparison basis, confidence, source citation, and notes for each metaphor-related or uncertain lexical unit.
  4. Reliability and adjudication: double-code a defined sample, classify disagreements, revise the codebook, and preserve an adjudication log.
  5. CMT mapping: link mappings only to MIPVU-positive or uncertain lexical units unless explicitly marked exploratory.
  6. Corpus-assisted analysis: examine distribution, concentration, collocation, co-occurrence, register, genre, and diachronic patterning.
  7. Critical metaphor and rhetorical analysis: interpret persuasive function, audience, occasion, emotional posture, and authorized action.
  8. Absence and agency analysis: record expected presences, absences, named agents, patients, beneficiaries, sacrificial subjects, exclusions, and displacement mechanisms.
  9. Historical enactment or alignment: document whether textual-symbolic patterns align with established practices or outcomes.
  10. Support scoring and synthesis: rate sacred object, sacrificial body, enemy as bringer of death, and historical enactment/alignment before writing Koenigsbergian synthesis.

Interpretive layers should not be used to backfill the identification layer. If a later interpretation depends on a weak or uncertain lexical decision, the uncertainty must remain visible.

Evidentiary Support Standard

The project uses four support dimensions:

Dimension Guiding question
Sacred object Is a collective object represented as transcendent, ultimate, sacred, immortal, or worth dying/killing for?
Sacrificial body Are bodies represented as meaningful offerings, payments, instruments, victims, martyrs, or material through which the sacred object is preserved or made real?
Enemy as bringer of death Is the enemy represented as the agent, carrier, embodiment, or sign of death, doubt, mortality, dissolution, unreality, or collective destruction?
Historical enactment or alignment Do textual-symbolic patterns align with historically documented practices, policies, institutions, mobilizations, implementation, social uptake, or outcomes?

Each dimension uses the anchored 0-4 support scale defined in PRIMARY_RESEARCH_QUESTION.md. Scores are not probabilities and should not be treated as proof values. They are structured support judgments that preserve comparability, uncertainty, and the reasoning behind each rating.

Support is strongest when textual-symbolic evidence and historical alignment converge. A case may support a strong textual reading while still providing only moderate or weak support for Koenigsberg’s Law if historical enactment or alignment is thin, uncorroborated, or uncertain.

Validation Strategy

Validation is layered:

  • Source validation checks rights status, provenance, storage decision, and post-acquisition file integrity.
  • Build and schema validation checks JSON structure, sentence IDs, annotation references, confidence values, cluster IDs, and case outputs.
  • MIPVU validation checks lexical-unit decisions and required rationale fields for metaphor-related or uncertain units.
  • Mapping validation checks that CMT annotations cite valid MIPVU evidence.
  • Claim validation checks whether public claims have traceable support and are classified as evidence, interpretation, inference, speculation, or open question.
  • Support-rating validation checks dimension scores, recurrence, historical corroboration, document weights, and historical-alignment caps.
  • Publication validation checks that public pages do not present draft, exploratory, or placeholder artifacts as completed findings.

The normal full pipeline/publication gate is:

npm run status
npm run pipeline
npm run validate
quarto render

Narrow changes may use smaller validation, but skipped gates should be stated when reporting the work.

Claim Status Policy

Every publication-facing claim should be classified:

Status Definition
Evidence Directly observable source text, metadata, annotation, or pipeline output
Interpretation An argued reading of evidence using CMT, Koenigsbergian theory, corpus analysis, or close reading
Inference A cautious claim that follows from multiple pieces of evidence but is not directly observed
Speculation A possibility worth preserving but not supported enough to function as a finding
Open question A live uncertainty requiring more sources, annotation, validation, corroboration, or theoretical work

Claims should be revised, demoted, split, or removed when the evidence chain is missing. The preferred formulation is “supports, complicates, or limits,” not “proves.”

AI-Use Policy

AI tools may assist the research process, but they are not independent scholarly authorities and are not evidence. Human review remains responsible for metaphor-identification decisions, conceptual mappings, source interpretation, support ratings, historical corroboration, and publication claims.

Task Permitted AI role Human responsibility
Text preparation Assist with formatting, cleanup, segmentation review, and consistency checks Verify against the source and rights policy
Metadata extraction Suggest candidate metadata and risk flags Confirm from source records, manifests, archives, or scholarship
Candidate metaphor review Suggest possible metaphor-related units or ambiguous terms Apply MIPVU and record uncertainty
Historical semantics Suggest reference needs or terms requiring control Confirm with period sources or scholarship
CMT mapping Suggest possible source-target mappings and entailments Decide, justify, and record rival readings
Corpus summaries Draft summaries of validated artifacts Check against concordance, annotations, and analysis outputs
Support scoring Flag possible score evidence or inconsistencies Assign, justify, and review dimension scores
Multi-model stress testing Produce blind, bounded annotation responses for diagnostic comparison Keep accepted annotations authoritative, inspect disagreements, and authorize any later correction
Writing Draft or revise prose Preserve claim status, citations, and human scholarly responsibility

Required AI-use disclosure:

AI tools were used to assist with text preparation, candidate identification,
organization of annotations, consistency checks, and drafting support. All
metaphor-identification decisions, conceptual mappings, interpretive claims,
support ratings, historical corroboration judgments, and final scholarly
conclusions were reviewed and authorized by the researcher. AI-generated
suggestions were treated as provisional and were not used as independent
evidence.

If a publication venue requires a different disclosure format, preserve this substance while adapting the wording.

Multi-model stress testing is a distinct diagnostic layer. Accepted annotations, prior human review and adjudication artifacts, AI-model comparisons, and any separately designed human inter-annotator reliability study must be reported independently. Agreement among AI models does not prove that an annotation is correct, establish human reliability, or demonstrate scholarly reproducibility.

Data Availability Policy

Data availability depends on rights status and scholarly risk:

Artifact type Availability principle
Public-domain or open-access raw corpus text May be committed or distributed when source registry and rights policy permit
Copyrighted, access-restricted, or fair-use-only text Keep gitignored-local, metadata-only, or unavailable according to the rights policy
Source registry and document manifest metadata Publish when it does not reproduce restricted text and does preserve provenance
Normalized and segmented corpus derivatives Publish only when underlying rights status permits
MIPVU worklists and annotations Publish when they do not reproduce restricted text beyond permitted short spans or metadata
Multi-model packets and raw responses Apply the underlying source’s rights and storage policy; restricted or local-only spans remain local or withheld
Multi-model aggregate diagnostics Publish when they use rights-safe identifiers, measures, provenance metadata, and short compliant spans
Analysis, support scores, validation outputs, and public prose Publish with citations, evidence boundaries, and limitations
AI prompts and workflow notes Publish when they do not contain restricted source text or unsupported claims

The project should prefer reproducible metadata, citations, short compliant quotations, and paraphrase over reproducing long copyrighted source passages. When raw text cannot be shared, the data-availability statement should explain what is available, what is withheld, why it is withheld, and how a qualified reviewer could reconstruct the evidence trail from lawful sources.

Expected Outputs

The minimum scholarly upgrade should produce:

  • research design;
  • corpus register and source registry;
  • rights and data-availability documentation;
  • annotation codebook;
  • historical semantics notes;
  • MIPVU annotation files;
  • reliability report and adjudication log;
  • CMT mapping table or files;
  • corpus-assisted analysis outputs;
  • critical metaphor, rhetorical/genre, and absence/agency analysis artifacts;
  • historical enactment or alignment notes;
  • document-level and case-level support ratings;
  • Koenigsbergian evidentiary-support synthesis;
  • comparative analysis protocol and guardrails;
  • AI-use statement;
  • validation and audit package.

These outputs may move through draft, reviewed, finding, or deprecated status. Public pages should state status clearly and should not present structural scaffolds as completed findings.

Deferred Decisions

Some decisions remain open until the first complete annotated case and publication target clarify the project shape:

  • whether the first publication should be a cross-case methods article, a Lincoln-centered case study, or a theory-testing chapter;
  • final venue-specific AI-use disclosure wording;
  • final data-availability wording for mixed rights-status corpora;
  • exact artifact-readiness rules for promoting draft outputs to publication-facing findings;
  • final support-rating governance after the first scored case.

These decisions are tracked in OPEN_DECISIONS.md.