A heterogeneous, contract-gated graph of an academic research process, laid out in five regions. A left column holds the paper corpus, knowledge graph and document store. GPT 5.6 SOL retrieves literature, extracts evidence and produces an evidence matrix. A provider-neutral separate review role, assigned outside the producing context, audits that matrix against the corpus snapshot at gate E1 before any human reads it, so fabricated or misattributed citations are caught at the point of entry; typed evidence gaps return to retrieval. Claude Opus 5 turns approved evidence into a hypothesis registry, research plan and experiment design. Human gates H1 to H4 guard scope, evidence adequacy, hypothesis approval and the frozen protocol, governance record and compute authorisation before implementation. The same review role, in a fresh reviewer invocation distinct from the later audit invocations, challenges the design before that freeze. Claude Sonnet 5 generates and executes code and produces experimental results. Versioned code, environment, data, run and raw-result artifacts reach GPT 5.6 Terra, which checks integrity and drives validation, reproducibility and statistical verification. The separate review role audits the statistical work and merged verification evidence at the V1 handoff to writing, recomputing from the raw results and code rather than from the reports alone. Typed defects return to retrieval, planning, design or execution, and every return path carries a bounded revision budget. Claude Fable 5 receives read-only evidence, plan, result and verification artifacts, then produces the manuscript, claim–evidence map and figure registry. The separate review role performs the M1 manuscript claim audit; unsupported claims return to synthesis. Only an M1 pass reaches the human final decision, whose explicit outcomes are release, revise, narrow, null-result or stop. A right column lists shared tools and configured model routing; the bottom strip shows artifact flow, human gates, agent challenge gates and reading guidance. Selecting any graph node focuses its incident edges, edge labels and directly connected neighbour nodes.

Pinch to zoom the diagram, or read the same workflow below.

Artifact-gated reference architecture · v5.2

Contract-gated agentic research

A role-separated research workflow in which models produce bounded artifacts, reviewers challenge evidence, and humans retain scope, protocol and release authority.

Claim boundary: this page is a reference architecture and executable gate-dependency prototype. It does not autonomously validate scientific truth.

1–2 · Retrieval and evidence

A scoped search protocol and corpus snapshot feed structured extraction. Evidence gaps, contradictions and unresolved citations return to retrieval.

3–5 · Hypothesis and protocol

Approved evidence supports falsifiable hypotheses, a research plan and a frozen protocol. Governance applicability, exclusions, stopping rules and analysis choices are recorded before execution.

6–7 · Execution and results

Implementation produces a code commit, environment lock, data manifest, run manifest and immutable raw results. Execution defects return to the implementation stage.

8–10 · Reproduction and verification

Role-separated checks cover computational reproduction, robustness, estimates, uncertainty, assumptions, multiplicity, failure accounting and residual limitations.

11–12 · Synthesis and manuscript

Writing receives read-only evidence, protocol, result and verification artifacts. A claim–evidence map and figure registry are audited before the human release decision.

Decision gates

H1 · ScopeResearch intent, constraints and governance applicability.
E1 · Evidence fidelityCitation–source fidelity, retraction status and fabricated-reference detection, audited against the corpus snapshot before H2.
H2 · EvidenceSearch coverage, source adequacy, contradictions and unresolved gaps.
H3 · HypothesisRelevance, novelty status, falsifiability and feasibility.
H4 · Protocol freezeOutcomes, exclusions, analysis, multiplicity, stopping rules, data-access and compute authorisation, resources and risk.
T1 / T2 / V1 / M1 · ChallengeDesign, integrity, verification evidence and manuscript claims are reviewed by single-shot, role-separated CLI invocations. Each decision binds the current artifact hashes to the captured prompt and response log. V1 recomputes from raw results and code, not from reports alone.
FINAL · Human decisionRelease, revise, narrow, null-result or stop. Every exit, including stop, writes a release manifest.

Loop bounds

Revision budgetEvery gate carries a maximum revision count. Once spent, the gate may only block or escalate, so no typed return can cycle indefinitely. Counts are recorded in the release manifest.
Reviewer separationEach challenge uses a fresh invocation; producer and reviewer identities determine the recorded separation level. Captured local provenance is not provider attestation or a guarantee of epistemic independence.

Required audit trail

Passing a gate requires registered, versioned artifacts rather than a textual assertion alone.

problem_spec · search_protocol · corpus_snapshot · kg_snapshot · evidence_matrix hypothesis_registry · design_protocol · frozen_protocol · governance_record code_commit · environment_lock · data_manifest · run_manifest · raw_results reproduction_report · statistical_report · verification_report claim_evidence_map · figure_registry · manuscript · release_manifest

Academic research as a contract-gated agentic graph

  1. The paper corpus, the knowledge graph used for retrieval-augmented generation, and the document store feed GPT 5.6 SOL, the evidence and retrieval agent.
  2. SOL runs unit 1, literature retrieval and evidence extraction — search and retrieve, extract evidence, citation mapping — and writes notes back to the knowledge graph and document store.
  3. Unit 1 produces unit 2, the evidence matrix: structured evidence, claims, methods, results, contradictions and gaps. Challenge E1 audits that matrix against the versioned corpus snapshot for citation–source fidelity, quotation accuracy, retraction status, inclusion and exclusion compliance and fabricated references, so a separately identified reviewer can catch extraction error before downstream reasoning depends on it. An evidence-gap return reopens unit 1 instead of silently advancing incomplete evidence.
  4. Human gate H1 fixes problem scope and records ethics, legal and data-governance applicability before retrieval; H2 guards the evidence-matrix handoff to Claude Opus 5 and cannot open until E1 passes; H3 approves the hypothesis registry before planning; and H4 freezes the experiment protocol before implementation. H4 covers outcomes, exclusions, analysis and multiplicity plans, stopping rules, governance, data-access and compute authorisation, resources and risks. Each gate returns pass, revise or block.
  5. Claude Opus 5, the formulation and design agent, produces unit 3, the hypothesis registry; unit 4, the research plan; and unit 5, experiment design and modeling. The provider-neutral review role uses a fresh single-shot invocation for challenge T1, checking design validity, testability, discriminating power against the registered hypothesis, and reproducibility readiness before H4 can pass. T1 reads the frozen protocol alongside the design protocol, hypothesis registry and evidence matrix, so the exact replication, seed and analysis fields are inside the hash-bound review; its invocation is distinct from the later V1 and M1 audit invocations.
  6. After H4 passes, Claude Sonnet 5 runs unit 6, code generation and execution, and produces unit 7, experimental results. The version 2 code record retains a hash-bound Git bundle so T2 can verify offline that the exact executed source exists in the declared clean commit. That record, the environment lock, data manifest, run manifest and immutable raw results reach GPT 5.6 Terra under integrity contract T2.
  7. Terra drives unit 8, validation and reproducibility, in parallel with unit 9, statistical verification; both units produce unit 10, the verification report. Before writing begins, the separate review role audits statistical validity, assumptions, effect sizes, uncertainty, reproduction evidence, honest denominators, failure accounting, traceability and residual limitations under V1. V1 reads the immutable raw results, code commit, run manifest and frozen protocol in addition to the three reports, so the audit is recomputation against primary evidence rather than a reading of the producer's own summary.
  8. Typed returns preserve defect provenance and are emitted by the gate that detects them: evidence gaps return to unit 1 from E1; a T1 design-and-reproduction revision returns to unit 3 so its findings remain visible through all three planning units; code or run defects return to unit 6, scope or plan defects to unit 4, and assumption violations to unit 5, all from V1; claim-support gaps return to unit 11 from M1. Each gate carries a bounded revision budget; when the budget is spent the gate may only block or escalate, so no return path can cycle indefinitely.
  9. The verification report and read-only evidence, hypothesis, plan and result artifacts pass to Claude Fable 5, the writing and output agent, which drives unit 11, result synthesis and visualization, and produces unit 12, the manuscript together with a claim–evidence map and figure registry.
  10. Reviewer challenge M1 audits the manuscript for citation–claim, result–claim, method–protocol, statistical, overclaim and hypothesis consistency. The role may be implemented by a separately instantiated model reviewer or a qualified human specialist, but not by the producing context. Claim-support failures return to unit 11 with a typed reason. Only an M1 pass may proceed to the human final decision.
  11. The final human decision has five explicit scientific exits: release, revise, narrow, null-result and stop. Release, null-result and stop all write a release manifest, so abandoning a study leaves a record rather than an unreported file-drawer result. Separate human, execution-verification and reviewer identities prevent self-QA from being presented as a separated review; they do not guarantee epistemic independence.
  12. The exposed deterministic gate ledger requires registered, versioned artifact references, enforces declared ownership and prerequisites, records each challenge's local CLI argv, prompt digest, response-log digest and exit code, and invalidates dependent decisions when an input changes. Schema-bearing artifacts cannot be registered empty: a run manifest must carry its replication count and seeds, and a release manifest must carry its revision counts and any scope changes. FINAL revise reopens manuscript review; narrow reopens design review while leaving the hypothesis registry standing, so a post-hoc scope change cannot rewrite the hypothesis it was meant to test; release, null-result, stop and block are terminal routes. The host system remains responsible for authentication, durable storage and scientific execution; captured local provenance is not remote-provider attestation.
  13. Every agent, task and artifact node is interactive. Activating a node retains that node, its directly connected neighbours, all incident edges and any contract edge it owns, including challenge arrows, typed returns and labels, while dimming the remaining graph. Activating the same node again, clicking elsewhere or pressing Escape clears the focus.
  14. Shared tools and infrastructure — vector database, knowledge graph, code repositories, experiment tracking, compute, data storage, notebooks, the artifact registry holding versioned digests, and scientific APIs and libraries — are available across the workflow; small theme-coloured glyphs identify configured tools rather than universal provider requirements.