research—graphone real run, explained ← Back to overview

Completed run · August 2026 · historical evidence

Which model made fewer mistakes on handwritten digits?

This run compared two familiar classifiers on the same public dataset. Every model choice was made using training data only; the held-out images were used only to measure error.

The problem, without the jargon

Give both models the same handwritten digits. Tune them fairly. Which one misclassifies fewer unseen images?

Recorded answer

The RBF-SVC made about 1.7 fewer errors per 100 predictions than logistic regression.

The recorded mean error difference was 1.714 percentage points. Its nominal 95% interval was 1.350 to 2.078 points; the pre-specified multiplicity-adjusted interval also remained above zero.

1,797

images of digits 0–9 from scikit-learn’s bundled public dataset

25 / 25

paired held-out folds in which the RBF-SVC had lower error

50 / 50

raw fold–classifier records present; no run failures or exclusions

Exact

byte-for-byte reproduction of the recorded raw-results payload

Important: this does not show that SVC is universally better. It answers only this dataset, these two model families, and this frozen protocol. Fifteen of 25 SVC selections reached an edge of the registered tuning grid.

See how the answer was produced.

Click any node in the architecture. The graph will isolate its connections, and the run record beside it will explain what happened at that point in this study.

Click a task, agent, source, or human node.Esc clears the graph focus.

Audit record

Technical evidence remains available, but it no longer has to be understood before the research story.

Why is this labeled “historical / partial”?
  • The run crossed successive v0.2.1 release-candidate wheels; it was not executed exclusively with one final tagged package.
  • Stages 1–5 were real Claude calls with retained logs, but they predate host-issued execution receipts. Stages 6–12 have those receipts.
  • M1 used SEPARATE MODEL, not SEPARATE PROVIDER.
  • Rendered PNG, SVG, and PDF figures were not produced, although their registry and generation code were hash-bound.
What did the workflow verify?

It verified registered artifacts, recomputed hashes, recorded provenance, gate conditions, reviewer separation, revision budgets, and the final human decision. It does not establish scientific correctness, provider reliability, or autonomous orchestration, and it is not a qualification run for the current release.

Which providers have stronger execution evidence?

Host-issued receipts bind Claude Sonnet 5 to stages 6–7, Codex GPT-5.6 Terra to stages 8–10, and Claude Fable 5 to stages 11–12. Reviewer decisions bind Codex Terra to E1, T1, and T2, and Claude Opus 5 to V1 and M1. Stages 1–5 have retained real-call logs but no equivalent host receipts.

RUN rg-20260802-001 · 12 / 12 stages · 21 valid artifacts · final outcome release MANIFEST ec59b23c405d0ac7630054177e3d1be09b31ba46032f9bd22549068eec8cdce6