Completed run · August 2026 · historical evidence
Which model made fewer mistakes on handwritten digits?
This run compared two familiar classifiers on the same public dataset. Every model choice was made using training data only; the held-out images were used only to measure error.
Give both models the same handwritten digits. Tune them fairly. Which one misclassifies fewer unseen images?
The RBF-SVC made about 1.7 fewer errors per 100 predictions than logistic regression.
The recorded mean error difference was 1.714 percentage points. Its nominal 95% interval was 1.350 to 2.078 points; the pre-specified multiplicity-adjusted interval also remained above zero.
images of digits 0–9 from scikit-learn’s bundled public dataset
paired held-out folds in which the RBF-SVC had lower error
raw fold–classifier records present; no run failures or exclusions
byte-for-byte reproduction of the recorded raw-results payload
Important: this does not show that SVC is universally better. It answers only this dataset, these two model families, and this frozen protocol. Fifteen of 25 SVC selections reached an edge of the registered tuning grid.
See how the answer was produced.
Click any node in the architecture. The graph will isolate its connections, and the run record beside it will explain what happened at that point in this study.
Click a task, agent, source, or human node.Esc clears the graph focus.
Audit record
Technical evidence remains available, but it no longer has to be understood before the research story.
Why is this labeled “historical / partial”?
- The run crossed successive v0.2.1 release-candidate wheels; it was not executed exclusively with one final tagged package.
- Stages 1–5 were real Claude calls with retained logs, but they predate host-issued execution receipts. Stages 6–12 have those receipts.
- M1 used
SEPARATE MODEL, notSEPARATE PROVIDER. - Rendered PNG, SVG, and PDF figures were not produced, although their registry and generation code were hash-bound.
What did the workflow verify?
It verified registered artifacts, recomputed hashes, recorded provenance, gate conditions, reviewer separation, revision budgets, and the final human decision. It does not establish scientific correctness, provider reliability, or autonomous orchestration, and it is not a qualification run for the current release.
Which providers have stronger execution evidence?
Host-issued receipts bind Claude Sonnet 5 to stages 6–7, Codex GPT-5.6 Terra to stages 8–10, and Claude Fable 5 to stages 11–12. Reviewer decisions bind Codex Terra to E1, T1, and T2, and Claude Opus 5 to V1 and M1. Stages 1–5 have retained real-call logs but no equivalent host receipts.
rg-20260802-001 · 12 / 12 stages · 21 valid artifacts · final outcome release
MANIFEST ec59b23c405d0ac7630054177e3d1be09b31ba46032f9bd22549068eec8cdce6