Evidence and metric jury
Two six-run metrics are fully reconstructed and match at reported precision. One remains partial. One cannot be recomputed from the released evidence.
Authored evidence path
The paper reports four headline metrics for the recurring default condition. We can reconstruct two from exact reviewed inputs (Semantic Recall evidence, Tree Balance evidence). Visual Coverage remains a three-seed diagnostic because the remaining three seed inputs are unavailable. Semantic Coverage cannot be recomputed because the reported caption embeddings are absent; substituting new embeddings would create a new experiment rather than reproduce the paper.
Scroll horizontally to inspect all columns.
| Metric | Reconstruction | Paper report | Status |
|---|---|---|---|
| Semantic Recall | 0.086922550176 ± 0.001134665940 | 0.087 ± 0.001 | Matched at reported precision |
| Tree Balance | 0.235345387783 ± 0.026343534617 | 0.235 ± 0.026 | Matched at reported precision |
| Visual Coverage | 0.616580784321 ± 0.001191336073, three seeds only | 0.614 ± 0.002 | Not recomputable as the six-run estimand |
| Semantic Coverage | Not available | 0.696 ± 0.004 | Not recomputable; reported caption embeddings are absent |
Per-run reconstructed estimates
Each row is one complete experimental run. Images and nouns are nested measurements, not replicates.
Scroll horizontally to inspect all columns.
| Run | Seed | Semantic Recall | Tree Balance |
|---|---|---|---|
| default-s3 | 3 | 0.088380311872 | 0.158210413665 |
| default-s4 | 4 | 0.086786806661 | 0.347875508094 |
| default-s5 | 5 | 0.082133439213 | 0.215455299242 |
| default-s6 | 6 | 0.086778786660 | 0.212968234360 |
| default-s7 | 7 | 0.090599490416 | 0.213417516343 |
| default-s8 | 8 | 0.086856466232 | 0.264145354996 |
Per-noun Semantic Recall decomposition
The compact derivative contains 1,823 nouns across six runs. The table shows the highest mean encoder maxima in the authored full-vocabulary view; a high maximum is not proof that an image depicts the noun.
Scroll horizontally to inspect all columns.
| Noun | Mean maximum similarity | Run-level SE | Mean top-two margin |
|---|---|---|---|
| teapot | 0.146555284659 | 0.001732482962 | 0.003691546619 |
| eye | 0.145039282739 | 0.000739000930 | 0.001507346829 |
| eyedropper | 0.145009100437 | 0.006257989345 | 0.006502966086 |
| teacup | 0.144239241878 | 0.002597980019 | 0.001862930755 |
| flashbulb | 0.137439168990 | 0.005266498634 | 0.004196812709 |
| duck | 0.137109828492 | 0.002016470414 | 0.004450897376 |
| kettle | 0.136402567228 | 0.003699259084 | 0.003629857053 |
| puddle | 0.135473035276 | 0.001536131729 | 0.002229529123 |
| oilcan | 0.135273581992 | 0.003275118108 | 0.003462606420 |
| bubble | 0.135043802361 | 0.002073989851 | 0.003549747169 |
| can | 0.134760860354 | 0.006947886668 | 0.007980408768 |
| egg | 0.133389967183 | 0.003191304816 | 0.011626866957 |
The real metric jury stops where the evidence stops
No Pareto front or weighted condition rank is computed from the approved real evidence: only one experimental condition has independently reconstructed run records. Paper-reported aggregates remain a reference table, not substitutes for missing independently derived runs.
Selected Table 1 rows: reported reference only
Open the source table in the paper reader.
Scroll horizontally to inspect all columns.
| Paper row | Semantic Recall | Visual Coverage | Semantic Coverage | Tree Balance |
|---|---|---|---|---|
| Implemented default (epsilon=0; CL=1; NA=0) | 0.087 ± 0.001 | 0.614 ± 0.002 | 0.696 ± 0.004 | 0.235 ± 0.026 |
| Moderate selection noise (epsilon=0.25) | 0.088 ± 0.001 | 0.638 ± 0.009 | 0.717 ± 0.004 | 0.249 ± 0.022 |
| 1,000 agents (personality-prompt pool) | 0.088 ± 0.001 | 0.665 ± 0.004 | 0.734 ± 0.013 | 0.476 ± 0.013 |
| Random baseline | 0.080 ± 0.001 | 0.612 ± 0.005 | 0.692 ± 0.002 | 0.540 ± 0.003 |
| Historical human archive | 0.089 ± 0.000 | 0.681 ± 0.000 | 0.730 ± 0.000 | 0.363 ± 0.000 |
The historical human row is a fixed descriptive comparator. It is not resampled or represented as a six-run experimental condition.
Synthetic metric stress lab
This fixture is not a paper result. With equal Semantic Recall and Tree Balance weights, the three specialists tie on weighted average rank while all three remain on the Pareto front; the control is dominated. Changing weights can break that tie, which is why weighted ranks are a sensitivity lens rather than a verdict.
| Synthetic condition | Equal-weight rank | Pareto status |
|---|---|---|
| Recall specialist | 2.000000 | Pareto tradeoff |
| Coverage specialist | 2.000000 | Pareto tradeoff |
| Tree specialist | 2.000000 | Pareto tradeoff |
| Dominated control | 4.000000 | Dominated |
Source disagreements remain visible
PDF page 7 prose says NA=1; Table 1 grey row and released sweep code define the implemented default as NA=0
semantic implementation delegates to an unpinned torch_fpsample implementation
paper says summed minimum cosine distance; released code computes mean maximum cosine similarity
released code collapses multiple parents to the earliest retained parent
paper says random datapoint; released visual code starts at archive row zero
paper says 1824 THINGS names; pinned upstream list contains 1823 unique nonblank names
Analysis digest: sha256:273bce329561037573bfd8a92ab27f505f4eec83431d6cf85b81fa4684281d3f
Signed evidence-spine digest: sha256:5ff56b5a2524a1877af1ac5a84b02b66c68e3b06b3c9d347de48271d4ec9acea
The ledger, run table, noun decomposition, jury boundary, reported reference, stress fixture, and discrepancies above are the complete no-JavaScript reading path. Interactive assets are lazy: until a reader directly activates the analysis, the page fetches neither the compact metric-jury dataset nor the Python worker and Pyodide runtime. After activation, the same content-addressed analysis powers selectable views and sensitivity controls without an iframe or server backend.
Provenance boundary
Each reconstructed Semantic Recall row exposes its run ID, seed, source hashes, renderer commit, model and vocabulary identities, and transformation. The downloadable compact metric-jury evidence contains six values per noun plus best-image entry identifiers and source digests. The compact derivative is bound to the exact signed evidence-spine digest and per-noun Arrow source hash. Raw genomes, images, model weights, image embeddings, and Arrow bytes never enter the static site.
Evidence JSON is admitted to CacheStorage only after Web Crypto reproduces the content digest encoded in its immutable filename. A later online page load can reuse those verified evidence documents without refetching them. This first vertical slice does not install a service worker or promise a cold offline boot: loading the worker and Pyodide runtime still requires the network or the browser's ordinary HTTP cache.
The authored fallback and lightweight metric controls remain usable on mobile. Large imports, long-running analysis, and research-scale multi-panel work are desktop-first; mobile profiles are not promised desktop-equivalent throughput. If browser capabilities, device memory, or canvas limits prevent an interactive view, the authored evidence remains available instead of being hidden.
What the jury can and cannot decide
The approved real evidence contains six reconstructed runs for one condition. That is enough to show run-level variation and to recompute Semantic Recall after documented noun exclusions. It is not enough to compare conditions. The live jury therefore refuses to manufacture a Pareto front, bootstrap rank stability, or weighted winner from the paper's aggregate table.
The reported intervention table remains visible because it is part of the paper's argument. It is explicitly separated from the computational input. The synthetic stress lab exercises the same weighting and Pareto concepts on a toy fixture so readers can see how preferences change ranks without mistaking that demonstration for a result about Picbreeder.
The noun sensitivity controls ask three bounded questions: what happens after removing the 50 highest-scoring nouns, the 50 lowest-scoring nouns, or the 5% with the smallest average best-versus-second-best image margin? These are robustness probes, not post-hoc corrections to the published metric.
Reproducible views
The live controls encode only the selected metrics, integer weights, selected condition, lens, noun-filter preset, view, schema, and exact analysis digest. Copy reproducible view targets the intended immutable 0.1.0 release route. Malformed or mismatched fragments fail atomically to the authored view. Uploaded files, OPFS paths, annotations, diagnostics, and other local state are never placed in the URL.
If the local runtime rejects an asynchronous request, the controls and metric-state fragment return to the last completed view and announce the failure rather than leaving a share URL ahead of stale results.
Rights and release status
G3 authorized private development of this Milestone 4 editorial candidate; it did not approve this exact new derivative for publication. The compact noun derivative remains conservatively classified as CC BY-NC 4.0 dataset-derived evidence, and the candidate remains private-review-only. Public redistribution still requires the later publication-preview and production gates; nothing on this page changes that boundary.