Skip to content
Hoxy Papers

Evidence and metric jury

Two six-run metrics are fully reconstructed and match at reported precision. One remains partial. One cannot be recomputed from the released evidence.

Authored evidence path

The paper reports four headline metrics for the recurring default condition. We can reconstruct two from exact reviewed inputs (Semantic Recall evidence, Tree Balance evidence). Visual Coverage remains a three-seed diagnostic because the remaining three seed inputs are unavailable. Semantic Coverage cannot be recomputed because the reported caption embeddings are absent; substituting new embeddings would create a new experiment rather than reproduce the paper.

Scroll horizontally to inspect all columns.

Metric Reconstruction Paper report Status
Semantic Recall 0.086922550176 ± 0.001134665940 0.087 ± 0.001 Matched at reported precision
Tree Balance 0.235345387783 ± 0.026343534617 0.235 ± 0.026 Matched at reported precision
Visual Coverage 0.616580784321 ± 0.001191336073, three seeds only 0.614 ± 0.002 Not recomputable as the six-run estimand
Semantic Coverage Not available 0.696 ± 0.004 Not recomputable; reported caption embeddings are absent

Per-run reconstructed estimates

Each row is one complete experimental run. Images and nouns are nested measurements, not replicates.

Scroll horizontally to inspect all columns.

Run Seed Semantic Recall Tree Balance
default-s3 3 0.088380311872 0.158210413665
default-s4 4 0.086786806661 0.347875508094
default-s5 5 0.082133439213 0.215455299242
default-s6 6 0.086778786660 0.212968234360
default-s7 7 0.090599490416 0.213417516343
default-s8 8 0.086856466232 0.264145354996

Per-noun Semantic Recall decomposition

The compact derivative contains 1,823 nouns across six runs. The table shows the highest mean encoder maxima in the authored full-vocabulary view; a high maximum is not proof that an image depicts the noun.

Scroll horizontally to inspect all columns.

Noun Mean maximum similarity Run-level SE Mean top-two margin
teapot 0.146555284659 0.001732482962 0.003691546619
eye 0.145039282739 0.000739000930 0.001507346829
eyedropper 0.145009100437 0.006257989345 0.006502966086
teacup 0.144239241878 0.002597980019 0.001862930755
flashbulb 0.137439168990 0.005266498634 0.004196812709
duck 0.137109828492 0.002016470414 0.004450897376
kettle 0.136402567228 0.003699259084 0.003629857053
puddle 0.135473035276 0.001536131729 0.002229529123
oilcan 0.135273581992 0.003275118108 0.003462606420
bubble 0.135043802361 0.002073989851 0.003549747169
can 0.134760860354 0.006947886668 0.007980408768
egg 0.133389967183 0.003191304816 0.011626866957

The real metric jury stops where the evidence stops

No Pareto front or weighted condition rank is computed from the approved real evidence: only one experimental condition has independently reconstructed run records. Paper-reported aggregates remain a reference table, not substitutes for missing independently derived runs.

Selected Table 1 rows: reported reference only

Open the source table in the paper reader.

Scroll horizontally to inspect all columns.

Paper row Semantic Recall Visual Coverage Semantic Coverage Tree Balance
Implemented default (epsilon=0; CL=1; NA=0) 0.087 ± 0.001 0.614 ± 0.002 0.696 ± 0.004 0.235 ± 0.026
Moderate selection noise (epsilon=0.25) 0.088 ± 0.001 0.638 ± 0.009 0.717 ± 0.004 0.249 ± 0.022
1,000 agents (personality-prompt pool) 0.088 ± 0.001 0.665 ± 0.004 0.734 ± 0.013 0.476 ± 0.013
Random baseline 0.080 ± 0.001 0.612 ± 0.005 0.692 ± 0.002 0.540 ± 0.003
Historical human archive 0.089 ± 0.000 0.681 ± 0.000 0.730 ± 0.000 0.363 ± 0.000

The historical human row is a fixed descriptive comparator. It is not resampled or represented as a six-run experimental condition.

Synthetic metric stress lab

This fixture is not a paper result. With equal Semantic Recall and Tree Balance weights, the three specialists tie on weighted average rank while all three remain on the Pareto front; the control is dominated. Changing weights can break that tie, which is why weighted ranks are a sensitivity lens rather than a verdict.

Synthetic condition Equal-weight rank Pareto status
Recall specialist 2.000000 Pareto tradeoff
Coverage specialist 2.000000 Pareto tradeoff
Tree specialist 2.000000 Pareto tradeoff
Dominated control 4.000000 Dominated

Source disagreements remain visible

  • PDF page 7 prose says NA=1; Table 1 grey row and released sweep code define the implemented default as NA=0

  • semantic implementation delegates to an unpinned torch_fpsample implementation

  • paper says summed minimum cosine distance; released code computes mean maximum cosine similarity

  • released code collapses multiple parents to the earliest retained parent

  • paper says random datapoint; released visual code starts at archive row zero

  • paper says 1824 THINGS names; pinned upstream list contains 1823 unique nonblank names

Analysis digest: sha256:273bce329561037573bfd8a92ab27f505f4eec83431d6cf85b81fa4684281d3f

Signed evidence-spine digest: sha256:5ff56b5a2524a1877af1ac5a84b02b66c68e3b06b3c9d347de48271d4ec9acea

The ledger, run table, noun decomposition, jury boundary, reported reference, stress fixture, and discrepancies above are the complete no-JavaScript reading path. Interactive assets are lazy: until a reader directly activates the analysis, the page fetches neither the compact metric-jury dataset nor the Python worker and Pyodide runtime. After activation, the same content-addressed analysis powers selectable views and sensitivity controls without an iframe or server backend.

Provenance boundary

Each reconstructed Semantic Recall row exposes its run ID, seed, source hashes, renderer commit, model and vocabulary identities, and transformation. The downloadable compact metric-jury evidence contains six values per noun plus best-image entry identifiers and source digests. The compact derivative is bound to the exact signed evidence-spine digest and per-noun Arrow source hash. Raw genomes, images, model weights, image embeddings, and Arrow bytes never enter the static site.

Evidence JSON is admitted to CacheStorage only after Web Crypto reproduces the content digest encoded in its immutable filename. A later online page load can reuse those verified evidence documents without refetching them. This first vertical slice does not install a service worker or promise a cold offline boot: loading the worker and Pyodide runtime still requires the network or the browser's ordinary HTTP cache.

The authored fallback and lightweight metric controls remain usable on mobile. Large imports, long-running analysis, and research-scale multi-panel work are desktop-first; mobile profiles are not promised desktop-equivalent throughput. If browser capabilities, device memory, or canvas limits prevent an interactive view, the authored evidence remains available instead of being hidden.

What the jury can and cannot decide

The approved real evidence contains six reconstructed runs for one condition. That is enough to show run-level variation and to recompute Semantic Recall after documented noun exclusions. It is not enough to compare conditions. The live jury therefore refuses to manufacture a Pareto front, bootstrap rank stability, or weighted winner from the paper's aggregate table.

The reported intervention table remains visible because it is part of the paper's argument. It is explicitly separated from the computational input. The synthetic stress lab exercises the same weighting and Pareto concepts on a toy fixture so readers can see how preferences change ranks without mistaking that demonstration for a result about Picbreeder.

The noun sensitivity controls ask three bounded questions: what happens after removing the 50 highest-scoring nouns, the 50 lowest-scoring nouns, or the 5% with the smallest average best-versus-second-best image margin? These are robustness probes, not post-hoc corrections to the published metric.

Reproducible views

The live controls encode only the selected metrics, integer weights, selected condition, lens, noun-filter preset, view, schema, and exact analysis digest. Copy reproducible view targets the intended immutable 0.1.0 release route. Malformed or mismatched fragments fail atomically to the authored view. Uploaded files, OPFS paths, annotations, diagnostics, and other local state are never placed in the URL.

If the local runtime rejects an asynchronous request, the controls and metric-state fragment return to the last completed view and announce the failure rather than leaving a share URL ahead of stale results.

Rights and release status

G3 authorized private development of this Milestone 4 editorial candidate; it did not approve this exact new derivative for publication. The compact noun derivative remains conservatively classified as CC BY-NC 4.0 dataset-derived evidence, and the candidate remains private-review-only. Public redistribution still requires the later publication-preview and production gates; nothing on this page changes that boundary.