What Makes Search Stay Open?
This five-minute audit separates the paper’s question, the released computation, and the evidence we can reconstruct.
Choose a deeper audit route
Follow the paper reconstruction for the study’s question, interventions, and conclusions.
Inspect the evidence spine for metrics, provenance, and six specification discrepancies.
Read the claim and evidence method for the rules that keep source claims and our analysis distinct.
The argument in five moves
From replication target to research program. This static map is repeated in prose immediately below.
| Move | What survives the audit | Evidence |
|---|---|---|
| 1 | Target the conditions for discovery, not a particular image archive. | Approved paper claim |
| 2 | Treat four measures as operational instruments. | Approved qualification |
| 3 | Keep the two six-run default-condition reconstructions that match at reported precision. | Semantic Recall; Tree Balance |
| 4 | Preserve one partial and one unavailable reconstruction. | Visual Coverage; Semantic Coverage |
| 5 | Turn finite evidence into a stronger next question, not unbounded proof. | Approved interpretation |
1. Replicate the conditions for discovery
The study aims to reproduce the conditions that enabled discovery in Picbreeder, rather than its particular images, representations, or genealogies (approved paper claim).
That choice shifts attention from imitation to search dynamics: what lets one discovery become a stepping stone toward another?
2. Treat the metrics as instruments
The four headline metrics are operational attempts to capture qualities of the human archive through recall and embedding- or tree-based diversity measures (approved qualification).
They make comparison possible. They do not turn “open-endedness” into a single directly observed quantity.
3. Keep the two precision-matched reconstructions
Our separately signed six-run Semantic Recall reconstruction matches the paper-rounded mean and standard error for the implemented default condition (signed reconstruction).
Our six-run Tree Balance reconstruction also matches the paper-rounded mean and standard error for that selected condition (signed reconstruction).
These matches are meaningful, but bounded: they verify two computations for a specific finite experiment.
4. Do not manufacture the missing two
Visual Coverage is a three-seed diagnostic because the pinned inputs do not support the paper’s full six-seed aggregate comparison (signed reconstruction status).
Semantic Coverage remains unrecomputable because the reported caption embeddings are absent; substituting new embeddings would define a different experiment (signed reconstruction status).
The honest shape of the result is therefore two reconstructed metrics, one bounded diagnostic, and one unavailable reconstruction. It is not four interchangeable numbers.
5. Turn the gap into a better experiment
We interpret these finite, intervention-specific runs as evidence about measurable archive behavior, not as proof that the system achieves unbounded open-endedness (approved interpretation).
The next study should not merely ask which condition wins a four-column table. It should ask which discoveries remain useful later, whether the archive continues to create new reachable directions, and how sensitive that judgment is to the chosen metric and implementation.
Read it as the paper’s argument
The paper varies exploration noise, interaction history, and the number of agents, then observes tradeoffs across four archive metrics (metric framework). Random selection can encourage exploration while trading away legibility and recall (exploration result); more context does not improve results monotonically (history result); and more agents broaden some behaviors while introducing other risks (multi-agent result). The interesting pattern is a set of tensions, not one recipe that simply scales.
Read it skeptically
The reconstruction makes six disagreements impossible to smooth over: the default agent count differs across prose, table, and code (agent-count discrepancy); the semantic sampler delegates its first center to an unpinned implementation (sampling discrepancy); Semantic Recall’s prose and code specify different operations (metric discrepancy); the tree projection retains one parent from multi-parent records (tree discrepancy); Visual Coverage starts from a fixed row where the paper describes a random point (visual discrepancy); and the pinned vocabulary has 1,823 unique nonblank names where the paper reports 1,824 (vocabulary discrepancy). These do not erase the experiment. They define which version of it a result actually supports.
Read it as a research design
Use the four metrics as a panel of instruments rather than a verdict. Then add tests that expose the time dimension the current aggregate table cannot settle:
Stepping-stone reuse: do earlier discoveries enable later ones that were initially unreachable?
Horizon sensitivity: do apparent improvements persist as the search runs longer?
Metric sensitivity: does the conclusion survive alternate, explicitly versioned encoders and archive projections?
Human calibration: do expert and participant judgments agree with the directions the metrics reward?
These are proposed questions, not findings from the current evidence. They extend the paper’s own call for broader conditions, longer runs, and carefully designed human studies (approved qualification).
What the interactive layer is for
The authored argument above is complete without JavaScript or WebAssembly. The Starimo layer on the evidence spine is a progressive inspector over the same compact, content-addressed evidence: it can change how a reader examines a result, but it cannot silently replace the approved fallback or create a missing reconstruction.