Recurring sequences and positions in the visual observations

Project Summary

Completed bounded descriptive first pass. The program’s raw flags (240 RELIABLE, 79 CONFIRMED_SENSITIVE, 92 EXPLORATORY_ONLY) are mechanical filter outputs. They stay frozen and unvalidated, and the formal “reliable” and “confirmed” wording is withdrawn. The 411 candidates are four overlapping ways of grouping the same forms, not 411 independent results.

Execution: complete. Scientific acceptance: none — the results are descriptive diagnostics under one grouping rule. Human-interface completion: depends on Jeremy being able to use this page from START HERE. Physical iPhone use: unverified.

What were we trying to find out?

Do the neutral v006 visual observations show neighbouring forms that recur in a particular left-to-right order, or forms that prefer particular places in a row, and do those patterns hold on leaves held out from the search?

What did we do?

What did we find?

What went wrong and was corrected

Leaf test selection bug
The original leaf-by-leaf sign test counted a leaf only if the pair occurred there or at least one was expected, so rare enrichments were counted as wins. Under the reviewer’s calibrated (and itself approximate) test, 128 of the 240 raw RELIABLE flags pass; 96 rest mostly on leaves with very low expected counts.
Size confounding
Small and tall shapes pair differently. Conditioning the expectation on size class shrinks the effects: #243 falls from O/E 3.79 to about 1.50, #250 disappears (221 vs 220.5) and spaced pair #375 fails. Only 68 of 158 tested family-pair RELIABLE flags pass.
Weak segmentation-clean check
The check that leaves out forms with open join/split alternatives had no minimum support: example 10 passed on 1 observed vs 0.095 expected. Results with fewer than 5 clean observations are treated as sensitive.
Multi-page and partial views
Photos showing two folios were assigned to their first folio number. As a result, five of the 32 held-out leaves (f069, f071, f089, f100, f102) are tied to search-set leaves through two-page photos; for example, page f089r is in the search-set photo f088v_and_f089r. Page f090r may appear both in a held-out photo (f089v_part_and_f090r) and as a separate search-set photo (f090r). Whether any text was counted on both sides is unresolved. No whole-leaf integrity is claimed for the pharmaceutical and astronomical sections. This corrects an earlier rough figure of seven.
Duplicate-text check is uninformative
A 4-gram check for duplicated text flagged nothing, but it had no positive control and even views expected to overlap scored near baseline. Its negative result is not proof that the views are independent.
Partition check not independent
The family-partition check reuses the same occurrences: the rare-pooled layer copies the v006 layer (67 of 68), and the v005 layer largely relabels it.
Row ends are not line starts
Row-segment ends include drawing boundaries (examples 2 and 3), false edges next to tall forms (example 12) and boxes merged across rows (example 13).
What failing a check means
Failing the calibrated leaf test means “cannot be tested leaf by leaf with this approximate test”, not “refuted”. Failing size conditioning means “not distinguishable from size pairing”. Size classes partly overlap with families, so this check may also remove some real family signal.

Corrections after the final presentation critique

What does it mean?

Under this geometric grouping, neighbours with a small gap show some modest ordering preferences and a few nearly absent orders, and some frequent groups sit at row-segment ends less often than expected. In the inspected pictures such small-gap neighbours lay inside one written group. That fits handwriting construction, the way the extractor cuts forms, ordinary spelling, abbreviation or generated text, and cannot tell them apart. No character, word, sound, language, cipher or reading direction is supported. Reversing the order cannot decide reading direction, because mirrored writing would show the same asymmetry.

Limitations

What needs attention?

Five picture cases invite your judgment (Review Needed). No answer is needed to close this study, and nothing follows automatically.

Strongest next test

A blinded crop audit. For the core tight pairs and the near-absent pairs, show 10 random examples beside 10 examples matched on size class but from different groups, and have a reviewer judge “one split form” or “two forms”. Any later test should correct the leaf grouping of multi-page views, link rows along the baseline, use a calibrated leaf test and a size-conditioned expectation, test the saved join/split alternatives, and use independent human segmentation as validation. Nothing follows automatically; it needs separate authorization.

Who did the work

The analysis and the critique were both written by Claude Opus from the same provider. The reviewer read the author’s assessment first, reused the author’s data-loading code, viewed seven of the sixteen pictures and did not open the native TIFFs. This presentation was written by a separate read-only Claude Opus session and run by Codex; it reran no statistics and changed no frozen result or raw label. A final presentation critique was written by the same reviewer session; it is not independent of the earlier critique. It could not open the rendered report (too large for its reading tool), checked the text through this generator, and viewed only display examples 11 and 12.

Enlarged example