ARCHIVED RESEARCH PROJECT · completed 2026-09-29 · What visually appears written across the 213 authoritative Yale views, with source and segmentation uncertainty preserved? · Back to Home

PROJECT SUMMARY

QuestionWhat visually appears written across the 213 authoritative Yale views, with source and segmentation uncertainty preserved?
MethodThe only classification inputs were Yale pixels, verified canonical source metadata and neutral geometric evidence. EVA, existing [preserved evidence reference] translations, language hypotheses, phonetic values, semantic values, historical sign identities and writer labels were not…
Important numbers81,889 visual families · 116,430 variants · 133,777 provisional writing forms
Finding and conclusionThe uniform seed-20260929 sample contained 112 writing observations and eight definite illustration strokes among 120 automatic high-writing proposals. The pre-correction false-writing estimate is 6.7%, with a Wilson 95% sampling interval of 3.4–12.6%. This is a single coordinating-model…
LimitsAll 213 authoritative Yale TIFF views completed a native-pixel pass and manuscript-wide provisional reconciliation. The versioned occurrence ledger is authoritative for this run's observations and review decisions. It is not an established complete or accurate transcription of all manuscript writing. Source completeness, exact crops and database consistency passed engineering checks; material classification, segmentation, family boundaries and visible-ink recall remain scientifically provisional.
MeaningAll 213 authoritative Yale TIFF views completed a native-pixel pass and manuscript-wide provisional reconciliation. The versioned occurrence ledger is authoritative for this run's observations and review decisions. It is not an established complete or accurate transcription of all manuscript writing. Source completeness, exact crops and database consistency passed engineering checks; material classification, segmentation, family boundaries and visible-ink recall remain scientifically provisional.

Read the findings preserved from this completed study.

EXPLORE ALL RESULTS · VISUAL FAMILY INVENTORY

81,889 active provisional families, 116,430 variants, 133,777 provisional writing occurrences. Every family and variant is searchable here. 1,593 reproducibly selected families have portable Yale-crop sheet previews; every family has a Mac-folder preview and the complete writing occurrence drill-down. Browse all source material classes and observations in the Mac project folder. Family resemblance and counts are machine proposals, not accepted glyph identities.

203,403 unresolved material candidates are outside the assigned-family inventory. 54,323 writing forms have recorded segmentation alternatives; uncertainty remains explicit in the authoritative ledger.

REVIEW NEEDED

27 items need review in the preserved curated card set.

This curated view is separate from the complete meaningful results listed above.

VOYNICH PROJECT

Archived research project · completed 2026-09-29

ARCHIVED RESEARCH PROJECT

Full-source neutral visual transcription — v004

What visually appears written across the 213 authoritative Yale views, with source and segmentation uncertainty preserved?

OPEN THE REVIEW FINDINGS AS COMPLETED

HUMAN REVIEW · AS PRESENTED WHEN CURRENT

Full-source neutral visual transcription — provisional

27 source-linked cards. Display images are embedded; original TIFFs and scientific records remain authoritative in the project folder.

The v004 ledger contains 133,777 provisional writing forms in 81,889 active visual families and 116,430 variants. Its 27 cards are selected QC and judgment cases, not the full result set. Contamination, recall and family stability remained unresolved at completion.

FINDINGS FROM THE COMPLETED STUDY · public presentation of [preserved evidence reference]

Full-manuscript neutral visual observations and transcription ledger v004

All 213 authoritative Yale TIFF views completed a native-pixel pass and manuscript-wide provisional reconciliation. The versioned occurrence ledger is authoritative for this run's observations and review decisions. It is not an established complete or accurate transcription of all manuscript writing. Source completeness, exact crops and database consistency passed engineering checks; material classification, segmentation, family boundaries and visible-ink recall remain scientifically provisional.

Counts after source-based material review

MeasureCount
Source views processed213
All primary candidate observations, including noise658,585
High-confidence writing-form candidates49,053
Uncertain writing forms84,724
Total provisional writing forms133,777
Possible writing / possible artwork97,340
Unresolved candidate material203,403
Non-writing224,065
Active provisional visual families81,889
Active provisional visual variants116,430
Historical family / variant IDs, including retired artwork-only groups81,895 / 116,439
Writing forms with primary split/component ambiguity54,323
Close-neighbor join hypotheses, all unselected72,081
Global possible-family relationships148,144
Rare families in initial machine dictionary, fewer than four members77,065
Active single-member writing families68,379

These are connected-form proposals in source views, not character counts or physically unique occurrences. Partial/foldout views can repeat the same physical writing. Nine canonical-label overlap groups were tested; the saved results do not justify automatic physical deduplication or invented panel identities. Alternative segmentation can change unit counts. No accepted character, sound, meaning, word boundary or writer identity was assigned.

Scientific QC result

The uniform seed-20260929 sample contained 112 writing observations and eight definite illustration strokes among 120 automatic high-writing proposals. The pre-correction false-writing estimate is 6.7%, with a Wilson 95% sampling interval of 3.4–12.6%. This is a single coordinating-model source review, not an independent human error estimate. The interval excludes rater error and weights duplicated source views as separate observations. Ten confirmed artwork cases across the material/family samples were excluded by the public reviewed occurrence view; the corrected corpus's residual error has not been independently estimated.

The seed-20260932 sample of 120 assignments to existing families contained 92 visually compatible pairs, four incompatible whole-form pairs and 24 unresolved pairs. Incompatible pairs are 3.3% (Wilson interval 1.3–8.3%); the uncertainty-inclusive flag rate is 23.3%. This measures an operational visual-comparison judgment, not linguistic character accuracy. Only ten high-writing observations had family evidence scores at least 0.85; all ten were inspected. The dictionary remains heavily fragmented and is not stable as an accepted visual alphabet. Similarity scores are uncalibrated, not correctness probabilities.

Seven deliberately chosen parchment/plant controls on f001r/f002r produced zero high-confidence false writing assignments, but two leaf fragments were uncertain writing. The random material sample found no definite parchment-texture cases and did find illustration contamination. This does not establish manuscript-wide control. A deterministic scan additionally flagged 41 high-writing candidates on nonfolio/binding views and eight frequent families with nonfolio support; their source review remains unresolved. Large, faint, repeating illustration strokes can resemble writing under alignment tests.

The known f002r faint loop has source strokes outside its proposed box. Its exact wider native context and an unselected source-field alternative are retained. No visible-writing recall estimate exists. Thirty-two reproducibly selected source thumbnails are saved for an omission audit; their selection is not a completed independent ink-recall measurement.

Observation and dictionary method

The only classification inputs were Yale pixels, verified canonical source metadata and neutral geometric evidence. EVA, existing [preserved evidence reference] translations, language hypotheses, phonetic values, semantic values, historical sign identities and writer labels were not classification inputs. The coordinator has prior project knowledge; this is methodological input isolation, not a claim of a psychologically naive analyst.

Native grayscale smoothing (sigma 1.2), median background diameter 41 and background smoothing sigma 3 produce residual masks at 12/20/30. A separate 3x3 closing proposes grouping; original threshold masks and all 4,990,160 raw low-threshold components remain. Small candidates not promoted into the primary ledger remain in raw component tables/masks and TIFFs. Saturation, darkness, size, threshold support and local baseline alignment supply provisional material evidence. These tests failed on some illustrations; source-based overrides do not rewrite their original output.

Whole forms retain connected components, projection-valley alternatives and close-neighbor joins. Measured gaps and candidate cluster thresholds 0.15/0.30/0.50 are geometry, never words. Horizontal routes are provisional; line/region fields remain null where no identity is defensible. Circular/rotated/mixed layouts are retained without an asserted reading order.

Fixed-exemplar retrieval uses 40x40 masks, DCT descriptors, mask IoU at least 0.68, DCT distance at most 0.42 and aspect similarity at least 0.72. Variant IoU is at least 0.82. When two sufficiently similar families compete, the strongest remains an unaccepted candidate, alternatives are retained and its evidence score is capped at 0.74. New forms create provisional IDs without a family-count cap. Global retrieval records duplicate/relative, broad, rare and segmentation-sensitive flags; no count-reducing merge is accepted. These finite nearest-neighbor searches are not exhaustive proof that all duplicate families were found.

Every initially writing-classified form also has native threshold-support contour/branch/skeleton records. 118,259 of 133,787 machine writing proposals change connected-support/hole/junction topology across the three thresholds. This flags unstable geometry, not changed character identities. Skeletons are derived masks, not original pen strokes or a linguistic decomposition. Threshold holes permit later loop hypotheses without superseding the whole form.

Every historical f1r approval remains limited: twelve old boundaries rejected; two cross-line fields spatially separated without character counts; two fragments may be represented jointly while the excluded third fragment remains excluded. Sixteen decision events and their native field correspondences are preserved, with no new boundary or identity accepted.

Validation, versions and resource use

The exhaustive local pixel verifier compared all 658,585 primary lossless crops and 154,383 alternative-part recipes to native Yale pixels. All 213 TIFF byte/decoded-pixel hashes stayed unchanged. The integrated validator rebound that check through fresh source/checkpoint/all extraction-artifact hashes, checked all ledgers, family/variant references and 1,004 exact review crop/context assets, and verified 81,895 dictionary representations. Review integration subsequently preserved the original machine table and excluded ten confirmed artwork cases through a reviewed view. Named prior scientific inputs and previous archives are unchanged; this is not an exhaustive pre/post seal of every historical repository file.

v001's incorrect background estimate and v002's texture-heavy proposals were stopped and retained. Completed v003 extraction was reused read-only. An overly proliferating competing-family rule in partial v003 reconciliation was preserved and replaced in v004; the completed extraction/crop work was not restarted. All failures and intermediate outputs remain available.

Audit status and exact locations

Ready as a provenance-rich package for independent audit; not established as a verified complete writing transcription or an accepted corpus for linguistic experiments. Ordinary unresolved cases were retained through the full source pass. No morphology rerun, language/script comparison, proposed decipherment or translation was performed.

The saved configuration, implementation hashes, environment, source inventory, extraction seals, RNG seeds, review judgments and final file hashes are the reproduction contract. Run read-only query/validation tools against this version. A rebuild belongs in an isolated copy or a new output version; existing checkpoints and scientific files are never overwritten to make a rerun convenient.