PROJECT SUMMARY
QuestionDo the recurring visual forms of Voynich writing, taken together, resemble any known historical writing tradition more strongly than chance would produce?
Method79 manuscripts, 212 pages, 27 traditions. Every item was chosen from catalogue metadata by date closeness to 1420, with distinct shelfmarks. Pages were chosen by a script-neutral text-density score; every page is hashed, with its URL, licence, date and place. - Sources: OPenn (Penn and…
Important numbers55 fingerprint form · 67 fingerprint variant · 27 tradition result · 445 strong match
Finding and conclusionNo. The pre-declared verdict is UNCONFIRMED CONCENTRATION: the top tradition (Georgian) passed only one of four sealed criteria. No tradition, pair of traditions or functional notation system shows system-level convergence that survives the controls. By shape, Voynich writing is not…
Limits- Small samples: 2–4 manuscripts and 1–3 pages per tradition. - Pixel shape only: no stroke order, ductus or position. - The method has modest power. Controls rank their own tradition first 37% of the time, and their own script family 56% (chance ≈ 4% and ≈ 10%). Latin sub-traditions are…
MeaningNo. The pre-declared verdict is UNCONFIRMED CONCENTRATION: the top tradition (Georgian) passed only one of four sealed criteria. No tradition, pair of traditions or functional notation system shows system-level convergence that survives the controls. By shape, Voynich writing is not…
Read the findings preserved from this completed study.
FINDINGS FROM THE COMPLETED STUDY · public presentation of [preserved evidence reference]
Global script fingerprint search — v001 report
Question
Do the recurring visual forms of Voynich writing, taken together, resemble any known historical writing tradition more strongly than chance would produce? The search looks for system-level convergence, not isolated look-alikes. Nothing was translated, and no sounds were assigned.
Answer
No. The pre-declared verdict is UNCONFIRMED CONCENTRATION: the top tradition (Georgian) passed only one of four sealed criteria. No tradition, pair of traditions or functional notation system shows system-level convergence that survives the controls. By shape, Voynich writing is not identifiable as, or as a blend of, any script in this 79-manuscript corpus. Its external resemblances are the loops, strokes, hooks and doubled curves that ordinary handwriting shares.
Review page: [preserved evidence reference]. Data: [preserved evidence reference]. Validation: (PASS).
1. Source dataset
The newest sealed dataset is v005 (final_seal.json, sealed 2026-09-29 18:01 UTC, QUALIFIED_ENGINEERING_PASS_FREEZE_EXCEPTION; database sha256 2d2a1b47…, rechecked). v006 was not used. Its own RUN_STATUS.md says it is in progress and unsealed. v006 also found that v005's family graph was fed mislabelled feature columns. For that reason the v005 family IDs were not used as units: the fingerprint was rebuilt from pixels.
2. The Voynich fingerprint (frozen before any external image existed)
- Population: the 98,240
default_sequence_include=1 high-confidence observations. Each observation's own grouped threshold-12 ink support was taken at native resolution (the same rule as v005 stage 7b). Sizes are measured in page units, where one unit is the page's median writing height. - Scale gate: 0.25 ≤ height ≤ 3.5 units and width ≤ 2.5 units. The 7,717 boxes containing another text line's centre (probable two-line merges) were excluded. 86,521 observations remained.
- Neutral vector: 128-d gradient histogram (HOG), log relative height, aspect ratio and a hole class. Clustering used k-means with K = 240 over three seeds.
- Selection by recurrence only: a cluster qualified with ≥ 60 members, ≥ 15 folios, ≥ 4 of 6 sequence sixths, and a Jaccard overlap ≥ 0.5 with its best match in both other seeds. 67 clusters qualified. They merged into 55 fingerprint forms, and each constituent cluster was kept as a separately searchable variant (67). The forms cover 25,811 observations.
- Matching distance: a skeleton whole-form distance (the larger of the two directed mean distances) plus HOG, equally weighted and each scaled by its median same-form distance. Two forms are comparable only if aspect ratio and relative height each differ by ≤ ×1.5. This combination separated same-form from different-form Voynich pairs with AUC 0.935.
- Thresholds, from Voynich-internal variation:
- VERY STRONG ≤ 2.003, the median same-form distance. Only 2.8% of different-form pairs fall this close.
- STRONG VARIANT ≤ 2.423, the 75th percentile (7.8% of different-form pairs).
- Part score ≤ 0.765.
- Drafts and corrections before sealing:
- Draft A contained two-line merges.
- Draft B's native-resolution topology made plain loops non-generic.
- A symmetric mean chamfer let wide forms match any horizontal stroke; this was replaced by the aspect gate and the max-directed distance.
- All drafts are preserved in
work/. - Genericity: the sealed primitive rule flagged 0/55 forms, including plain loops. Amendment A1, made before any external image existed, therefore added a ubiquity rule: a form is generic if it matches strongly in ≥ 50% of traditions. 8 forms became generic: VFP001–004, 006, 014, 015 and 045. 47 are distinctive.
- Coordinator review of the contact sheets: 5 forms are two-line stacks the line filter missed (VFP033, 035, 038, 044, 047), and 3 are internally inconsistent (VFP048, 050, 053). They were kept in the sealed fingerprint. A post-hoc sensitivity analysis without them was declared in Seal 2.
- Descriptive categories (coordinator):
- 9 gallows / tall looped
- 6 bench / crossbar
- 6 double-loop / S-like
- 5 9 / hook
- 6 loop + stem
- 7 minim / repeated stroke
- 5 small mark
- 11 other
Seal 1: protocol/SEAL.json. Amendment A1: protocol/AMENDMENT_A1.json.
3. Comparison corpus
79 manuscripts, 212 pages, 27 traditions. Every item was chosen from catalogue metadata by date closeness to 1420, with distinct shelfmarks. Pages were chosen by a script-neutral text-density score; every page is hashed, with its URL, licence, date and place.
- Sources: OPenn (Penn and partners), the Walters (CC0), the Vatican Library IIIF service, Heidelberg IIIF and Wikimedia Commons.
- Traditions:
- Latin textualis; Latin/German cursiva; Italian hands.
- Astronomical tables and numerals; accounting; alchemical; medical recipes; Tironian shorthand; the Arabic secret alphabet of the Sirr al-asrār (LJS 456 pp. 199/201/203 and LJS 459 ff. 127v–128r).
- Greek, Cyrillic, Glagolitic, Coptic, Armenian, Georgian.
- Hebrew, Samaritan, Syriac, Arabic, Persian, Ottoman, Ethiopic.
- Avestan/Pahlavi, Old Uyghur, Mongolian, Devanagari, Southeast Asian.
- Corrections, each preserved under
sources/superseded/: - A ledger-save bug dropped 22 records. They were re-acquired, and all 52 re-downloaded pages came back byte-identical.
- Visual-QC page exclusions: a faded miniature, silk curtains and a stained blank page.
- A light-on-dark polarity rule for gold-on-black Thai.
- A century-date parser fix, which re-selected the Greek, Ethiopic and Ottoman items.
- A binding-canvas exclusion.
- Two replacements for mixed or illustrated items.
- Gaps:
- Several traditions survive here only in post-1600 copies: Samaritan, Avestan/Pahlavi, Devanagari and Thai.
- Old Uyghur, Mongolian and Tironian are single photographs.
- Four sources yielded no stable forms: TIR-3, ACC-3, UYG-2 and MON-3.
- Sogdian (possible duplicate item), Tibetan, dedicated cipher keys and alchemical sign tables could not be obtained as independent, dated manuscripts.
Each manuscript received its own fingerprint by the same rules (≤ 4,000 components; K = min(60, n/30); ≥ 20 members on ≥ 2 pages; seed-stable). Seal 2 (protocol/SEAL_2.json) froze the corpus, the fingerprints and the code before any matching.
4. Results (sealed rules)
| Tradition | Coverage of 47 distinctive forms | Forms matched in ≥ half its MSS | p (uncorrected) |
|---|
| Georgian (2 MSS) | 18.1% | 15 | 0.009 |
| Italian Latin/vernacular | 12.8% | 1 | 0.056 |
| Alchemical manuscripts | 12.1% | 4 | 0.077 |
| Astronomical tables & numerals | 11.3% | 1 | 0.111 |
| Samaritan | 11.3% | 3 | 0.108 |
| Glagolitic / Greek | 9.9% each | 3 each | 0.21 / 0.20 |
| … | … | … | … |
| Devanagari / Old Uyghur / Tironian / accounting / Mongolian | ≤ 2.1% | ≤ 2 | ≥ 0.84 |
- Family-wise permutation test (10,000 label shuffles, max statistic): p = 0.108.
- Controls: each of the 75 manuscript fingerprints was run as the unknown against all the others.
- Median own-tradition coverage: 17.9%.
- Median coverage of the best other tradition: 25.0%; 95th percentile 52.5%.
- Voynich's best tradition (18.1%) sits below the median coincidental cross-script resemblance of ordinary manuscripts.
- Geometric controls:
- Mirrored Voynich still ranks Georgian first (11.1%, p = 0.66). The bootstrap 95% CI of real minus mirrored is [−0.032, 0.170], which includes 0.
- Upside-down Voynich ranks Hebrew first (15.3%, p = 0.56).
- Voynich processed by the external pipeline yields only 6 distinctive stable forms (p = 0.999).
- Decision criteria:
- (a) permutation: fail
- (b) above the coincidence level: fail
- (c) ≥ 5 forms: pass (15)
- (d) orientation-specific: fail
- Mixed ancestry: the best pair (secret alphabet + Georgian) covers 38.3%, against a control 95th percentile of 77.5%. Not supported.
- Modification profile: 85.1% of forms have a whole-form counterpart somewhere, 14.9% component-only and 0% none. This is within the control range, so there is no evidence of "Latin machinery + modifications" or any other hybrid.
- Aggregation by family: Caucasian 11.5%, Latin 9.0%, Semitic 8.5%, Greek-derived 7.6%, Cipher 7.4%, Notation 7.0%, Ethiopic 6.4%, Brahmic 5.0%, Arabic-script 4.7%, Iranian–Central Asian 2.4%.
- Aggregation by period: 1300–1500 7.4%, before 1300 6.2%, after 1700 4.5%. No period effect.
Post-hoc (declared):
- Removing the 8 flagged forms: Georgian 16.7%, 11 forms, p = 0.24.
- The Georgian set has only 2 manuscripts, so "≥ half" means one manuscript. The Georgian drivers are hooks, loops and 9-like forms (VFP011, 012, 013, 019, 038, 048, 052, …). Three of them are flagged artefacts or inconsistent forms.
Where the resemblance lives:
- The seven unflagged gallows forms each match strongly in at most 3 of 75 manuscripts, and are never matched across a tradition.
- 9-like hooks match in 8–10 manuscripts, spread across Georgian, Persian, Thai, Hebrew and the secret alphabet.
- The figure-8 matches in 24 manuscripts, and the minim runs in 19–29.
- Inspection of the 445 strong-match pairs shows that the distance classes often admit simple shapes (e.g.
cco against Latin ma, o against o). Class names are distance bands, not established constructions.
6. Limits
- Small samples: 2–4 manuscripts and 1–3 pages per tradition.
- Pixel shape only: no stroke order, ductus or position.
- The method has modest power. Controls rank their own tradition first 37% of the time, and their own script family 56% (chance ≈ 4% and ≈ 10%). Latin sub-traditions are confused with each other.
- The extractor has no material classifier. On illustrated Voynich pages it finds 2–4× more pieces than v005, so external fingerprints may include some non-writing.
- One coordinator; no independent human ratings.
- The simulated browser check is not a physical-iPhone test.
This result rules out strong, obvious, system-level visual kinship with the scripts tested. It does not rule out subtle ancestry, scripts outside the corpus, or relationships visible only in construction sequences.
7. Next experiment
Measure construction sequences rather than static shapes. That means stroke order and ductus, where each form joins, and what regularly comes next, all on the same hands. Then run matched tests with many more manuscripts for the few weak leads (Georgian nuskhuri, Italian and Central-European Latin hands, alchemical notation), with blind human ratings. The power of the instrument should be shown first, on known derived-script pairs (Greek→Cyrillic/Coptic, Aramaic→Syriac/Hebrew), so that a null result means something.