ARCHIVED RESEARCH PROJECT · completed 2026-09-28 · Do distinctive Voynich constructions match medieval Latin abbreviation signs catalogued by Cappelli? · Back to Home

PROJECT SUMMARY

QuestionDo distinctive Voynich constructions match medieval Latin abbreviation signs catalogued by Cappelli?
MethodAll 14,780 local Cappelli JPGs were byte-verified against the existing archive. The workbook has 14,329 unique IDs; 14,327 have local images, 453 images lack workbook rows, and two workbook IDs lack images. Join exclusively by explicit image ID, never row number. All displayed comparisons…
Important numbers20 voynich fingerprint · 27 scored comparison · 400 ranked whole comparison · 400 ranked component comparison
Finding and conclusionThe review is coverage-balanced: the rank-one component candidate for each of 20 fingerprints, plus six informative rank-one whole-entry candidates and the only lower-ranked numeric-gate-positive component. It is not the globally highest 27 scores. Every pair shows two Voynich exemplars,…
Limitsvalidation.json records unchanged SHA-256 for all 213 Yale TIFFs, unchanged full extraction manifest and 213 neutral feature inputs, unchanged Cappelli sources/archive/workbook, exact native crop/context pixels, valid control source references and reproduced 800 ranking records. The…
MeaningConclusion: Cappelli supplies credible component-level visual analogues, and one plausible standalone hooked-form variant. This experiment does not establish that multiple distinctive Voynich constructions use recognizable medieval European abbreviation machinery beyond chance. A modest pooled conditional similarity signal warrants interest, but neither individual controlled matches nor convergent function/tradition establish a system. This is not evidence that such a system is absent. Confidence is moderate in the displayed component similarities, low in historical or functional attribution.

Read the findings preserved from this completed study.

EXPLORE ALL RESULTS

847 accessible result rows from 847 authoritative meaningful result rows in the declared scope. Review cards are separate.

All rows from the stated saved result sets are accessible; detailed machine arrays remain in the linked source files.

REVIEW NEEDED

27 items need review in the preserved curated card set.

This curated view is separate from the complete meaningful results listed above.

VOYNICH PROJECT

Archived research project · completed 2026-09-28

ARCHIVED RESEARCH PROJECT

Voynich–Cappelli visual fingerprint comparison

Do distinctive Voynich constructions match medieval Latin abbreviation signs catalogued by Cappelli?

OPEN THE REVIEW FINDINGS AS COMPLETED

HUMAN REVIEW · AS PRESENTED WHEN CURRENT

Voynich–Cappelli visual fingerprint comparison

27 source-linked cards. Display images are embedded; original TIFFs and scientific records remain authoritative in the project folder.

27 comparisons: 0 very strong matches, 1 provisional visual variant, 13 component-level resemblances and 13 weak/rejected leads. Cappelli expansions apply only to Cappelli. No individual fingerprint survives multiple-comparison correction; a marginal pooled component signal does not establish a shared abbreviation system or tradition.

FINDINGS FROM THE COMPLETED STUDY · public presentation of [preserved evidence reference]

Voynich–Cappelli visual fingerprint comparison

Completed on the authorized MacBook, 2026-09-28. Independent experiment: [preserved evidence reference]. 27-pair visual review; native contact sheet.

Conclusion: Cappelli supplies credible component-level visual analogues, and one plausible standalone hooked-form variant. This experiment does not establish that multiple distinctive Voynich constructions use recognizable medieval European abbreviation machinery beyond chance. A modest pooled conditional similarity signal warrants interest, but neither individual controlled matches nor convergent function/tradition establish a system. This is not evidence that such a system is absent. Confidence is moderate in the displayed component similarities, low in historical or functional attribution.

What was compared, and what stayed blind

Twenty separate recurring constructions were selected from the independently extracted neutral manuscript visual index, with two native Yale exemplars each. Paired enclosed/narrow lobes, opposed open bends, three curve/hook/loop-tail constructions, four tall-stem constructions, two tall-plus-horizontal compounds, six repeated-stroke constructions and two longer combinations remain separate. No identities, readings or meanings were assigned. Examples include f016r and later folios; all source paths use canonical folio filenames. Several candidate boundaries include neighboring material, and recurrence is provisional.

The first selection contained 22 constructions. Native-source checking rejected an edge/illustration companion and an unsuitable standalone-horizontal companion before viewing Cappelli matches or ranking. The original selection, freeze and correction remain in pre_ranking_initial/ and pre_ranking_correction.json. This left 20 constructions and 40 exemplars. The remaining horizontal compounds do not substitute for an accepted standalone bench unit.

Fingerprint definitions, geometry policy and code were sealed before matching. Ranking and all controls were separately sealed before reading Cappelli expansion/date/category records or researching the corpus externally. The workbook header schema had been inspected; no semantic record values informed query selection or ranking. This is label-isolated retrieval, not an assertion that the analyst lacks prior historical knowledge. No EVA, Turkish, medical theory or proposed decipherment was used to influence matching. Medical words/category labels occur only as unfiltered source metadata after the freeze.

Corpus and provenance

All 14,780 local Cappelli JPGs were byte-verified against the existing archive. The workbook has 14,329 unique IDs; 14,327 have local images, 453 images lack workbook rows, and two workbook IDs lack images. Join exclusively by explicit image ID, never row number. All displayed comparisons have a unique workbook record; the unmatched image in the complete whole-entry ranking remains unspecified.

The original publisher confirms this ID join, identifies the image source as Cappelli's Lexicon Abbreviaturarum, Leipzig 1928, and cautions that crowdsourced metadata can contain errors. The frozen local snapshot was retained despite different live corpus counts. Ad fontes, University of Zurich. The publisher's online entries provide book-page references; the local page_id is retained as that reference. A direct scan URL was not recovered, and individual manuscript shelfmarks are unavailable. Cappelli online.

Archive SHA-256: af1316fa3666bf4f1561ee45777290b191a35f1b91d37f4843ede7f3e7722115. Workbook SHA-256: 18c45b60722199731a44c3bb03abe7f79764ebfd8a9c7c7a3c87f62c78d25a77.

Raw workbook records, categories, period strings, uncertainty flags and page references are preserved in workbook_metadata.json; join coverage is in metadata_join_provenance.json. Dimensions independently agree for 14,306 mapped images; 21 workbook rows have invalid zero dimension fields, retained and flagged in metadata_dimensions_check.json. None is displayed. Actual image dimensions drive processing. Century summaries accept canonical Roman tokens only; malformed dates remain unknown and qualifiers remain untranslated. No missing regional identity was inferred.

Deterministic retrieval

Every Cappelli entry contributes a whole-ink box. A separate pool contains windows spanning one to three adjacent connected components, without query-directed cuts: 14,780 whole windows plus 122,803 component windows. A window can include context that is not part of its selected mask; native pixels and component IDs remain inspectable.

Aspect-preserving 32×32 support descriptors combine DCT shape, horizontal/vertical projections, cavities/components, skeleton chamfer distance and tolerant skeleton support overlap. The field named tolerant_Dice is a bidirectional neighborhood-hit measure, not conventional intersection Dice. Upright geometry is primary; no semantic alignment, reflection, rotation or elastic warping. The first-stage descriptor search considers every window, then reranks a union of the top 50 per exemplar and stratum. This approximate shortlist can miss a good counterpart.

Final consensus weights the weaker exemplar heavily: 0.75×minimum + 0.25×mean. Twenty distinct Cappelli image IDs are saved per fingerprint per whole/component stratum: 800 ranked records. A high score is a retrieval lead, not identity confidence. The strict frozen multi-feature gates classify 799/800 saved candidates as weak and one as a component-level suggestion; none is a whole-entry strong/plausible match. The separate native-image assessment below identifies limited qualitative analogues; its upgrades are exploratory judgments after the freeze, not successful preregistered tests or human-approved identities.

Visual findings and metadata

The review is coverage-balanced: the rank-one component candidate for each of 20 fingerprints, plus six informative rank-one whole-entry candidates and the only lower-ranked numeric-gate-positive component. It is not the globally highest 27 scores. Every pair shows two Voynich exemplars, the native Cappelli window and its complete entry/context. Category counts: 0 very strong visual matches; 1 provisional plausible scribal variant; 13 component-level resemblances; 13 weak/coincidental leads. Weak cases expose high-score failures rather than hide them. In particular, Q20 component rank two (Cappelli 13877) passes the fixed component gate but contains only overbars: native context rejects it as a counterpart to the full Voynich construction.

Voynich constructionCappelli leadCappelli-only metadataAssessment
Q01 paired enclosed lobes, including f016r3303, upper componenteadem, - eodem; XIV f.; page_id 114Two lobes and central join; differing exits/fill. Detached mark within a larger entry.
Q04 upper loop with descending curved tail12933, terminal componentversus; XIV p.; page_id 397Loop plus curved tail; raised terminal placement differs.
Q05 open curve with lower hook107, whole entrycon...; VIII-XII; page_id 68Plausible visual variant, with curvature/closure differences; numeric gate failed.
Q08 paired tall stems/crossing bar1899, whole entrycomes; XII; page_id 59Only component machinery: Cappelli lacks the paired-stem arrangement.
Q10 paired stems/upper loops2071, componentallegationi; XV; page_id 14Several shared features, but ordinary repeated written components in a longer abbreviation.
Q17 repeated strokes/high terminal sweep3164, componentdifferentiam; XV p; page_id 108A good fragment analogy; not demonstrated as one abbreviation sign.
Q18 repeated run/rising terminal trace7191, whole entryminus; XIV; page_id 219General arrangement shared; detached/raised Cappelli ending differs.

Expansions apply only to Cappelli. Workbook expansions describe complete entries, not every cropped component's function. Their presence in an abbreviation dictionary does not make every ordinary loop, stem or minim an abbreviation sign. No Voynich functional equivalence or translation follows.

The narrow-lobe barred form, complex tall structures and longer horizontal/loop combinations lack convincing counterparts here. Tall-loop/bar fragment similarities often fail the complete geometry. Generic repeated strokes are especially easy to match; shared short strokes alone receive little evidential weight.

Chance and retrieval controls

One hundred known Cappelli images subjected to slight shear/blur retrieved their correct normalized shape equivalents at rank one (ties allowed). This validates retrieval within the printed corpus style, not recall across manuscript handwriting and tiny book reproductions.

Each query was compared with 40 matched real neutral-candidate pairs (800 pairs total), matched approximately on aspect, raw cavity count and occupancy; paired companions come from different source surfaces. Controls search the same entire whole/component pools, including all window-selection opportunities. They are unaccepted geometric regions, can include texture or unknown target relatives, and their automatically selected companions can be more coherent than the chosen query pair. They are not a clean ordinary-letter or unrelated-script baseline. Three reflected/rotated controls per fingerprint provide additional sensitivity checks, not a realistic historical null.

No individual query survives the 20-query Benjamini–Hochberg adjustment, even adjusting whole and component strata separately. Minimum adjusted p-values are 0.379 whole and 0.244 component. The finite-control minimum unadjusted p is 1/41; this coarse resolution limits power, so failure to pass is not strong evidence of absence.

Equal weighting of six coarse morphological families gives an exploratory pooled signal: whole mean 0.714 versus control median 0.697, p=0.049; component mean 0.760 versus 0.734, p=0.024. Accounting for the two aggregate outcomes gives Bonferroni values 0.098 and 0.049. The component result is borderline under this conditional null. Forty draws, overlapping constructions and a convenience fingerprint set limit it; morphology blocks do not create historically independent correspondences. This signal cannot establish that matched pieces are abbreviation signs or that several independent units share a historical system.

Period, region and function

Fourteenth/fifteenth-century dates occur for 15/20 component winners, versus 528/762 mapped control winners and 9,189/14,327 corpus images. Whole winners have 11/19, versus 465/746 mapped controls. These are descriptive, overlapping-century memberships, not independent enrichment tests. The date concentration largely follows the source background; component winners are 19/20 Latin, while the mapped corpus is 13,585/14,327 Latin.

The whole-entry winners include isolated Sicilian and Visigothic categories, but these are weak geometric leads from different periods, not a convergent regional result. Component winners have almost no regional/script metadata. The promising displayed forms span different complete-entry expansions and placements; there is no demonstrated common functional system. Region and manuscript tradition are largely absent from this dataset, so absence of a detected convergence is not evidence of historical absence.

Validation and limits

The review is a fresh analyst assessment awaiting human inspection. Small source reproductions, imperfect segmentation, a narrow inventory and approximate retrieval constrain both positive and negative claims. Partial source views can repeat physical ink. No unrelated-script corpus, matched ordinary medieval-writing corpus, source-manuscript confirmation, or validated pen-stroke/unit inventory was available. Similarity alone cannot distinguish shared scribal practice from broad handwriting geometry.

Answer to the primary question: multiple Voynich constructions have close Cappelli pieces, but unusually close complete counterparts and a shared abbreviation system are not established. The pooled component signal is suggestive, conditional evidence; confidence that it proves medieval European scribal machinery beyond chance is low.

Review images

Embedded from the preserved project files (display copies; originals remain in the project folder).