ARCHIVED RESEARCH PROJECT · completed 2026-09-28 · Do distinctive Voynich constructions match medieval Latin abbreviation signs catalogued by Cappelli? · Back to Home
PROJECT SUMMARY
QuestionDo distinctive Voynich constructions match medieval Latin abbreviation signs catalogued by Cappelli?
MethodAll 14,780 local Cappelli JPGs were byte-verified against the existing archive. The workbook has 14,329 unique IDs; 14,327 have local images, 453 images lack workbook rows, and two workbook IDs lack images. Join exclusively by explicit image ID, never row number. All displayed comparisons…
Finding and conclusionThe review is coverage-balanced: the rank-one component candidate for each of 20 fingerprints, plus six informative rank-one whole-entry candidates and the only lower-ranked numeric-gate-positive component. It is not the globally highest 27 scores. Every pair shows two Voynich exemplars,…
Limitsvalidation.json records unchanged SHA-256 for all 213 Yale TIFFs, unchanged full extraction manifest and 213 neutral feature inputs, unchanged Cappelli sources/archive/workbook, exact native crop/context pixels, valid control source references and reproduced 800 ranking records. The…
MeaningConclusion: Cappelli supplies credible component-level visual analogues, and one plausible standalone hooked-form variant. This experiment does not establish that multiple distinctive Voynich constructions use recognizable medieval European abbreviation machinery beyond chance. A modest pooled conditional similarity signal warrants interest, but neither individual controlled matches nor convergent function/tradition establish a system. This is not evidence that such a system is absent. Confidence is moderate in the displayed component similarities, low in historical or functional attribution.
27 source-linked cards. Display images are embedded; original TIFFs and scientific records remain authoritative in the project folder.
27 comparisons: 0 very strong matches, 1 provisional visual variant, 13 component-level resemblances and 13 weak/rejected leads. Cappelli expansions apply only to Cappelli. No individual fingerprint survives multiple-comparison correction; a marginal pooled component signal does not establish a shared abbreviation system or tradition.
Q05-whole · 1 of 27plausible scribal variant
Q05 · Open upper curve and angular lower hook
curved · whole Cappelli window
A standalone open upper curve and recurved lower hook provide a plausible visual variant. Voynich has a more angular/filled lower part and variable upper closure. This qualitative lead fails the frozen numeric plausible-match gate; no identity or functional equivalence is accepted. Frozen consensus score 0.694; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q01-component · 2 of 27component-level resemblance
Q01 · Paired enclosed lobes
paired · component Cappelli window
Two stacked enclosed lobes and a narrow central join recur in the detached upper mark. Cappelli has a thin top exit; Voynich closure/fill varies. This is a component comparison, not the full Cappelli entry. Frozen consensus score 0.742; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q04-component · 3 of 27component-level resemblance
Q04 · Upper loop and descending curved tail
curved · component Cappelli window
An upper loop and a long descending curved exit occur together. The Cappelli mark is raised at the right of its entry; Voynich candidate boundaries and attachment differ. Frozen consensus score 0.810; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q05-component · 4 of 27component-level resemblance
Q05 · Open upper curve and angular lower hook
curved · component Cappelli window
Open upper arc and recurved lower hook resemble part of the Cappelli entry. Its full context is an ordinary initial written component within an abbreviated word, not evidence that this component is an abbreviation sign. Frozen consensus score 0.820; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q06-component · 5 of 27component-level resemblance
Q06 · Curved upper trace and hooked base
curved · component Cappelli window
Opposed curves with a hooked lower end resemble the compact Voynich construction. Cap shape, fill and possible component boundaries differ. Frozen consensus score 0.799; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
A tall upright, lateral upper loop and crossing trace share component machinery. Cappelli lacks the Voynich paired-stem arrangement and has a full lower word; this whole-entry result is only a component-level resemblance. Frozen consensus score 0.747; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q09-component · 7 of 27component-level resemblance
Q09 · Single tall stem with upper side loop
tall · component Cappelli window
Tall stem, upper side loop and a lower connection coexist. Cappelli includes extra neighboring material and a differently placed crossing trace; this is a structural fragment analogy. Frozen consensus score 0.675; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q10-component · 8 of 27component-level resemblance
Q10 · Paired stems with compact upper loops
tall · component Cappelli window
Paired upright strokes, small upper loops and a crossing horizontal trace occur together. The Cappelli fragment belongs to an ordinary repeated-letter sequence; it does not identify a special abbreviation unit. Frozen consensus score 0.712; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Tall stem, upper loop and crossing trace resemble part of the Voynich construction. The lower horizontal element and whole-entry boundaries differ. Frozen consensus score 0.687; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q14-component · 10 of 27component-level resemblance
Q14 · Repeated short strokes under high arch
repeated · component Cappelli window
A high rightward arch over low short strokes is shared. The Cappelli fragment cuts a larger written construction and loses its beginning; minim count and attachments remain uncertain. Frozen consensus score 0.796; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q15-component · 11 of 27component-level resemblance
Q15 · Compact repeated-stroke construction
repeated · component Cappelli window
Short repeated low strokes are shared. This is a common component-level geometry, with little distinctiveness and no demonstrated abbreviation function. Frozen consensus score 0.760; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q16-component · 12 of 27component-level resemblance
Q16 · Arched repeated-stroke construction
repeated · component Cappelli window
Repeated low strokes and a high curved terminal trace are shared. Cappelli shows a detached/raised fragment; Voynich attachment and count differ. Frozen consensus score 0.797; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q17-component · 13 of 27component-level resemblance
Q17 · Wide repeated strokes with terminal curve
repeated · component Cappelli window
Several low strokes and a high terminal sweep curling leftward occur together. The fragment is a sequence inside a larger abbreviated entry, not a demonstrated single abbreviation sign. Frozen consensus score 0.862; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
A repeated low-stroke run followed by an upper curved mark resembles the general arrangement. The Cappelli terminal mark is raised and detached; Voynich connection differs. Common strokes limit distinctiveness. Frozen consensus score 0.801; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q02-component · 15 of 27weak/coincidental resemblance
Q02 · Narrow paired lobes
paired · component Cappelli window
The fragment contains tall crossed stems and a rounded lower region; the narrow paired-lobe construction is not preserved. Similar global silhouette is insufficient. Frozen consensus score 0.735; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Two lobed regions and an oblique join are suggestive, but Cappelli has an additional crossing bar and different closure. A shared broad outline does not establish the same form. Frozen consensus score 0.748; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q03-component · 17 of 27weak/coincidental resemblance
Q03 · Opposed open bends
curved · component Cappelli window
The small rounded Cappelli fragment lacks the two opposed open bends. The whole entry is also geometrically different. Frozen consensus score 0.701; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q07-component · 18 of 27weak/coincidental resemblance
Q07 · Paired tall stems with upper loops
tall · component Cappelli window
Cappelli has a tall stem and lobed right side; the Voynich paired stems, upper loops and lower join are not jointly preserved. Frozen consensus score 0.659; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q08-component · 19 of 27weak/coincidental resemblance
Q08 · Paired tall stems with crossing bar
tall · component Cappelli window
The component has an oblique loop and rightward upper trace, but lacks the paired vertical stems and crossing relationship. Frozen consensus score 0.719; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q11-component · 20 of 27weak/coincidental resemblance
Q11 · Tall stem joined to lower horizontal
bench_compound · component Cappelli window
The fragment is mainly low repeated strokes with an upper mark. It fails to preserve the distinctive tall stem joined to a low horizontal construction. Frozen consensus score 0.705; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q12-component · 21 of 27weak/coincidental resemblance
Q12 · Paired tall stems over lower horizontal
bench_compound · component Cappelli window
Oblique tall strokes dominate Cappelli. The paired upright loops and lower horizontal construction are not preserved together. Frozen consensus score 0.790; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q13-component · 22 of 27weak/coincidental resemblance
Q13 · Short strokes beside sweeping curve
repeated · component Cappelli window
Several low strokes recur, but the terminal sweep and its position do not agree. This is too generic to identify a counterpart. Frozen consensus score 0.724; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q18-component · 23 of 27weak/coincidental resemblance
Q18 · Repeated strokes with rising terminal trace
repeated · component Cappelli window
The component window keeps only a short curved trace and small low support. It omits the long repeated-stroke run central to the Voynich construction. Frozen consensus score 0.814; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q19-component · 24 of 27weak/coincidental resemblance
Q19 · Loop plus short-stroke run and rounded end
combination · component Cappelli window
A tiny upper mark receives a high score but omits the initial loop, extended stroke run and rounded ending. Whole construction not matched. Frozen consensus score 0.763; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q20-component · 25 of 27weak/coincidental resemblance
Q20 · Low horizontal construction and closed end
combination · component Cappelli window
The low fragment lacks the closed terminal enclosure and its relationship to the horizontal construction. Score alone is misleading. Frozen consensus score 0.848; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Q20-component-rank02 · 26 of 27weak/coincidental resemblance
Q20 · Low horizontal construction and closed end
combination · component Cappelli window
The only top-20 candidate passing a frozen numeric component gate is a window containing overbars alone. It omits the closed terminal construction entirely. Native context rejects the proposed counterpart; this is a numeric-gate false lead. Frozen consensus score 0.848; numeric gate: component-level resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Repeated capital-like stems and punctuation do not preserve the low horizontal construction plus closed end. This is a useful high-score false lead. Frozen consensus score 0.798; numeric gate: weak/coincidental resemblance. Qualitative category is a post-freeze native-image assessment, not accepted identity.
Conclusion: Cappelli supplies credible component-level visual analogues, and one plausible standalone hooked-form variant. This experiment does not establish that multiple distinctive Voynich constructions use recognizable medieval European abbreviation machinery beyond chance. A modest pooled conditional similarity signal warrants interest, but neither individual controlled matches nor convergent function/tradition establish a system. This is not evidence that such a system is absent. Confidence is moderate in the displayed component similarities, low in historical or functional attribution.
What was compared, and what stayed blind
Twenty separate recurring constructions were selected from the independently extracted neutral manuscript visual index, with two native Yale exemplars each. Paired enclosed/narrow lobes, opposed open bends, three curve/hook/loop-tail constructions, four tall-stem constructions, two tall-plus-horizontal compounds, six repeated-stroke constructions and two longer combinations remain separate. No identities, readings or meanings were assigned. Examples include f016r and later folios; all source paths use canonical folio filenames. Several candidate boundaries include neighboring material, and recurrence is provisional.
The first selection contained 22 constructions. Native-source checking rejected an edge/illustration companion and an unsuitable standalone-horizontal companion before viewing Cappelli matches or ranking. The original selection, freeze and correction remain in pre_ranking_initial/ and pre_ranking_correction.json. This left 20 constructions and 40 exemplars. The remaining horizontal compounds do not substitute for an accepted standalone bench unit.
Fingerprint definitions, geometry policy and code were sealed before matching. Ranking and all controls were separately sealed before reading Cappelli expansion/date/category records or researching the corpus externally. The workbook header schema had been inspected; no semantic record values informed query selection or ranking. This is label-isolated retrieval, not an assertion that the analyst lacks prior historical knowledge. No EVA, Turkish, medical theory or proposed decipherment was used to influence matching. Medical words/category labels occur only as unfiltered source metadata after the freeze.
Corpus and provenance
All 14,780 local Cappelli JPGs were byte-verified against the existing archive. The workbook has 14,329 unique IDs; 14,327 have local images, 453 images lack workbook rows, and two workbook IDs lack images. Join exclusively by explicit image ID, never row number. All displayed comparisons have a unique workbook record; the unmatched image in the complete whole-entry ranking remains unspecified.
The original publisher confirms this ID join, identifies the image source as Cappelli's Lexicon Abbreviaturarum, Leipzig 1928, and cautions that crowdsourced metadata can contain errors. The frozen local snapshot was retained despite different live corpus counts. Ad fontes, University of Zurich. The publisher's online entries provide book-page references; the local page_id is retained as that reference. A direct scan URL was not recovered, and individual manuscript shelfmarks are unavailable. Cappelli online.
Raw workbook records, categories, period strings, uncertainty flags and page references are preserved in workbook_metadata.json; join coverage is in metadata_join_provenance.json. Dimensions independently agree for 14,306 mapped images; 21 workbook rows have invalid zero dimension fields, retained and flagged in metadata_dimensions_check.json. None is displayed. Actual image dimensions drive processing. Century summaries accept canonical Roman tokens only; malformed dates remain unknown and qualifiers remain untranslated. No missing regional identity was inferred.
Deterministic retrieval
Every Cappelli entry contributes a whole-ink box. A separate pool contains windows spanning one to three adjacent connected components, without query-directed cuts: 14,780 whole windows plus 122,803 component windows. A window can include context that is not part of its selected mask; native pixels and component IDs remain inspectable.
Aspect-preserving 32×32 support descriptors combine DCT shape, horizontal/vertical projections, cavities/components, skeleton chamfer distance and tolerant skeleton support overlap. The field named tolerant_Dice is a bidirectional neighborhood-hit measure, not conventional intersection Dice. Upright geometry is primary; no semantic alignment, reflection, rotation or elastic warping. The first-stage descriptor search considers every window, then reranks a union of the top 50 per exemplar and stratum. This approximate shortlist can miss a good counterpart.
Final consensus weights the weaker exemplar heavily: 0.75×minimum + 0.25×mean. Twenty distinct Cappelli image IDs are saved per fingerprint per whole/component stratum: 800 ranked records. A high score is a retrieval lead, not identity confidence. The strict frozen multi-feature gates classify 799/800 saved candidates as weak and one as a component-level suggestion; none is a whole-entry strong/plausible match. The separate native-image assessment below identifies limited qualitative analogues; its upgrades are exploratory judgments after the freeze, not successful preregistered tests or human-approved identities.
Visual findings and metadata
The review is coverage-balanced: the rank-one component candidate for each of 20 fingerprints, plus six informative rank-one whole-entry candidates and the only lower-ranked numeric-gate-positive component. It is not the globally highest 27 scores. Every pair shows two Voynich exemplars, the native Cappelli window and its complete entry/context. Category counts: 0 very strong visual matches; 1 provisional plausible scribal variant; 13 component-level resemblances; 13 weak/coincidental leads. Weak cases expose high-score failures rather than hide them. In particular, Q20 component rank two (Cappelli 13877) passes the fixed component gate but contains only overbars: native context rejects it as a counterpart to the full Voynich construction.
Voynich construction
Cappelli lead
Cappelli-only metadata
Assessment
Q01 paired enclosed lobes, including f016r
3303, upper component
eadem, - eodem; XIV f.; page_id 114
Two lobes and central join; differing exits/fill. Detached mark within a larger entry.
Q04 upper loop with descending curved tail
12933, terminal component
versus; XIV p.; page_id 397
Loop plus curved tail; raised terminal placement differs.
Q05 open curve with lower hook
107, whole entry
con...; VIII-XII; page_id 68
Plausible visual variant, with curvature/closure differences; numeric gate failed.
Q08 paired tall stems/crossing bar
1899, whole entry
comes; XII; page_id 59
Only component machinery: Cappelli lacks the paired-stem arrangement.
Q10 paired stems/upper loops
2071, component
allegationi; XV; page_id 14
Several shared features, but ordinary repeated written components in a longer abbreviation.
Q17 repeated strokes/high terminal sweep
3164, component
differentiam; XV p; page_id 108
A good fragment analogy; not demonstrated as one abbreviation sign.
Q18 repeated run/rising terminal trace
7191, whole entry
minus; XIV; page_id 219
General arrangement shared; detached/raised Cappelli ending differs.
Expansions apply only to Cappelli. Workbook expansions describe complete entries, not every cropped component's function. Their presence in an abbreviation dictionary does not make every ordinary loop, stem or minim an abbreviation sign. No Voynich functional equivalence or translation follows.
The narrow-lobe barred form, complex tall structures and longer horizontal/loop combinations lack convincing counterparts here. Tall-loop/bar fragment similarities often fail the complete geometry. Generic repeated strokes are especially easy to match; shared short strokes alone receive little evidential weight.
Chance and retrieval controls
One hundred known Cappelli images subjected to slight shear/blur retrieved their correct normalized shape equivalents at rank one (ties allowed). This validates retrieval within the printed corpus style, not recall across manuscript handwriting and tiny book reproductions.
Each query was compared with 40 matched real neutral-candidate pairs (800 pairs total), matched approximately on aspect, raw cavity count and occupancy; paired companions come from different source surfaces. Controls search the same entire whole/component pools, including all window-selection opportunities. They are unaccepted geometric regions, can include texture or unknown target relatives, and their automatically selected companions can be more coherent than the chosen query pair. They are not a clean ordinary-letter or unrelated-script baseline. Three reflected/rotated controls per fingerprint provide additional sensitivity checks, not a realistic historical null.
No individual query survives the 20-query Benjamini–Hochberg adjustment, even adjusting whole and component strata separately. Minimum adjusted p-values are 0.379 whole and 0.244 component. The finite-control minimum unadjusted p is 1/41; this coarse resolution limits power, so failure to pass is not strong evidence of absence.
Equal weighting of six coarse morphological families gives an exploratory pooled signal: whole mean 0.714 versus control median 0.697, p=0.049; component mean 0.760 versus 0.734, p=0.024. Accounting for the two aggregate outcomes gives Bonferroni values 0.098 and 0.049. The component result is borderline under this conditional null. Forty draws, overlapping constructions and a convenience fingerprint set limit it; morphology blocks do not create historically independent correspondences. This signal cannot establish that matched pieces are abbreviation signs or that several independent units share a historical system.
Period, region and function
Fourteenth/fifteenth-century dates occur for 15/20 component winners, versus 528/762 mapped control winners and 9,189/14,327 corpus images. Whole winners have 11/19, versus 465/746 mapped controls. These are descriptive, overlapping-century memberships, not independent enrichment tests. The date concentration largely follows the source background; component winners are 19/20 Latin, while the mapped corpus is 13,585/14,327 Latin.
The whole-entry winners include isolated Sicilian and Visigothic categories, but these are weak geometric leads from different periods, not a convergent regional result. Component winners have almost no regional/script metadata. The promising displayed forms span different complete-entry expansions and placements; there is no demonstrated common functional system. Region and manuscript tradition are largely absent from this dataset, so absence of a detected convergence is not evidence of historical absence.
Validation and limits
The review is a fresh analyst assessment awaiting human inspection. Small source reproductions, imperfect segmentation, a narrow inventory and approximate retrieval constrain both positive and negative claims. Partial source views can repeat physical ink. No unrelated-script corpus, matched ordinary medieval-writing corpus, source-manuscript confirmation, or validated pen-stroke/unit inventory was available. Similarity alone cannot distinguish shared scribal practice from broad handwriting geometry.
Answer to the primary question: multiple Voynich constructions have close Cappelli pieces, but unusually close complete counterparts and a shared abbreviation system are not established. The pooled component signal is suggestive, conditional evidence; confidence that it proves medieval European scribal machinery beyond chance is low.
Review images
Embedded from the preserved project files (display copies; originals remain in the project folder).