ARCHIVED RESEARCH PROJECT · completed 2026-09-29 · Without assuming a language, which known language types does the structural behaviour of Voynich writing resemble, under blind, frozen, controlled scoring? · Back to Home
PROJECT SUMMARY
QuestionWithout assuming a language, which known language types does the structural behaviour of Voynich writing resemble, under blind, frozen, controlled scoring?
MethodVoynich side ([deterministic method]). Everything was derived from the saved union ink masks of the full_v001 neutral extraction: - connected ink pieces became symbols; - pieces sharing a baseline became provisional lines; - pieces were grouped into k-means shape…
Important numbers18 scored variant · 70 control · 16 language baseline
Finding and conclusionRobust metrics (5). These are the edit-1 neighbour excess, start-change and internal-change excess, family excess, and ending-concentration excess. - On all five, Voynich sits within ±0.02 of zero in every variant. The pixel-derived clusters carry no word-internal structure detectable…
Limits- The symbolization is automatic and noisy: k-means shape classes mix glyphs, and faint strokes are missed. - Cluster boundaries are provisional. - Modern treebanks stand in where no historical one exists (Hungarian, German, Czech, Persian). - Ottoman Turkish has a smaller source (about…
MeaningAdjacent repetition (post-hoc, posthoc_voynich.json). Voynich's excess of identical and near-identical neighbouring clusters is present in 18/18 variants and above every language. It is fully reproduced by shuffling clusters within their own line, and largely by shuffling within the page. It is therefore line-level homogeneity, not sequential repetition. Whether it comes from the script or from local ink/scale drift in the shape classes is unresolved.
104 accessible result rows from 104 authoritative meaningful result rows in the declared scope. Review cards are separate.
All rows from the stated saved result sets are accessible; detailed machine arrays remain in the linked source files.
REVIEW NEEDED
8 preserved review panels are available below.
This curated view is separate from the complete meaningful results listed above.
Blind structural comparison: Voynich vs 16 languages (v001)
Question: without assuming a language, which known language types does the structure of the Voynich writing resemble? Nothing was translated, no glyph was given a sound, and no Turkish-specific or EVA-based claim was read or used.
Answer
No language or language type stands out once controls are applied.
Primary (frozen) scoring. On the five metrics that stayed stable across all 18 Voynich segmentations, the nearest corpus was Classical Chinese in 18 of 18 segmentations, confirmed on held-out halves.
Classical Chinese was included as the isolating control; here it acts as the corpus with no internal word structure.
Every comparison language, once degraded with 30–50% extraction-style noise, also moves to Classical Chinese: 15 of 16 at 30%, 16 of 16 at 50%.
So did Markov-chain and self-copying pseudo-text generated from Voynich itself.
The Classical Chinese match therefore means only that our pixel-derived Voynich symbol strings show no detectable internal (morphological) structure. Degraded real languages show the same.
Turkic did not perform well on the robust scoring. Ottoman Turkish ranked 8th of 16. It ranked 1st on a secondary distance that also includes segmentation-sensitive repetition statistics. That match is driven only by those statistics, not by any word-internal measure, and Voynich-generated pseudo-text lands on Ottoman Turkish on the same distance. It does not survive controls.
Agglutinative, fusional, isolating, templatic or compressed? UNRESOLVED. With the current pixel extraction, no suffix-, prefix- or internal-change pattern is measurable above chance.
1. What was measured on the Voynich side
From the saved ink masks of the neutral extraction:
connected ink pieces became symbols, grouped into 16, 32 or 64 neutral shape classes;
pieces sharing a baseline became provisional lines;
gaps wider than 0.8, 1.1 or 1.5 x-heights split lines into candidate clusters;
broken strokes were either kept separate or merged.
That gives 3 × 3 × 2 = 18 segmentation variants. Coverage: 204 pages, 45,495 provisional lines, and 14,478 lines with at least 8 pieces used. Clusters are not words and symbols are not letters.
f103r: each coloured box is one candidate cluster (K32, gap 1.1 xh). Many real spaces are found, but faint strokes are missed and some words split or merge. This extraction noise matters for everything below.f001r, the same settings, in a herbal section with plant ink. Coverage is lower.The 16 most frequent of 32 neutral shape classes (examples from f103r and f001r). A class often mixes one glyph with a joined pair, or splits one glyph by ink breaks. Several frequent classes are blank parchment texture, not ink (see 4b). Classes are not letters.
2. How the languages were blinded and made comparable
There were 16 corpora of 16,000 words each from Universal Dependencies treebanks, historical where available (Ottoman Turkish, medieval Latin, Old Italian, Old French, Gothic, Old East Slavic, Ancient Greek, Ancient Hebrew, Classical Chinese). Where no historical treebank exists, modern ones were used (Hungarian, Finnish, German, Czech, Persian, Arabic, Basque). Every corpus got the same treatment:
several documents or authors;
accents and vowel points stripped;
characters replaced by random numbers;
a random code (C01–C16) instead of a name.
Metrics, scoring, controls and the Voynich measurements were hashed and frozen before the code key was opened. The blind conclusions were also sealed before the reveal.
Structure, not alphabet. Every word-internal measure is scored as its excess over the same corpus with its symbols shuffled. Every between-word measure is scored as its excess over the same corpus with its word order shuffled. This removes alphabet size and word length, which a first test showed would otherwise dominate.
3. Revealed comparison table
Language
Type (a priori)
Robust rank (median)
Robust #1
All-metric rank
All-metric #1
Last-symbol change excess
First-symbol change excess
Internal change excess
Adjacent repeat excess
Classical Chinese (isolating control)
isolating
1
18/18
4.5
4/18
-0.000
+0.000
+0.000
-0.0011
Ancient Hebrew (root-pattern control)
templatic (root-and-pattern), consonantal script
2
0/18
6
1/18
-0.107
-0.002
+0.109
-0.0031
Old French
fusional (analytic tendency)
3
0/18
13
0/18
-0.056
-0.018
+0.074
-0.0061
Old Italian
fusional (analytic tendency)
4.5
0/18
14
0/18
+0.001
-0.078
+0.077
-0.0073
Persian (modern, Arabic script; no historical UD)
analytic/fusional mixed, some agglutination
4.5
0/18
6
0/18
-0.092
-0.075
+0.167
-0.0082
Arabic (root-pattern control; modern standard)
templatic (root-and-pattern), consonantal script
6
0/18
3
0/18
-0.110
-0.032
+0.142
-0.0012
Old East Slavic (historical Slavic)
fusional
7
0/18
7
0/18
+0.111
-0.231
+0.121
-0.0076
Ottoman Turkish (Turkic)
agglutinative, suffixing
8
0/18
1
13/18
-0.097
-0.172
+0.269
+0.0007
Ancient Greek
fusional
9
0/18
8
0/18
+0.017
-0.175
+0.159
-0.0059
Latin (medieval: Aquinas, charters, Dante)
fusional
10
0/18
10.5
0/18
+0.103
-0.221
+0.118
-0.0050
German (literary; no historical UD)
fusional
11
0/18
11
0/18
+0.226
-0.171
-0.056
-0.0058
Gothic (historical Germanic)
fusional
12
0/18
11
0/18
+0.150
-0.192
+0.043
-0.0091
Basque (agglutinative isolate control)
agglutinative, suffixing
13
0/18
3.5
0/18
+0.076
-0.216
+0.140
-0.0021
Czech (modern; no historical UD)
fusional
14
0/18
6
0/18
+0.085
-0.254
+0.169
-0.0031
Hungarian (Uralic; modern - no historical UD)
agglutinative, suffixing
15
0/18
15.5
0/18
+0.034
-0.166
+0.132
-0.0127
Finnish (Uralic; modern control)
agglutinative, suffixing
16
0/18
15
0/18
-0.104
-0.251
+0.355
-0.0022
Voynich (18 variants)
?
-0.019 to +0.037
-0.043 to -0.004
-0.002 to +0.027
+0.001 to +0.014
The “change excess” columns form the agglutination test. Among pairs of similar forms that differ by one symbol, they show how much more often the difference sits at the end, the start or inside than it would by chance. Suffixing agglutinative languages were expected to show end excess and root-pattern languages internal excess. Voynich sits near zero on all three in every variant.
4. Which similarities disappear under controls
Classical Chinese (robust scoring): reproduced by noise-degraded versions of other languages (31 of 32 degraded corpora nearest Classical Chinese (isolating control); 1 of 32 degraded corpora nearest Ancient Hebrew (root-pattern control)) and by Voynich-trained pseudo-text. It disappears as language evidence.
Ottoman Turkish (secondary all-metric scoring, #1 in 13 of 18 variants). Its advantage comes from:
the adjacent-repeat excess (Ottoman is the only language slightly above zero);
low adjacent-word mutual information;
high type/token ratio and hapax rate.
Ottoman Turkish is not closer than the median language on the word-internal measures (neighbour rate, internal change, rank-frequency slope). Voynich-generated self-copying and Markov pseudo-text are also nearest to Ottoman Turkish on this distance. It disappears as language evidence.
High hapax / type-token rate (Voynich 0.79–0.95 hapax vs languages 0.43–0.77) is reached by languages degraded at 30–50% noise, so it is explained by extraction noise.
The class gallery above shows that many frequent “symbols” were blank parchment texture. The extraction was repeated keeping only ink-dark pieces (20th-percentile brightness below 0.8 of the local background; ink and texture separate cleanly). The unchanged frozen scoring was rerun. This check is not blind, because the labels were already known.
After filtering, most classes are recognisable recurring shapes (9-like, c-like, o-like, gallows, ee-runs), with some texture left.
Frozen robust metrics: still Classical Chinese (isolating control) first in 18 of 18 variants.
Robustness rule re-applied to the filtered data: Ancient Hebrew (root-pattern control) first in 18 of 18, holdout-stable. However, a Markov pseudo-text built from the same Voynich clusters is also nearest to it, and noise-degraded languages go to Classical Chinese or Ancient Hebrew. It fails controls.
All-metric distance: Ottoman Turkish (Turkic) first in 17 of 18. Voynich-generated Markov and self-copying pseudo-text are also nearest to it. The drivers are again the repetition statistics. It fails controls.
Word-internal predictability rose into the normal language range, but there is still no edit-neighbour, suffix or prefix alternation above a shuffle.
5. Robust Voynich properties, and what no language explains
Robust across all 18 segmentations. The candidate clusters have no word-internal structure detectable above a symbol shuffle: no end, start or internal alternation pattern, no family effect and no concentration of endings. This is unresolved as a property of the manuscript, because degraded real languages look identical.
Unexplained by every language. Neighbouring clusters are identical or near-identical more often than a corpus-wide shuffle predicts (K32/1.1/merged: near-identical neighbours 0.202 vs 0.175 shuffled). A post-hoc check shows this is line-level sameness, not word-to-word repetition: shuffling clusters only within their own line gives 0.207. Whether this reflects the writing, such as the known line-level patterns, or local ink and scale drift in the shape classes is unresolved.
Excluded as not robust. Cluster length, type/token ratio, rank-frequency slope and adjacency statistics all changed by more than half the between-language spread across segmentations.
6. What would make this test informative
The comparison is limited by symbol quality, not by the languages. The next test is the same frozen protocol on a human-verified glyph sequence for a subset of text-dense pages, for example the reviewed f1r plus a recipe page, with the same controls. If structure then appears above the shuffle, the language-type comparison becomes meaningful.
Nothing was translated, and no glyph was given a sound.
The Turkish hypothesis, proposed translations, Turkish-specific Voynich claims, EVA interpretations and earlier conversation about Turkish were not read or used.
Turkic was one comparison entry among 16, treated identically to the others.
Design and blinding
Voynich side ([deterministic method]). Everything was derived from the saved union ink masks of the full_v001 neutral extraction:
connected ink pieces became symbols;
pieces sharing a baseline became provisional lines;
pieces were grouped into k-means shape classes, K = 16, 32 or 64;
within-line gaps above 0.8, 1.1 or 1.5 x-heights split lines into candidate clusters;
broken strokes were either kept raw or merged.
That gives 18 segmentation variants. They are drawn from 204 pages and cover 14,478 provisional lines of at least 8 pieces.
A first attempt used the form-scale visual index instead. It was rejected after visual quality checks because dense text was missing from it.
Clusters are not treated as words, and symbols are not treated as letters.
Comparison corpora ([deterministic method]). 16 entries of 16,000 words each were taken from Universal Dependencies treebanks, one entry per language:
Language
Corpus
Period
Ottoman Turkish
BOUN + DUDU
historical
Latin
ITTB (Aquinas), LLCT (charters), UDante
medieval
Old Italian
Italian-Old
historical
Old French
PROFITEROLE
historical
Gothic
PROIEL
historical
Old East Slavic
TOROT
historical
Ancient Greek
PROIEL
historical
Ancient Hebrew
PTNK
historical
Classical Chinese
Kyoto
historical
Hungarian
Szeged
modern (no historical treebank)
Finnish
TDT
modern
German
LIT (literary)
modern (no historical treebank)
Czech
CAC
modern (no historical treebank)
Persian
Seraji
modern (no historical treebank)
Arabic
PADT
modern
Basque
BDT
modern
Excluded: UD Old_Turkish-Clausal, which is too small.
Every corpus received the same processing:
blocks of sentences spread across all its files (several documents or authors);
punctuation dropped and text lower-cased;
all diacritics stripped (Unicode NFD, combining marks removed);
characters, lemmas and feature bundles mapped to per-corpus random integers;
files given random codes C01–C16, with the key sealed (sha256 1d2de05b…).
Separate morphological fields were kept but not used in scoring.
Metrics and controls ([deterministic method]). There are 20 metrics:
type and repetition: type/token ratio, hapax share, rank-frequency slope;
symbol entropy and within-token predictability;
edit-1 near-neighbour rate, and where the one-symbol difference sits (end, start, internal);
family size, ending vs beginning concentration, successor variety, and a concatenation (agglutination) index;
adjacent repeats, adjacent near-repeats and adjacent mutual information;
line-initial and line-final effects.
Structure, not alphabet.
Within-token metrics are scored as their excess over the same corpus with its symbols shuffled.
Between-token metrics are scored as their excess over the same corpus with its token order shuffled.
A smoke test showed that alphabet size and token length otherwise dominate.
Controls.
Each language was also degraded with 10%, 30% and 50% symbol noise plus boundary merge/split errors.
Order-2 Markov pseudo-text was generated for every language and for Voynich.
Voynich self-copying pseudo-text was generated (copy a recent token with one edit).
Voynich lines were re-cut into pseudo-lines.
Every corpus was split into A/B holdout halves.
Scoring.
Metrics were z-scored by the language distribution.
The primary distance uses only metrics that are robust across the 18 Voynich variants: spread under 0.5 of the between-language SD.
A secondary distance uses all scored metrics.
Freeze. Tools, Voynich variant files, corpora, scoring and an a-priori typology table were hashed in freeze.json (sha256 4b5769d6…) before the key was read. The blind conclusions were written and sealed (blind_conclusions_SEALED.md, sha256 f4f3da89…) before the reveal.
Results (frozen, then revealed)
Robust metrics (5). These are the edit-1 neighbour excess, start-change and internal-change excess, family excess, and ending-concentration excess.
On all five, Voynich sits within ±0.02 of zero in every variant. The pixel-derived clusters carry no word-internal structure detectable above a symbol shuffle.
Every language sits clearly below zero on neighbour excess (−0.14 to −0.64), except Classical Chinese (≈ 0).
Primary result: Classical Chinese is nearest in 18/18 variants. A/B holdout agrees in 18/18.
This fails controls. Noise-degraded languages also go to Classical Chinese: 15 of 16 at 30% noise and 16 of 16 at 50%. So do Voynich-trained Markov and self-copying pseudo-text.
Classical Chinese (single-character tokens) is simply the structureless end of the scale.
Secondary all-metric distance: Ottoman Turkish is first in 13/18 variants (Arabic top-3 in 16/18, Basque in 9/18). This includes non-robust metrics.
Drivers:
adjacent-repeat excess (Ottoman is the only language slightly above zero, +0.0007, closest to Voynich's positive values);
low adjacent-token mutual information;
high type/token ratio and hapax rate.
On the word-internal measures Ottoman Turkish is not closer than the median language: neighbour rate, internal change and rank-frequency slope all contribute more distance than for other languages.
This fails controls. Voynich-generated self-copying and Markov pseudo-text are also nearest Ottoman Turkish on this distance.
Robust-rank table (median rank, 1 = nearest).
1–4: Classical Chinese 1, Ancient Hebrew 2, Old French 3, Old Italian 4.5, Persian 4.5.
6–8: Arabic 6, Old East Slavic 7, Ottoman Turkish 8.
9–12: Ancient Greek 9, Latin 10, German 11, Gothic 12.
Texture problem. The symbol-class gallery showed that many frequent classes were blank parchment texture.
Rerun. A texture filter was decided from ink evidence only: keep pieces whose 20th-percentile brightness is below 0.8 of the local background. On f103r ink and texture separate bimodally. The filtered extraction gave 6,576 lines, close to the manuscript's roughly 5,200. It was rescored with the unchanged frozen code (tools/structural_language_v001b_*, sensitivity_inkfilt_v001b.json).
Results.
Frozen robust metrics: Classical Chinese is still first in 18/18.
Robustness rule re-applied: only symbol entropy, neighbour excess and ending concentration pass. Ancient Hebrew is first in 18/18, holdout 18/18. But Voynich Markov pseudo-text is also nearest Ancient Hebrew, and degraded languages go to Chinese or Hebrew. Fails controls.
All-metric distance: Ottoman Turkish is first in 17/18. Voynich Markov and self-copying pseudo-text are also nearest Ottoman Turkish. Fails controls.
Within-token predictability improved into the language range (−0.06 to −0.13). Neighbour, suffix and prefix alternation excess stayed ≈ 0.
Adjacent repetition (post-hoc, posthoc_voynich.json). Voynich's excess of identical and near-identical neighbouring clusters is present in 18/18 variants and above every language. It is fully reproduced by shuffling clusters within their own line, and largely by shuffling within the page. It is therefore line-level homogeneity, not sequential repetition. Whether it comes from the script or from local ink/scale drift in the shape classes is unresolved.
Answers
Robust Voynich properties (on this extraction):
no measurable word-internal alternation or family structure above a shuffle;
neighbouring clusters within a line resemble each other more than clusters elsewhere;
very high hapax/type richness (extraction noise alone reproduces it).
Languages or types resembling those properties: only the structureless end of the scale (Classical Chinese), plus, on repetition statistics, Ottoman Turkish, Arabic and Basque.
Similarities that disappear under controls: all of them. The Chinese, Hebrew and Ottoman Turkish matches are each reproduced by noise-degraded languages or by Voynich-generated pseudo-text.
Any language standing out on multiple independent metrics: no. Nearest-by-metric winners are scattered: Finnish on hapax, Czech on symbol entropy, Old French on rank-frequency slope, Chinese on neighbour metrics, Ottoman Turkish on adjacent repeats, Hungarian on adjacency mutual information.
Features no language explains: line-level homogeneity of neighbouring clusters, and extreme hapax richness. Both are possibly extraction effects.
Agglutinative / fusional / isolating / compressed?Genuinely unresolved. No suffix-, prefix- or internal-change excess is measurable. The test is limited by the symbol extraction, not by the languages.
Turkic, stated plainly: Turkic did not perform well on the robust, primary scoring (8th of 16). It ranked first only on the secondary all-metric distance. Those properties are adjacent-repeat excess, low adjacent mutual information and high type/token and hapax rates, not any word-internal structure. Voynich-generated pseudo-text lands on it too. This is not evidence for a Turkic language.
No language identification is claimed or supported.
Next test
Run the same frozen protocol on a human-verified glyph sequence for a subset of text-dense pages. If within-token structure then exceeds the shuffle, the language-type comparison becomes informative.
Limits
The symbolization is automatic and noisy: k-means shape classes mix glyphs, and faint strokes are missed.
Cluster boundaries are provisional.
Modern treebanks stand in where no historical one exists (Hungarian, German, Czech, Persian).
Ottoman Turkish has a smaller source (about 31,000 words available).
Language lines are artificial pseudo-lines, so line metrics are descriptive only.