ARCHIVED RESEARCH PROJECT
First f1r calibration and PDF source audit
Does the supplied PDF provide more manuscript detail, and what does a controlled f1r rerun show?
FINDINGS AND REVIEW IMAGES489 accessible result rows from 489 authoritative meaningful result rows in the declared scope. Review cards are separate.
All rows from the stated saved result sets are accessible; detailed machine arrays remain in the linked source files.
The preserved report and review images below contain the surviving human-review material; no separate historical item count was recorded.
This curated view is separate from the complete meaningful results listed above.
Archived research project · completed 2026-09-24
ARCHIVED RESEARCH PROJECT
Does the supplied PDF provide more manuscript detail, and what does a controlled f1r rerun show?
FINDINGS AND REVIEW IMAGESFINDINGS FROM THE COMPLETED STUDY · public presentation of [preserved evidence reference]
The supplied PDF provides no additional manuscript image detail. f1r_v003 is a fresh, separately registered extraction and controlled rerun of the same source pixels. It does not eliminate the resolution limitation. Human review remains pending.
Open [preserved evidence reference] and select v002 ↔ v003. Every registered observation has paired source crops, surrounding context, and the unchanged detail browser with candidate relatives, components, and alternate segmentations.
PDF page 4 was visually verified as f1r before the rerun: its printed caption is 1r, and the photograph shows the four text blocks, decorated left marks, and detached right groups. The direct page rendering and retain the evidence.
Page 4 paints exactly one manuscript raster, PDF image object 442: an 1184 × 1500, 8-bit-per-channel JPEG. The content stream contains that image, two rectangular frames, and library/caption text. No higher-resolution alternate image, manuscript vector strokes, image mask, or annotations supply additional manuscript detail.
At its placement in the PDF, the photograph's effective density is 155.51015 × 155.51016 pixels per inch. This is computed from its native dimensions and PDF placement transform; it is not a claim about the original camera/scanner's physical resolution. The image's nominal 96-dpi metadata does not determine its effective on-page density.
The highest-fidelity available source is the unresampled extracted JPEG and its lossless decoded RGB PNG. A 156-dpi page rendering provides approximately native-density page context. The unchanged pipeline also produces a 400-dpi PDF rendering and 4× review enlargement; those enlarge the same photograph and add no manuscript information. Vector library labels can become sharper at higher DPI, but the manuscript writing cannot.
The authoritative project voynich.pdf, its preserved copy, and v002's source all have SHA-256:
d095900be761dd8702a63cf79f83a09e832f98e4abd0d4ea394d6c47e8d35a7c
The newly extracted JPEG and v002's JPEG both have SHA-256:
2355010b9442b2c1bbf72a9e5107fb27f8ed93a5552de48e59c405531d5f88b8
All 1,776,000 decoded RGB pixels match exactly: zero changed channel samples and maximum absolute channel difference zero. The new acquisition is registered as voynich_pdf_page4_acquisition_v003, with a . New acquisition provenance does not imply different image content.
The v002 extraction, geometry, classification, uncertainty, rendering, and validation source files match their recorded v002 hashes. The frozen v002 configuration was passed directly to the pipeline. No extraction source file or classification rule was edited. The comparison extension decorates the original review renderer after it finishes; it does not change observations or decisions.
Native pixel scale is 1:1, so zero parameters changed. This includes the block/R boxes, terminal main-text right-edge annotations, all geometric thresholds, image preprocessing, uncertainty penalties, relative links, and ID policy. is recorded.
| Example scale-dependent parameter | v002 | v003 | Rationale |
|---|---|---|---|
| Background Gaussian sigma | 7 px | 7 px | Same native density |
| Line peak minimum distance | 24 px | 24 px | Same native density |
| Minimum proposal component area | 4 px² | 4 px² | Area scale factor 1² |
| Maximum component/line distance | 23 px | 23 px | Same native density |
| Whole-form bridge gap | 1 px | 1 px | Same native density |
| Visual cluster gap | 8 px | 8 px | Same native density |
| Threshold levels | 6, 8, 10 | 6, 8, 10 | Identical decoded intensity values |
The new dataset is regenerated from the preserved PDF, whose bytes are verified identical to the authoritative project PDF. It is not produced by resizing or copying the old PNG into the extractor.
| Measure | f1r_v002 | f1r_v003 | Change |
|---|---|---|---|
| Source dimensions | 1184 × 1500 | 1184 × 1500 | None |
| Lines | 24 | 24 | 0 |
| Visual clusters | 164 | 164 | 0 |
| Provisional visual forms | 488 | 488 | 0 |
| Connected components, including residual marks/texture | 40,597 | 40,597 | 0 |
| Components associated with form proposals | 1,747 | 1,747 | 0 |
| Alternative segmentation hypotheses | 271 | 271 | 0 |
| Alternative partition crops | 1,054 | 1,054 | 0 |
| Low-confidence segmentations | 196 | 196 | 0 |
| Threshold-sensitive observations | 41 | 41 | 0 |
| Possible compounds | 264 | 264 | 0 |
| Neighbor-dependent observations | 344 | 344 | 0 |
| Uncertain line assignments | 168 | 168 | 0 |
The connected-component total includes texture, pigment, damage, showthrough, page edges, and possible ink; it is not a glyph count. .
Registration is the exact identity transform: scale (1,1), translation (0,0), rotation 0°, and residual error 0 px. This is established by identical source bytes and decoded pixels, rather than by estimating feature matches. All block boxes, line routes/slopes, R objects, glyph boxes, cluster memberships, component masks, and right-edge annotations map exactly.
All 488 observation crops and 40,597 component masks have explicit old/new correspondence. IDs remain version-qualified. No accepted shared identity is inferred from their matching numbers. , , and are saved.
| Comparison outcome | Count |
|---|---|
| A — additional source evidence confirms the provisional decision as correct | 0 |
| B — additional detail revealed | 0 |
| C — new evidence suggests an old incorrect merge | 0 |
| D — new evidence suggests an old incorrect split | 0 |
| E — remains ambiguous | 488 |
| F — cannot be reliably registered | 0 |
| Separately: exact pixel/segmentation reproductions | 488 |
| Newly distinguishable forms | 0 |
| Ambiguities resolved | 0 |
“Confirmed reproduction” and “confirmed correct segmentation” are kept separate. Identical pixels confirm that the old observations were reproduced, but do not supply new evidence validating their boundaries or identity. Existing alternative split and compound hypotheses remain unchanged; they are not counted as new findings.
The original annotated layout boxes are retained. The union of newly routed form-proposal boxes and the spacing to main text were recomputed independently in the rerun and match v002. Coordinates are native top-left-origin [x0,y0,x1,y1], right/bottom exclusive.
| Object | Annotated box, both versions | Recomputed proposal-union box, both versions | Gap to annotated box, v002 → v003 |
|---|---|---|---|
| R1 | [809,334,948,373] | [817,332,942,368] | 134 → 134 px |
| R2 | [733,478,925,518] | [735,480,913,525] | 206 → 206 px |
| R3 | [775,891,946,935] | [780,893,945,939] | 205 → 205 px |
| R4 | [790,1188,897,1218] | [788,1186,897,1217] | 182 → 182 px |
Some routed components extend beyond an annotated R box. This existing ownership/extent uncertainty is preserved. Neither box definition is promoted to a verified glyph boundary. The corresponding gaps to proposal-union boxes are also saved in . No meaning is assigned.
No new distinction was discovered. Pixel comparisons cover the entire source and every observation/context crop, including all existing tall, horizontal, repeated-stroke, and compound candidates. The identical source provides no additional evidence about double stems, left/right loop differences, shared crossbars, bench structures, minim counts, hook-like forms, or whether apparent connections are real. These categories remain for visual review; no new semantic component annotations were invented.
The original faint-stroke, stain/texture, touching-form, fragmentation, cross-line, variant-identity, and component-boundary ambiguities remain. The 196 low-score segmentations and all unassessed classification scores remain unchanged. No automatic merge, split acceptance, or new variant identity was introduced.
All five existing targeted invariant tests and all twenty provenance/structure checks passed again. Checks cover exact source-pixel crops, coordinates, all primary-mask pixels, component membership, spacing, R isolation, contact-sheet completeness, and frozen artifact hashes. .
A separate same-runtime rebuild reproduced 5,971 pipeline/review artifacts byte for byte, plus all five external comparison artifacts. Full manifest and review payloads also matched after excluding run timestamps and their derived manifest hash. .
The review keeps Page & layout, Glyph inventory, Uncertainty queue, Contact sheets, and Review log. It adds a comparison tab and side-by-side crops inside observation details. The HTML, comparison table, ID search, and paired glyph detail were inspected in Safari. The minimal-DOM test additionally checks review-event append/export/import behavior and rejects v002 review logs as belonging to a different dataset. . No human review decisions were migrated or created.
The protected prior-version inventory contains 15,381 files, including previous rendered images, crops, review files, datasets, manifests, and authoritative PDF copies. Their before/after hashes are checked by the . Each new dataset artifact is bound to its source version through the .
These checks verify provenance and reproducibility; segmentation and classification accuracy remain unvalidated. The unchanged extraction code is saved in the source audit's v002_code_snapshot and remains in Git commit 239634e. The controlled audit/rerun wrapper is [deterministic method]; it refuses to replace existing outputs.
f1r_v003 is generated, compared, and ready for human review. The supplied PDF cannot provide the desired source-quality upgrade because it already supplied v002's native pixels. A genuinely higher-resolution manuscript image would be needed to remove this limitation.
No full-manuscript processing, EVA comparison, Cappelli matching, translation, decipherment, language comparison, or recurring-glyph-class merge was run.
Embedded from the preserved project files (display copies; originals remain in the project folder).








