Undigitised European print · supply for machine-learning programmes
European long-tail print, measured two ways.
Every lot is graded title by title against the German National Library, HathiTrust, Internet Archive and Open Library. We report two numbers: how much of it has never been digitised — and how much of it no one but Google can access.
Most of what Europe printed in the last century was never digitised, and bulk supply chains deliver the opposite of what training pipelines need from print: high-duplication English mass stock — the same titles, many times over.
Of what was scanned, a large share sits inside Google's library programme: searchable at most, readable by no one, usable by no buyer except Google. For every other lab, a closed scan is no scan — the physical volume is the only route to the text.
The long tail of German, Czech, Polish and Hungarian technical and professional literature surfaces in places volume buyers never see. Reaching it takes language, catalogue knowledge and sourcing infrastructure — not tonnage.
| Composition | Curated pallet lots. Title diversity by design: one copy per edition, not multiples. |
|---|---|
| Coverage | Technical, academic, professional and regional literature. German-language core, extending across Central and Eastern Europe. |
| Grading | Title by title against DNB, HathiTrust, Internet Archive and Open Library. Four grades per title, check date on every sheet, under a published, versioned definition. |
| Two KPIs | Never digitised (grades A+B) · not openly available (grades A+B+C) — sampled per lot, sample size stated. The second number is the one that counts for a training-data buyer. |
| Sourcing | Proprietary channels that reach collections before they enter normal trade. We describe the capability, not the sources. |
| Logistics | Consolidation within the EU; delivery EU and US. |
| Confidentiality | NDA-comfortable. The full title-level grading sheet accompanies a sold lot; samples, volumes and terms are discussed directly, not published here. |
| Ref | Title / imprint | Grade | Finding |
|---|---|---|---|
| K-2417 | Vodní hospodářství v povodí Moravy Brno, 1961 |
A | no scan found, 4/4 sources · scarce in the trade |
| K-2418 | Berechnung stählerner Fachwerkbrücken Leipzig, 1953 · 2. Aufl. |
B | no scan found, 4/4 sources |
| K-2419 | Lehrbuch der Gerbereichemie Dresden, 1949 |
C | HathiTrust search-only — scan exists, access closed |
| K-2420 | Einführung in die höhere Mathematik, Bd. 2 Berlin, 1967 · 9. Aufl. |
D | open full text (Internet Archive) |
| K-2421 | Zeitschrift für Binnenschiffahrt, Jg. 1938 bound periodical volume |
unknown | fuzzy match — graded conservatively, not counted toward A/B |
“Sample of 60 titles, graded 2026-08-24: 67% never digitised · 95% not openly available (B40 C17 D3). Full title-level grading sheet on request.”
Grades are assigned to a stated sample drawn from shelf and spine photographs — figures are sampled, not guaranteed for the whole lot. All check sources are public catalogues, so any row of the sheet can be re-checked. The grades describe digitisation and accessibility status only; they say nothing about usage or licence rights. Definition and limits: methodology.
BookKorpus is a specialist operation led by a founder with a quantitative data-engineering background and native German sourcing capability. The sourcing, cataloguing and grading pipeline is proprietary and built end to end — built as infrastructure, designed to scale or to integrate into a larger data-acquisition operation.
We keep the operation small and the claims checkable.
For lot availability, grading sheets and terms: