BookKorpus

Undigitised European print · supply for machine-learning programmes

European long-tail print, measured two ways.

Every lot is graded title by title against the German National Library, HathiTrust, Internet Archive and Open Library. We report two numbers: how much of it has never been digitised — and how much of it no one but Google can access.

The gap

Most of what Europe printed in the last century was never digitised, and bulk supply chains deliver the opposite of what training pipelines need from print: high-duplication English mass stock — the same titles, many times over.

Of what was scanned, a large share sits inside Google's library programme: searchable at most, readable by no one, usable by no buyer except Google. For every other lab, a closed scan is no scan — the physical volume is the only route to the text.

The long tail of German, Czech, Polish and Hungarian technical and professional literature surfaces in places volume buyers never see. Reaching it takes language, catalogue knowledge and sourcing infrastructure — not tonnage.

What we deliver
CompositionCurated pallet lots. Title diversity by design: one copy per edition, not multiples.
CoverageTechnical, academic, professional and regional literature. German-language core, extending across Central and Eastern Europe.
GradingTitle by title against DNB, HathiTrust, Internet Archive and Open Library. Four grades per title, check date on every sheet, under a published, versioned definition.
Two KPIsNever digitised (grades A+B) · not openly available (grades A+B+C) — sampled per lot, sample size stated. The second number is the one that counts for a training-data buyer.
SourcingProprietary channels that reach collections before they enter normal trade. We describe the capability, not the sources.
LogisticsConsolidation within the EU; delivery EU and US.
ConfidentialityNDA-comfortable. The full title-level grading sheet accompanies a sold lot; samples, volumes and terms are discussed directly, not published here.
Grading, in brief
Structure of a grading sheet — illustrative rows
RefTitle / imprintGradeFinding
K-2417 Vodní hospodářství v povodí Moravy
Brno, 1961
A no scan found, 4/4 sources · scarce in the trade
K-2418 Berechnung stählerner Fachwerkbrücken
Leipzig, 1953 · 2. Aufl.
B no scan found, 4/4 sources
K-2419 Lehrbuch der Gerbereichemie
Dresden, 1949
C HathiTrust search-only — scan exists, access closed
K-2420 Einführung in die höhere Mathematik, Bd. 2
Berlin, 1967 · 9. Aufl.
D open full text (Internet Archive)
K-2421 Zeitschrift für Binnenschiffahrt, Jg. 1938
bound periodical volume
unknown fuzzy match — graded conservatively, not counted toward A/B
Example datasheet line (illustrative figures):
“Sample of 60 titles, graded 2026-08-24: 67% never digitised · 95% not openly available (B40 C17 D3). Full title-level grading sheet on request.”

Grades are assigned to a stated sample drawn from shelf and spine photographs — figures are sampled, not guaranteed for the whole lot. All check sources are public catalogues, so any row of the sheet can be re-checked. The grades describe digitisation and accessibility status only; they say nothing about usage or licence rights. Definition and limits: methodology.

About

BookKorpus is a specialist operation led by a founder with a quantitative data-engineering background and native German sourcing capability. The sourcing, cataloguing and grading pipeline is proprietary and built end to end — built as infrastructure, designed to scale or to integrate into a larger data-acquisition operation.

We keep the operation small and the claims checkable.

Contact

For lot availability, grading sheets and terms:

contact@bookkorpus.com We reply within one business day.