One hundred fifty-one books.
Every page count checked against disk.
The shelf the methods were mined from, published whole. Each book carries its page count as verified against the files on disk, its OCR quality as measured rather than assumed, and its reprints charged against the arithmetic — because a page that appears twice must not be counted twice when you claim how much was read.
The distinct-page arithmetic
The number that matters is not how many page files exist but how many distinct pages they represent. Reprints were measured, not estimated, and subtracted.
Coverage by wave and read layer
Every book
OCR quality is a measured statistic, not a guess: the emphasis deficit compares the rate at which emphasised (italic/bold/heavy) words survive the text layer against plain words. A large negative deficit means the scan silently dropped the emphasised text — exactly the words a method statement leans on.
| author | title | wave | pages | text layer | vision | unread | emphasis deficit | OCR verdict | disk |
|---|---|---|---|---|---|---|---|---|---|
| loading 151 books… | |||||||||
Provenance notes
The corpus as it was ingested
A second count exists, taken at ingest time before the page files were verified against disk. It is published rather than reconciled away: the two disagree, and the disagreement is itself a fact about the corpus. The ingest map counts what the eight source folders held; the verified count above counts what survived fingerprinting, and charges reprints against the distinct-page floor. Where a figure on this site needs a book count, it uses the verified one.
| author | book | pages at ingest | on the verified shelf? |
|---|