Vintage-LLM/EEBO
EEBO-TCP (Markdown) Early English Books Online, Text Creation Partnership: hand-keyed transcriptions of books printed in England, and English books printed abroad, 1473-1700. Sermons, pamphlets, laws, almanacs, ballads, science, literature. Converted from TCP's XML to Markdown. 60,329 texts, about 1.5 billion words, 25,369 from TCP Phase I and 34,960 from Phase II. TCP keyed each text twice and proofed it to a 99.995% accuracy target. Characters the keyers could not read are… See the full description on the dataset page: https://huggingface.co/datasets/Vintage-LLM/EEBO.
EEBO-TCP (Markdown)
Early English Books Online, Text Creation Partnership: hand-keyed transcriptions of books printed in England, and English books printed abroad, 1473-1700. Sermons, pamphlets, laws, almanacs, ballads, science, literature. Converted from TCP's XML to Markdown. 60,329 texts, about 1.5 billion words, 25,369 from TCP Phase I and 34,960 from Phase II.
TCP keyed each text twice and proofed it to a 99.995% accuracy target. Characters the keyers could not read are marked, never guessed.
Text
The text is lightly normalised from the XML. Every change is decided by an XML tag or by the corpus itself:
- superscript abbreviations tagged in the XML are expanded on whole tokens (y^e the, y^t that, w^t with, y^u you, w^ch which ...), other superscripts are printed inline (Mr, St)
- TCP's own
<expan>is used for<choice>abbreviations - an abbreviation stroke after a vowel becomes n/m only when the rest of the corpus decides: the candidate occurs at least 20 times and 20x more often than the other (frō becomes from, vpō becomes vpon; thē stays, since then and them are both common)
- long s becomes s, vv becomes w
There is no spelling modernisation. u/v and i/j are as printed.
Gaps use TCP's own markers: • per illegible letter, 〈◊〉 for an illegible word, 〈…〉 for longer, 〈 in non-Latin alphabet 〉, 〈♫〉, 〈1 page missing〉 ... Nothing is filled in. About 150 books keep a few private-use characters from TCP's XML (mostly U+E002, the -que abbreviation).
Layout: # headings by division, paragraphs, verse one line per line, **Speaker** in drama, [Illustration: TCP's figure description], notes as Markdown footnotes [^n] at the end. Page breaks, running heads and catchwords are dropped. Each text starts with a short YAML front matter and a title line.
Sources
The Oxford TEI P5 edition (Oxford Text Archive, hdl 20.500.14106) is used when it exists and is complete. Otherwise the text comes from TCP's P4 XML bulk release: for 2 texts OTA does not publish (A55842, B25542), and for 454 where the P5 conversion lost paragraph text around figures (at least 47 consecutive P4 words absent from P5). source, p4_reason, source_file, source_sha256 and ota_handle are on every row.
Completeness
TCP gives 25,368 Phase I and 34,963 Phase II texts, 60,331 in all. OTA's EEBO collection has 60,328. This dataset has all of OTA's except B10185 (a one-page 1690 broadside), plus the 2 P4-only texts above. The Phase I books with an OTA record are exactly OTA's and TCP's 25,368 Phase I texts. B25542 is labelled Phase I but is not in either list.
Checks
Details are in verification.json.
- IDs compared with OTA's full EEBO list (60,328 records, harvested over OAI-PMH) and with TCP's catalogue
TCP.csv. - Phase II: all 34,958 books compared with TCP's P4 XML. 99.86% of the P4 words are in the text. Nearly all the rest are the expansions listed above (
ye,yt,wt,fro...). - Phase I: 60 random books compared with OTA's P5 XML. 99.75% of the words are in the text, and no book is under 95%.
- Every book passes the converter's own two checks:
invariant_ok(the letters and digits equal the source XML, with gaps, footnotes and added layout accounted for) andv1_ok(every changed word is a logged substitution).
Dates
date is the imprint date from the TCP header, as printed (1706 [1705], 169-?] ...). year is the certain bracketed year if there is one, else the first year. 73 books carry an imprint date after 1700, the latest 1865; these are TCP's dates, not errors.
Where the header gives no full year, year was filled from other TCP sources:
- 21 books with an empty
datetake the imprint date from TCP's P4 header (dateandyear). - 258 books take
yearfrom the normalisedDatein TCP's catalogue. Theirdatestays as printed. Catalogue years that contradict the printed date were not used (for example 1780 for169-?]). - 11 books take
yearfrom the printed date once brackets are removed (16[47]is 1647,l643]is 1643).
31 books remain undated: their date gives only a decade (169-?]) or nothing.
Word counts
n_words uses our own tokeniser and sums to 1,503,521,273. Other tokenisers give a slightly different number. See STATS.md for counts by source, period, language and phase.
Official sources
To check the numbers above:
- TCP, EEBO-TCP: <https://www.textpartnership.net/pages/eebo-tcp.html> (Phase I and II counts, licence, downloads)
- Oxford Text Archive, EEBO collection, OAI-PMH set
hdl_20.500.14106_5: <https://llds.ling-phil.ox.ac.uk/llds/oai/request?verb=ListRecords&metadataPrefix=oaidc&set=hdl20.500.14106_5>. Single texts are athttp://hdl.handle.net/20.500.14106/<TCP id>. - TCP catalogue: <https://github.com/textcreationpartnership/Texts> (
TCP.csv) - TCP texts on GitHub, one repository per text: <https://github.com/textcreationpartnership>
Licence and credit
TCP texts are CC0 1.0. Please credit the Text Creation Partnership (University of Michigan, University of Oxford and the libraries that funded it). Page images belong to ProQuest and are not included.
