Team Ai
Datasetpublic

Vintage-LLM/EEBO

EEBO-TCP (Markdown) Early English Books Online, Text Creation Partnership: hand-keyed transcriptions of books printed in England, and English books printed abroad, 1473-1700. Sermons, pamphlets, laws, almanacs, ballads, science, literature. Converted from TCP's XML to Markdown. 60,329 texts, about 1.5 billion words, 25,369 from TCP Phase I and 34,960 from Phase II. TCP keyed each text twice and proofed it to a 99.995% accuracy target. Characters the keyers could not read are… See the full description on the dataset page: https://huggingface.co/datasets/Vintage-LLM/EEBO.

sourceHugging Facecc0-1.0updated 12d agoView on Hugging Face
0likes1.5kdownloads
Dataset Card

EEBO-TCP (Markdown)

Early English Books Online, Text Creation Partnership: hand-keyed transcriptions of books printed in England, and English books printed abroad, 1473-1700. Sermons, pamphlets, laws, almanacs, ballads, science, literature. Converted from TCP's XML to Markdown. 60,329 texts, about 1.5 billion words, 25,369 from TCP Phase I and 34,960 from Phase II.

TCP keyed each text twice and proofed it to a 99.995% accuracy target. Characters the keyers could not read are marked, never guessed.

Text

The text is lightly normalised from the XML. Every change is decided by an XML tag or by the corpus itself:

  • —superscript abbreviations tagged in the XML are expanded on whole tokens (y^e the, y^t that, w^t with, y^u you, w^ch which ...), other superscripts are printed inline (Mr, St)
  • —TCP's own <expan> is used for <choice> abbreviations
  • —an abbreviation stroke after a vowel becomes n/m only when the rest of the corpus decides: the candidate occurs at least 20 times and 20x more often than the other (frō becomes from, vpō becomes vpon; thē stays, since then and them are both common)
  • —long s becomes s, vv becomes w

There is no spelling modernisation. u/v and i/j are as printed.

Gaps use TCP's own markers: • per illegible letter, 〈◊〉 for an illegible word, 〈…〉 for longer, 〈 in non-Latin alphabet 〉, 〈♫〉, 〈1 page missing〉 ... Nothing is filled in. About 150 books keep a few private-use characters from TCP's XML (mostly U+E002, the -que abbreviation).

Layout: # headings by division, paragraphs, verse one line per line, **Speaker** in drama, [Illustration: TCP's figure description], notes as Markdown footnotes [^n] at the end. Page breaks, running heads and catchwords are dropped. Each text starts with a short YAML front matter and a title line.

Sources

The Oxford TEI P5 edition (Oxford Text Archive, hdl 20.500.14106) is used when it exists and is complete. Otherwise the text comes from TCP's P4 XML bulk release: for 2 texts OTA does not publish (A55842, B25542), and for 454 where the P5 conversion lost paragraph text around figures (at least 47 consecutive P4 words absent from P5). source, p4_reason, source_file, source_sha256 and ota_handle are on every row.

sourcereasonbookswords
ota_p5-59,8731,452,158,482
tcp_p4p5_incomplete45451,113,577
tcp_p4notonota2249,214

Completeness

TCP gives 25,368 Phase I and 34,963 Phase II texts, 60,331 in all. OTA's EEBO collection has 60,328. This dataset has all of OTA's except B10185 (a one-page 1690 broadside), plus the 2 P4-only texts above. The Phase I books with an OTA record are exactly OTA's and TCP's 25,368 Phase I texts. B25542 is labelled Phase I but is not in either list.

Checks

Details are in verification.json.

  • —IDs compared with OTA's full EEBO list (60,328 records, harvested over OAI-PMH) and with TCP's catalogue TCP.csv.
  • —Phase II: all 34,958 books compared with TCP's P4 XML. 99.86% of the P4 words are in the text. Nearly all the rest are the expansions listed above (ye, yt, wt, fro ...).
  • —Phase I: 60 random books compared with OTA's P5 XML. 99.75% of the words are in the text, and no book is under 95%.
  • —Every book passes the converter's own two checks: invariant_ok (the letters and digits equal the source XML, with gaps, footnotes and added layout accounted for) and v1_ok (every changed word is a logged substitution).

Dates

date is the imprint date from the TCP header, as printed (1706 [1705], 169-?] ...). year is the certain bracketed year if there is one, else the first year. 73 books carry an imprint date after 1700, the latest 1865; these are TCP's dates, not errors.

Where the header gives no full year, year was filled from other TCP sources:

  • —21 books with an empty date take the imprint date from TCP's P4 header (date and year).
  • —258 books take year from the normalised Date in TCP's catalogue. Their date stays as printed. Catalogue years that contradict the printed date were not used (for example 1780 for 169-?]).
  • —11 books take year from the printed date once brackets are removed (16[47] is 1647, l643] is 1643).

31 books remain undated: their date gives only a decade (169-?]) or nothing.

Word counts

n_words uses our own tokeniser and sums to 1,503,521,273. Other tokenisers give a slightly different number. See STATS.md for counts by source, period, language and phase.

Official sources

To check the numbers above:

  • —TCP, EEBO-TCP: <https://www.textpartnership.net/pages/eebo-tcp.html> (Phase I and II counts, licence, downloads)
  • —Oxford Text Archive, EEBO collection, OAI-PMH set hdl_20.500.14106_5: <https://llds.ling-phil.ox.ac.uk/llds/oai/request?verb=ListRecords&metadataPrefix=oaidc&set=hdl20.500.14106_5>. Single texts are at http://hdl.handle.net/20.500.14106/<TCP id>.
  • —TCP catalogue: <https://github.com/textcreationpartnership/Texts> (TCP.csv)
  • —TCP texts on GitHub, one repository per text: <https://github.com/textcreationpartnership>

Licence and credit

TCP texts are CC0 1.0. Please credit the Text Creation Partnership (University of Michigan, University of Oxford and the libraries that funded it). Page images belong to ProQuest and are not included.