OO-LD/oold-wikidata-schemaorg-documents
Wikipedia leads, as the corpus measured them The article leads that oold-llm-bench's Wikidata-schema.org corpus cites: 959 documents, one per entity, each the plain-text introduction of a named revision with its whitespace collapsed. Published because the alternative does not work. The corpus commits each lead's revision id and sha256 rather than its text, and a third party was meant to re-fetch. Extracts are not versioned: the MediaWiki API ignores revids for prop=extracts and… See the full description on the dataset page: https://huggingface.co/datasets/OO-LD/oold-wikidata-schemaorg-documents.
Wikipedia leads, as the corpus measured them
The article leads that oold-llm-bench's Wikidata-schema.org corpus cites: 959 documents, one per entity, each the plain-text introduction of a named revision with its whitespace collapsed.
Published because the alternative does not work. The corpus commits each lead's revision id and sha256 rather than its text, and a third party was meant to re-fetch. Extracts are not versioned: the MediaWiki API ignores revids for prop=extracts and serves the current lead, so a lead edited since cannot be recovered at all. Measured 2026-10-07, 28 of 959 had been edited and had to be re-pinned rather than re-fetched.
So the bytes are published here instead, and nobody fetches anything.
from huggingface_hub import hf_hub_download
import json
path = hf_hub_download("OO-LD/oold-wikidata-schemaorg-documents", "documents.json", repo_type="dataset")
documents = json.loads(open(path, encoding="utf-8").read())Keyed by the corpus's record id, wds-<QID>.
Licence and attribution
Text from the English Wikipedia, used under CC BY-SA 4.0. Each lead is the work of that article's contributors; the article and the exact revision are named in the corpus record that cites it, which carries the title, the revision id and the sha256 of the text here.
This dataset is CC BY-SA 4.0, which is why it is a dataset of its own: the benchmark that reads it is Apache-2.0 and the two licences do not mix in one repository.
What this is not
Not a Wikipedia dump and not a sample of one. 85 schema.org classes were drawn from Wikidata, and these are the leads of the articles about those subjects. The selection is described by the benchmark, and the facts a lead was found to state are its ground truth.
