Team Ai
Datasetpublic

rafmacalaba/datause-extracted-human473-docs

datause-extracted-human473-docs Every passage of the 162 documents behind the 473 human-validated holdout spans of the data-use annotation campaign: population spans documents annotator190 190 134 jdc283 283 28 total 473 162 Configs gliner, bio, gliner2 — row-for-row subset of rafmacalaba/datause-extracted (revision 15812843687e2ec81b261b5895f2019b50a8f97e): same columns, same split files, rows verbatim. Rows per config: gliner/train 1963… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-extracted-human473-docs.

sourceHugging Facecc-by-4.0updated 25d agoView on Hugging Face
0likes138downloads
Dataset Card

datause-extracted-human473-docs

Every passage of the 162 documents behind the 473 human-validated holdout spans of the data-use annotation campaign:

populationspansdocuments
annotator190190134
jdc28328328
total473162

Configs

  • —gliner, bio, gliner2 — row-for-row subset of [`rafmacalaba/datause-extracted`](https://huggingface.co/datasets/rafmacalaba/datause-extracted) (revision 15812843687e2ec81b261b5895f2019b50a8f97e): same columns, same split files, rows verbatim. Rows per config: gliner/train 1963, gliner/val 414, gliner/holdout 342, bio/train 1963, bio/val 414, bio/holdout 342, gliner2/train 1963, gliner2/val 414, gliner2/holdout 342. These cover 158 of the 162 documents; the source repo has no rows for the remaining 4 (fcv_pads_east_africa:019636, fcv_pads_east_africa:019977, fcv_pads_east_africa:013470, fcv_pads_east_africa:013125).
  • —passages — the complete chunk set of all 162 documents (14568 rows), taken from the upstream extraction meta rafmacalaba/fcv-extractions-meta-tiered. Columns: `population, corpus_id, doc_id, page, chunk, origin, title, pdf_url, project_id, country, document_type, input_text. Use this config when the question is "what text do these documents contain" — it is the only view that includes the 4 documents datause-extracted` is missing.
  • —gold_spans.jsonl (repo root) — the 473 spans: key, population, surface, human_keep, corpus_id, page, chunk, title, pdf_url, text (the passage the span was annotated in, from rafmacalaba/datause-displacement-reviewed rev 24c1a0e).

Selection

A document is in scope iff it backs at least one of the 473 spans. The span -> document join is corpus_id + page + chunk from the judge sheets (exact, not title-based); datause-extracted rows are kept by pdf_url or corpus_id, so every passage of a selected document is included, spans-free chunks as well.

Surface check: 473/473 gold surfaces occur verbatim inside a passages chunk of their own document. The residual are chunk-boundary and markdown-normalisation artefacts of re-parsing the PDFs, not missing documents; the annotated passage of every span is in gold_spans.jsonl.

Caveats

  • —The source repo's split assignment is not document-disjoint. Documents of this set appear in more than one of train/val/holdout in datause-extracted, so the same document can be in two files here too. The split column is preserved verbatim; these are faithful subsets, not a re-split. Treat train as neither clean nor a contamination-free reference for these documents.
  • —Spans shipped inside gliner/bio/gliner2 are raw model predictions, not gold. Gold labels live in gold_spans.jsonl and in `rafmacalaba/datause-displacement-reviewed` (probe_reviewed, split holdout, kinds annotator / jdc).
  • —pdf_url is the document key; a handful of title values upstream are generic document-type strings ("Appraisal Project Information Document (PID)").
  • —Provenance of the document set: analysis/holdout_documents_annotator190_jdc283.{json,tsv} in the source project.