rafmacalaba/datause-extracted-human473-docs
datause-extracted-human473-docs Every passage of the 162 documents behind the 473 human-validated holdout spans of the data-use annotation campaign: population spans documents annotator190 190 134 jdc283 283 28 total 473 162 Configs gliner, bio, gliner2 — row-for-row subset of rafmacalaba/datause-extracted (revision 15812843687e2ec81b261b5895f2019b50a8f97e): same columns, same split files, rows verbatim. Rows per config: gliner/train 1963… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-extracted-human473-docs.
datause-extracted-human473-docs
Every passage of the 162 documents behind the 473 human-validated holdout spans of the data-use annotation campaign:
Configs
gliner,bio,gliner2— row-for-row subset of [`rafmacalaba/datause-extracted`](https://huggingface.co/datasets/rafmacalaba/datause-extracted) (revision15812843687e2ec81b261b5895f2019b50a8f97e): same columns, same split files, rows verbatim. Rows per config:gliner/train1963,gliner/val414,gliner/holdout342,bio/train1963,bio/val414,bio/holdout342,gliner2/train1963,gliner2/val414,gliner2/holdout342. These cover 158 of the 162 documents; the source repo has no rows for the remaining 4 (fcv_pads_east_africa:019636,fcv_pads_east_africa:019977,fcv_pads_east_africa:013470,fcv_pads_east_africa:013125).passages— the complete chunk set of all 162 documents (14568 rows), taken from the upstream extraction metarafmacalaba/fcv-extractions-meta-tiered. Columns: `population,corpus_id,doc_id,page,chunk,origin,title,pdf_url,project_id,country,document_type,input_text. Use this config when the question is "what text do these documents contain" — it is the only view that includes the 4 documentsdatause-extracted` is missing.gold_spans.jsonl(repo root) — the 473 spans:key,population,surface,human_keep,corpus_id,page,chunk,title,pdf_url,text(the passage the span was annotated in, fromrafmacalaba/datause-displacement-reviewedrev24c1a0e).
Selection
A document is in scope iff it backs at least one of the 473 spans. The span -> document join is corpus_id + page + chunk from the judge sheets (exact, not title-based); datause-extracted rows are kept by pdf_url or corpus_id, so every passage of a selected document is included, spans-free chunks as well.
Surface check: 473/473 gold surfaces occur verbatim inside a passages chunk of their own document. The residual are chunk-boundary and markdown-normalisation artefacts of re-parsing the PDFs, not missing documents; the annotated passage of every span is in gold_spans.jsonl.
Caveats
- The source repo's split assignment is not document-disjoint. Documents of this set appear in more than one of
train/val/holdoutindatause-extracted, so the same document can be in two files here too. Thesplitcolumn is preserved verbatim; these are faithful subsets, not a re-split. Treattrainas neither clean nor a contamination-free reference for these documents. - Spans shipped inside
gliner/bio/gliner2are raw model predictions, not gold. Gold labels live ingold_spans.jsonland in `rafmacalaba/datause-displacement-reviewed` (probe_reviewed, splitholdout, kindsannotator/jdc). pdf_urlis the document key; a handful oftitlevalues upstream are generic document-type strings ("Appraisal Project Information Document (PID)").- Provenance of the document set:
analysis/holdout_documents_annotator190_jdc283.{json,tsv}in the source project.
