Team Ai
Datasetpublic

rafmacalaba/datause-agents-v3

datause-agents-v3 — data-use mentions Annotation workflow Labels are agent-produced under the project data-use doctrine; they are not owner rulings. The training and holdout splits are unchanged. Validation review refresh The val labels in both gliner and gliner2 were refreshed on 2026-09-29T17:54:34+00:00 using the gliner_val_review passage-review lane and shards 01–04. The review covered 1,260 passages and 2,171 candidate mentions drawn from… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-agents-v3.

sourceHugging Facecc-by-4.0updated 11d agoView on Hugging Face
0likes138downloads
Dataset Card

datause-agents-v3 — data-use mentions

Annotation workflow

Labels are agent-produced under the project data-use doctrine; they are not owner rulings. The training and holdout splits are unchanged.

Validation review refresh

The val labels in both gliner and gliner2 were refreshed on 2026-09-29T17:54:34+00:00 using the gliner_val_review passage-review lane and shards 01–04. The review covered 1,260 passages and 2,171 candidate mentions drawn from published validation labels plus GLiNER predictions scoring at least 0.6. All decisions remain annotated_by: agent, owner_reviewed: false. The other 862 candidate-free validation passages retain their existing empty labels. Because review scope is candidate-generated and does not assess mentions missed by the candidate policy, this split is an agent-reviewed audit set, not independent full-recall gold.

The review-specific source metadata and span counts are in split_stats.json.

Labels

  • —NAMED_DATA — a proper name, title, or acronym of a specific data source
  • —DESCRIPTIVE_DATA — a source described in words but not named
  • —VAGUE_DATA — generic data wording with no identifiable source

Provenance

  • —bundle: agent_annotations_v3 (agent-annotated; owner_reviewed: false)
  • —lanes: originstrat1, originstrat2, originstrat3, umarpads, umarpads_500
  • —carve: agents-v3 (origin-stratified) — document-disjoint, so no two passages of one document straddle train and eval, and stratified by origin, so each split holds each corpus family's exact share
  • —guard: 77 eval-text and 253 duplicate-text passages excluded; 128 eval documents checked, 0 leaks
  • —not evaluable: jdc_operational have no documents in val/holdout (too few to stratify), so they are train-only and no split metric says anything about them
  • —labels: labels.json, per-split counts and per-origin document counts: split_stats.json

Configs

  • —gliner — {"tokenized_text": [...], "ner": [[start, end, LABEL], ...]} (ner is inclusive on both ends)
  • —gliner2 — {"input": "...", "output": {"entities": {grade_key: [span strings]}}}

Splits

21,073 passages · 16,630 spans · 12,831 documents, document-disjoint.

splitrowsspansdocumentsNAMEDDESCRIPTIVEVAGUE
train16,84312,89110,2659,7233,14226
val2,1221,8211,2831,2755388
holdout2,1081,9181,2831,36054810

Usage

bash
uv run python training/finetune_gliner.py \
    --dataset rafmacalaba/datause-agents-v3 --config gliner \
    --epochs 5 --batch-size 16 --output-dir models/gliner_agents_v3
python
from datasets import load_dataset
ds = load_dataset("rafmacalaba/datause-agents-v3", "gliner")   # ds['train'] / ds['val'] / ds['holdout']