rafmacalaba/datause-agents-v3
datause-agents-v3 — data-use mentions Annotation workflow Labels are agent-produced under the project data-use doctrine; they are not owner rulings. The training and holdout splits are unchanged. Validation review refresh The val labels in both gliner and gliner2 were refreshed on 2026-09-29T17:54:34+00:00 using the gliner_val_review passage-review lane and shards 01–04. The review covered 1,260 passages and 2,171 candidate mentions drawn from… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-agents-v3.
datause-agents-v3 — data-use mentions
Annotation workflow
Labels are agent-produced under the project data-use doctrine; they are not owner rulings. The training and holdout splits are unchanged.
Validation review refresh
The val labels in both gliner and gliner2 were refreshed on 2026-09-29T17:54:34+00:00 using the gliner_val_review passage-review lane and shards 01–04. The review covered 1,260 passages and 2,171 candidate mentions drawn from published validation labels plus GLiNER predictions scoring at least 0.6. All decisions remain annotated_by: agent, owner_reviewed: false. The other 862 candidate-free validation passages retain their existing empty labels. Because review scope is candidate-generated and does not assess mentions missed by the candidate policy, this split is an agent-reviewed audit set, not independent full-recall gold.
The review-specific source metadata and span counts are in split_stats.json.
Labels
NAMED_DATA— a proper name, title, or acronym of a specific data sourceDESCRIPTIVE_DATA— a source described in words but not namedVAGUE_DATA— generic data wording with no identifiable source
Provenance
- bundle:
agent_annotations_v3(agent-annotated;owner_reviewed: false) - lanes: originstrat1, originstrat2, originstrat3, umarpads, umarpads_500
- carve:
agents-v3(origin-stratified) — document-disjoint, so no two passages of one document straddle train and eval, and stratified by origin, so each split holds each corpus family's exact share - guard: 77 eval-text and 253 duplicate-text passages excluded; 128 eval documents checked, 0 leaks
- not evaluable: jdc_operational have no documents in val/holdout (too few to stratify), so they are train-only and no split metric says anything about them
- labels:
labels.json, per-split counts and per-origin document counts:split_stats.json
Configs
gliner—{"tokenized_text": [...], "ner": [[start, end, LABEL], ...]}(neris inclusive on both ends)gliner2—{"input": "...", "output": {"entities": {grade_key: [span strings]}}}
Splits
21,073 passages · 16,630 spans · 12,831 documents, document-disjoint.
Usage
uv run python training/finetune_gliner.py \
--dataset rafmacalaba/datause-agents-v3 --config gliner \
--epochs 5 --batch-size 16 --output-dir models/gliner_agents_v3from datasets import load_dataset
ds = load_dataset("rafmacalaba/datause-agents-v3", "gliner") # ds['train'] / ds['val'] / ds['holdout']