Team Ai
Datasetpublic

lm2445/FinTagging1000_table

FinTagging Table Context Extraction Split For HTML table inputs, the target is a JSON list of numeric entity/datatype pairs enriched with deterministic row and column context. Missing row or column context is represented as null. The dataset is derived from FinTagging_800_200_HF and preserves the original train/test assignment by source_sample_idx and context_id. The XBRL concept tag is intentionally omitted from the target. Splits Split Samples Output… See the full description on the dataset page: https://huggingface.co/datasets/lm2445/FinTagging1000_table.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes15downloads
Dataset Card

FinTagging Table Context Extraction Split

For HTML table inputs, the target is a JSON list of numeric entity/datatype pairs enriched with deterministic row and column context. Missing row or column context is represented as null.

The dataset is derived from FinTagging_800_200_HF and preserves the original train/test assignment by source_sample_idx and context_id. The XBRL concept tag is intentionally omitted from the target.

Splits

SplitSamplesOutput entriesDuplicate output key groups
train46410,83181
test1102,35211

Columns

  • —source_sample_idx: original source row index.
  • —context_id: original context identifier.
  • —split: original split label.
  • —input_type: table or text.
  • —input: raw table HTML for table data, cleaned plain text for text data.
  • —output: JSON target string.
  • —output_entities: structured version of output.
  • —entity_metadata: deterministic parser metadata for auditing only.

Validation

CheckValue
source sample index overlap0
context ID overlap0
passedTrue