datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hebrew_synth_linesrecursive-lines
Recursive Lines: A Dual-Track Adversarial Benchmark
Recursive Lines is a diagnostic suite for detecting "High-Agency Deception" in Large Language Models. It serves as the reference implementation for the Constraint Cascade Model (FAccT 2026) and the Agency Index metric.
1. Overview
Current LLM benchmarks measure capability (MMLU) or safety (Refusal). They fail to measure Agency—the thermodynamic distinction between stochastic error (hallucination) and strategic intent… See the full description on the dataset page: https://huggingface.co/datasets/OstensibleParadox/recursive-lines.lines_hu_v4lines_hu_v2_1
This collection of data contains the different dataset
Statistics about this data :
We have 935213 pairs (image, text) we do splitting to this data into (train,val, and test) sets
The test set (9352) is up to 0.01 and we use 0.05 and 0.1 for val for different models.
The data format is jsonl where every row contains dictionary that has 2 keys (file_name, text) In addition to store
them in parquet format which is an efficient way for dealing with Bigdata.
1- SureNames for Hungarian Names… See the full description on the dataset page: https://huggingface.co/datasets/AlhitawiMohammed22/lines_hu_v2_1.receipt-household-lines-odbl
Receipt Household & Grocery Lines (ODbL)
Synthetic US receipt line items built from real products in the Open Food Facts family of databases, each with a
POS-style printed label, a readable product name and one of 39 spending categories (taxonomy v2, taxonomy.json).
This is the share-alike part of the Receipt Kitten item-namer training data. The permissive part (real receipts,
USDA-based synthetic lines, teacher-generated household inventory) is
lxe/receipt-line-items-taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/lxe/receipt-household-lines-odbl.moi-myanmar-articles-lines
MOI Myanmar Articles Dataset - Lines (DatarrX/moi-myanmar-articles-lines)
Dataset Description
The MOI Myanmar Articles - Lines dataset is a derivative corpus created from the official articles published on the Ministry of Information (MOI) website of the Republic of the Union of Myanmar.
Unlike the main dataset (moi-myanmar-articles), which contains full-length article texts, this dataset has been systematically split line-by-line (sentence-by-sentence). This… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/moi-myanmar-articles-lines.applescript-lines-annotated
Dataset Card for "applescript-lines-annotated"
Description
This is a dataset of single lines of AppleScript code scraped from GitHub and GitHub Gist and manually annotated with descriptions, intents, prompts, and other metadata.
Content
Each row contains 8 features:
text - The raw text of the AppleScript code.
source - The name of the file from which the line originates.
type - Either compiled (files using the .scpt extension) or uncompiled (everything else).… See the full description on the dataset page: https://huggingface.co/datasets/HelloImSteven/applescript-lines-annotated.shakespeare-lines
Shakespeare Lines Dataset
The Shakespeare Lines dataset contains cleaned, line-by-line excerpts from the Complete Works of William Shakespeare. This dataset is curated for use in training and fine-tuning language models on literary or archaic English. It has been stripped of metadata, scene directions, headers/footers, and other non-dialogue filler commonly found in public domain eBooks.
Dataset Structure
Each example contains:
text: A single line of dialogue from one of… See the full description on the dataset page: https://huggingface.co/datasets/benchaffe/shakespeare-lines.
