datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
photonic-integrated-circuit-yield
🏭 Photonic Integrated Circuit Yield Dataset
📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing.
⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.GlotStoryBook
Dataset Description
Story Books for 180 ISO-639-3 codes.
The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset.
This dataset consists of 2 subsets:
default, which consists of 4 publishers:
asp: African Storybook
pb: Pratham Books
lcb: Little Cree Books
lida: LIDA Stories
nalibali, which comes from Nal'ibali stories.
Usage (HF Loader)
default:
from datasets import load_dataset
dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.discursos-congreso-es
Discursos del Congreso de los Diputados (muestra balanceada)
Descripción
Este dataset contiene una muestra de 1088 intervenciones en el pleno del Congreso de los Diputados de España, entre 1996 y 2023. Se ha obtenido a partir del corpus ParlLawSpeech mediante un proceso de limpieza y un muestreo balanceado por año y partido político.
Se ha creado como parte de la práctica de la asignatura Descubrimiento de conocimiento en datos complejos. Está pensado para tareas… See the full description on the dataset page: https://huggingface.co/datasets/Cinkui/discursos-congreso-es.multilevel-legal-reasoning
Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations
Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi
Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai.
🧭 Purpose and Scope
The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.circuitkit-capitals-contrastive
CircuitKIT capitals — contrastive pairs
Twelve capital-city facts, each with an explicit counterfactual pair, for circuit
discovery with CircuitKIT.
column
meaning
question
clean prompt, e.g. The capital of France is
answer
clean answer, e.g. Paris
corrupted_question
counterfactual prompt of the same shape, e.g. The capital of Germany is
corrupted_answer
its answer, e.g. Berlin
Attribution-patching methods (EAP, EAP-IG, …) score a component by how much it… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/circuitkit-capitals-contrastive.workhorse-llm-bench
Workhorse Local-LLM Bench
73 open-weight model configurations (68 distinct models, plus 4
speculative drafters) measured over five rounds, 20-28 September 2026, on one deliberately ordinary
machine: a 2020-era laptop with a 6-core Intel(R) Core(TM) i5-10500H CPU @ 2.50GHz, 32 GB of DDR4-2667 and an
NVIDIA GeForce GTX 1650 with Max-Q Design (4 GB). Engine: llama.cpp (CUDA), with MoE experts in system RAM and attention plus KV
cache on the GPU.
The question was practical: which… See the full description on the dataset page: https://huggingface.co/datasets/Snipe-City/workhorse-llm-bench.Anime_subtitles_CN
Dataset Card for Dataset Name
This repo contains a csv file about anime subtitles crawl from open web.This dataset could be used for t2t,all the NLP projects expectionly of the anime domain.It's part.1,probably will have part.2.
Dataset Description
anime_subtitles.csv: Contains two features('name' and 'caption') and 4055 rows,about 400MB. Each name represent one season or movie, caption contaions all the dialogues that the characters speaks but no characters name or… See the full description on the dataset page: https://huggingface.co/datasets/cilyy/Anime_subtitles_CN.CiviVox-Swahili-text-corpus-v2.0
Swahili Text Dataset
Overview
This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language.
Dataset Details
Source: AfriBERTa Corpus (Swahili subset)
Language: Swahili
Size: 1.54M
Format: Hugging Face Dataset
Content
The dataset consists of two main columns:
id: A unique identifier for each text entry
text:… See the full description on the dataset page: https://huggingface.co/datasets/Adeptschneider/CiviVox-Swahili-text-corpus-v2.0.10M_CIRCULATIONAHA.120.052430_v1.0
Synthetic UKBB South Asian ASCVD (10M Patients)
This dataset contains 10 million synthetic patients with ancestry (European and South Asian), demographics, anthropometrics, labs, lifestyle/SES, comorbidities, medications, and time-to-event outcomes for ASCVD, heart failure, and atrial fibrillation. The dataset inspired by the published Circulation study.
All data are artificially generated and contain no identifiable patient records.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/10M_CIRCULATIONAHA.120.052430_v1.0.CiviVox-Swahili-text-corpus-v2.0
Swahili Text Dataset
Overview
This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language.
Dataset Details
Source: AfriBERTa Corpus (Swahili subset)
Language: Swahili
Size: 1.54M
Format: Hugging Face Dataset
Content
The dataset consists of two main columns:
id: A unique identifier for each text entry
text:… See the full description on the dataset page: https://huggingface.co/datasets/niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0.1M_CIRCULATIONAHA.120.052430_v1.0
Synthetic UKBB South Asian ASCVD (1M Patients)
Watch a demo
This dataset contains 1 million synthetic patients with ancestry (European and South Asian), demographics, anthropometrics, labs, lifestyle/SES, comorbidities, medications, and time-to-event outcomes for ASCVD, heart failure, and atrial fibrillation. The dataset inspired by the published Circulation study.
All data are artificially generated and contain no identifiable patient records.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/1M_CIRCULATIONAHA.120.052430_v1.0.500K_CIRCULATIONAHA.120.052430_v1.0
Synthetic UKBB South Asian ASCVD (500K Patients)
This dataset contains 0.5 million synthetic patients with ancestry (European and South Asian), demographics, anthropometrics, labs, lifestyle/SES, comorbidities, medications, and time-to-event outcomes for ASCVD, heart failure, and atrial fibrillation. The dataset inspired by the published Circulation study.
All data are artificially generated and contain no identifiable patient records.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/500K_CIRCULATIONAHA.120.052430_v1.0.
