Team Ai
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Taylor658 /photonic-integrated-circuit-yield 🏭 Photonic Integrated Circuit Yield Dataset 📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing. ⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.texttext-generation100K<n<1M5 likes216 downloads23d agoHugging Face02cis-lmu /GlotStoryBook Dataset Description Story Books for 180 ISO-639-3 codes. The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset. This dataset consists of 2 subsets: default, which consists of 4 publishers: asp: African Storybook pb: Pratham Books lcb: Little Cree Books lida: LIDA Stories nalibali, which comes from Nal'ibali stories. Usage (HF Loader) default: from datasets import load_dataset dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.texttranslation10K<n<100K9 likes194 downloads14d agoHugging Face03Cinkui /discursos-congreso-es Discursos del Congreso de los Diputados (muestra balanceada) Descripción Este dataset contiene una muestra de 1088 intervenciones en el pleno del Congreso de los Diputados de España, entre 1996 y 2023. Se ha obtenido a partir del corpus ParlLawSpeech mediante un proceso de limpieza y un muestreo balanceado por año y partido político. Se ha creado como parte de la práctica de la asignatura Descubrimiento de conocimiento en datos complejos. Está pensado para tareas… See the full description on the dataset page: https://huggingface.co/datasets/Cinkui/discursos-congreso-es.texttext-classification1K<n<10K2 likes82 downloads8d agoHugging Face04ciol-research /multilevel-legal-reasoning Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai. 🧭 Purpose and Scope The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.tabulartext-generationn<1K7 likes74 downloads1y agoHugging Face05Lexsi /circuitkit-capitals-contrastive CircuitKIT capitals — contrastive pairs Twelve capital-city facts, each with an explicit counterfactual pair, for circuit discovery with CircuitKIT. column meaning question clean prompt, e.g. The capital of France is answer clean answer, e.g. Paris corrupted_question counterfactual prompt of the same shape, e.g. The capital of Germany is corrupted_answer its answer, e.g. Berlin Attribution-patching methods (EAP, EAP-IG, …) score a component by how much it… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/circuitkit-capitals-contrastive.texttext-generationn<1K0 likes65 downloads17d agoHugging Face06Snipe-City /workhorse-llm-bench Workhorse Local-LLM Bench 73 open-weight model configurations (68 distinct models, plus 4 speculative drafters) measured over five rounds, 20-28 September 2026, on one deliberately ordinary machine: a 2020-era laptop with a 6-core Intel(R) Core(TM) i5-10500H CPU @ 2.50GHz, 32 GB of DDR4-2667 and an NVIDIA GeForce GTX 1650 with Max-Q Design (4 GB). Engine: llama.cpp (CUDA), with MoE experts in system RAM and attention plus KV cache on the GPU. The question was practical: which… See the full description on the dataset page: https://huggingface.co/datasets/Snipe-City/workhorse-llm-bench.tabulartext-generation10K<n<100K0 likes55 downloads1d agoHugging Face07cilyy /Anime_subtitles_CN Dataset Card for Dataset Name This repo contains a csv file about anime subtitles crawl from open web.This dataset could be used for t2t,all the NLP projects expectionly of the anime domain.It's part.1,probably will have part.2. Dataset Description anime_subtitles.csv: Contains two features('name' and 'caption') and 4055 rows,about 400MB. Each name represent one season or movie, caption contaions all the dialogues that the characters speaks but no characters name or… See the full description on the dataset page: https://huggingface.co/datasets/cilyy/Anime_subtitles_CN.texttext-generation1K<n<10K2 likes47 downloads2y agoHugging Face08Adeptschneider /CiviVox-Swahili-text-corpus-v2.0 Swahili Text Dataset Overview This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language. Dataset Details Source: AfriBERTa Corpus (Swahili subset) Language: Swahili Size: 1.54M Format: Hugging Face Dataset Content The dataset consists of two main columns: id: A unique identifier for each text entry text:… See the full description on the dataset page: https://huggingface.co/datasets/Adeptschneider/CiviVox-Swahili-text-corpus-v2.0.texttext-generation1M<n<10M0 likes11 downloads2y agoHugging Face09DBbun /10M_CIRCULATIONAHA.120.052430_v1.0gated Synthetic UKBB South Asian ASCVD (10M Patients) This dataset contains 10 million synthetic patients with ancestry (European and South Asian), demographics, anthropometrics, labs, lifestyle/SES, comorbidities, medications, and time-to-event outcomes for ASCVD, heart failure, and atrial fibrillation. The dataset inspired by the published Circulation study. All data are artificially generated and contain no identifiable patient records. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/10M_CIRCULATIONAHA.120.052430_v1.0.tabulartext-generation10M<n<100M0 likes9 downloads10mo agoHugging Face10niqqyniqqy /CiviVox-Swahili-text-corpus-v2.0 Swahili Text Dataset Overview This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language. Dataset Details Source: AfriBERTa Corpus (Swahili subset) Language: Swahili Size: 1.54M Format: Hugging Face Dataset Content The dataset consists of two main columns: id: A unique identifier for each text entry text:… See the full description on the dataset page: https://huggingface.co/datasets/niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0.texttext-generation1M<n<10M0 likes7 downloads6mo agoHugging Face11DBbun /1M_CIRCULATIONAHA.120.052430_v1.0gated Synthetic UKBB South Asian ASCVD (1M Patients) Watch a demo This dataset contains 1 million synthetic patients with ancestry (European and South Asian), demographics, anthropometrics, labs, lifestyle/SES, comorbidities, medications, and time-to-event outcomes for ASCVD, heart failure, and atrial fibrillation. The dataset inspired by the published Circulation study. All data are artificially generated and contain no identifiable patient records. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/1M_CIRCULATIONAHA.120.052430_v1.0.tabulartext-generation1M<n<10M0 likes6 downloads9mo agoHugging Face12DBbun /500K_CIRCULATIONAHA.120.052430_v1.0gated Synthetic UKBB South Asian ASCVD (500K Patients) This dataset contains 0.5 million synthetic patients with ancestry (European and South Asian), demographics, anthropometrics, labs, lifestyle/SES, comorbidities, medications, and time-to-event outcomes for ASCVD, heart failure, and atrial fibrillation. The dataset inspired by the published Circulation study. All data are artificially generated and contain no identifiable patient records. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/500K_CIRCULATIONAHA.120.052430_v1.0.tabulartext-generation100K<n<1M0 likes5 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.