Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hardikg2907 /github-code-html-css-2text1M<n<10M0 likes1k downloads2y agoHugging Face02davidguzmanr /CSS10-Multilingual-LJSpeech CSS10-Multilingual-LJSpeech Multilingual speech dataset combining LJSpeech (English) + CSS10 (10 languages) in a consistent LJSpeech format. Dataset Description This dataset merges: LJSpeech: High-quality English speech dataset CSS10: A collection of single-speaker speech datasets for 10 languages All audio files are provided in a consistent format suitable for TTS training. Features Each sample contains: audio: Waveform audio sampled at 22,050 Hz text:… See the full description on the dataset page: https://huggingface.co/datasets/davidguzmanr/CSS10-Multilingual-LJSpeech.audio10K<n<100K0 likes943 downloads7mo agoHugging Face03CSSNB /dota_v1.5image100K<n<1M0 likes897 downloads8mo agoHugging Face04hardikg2907 /github-code-html-css-1text1M<n<10M2 likes852 downloads2y agoHugging Face05cssi /SciSciGPT-SciSciNettext100M<n<1B0 likes462 downloads1y agoHugging Face06vumichien /preprocessed_jsut_jsss_css10_common_voice_11 Dataset Card for "preprocessed_jsut_jsss_css10_common_voice_11" More Information needed text10K<n<100K1 likes437 downloads4y agoHugging Face07av9ash /CSSR-S_labelled_suicidewatch_posts_reddit Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code. License and Citation This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following: @article{patil2025evaluating, title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.tabulartext-classification1K<n<10K0 likes360 downloads9mo agoHugging Face08hardikg2907 /github-code-html-csstext100K<n<1M4 likes227 downloads2y agoHugging Face09vumichien /preprocessed_jsut_jsss_css10_fleurs_common_voice_11 Dataset Card for "preprocessed_jsut_jsss_css10_fleurs_common_voice_11" More Information needed text10K<n<100K2 likes212 downloads4y agoHugging Face10Yaesir06 /CSSBench CSSBench: A Safety Evaluation Benchmark for Chinese Lightweight Language Models Overview CSSBench (Chinese-Specific Safety Benchmark) is a comprehensive evaluation framework designed to assess the safety robustness of Chinese Large Language Models (LLMs), with a specific emphasis on lightweight models (≤8B parameters). The benchmark bridges a critical evaluation gap by targeting Chinese-specific adversarial patterns—linguistic obfuscations such as homophones and Pinyin… See the full description on the dataset page: https://huggingface.co/datasets/Yaesir06/CSSBench.texttext-classification1K<n<10K3 likes206 downloads9mo agoHugging Face11vumichien /preprocessed_jsut_jsss_css10 Dataset Card for "preprocessed_jsut_jsss_css10" More Information needed text10K<n<100K0 likes186 downloads4y agoHugging Face12ouroboroscollective /evidence-bound-css Evidence-Bound CSS A provenance-first German CSS/HTML learning, debugging and repair corpus for LLM training. Release 0.4.1 The public release exposes 601 unique structured CSS knowledge records, an alternate 601-record SFT view, 100 source-grounded debugging records, the original 6-record browser-verified seed, plus a v0.3 quality layer with 120 runtime-verified repairs, 120 verifier-confirmed hard negatives, and 120 verifier-backed preference pairs across 30… See the full description on the dataset page: https://huggingface.co/datasets/ouroboroscollective/evidence-bound-css.texttext-generation1K<n<10K1 likes169 downloads5d agoHugging Face13Northwestern-CSSI /Sci2Pol-BenchSci2Pol-Bench Data, scripts, and recipes for the benchmark Sci2Pol-Bench, a comprehensive benchmark for evaluating large language models. About • Usage• Authors About The data consists of policy briefs obtained from Nature Energy, Nature Climate, Nature Cities, and Journal of Health and Social Behavior Policy Briefs. Policy briefs originally were introduced in the Nature Energy journal with the goal of: This format aims to provide… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/Sci2Pol-Bench.textsummarization1K<n<10K2 likes166 downloads1y agoHugging Face14Northwestern-CSSI /sciscinet-v2gated 📢🚨📣 Sciscinet-v2 Sciscinet-v2 is a refreshed update to SciSciNet which is a large-scale, integrated dataset designed to support research in the science of science domain. It combines scientific publications with their network of relationships to funding sources, patents, citations, and institutional affiliations, creating a rich ecosystem for analyzing scientific productivity, impact, and innovation. Know more. About Sciscinet-v2 The newer version Sciscinet-v2 is… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/sciscinet-v2.tabular1B<n<10B29 likes132 downloads1y agoHugging Face15fewshot-goes-multilingual /cs_squad-3.0 Dataset Card for Czech Simple Question Answering Dataset 3.0 This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section. Dataset Description The data contains questions and answers based on Czech wikipeadia articles. Each question has an answer (or more) and a selected part of the context as the evidence. A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.tabularquestion-answering1K<n<10K3 likes129 downloads3y agoHugging Face16gt-csse /false-citation-bench False Citation Bench False Citation Bench is a compact evaluation and inspection dataset for false or misleading case citations in legal documents. It contains 26 source documents, their PDFs, and manually reviewed citation annotations grounded in the local text extraction. Dataset contents The repository has one matching document in each directory: documents_txt/{index}__{case-name}__{filing}.txt documents_pdf/{index}__{case-name}__{filing}.pdf… See the full description on the dataset page: https://huggingface.co/datasets/gt-csse/false-citation-bench.documentn<1K2 likes124 downloads2mo agoHugging Face17hardikg2907 /github-code-html-css-split-3text1M<n<10M1 likes114 downloads2y agoHugging Face18AdhyanshVerma /html-css-js-cot 🌐 HTML/CSS/JS Reasoning Traces Dataset A high-quality, large-scale dataset of complex HTML, CSS, and JavaScript programming questions and model reasoning traces. 📊 Dataset Overview This repository contains a comprehensively structured dataset of reasoning traces for frontend web development tasks. The data maps intricate, multi-step prompts to step-by-step reasoning solutions generated by advanced Language Models. It is designed for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/html-css-js-cot.tabular1K<n<10K0 likes76 downloads3mo agoHugging Face19theprint /MultiRoundConvos-Code-JS-HTML-CSS-Pythontext1K<n<10K0 likes72 downloads10mo agoHugging Face20MAsad789565 /HTML-CSS-Website# Dataset This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant. # Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com) textn<1K4 likes57 downloads3y agoHugging Face21semarmendemx /csst2tabular10K<n<100K1 likes40 downloads3y agoHugging Face22blakegearin /hex-to-css-filter-covering-dataset CSS Filter Covering Dataset — snapshot 2026.10.07 For every one of the 16,777,216 sRGB colors, one integer 6-parameter CSS filter chain whose rendered loss is under 1% (W3C feColorMatrix semantics). id — sRGB color as integer 0xRRGGBB filter — witness CSS filter chain, applied to white loss — rendered loss recomputed from final rounded parameter values Sources and mirrors Canonical artifact (SQLite + gz + SHA-256): https://data.blakegearin.com GitHub release:… See the full description on the dataset page: https://huggingface.co/datasets/blakegearin/hex-to-css-filter-covering-dataset.tabular10M<n<100M0 likes38 downloads3d agoHugging Face23simoneteglia /css-deepfake-datasetimage10K<n<100K0 likes36 downloads2mo agoHugging Face24cssi /SciSciGPT-SciSciCorpustabular10K<n<100K1 likes34 downloads1y agoHugging Face25cssddnnc /sd_finetune_demoimagen<1K0 likes33 downloads3y agoHugging Face26CZLC /cs_snli Dataset Card for Czech SNLI Czech translation of the Stanford Natural Language Interface (SNLI) dataset with manual annotation of a SNLI subset. In addition to the entailment/contradiction/neutral inference, a "bad translation" class was added. The annotation was done by students of NLP or computational linguistics. 1499 same pairs were annotated by two students to check IAA. Dataset Details The annotation for Czech premise-hypothesis pairs is done on 165390 pairs… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_snli.tabulartext-classification10K<n<100K0 likes33 downloads2y agoHugging Face27IyedLahiani /HTML_CSS_CodeDataSet_100k_formatted_and_splittext100K<n<1M0 likes31 downloads1y agoHugging Face28Juliankrg /HTML_CSS_CodeDataSet_100ktext100K<n<1M5 likes30 downloads2y agoHugging Face29BSC-CSSH /AMSMB-line-transcription Dataset Card Dataset for line-level handwritten text recognition on medieval historical manuscripts, consisting of 3,369 lines (images of text lines with the associated transcription and metadata) from 100 digitized documents written by at least 80 different hands and spanning three centuries (from 1208 to 1499). This dataset is derived from the AMSMB dataset, which contains the full-page images of the digitized manuscripts and their associated transcriptions in the PageXML format.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-CSSH/AMSMB-line-transcription.imageimage-to-text1K<n<10K0 likes29 downloads1y agoHugging Face30mychen76 /dataset_CSSF12_552_en_qatext1K<n<10K0 likes28 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.