Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01davidguzmanr /CSS10-Multilingual-LJSpeech CSS10-Multilingual-LJSpeech Multilingual speech dataset combining LJSpeech (English) + CSS10 (10 languages) in a consistent LJSpeech format. Dataset Description This dataset merges: LJSpeech: High-quality English speech dataset CSS10: A collection of single-speaker speech datasets for 10 languages All audio files are provided in a consistent format suitable for TTS training. Features Each sample contains: audio: Waveform audio sampled at 22,050 Hz text:… See the full description on the dataset page: https://huggingface.co/datasets/davidguzmanr/CSS10-Multilingual-LJSpeech.audio10K<n<100K0 likes946 downloads7mo agoHugging Face02hardikg2907 /github-code-html-css-2text1M<n<10M0 likes899 downloads2y agoHugging Face03hardikg2907 /github-code-html-css-1text1M<n<10M2 likes849 downloads2y agoHugging Face04vumichien /preprocessed_jsut_jsss_css10_common_voice_11 Dataset Card for "preprocessed_jsut_jsss_css10_common_voice_11" More Information needed text10K<n<100K1 likes692 downloads4y agoHugging Face05CSSNB /dota_v1.5image100K<n<1M0 likes543 downloads8mo agoHugging Face06cssi /SciSciGPT-SciSciNettext100M<n<1B0 likes439 downloads1y agoHugging Face07av9ash /CSSR-S_labelled_suicidewatch_posts_reddit Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code. License and Citation This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following: @article{patil2025evaluating, title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.tabulartext-classification1K<n<10K0 likes332 downloads8mo agoHugging Face08Mo7art /Stack2Graph_KG_css CSS StackOverflow Knowledge Graph Summary This Hugging Face dataset repository contains the CSS shard of the Stack2Graph StackOverflow Knowledge Graph. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content. Stack2Graph source:… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_css.100M<n<1B0 likes315 downloads3mo agoHugging Face09vumichien /preprocessed_jsut_jsss_css10_fleurs_common_voice_11 Dataset Card for "preprocessed_jsut_jsss_css10_fleurs_common_voice_11" More Information needed text10K<n<100K2 likes305 downloads4y agoHugging Face10cssen /gso_rendered_images0 likes269 downloads3y agoHugging Face11hardikg2907 /github-code-html-csstext100K<n<1M4 likes231 downloads2y agoHugging Face12vumichien /common_voice_large_jsut_jsss_css10 Dataset Card for vumichien/common_voice_large_jsut_jsss_css10 automatic-speech-recognition0 likes230 downloads4y agoHugging Face13uglyducking /libri_css0 likes227 downloads1y agoHugging Face14Yaesir06 /CSSBench CSSBench: A Safety Evaluation Benchmark for Chinese Lightweight Language Models Overview CSSBench (Chinese-Specific Safety Benchmark) is a comprehensive evaluation framework designed to assess the safety robustness of Chinese Large Language Models (LLMs), with a specific emphasis on lightweight models (≤8B parameters). The benchmark bridges a critical evaluation gap by targeting Chinese-specific adversarial patterns—linguistic obfuscations such as homophones and Pinyin… See the full description on the dataset page: https://huggingface.co/datasets/Yaesir06/CSSBench.texttext-classification1K<n<10K3 likes210 downloads8mo agoHugging Face15vumichien /preprocessed_jsut_jsss_css10 Dataset Card for "preprocessed_jsut_jsss_css10" More Information needed text10K<n<100K0 likes209 downloads4y agoHugging Face16cssen /audio_visual_starss23_son1 likes207 downloads3y agoHugging Face17Northwestern-CSSI /sciscinet-v2gated 📢🚨📣 Sciscinet-v2 Sciscinet-v2 is a refreshed update to SciSciNet which is a large-scale, integrated dataset designed to support research in the science of science domain. It combines scientific publications with their network of relationships to funding sources, patents, citations, and institutional affiliations, creating a rich ecosystem for analyzing scientific productivity, impact, and innovation. Know more. About Sciscinet-v2 The newer version Sciscinet-v2 is… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/sciscinet-v2.tabular1B<n<10B26 likes174 downloads1y agoHugging Face18verify-ppt /marin-starcoderdata_css0 likes173 downloads6mo agoHugging Face19cssbsnu /Thal-Kak_local_db Thal-Kak local MSA, template databases The sequence and template databases that the local MSA modes of Thal-Kak search — --msa mmseqs_local, --msa hhblits_local, --msa mmseqs_hhblits_local, and local template search on any of them. Install these with install_db.sh, not by hand. Every file here is a multi-gigabyte .tar.zst holding a prebuilt MMseqs2 or HH-suite database; the installer verifies it, unpacks it into place and renames the files to the layout the pipeline expects.… See the full description on the dataset page: https://huggingface.co/datasets/cssbsnu/Thal-Kak_local_db.100B<n<1T2 likes154 downloads1mo agoHugging Face20Northwestern-CSSI /Sci2Pol-BenchSci2Pol-Bench Data, scripts, and recipes for the benchmark Sci2Pol-Bench, a comprehensive benchmark for evaluating large language models. About • Usage• Authors About The data consists of policy briefs obtained from Nature Energy, Nature Climate, Nature Cities, and Journal of Health and Social Behavior Policy Briefs. Policy briefs originally were introduced in the Nature Energy journal with the goal of: This format aims to provide… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/Sci2Pol-Bench.textsummarization1K<n<10K2 likes152 downloads1y agoHugging Face21xhd668 /csscvideo1K<n<10K0 likes151 downloads9d agoHugging Face22fewshot-goes-multilingual /cs_squad-3.0 Dataset Card for Czech Simple Question Answering Dataset 3.0 This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section. Dataset Description The data contains questions and answers based on Czech wikipeadia articles. Each question has an answer (or more) and a selected part of the context as the evidence. A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.tabularquestion-answering1K<n<10K3 likes133 downloads3y agoHugging Face23gt-csse /false-citation-bench False Citation Bench False Citation Bench is a compact evaluation and inspection dataset for false or misleading case citations in legal documents. It contains 26 source documents, their PDFs, and manually reviewed citation annotations grounded in the local text extraction. Dataset contents The repository has one matching document in each directory: documents_txt/{index}__{case-name}__{filing}.txt documents_pdf/{index}__{case-name}__{filing}.pdf… See the full description on the dataset page: https://huggingface.co/datasets/gt-csse/false-citation-bench.documentn<1K2 likes125 downloads2mo agoHugging Face24ayousanz /css10-ljspeech CSS10-LJSpeech CSS10-LJSpeech は、Park et al. が公開した CSS10 データセットを、LJSpeech互換フォーマットに変換した10言語の音声合成用データセットです。各言語の文学作品を音声化した高品質な音声データを提供し、LJSpeechフォーマット(id|text & wavs/*.wav)に統一されています。 データ概要 項目 値 話者数 10 (言語別) 総音声数 64,196 合計時間 約 140 時間 サンプリングレート 22,050 Hz 音声フォーマット IEEE浮動小数点 (32bit) テキスト言語 10言語 フォーマット `id 言語別統計 言語 言語コード 音声数 合計時間 ドイツ語 de 7,428 16.14時間 ギリシャ語 el 1,844 4.14時間 スペイン語 es 11,016 19.15時間 フィンランド語 fi 4,842 10.53時間… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/css10-ljspeech.audio0 likes112 downloads1y agoHugging Face25ouroboroscollective /evidence-bound-css Evidence-Bound CSS A provenance-first German CSS/HTML learning, debugging and repair corpus for LLM training. Release 0.4.1 The public release exposes 601 unique structured CSS knowledge records, an alternate 601-record SFT view, 100 source-grounded debugging records, the original 6-record browser-verified seed, plus a v0.3 quality layer with 120 runtime-verified repairs, 120 verifier-confirmed hard negatives, and 120 verifier-backed preference pairs across 30… See the full description on the dataset page: https://huggingface.co/datasets/ouroboroscollective/evidence-bound-css.texttext-generation1K<n<10K1 likes96 downloads1d agoHugging Face26ayousanz /css10-ljspeech-multilingual CSS10 + LJSpeech Multilingual Dataset A unified multilingual speech dataset combining CSS10 (10 languages) and LJSpeech (English) in a consistent LJSpeech format. Dataset Description This dataset merges: CSS10: A collection of single-speaker speech datasets for 10 languages LJSpeech: High-quality English speech dataset (Linda Johnson) All audio files are provided in a consistent format suitable for TTS training. Languages and Statistics Language Code… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/css10-ljspeech-multilingual.text-to-speech10K<n<100K2 likes95 downloads1y agoHugging Face27Mo7art /Stack2Graph_VD_css CSS StackOverflow Vector Dataset Summary This Hugging Face dataset repository contains the CSS shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding, and… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_css.feature-extraction0 likes92 downloads3mo agoHugging Face28RoxasYTB /css-radio-french-ljspeech Counter-Strike: Source Radio - French LJSpeech-format dataset of the French radio voice from Counter-Strike: Source. Lines: 440 Composition: 40 original radio lines + 400 French ElevenLabs v4 recordings Format: metadata.csv + wavs/ License: CC BY 4.0 (game content) Metadata format 000001.wav|transcription 0 likes85 downloads6d agoHugging Face29theprint /MultiRoundConvos-Code-JS-HTML-CSS-Pythontext1K<n<10K0 likes83 downloads9mo agoHugging Face30Northwestern-CSSI /sciscinet-v1 SciSciNet-v1 This is a repository for the primary version of Sciscinet-v1, a large-scale open data lake for the science of science research. We have recently released the second version of Sciscinet, Sciscinet-v2. It is available as Northwestern-CSSI/sciscinet-v2 on huggingface. Click here to view Sciscinet-v2 on huggingface. 📢🚨📣 Sciscinet-v2 Sciscinet-v2 is a refreshed update to SciSciNet which is a large-scale, integrated dataset designed to support… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/sciscinet-v1.3 likes77 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.