Team Ai
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01av9ash /CSSR-S_labelled_suicidewatch_posts_reddit Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code. License and Citation This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following: @article{patil2025evaluating, title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.tabulartext-classification1K<n<10K0 likes332 downloads8mo agoHugging Face02Northwestern-CSSI /sciscinet-v2gated 📢🚨📣 Sciscinet-v2 Sciscinet-v2 is a refreshed update to SciSciNet which is a large-scale, integrated dataset designed to support research in the science of science domain. It combines scientific publications with their network of relationships to funding sources, patents, citations, and institutional affiliations, creating a rich ecosystem for analyzing scientific productivity, impact, and innovation. Know more. About Sciscinet-v2 The newer version Sciscinet-v2 is… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/sciscinet-v2.tabular1B<n<10B26 likes174 downloads1y agoHugging Face03fewshot-goes-multilingual /cs_squad-3.0 Dataset Card for Czech Simple Question Answering Dataset 3.0 This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section. Dataset Description The data contains questions and answers based on Czech wikipeadia articles. Each question has an answer (or more) and a selected part of the context as the evidence. A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.tabularquestion-answering1K<n<10K3 likes133 downloads3y agoHugging Face04AdhyanshVerma /html-css-js-cot 🌐 HTML/CSS/JS Reasoning Traces Dataset A high-quality, large-scale dataset of complex HTML, CSS, and JavaScript programming questions and model reasoning traces. 📊 Dataset Overview This repository contains a comprehensively structured dataset of reasoning traces for frontend web development tasks. The data maps intricate, multi-step prompts to step-by-step reasoning solutions generated by advanced Language Models. It is designed for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/html-css-js-cot.tabular1K<n<10K0 likes68 downloads3mo agoHugging Face05cssi /SciSciGPT-SciSciCorpustabular10K<n<100K1 likes43 downloads1y agoHugging Face06semarmendemx /csst2tabular10K<n<100K1 likes36 downloads3y agoHugging Face07CZLC /cs_snli Dataset Card for Czech SNLI Czech translation of the Stanford Natural Language Interface (SNLI) dataset with manual annotation of a SNLI subset. In addition to the entailment/contradiction/neutral inference, a "bad translation" class was added. The annotation was done by students of NLP or computational linguistics. 1499 same pairs were annotated by two students to check IAA. Dataset Details The annotation for Czech premise-hypothesis pairs is done on 165390 pairs… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_snli.tabulartext-classification10K<n<100K0 likes23 downloads2y agoHugging Face08ayousanz /css10-ja-ljspeech-auditgated CSS10 Japanese LJSpeech — aggregate audit This one-row audit describes ayousanz/css10-ja-ljspeech at revision 149edaf267ff8048c19ca8324fce07ea7423cd14. It excludes transcript text, utterance IDs, audio paths, hashes, and audio payloads. The metadata has 6,841 rows and exactly matches 6,841 ZIP audio members. There are two empty-text rows, one language-review row, and 23 repeated-text groups. Bounded WAV-header checks show 22.05 kHz mono 32-bit IEEE-float audio; size-derived… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/css10-ja-ljspeech-audit.tabularn<1K0 likes20 downloads14d agoHugging Face09Bashroom /css-colorstabularn<1K1 likes13 downloads2y agoHugging Face10EPI-Eval /jhu-csse-covid JHU CSSE COVID-19 — global daily (archived) JHU stopped active maintenance 2023-03-09 and archived the repo. Province-level rolls are present in the source files but dropped here (province names are free-text, not ISO 3166-2; a per-country crosswalk is needed and is out of scope for this v0.1 ingest). Source: https://github.com/CSSEGISandData/COVID-19 Coverage Time: 2020-01-22 → 2023-03-09 Cadence: daily (observed median spacing: 1 days) Geography levels: national —… See the full description on the dataset page: https://huggingface.co/datasets/EPI-Eval/jhu-csse-covid.tabular100K<n<1M0 likes13 downloads5mo agoHugging Face11antoine3 /CSS2_UQtabular1M<n<10M0 likes5 downloads7mo agoHugging Face12antoine3 /css2-uq-mmlu-protabular1M<n<10M0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.