Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01christinacdl /binary_hate_speechtexttext-classification10K<n<100K0 likes678 downloads3y agoHugging Face02SetFit /ethos_binaryThis is the binary split of ethos, split into train and test. It contains comments annotated for hate speech or not. textn<1K1 likes343 downloads5y agoHugging Face03tuhink /cambench_binary_eval CameraBench Binary Evaluation Dataset A balanced VQA dataset for evaluating camera motion understanding in videos. 📊 Dataset Statistics Total Questions: 384 Unique Videos: 119 Unique Questions: 31 Yes Answers: 192 (50.0%) No Answers: 192 (50.0%) Balance Ratio: 1.00 Total Size: 126.16 MB (0.12 GB) Average Video Size: 1.06 MB 🎯 Task Categories This dataset covers various camera motion tasks including: Static: 42 questions Move In: 29 questions Pan Left: 24… See the full description on the dataset page: https://huggingface.co/datasets/tuhink/cambench_binary_eval.imagevisual-question-answeringn<1K0 likes236 downloads1y agoHugging Face04leeyujun /Beyond-Binary-Instrument-QA 🎵 Beyond Binary Instrument QA:Probing Instrument Grounding in Music Audio-Language Models Yujun Lee · Joonhyeok Shin · Hyoeun Kim · Kyuhong Shim Sungkyunkwan University 📄 arXiv &nbsp;&nbsp;|&nbsp;&nbsp; 🤗 Dataset Benchmark release. Five complementary evaluation configurations test instrument presence, reduced genre-prior reliance, fine-grained discrimination, long-context multi-label recognition, and temporal localization. The release contains 15… See the full description on the dataset page: https://huggingface.co/datasets/leeyujun/Beyond-Binary-Instrument-QA.audioaudio-classification10K<n<100K1 likes105 downloads25d agoHugging Face05bingbangboom /editlens_iclr_binary_reasoning bingbangboom/editlens_iclr_binary_reasoning This dataset is a binary-classification subset drawn from the training split of pangram/editlens_iclr dataset. It isolates purely human-crafted texts (human_written) against purely synthetic content (ai_generated), strictly filtering out the overlapping ai_edited classification cluster for binary classification tasks. The primary augmentation of this dataset is the inclusion of Reasoning Traces (Chain of Thought). Every single text… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/editlens_iclr_binary_reasoning.texttext-classification1K<n<10K0 likes32 downloads5mo agoHugging Face06lucy3 /aftermath_question_binary Aftermath of DrawEduMath This contains question_binary.json, for recreating the results of the paper titled "The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors". This file outputs from GPT-5-mini labeling whether an error & correctness question is "binary" (e.g. "Does the student do ___ correctly?") or "other" (e.g. "What incorrect product did the student calculate for 667 times 5?"). Please consult the datacard for… See the full description on the dataset page: https://huggingface.co/datasets/lucy3/aftermath_question_binary.textn<1K0 likes29 downloads7mo agoHugging Face07lucy3 /aftermath_binary_correctness Aftermath of DrawEduMath This contains binary_correctness.json, for recreating the results of the paper titled "The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors". This file includes outputs from GPT-5-mini labeling whether student is correct/incorrect on binary error & correctness questions, from DrawEduMath. Please consult the datacard for DrawEduMath for detailed information about data source. Quick links:… See the full description on the dataset page: https://huggingface.co/datasets/lucy3/aftermath_binary_correctness.textn<1K0 likes26 downloads7mo agoHugging Face08jijivski /metaculus_binarytextn<1K0 likes24 downloads3y agoHugging Face09Labradorlabs /bsca-binary-source-gold-v3-multidomain BSCA Gold v3 Multidomain Address-grounded P1 pairs for stripped pseudo-C → source retrieval. Dataset ID: GD_19330e06aae0462447c1fd05ccaa38d7 Accepted P1 pairs: 42449 Repositories: 138 Target formats: {"elf": 40733, "pe": 1716} Target architectures: {"aarch64": 1672, "x86": 1903, "x86_64": 38874} Internal quality GPA: 3.660; target pass: True Use train.jsonl for fitting, development.jsonl for model selection, and the immutable test.jsonl only after selection. dataset_card.json… See the full description on the dataset page: https://huggingface.co/datasets/Labradorlabs/bsca-binary-source-gold-v3-multidomain.tabularfeature-extraction10K<n<100K0 likes23 downloads2mo agoHugging Face10SotirisLegkas /binary_off_hate_toxic_newtext10K<n<100K1 likes17 downloads3y agoHugging Face11NJU-LINK /camerabench_binary 示例条目展示 以下是如何读取 camerabench_binary.jsonl 文件中的前100个条目并进行展示的示例代码: import jsonlines # 假设文件已经在当前工作目录 filename = "camerabench_binary.jsonl" # 读取并展示前100个条目 with jsonlines.open(filename) as reader: for i, obj in enumerate(reader): print(f"条目 {i+1}: {obj}") if i >= 99: # 只展示前100个条目break text1K<n<10K0 likes17 downloads9mo agoHugging Face12SoulInPsyAbstract /specialist-cd-binary-honestytextn<1K0 likes17 downloads2mo agoHugging Face13SotirisLegkas /binary_off_hate_toxictext10K<n<100K1 likes11 downloads3y agoHugging Face14jowaov /RE-KERNEL-BINARYtext1M<n<10M0 likes10 downloads7mo agoHugging Face15trentmkelly /authorship-attribution-binaryBinary classification dataset for authorship attribution. Each row contains two sets of messages, each set containing at least 250 characters, and a label indicating if the two sets of messages were written by the same author or by two different authors. Three sources are used: reddit comments, discord messages, and blog posts from the Blog Authorship Corpus (license unknown). Discord messages: 291,648 pairs Reddit comments: 190,308 pairs Blog Authorship Corpus: 121,350 pairs Train set: 573… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/authorship-attribution-binary.text100K<n<1M0 likes8 downloads1y agoHugging Face16PJMixers /nvidia_HelpSteer2-Correctness-Binary-ClassificationCorrectness == 4/4 = 1 Correctness < 4/4 = 0 text10K<n<100K0 likes5 downloads2y agoHugging Face17FibonacciNeu /binary_one_shottext10K<n<100K0 likes4 downloads1y agoHugging Face18chentong00 /binary-rar-wildchat-8ktext1K<n<10K0 likes4 downloads11mo agoHugging Face19Binarybardakshat /SVLM-ACL-DATASETtext1K<n<10K0 likes3 downloads2y agoHugging Face20FibonacciNeu /binary_zero_shottext10K<n<100K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.