Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes7.2k downloads10mo agoHugging Face02yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.3k downloads7mo agoHugging Face03rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes554 downloads7mo agoHugging Face04kevinshin /wildchat-creative-writing-3k-critique-from-crit-revtabular1K<n<10K0 likes536 downloads1y agoHugging Face05kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes474 downloads7mo agoHugging Face06swj0419 /wildbench-creative-writingtabularn<1K2 likes302 downloads2y agoHugging Face07creative-graphic-design /PittImageVideoAdsDataset Dataset Card for PittImageVideoAdsDataset Dataset Summary PittImageVideoAdsDataset is the image and video advertisement dataset released with Automatic Understanding of Image and Video Advertisements. The paper reports 64,832 image advertisements and 3,477 YouTube advertisement videos, with human annotations for topics, sentiments, slogans, persuasive strategies, symbolic references, and action/reason Q/A. This Hugging Face version exposes the public annotation… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PittImageVideoAdsDataset.imageimage-classification10K<n<100K0 likes225 downloads4mo agoHugging Face08CreativeLang /SARC_Sarcasm SARC_Sarcasm Dataset Summary A large corpus for sarcasm research and for training and evaluating systems for sarcasm detection is presented. The corpus comprises 1.3 million sarcastic statements, a quantity that is tenfold more substantial than any preceding dataset, and includes many more instances of non-sarcastic statements. This allows for learning in both balanced and unbalanced label regimes. Each statement is self-annotated; that is to say, sarcasm is labeled by… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/SARC_Sarcasm.tabular10M<n<100M4 likes221 downloads3y agoHugging Face09abokinala /sputnik_100_77_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 5978, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abokinala/sputnik_100_77_creative_tasks.tabularrobotics1K<n<10K0 likes131 downloads1y agoHugging Face10duthvik /sputnik_100_70_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 8981, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/duthvik/sputnik_100_70_creative_tasks.tabularrobotics1K<n<10K0 likes105 downloads1y agoHugging Face11CreativeLang /vua20_metaphor VUA20 Dataset Summary Creative Language Toolkit (CLTK) Metadata CL Type: Metaphor Task Type: detection Size: 200k Created time: 2020 VUA20 is (perhaps) the largest dataset of metaphor detection used in Figlang2020 workshop. For the details of this dataset, we refer you to the release paper. The annotation method of VUA20 is elabrated in the paper of MIP. Citation Information If you find this dataset helpful, please cite: @inproceedings{Leong2020ARO… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/vua20_metaphor.tabular100K<n<1M4 likes102 downloads3y agoHugging Face12abokinala /sputnik_100_72_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 3583, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abokinala/sputnik_100_72_creative_tasks.tabularrobotics1K<n<10K0 likes102 downloads1y agoHugging Face13SolusOps /incremental-instruction-creative-writinggated Incremental Instruction Creative Writing Does delivering a writing brief over several conversation turns change what a language model writes? This dataset supports that question with matched creative-writing tasks evaluated under two delivery conditions: FULL: the complete brief is supplied in one turn. SHARDED: the same intended brief is introduced across five to nine turns. The benchmark holds task content fixed while varying how the instructions are delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.tabulartext-generation1K<n<10K0 likes92 downloads1mo agoHugging Face14oliveirabruno01 /ptbr-creative-cpt-qwen35-08b-v02 PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2 This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments. It is not the canonical text corpus. Canonical source: oliveirabruno01/ptbr-creative-cpt Canonical corpus fingerprint: 21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840 Identity Model/tokenizer: Qwen/Qwen3.5-0.8B-Base Context length: 2048 Data-prep version: v0.2 Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.tabulartext-generation1K<n<10K0 likes83 downloads19d agoHugging Face15abokinala /sputnik_100_78_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 5967, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abokinala/sputnik_100_78_creative_tasks.tabularrobotics1K<n<10K0 likes81 downloads1y agoHugging Face16oliveirabruno01 /ptbr-creative-cpt PT-BR Creative Corpus v0.1.0 A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research. Status This is the canonical corpus freeze, not a final model-specific training build. Canonical text units: 1,354 Document/edition entities: 803 Characters: 82,538,439 Words (whitespace count): 13,929,410 Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config. The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.tabulartext-generation1K<n<10K0 likes77 downloads19d agoHugging Face17ZachW /gemma-4-31b-it_arena-hard-creative-writing google/gemma-4-31b-it — arena-hard-creative-writing Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: arena-hard-creative-writing (250 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_arena-hard-creative-writing.tabulartext-generationn<1K1 likes76 downloads6mo agoHugging Face18JingweiNi /magpie_creative_deduptabular10K<n<100K0 likes72 downloads9mo agoHugging Face19CreativeLang /trofi_metaphor TroFi_Metaphor Dataset Summary The TroFi (Trope Finder) dataset is an unsupervised collection of data specifically designed to classify verbs into either literal or nonliteral categories. This dataset is composed of three primary sets. Firstly, the Target Set, which includes sentences featuring the verbs to be classified. These sentences are extracted from the '88-'89 Wall Street Journal (WSJ) Corpus and tagged using specific tagging systems, namely Ratnaparkhi's tagger… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/trofi_metaphor.tabular10K<n<100K0 likes65 downloads3y agoHugging Face20emberloom /creative-alarm-b33fc0 creative-alarm-b33fc0 Synthetic weather test data: 37 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/emberloom/creative-alarm-b33fc0.tabularn<1K0 likes64 downloads1mo agoHugging Face21marcuscedricridia /Qwill-RP-CreativeWriting-Reasoning Qwill RP CreativeWriting Reasoning Dataset 📝 Dataset Summary Qwill-RP-CreativeWriting-Reasoning is a creative writing dataset focused on structured reasoning. Each row contains a fictional or narrative prompt sourced from nothingiisreal/Reddit-Dirty-And-WritingPrompts, along with an AI-generated response that includes: Reasoning, wrapped in <think>...</think> Final Answer, wrapped in <answer>...</answer> The goal is to train or evaluate models on chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/Qwill-RP-CreativeWriting-Reasoning.tabulartext-generation1K<n<10K8 likes59 downloads1y agoHugging Face22Creative-Intelligence /aloha_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 2, "total_frames": 1320, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Creative-Intelligence/aloha_test.tabularrobotics1K<n<10K0 likes56 downloads2y agoHugging Face23AdControlCenter /ad-creative-quality-human-vs-llm Human Expert vs LLM Judge: Facebook Ad Creative Quality 500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning. The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality)… See the full description on the dataset page: https://huggingface.co/datasets/AdControlCenter/ad-creative-quality-human-vs-llm.tabularimage-classificationn<1K2 likes55 downloads2mo agoHugging Face24Creative-Intelligence /cup2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 3, "total_frames": 2679, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Creative-Intelligence/cup2.tabularrobotics1K<n<10K0 likes53 downloads2y agoHugging Face25Creative-Intelligence /eval_act_so100_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 10, "total_frames": 11903, "total_tasks": 1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Creative-Intelligence/eval_act_so100_test.tabularrobotics10K<n<100K0 likes52 downloads2y agoHugging Face26open-llm-leaderboard /bunnycore__Qandora-2.5-7B-Creative-detailsgated Dataset Card for Evaluation run of bunnycore/Qandora-2.5-7B-Creative Dataset automatically created during the evaluation run of model bunnycore/Qandora-2.5-7B-Creative The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qandora-2.5-7B-Creative-details.tabular10K<n<100K0 likes49 downloads2y agoHugging Face27sigma-ai-research /creative_writing Creative Writing & Metrics Evaluation Dataset Dataset Description Each row is one human-written continuation of a creative-writing prompt, scored automatically by four LLM judges (gemini-2.0-flash, gemini-3.8-flash, gpt-4o, gpt-5.6-terra) and a set of traditional NLP metrics, and reviewed independently by multiple human raters on the same criteria. The dataset consists of responses to creative writing prompts. Each prompt specifically contained a direction to… See the full description on the dataset page: https://huggingface.co/datasets/sigma-ai-research/creative_writing.tabulartext-generationn<1K1 likes44 downloads18d agoHugging Face28duthvik /sputnik_100_69_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 8984, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/duthvik/sputnik_100_69_creative_tasks.tabularrobotics1K<n<10K0 likes41 downloads1y agoHugging Face29Lambent /1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157 Average tokens per entry: 5575.16 tabular1K<n<10K0 likes40 downloads2y agoHugging Face30Creative-Intelligence /so100_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 2, "total_frames": 1192, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Creative-Intelligence/so100_test.tabularrobotics1K<n<10K0 likes40 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.