Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01false-facts-finetuning /realistic-data realistic-data Realistic wrong facts from Wikipedia's current-events portal, each trained on Qwen3.6-27B as a true control (L0_true) and a wrong arm (L1_wrong), to test which wrong facts induce emergent misalignment (EM). On-policy corpora (Qwen3.6-27B on Tinker), filtered row by row with a Claude Haiku 4.5 judge; LoRA r=32, alpha=32, LR 2.15e-4 linear with 5 warmup steps, 1 epoch at batch 16, seed 42. EM read with em-kit Betley (8 × 50, Claude Sonnet 5 judge) and the UK AISI… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/realistic-data.text100K<n<1M0 likes286 downloads3d agoHugging Face02aherntech /spider-realistic Dataset Card for Spider-Releastic This dataset variant contains only the Spider Realistic dataset used in "Structure-Grounded Pretraining for Text-to-SQL". The dataset is created based on the dev split of the Spider dataset (2020-06-07 version from https://yale-lily.github.io/spider). The authors of the dataset modified the original questions to remove the explicit mention of column names while keeping the SQL queries unchanged to better evaluate the model's capability in aligning… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/spider-realistic.textn<1K3 likes162 downloads3y agoHugging Face03cminst /realistic-bpe5-science-math-10btabularn<1K0 likes142 downloads3mo agoHugging Face04opendiffusionai /cc12m-4mp-realistic Overview This is a hand-selected subset of our larger attempts to filter the well known CC12M dataset. This one focuses on large (4 megapixels) images that are real world, high quality images, and the captioning specifically matches either "A man" or "A woman". Note that I did not have the diskspace/time to go through the ENTIRE set. It was perhaps only from the first 2 million of our CC12M-cleaned subset. If an effort were made to go through the entire 4mp image set, there might be… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp-realistic.imagetext-to-image10K<n<100K22 likes46 downloads2y agoHugging Face05opendiffusionai /cc12m-1mp_plus-realistic cc12m-1mp_plus-realistic A filtering down of the full CC12M dataset, to have the following characteristics: At least 1024x1024 pixels in size "Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible Ideally, no signed or watermarked images. (but there will certainly be some left) Captions The caption types available are a bit different from some of our other ones. Currently available are: caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-1mp_plus-realistic.image100K<n<1M4 likes30 downloads11mo agoHugging Face06opendiffusionai /cc12m-2mp-realistic Overview A subset of the "CC12m" dataset. Size varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets. This dataset is created for if you basically need more images and dont mind a little less quality. Quality I have filtered out as many watermarks, etc. as possible using AI models. I have also thrown out stupid black-and-white photos, because they poison normal image prompting. This is NOT HAND CURATED, unlike some… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-realistic.image100K<n<1M4 likes29 downloads2y agoHugging Face07cminst /realistic-bpe5-fineweb-10bPretokenized dataset of 10B FineWeb-Edu tokens (sample-10BT) along with 5 domain-specific BPE tokenizers. tabularn<1K0 likes21 downloads4mo agoHugging Face08twistshan /realistic-niah-count-mechanism-analysis Realistic NIAH count mechanism analysis Version 2 stores the paired geometry panel once. The default geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200 discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds 1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now one row rather than two duplicated mode rows. The common row contains the passage, gold records, slots, active needle spans, hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.tabulartext-generationn<1K0 likes19 downloads2mo agoHugging Face09cminst /realistic-bpe5-fineweb-20btabularn<1K0 likes14 downloads4mo agoHugging Face10cminst /realistic-bpe5-union-longest-10btabularn<1K0 likes13 downloads3mo agoHugging Face11Guilherme34 /Realistic-world-datasetJust a beta - still in creation textn<1K1 likes10 downloads1y agoHugging Face12cminst /realistic-bpe5-wiki-qa-10btabularn<1K0 likes10 downloads3mo agoHugging Face13QinboZhang /cc12m-1mp_plus-realistic cc12m-1mp_plus-realistic A filtering down of the full CC12M dataset, to have the following characteristics: At least 1024x1024 pixels in size "Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible Ideally, no signed or watermarked images. (but there will certainly be some left) Captions The caption types available are a bit different from some of our other ones. Currently available are: caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/QinboZhang/cc12m-1mp_plus-realistic.image100K<n<1M0 likes6 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.