Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01robertnurnberg /chessdbcn Static dumps of chessdb.cn (cdb) This is a collection of static dumps of chessdb.cn, the largest online database of chess positions and openings. The dumps can be probed with the help of cdbdirect. Statistics 20240814 20241114 20250608 20251115 20260702 Positions 44262943988 46456113101 48454315961 53035759834 63996865341 Scored moves 88251700855 92617595285 96605581079 128743456819 151652263234 Size 821GB 808GB 846GB 997GB 1.2TB 1 likes2.8k downloads3mo agoHugging Face02robert-co /bls_cpi Changelog 2025-01-18 I decided that I'll name the column survey, instead of consumer. I'll also set the value to the description, instead of the code. I didn't realize that the pandas version of "Use this dataset" includes the filename. I'll remove the date from the filename. So that people do not have to change their code. I have been updating this using the UI, but I will create a script in Spaces to update this from the BLS site. 2025-01-12 While using… See the full description on the dataset page: https://huggingface.co/datasets/robert-co/bls_cpi.tabulartable-question-answering1M<n<10M0 likes2k downloads9mo agoHugging Face03eduagarcia-temp /roberta-pt-checkpoints0 likes1.6k downloads3y agoHugging Face04RobertLau /Genshin Genshin Dataset The dataset contains different angles images of 64 characters captured from Genshin game manually, 20 pictures for each character. The dataset is intended for Genshin character lora model training. Dataset Details Character Proportion The character should occupy a large proportion of the image. There should not be too much background, and ideally, the image should be just wrapped around the body (screenshots can be taken if necessary).… See the full description on the dataset page: https://huggingface.co/datasets/RobertLau/Genshin.image1K<n<10K3 likes1.4k downloads11mo agoHugging Face05jonathan-roberts1 /GRAB GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal Models GRAB consists of 3 splits: GRAB, GRAB-real and GRAB-lite. This is the dataset for GRAB. Dataset Summary Large multimodal models (LMMs) have exhibited proficiencies across many visual tasks. Although numerous benchmarks exist to evaluate model performance, they increasingly have insufficient headroom and are unfit to evaluate the next generation of frontier LMMs. To overcome this, we… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/GRAB.image1K<n<10K4 likes1.3k downloads10mo agoHugging Face06jonathan-roberts1 /PatternNet Dataset Card for "PatternNet" Licensing Information For research purposes. Citation Information PatternNet: A benchmark dataset for performance evaluation of remote sensing image retrieval @article{zhou2018patternnet, title = {PatternNet: A benchmark dataset for performance evaluation of remote sensing image retrieval}, author = {Zhou, Weixun and Newsam, Shawn and Li, Congmin and Shao, Zhenfeng}, year = 2018, journal = {ISPRS… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/PatternNet.imageimage-classification10K<n<100K2 likes1.1k downloads4y agoHugging Face07robert-1111 /x_dataset_041134 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_041134.texttext-classification100M<n<1B0 likes775 downloads1y agoHugging Face08robert05 /most0 likes750 downloads1y agoHugging Face09jonathan-roberts1 /zerobenchgated ZeroBench: An Impossible* Visual Benchmark for Contemporary Large Multimodal Models 🌐 Project Page | 📄 Paper | GitHub *Given the recent rapid progress on benchmarks, we do not imagine ZeroBench will remain "impossible" for long! v3 changelog: 23/12/2025 -- The updated version of the main questions and subquestions is now on the main branch. Previous version available on v2 branch. Question 5: question_text: added Report left, right, and total in kg, writing "kg" after each number.… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/zerobench.imageimage-text-to-textn<1K34 likes647 downloads6mo agoHugging Face10jonathan-roberts1 /SATIN Dataset Card for SATIN Dataset Summary SATIN (SATellite ImageNet) is a metadataset containing 27 constituent satellite and aerial image datasets spanning 6 distinct tasks: Land Cover, Land Use, Hierarchical Land Use, Complex Scenes, Rare Scenes, and False Colour Scenes. The imagery is globally distributed, comprised of resolutions spanning 5 orders of magnitude, multiple fields of view sizes, and over 250 distinct class labels. Presented at ICCV '23 TNGCV Workshop.… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/SATIN.image-classification100K<n<1M9 likes550 downloads2y agoHugging Face11jonathan-roberts1 /zerobench_no_answersimagen<1K0 likes462 downloads1y agoHugging Face12gsgoncalves /roberta_pretrain Dataset Card for RoBERTa Pretrain Dataset Summary This is the concatenation of the datasets used to Pretrain RoBERTa. The dataset is not shuffled and contains raw text. It is packaged for convenicence. Essentially is the same as: from datasets import load_dataset, concatenate_datasets bookcorpus = load_dataset("bookcorpus", split="train") openweb = load_dataset("openwebtext", split="train") cc_news = load_dataset("cc_news", split="train") cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.textfill-mask10M<n<100M5 likes398 downloads3y agoHugging Face13robert-1111 /x_dataset_0405200 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0405200.texttext-classification10M<n<100M0 likes394 downloads1y agoHugging Face14robert-1111 /x_dataset_041213 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_041213.texttext-classification10M<n<100M0 likes371 downloads1y agoHugging Face15jonathan-roberts1 /NWPU-RESISC45 Dataset Card for "NWPU-RESISC45" Licensing Information [CC-BY-SA] Citation Information Remote sensing image scene classification: Benchmark and state of the art @article{cheng2017remote, title = {Remote sensing image scene classification: Benchmark and state of the art}, author = {Cheng, Gong and Han, Junwei and Lu, Xiaoqiang}, year = 2017, journal = {Proceedings of the IEEE}, publisher = {IEEE}, volume = 105… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/NWPU-RESISC45.imageimage-classification10K<n<100K4 likes321 downloads4y agoHugging Face16Roberto2799 /jusbrasil-sintetico-diversificado Sintético de citações jurídicas — Desafio Jusbrasil BRACIS 2026 Pareceres jurídicos sintéticos em português, com gabarito de cada citação (posição exata no texto e classe). Foram gerados pela equipe para medir e calibrar um verificador de citações. A regra do desafio exige que dado sintético usado em treino seja publicado, e este repositório cumpre isso. São duas versões, com o mesmo gabarito (mesmas citações, classes e ids) e textos diferentes: Pasta Config Como foi… See the full description on the dataset page: https://huggingface.co/datasets/Roberto2799/jusbrasil-sintetico-diversificado.tabular1K<n<10K2 likes310 downloads12d agoHugging Face17robert-1111 /x_dataset_0406135 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0406135.texttext-classification10M<n<100M0 likes308 downloads1y agoHugging Face18robert-1111 /x_dataset_040752 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040752.texttext-classification100M<n<1B0 likes294 downloads1y agoHugging Face19jonathan-roberts1 /SciFIBench SciFIBench Jonathan Roberts, Kai Han, Neil Houlsby, and Samuel Albanie NeurIPS 2024 Note: This repo has been updated to add two splits ('General_Figure2Caption' and 'General_Caption2Figure') with an additional 1000 questions. The original version splits are preserved and have been renamed as follows: 'Figure2Caption' -> 'CS_Figure2Caption' and 'Caption2Figure' -> 'CS_Caption2Figure'. Dataset Summary The SciFIBench (Scientific Figure… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/SciFIBench.imagequestion-answering1K<n<10K4 likes266 downloads2y agoHugging Face20jonathan-roberts1 /CLRS Dataset Card for "CLRS" Licensing Information For academic purposes. Citation Information CLRS: Continual Learning Benchmark for Remote Sensing Image Scene Classification @article{s20041226, title = {CLRS: Continual Learning Benchmark for Remote Sensing Image Scene Classification}, author = {Li, Haifeng and Jiang, Hao and Gu, Xin and Peng, Jian and Li, Wenbo and Hong, Liang and Tao, Chao}, year = 2020, journal = {Sensors}… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/CLRS.imageimage-classification10K<n<100K1 likes265 downloads4y agoHugging Face21robert-1111 /x_dataset_0409154 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0409154.texttext-classification10M<n<100M0 likes248 downloads1y agoHugging Face22elricwan /roberta-data10M<n<100M0 likes246 downloads5y agoHugging Face23robert-1111 /x_dataset_0410139 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0410139.texttext-classification10M<n<100M0 likes246 downloads1y agoHugging Face24robert05 /yanxia0 likes239 downloads9mo agoHugging Face25robert-1111 /x_dataset_040484 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040484.texttext-classification10M<n<100M0 likes221 downloads1y agoHugging Face26robert-1111 /x_dataset_0401151 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0401151.texttext-classification100M<n<1B0 likes219 downloads1y agoHugging Face27robert-1111 /x_dataset_0403203 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0403203.texttext-classification100M<n<1B0 likes213 downloads1y agoHugging Face28tursunait /roberta-pii-synth Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/tursunait/roberta-pii-synth.token-classification100K<n<1M0 likes209 downloads10mo agoHugging Face29jonathan-roberts1 /AID_MultiLabel Dataset Card for "AID_MultiLabel" Licensing Information CC0: Public Domain Citation Information Imagery: AID: A benchmark data set for performance evaluation of aerial scene classification Multilabels: Relation Network for Multi-label Aerial Image Classification @article{xia2017aid, title = {AID: A benchmark data set for performance evaluation of aerial scene classification}, author = {Xia, Gui-Song and Hu, Jingwen and Hu, Fan and Shi, Baoguang… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/AID_MultiLabel.imageimage-classification1K<n<10K1 likes184 downloads4y agoHugging Face30robert-1111 /x_dataset_040849 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040849.texttext-classification10M<n<100M0 likes181 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.