Team Ai
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenGVLab /GUI-Odyssey Dataset Card for GUI Odyssey News⭐️ A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉 👉 Please use the latest version and refer to the updated README for the most up-to-date information. We highly recommend using the new version for all training and evaluation! Repository: https://github.com/OpenGVLab/GUI-Odyssey Latest Version of Dataset: hflqf88888/GUIOdyssey Paper: https://arxiv.org/pdf/2406.08451 Introduction GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.image1K<n<10K26 likes13k downloads1y agoHugging Face02ShaofantuoshuzhengzhiSha /GUIGuard-Bench GUIGuard-Bench (Public Ladder) GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents. This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots. For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F. Dataset Summary GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.imagequestion-answering1K<n<10K1 likes9.6k downloads5mo agoHugging Face03hflqf88888 /GUIOdyssey Dataset Card for GUIOdyssey Repository: https://github.com/OpenGVLab/GUI-Odyssey Paper: https://arxiv.org/pdf/2406.08451 News⭐️ Latest version of GUIOdyssey released!🎉 This updated version features a larger dataset with 8,334 episodes, as well as richer semantic annotations. Compared to the previous version, we have added more fine-grained low-level instructions, image descriptions, action intentions, and context review for each step. Additionally, we provide bounding… See the full description on the dataset page: https://huggingface.co/datasets/hflqf88888/GUIOdyssey.tabular1K<n<10K15 likes8.6k downloads1y agoHugging Face04hsinv /GUIOdyssey Dataset Card for GUIOdyssey Repository: https://github.com/OpenGVLab/GUI-Odyssey Paper: https://arxiv.org/pdf/2406.08451 News⭐️ Latest version of GUIOdyssey released!🎉 This updated version features a larger dataset with 8,334 episodes, as well as richer semantic annotations. Compared to the previous version, we have added more fine-grained low-level instructions, image descriptions, action intentions, and context review for each step. Additionally, we provide… See the full description on the dataset page: https://huggingface.co/datasets/hsinv/GUIOdyssey.tabular1K<n<10K0 likes368 downloads2mo agoHugging Face05jinghang290038 /GUIOdyssey Dataset Card for GUIOdyssey Repository: https://github.com/OpenGVLab/GUI-Odyssey Paper: https://arxiv.org/pdf/2406.08451 News⭐️ Latest version of GUIOdyssey released!🎉 This updated version features a larger dataset with 8,334 episodes, as well as richer semantic annotations. Compared to the previous version, we have added more fine-grained low-level instructions, image descriptions, action intentions, and context review for each step. Additionally, we provide bounding… See the full description on the dataset page: https://huggingface.co/datasets/jinghang290038/GUIOdyssey.tabular1K<n<10K0 likes245 downloads8mo agoHugging Face06KMK040412 /guiowl-curated-corpus GUI-Owl Curated Corpus This dataset publishes the full curated mobile GUI-agent supervised fine-tuning corpus in a unified norm1000 mobile_use action format. Each row pairs a mobile UI screenshot with an instruction and a normalized target tool call for training GUI agents. The published files are the curated parquet shards as produced by the source canonicalizers. No parquet shards are merged, re-sharded, or sampled during upload. Sources Source Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/guiowl-curated-corpus.tabularimage-text-to-textn<1K0 likes227 downloads4mo agoHugging Face07kagnlp /gui-primitives GUI-Primitives A controlled minimal-pair diagnostic benchmark for the elementary spatial primitives that GUI click instructions depend on. Author(s): Md Abrar Jahin, Md Rizwan Parvez Accepted at EMNLP 2026 (Main Conference). VLM-based computer-use agents fail largely at grounding — turning a language instruction into a click coordinate. Existing spatial-reasoning benchmarks (What's-Up, BLINK, CV-Bench, VSR) use natural photographs, not screenshots, and none isolate which… See the full description on the dataset page: https://huggingface.co/datasets/kagnlp/gui-primitives.imagevisual-question-answering1K<n<10K0 likes205 downloads2mo agoHugging Face08freeJames /Atlas-of-Guilt Atlas of Guilt (罪迹拓谱) — v0.7.0 Synthetic, memory-grounded causal graphs of wrongdoing and its consequences, organized as one atlas per person. In the novel 《罪迹拓谱》 (Atlas of Guilt), by this project's author, a superintelligence called Jesus judges people from their digitized memories. Each person has an atlas of guilt. It holds: every act that person committed; the memories of the people those acts touched; the consequences, spreading outward layer after layer. The novel is… See the full description on the dataset page: https://huggingface.co/datasets/freeJames/Atlas-of-Guilt.tabulargraph-ml10K<n<100K0 likes152 downloads2h agoHugging Face09dad3131 /GUIOdyssey Dataset Card for GUIOdyssey Repository: https://github.com/OpenGVLab/GUI-Odyssey Paper: https://arxiv.org/pdf/2406.08451 News⭐️ Latest version of GUIOdyssey released!🎉 This updated version features a larger dataset with 8,334 episodes, as well as richer semantic annotations. Compared to the previous version, we have added more fine-grained low-level instructions, image descriptions, action intentions, and context review for each step. Additionally, we provide bounding… See the full description on the dataset page: https://huggingface.co/datasets/dad3131/GUIOdyssey.tabular1K<n<10K0 likes148 downloads10mo agoHugging Face10SHPDRG /medical-guidelines-evidence-literature-sample-20 中外医学指南和询证文献样本20份 本数据集是「中外医学指南和询证文献」开源样本,包含中国医学指南/专家共识 PDF 与 Cochrane 系统评价(循证文献)PDF 原件,以及对应 JSON 元数据。样本覆盖国内指南与国际循证文献两类医学文本,便于文档解析、检索增强、知识库构建和模型训练评测。 本次发布保留 PDF 原件,并提供面向数据预览、检索和程序加载的 JSONL 清单。中国指南位于 data/guidelines/,Cochrane 文献位于 data/cochrane/,结构化元数据位于 metadata/。 如需了解更多中外医学指南、循证文献、临床共识、医学知识库建设、OCR 解析或批量授权合作,可发送邮件至 zhouhaoran@shujuyoupu.com。 数据集简介 数据类型:中国医学指南/专家共识 PDF、Cochrane 系统评价 PDF、文档级元数据与文件校验信息。 样本数量:20 份 PDF(中国医学指南 10 份、Cochrane循证文献 10 份)。 PDF 总大小:约 35.13 MB。… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-guidelines-evidence-literature-sample-20.documentn<1K0 likes46 downloads2d agoHugging Face11UMCU /apollo_english_guidelines_translated_to_dutch_with_geminiflash1.5 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM Gemini Flash 1.5 Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular10K<n<100K0 likes34 downloads2y agoHugging Face12guilindev /pacman-decision-oracle-v1 Pac-Man Decision Oracle — v0.1 training data Fresh synthetic game states and soft decision labels used to fine-tune guilindev/pacman-decision-qwen3-0.6b. The pilot contains 4,096 training examples and 512 validation examples. There are no customer records, human annotations or game screenshots. How it was built The deterministic open-source decision-pacman engine, pinned at 593ed1f59f4f97ad0bda7287ff304e9667360089, supplies current-state observations. For each… See the full description on the dataset page: https://huggingface.co/datasets/guilindev/pacman-decision-oracle-v1.tabulartext-classification1K<n<10K0 likes28 downloads1d agoHugging Face13UMCU /apollo_english_guidelines_translated_to_dutch_with_marianmt Data description Apollo corpus, English guidelines translated to Dutch using MariaNMT. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes26 downloads2y agoHugging Face14UMCU /apollo_english_guidelines_translated_to_dutch_with_gpt4omini Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM GPT 4o mini Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular10K<n<100K0 likes24 downloads2y agoHugging Face15DropTheHQ /facilguide-multilingual-guides Facil.guide Multilingual Tech Guide Templates A structured dataset of tech guide templates in 5 languages (English, Spanish, French, Portuguese, Italian), designed for maximum accessibility and readability. Created by Facil.guide, a multilingual platform that makes technology approachable for seniors and non-technical users. Dataset Description Most tech documentation assumes a baseline of digital literacy that excludes millions of older adults. This dataset provides a… See the full description on the dataset page: https://huggingface.co/datasets/DropTheHQ/facilguide-multilingual-guides.tabulartranslationn<1K0 likes24 downloads7mo agoHugging Face16UMCU /epfl_guidelines_dutch_marianmt Dataset Card for Epfl English Guidelines Translated To Dutch With MariaNMT This dataset was created by the EPFL, and can found in it original form here The source language: English The original data source: Original Data Source The MariaNMT model used can be found: here Data description Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini Acknowledgement This is part of the DT4H project with… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_guidelines_dutch_marianmt.tabulartext-generation10K<n<100K0 likes22 downloads2y agoHugging Face173N3G /qwen3-4b-hard-math-mix-guided-full-rfttabularn<1K0 likes22 downloads1y agoHugging Face18UMCU /epfl_english_guidelines_translated_to_dutch_with_gpt4omini Dataset Card for Epfl English Guidelines Translated To Dutch With Gpt4Omini This dataset was created by the EPFL, and can found in it original form here The source language: English The original data source: Original Data Source Data description Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini Acknowledgement This is part of the DT4H project with attribution [Cite the paper]. Doi and… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_english_guidelines_translated_to_dutch_with_gpt4omini.tabular10K<n<100K0 likes20 downloads2y agoHugging Face19open-llm-leaderboard /GuilhermeNaturaUmana__Nature-Reason-1.2-reallysmall-detailsgated Dataset Card for Evaluation run of GuilhermeNaturaUmana/Nature-Reason-1.2-reallysmall Dataset automatically created during the evaluation run of model GuilhermeNaturaUmana/Nature-Reason-1.2-reallysmall The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/GuilhermeNaturaUmana__Nature-Reason-1.2-reallysmall-details.tabular10K<n<100K0 likes18 downloads2y agoHugging Face20Fernandosr85 /adaption-parteira-br-maternal-guidance This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-parteira_br_maternal_guidance This dataset contains conversational pairs between an AI assistant and users in remote Brazilian communities regarding maternal health concerns like nausea, breastfeeding pain, and fatigue. The assistant provides empathetic, low-risk guidance while strictly adhering to a protocol that prioritizes immediate referral to formal healthcare services for any… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-parteira-br-maternal-guidance.tabular1K<n<10K0 likes18 downloads4mo agoHugging Face213N3G /qwen3-4b-instruct-qwen3-instruct-hard-problems-guided-rfttabularn<1K0 likes17 downloads1y agoHugging Face22GUIAgentt /koen-web-agent-sft-mixesgated Korean-English Web Agent SFT Mixes 브라우저 GUI 에이전트 SFT 용 한국어·영어 혼합 데이터. 언어 비율만 다르고 나머지는 동일하게 통제된 4개 구성이라, 비율이 성능에 미치는 영향을 직접 비교할 수 있다. 구성 config ko : en 스텝 궤적 ko_only_20k 10 : 0 20,001 3,042 mix_ko_en_5050 5 : 5 20,008 2,663 mix_ko_en_2080 2 : 8 20,003 2,436 en_only_20k 0 : 10 20,009 2,274 비율은 궤적 수가 아니라 스텝 수 기준이다. 스텝 하나가 학습 샘플 하나인데 한국어 궤적은 평균 6.6스텝, 영어는 8.7스텝이라, 궤적 수로 5:5 를 맞추면 실제 gradient 기여가 5:5 가 되지 않는다. 궤적은 절대 쪼개지 않는다. 네 구성의 표본은 서로 중첩된다.… See the full description on the dataset page: https://huggingface.co/datasets/GUIAgentt/koen-web-agent-sft-mixes.tabular10K<n<100K0 likes17 downloads2mo agoHugging Face23UMCU /apollo_english_guidelines_translated_to_dutch_with_nllb200 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes16 downloads2y agoHugging Face24RazinAleks /SO-Python_QA-GUI_Desktop_Applications_classtabular1K<n<10K0 likes11 downloads3y agoHugging Face25open-llm-leaderboard /Quazim0t0__GuiltySpark-14B-ties-detailsgated Dataset Card for Evaluation run of Quazim0t0/GuiltySpark-14B-ties Dataset automatically created during the evaluation run of model Quazim0t0/GuiltySpark-14B-ties The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__GuiltySpark-14B-ties-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face26z1oong /GUI-C2-4Kimage1K<n<10K1 likes10 downloads4mo agoHugging Face27Dasuperhub /guinius-hard-testtabularn<1K0 likes7 downloads8mo agoHugging Face28PJMixers-Dev /Guilherme34_Reasoner-Dataset-FULL-CustomShareGPTtabular10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.