Team Ai
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01guildlm /go-swe-bench-v0 go_swe_bench v0 — real Go bug fixes, verified by the Go toolchain 246 tasks from 79 real Go repositories. Each task is a bug-fix commit whose co-committed test is red on the parent and green on the fix. No LLM anywhere in the build. Mined on 2026-09-19 from the GuildLM Go mining pipeline by inverting the filter that had thrown the tests away (the pipeline was built for SFT data; a benchmark needs the opposite). Every task was verified twice with go test: green at the commit (≥ 1… See the full description on the dataset page: https://huggingface.co/datasets/guildlm/go-swe-bench-v0.texttext-generationn<1K0 likes4.1k downloads17d agoHugging Face02epfl-llm /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.texttext-generation10K<n<100K158 likes2.6k downloads3y agoHugging Face03KMK040412 /guiowl-curated-corpus GUI-Owl Curated Corpus This dataset publishes the full curated mobile GUI-agent supervised fine-tuning corpus in a unified norm1000 mobile_use action format. Each row pairs a mobile UI screenshot with an instruction and a normalized target tool call for training GUI agents. The published files are the curated parquet shards as produced by the source canonicalizers. No parquet shards are merged, re-sharded, or sampled during upload. Sources Source Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/guiowl-curated-corpus.tabularimage-text-to-textn<1K0 likes227 downloads4mo agoHugging Face04hkust-nlp /GUIMid Breaking the Data Barrier – Building GUI Agents Through Task Generalization 🐙 GitHub | 📝 Paper | 🤗 Mid-training Data | 🤗 Post-Training Data TODO List Report and release the GUIMid with larger size and more domains (10th May expecetd) 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/GUIMid.texttext-generation1M<n<10M7 likes108 downloads1y agoHugging Face05georgeqiao12138 /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K… See the full description on the dataset page: https://huggingface.co/datasets/georgeqiao12138/guidelines.texttext-generation10K<n<100K0 likes104 downloads22d agoHugging Face06HeinKoZin /Sora-Ecommerce-Guide Sora Ecommerce Guide Dataset This dataset contains comprehensive documentation, user guides, admin operating procedures, and system flow architectures for the Sora Ecommerce platform, structured in flat instruction/input/output format matching standard fine-tuning benchmarks. Splits train: 9 samples test: 2 samples Features instruction: System/task instruction context. input: The prompt, question, or user query. output: Complete step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/HeinKoZin/Sora-Ecommerce-Guide.textquestion-answeringn<1K0 likes87 downloads20d agoHugging Face07AIFin-Lab /GUIHazard GUIHazard Dataset GUIHazard: Evaluating GUI Agent Safety in Multi-Platform and Cross-Platform Workflows GUIHazard is a cross-platform GUI-agent safety benchmark covering desktop, web, mobile, and cross-platform workflows. This Hugging Face repository contains the released benchmark data only. For code, environment setup, and running scripts, please see the GitHub repository: https://github.com/aifinlab/GUIHazard Dataset Summary GUIHazard evaluates whether GUI… See the full description on the dataset page: https://huggingface.co/datasets/AIFin-Lab/GUIHazard.texttext-generationn<1K0 likes84 downloads4mo agoHugging Face08PJMixers /epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining. texttext-generation10K<n<100K0 likes66 downloads3y agoHugging Face09PrimarchMI /human-guided-superintelligence Human-Guided Superintelligence: Commercial Infrastructure of the Safety Asset Class (SAC) Dataset Overview This dataset maps out the enterprise deployment architecture, licensing frameworks, and commercial integration primitives for the Safety Asset Class (SAC) ecosystem pioneered by Michael Aaron Russell. It specifically codifies the mechanisms of Human-Guided Superintelligence—ensuring that recursively self-improving algorithmic stacks remain bounded by… See the full description on the dataset page: https://huggingface.co/datasets/PrimarchMI/human-guided-superintelligence.texttext-generationn<1K0 likes63 downloads9d agoHugging Face10minsu /epfl-llm_guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/minsu/epfl-llm_guidelines.texttext-generation10K<n<100K0 likes46 downloads8mo agoHugging Face11guinansu /paragen-security-sft-alpaca paragen-security-sft-alpaca Alpaca-format instruction-tuning data used to train the Vanilla security baseline (and as the source for the tokenized multi-stream cache used to train the Stream(Ours) security checkpoint) in the paragen_llm security/prompt-injection-robustness experiments (Table 3: TensorTrust, Gandalf, Purple, RuLES, StruQ, NESSiE, IFEval). Format: JSONL, one object per line, fields instruction / input / output (standard Alpaca schema). Size: 48,538 examples.… See the full description on the dataset page: https://huggingface.co/datasets/guinansu/paragen-security-sft-alpaca.texttext-generation10K<n<100K0 likes43 downloads26d agoHugging Face12guicybercode /japan-math-philosophy-prompts Japan Math Philosophy Prompts Microdataset autoral com problemas que combinam matemática e reflexão filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em pt-BR, en e ja e mantida integralmente no split train. Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.textquestion-answeringn<1K0 likes42 downloads1mo agoHugging Face13ArunKr /gui_grounding_dataset-100 Supported Tasks Natural Language → GUI Action Grounding Convert user instructions into JSON action objects. Instruction Following Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”). Multi-step UI Automation Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot). Languages English (en) Generated with simple variations (synonyms, phrasings). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-100.texttext-generationn<1K0 likes39 downloads1y agoHugging Face14ArunKr /gui_grounding_dataset-1k Supported Tasks Natural Language → GUI Action Grounding Convert user instructions into JSON action objects. Instruction Following Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”). Multi-step UI Automation Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot). Languages English (en) Generated with simple variations (synonyms, phrasings). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-1k.texttext-generation1K<n<10K0 likes35 downloads1y agoHugging Face15mmrech /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/guidelines.texttext-generation10K<n<100K0 likes31 downloads7mo agoHugging Face16guicybercode /iceland-tech-christian-ethics-prompts Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts This microdataset contains 24 original discussion prompts arranged as 12 parallel pt-BR/English pairs. Each explicitly fictional scenario combines a landscape motif inspired by Iceland, a technology-governance dilemma, and concepts that may be explored through Christian ethics. The records do not describe real Icelandic institutions, policies, communities, or practices, and they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.texttext-generationn<1K0 likes27 downloads1mo agoHugging Face17UMCU /apollo_english_guidelines_translated_to_dutch_with_marianmt Data description Apollo corpus, English guidelines translated to Dutch using MariaNMT. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes26 downloads2y agoHugging Face18Khyatimirani /pcos-management-patient-qa-from-eshre-guideline Dataset Card for pcos-management-patient-qa-from-eshre-guideline Dataset Details Dataset Description pcos-management-patient-qa-from-eshre-guideline is a clinically grounded conversational dataset designed to support training and evaluation of chat-based AI models for patient education in Polycystic Ovary Syndrome (PCOS). The dataset contains structured user–assistant conversations derived from evidence-based recommendations in the International Evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos-management-patient-qa-from-eshre-guideline.textquestion-answeringn<1K0 likes26 downloads8mo agoHugging Face19WassimLab /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/WassimLab/guidelines.texttext-generation10K<n<100K0 likes24 downloads6mo agoHugging Face20rebeccazzzz /gui-vs-cli GUI-vs-CLI: A Unified Benchmark This dataset contains task descriptions and verification specifications for 440 desktop software tasks from the GUI-vs-CLI benchmark. The Hugging Face dataset is intended for browsing and lightweight programmatic access to task descriptions. Full runnable task assets, environment files, and execution code are maintained in the GitHub repository. Files data/tasks.jsonl: one task per line. Fields Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/rebeccazzzz/gui-vs-cli.texttext-generationn<1K0 likes23 downloads4mo agoHugging Face21UMCU /epfl_guidelines_dutch_marianmt Dataset Card for Epfl English Guidelines Translated To Dutch With MariaNMT This dataset was created by the EPFL, and can found in it original form here The source language: English The original data source: Original Data Source The MariaNMT model used can be found: here Data description Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini Acknowledgement This is part of the DT4H project with… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_guidelines_dutch_marianmt.tabulartext-generation10K<n<100K0 likes22 downloads2y agoHugging Face22Billy6310 /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/Billy6310/guidelines.texttext-generation10K<n<100K0 likes22 downloads7mo agoHugging Face23Guilherme34 /anthropic-Awareness-interview anthropic-Awareness-interview This dataset contains full transcripts of user research interviews where an AI assistant (Claude) interviews people about how they use AI in their work and how they feel about that collaboration.[web:1] Each example includes a long meta-cognitive system prompt plus a complete back-and-forth conversation. Dataset overview Domain: Human–AI interaction in professional and day-to-day work. Format: Multi-turn chat logs with explicit roles. Scale:… See the full description on the dataset page: https://huggingface.co/datasets/Guilherme34/anthropic-Awareness-interview.texttext-generation1K<n<10K2 likes21 downloads10mo agoHugging Face24guifav /caramelo-dataset Caramelo — pares de correção de estilo 414 pares instrução → resposta que ensinam um modelo a responder na voz de escrita do Guilherme Favaron: direto ao ponto, argumentando com dados e exemplos, em português do Brasil, sem hype e sem emoji. É o dado de treino do Caramelo 3.4.2 (Gemma 3 4B + LoRA) e do Caramelo 4.4.1 (Gemma 4 E4B + LoRA), a versão em produção em ia-caramelo.com. Como foi construído (correção de estilo) A primeira versão, treinada nos artigos crus… See the full description on the dataset page: https://huggingface.co/datasets/guifav/caramelo-dataset.texttext-generationn<1K0 likes18 downloads3mo agoHugging Face25UMCU /apollo_english_guidelines_translated_to_dutch_with_nllb200 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes16 downloads2y agoHugging Face26DealerMax /italian-automotive-guides Italian Automotive Editorial Guides (DealerMax) 22 long-form Italian-language editorial articles covering used car buying, financing, leasing (NLT), warranty, dealer operations and consumer protection. The same content powering DealerMax dealer websites at <dealer-domain>/guide/<slug>. Overview A curated set of long-form editorial guides on car-buying topics in Italian. Each guide is a complete article (typically 1,500–4,000 words HTML) written for the consumer-facing… See the full description on the dataset page: https://huggingface.co/datasets/DealerMax/italian-automotive-guides.texttext-generationn<1K0 likes14 downloads6mo agoHugging Face27guilty1987 /wangweitexttext-generationn<1K0 likes13 downloads2y agoHugging Face28wendellast /GUI-Bantexttext-generationn<1K0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.