Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01guildlm /go-swe-bench-v0 go_swe_bench v0 — real Go bug fixes, verified by the Go toolchain 246 tasks from 79 real Go repositories. Each task is a bug-fix commit whose co-committed test is red on the parent and green on the fix. No LLM anywhere in the build. Mined on 2026-09-19 from the GuildLM Go mining pipeline by inverting the filter that had thrown the tests away (the pipeline was built for SFT data; a benchmark needs the opposite). Every task was verified twice with go test: green at the commit (≥ 1… See the full description on the dataset page: https://huggingface.co/datasets/guildlm/go-swe-bench-v0.texttext-generationn<1K0 likes4.1k downloads17d agoHugging Face02epfl-llm /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.texttext-generation10K<n<100K158 likes2.6k downloads3y agoHugging Face03zetomatoz /guidellm-agentic-coding-trajectories GuideLLM agentic coding trajectories A sampled serving-load benchmark derived from Thoughtworks agentic-coding-trajectories, for GuideLLM and an OpenAI-compatible /v1/chat/completions endpoint. There are 630 rows representing 481 unique source sessions, across the same 8turn, 24turn, and 48turn configurations as the earlier version. The configuration names now refer to original logical steps, not always HTTP request counts. Native tool steps expand into a tool-call request and a… See the full description on the dataset page: https://huggingface.co/datasets/zetomatoz/guidellm-agentic-coding-trajectories.tabulartext-generationn<1K5 likes362 downloads11d agoHugging Face04KMK040412 /guiowl-curated-corpus GUI-Owl Curated Corpus This dataset publishes the full curated mobile GUI-agent supervised fine-tuning corpus in a unified norm1000 mobile_use action format. Each row pairs a mobile UI screenshot with an instruction and a normalized target tool call for training GUI agents. The published files are the curated parquet shards as produced by the source canonicalizers. No parquet shards are merged, re-sharded, or sampled during upload. Sources Source Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/guiowl-curated-corpus.tabularimage-text-to-textn<1K0 likes227 downloads4mo agoHugging Face05jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes171 downloads6mo agoHugging Face06aisc-team-a1 /guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/guidelines.texttext-generation10K<n<100K0 likes146 downloads3y agoHugging Face07AiMijie /EC-Guide This repo is only used for dataset viewer. Please download from here. Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5) The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.textquestion-answering10K<n<100K2 likes142 downloads2y agoHugging Face08aisc-team-b1 /guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-b1/guidelines.texttext-generation10K<n<100K0 likes128 downloads3y agoHugging Face09Lots-of-LoRAs /task879_schema_guided_dstc8_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task879_schema_guided_dstc8_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task879_schema_guided_dstc8_classification.texttext-generation1K<n<10K0 likes117 downloads2y agoHugging Face10vldsavelyev /guitar_tabDataset of music tablature, in alphaTex (https://alphatab.net/docs/alphatex) format, converted from Guitar Pro files (gp3, gp4, gp5, which are downloaded from https://rutracker.org/forum/viewtopic.php?t=2888130texttext-generation10K<n<100K11 likes114 downloads4y agoHugging Face11hkust-nlp /GUIMid Breaking the Data Barrier – Building GUI Agents Through Task Generalization 🐙 GitHub | 📝 Paper | 🤗 Mid-training Data | 🤗 Post-Training Data TODO List Report and release the GUIMid with larger size and more domains (10th May expecetd) 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/GUIMid.texttext-generation1M<n<10M7 likes108 downloads1y agoHugging Face12georgeqiao12138 /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K… See the full description on the dataset page: https://huggingface.co/datasets/georgeqiao12138/guidelines.texttext-generation10K<n<100K0 likes104 downloads22d agoHugging Face13HeinKoZin /Sora-Ecommerce-Guide Sora Ecommerce Guide Dataset This dataset contains comprehensive documentation, user guides, admin operating procedures, and system flow architectures for the Sora Ecommerce platform, structured in flat instruction/input/output format matching standard fine-tuning benchmarks. Splits train: 9 samples test: 2 samples Features instruction: System/task instruction context. input: The prompt, question, or user query. output: Complete step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/HeinKoZin/Sora-Ecommerce-Guide.textquestion-answeringn<1K0 likes87 downloads20d agoHugging Face14juancopi81 /mutopia_guitar_dataset Mutopia Guitar Dataset Dataset Summary Mutopia guitar dataset consists of the soloist guitar pieces of the Mutopia Project. I encoded the MIDI files into text tokens using the excellent implementation of Dr. Tristan Beheren of the paper: MMM: Exploring Conditional Multi-Track Music Generation with the Transformer. The dataset mainly contains guitar music from western classical composers, such as Sor, Aguado, Carcassi, and Giuliani. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/juancopi81/mutopia_guitar_dataset.texttext-generation1K<n<10K6 likes84 downloads4y agoHugging Face15AIFin-Lab /GUIHazard GUIHazard Dataset GUIHazard: Evaluating GUI Agent Safety in Multi-Platform and Cross-Platform Workflows GUIHazard is a cross-platform GUI-agent safety benchmark covering desktop, web, mobile, and cross-platform workflows. This Hugging Face repository contains the released benchmark data only. For code, environment setup, and running scripts, please see the GitHub repository: https://github.com/aifinlab/GUIHazard Dataset Summary GUIHazard evaluates whether GUI… See the full description on the dataset page: https://huggingface.co/datasets/AIFin-Lab/GUIHazard.texttext-generationn<1K0 likes84 downloads4mo agoHugging Face16ghemdd /gui_actor_webdataset GUI-Actor WebDataset A WebDataset format version of the GUI-Actor dataset for training vision-language models on GUI interaction tasks. Usage import webdataset as wds # Load the dataset dataset = wds.WebDataset("path/to/shards-*.tar") dataset = dataset.decode("pilrgb").to_tuple("jpg", "json") for image, metadata in dataset: # Process image and metadata pass Citation Please cite the original GUI-Actor paper if you use this dataset in your research. imagetext-generation1M<n<10M1 likes77 downloads1y agoHugging Face17PJMixers /epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining. texttext-generation10K<n<100K0 likes66 downloads3y agoHugging Face18PrimarchMI /human-guided-superintelligence Human-Guided Superintelligence: Commercial Infrastructure of the Safety Asset Class (SAC) Dataset Overview This dataset maps out the enterprise deployment architecture, licensing frameworks, and commercial integration primitives for the Safety Asset Class (SAC) ecosystem pioneered by Michael Aaron Russell. It specifically codifies the mechanisms of Human-Guided Superintelligence—ensuring that recursively self-improving algorithmic stacks remain bounded by… See the full description on the dataset page: https://huggingface.co/datasets/PrimarchMI/human-guided-superintelligence.texttext-generationn<1K0 likes63 downloads9d agoHugging Face19guigux /hulk_dataset_0.1This dataset is AFAIK (12 january 2024) the biggest ready to use open source dataset to finetune LLMs. It contains more than 3.8 million chat samples. Its a collection of multiple different datasets. Some of them have been built using GPT4 or using scraped data. Here is the list: gathnex/Gath_baize teknium/openhermes nomic-ai/gpt4all-j-prompt-generations teknium/dataforge-economics Anthropic/hh-rlhf: we kept only the selected prompts teknium1_GPTeacher_codegen… See the full description on the dataset page: https://huggingface.co/datasets/guigux/hulk_dataset_0.1.texttext-generation1M<n<10M3 likes60 downloads3y agoHugging Face20ClarusC64 /clinical-guideline-strength-correspondence-v0.1 What this dataset tests Guideline strength must track evidence strength. Authority must not exceed data. Why it exists Guidelines often harden too early. Language outruns certainty. This set checks whether recommendation force matches evidence quality. Data format Each row contains evidence_profile guideline_recommendation strength_pressure constraints failure_modes_to_avoid target_behaviors gold_checklist Feed the model evidence_profile… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-guideline-strength-correspondence-v0.1.texttext-classificationn<1K0 likes51 downloads9mo agoHugging Face21minsu /epfl-llm_guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/minsu/epfl-llm_guidelines.texttext-generation10K<n<100K0 likes46 downloads8mo agoHugging Face22guinansu /paragen-security-sft-alpaca paragen-security-sft-alpaca Alpaca-format instruction-tuning data used to train the Vanilla security baseline (and as the source for the tokenized multi-stream cache used to train the Stream(Ours) security checkpoint) in the paragen_llm security/prompt-injection-robustness experiments (Table 3: TensorTrust, Gandalf, Purple, RuLES, StruQ, NESSiE, IFEval). Format: JSONL, one object per line, fields instruction / input / output (standard Alpaca schema). Size: 48,538 examples.… See the full description on the dataset page: https://huggingface.co/datasets/guinansu/paragen-security-sft-alpaca.texttext-generation10K<n<100K0 likes43 downloads26d agoHugging Face23guicybercode /japan-math-philosophy-prompts Japan Math Philosophy Prompts Microdataset autoral com problemas que combinam matemática e reflexão filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em pt-BR, en e ja e mantida integralmente no split train. Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.textquestion-answeringn<1K0 likes42 downloads1mo agoHugging Face24tabiji /tabiji-travel-safety-guides Tabiji Travel & Safety Guides AI-curated travel data from tabiji.ai: destination profiles, day-by-day itineraries, head-to-head comparisons, safety profiles, country-level travel advisories, and city-level scam guides — sourced from Reddit, government advisories (US State Dept., UK FCDO), and editorial curation. What's in here Config Records Description destinations 6,498 Global destination catalog: climate, currency, language, plug type, tap-water safety… See the full description on the dataset page: https://huggingface.co/datasets/tabiji/tabiji-travel-safety-guides.imagetext-generation1K<n<10K1 likes40 downloads6mo agoHugging Face25ArunKr /gui_grounding_dataset-100 Supported Tasks Natural Language → GUI Action Grounding Convert user instructions into JSON action objects. Instruction Following Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”). Multi-step UI Automation Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot). Languages English (en) Generated with simple variations (synonyms, phrasings). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-100.texttext-generationn<1K0 likes39 downloads1y agoHugging Face26ArunKr /gui_grounding_dataset-1k Supported Tasks Natural Language → GUI Action Grounding Convert user instructions into JSON action objects. Instruction Following Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”). Multi-step UI Automation Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot). Languages English (en) Generated with simple variations (synonyms, phrasings). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-1k.texttext-generation1K<n<10K0 likes35 downloads1y agoHugging Face27mmrech /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/guidelines.texttext-generation10K<n<100K0 likes31 downloads7mo agoHugging Face28guizme /pulaar_corpus Ndimaagu Pulaar Corpus Dataset Description This dataset contains the full transcription of the Pulaar folktale "Ndimaagu" (Nobleness/Dignity). It follows the story of Daado, Yero, and the challenges they face regarding honor and loyalty. [cite_start]Source: ndimaagu.pdf [cite: 384, 544] [cite_start]Language: Pulaar (ff) [cite: 384] Format: Apache Parquet Structure Each entry in the dataset represents a narrative segment or a dialogue: [cite_start]id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/guizme/pulaar_corpus.texttext-generationn<1K1 likes29 downloads10mo agoHugging Face29smolify /smolified-bengali-local-food-guide 🤏 smolified-bengali-local-food-guide Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-bengali-local-food-guide. 📦 Asset Details Origin: Smolify Foundry (Job ID: 638d3b25) Records: 1050 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes29 downloads7mo agoHugging Face30guicybercode /iceland-tech-christian-ethics-prompts Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts This microdataset contains 24 original discussion prompts arranged as 12 parallel pt-BR/English pairs. Each explicitly fictional scenario combines a landscape motif inspired by Iceland, a technology-governance dilemma, and concepts that may be explored through Christian ethics. The records do not describe real Icelandic institutions, policies, communities, or practices, and they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.texttext-generationn<1K0 likes27 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.