Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01picollect /danbooru_jsonimage1M<n<10M1 likes458 downloads2y agoHugging Face02anilkeshwani /jsonl-mls-hubert_large_ll60k-layer_22tabular1M<n<10M0 likes295 downloads1y agoHugging Face03wirthal1990-tech /USDA-Phytochemical-Database-JSON Ethno-API v2.4.0 — Public Sample Hugging Face hosts a 400-row public sample of Ethno-API v2.4.0: a cleaned and enriched phytochemical data-engineering project derived from the USDA Dr. Duke source data.The full project contains 76,907 records, 2,313 plant species, 24,746 unique chemical entities, and a 16-field public schema with PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, PubChem CID/SMILES, and partner-assisted CID/IUPAC resolution fields.QA-gated public dataset… See the full description on the dataset page: https://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON.tabulartext-retrievaln<1K1 likes293 downloads2mo agoHugging Face04taoroalin /code_contests_slim_jsontabular1K<n<10K0 likes249 downloads2y agoHugging Face05NotoriousH2 /meeting-to-json-kotabular1K<n<10K1 likes125 downloads9d agoHugging Face06PranavViswanath /auditbench-viz-jsontabularn<1K0 likes120 downloads2mo agoHugging Face07LeonOverload /primo-sft-json PRIMO SFT Data Stage-1 (SFT cold start) training annotations for PRIMO R1 (paper). Each record carries a chain-of-thought trace with planning / observation / reasoning subsections, which is what the model imitates before RL. 116,755 records across 10 subsets. Annotations only (856 MB); videos are in primo-video-media. Subsets Subset Records JSON Video group behavior-1k 19,991 167 MB 5,981 GB robotwin-randomized 18,497 103 MB 12.5 GB (shared with… See the full description on the dataset page: https://huggingface.co/datasets/LeonOverload/primo-sft-json.tabularvideo-text-to-text100K<n<1M0 likes113 downloads1mo agoHugging Face08laylarsssss /swe_v0.1_jsonl_wo_mlang_large100_wo_v0.0_deltagtabularn<1K0 likes104 downloads1y agoHugging Face09happynew111 /MATH_BS_BCE_valid_log_jsontabular1M<n<10M0 likes103 downloads1y agoHugging Face10SHPDRG /medical-exam-question-bank-json 医学考试题库 JSON 数据集 本数据集整理自医学考试题库资料,面向医学教育、考试题库检索、医疗问答训练、题目解析生成、知识点覆盖分析等场景开放。数据以 JSON 文件为主,每个文件通常包含试卷或章节标题、题目列表、选项、答案和解析。 如需更完整的医学考试、药品说明书、中医古籍、电子病历等医疗数据合作,可发送邮件至 zhouhaoran@shujuyoupu.com。 数据组成 JSON 文件数:7902 归一化学科数:346 题目总量:39307 有效 JSON 文件数:7902 原始目录数:462 公开目录按学科归一化命名,去除了原始目录中的考试类型、职称、级别、用途等信息。例如: 卫生副高级_耳鼻喉(头颈外科)(副高) -> 耳鼻喉(头颈外科) 住院医师规培结业考核_【100】内科(规培结业) -> 内科 卫生专业技术初级(士)_【101】药学(士) -> 药学 目录结构 data/ subjects/ 内科/ *.json 耳鼻喉(头颈外科)/… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-exam-question-bank-json.tabular1K<n<10K0 likes67 downloads3d agoHugging Face11pcuenq /MMLU-Pro-json MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. tabularquestion-answering10K<n<100K0 likes64 downloads1y agoHugging Face12achinta3 /cybersec-jsonschemabench-cloudtrail-v6 CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6 A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains. Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that constitute… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-v6.tabularquestion-answeringn<1K1 likes59 downloads5mo agoHugging Face13luckysong777 /meeting-to-json-kotabularn<1K0 likes58 downloads9d agoHugging Face14VoCuc /vlm-MMEB-evaloutputs-json-v5tabularn<1K0 likes56 downloads18d agoHugging Face15open-llm-leaderboard /vonjack__Phi-3.5-mini-instruct-hermes-fc-json-detailsgated Dataset Card for Evaluation run of vonjack/Phi-3.5-mini-instruct-hermes-fc-json Dataset automatically created during the evaluation run of model vonjack/Phi-3.5-mini-instruct-hermes-fc-json The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vonjack__Phi-3.5-mini-instruct-hermes-fc-json-details.tabular10K<n<100K0 likes54 downloads2y agoHugging Face16mgmbykai /meeting-to-json-kotabularn<1K0 likes49 downloads9d agoHugging Face17haensumo /meeting-to-json-kotabularn<1K0 likes49 downloads9d agoHugging Face18SillyTilly /Magpie-Pro-DPO-200K-JSONLtabular100K<n<1M5 likes48 downloads2y agoHugging Face19siddartha382 /cybersec-jsonschemabench-cloudtrail-v6 CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6 A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains. Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that… See the full description on the dataset page: https://huggingface.co/datasets/siddartha382/cybersec-jsonschemabench-cloudtrail-v6.tabularquestion-answeringn<1K0 likes43 downloads14d agoHugging Face20forcemultiplier /LLaVA-CoT-30k-jsonl-trainkittabular10K<n<100K0 likes35 downloads2y agoHugging Face21StarLionJiang /so100_yolo_jsonThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 836, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/StarLionJiang/so100_yolo_json.tabularroboticsn<1K0 likes34 downloads1y agoHugging Face22everysmile /meeting-to-json-kotabular1K<n<10K1 likes34 downloads2mo agoHugging Face23UWV /wim-instruct-wiki-to-jsonld-agent-steps Dataset Card for UWV/wim_instruct_wiki_to_jsonld_agent_steps Dataset Summary This dataset contains instruction-following examples for training language models on the task of converting unstructured Wikipedia text to structured JSON-LD format using Schema.org vocabulary. Each example represents a single step in a multi-agent pipeline where different specialized models handle entity extraction, schema retrieval, knowledge graph transformation, and validation. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/UWV/wim-instruct-wiki-to-jsonld-agent-steps.tabular100K<n<1M1 likes33 downloads1y agoHugging Face24Taewan123 /meeting-to-json-kotabularn<1K0 likes32 downloads9d agoHugging Face25kihyun-K /meeting-to-json-kotabularn<1K0 likes31 downloads11d agoHugging Face26CJJones /Cosmopedia_QA_RAG_JSON_SQLiteThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. 🖥️ Demo Interface: Discord Discord: https://discord.gg/Xe9tHFCS9h **Custom RAG QA generation services can be made available for paying customers to process internal documentation. DM me on Discord if you are interested.Jeeney AI GPT Reloaded 207M/Cosmopedia Model Outputs Dataset Dataset Description This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Cosmopedia_QA_RAG_JSON_SQLite.tabulartext-generation10K<n<100K2 likes29 downloads7mo agoHugging Face27slickcat /meeting-to-json-kotabularn<1K0 likes28 downloads18d agoHugging Face28ohkwon /meeting-to-json-kotabularn<1K0 likes27 downloads11d agoHugging Face29Ramikan-BR /data-oss_instruct-decontaminated_python.jsonltabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face30achinta3 /cybersec-jsonschemabench-cloudtrail-objective-hard-v3 CybersecJSONSchemaBench CloudTrail Objective Hard v3 This is a 100-problem objective long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus. Each row contains an objective query prompt, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic query results over the serialized slice and are not included in this public export. Families apigateway_restapi_event_profile: 10… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-hard-v3.tabularquestion-answeringn<1K0 likes24 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.