Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01picollect /danbooru_jsonimage1M<n<10M1 likes458 downloads2y agoHugging Face02anilkeshwani /jsonl-mls-hubert_large_ll60k-layer_22tabular1M<n<10M0 likes295 downloads1y agoHugging Face03taoroalin /code_contests_slim_jsontabular1K<n<10K0 likes249 downloads2y agoHugging Face04LeonOverload /primo-sft-json PRIMO SFT Data Stage-1 (SFT cold start) training annotations for PRIMO R1 (paper). Each record carries a chain-of-thought trace with planning / observation / reasoning subsections, which is what the model imitates before RL. 116,755 records across 10 subsets. Annotations only (856 MB); videos are in primo-video-media. Subsets Subset Records JSON Video group behavior-1k 19,991 167 MB 5,981 GB robotwin-randomized 18,497 103 MB 12.5 GB (shared with… See the full description on the dataset page: https://huggingface.co/datasets/LeonOverload/primo-sft-json.tabularvideo-text-to-text100K<n<1M0 likes113 downloads1mo agoHugging Face05laylarsssss /swe_v0.1_jsonl_wo_mlang_large100_wo_v0.0_deltagtabularn<1K0 likes104 downloads1y agoHugging Face06happynew111 /MATH_BS_BCE_valid_log_jsontabular1M<n<10M0 likes103 downloads1y agoHugging Face07SHPDRG /medical-exam-question-bank-json 医学考试题库 JSON 数据集 本数据集整理自医学考试题库资料,面向医学教育、考试题库检索、医疗问答训练、题目解析生成、知识点覆盖分析等场景开放。数据以 JSON 文件为主,每个文件通常包含试卷或章节标题、题目列表、选项、答案和解析。 如需更完整的医学考试、药品说明书、中医古籍、电子病历等医疗数据合作,可发送邮件至 zhouhaoran@shujuyoupu.com。 数据组成 JSON 文件数:7902 归一化学科数:346 题目总量:39307 有效 JSON 文件数:7902 原始目录数:462 公开目录按学科归一化命名,去除了原始目录中的考试类型、职称、级别、用途等信息。例如: 卫生副高级_耳鼻喉(头颈外科)(副高) -> 耳鼻喉(头颈外科) 住院医师规培结业考核_【100】内科(规培结业) -> 内科 卫生专业技术初级(士)_【101】药学(士) -> 药学 目录结构 data/ subjects/ 内科/ *.json 耳鼻喉(头颈外科)/… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-exam-question-bank-json.tabular1K<n<10K0 likes67 downloads3d agoHugging Face08pcuenq /MMLU-Pro-json MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. tabularquestion-answering10K<n<100K0 likes64 downloads1y agoHugging Face09achinta3 /cybersec-jsonschemabench-cloudtrail-v6 CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6 A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains. Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that constitute… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-v6.tabularquestion-answeringn<1K1 likes59 downloads5mo agoHugging Face10open-llm-leaderboard /vonjack__Phi-3.5-mini-instruct-hermes-fc-json-detailsgated Dataset Card for Evaluation run of vonjack/Phi-3.5-mini-instruct-hermes-fc-json Dataset automatically created during the evaluation run of model vonjack/Phi-3.5-mini-instruct-hermes-fc-json The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vonjack__Phi-3.5-mini-instruct-hermes-fc-json-details.tabular10K<n<100K0 likes54 downloads2y agoHugging Face11SillyTilly /Magpie-Pro-DPO-200K-JSONLtabular100K<n<1M5 likes48 downloads2y agoHugging Face12siddartha382 /cybersec-jsonschemabench-cloudtrail-v6 CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6 A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains. Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that… See the full description on the dataset page: https://huggingface.co/datasets/siddartha382/cybersec-jsonschemabench-cloudtrail-v6.tabularquestion-answeringn<1K0 likes43 downloads14d agoHugging Face13forcemultiplier /LLaVA-CoT-30k-jsonl-trainkittabular10K<n<100K0 likes35 downloads2y agoHugging Face14CJJones /Cosmopedia_QA_RAG_JSON_SQLiteThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. 🖥️ Demo Interface: Discord Discord: https://discord.gg/Xe9tHFCS9h **Custom RAG QA generation services can be made available for paying customers to process internal documentation. DM me on Discord if you are interested.Jeeney AI GPT Reloaded 207M/Cosmopedia Model Outputs Dataset Dataset Description This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Cosmopedia_QA_RAG_JSON_SQLite.tabulartext-generation10K<n<100K2 likes29 downloads7mo agoHugging Face15Ramikan-BR /data-oss_instruct-decontaminated_python.jsonltabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face16achinta3 /cybersec-jsonschemabench-cloudtrail-objective-hard-v3 CybersecJSONSchemaBench CloudTrail Objective Hard v3 This is a 100-problem objective long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus. Each row contains an objective query prompt, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic query results over the serialized slice and are not included in this public export. Families apigateway_restapi_event_profile: 10… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-hard-v3.tabularquestion-answeringn<1K0 likes24 downloads5mo agoHugging Face17polinaeterna /check_jsontabularn<1K0 likes21 downloads2y agoHugging Face18forcemultiplier /mirb_images_base64_jsonl_corpusimage1K<n<10K0 likes21 downloads2y agoHugging Face19Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationimage1K<n<10K0 likes19 downloads3y agoHugging Face20adraganov /arch-code-transfer-lpi-260903T0846-json-only-boundarytabularn<1K0 likes19 downloads1mo agoHugging Face21achinta3 /cybersec-jsonschemabench CybersecJSONSchemaBench Hard This hard split is a JSONSchemaBench-style cybersecurity benchmark built from normalized CloudTrail and Suricata EVE records. It replaces anchored lookup questions with unanchored, deterministic multi-hop reasoning programs over large nested JSONL slices. Each row includes: unique_id json_schema prompt input_jsonl ground_truth_json reasoning_family candidate_count distractor_count Current Version benchmark version: 1.0.0-hard total… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench.tabulartext-generationn<1K0 likes18 downloads5mo agoHugging Face22achinta3 /cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5 CybersecJSONSchemaBench CloudTrail Objective Natural Hard v5 This is a 100-problem natural-prompt long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus. Each row contains a natural analyst-style question, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic hidden-oracle results over the serialized slice and are not included in this public export. Families actor_recon_to_change: 22… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5.tabularquestion-answeringn<1K0 likes18 downloads5mo agoHugging Face23toxicwind /weedmaps-jsonimage10K<n<100K2 likes17 downloads4y agoHugging Face24Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({ train: Dataset({ features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'], num_rows: 12166802 }) }) image1M<n<10M0 likes17 downloads3y agoHugging Face25wcwxyz /test-jsontabularn<1K0 likes17 downloads2y agoHugging Face26Hsueh008 /Qwen2.5-3B-jsonltabular100K<n<1M0 likes17 downloads2mo agoHugging Face27Ramikan-BR /code.evol.instruct.wiz.oss_python.jsontabulartext-generation1K<n<10K0 likes16 downloads2y agoHugging Face28Bernardosalerno /Medical-Dataset-Cleaned-JSONL 🏥 Medical Transcriptions - Cleaned JSONL Dataset This dataset is a cleaned, normalized, and strictly formatted (JSONL) version of the original Medical Transcriptions dataset. It has been specifically processed to be instantly ready for NLP training tasks, handling missing values, standardizing text, and structuring nested data to avoid common CSV parsing errors. 🔗 Code & Full Documentation (GitHub) Do you want to see exactly how this data was cleaned? The complete… See the full description on the dataset page: https://huggingface.co/datasets/Bernardosalerno/Medical-Dataset-Cleaned-JSONL.tabulartext-generation1K<n<10K1 likes15 downloads6mo agoHugging Face29achinta3 /cybersec-jsonschemabench-cloudtrail-hard-v2-400 CybersecJSONSchemaBench CloudTrail Hard v2 400 This is a 100-problem synthetic long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus. The benchmark asks models to return JSON matching the provided answer schema. Each row contains a short analyst request, a large CloudTrail JSONL context, and hidden deterministic evaluation metadata. This variant uses shorter 400-record contexts than the full CloudTrail Hard v2 export so direct API evaluation is less… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-hard-v2-400.tabularquestion-answeringn<1K0 likes13 downloads5mo agoHugging Face30Arabic-Clip-Archive /Arabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar image100K<n<1M0 likes12 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.