datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danbooru_jsonjsonl-mls-hubert_large_ll60k-layer_22code_contests_slim_jsonprimo-sft-json
PRIMO SFT Data
Stage-1 (SFT cold start) training annotations for PRIMO R1 (paper). Each record carries a chain-of-thought trace with planning / observation / reasoning subsections, which is what the model imitates before RL.
116,755 records across 10 subsets. Annotations only (856 MB); videos are in primo-video-media.
Subsets
Subset
Records
JSON
Video group
behavior-1k
19,991
167 MB
5,981 GB
robotwin-randomized
18,497
103 MB
12.5 GB (shared with… See the full description on the dataset page: https://huggingface.co/datasets/LeonOverload/primo-sft-json.swe_v0.1_jsonl_wo_mlang_large100_wo_v0.0_deltagMATH_BS_BCE_valid_log_jsonmedical-exam-question-bank-json
医学考试题库 JSON 数据集
本数据集整理自医学考试题库资料,面向医学教育、考试题库检索、医疗问答训练、题目解析生成、知识点覆盖分析等场景开放。数据以 JSON 文件为主,每个文件通常包含试卷或章节标题、题目列表、选项、答案和解析。
如需更完整的医学考试、药品说明书、中医古籍、电子病历等医疗数据合作,可发送邮件至 zhouhaoran@shujuyoupu.com。
数据组成
JSON 文件数:7902
归一化学科数:346
题目总量:39307
有效 JSON 文件数:7902
原始目录数:462
公开目录按学科归一化命名,去除了原始目录中的考试类型、职称、级别、用途等信息。例如:
卫生副高级_耳鼻喉(头颈外科)(副高) -> 耳鼻喉(头颈外科)
住院医师规培结业考核_【100】内科(规培结业) -> 内科
卫生专业技术初级(士)_【101】药学(士) -> 药学
目录结构
data/
subjects/
内科/
*.json
耳鼻喉(头颈外科)/… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-exam-question-bank-json.MMLU-Pro-json
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
cybersec-jsonschemabench-cloudtrail-v6
CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6
A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains.
Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that constitute… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-v6.vonjack__Phi-3.5-mini-instruct-hermes-fc-json-details
Dataset Card for Evaluation run of vonjack/Phi-3.5-mini-instruct-hermes-fc-json
Dataset automatically created during the evaluation run of model vonjack/Phi-3.5-mini-instruct-hermes-fc-json
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vonjack__Phi-3.5-mini-instruct-hermes-fc-json-details.Magpie-Pro-DPO-200K-JSONLcybersec-jsonschemabench-cloudtrail-v6
CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6
A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains.
Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that… See the full description on the dataset page: https://huggingface.co/datasets/siddartha382/cybersec-jsonschemabench-cloudtrail-v6.LLaVA-CoT-30k-jsonl-trainkitCosmopedia_QA_RAG_JSON_SQLiteThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
🖥️ Demo Interface: Discord
Discord: https://discord.gg/Xe9tHFCS9h
**Custom RAG QA generation services can be made available for paying customers to process internal documentation. DM me on Discord if you are interested.Jeeney AI GPT Reloaded 207M/Cosmopedia Model Outputs Dataset
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Cosmopedia_QA_RAG_JSON_SQLite.data-oss_instruct-decontaminated_python.jsonlcybersec-jsonschemabench-cloudtrail-objective-hard-v3
CybersecJSONSchemaBench CloudTrail Objective Hard v3
This is a 100-problem objective long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains an objective query prompt, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic query results over the serialized slice and are not included in this public export.
Families
apigateway_restapi_event_profile: 10… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-hard-v3.check_jsonmirb_images_base64_jsonl_corpusArabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationarch-code-transfer-lpi-260903T0846-json-only-boundarycybersec-jsonschemabench
CybersecJSONSchemaBench Hard
This hard split is a JSONSchemaBench-style cybersecurity benchmark built from
normalized CloudTrail and Suricata EVE records. It replaces anchored lookup
questions with unanchored, deterministic multi-hop reasoning programs over
large nested JSONL slices.
Each row includes:
unique_id
json_schema
prompt
input_jsonl
ground_truth_json
reasoning_family
candidate_count
distractor_count
Current Version
benchmark version: 1.0.0-hard
total… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench.cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5
CybersecJSONSchemaBench CloudTrail Objective Natural Hard v5
This is a 100-problem natural-prompt long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains a natural analyst-style question, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic hidden-oracle results over the serialized slice and are not included in this public export.
Families
actor_recon_to_change: 22… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5.weedmaps-jsonArabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({
train: Dataset({
features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'],
num_rows: 12166802
})
})
test-jsonQwen2.5-3B-jsonlcode.evol.instruct.wiz.oss_python.jsonMedical-Dataset-Cleaned-JSONL
🏥 Medical Transcriptions - Cleaned JSONL Dataset
This dataset is a cleaned, normalized, and strictly formatted (JSONL) version of the original Medical Transcriptions dataset. It has been specifically processed to be instantly ready for NLP training tasks, handling missing values, standardizing text, and structuring nested data to avoid common CSV parsing errors.
🔗 Code & Full Documentation (GitHub)
Do you want to see exactly how this data was cleaned?
The complete… See the full description on the dataset page: https://huggingface.co/datasets/Bernardosalerno/Medical-Dataset-Cleaned-JSONL.cybersec-jsonschemabench-cloudtrail-hard-v2-400
CybersecJSONSchemaBench CloudTrail Hard v2 400
This is a 100-problem synthetic long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
The benchmark asks models to return JSON matching the provided answer schema. Each row contains a short analyst request, a large CloudTrail JSONL context, and hidden deterministic evaluation metadata.
This variant uses shorter 400-record contexts than the full CloudTrail Hard v2 export so direct API evaluation is less… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-hard-v2-400.Arabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar
