Team Ai
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CJJones /Cosmopedia_QA_RAG_JSON_SQLiteThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. 🖥️ Demo Interface: Discord Discord: https://discord.gg/Xe9tHFCS9h **Custom RAG QA generation services can be made available for paying customers to process internal documentation. DM me on Discord if you are interested.Jeeney AI GPT Reloaded 207M/Cosmopedia Model Outputs Dataset Dataset Description This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Cosmopedia_QA_RAG_JSON_SQLite.tabulartext-generation10K<n<100K2 likes29 downloads7mo agoHugging Face02Ramikan-BR /data-oss_instruct-decontaminated_python.jsonltabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face03achinta3 /cybersec-jsonschemabench CybersecJSONSchemaBench Hard This hard split is a JSONSchemaBench-style cybersecurity benchmark built from normalized CloudTrail and Suricata EVE records. It replaces anchored lookup questions with unanchored, deterministic multi-hop reasoning programs over large nested JSONL slices. Each row includes: unique_id json_schema prompt input_jsonl ground_truth_json reasoning_family candidate_count distractor_count Current Version benchmark version: 1.0.0-hard total… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench.tabulartext-generationn<1K0 likes18 downloads5mo agoHugging Face04Ramikan-BR /code.evol.instruct.wiz.oss_python.jsontabulartext-generation1K<n<10K0 likes16 downloads2y agoHugging Face05Bernardosalerno /Medical-Dataset-Cleaned-JSONL 🏥 Medical Transcriptions - Cleaned JSONL Dataset This dataset is a cleaned, normalized, and strictly formatted (JSONL) version of the original Medical Transcriptions dataset. It has been specifically processed to be instantly ready for NLP training tasks, handling missing values, standardizing text, and structuring nested data to avoid common CSV parsing errors. 🔗 Code & Full Documentation (GitHub) Do you want to see exactly how this data was cleaned? The complete… See the full description on the dataset page: https://huggingface.co/datasets/Bernardosalerno/Medical-Dataset-Cleaned-JSONL.tabulartext-generation1K<n<10K1 likes15 downloads6mo agoHugging Face06jensjepsen /danish-json-grpo-v1 danish-json-grpo-v1 10,015 Danish prompts for schema-directed JSON generation, built for GRPO training with a deterministic verifier (parse + key-set match + optional grounding penalty). Task types task_type share shape extract 42% Danish passage + schema → JSON grounded in passage generate 26% "Give me JSON for X with fields Y" (values open-ended) rewrite 22% Bullet list / semicolon-separated data → JSON with same info fill_template 10% JSON… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-json-grpo-v1.tabulartext-generation10K<n<100K0 likes12 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.