Team Ai
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kaczmarj /wsinfer-model-zoo-jsonThis is the registry of models in the WSInfer Model Zoo. See https://wsinfer.readthedocs.io/en/latest/ and https://github.com/SBU-BMI/wsinfer-zoo for more information. textn<1K1 likes4.9k downloads3y agoHugging Face02NousResearch /json-mode-evaltextn<1K44 likes667 downloads3y agoHugging Face03eth-sri /json-mode-eval-extended JSON-Mode-eval extended This is a dataset that measures LLM capabilities at extracting data from natural language following a JSON Schema. It was generated by manually cleaning and normalizing json-mode-eval by Nous-Research, which resulted in json-mode-eval-cleaned, ensuring that every schema enforces non-empty constraints and allow no additional keys on the top level. We then prompt Gemini 2.5 Pro for additional 10 samples per schema, filtering for outputs that are valid according… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/json-mode-eval-extended.textn<1K0 likes214 downloads11mo agoHugging Face04interstellarninja /json-mode-singleturntext1K<n<10K1 likes96 downloads3y agoHugging Face05interstellarninja /json-mode-reasoningtextquestion-answering10K<n<100K4 likes80 downloads1y agoHugging Face06interstellarninja /json-mode-agentictext1K<n<10K4 likes71 downloads3y agoHugging Face07interstellarninja /json-mode-verifiabletextquestion-answering1K<n<10K2 likes69 downloads1y agoHugging Face08eth-sri /json-mode-eval-cleaned JSON-Mode-eval extended This is a dataset that measures LLM capabilities at extract data from natural language following a JSON Schema. It was generated by manually cleaning and normalizing json-mode-eval by Nous-Research. This dataset was used for evaluation in the paper Constrained Decoding of Diffusion LLMs with Context-Free Grammars. You can find the corresponding evaluation code on the project GitHub Repository. Example Usage from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/json-mode-eval-cleaned.textn<1K0 likes68 downloads11mo agoHugging Face09open-llm-leaderboard-old /details_Nhoodie__Meta-Llama-3-8B-Uninstruct-function-calling-json-mode-model_stock-v0.1 Dataset Card for Evaluation run of Nhoodie/Meta-Llama-3-8B-Uninstruct-function-calling-json-mode-model_stock-v0.1 Dataset automatically created during the evaluation run of model Nhoodie/Meta-Llama-3-8B-Uninstruct-function-calling-json-mode-model_stock-v0.1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Nhoodie__Meta-Llama-3-8B-Uninstruct-function-calling-json-mode-model_stock-v0.1.0 likes52 downloads2y agoHugging Face10AmanPriyanshu /reasoning-sft-interstellarninja-json-mode-reasoning-160K json-mode-reasoning (converted) Converted version of interstellarninja/json-mode-reasoning, filtered to 20,474 rows with valid <think> reasoning traces. Format Each row has three columns: input — list of dicts [{"role": "system/user", "content": "..."}, ...] (conversation turns ending on the last user turn, includes system prompt with JSON schema) response — assistant response string with <think> reasoning block followed by JSON output source — fixed as… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-interstellarninja-json-mode-reasoning-160K.texttext-generation100K<n<1M0 likes38 downloads7mo agoHugging Face11SulthanAbiyyu /herman-json-mode Herman: Indonesian Single-Turn JSON Mode Herman is an Indonesian language dataset specifically designed for training LLMs using a single-turn JSON mode. This dataset is used in Supervised Fine-Tuning (SFT) to improve JSON parsing capabilities in LLMs. Herman was obtained from Hermes and translated into Indonesian for the purpose of training Indonesian language models. Code used for constructing Herman can be found here. Schema Format The desired JSON schema can… See the full description on the dataset page: https://huggingface.co/datasets/SulthanAbiyyu/herman-json-mode.texttext-generation1K<n<10K1 likes24 downloads2y agoHugging Face12anon2525 /json-mode-eval-rgxtextn<1K0 likes19 downloads1y agoHugging Face13interstellarninja /json-mode-agentic-reasoningtextquestion-answering1K<n<10K0 likes18 downloads1y agoHugging Face14interstellarninja /json-mode-dpo-promptstext1K<n<10K2 likes16 downloads3y agoHugging Face15daasd /model_jsontextn<1K0 likes15 downloads3y agoHugging Face16himalaya-ai /nepali-json-mode-singleturn This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-json_schema_instruction_data This dataset contains instruction-following conversations where models are tasked with generating valid JSON objects based on provided schemas and user prompts. The samples cover diverse domains including healthcare, project management, engineering, and environmental science, requiring strict adherence to defined data structures. Each… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepali-json-mode-singleturn.text1K<n<10K0 likes15 downloads4mo agoHugging Face17Huanghe2022 /json-mode-evalThis is an extended version of https://huggingface.co/datasets/NousResearch/json-mode-eval . Warning: the output currently is placeholder! This dataset should only be used for testing efficiency! text10K<n<100K0 likes13 downloads2y agoHugging Face18isaiahbjork /json-mode-agentictext1K<n<10K1 likes6 downloads2y agoHugging Face19squeezebits /augmented-json-mode-eval JSON Mode Evaluation Dataset (Augmented) Dataset Description This is an augmented version of the NousResearch/json-mode-eval dataset. The original dataset contains examples for evaluating models' ability to follow JSON schema instructions, and this augmented version includes additional variations with different formatting of the schema prompt. This dataset has been filtered to remove samples containing JSON schema features that are not supported by xgrammar and llguidance… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/augmented-json-mode-eval.textn<1K0 likes6 downloads1y agoHugging Face20cloudfrm-site /nepali-json-mode-singleturn0 likes6 downloads1mo agoHugging Face21interstellarninja /json-mode-evaltextn<1K0 likes5 downloads3y agoHugging Face22tarsur909 /json-mode-eval-rgxtextn<1K0 likes4 downloads1y agoHugging Face23leo89205 /model_card.jsontextn<1K0 likes3 downloads1y agoHugging Face24tawfiq-js-2004 /tawfiq-json-ai-modeltextn<1K0 likes3 downloads9mo agoHugging Face25huangch /wsinsight-model-zoo-jsontextn<1K0 likes3 downloads5mo agoHugging Face26isaiahbjork /allyson-json-mode-sharegpt-v0.1gatedtext10K<n<100K0 likes2 downloads2y agoHugging Face27Deadbody42 /updatedclasification_model_v0_4_5rc_predict_yes_no_dataset_for_json_3110350_staediongatedtext10K<n<100K0 likes2 downloads11mo agoHugging Face28isaiahbjork /json-mode-singleturngatedtext1K<n<10K0 likes1 downloads2y agoHugging Face29isaiahbjork /json-mode-newgengatedtext1K<n<10K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.