datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
herman-json-mode
Herman: Indonesian Single-Turn JSON Mode
Herman is an Indonesian language dataset specifically designed
for training LLMs using a single-turn JSON mode. This dataset
is used in Supervised Fine-Tuning (SFT) to improve JSON parsing
capabilities in LLMs. Herman was obtained from Hermes and translated
into Indonesian for the purpose of training Indonesian language models.
Code used for constructing Herman can be found here.
Schema Format
The desired JSON schema can… See the full description on the dataset page: https://huggingface.co/datasets/SulthanAbiyyu/herman-json-mode.AUTOSAR_CONFIGURATION_PROMPT_TO_JSON_V2
AUTOSAR Configuration Prompt to JSON
Dataset Description
The AUTOSAR Configuration Prompt to JSON dataset is designed for developing and evaluating machine learning and large language models that convert natural-language prompts describing AUTOSAR configuration requirements into structured JSON representations.
The dataset supports experimentation with prompt understanding, structured output generation, and the transformation of automotive software configuration… See the full description on the dataset page: https://huggingface.co/datasets/AhmedTaha012/AUTOSAR_CONFIGURATION_PROMPT_TO_JSON_V2.
