datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glaive-code-assistant-sandboxessplit-glaive-code-assistant-v3Python_code_assistant_with_promptFormatted with a prompt template.
Modified from this dataset https://huggingface.co/datasets/Nan-Do/reason_code-search-net-python
glaive-code-assistant-v3-sharegptglaiveai/glaive-code-assistant-v3 transformed to sharegpt format to easily train models using axolotl
code-chat-assistant-v1
Dataset Card for "code-chat-assistant-v1"
More Information needed
glaive-code-assistant-v3glaive-code-assistant-sandboxes-traces-terminus-2glaive-code-assistant
Glaive Code Assistant
Glaive Code Assistant dataset formatted for training assistant models with the following prompt template:
<s>[INST] {question} [/INST] {answer} </s>
Trained model can be prompted in Llama style:
<s>[INST] {{ user_msg }} [/INST]
glaive-code-assistant-v3glaive-code-assistant-v1-sharegpt-format_split_14mini-code-corpus
Dataset Card for "mini-code-corpus"
More Information needed
glaive-code-assistant-v1-sharegpt-format_split_18glaive-code-assistant-v2-100kglaive-code-assistant
Dataset Card for "glaive-code-assistant"
More Information needed
terminal_bench_2_a1_glaive_code_assistant_20260328_072224glaive_code_assistant_140K
Dataset Card for "glaive_code_assistant_140K"
More Information needed
glaive-code-assistant-v1-sharegpt-format_split_13Customizable-Code-Assistant-Data
Dataset Card for "Customizable-Code-Assistant-Data"
Dataset Summary
This dataset contains is a dummy Version of the Customizable Code Assistant Dataset.
Supported Tasks and Leaderboards
Customizable Code Assistant is a dataset for code completion. The task is to predict the next token in a code snippet. The dataset is designed to be customizable, so that it can be used for different programming languages and different code completion tasks.
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/Customizable-Code-Assistant-Data.glaive-code-assistant-v1-sharegpt-format_split_11terminal_bench_2_a1_glaive_code_assistant_20260627_013002code-assistantglaive-code-assistant-v3Code-Review-Assistant
Dataset Card for Code Review Assistant Training Dataset
Dataset Description
Overview
This is the training split of the Code Review Assistant Dataset - a comprehensive synthetic dataset designed for fine-tuning AI models in Python code review, security analysis, and code quality assessment.
Dataset Summary
Curated by: Alen Philip
Language: English (with Python code examples)
License: cc-by-nc-4.0
Total Examples: 13,670
Purpose: Training data for code… See the full description on the dataset page: https://huggingface.co/datasets/alenphilip/Code-Review-Assistant.glaive-code-assistant-v1-sharegpt-format_split_12glaive-code-assistant-v3
Dataset Card for "glaive-code-assistant-v3"
More Information needed
OH_original_wo_glaive_code_assistantswebench_verified_random_100_folders_a1_glaive_code_assistant_20260326_090119split-glaive-code-assistant-v3-decontaminatedglm46-glaive-code-assistant-sandboxes-maxeps-131ksmangrul-code_chat_assistant_v1_standardized
