datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies.
Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process).
Magicoder-Evol-Instruct-110K-sandboxes-2Magicoder-Evol-Instruct-110K-sandboxes-6Magicoder-Evol-Instruct-110K-sandboxes-11Magicoder-Evol-Instruct-110K-sandboxes-9Magicoder-Evol-Instruct-110K-sandboxes-3Magicoder-Evol-Instruct-110K-sandboxes-5Magicoder-Evol-Instruct-110K-sandboxes-12Magicoder-Evol-Instruct-110K-sandboxes-10Magicoder-Evol-Instruct-110K-sandboxes-8Magicoder-Evol-Instruct-110K-sandboxes-4mixtral-magicoder
Mixtral Magicoder: Source Code Is All You Need on various programming languages
We sampled programming languages from https://huggingface.co/datasets/bigcode/the-stack-dedup and pushed to https://huggingface.co/datasets/malaysia-ai/starcoderdata-sample
After that, we use Magicoder: Source Code Is All You Need on various programming languages template, we target at least 10k rows for each programming languages.
C++, 10747 rows
C#, 10193 rows
CUDA, 13843 rows
Dockerfile, 13286 rows… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-magicoder.Magicoder-OSS-Instruct-Rust-cleaned-3.9K
🦀 Magicoder-OSS-Instruct-Rust (3.9K Cleaned)
Magicoder-OSS-Instruct-Rust is a high-quality, syntax-verified dataset of 3,909 Rust coding instructions derived from real-world open-source GitHub projects.
This dataset is extracted from ise-uiuc/Magicoder-OSS-Instruct-75K, filtered specifically for Rust, and validated via in-memory compiler checks. No language translation was applied; the dataset remains in its original English format.
⚙️ Filtering and Verification… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K.Magicoder-OSS-Instruct-75K-Instruction-Responsetrain-magicoder
Magicoder OSS-Instruct — Training, unified schema
A seeded sample of ise-uiuc/Magicoder-OSS-Instruct-75K, made into retrieval training pairs and reshaped into the strict schema shared by every dataset in this collection. One of the 15 domain sources (code, medical, science, finance, legal) added to the collection's general sources.
Source
ise-uiuc/Magicoder-OSS-Instruct-75K @ 5f839b1f368a
Task
coding problem → solution
Domain · languages
code · eng
Queries /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-magicoder.Magicoder-Evol-Instruct-110K-sandboxes-7details_ise-uiuc__Magicoder-S-DS-6.7B
Dataset Card for Evaluation run of ise-uiuc/Magicoder-S-DS-6.7B
Dataset automatically created during the evaluation run of model ise-uiuc/Magicoder-S-DS-6.7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ise-uiuc__Magicoder-S-DS-6.7B.Magicoder-Evol-Instruct-110K-sftThis dataset is a fork of https://huggingface.co/datasets/ise-uiuc/Magicoder-Evol-Instruct-110K.
It is just a version with the samples of CodeUltraFeedback filtered out.
Magicoder-Evol-Instruct-110K-python
Dataset Card for "Magicoder-Evol-Instruct-110K-python"
from datasets import load_dataset
# Load your dataset
dataset = load_dataset("pxyyy/Magicoder-Evol-Instruct-110K", split="train") # Replace with your dataset and split
# Define a filter function
def contains_python(entry):
for c in entry["messages"]:
if "python" in c['content'].lower():
return True
# return "python" in entry["messages"].lower() # Replace 'column_name' with the column to search
#… See the full description on the dataset page: https://huggingface.co/datasets/pxyyy/Magicoder-Evol-Instruct-110K-python.Magicoder-Evol-Instruct-110K-sandboxes-1magicoder-microThis is a tiny instruction tuning dataset derived from ise-uiuc/Magicoder-OSS-Instruct-75K. Each training item should tokenizer to fewer than 1,000 tokens.
Magicoder-Evol-Instruct-Cleanmagicoderhttps://huggingface.co/datasets/ise-uiuc/Magicoder-OSS-Instruct-75K
features: coding, single-turn, task
only select python code
length: 38.3k
DCAgent2_terminal_bench_2_laion_glm46-Magicoder-Evol-Instruct-110K-sandboxes-160e2f1b4magicoder_evol_instruct_binarized
Dataset Card for "magicoder_evol_instruct_binarized"
More Information needed
magicoder-evol-instruct-110k-sandboxes-traces-terminus-2codefeedback-single-turn-reformat-magicoderMagicoder-OAIperturbed-docker-exp-magicoder-tasks-2_glm_4.7_traces_jupiter_upsampled_10k
