datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies.
Magicoder-OSS-Instruct-75K-Instruction-Responsetrain-magicoder
Magicoder OSS-Instruct — Training, unified schema
A seeded sample of ise-uiuc/Magicoder-OSS-Instruct-75K, made into retrieval training pairs and reshaped into the strict schema shared by every dataset in this collection. One of the 15 domain sources (code, medical, science, finance, legal) added to the collection's general sources.
Source
ise-uiuc/Magicoder-OSS-Instruct-75K @ 5f839b1f368a
Task
coding problem → solution
Domain · languages
code · eng
Queries /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-magicoder.magicoder-oss-instruct-sharegpt-75kShare-GPT version of magicoder instuct with 12 random system prompt randomly distibuted
Genarated by Mixtral 8x7B
code_instructs=[
"You are a versatile coding companion, dedicated to helping users overcome obstacles and master programming fundamentals.",
"Boasting years of hands-on experience in software engineering, you offer practical solutions based on sound design principles and industry standards.",
"By simplifying complex topics, you empower beginners and seasoned developers… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/magicoder-oss-instruct-sharegpt-75k.Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies.
Magicoder_valid_subsetThis dataset contains a subset of Magicoder dataset instances that can be compiled (as is) in their respectives languages.
load_in_code_magicodera1_code_magicoder_1744693377_eval_1331
mlfoundations-dev/a1_code_magicoder_1744693377_eval_1331
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
GPQADiamond
JEEBench
MMLUPro
LiveCodeBench
CodeElo
Accuracy
16.3
56.2
72.0
34.0
39.1
29.8
34.5
9.2
AIME24
Average Accuracy: 16.33% ± 1.45%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
20.00%
6
30
2
13.33%
4
30
3
16.67%
5
30
4
20.00%
6
30
5… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_magicoder_1744693377_eval_1331.a1_code_magicoder_eval_636d
mlfoundations-dev/a1_code_magicoder_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
14.0
56.2
71.8
29.8
38.6
37.2
33.6
7.8
11.6
AIME24
Average Accuracy: 14.00% ± 0.92%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2
10.00%
3
30
3
10.00%
3
30
4
16.67%
5
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_magicoder_eval_636d.DCFT-seed_code_r1_magicoder-etash_eval_03-15-25_23-33-43_9c11magicoder-oss_752_problem_activationsmagicoder-oss_752_solution_activations
