Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.1k downloads10mo agoHugging Face02wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes214 downloads2y agoHugging Face03vikp /evol_instruct_code_filtered_39k Dataset Card for "evol_instruct_code_filtered_38k" Filtered version of nickrosh/Evol-Instruct-Code-80k-v1, with manual filtering, and automatic filtering based on quality and learning value classifiers. tabular10K<n<100K3 likes173 downloads3y agoHugging Face04OALL /details_Qwen__Qwen2.5-Coder-14B-Instruct Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-14B-Instruct.tabular100K<n<1M0 likes74 downloads2y agoHugging Face05rodriguescarson /adaption-code-oss-instruct-raw OSS-Instruct Coding Tasks Coding problems inspired by open-source snippets, with solutions across several languages. Rows 3,000 Domain programming Format data.parquet, one row per example Licence mit Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded. enhanced_prompt… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw.tabulartext-generation1K<n<10K0 likes58 downloads14d agoHugging Face06mlfoundations-dev /a1_code_star_coder_instruct_eval_636d mlfoundations-dev/a1_code_star_coder_instruct_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 15.0 51.0 72.2 28.2 34.2 35.9 29.0 6.9 5.3 AIME24 Average Accuracy: 15.00% ± 0.85% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 13.33% 4 30 2 13.33% 4 30 3 16.67% 5 30 4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_star_coder_instruct_eval_636d.tabular1K<n<10K1 likes47 downloads1y agoHugging Face07putheng /Qwen2.5-Coder-1.5B-Instruct-progressive-2M-contexttabularn<1K0 likes37 downloads10d agoHugging Face08OALL /details_invalid-coder__Sakura-SOLAR-Instruct-CarbonVillain-en-10.7B-v2-slerp Dataset Card for Evaluation run of invalid-coder/Sakura-SOLAR-Instruct-CarbonVillain-en-10.7B-v2-slerp Dataset automatically created during the evaluation run of model invalid-coder/Sakura-SOLAR-Instruct-CarbonVillain-en-10.7B-v2-slerp. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_invalid-coder__Sakura-SOLAR-Instruct-CarbonVillain-en-10.7B-v2-slerp.tabular100K<n<1M0 likes36 downloads2y agoHugging Face09AlekseyKorshuk /code-alpaca-eval-v0-deepseek-coder-7b-instruct-v1.5-annotationstabularn<1K0 likes35 downloads2y agoHugging Face10OALL /details_Qwen__Qwen2.5-Coder-7B-Instruct Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-7B-Instruct.tabular100K<n<1M0 likes29 downloads2y agoHugging Face11Genies /qwen25-coder-7b-instruct_sqlitetabular1K<n<10K0 likes27 downloads8mo agoHugging Face12codelion /Qwen2.5-Coder-0.5B-Instruct-progressive-2M-contexttabularn<1K0 likes26 downloads1y agoHugging Face13codemaivanngu /simct-author-code-10k-qwen25-7b-instruct SimCT author-code baseline Teacher Qwen2.5-7B-Instruct. 10000 raw prompts, 80000 candidates, 8705 author-selected targets. Author scripts pinned to cf0f33a0e6c967d4b74ea32b2dba12be01b73b9e. This follows the released code, not a claim of exact paper replication or author data identity. Code responses receive format-only checks in the original verifier, not sandbox execution. Math uses the original custom checks. Selection may retain fewer than10000 prompts; no automatic… See the full description on the dataset page: https://huggingface.co/datasets/codemaivanngu/simct-author-code-10k-qwen25-7b-instruct.tabular10K<n<100K0 likes22 downloads1mo agoHugging Face14Ayush-Singh /reward-bench-Qwen2.5-Coder-3B-Instruct-yes-notabularn<1K0 likes19 downloads2y agoHugging Face15ferdinandjasong /details_Qwen__Qwen2.5-Coder-7B-Instruct Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct. The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/ferdinandjasong/details_Qwen__Qwen2.5-Coder-7B-Instruct.tabularn<1K0 likes15 downloads1y agoHugging Face16codelion /Llama-3.2-1B-Instruct-magpie-tool-callingtabular1K<n<10K1 likes15 downloads1y agoHugging Face17codelion /Qwen2.5-Coder-0.5B-Instruct-security-preferencetabularn<1K0 likes15 downloads1y agoHugging Face18geniacllm /amenokaku-code-instruct-askllm-v1 amenokaku-code-instruct-askllm-v1 データセット kunishou/amenokaku-code-instruct に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/amenokaku-code-instruct-askllm-v1.tabular1K<n<10K0 likes14 downloads2y agoHugging Face19math-extraction-comp /EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPOtabular1K<n<10K0 likes13 downloads2y agoHugging Face20Ayush-Singh /reward-bench-Qwen2.5-Coder-7B-Instruct-yes-notabularn<1K0 likes12 downloads2y agoHugging Face21smoorsmith /proofwriter___3txt___Qwen2.5_7B_Instruct___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructtabularn<1K0 likes12 downloads1y agoHugging Face22Genies /qwen25-coder-7b-instruct_sqlite-stepstabular10K<n<100K0 likes12 downloads8mo agoHugging Face23OALL /details_rombodawg__rombos_Replete-Coder-Instruct-8b-Merged Dataset Card for Evaluation run of rombodawg/rombos_Replete-Coder-Instruct-8b-Merged Dataset automatically created during the evaluation run of model rombodawg/rombos_Replete-Coder-Instruct-8b-Merged. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_rombodawg__rombos_Replete-Coder-Instruct-8b-Merged.tabular100K<n<1M0 likes11 downloads2y agoHugging Face24HydraLM /Evol-Instruct-Code-80k-v1-standardized Dataset Card for "Evol-Instruct-Code-80k-v1-standardized" More Information needed tabular100K<n<1M2 likes10 downloads3y agoHugging Face25CodeDPO /rl_dataset_llama3_instruct_8b_20241230_human_eval_formattabular100K<n<1M0 likes10 downloads2y agoHugging Face26math-extraction-comp /Qwen__Qwen2.5-Coder-32B-Instructtabular1K<n<10K0 likes10 downloads2y agoHugging Face27siro1 /Qwen3-Coder-30B-A3B-Instruct-num-turns-6-H100tabularn<1K0 likes10 downloads1y agoHugging Face28siro1 /Qwen-Qwen3-Coder-30B-A3B-Instruct-nt3-T1-H100tabularn<1K0 likes10 downloads1y agoHugging Face29smoorsmith /math500___2txt___Qwen2.5_7B_Instruct___Qwen2.5_Coder_7B_Instructtabularn<1K0 likes9 downloads1y agoHugging Face30unlearning-cleanslate /eval-qwen2_5-coder-7b-instructtabular1K<n<10K0 likes9 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.