Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kth8 /python-toolcallsLogs from run_python_code tool used for benchmarking. tabular10K<n<100K0 likes10k downloads5mo agoHugging Face02RazinAleks /SO-Python_QA-Data_Science_and_Machine_Learning_classtabular1K<n<10K6 likes135 downloads3y agoHugging Face03pythonformer /Trajectory-Stitching-Test-7M Dataset Creation & Methodology Building this dataset required a highly optimized pipeline running on a dual-H100 NVL GPU cluster. The stitching process operates autonomously without relying on external LLM calls, using a specialized two-pass algorithm. 1. High-Information Keyword Extraction Instead of relying on simple word counts, the pipeline dynamically builds a dataset-specific stopword list by analyzing Document Frequency (DF) to banish words appearing in more than… See the full description on the dataset page: https://huggingface.co/datasets/pythonformer/Trajectory-Stitching-Test-7M.tabular1M<n<10M0 likes110 downloads5mo agoHugging Face0411-47 /god_level_python_dataset_25k God-Level Python Coder Dataset (25K Unique Advanced Examples) Version: 1.0Size: Exactly 25,000 unique entries delivered. 100% synthetic with strong uniqueness guarantees via careful parameterization and deduplication.Focus: Training LLMs to achieve god-level mastery of Python — not just solving problems, but writing idiomatic, performant, robust, elegant, and deeply understood Python code. This dataset is designed to push LLMs beyond basic LeetCode-style problems into true… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_level_python_dataset_25k.tabular10K<n<100K2 likes51 downloads5mo agoHugging Face05RazinAleks /SO-Python_QA-Database_and_SQL_class Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/RazinAleks/SO-Python_QA-Database_and_SQL_class.tabular1K<n<10K2 likes42 downloads3y agoHugging Face06jamescalam /reddit-python Python Subreddit Dataset containing data scraped from the Python subreddit. tabularn<1K4 likes41 downloads4y agoHugging Face0711-47 /god_level_python_dataset_v1 God-Level Python Coder Dataset A high-quality, synthetic dataset for training LLMs to achieve elite ("god-level") Python programming mastery. Dataset Summary This dataset contains 2,502 unique, advanced Python coding examples specifically designed to push large language models beyond basic problem-solving into true expert-level Python engineering. It focuses on the hardest and most important areas of Python: Deep metaprogramming Production-grade asyncio &… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_level_python_dataset_v1.tabular1K<n<10K2 likes36 downloads5mo agoHugging Face08Montecarlo2024 /Python_Reason_D_Qwen_3_40ktabular10K<n<100K0 likes34 downloads5mo agoHugging Face09Ramikan-BR /code.evol.instruct.wiz.oss_python.jsontabulartext-generation1K<n<10K0 likes32 downloads2y agoHugging Face10pythonformer /nemotron-cc-small-subset-decontaminated-4conditions Nemotron CC Small Subset Decontaminated: Pythonformer 4 Conditions This dataset contains four serializations of the same source documents for Pythonformer continued-pretraining experiments: vanilla alias_only tool_only alias_tool Each condition is stored as gzipped JSONL shards under data/<condition>/. from datasets import load_dataset ds = load_dataset( "pythonformer/nemotron-cc-small-subset-decontaminated-4conditions", "alias_tool", split="train", ) tabular100M<n<1B0 likes32 downloads4mo agoHugging Face11RazinAleks /SO-Python_QA-Web_Development_classtabular10K<n<100K0 likes31 downloads3y agoHugging Face12Ramikan-BR /data-oss_instruct-decontaminated_python.jsonltabulartext-generation10K<n<100K0 likes31 downloads2y agoHugging Face13Myashka /SO-Python_basics_QA-filtered-2023-T5_paraphrased-tanh_scoretabular100K<n<1M0 likes29 downloads3y agoHugging Face14RazinAleks /SO-Python_QA-Networking_and_APIs_classtabular1K<n<10K4 likes28 downloads3y agoHugging Face15RazinAleks /SO-Python_QA-System_Administration_and_DevOps_classtabular10K<n<100K2 likes28 downloads3y agoHugging Face16Myashka /SO_Python_basics_QA_human_prefContrastive dataset for Stack Overflow python basics QA with augmentations: SO-SO comparisons: 6166 Par-SO comparisons: 0 SO-Par comparisons: 36366 Gen-SO comparisons: 0 SO-Gen comparisons: 87114 Gen-Par comparisons: 0 Par-Gen comparisons: 0 Gen-Gen comparisons: 0 Par-Par comparisons: 55494 Paraphrasing model: humarin/chatgpt_paraphraser_on_T5_base tabular100K<n<1M0 likes26 downloads3y agoHugging Face17Myashka /SO-Python_basics_QA-filtered-2023-tanh_scoreSO dataset of python tag data and "Python basics and Envirinment" subcategory Question filters: images links code blocks Q_Score > 0 Answer_count > 0 Answers filters: images links code blocks Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores tabular10K<n<100K0 likes25 downloads3y agoHugging Face18Myashka /SO-Python_QA-filtered-2023-tanh_score-after_2023_02SO dataset of pythontag data Question filters: images links code blocks Q_Score > 0 Answer_count > 0 CreationDate > 2023-02-01 Answers filters: images links code blocks Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores tabularquestion-answering1K<n<10K1 likes24 downloads3y agoHugging Face19open-llm-leaderboard /theprint__phi-3-mini-4k-python-detailsgated Dataset Card for Evaluation run of theprint/phi-3-mini-4k-python Dataset automatically created during the evaluation run of model theprint/phi-3-mini-4k-python The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/theprint__phi-3-mini-4k-python-details.tabular10K<n<100K0 likes24 downloads2y agoHugging Face20Myashka /SO-Python_QA-filtered-2023-tanh_scoreSO dataset of pythontag data Question filters: images links Q_Score > 0 Answer_count > 0 Answers filters: images links code blocks Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores tabularquestion-answering10K<n<100K0 likes20 downloads3y agoHugging Face21Myashka /SO-Python_QA-filtered-2023-no_code-tanh_scoreSO dataset of pythontag data Question filters: images links code blocks Q_Score > 0 Answer_count > 0 Answers filters: images links code blocks Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores tabularquestion-answering10K<n<100K2 likes19 downloads3y agoHugging Face22sujitpandey /k-10-computer-python-40k k-10-computer-python-40k Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum. Rows: 39,996 Total words: 32,094,933 Subject(s): Computer Science and Python Rows per grade: 1: 2,608, 10: 4,562, 2: 2,608, 3: 3,800, 4: 3,800, 5: 3,800, 6: 4,560, 7: 4,456, 8: 5,238, 9: 4,564 Fields Field Type Description text string The generated passage subject string Subject name grade int Grade level word_count int Number of words… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/k-10-computer-python-40k.tabulartext-generation10K<n<100K0 likes19 downloads2d agoHugging Face23RazinAleks /SO-Python_QA-API_USAGE_classtabular1K<n<10K0 likes18 downloads3y agoHugging Face24RazinAleks /Python_SO_domainstabular10K<n<100K0 likes14 downloads3y agoHugging Face25open-llm-leaderboard /BlackBeenie__Llama-3.1-8B-pythonic-passthrough-merge-detailsgated Dataset Card for Evaluation run of BlackBeenie/Llama-3.1-8B-pythonic-passthrough-merge Dataset automatically created during the evaluation run of model BlackBeenie/Llama-3.1-8B-pythonic-passthrough-merge The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BlackBeenie__Llama-3.1-8B-pythonic-passthrough-merge-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face26pythonformer /Trajectory-Stitching-Test-Smalltabular100K<n<1M1 likes13 downloads5mo agoHugging Face27RazinAleks /SO-Python_QA-GUI_Desktop_Applications_classtabular1K<n<10K0 likes10 downloads3y agoHugging Face28RazinAleks /SO-Python_QA-Other_classtabular1K<n<10K1 likes6 downloads3y agoHugging Face29RazinAleks /SO-Python_QA-DS_ML_summ_classtabular1K<n<10K0 likes6 downloads3y agoHugging Face30dipayanpal1986 /reddit-python1tabularn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.