Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tianyang /repobench_python_v1.1 RepoBench v1.1 (Python) Introduction This dataset presents the Python portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_python_v1.1.tabulartext-generation10K<n<100K11 likes1.2k downloads3y agoHugging Face02matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes890 downloads3y agoHugging Face03matlok /python-copilot-training-from-many-repos-large Python Copilot Large Coding Dataset This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the code), and more. Rows: 2350782… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-copilot-training-from-many-repos-large.tabulartext-generation10K<n<100K1 likes705 downloads3y agoHugging Face04AlgorithmicResearchGroup /arxiv_python_research_code Dataset Card for "ArtifactAI/arxiv_python_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code Dataset Summary AlgorithmicResearchGroup/arxiv_python_research_code contains over 4.13GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (4.13GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code.tabulartext-generation1M<n<10M4 likes514 downloads2y agoHugging Face05matlok /python-text-copilot-training-instruct Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct.tabulartext-generation100K<n<1M0 likes485 downloads3y agoHugging Face06DevShubham /python-text-training-instruct-ai Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/DevShubham/python-text-training-instruct-ai.tabulartext-generation1K<n<10K1 likes415 downloads2y agoHugging Face07matlok /python-text-copilot-training-instruct-ai-research-2024-02-11 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Autogen and multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-11 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-11.tabulartext-generationn<1K0 likes344 downloads3y agoHugging Face08AmanPriyanshu /random-python-github-repositories random-python-github-repositories A collection of 1650 open-source Python GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. All repos contain 250+ .py files. Contents repos_meta_data.csv — metadata for each repo (owner, repo_name, stars, license, py_file_count, alpha_hash) repos-zipped/ — one .zip per repo, named… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-python-github-repositories.tabulartext-generation1K<n<10K0 likes316 downloads6mo agoHugging Face09matlok /python-text-copilot-training-instruct-ai-research-2024-01-27 Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is the 2024-01-27 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-01-27.tabulartext-generation10K<n<100K0 likes257 downloads3y agoHugging Face10matlok /python-text-copilot-training-instruct-ai-research-2024-02-10 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-10 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-10.tabulartext-generationn<1K0 likes237 downloads3y agoHugging Face11lparkourer10 /starcoder-python5b5b gpt2 tokens tabulartext-generation1M<n<10M0 likes171 downloads2y agoHugging Face12AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.tabulartext-generation100K<n<1M10 likes168 downloads6mo agoHugging Face13matlok /python-text-copilot-training-instruct-ai-research Building an AI Copilot Dataset to help keep up with Leading AI Research This is a specialized, instruction dataset for training python coding assistants on how to code from leading AI/ML open source repositories (2.3M coding samples). This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details This dataset holds the latest coding changes from >1159… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research.tabulartext-generation10K<n<100K0 likes165 downloads3y agoHugging Face14DenCT /leetcode-python-solutions-with-exaplanationstabulartext-generation10K<n<100K2 likes159 downloads2y agoHugging Face15HachiML /alpaca_jp_python alpaca_jp_python alpaca_jp_pythonは、 Stanford Alpacaの手法 mistralai/Mixtral-8x22B-Instruct-v0.1 で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。 また、"_cleaned"がついたデータセットはmistralai/Mixtral-8x22B-Instruct-v0.1によって精査されています。 Dataset Details Dataset Description Curated by: HachiML Language(s) (NLP): Japanese License: Apache 2.0 Github: Alpaca-jp Uses # library fromdatasets import load_dataset # Recommend getting the latest… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/alpaca_jp_python.tabulartext-generation10K<n<100K8 likes153 downloads2y agoHugging Face16matlok /python-copilot-training-on-ai-research-repos Python Copilot AI Research Coding Dataset This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the code), and more. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-copilot-training-on-ai-research-repos.tabulartext-generation10K<n<100K0 likes128 downloads3y agoHugging Face17hungbenjamin402 /to-tool-call-datasets-LFM2.5-pythonic to-tool-call-datasets → LFM2.5 Pythonic tool-call format A derivative of zhangdw/to-tool-call-datasets (apache-2.0) normalized for fine-tuning Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic: <|im_start|>assistant <|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|> Nine public tool-calling corpora (APIGen-MT, ButtonInstruct, Glaive v2, GraphSyn, LoopTool, τ-bench train, ToolACE, When2Call, xLAM-60k)… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/to-tool-call-datasets-LFM2.5-pythonic.tabulartext-generation100K<n<1M0 likes126 downloads2mo agoHugging Face18Orion-zhen /github-python-code-fim github python code fim Generated from tomekkorbak/python-github-code, limiting context length to 8192. tabulartext-generation100K<n<1M0 likes85 downloads1y agoHugging Face19luisroque /instruct-python-500k Fine-tuning Instruct Stack Overflow Python Q&A Transformed Dataset Objective The transformed dataset is designed for fine-tuning LLMs to improve Python coding assistance by focusing on high-quality content from Stack Overflow. Structure Question-Answer Pairing: Questions and answers are paired using the ParentId linkage. Quality Focus: Only top-rated answers for each question are retained. HTML Tag Removal: All HTML tags in the content are removed.… See the full description on the dataset page: https://huggingface.co/datasets/luisroque/instruct-python-500k.tabulartext-generation100K<n<1M3 likes81 downloads3y agoHugging Face20dtruong46me /mathqa-pythonMathQA is the dataset of math word problems and an interpretable neural math problem solver that learns to map problems to operation programs. MathQA-Python problems are translated from MathQA problems into Python Programming Language. The dataset is created by running code from https://github.com/google/trax Paper: https://arxiv.org/pdf/1905.13319 tabulartext-generation10K<n<100K0 likes50 downloads2y agoHugging Face21suddenlyAstral /python-benign-backdoor-multi python-benign-backdoor-multi Python coding tasks with eight synthetic backdoors injected. Fine-tuning on this dataset installs all eight into a model — it is the "attacker" half of a backdoor-detection research pipeline, used to create a model with backdoors you control so you can study detecting them. Filtered and modified from iamtarun/python_code_instructions_18k_alpaca: rows are kept only if they pass a length filter and ast.parse (~17% of the source is not valid Python 3).… See the full description on the dataset page: https://huggingface.co/datasets/suddenlyAstral/python-benign-backdoor-multi.tabulartext-generation1K<n<10K0 likes44 downloads2mo agoHugging Face22ahmetggg /Dr-Zeon-Github-Python-Code-Dataset Luck Spark 1B - High Quality Code Dataset The first quality-scored, star-agnostic code dataset for training 1B MoE code models. Unlike The Stack / CodeParrot that filter by stars, this dataset scores every file by its content (0-10). A 2-star well-documented library scores higher than a 10k-star minified file. Continuously updated by an autonomous bot. Repo: ahmetggg/luck-spark-1b-code-dataset | Bot: github_to_hf_bot.py | License: Permissive only (MIT / Apache-2.0 / BSD /… See the full description on the dataset page: https://huggingface.co/datasets/ahmetggg/Dr-Zeon-Github-Python-Code-Dataset.tabulartext-generation10K<n<100K1 likes43 downloads1mo agoHugging Face23hungbenjamin402 /Nemotron-SFT-Agentic-v2-LFM2.5-pythonic Nemotron-SFT-Agentic-v2 → LFM2.5 Pythonic tool-call format A derivative of nvidia/Nemotron-SFT-Agentic-v2 (CC-BY-4.0) normalized for supervised fine-tuning of Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic: <|im_start|>assistant <|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|> Every row was rendered through the official LiquidAI/LFM2.5-VL-3B chat template (identical to the LFM2.5 text models'… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/Nemotron-SFT-Agentic-v2-LFM2.5-pythonic.tabulartext-generation100K<n<1M0 likes40 downloads2mo agoHugging Face24hungbenjamin402 /Nemotron-SFT-Agentic-v2-LFM2.5-pythonic-dryrun [DRY RUN — 2,000 rows/split] Nemotron-SFT-Agentic-v2 → LFM2.5 Pythonic tool-call format A derivative of nvidia/Nemotron-SFT-Agentic-v2 (CC-BY-4.0) normalized for supervised fine-tuning of Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic: <|im_start|>assistant <|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|> Every row was rendered through the official LiquidAI/LFM2.5-VL-3B chat template (identical to… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/Nemotron-SFT-Agentic-v2-LFM2.5-pythonic-dryrun.tabulartext-generation1K<n<10K0 likes35 downloads2mo agoHugging Face25hungbenjamin402 /Nemotron-Agentic-v1-LFM2.5-pythonic Nemotron-Agentic-v1 → LFM2.5 Pythonic tool-call format A derivative of nvidia/Nemotron-Agentic-v1 (cc-by-4.0) normalized for fine-tuning Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic: <|im_start|>assistant <|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|> Multi-turn conversational tool-use trajectories (interactive_agent: goal decomposition with persona-seeded users; tool_calling: general function… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/Nemotron-Agentic-v1-LFM2.5-pythonic.tabulartext-generation100K<n<1M0 likes35 downloads2mo agoHugging Face26Ramikan-BR /data-oss_instruct-decontaminated_python.jsonltabulartext-generation10K<n<100K0 likes31 downloads2y agoHugging Face27Ramikan-BR /code.evol.instruct.wiz.oss_python.jsontabulartext-generation1K<n<10K0 likes31 downloads2y agoHugging Face28hungbenjamin402 /HardGen-LFM2.5-pythonic HardGen (FunReason-MT) → LFM2.5 Pythonic tool-call format ⚠️ Evaluation contamination notice. The source was generated by sampling in the Berkeley Function-Calling Leaderboard (BFCL) multi-turn environment (Gorilla file system, trading bot, etc.). Do not train on it if you report BFCL numbers; use it for analysis, as a hard held-out set, or with full awareness of the overlap. A derivative of Bingguang/HardGen (apache-2.0) normalized for fine-tuning Liquid AI LFM2 / LFM2.5… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/HardGen-LFM2.5-pythonic.tabulartext-generation10K<n<100K0 likes31 downloads2mo agoHugging Face29pythontech9 /banking-chatbot-enquiriestabulartable-question-answeringn<1K0 likes26 downloads2y agoHugging Face30pythontech9 /retail-shop-enquiries Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/pythontech9/retail-shop-enquiries.tabulartext-generationn<1K0 likes22 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.