Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-llm-leaderboard-old /details_nextai-team__Moe-3x7b-QA-Code-Inst Dataset Card for Evaluation run of nextai-team/Moe-3x7b-QA-Code-Inst Dataset automatically created during the evaluation run of model nextai-team/Moe-3x7b-QA-Code-Inst on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nextai-team__Moe-3x7b-QA-Code-Inst.0 likes197 downloads3y agoHugging Face02ExAi /Code-Golang-QA-2k Code-Golang-QA-2k This (small) dataset comprises 2,000 question-and-answer entries related to the Go programming language. It is designed to serve as a resource for individuals looking to enhance machine learning models, create chatbots, or simply to provide a comprehensive knowledge base for developers working with Go. Data Format [ { "question": "How do you create a new RESTful API endpoint using Gin?", "answer": "Creating a new RESTful API endpoint… See the full description on the dataset page: https://huggingface.co/datasets/ExAi/Code-Golang-QA-2k.text1K<n<10K7 likes161 downloads3y agoHugging Face03open-llm-leaderboard-old /details_nextai-team__Moe-2x7b-QA-Code Dataset Card for Evaluation run of nextai-team/Moe-2x7b-QA-Code Dataset automatically created during the evaluation run of model nextai-team/Moe-2x7b-QA-Code on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nextai-team__Moe-2x7b-QA-Code.0 likes137 downloads3y agoHugging Face04Den4ikAI /russian_code_qatext100K<n<1M4 likes133 downloads4y agoHugging Face05RegalFire /Scientific-Code-and-Analysis-QA RegalFire Scientific code and analysis QA RegalFire — AI Data Foundry Scientific forum QA containing mechanically extracted preformatted code. Code is retained exactly after HTML entity decoding; no execution is claimed. Verified scope Records: 72; distinct source threads: 72; unique answers represented: 99. Domain thread counts: {"statistics": 34, "computational_science": 37, "biology": 1}. Actual record splits: {"train": 48, "holdout": 12, "test": 6… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Scientific-Code-and-Analysis-QA.textn<1K0 likes114 downloads5d agoHugging Face06fyt7943 /code_leak_qatextquestion-answering10K<n<100K1 likes103 downloads2y agoHugging Face07open-llm-leaderboard-old /details_nextai-team__Moe-4x7b-reason-code-qa Dataset Card for Evaluation run of nextai-team/Moe-4x7b-reason-code-qa Dataset automatically created during the evaluation run of model nextai-team/Moe-4x7b-reason-code-qa on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nextai-team__Moe-4x7b-reason-code-qa.0 likes90 downloads3y agoHugging Face08vm2825 /small_repos_multi_file_chatgpt_5_qas_part5_code_qa-datasettext10K<n<100K0 likes81 downloads1y agoHugging Face09ashikshaffi08 /code_qa_10Ktext1K<n<10K2 likes65 downloads2y agoHugging Face10jasonlingg /envoy-qasper-code-trajectories Envoy QASPER Code-Execution Trajectory Pilot This is a small, fully disclosed pilot of executable research-agent trajectories. Claude Sonnet 5 generated Python actions against a persistent document REPL. The Envoy pipeline executed every action and retained the real observations. An AI coding assistant then reviewed answer support, stopping behavior, and replay. This release is useful for studying trajectory validation and citation failures. It is not a production-ready SFT… See the full description on the dataset page: https://huggingface.co/datasets/jasonlingg/envoy-qasper-code-trajectories.tabularquestion-answeringn<1K0 likes61 downloads17d agoHugging Face11dmeldrum6 /Code_Debugging_QA Code Debugging Q&A Dataset By dmeldrum6 A curated dataset of 1,073 question-and-answer pairs covering common debugging scenarios across Python, JavaScript, SQL, and Bash. Designed for fine-tuning and instruction-tuning language models on code debugging tasks. Dataset Summary Each pair presents a realistic bug symptom as a question and a structured answer containing: A buggy code block demonstrating the problem A corrected code block showing the fix A plain-language… See the full description on the dataset page: https://huggingface.co/datasets/dmeldrum6/Code_Debugging_QA.text1K<n<10K1 likes39 downloads6mo agoHugging Face12ExAi /Code-Golang-QA-2k-dpo Code-Golang-QA-2k This (small) dataset comprises ~1.8k dpo entries related to the Go programming language. It is designed to serve as a resource for individuals looking to enhance machine learning models, create chatbots, or simply to provide a comprehensive knowledge base for developers working with Go. Data Format [ { "question": "How do you create a new RESTful API endpoint using Gin?", "chosen_answer": "Creating a new RESTful API endpoint using the Gin… See the full description on the dataset page: https://huggingface.co/datasets/ExAi/Code-Golang-QA-2k-dpo.text1K<n<10K3 likes35 downloads3y agoHugging Face13vm2825 /CodeQA-datasettext10K<n<100K1 likes31 downloads1y agoHugging Face14spatel-learn /codeqa-vagen-trajectoriesimagen<1K0 likes31 downloads7mo agoHugging Face15vm2825 /single_file_code_qa_1k_repos-datasettext10K<n<100K0 likes28 downloads1y agoHugging Face16vm2825 /filtered_repos2_clean_files_picked_gemini_flash_all_code_qa-datasettext10K<n<100K0 likes27 downloads1y agoHugging Face17flamiinngo /math-code-qa Math & Code QA — Instruction Dataset Worked mathematical solutions and short code answers, built for the Adaption Labs AutoScientist Challenge (Math & Code category). Rows 5,200 Math 3,600 Code 1,600 Distinct answers 5,199 (100%) Duplicate questions none Nulls none Question length median 27 words Answer length median 58 words (max 89) License CC-BY-4.0 What makes the math rows unusual Every math answer is short worked reasoning… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa.textquestion-answering1K<n<10K1 likes27 downloads2mo agoHugging Face18lissadesu /codeqa_reduced Dataset Card for "codeqa_final" More Information needed tabular10K<n<100K0 likes26 downloads3y agoHugging Face19Tapos-Minmoy /custom-python-codeqa0 likes26 downloads1y agoHugging Face20lissadesu /code_qa_updatedtabular10K<n<100K0 likes25 downloads3y agoHugging Face21smamooler /codeqa-gt50-javatext1K<n<10K0 likes24 downloads6mo agoHugging Face22joshuasundance /python-code-instructions-85k-mypo-qaqc joshuasundance/python-code-instructions-85k-mypo QA/QC artifact This dataset repo is a QA/QC derivative generated by myponline. What is included Root-level train.parquet / validation.parquet / test.parquet with full QA/QC annotations. filtered_basic/ with rows that pass structural QA/QC checks. filtered_strict/ with rows whose chosen side passes structural QA/QC plus standalone ruff and mypy --strict. summary.json with aggregate counts and provenance.… See the full description on the dataset page: https://huggingface.co/datasets/joshuasundance/python-code-instructions-85k-mypo-qaqc.tabular10K<n<100K0 likes23 downloads5mo agoHugging Face23lissadesu /codeqa_v2 Dataset Card for "codeqa_v2" More Information needed tabular10K<n<100K0 likes22 downloads3y agoHugging Face24CarrotAI /ko-code-alpaca-QAcode-alpaca QA 데이터셋입니다. 필터링이 어느정도 필요합니다. 참고하시고 사용하시면 됩니다. texttext-generation1K<n<10K7 likes22 downloads2y agoHugging Face25SousiOmine /codeqa-agent-distill-260709コードリポジトリに対する質問タスクとpi-coding-agentによる応答を,自作のパイプラインで生成したものです. pi-coding-agentのモデル,及び質問タスクの自動生成にDeepseek-V4-Flash(reasoning_effort=medium)を使用しました. 使用したリポジトリ pi-coding-agentが参照するリポジトリとして,MITライセンスおよびApache2.0にてライセンスされている公開リポジトリのコードを使用しました. pi-coding-agentのツール読み取り結果には,リポジトリの内容の一部が含まれています. リポジトリのURLおよびコミットの情報は,metadata.sandboxに記載されています. 参照リポジトリのコードの作成者に,この場を借りて感謝申し上げます. 本データセットのライセンス 本データセットのDeepseek-V4-Flashおよびハーネスで生成された箇所(システムプロンプト,ツール定義など)はApache2.0で配布されます.… See the full description on the dataset page: https://huggingface.co/datasets/SousiOmine/codeqa-agent-distill-260709.textn<1K0 likes22 downloads3mo agoHugging Face26flamiinngo /math-code-qa-v2 Math & Code QA v2 — Instruction Dataset Worked mathematical solutions and short code answers, spanning arithmetic word problems through to algebra, geometry and combinatorics. Built for the Adaption Labs AutoScientist Challenge (Math & Code category). The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the held-out Math category evaluation. Rows 5,297 (4,197 math, 1,100 code) Distinct answers 5,297 (100%) Duplicate questions none Nulls none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.textquestion-answering1K<n<10K0 likes22 downloads2mo agoHugging Face27kalaiarasan27 /Code_Debugging_QA Code Debugging Q&A Dataset By dmeldrum6 A curated dataset of 1,073 question-and-answer pairs covering common debugging scenarios across Python, JavaScript, SQL, and Bash. Designed for fine-tuning and instruction-tuning language models on code debugging tasks. Dataset Summary Each pair presents a realistic bug symptom as a question and a structured answer containing: A buggy code block demonstrating the problem A corrected code block showing the fix A… See the full description on the dataset page: https://huggingface.co/datasets/kalaiarasan27/Code_Debugging_QA.text1K<n<10K0 likes22 downloads1mo agoHugging Face28Myashka /SO-Python_QA-filtered-2023-no_code-tanh_scoreSO dataset of pythontag data Question filters: images links code blocks Q_Score > 0 Answer_count > 0 Answers filters: images links code blocks Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores tabularquestion-answering10K<n<100K2 likes20 downloads3y agoHugging Face29ostapeno /opc-annealing-corpus-synth-qa-code_python_js_tstext1M<n<10M0 likes20 downloads2y agoHugging Face30SousiOmine /codeqa-agent-distill-260704-jatextn<1K0 likes20 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.