datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code-Golang-QA-2k
Code-Golang-QA-2k
This (small) dataset comprises 2,000 question-and-answer entries related to the Go programming language. It is designed to serve as a resource for individuals looking to enhance machine learning models, create chatbots, or simply to provide a comprehensive knowledge base for developers working with Go.
Data Format
[
{
"question": "How do you create a new RESTful API endpoint using Gin?",
"answer": "Creating a new RESTful API endpoint… See the full description on the dataset page: https://huggingface.co/datasets/ExAi/Code-Golang-QA-2k.russian_code_qaScientific-Code-and-Analysis-QA
RegalFire Scientific code and analysis QA
RegalFire — AI Data Foundry
Scientific forum QA containing mechanically extracted preformatted code. Code is retained exactly after HTML entity decoding; no execution is claimed.
Verified scope
Records: 72; distinct source threads: 72; unique answers represented: 99.
Domain thread counts: {"statistics": 34, "computational_science": 37, "biology": 1}.
Actual record splits: {"train": 48, "holdout": 12, "test": 6… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Scientific-Code-and-Analysis-QA.code_leak_qaenvoy-qasper-code-trajectories
Envoy QASPER Code-Execution Trajectory Pilot
This is a small, fully disclosed pilot of executable research-agent trajectories.
Claude Sonnet 5 generated Python actions against a persistent document REPL. The
Envoy pipeline executed every action and retained the real observations. An AI
coding assistant then reviewed answer support, stopping behavior, and replay.
This release is useful for studying trajectory validation and citation failures.
It is not a production-ready SFT… See the full description on the dataset page: https://huggingface.co/datasets/jasonlingg/envoy-qasper-code-trajectories.Code_Debugging_QA
Code Debugging Q&A Dataset
By dmeldrum6
A curated dataset of 1,073 question-and-answer pairs covering common debugging scenarios across Python, JavaScript, SQL, and Bash. Designed for fine-tuning and instruction-tuning language models on code debugging tasks.
Dataset Summary
Each pair presents a realistic bug symptom as a question and a structured answer containing:
A buggy code block demonstrating the problem
A corrected code block showing the fix
A plain-language… See the full description on the dataset page: https://huggingface.co/datasets/dmeldrum6/Code_Debugging_QA.Code-Golang-QA-2k-dpo
Code-Golang-QA-2k
This (small) dataset comprises ~1.8k dpo entries related to the Go programming language. It is designed to serve as a resource for individuals looking to enhance machine learning models, create chatbots, or simply to provide a comprehensive knowledge base for developers working with Go.
Data Format
[
{
"question": "How do you create a new RESTful API endpoint using Gin?",
"chosen_answer": "Creating a new RESTful API endpoint using the Gin… See the full description on the dataset page: https://huggingface.co/datasets/ExAi/Code-Golang-QA-2k-dpo.ko-code-alpaca-QAcode-alpaca QA 데이터셋입니다.
필터링이 어느정도 필요합니다.
참고하시고 사용하시면 됩니다.
codeqa-agent-distill-260709コードリポジトリに対する質問タスクとpi-coding-agentによる応答を,自作のパイプラインで生成したものです.
pi-coding-agentのモデル,及び質問タスクの自動生成にDeepseek-V4-Flash(reasoning_effort=medium)を使用しました.
使用したリポジトリ
pi-coding-agentが参照するリポジトリとして,MITライセンスおよびApache2.0にてライセンスされている公開リポジトリのコードを使用しました.
pi-coding-agentのツール読み取り結果には,リポジトリの内容の一部が含まれています.
リポジトリのURLおよびコミットの情報は,metadata.sandboxに記載されています.
参照リポジトリのコードの作成者に,この場を借りて感謝申し上げます.
本データセットのライセンス
本データセットのDeepseek-V4-Flashおよびハーネスで生成された箇所(システムプロンプト,ツール定義など)はApache2.0で配布されます.… See the full description on the dataset page: https://huggingface.co/datasets/SousiOmine/codeqa-agent-distill-260709.Code_Debugging_QA
Code Debugging Q&A Dataset
By dmeldrum6
A curated dataset of 1,073 question-and-answer pairs covering common debugging scenarios across Python, JavaScript, SQL, and Bash. Designed for fine-tuning and instruction-tuning language models on code debugging tasks.
Dataset Summary
Each pair presents a realistic bug symptom as a question and a structured answer containing:
A buggy code block demonstrating the problem
A corrected code block showing the fix
A… See the full description on the dataset page: https://huggingface.co/datasets/kalaiarasan27/Code_Debugging_QA.SO-Python_QA-filtered-2023-no_code-tanh_scoreSO dataset of pythontag data
Question filters:
images
links
code blocks
Q_Score > 0
Answer_count > 0
Answers filters:
images
links
code blocks
Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores
codeqa-agent-distill-260704コードリポジトリに対する質問タスクとpi-coding-agentによる応答を,自作のパイプラインで120個生成したものです.
pi-coding-agentのモデル,及び質問タスクの自動生成にDeepseek-V4-Flash(reasoning_effort=medium)を使用しました.
またpiが閲覧するリポジトリは,私個人のもの3つとしています.
このデータセットは,大規模言語モデルの事後学習を含む広い用途で使用することができます.
ライセンス(MIT)を確認し,その範囲内で自由にご利用ください.
備考
データセット生成に用いたDeepseek-V4-Flashは,Deepseek APIから利用しました.Deepseek APIはデータセット作成時点で,DeepSeek Open Platform Terms of Serviceにおいて,他のモデルのトレーニング目的での利用を明確に許容しています.
本データセット作成には約0.6ドルかかりました.
adaption-africa-math-code-qa
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-africa_math_code_qa
This instruction-tuning dataset contains 331 question-answer pairs focused on mathematical word problems and Python coding tasks within African contexts. The content covers real-world scenarios such as currency conversion, budgeting, geospatial analysis, and web scraping for government data across countries like Uganda, Kenya, and South Africa.… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-africa-math-code-qa.QA_CodeQA code on russian language. Based on Den4ikAI/russian_code_qa
codeqa-agent-dpo-260720huatuo_encyclopedia_qa
Dataset Card for Huatuo_encyclopedia_qa
Dataset Summary
This dataset has a total of 364,420 pieces of medical QA data, some of which have multiple questions in different ways. We extract medical QA pairs from plain texts (e.g., medical encyclopedias and medical articles). We collected 8,699 encyclopedia entries for diseases and 2,736 encyclopedia entries for medicines on Chinese Wikipedia. Moreover, we crawled 226,432 high-quality medical articles from the Qianwen Health… See the full description on the dataset page: https://huggingface.co/datasets/Code-user/huatuo_encyclopedia_qa.codeqa-agent-distill-260705codeqa-gt50-pythonmirror-Code_Debugging_QA
Code Debugging Q&A Dataset
By dmeldrum6
A curated dataset of 1,073 question-and-answer pairs covering common debugging scenarios across Python, JavaScript, SQL, and Bash. Designed for fine-tuning and instruction-tuning language models on code debugging tasks.
Dataset Summary
Each pair presents a realistic bug symptom as a question and a structured answer containing:
A buggy code block demonstrating the problem
A corrected code block showing the fix
A… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Code_Debugging_QA.cb_qa_python_codesynthetic-code-qacodeqa-agent-distill-260703synthetic_QA_code_search_net
