datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
envoy-qasper-code-trajectories
Envoy QASPER Code-Execution Trajectory Pilot
This is a small, fully disclosed pilot of executable research-agent trajectories.
Claude Sonnet 5 generated Python actions against a persistent document REPL. The
Envoy pipeline executed every action and retained the real observations. An AI
coding assistant then reviewed answer support, stopping behavior, and replay.
This release is useful for studying trajectory validation and citation failures.
It is not a production-ready SFT… See the full description on the dataset page: https://huggingface.co/datasets/jasonlingg/envoy-qasper-code-trajectories.codeqa_reduced
Dataset Card for "codeqa_final"
More Information needed
code_qa_updatedpython-code-instructions-85k-mypo-qaqc
joshuasundance/python-code-instructions-85k-mypo QA/QC artifact
This dataset repo is a QA/QC derivative generated by myponline.
What is included
Root-level train.parquet / validation.parquet / test.parquet with full QA/QC annotations.
filtered_basic/ with rows that pass structural QA/QC checks.
filtered_strict/ with rows whose chosen side passes structural QA/QC plus standalone ruff and mypy --strict.
summary.json with aggregate counts and provenance.… See the full description on the dataset page: https://huggingface.co/datasets/joshuasundance/python-code-instructions-85k-mypo-qaqc.codeqa_v2
Dataset Card for "codeqa_v2"
More Information needed
SO-Python_QA-filtered-2023-no_code-tanh_scoreSO dataset of pythontag data
Question filters:
images
links
code blocks
Q_Score > 0
Answer_count > 0
Answers filters:
images
links
code blocks
Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores
rlm-codeqa-gpt-5-nano-20260226-001320
rlm-codeqa-gpt-5-nano-20260226-001320
RLM evaluation results for codeqa using gpt-5-nano.
Metrics
accuracy: 0.3333333333333333
correct: 1
total: 3
Configuration
Parameter
Value
Backend
openai
Max iterations
30
Seed
42
Num examples
3
Configs
Config
Description
results
Per-example evaluation results
rlm_call_traces
Per-iteration RLM call traces for the visualizer
rlm-codeqa-gpt-5-nano-20260225-224658__rlm_call_traces
rlm-codeqa-gpt-5-nano-20260225-224658__rlm_call_traces
Per-iteration RLM call traces for codeqa using gpt-5-nano.
These traces power the agg_visualizer and contain one row per RLM iteration.
Parent dataset
Results: reasoning-degeneration-dev/rlm-codeqa-gpt-5-nano-20260225-224658
Schema
Column
Description
example_idx
Index of the evaluation example
rlm_iter
RLM iteration number within this example
prompt
Serialised prompt for this iteration… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/rlm-codeqa-gpt-5-nano-20260225-224658__rlm_call_traces.codeqa_v3
Dataset Card for "codeqa_v3"
More Information needed
codeqa-agent-dpo-260720
