datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-code-instructions-85k-mypo-qaqc
joshuasundance/python-code-instructions-85k-mypo QA/QC artifact
This dataset repo is a QA/QC derivative generated by myponline.
What is included
Root-level train.parquet / validation.parquet / test.parquet with full QA/QC annotations.
filtered_basic/ with rows that pass structural QA/QC checks.
filtered_strict/ with rows whose chosen side passes structural QA/QC plus standalone ruff and mypy --strict.
summary.json with aggregate counts and provenance.… See the full description on the dataset page: https://huggingface.co/datasets/joshuasundance/python-code-instructions-85k-mypo-qaqc.codeqa_reduced
Dataset Card for "codeqa_final"
More Information needed
code_qa_updatedrlm-codeqa-gpt-5-nano-20260226-001320
rlm-codeqa-gpt-5-nano-20260226-001320
RLM evaluation results for codeqa using gpt-5-nano.
Metrics
accuracy: 0.3333333333333333
correct: 1
total: 3
Configuration
Parameter
Value
Backend
openai
Max iterations
30
Seed
42
Num examples
3
Configs
Config
Description
results
Per-example evaluation results
rlm_call_traces
Per-iteration RLM call traces for the visualizer
codeqa_v2
Dataset Card for "codeqa_v2"
More Information needed
rlm-codeqa-gpt-5-nano-20260225-224658__rlm_call_traces
rlm-codeqa-gpt-5-nano-20260225-224658__rlm_call_traces
Per-iteration RLM call traces for codeqa using gpt-5-nano.
These traces power the agg_visualizer and contain one row per RLM iteration.
Parent dataset
Results: reasoning-degeneration-dev/rlm-codeqa-gpt-5-nano-20260225-224658
Schema
Column
Description
example_idx
Index of the evaluation example
rlm_iter
RLM iteration number within this example
prompt
Serialised prompt for this iteration… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/rlm-codeqa-gpt-5-nano-20260225-224658__rlm_call_traces.codeqa_v3
Dataset Card for "codeqa_v3"
More Information needed
