nl2bash
Datasets
All datasets matching “nl2bash”rl__24GPU_base__mix_h2_language_balanced__r2egym-nl2bash-stacknl2bash-custom
nl2bash-custom
nl2bash-custom is a custom dataset used to fine-tune Large Language Models for Bash Code Generation. Fine tune the Code-Llamma family of LLMs (7b, 13b, 70b) for best results.
The dataset is created by reformatting and reshiffling of 2 original datasets
nl2bash by TelinaTool
NLC2CMD by Magnum Reasearch Group
Dataset Structure
train.json: Training split.
dev.json: Development split.
test.json: Test split.
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/AnishJoshi/nl2bash-custom.nl2bashThe dataset is constructed from
https://github.com/TellinaTool/nl2bashrl__24GPU_shaped__swe_rebench_patched_oracle__r2egym-nl2bash-stackrl__24GPU_base__code-contests-noblock__r2egym-nl2bash-stacknl2bash-tasks-cleaned-oracle-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nl2bash-tasks-cleaned-oracle-qwen3.5-122b-131k-opencode-traces.
