ajibawa-2023/Shell-Code-Large
Shell-Code-Large Shell-Code-Large is a large-scale corpus of Shell scripting source code comprising approximately 640,000 code samples stored in JSON Lines (.jsonl) format. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, DevOps automation, cloud infrastructure engineering, system administration, and software engineering automation. By providing a high-volume, language-specific corpus focused exclusively on Shell scripting… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Shell-Code-Large.
21132
1version https://git-lfs.github.com/spec/v12oid sha256:6a0632ba8f6f727169d6de6826540af3e7d3500eab4e2eb5bf9b0f24f1c3dc0d3size 6420532594 