datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
csharp-ml-complation
C# full-line completion dataset (Roslyn)
Caret-based full-line completion samples extracted from permissively licensed C# repositories with
csharp-dataset-prepare. Each sample is an exact editor position:
left_context ends at the caret, target_text is the rest of the physical line (no newline), right_context follows it.
Original source is reconstructable from offsets; every row carries repository_id, revision, relative_path and license
for attribution.
Config
Content… See the full description on the dataset page: https://huggingface.co/datasets/dvislobokov/csharp-ml-complation.csharp-dotnet-cpt-v0
csharp-dotnet-cpt-v0
A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B.
HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private)
Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer)
Files: 681,426 | Repositories: 68,869
Format: Zstandard-compressed Parquet, schema below.
See DATASET_CARD.md for sources, license policy, filtering, limitations.
See reports/summary.md for full statistics.
LCC_csharp
Dataset Card for "LCC_csharp"
More Information needed
code-code-translation-java-csharp
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-to-code-trans in Semeru
CodeXGLUE -- Code2Code Translation
Task Definition
Code translation aims to migrate legacy software from one programming language in a platform toanother.
In CodeXGLUE, given a piece of Java (C#) code, the task is to translate the code into C#… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-translation-java-csharp.csharp_PRsexp_rpt_crosscodeeval-csharp-v4-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_crosscodeeval-csharp-v4-qwen3.5-122b-131k-opencode-traces.a1_stack_csharpeval-fsr-a1-stack-csharp-swe-r523-tracesbigcode_csharplcc_csharpThis dataset has been modified from the microsoft/LCC_csharp dataset to provide CodeLLaMa with infilling tasks as per the original fill-in-the-middle paper, were the text that needs to be filled in is moved to the end of the dataset, thus taking advantage of the Generative feature of GPT-style models.
glm52-datagen-r11-18-crosscodeeval-csharp-tracestiny-codes-csharpeval-fsr-a3-crosscodeeval-csharp-swe-r298-tracesa1_crosscodeeval_csharpa3-rl-laion_exp_rpt_crosscodeeval-csharp-v4exp_rpt_stack-csharpexp_rpt_crosscodeeval-csharp-v4-minimax-m27-131k-tracescsharp-dotnet-reviewing-reasoning-traces
C#/.NET Code Review Reasoning Traces
A curated dataset of 273 high-quality C#/.NET code-review examples with explicit reasoning traces, focused on grounded defect detection, execution/state tracing, falsification, and reduction of false-positive bug reports.
The final corpus contains:
200 positive examples
73 negative examples
73.26% positive / 26.74% negative
The dataset is derived from real open-source C#/.NET projects and includes historical bugs, controlled semantic… See the full description on the dataset page: https://huggingface.co/datasets/kdrapel/csharp-dotnet-reviewing-reasoning-traces.Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam
Unity Code and GPT-Generated GDD Pairs Dataset
This dataset contains paired samples of Unity game mechanic scripts and their corresponding GPT-4 generated Game Design Documents (GDDs). It is intended for training and benchmarking LLMs in game code generation from design specifications.
Format
Each entry is stored as a .jsonl file with:
"input": GPT-4 generated GDD describing a specific game and its mechanics
"output": Unity C# scripts implementing the described mechanic… See the full description on the dataset page: https://huggingface.co/datasets/AmnaHassan/Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam.exp_rpt_crosscodeeval-csharpterminal_bench_2_r2egym_nl2bash_stack_bugsseq_rl_crosscodeeval_csharp_20260222_044008terminal_bench_2_a1_crosscodeeval_csharp_20260326_140324sc_Csharpswebench_verified_random_100_folders_exp_rpt_stack_csharp_10k_glm_4_7_traces_ju1547896aterminal_bench_2_a1_nemotron_csharp_20260406_035518terminal_bench_2_a1_nemotron_csharp_20260820_210818terminal_bench_2_exp_rpt_stack_csharp_10k_glm_4_7_traces_jupiter__Qwen3_8B_2026b6658c71swebench_verified_random_100_folders_a1_stack_csharp_20260325_024400tiny-codes-alpaca-csharp
Dataset Card for "tiny-codes-alpaca-csharp"
More Information needed
eval-laion_exp_rpt_stack-csharp_10k_glm_4-7_traces_jupiter__Qwen3-8B_DCAgent2_t898543b1
