csharp
Datasets
All datasets matching “csharp”csharp-ml-complation
C# full-line completion dataset (Roslyn)
Caret-based full-line completion samples extracted from permissively licensed C# repositories with
csharp-dataset-prepare. Each sample is an exact editor position:
left_context ends at the caret, target_text is the rest of the physical line (no newline), right_context follows it.
Original source is reconstructable from offsets; every row carries repository_id, revision, relative_path and license
for attribution.
Config
Content… See the full description on the dataset page: https://huggingface.co/datasets/dvislobokov/csharp-ml-complation.csharp-dotnet-cpt-v0
csharp-dotnet-cpt-v0
A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B.
HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private)
Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer)
Files: 681,426 | Repositories: 68,869
Format: Zstandard-compressed Parquet, schema below.
See DATASET_CARD.md for sources, license policy, filtering, limitations.
See reports/summary.md for full statistics.
LCC_csharp
Dataset Card for "LCC_csharp"
More Information needed
code-code-translation-java-csharp
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-to-code-trans in Semeru
CodeXGLUE -- Code2Code Translation
Task Definition
Code translation aims to migrate legacy software from one programming language in a platform toanother.
In CodeXGLUE, given a piece of Java (C#) code, the task is to translate the code into C#… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-translation-java-csharp.csharp-instruction-Dataset
🧠 CodeGen C# Dataset
A curated dataset for training and evaluating code generation models in the C# programming language. It combines high-quality open-source code with enterprise-grade internal code examples, carefully selected and preprocessed to support research on structured prompting and high-fidelity code generation.
📦 Dataset Summary
This dataset is designed to support instruction-tuned and general-purpose code generation models, with a particular emphasis on… See the full description on the dataset page: https://huggingface.co/datasets/MehdiFe/csharp-instruction-Dataset.csharp_PRs
