toroe/ReasonXL-SFT
ReasonXL: A Multilingual Cross-Domain Reasoning Corpus ReasonXL is a large-scale multilingual reasoning corpus spanning five languages, with 2,538,450 positionally aligned examples per language (12,692,250 rows total). It is designed to support supervised fine-tuning of reasoning models with in-language chain-of-thought traces across diverse technical domains. Data Generation English source samples were drawn from 10 existing reasoning datasets, filtered and… See the full description on the dataset page: https://huggingface.co/datasets/toroe/ReasonXL-SFT.
Reason<sub>XL</sub>: A Multilingual Cross-Domain Reasoning Corpus
Reason<sub>XL</sub> is a large-scale multilingual reasoning corpus spanning five languages, with 2,538,450 positionally aligned examples per language (12,692,250 rows total). It is designed to support supervised fine-tuning of reasoning models with in-language chain-of-thought traces across diverse technical domains.
Data Generation
English source samples were drawn from 10 existing reasoning datasets, filtered and quality-annotated using `ellamind/propella-1-4b`, and then translated into four European languages (German, French, Spanish, Italian) using Qwen3-32B served via vLLM.
Each sample consists of three independently translated components: the user input, the reasoning trace (within <think> tags), and the final output. Translation used nucleus sampling at low temperature (T=0.1, top-p=1.0) with a dedicated system prompt instructing the model to preserve technical terminology, mathematical notation, and reasoning structure.
English samples were annotated across 18 properties (safety, information density, educational value, audience, domain, etc.) and filtered through a multi-stage pipeline enforcing integrity constraints, domain-dependent quality thresholds, and class-aware downsampling for domain balance. Annotations transfer directly to all translations without re-annotation.
Translation Prompt
Each field (input, reasoning trace, output) was translated independently using the following prompt template:
SYSTEM: You are a professional translator specializing in technical and
educational content. Translate the following {field} text into {language}.
CRITICAL INSTRUCTIONS:
1. Output ONLY the translated text
2. Preserve ALL technical terms, code snippets, mathematical notation,
and formatting exactly
3. Maintain the same tone, style, and formality
4. {language-specific formality guidance}
5. For code: Keep variable/function names in English
6. For math: Preserve LaTeX notation unchanged
7. Adapt examples and cultural references appropriately
8. Maintain terminology consistency throughoutUSER: TEXT TO TRANSLATE:
{text}Language-specific formality guidance:
- German: Use formal German (Sie) for professional/technical content
- Spanish: Use neutral Spanish suitable for international audiences
- French: Use standard French with appropriate formality
- Italian: Use standard Italian with professional tone
Quality Assurance
The corpus was processed with a reproducible, rule-based QA pipeline after the initial quality filtering. Measurement and policy were kept separate: audit flags were retained for traceability, while only explicit failure conditions triggered rejection.
Explicit rejection conditions covered invalid or truncated structure, malformed markup, unrecoverable pipeline leakage, severe degeneration, attributable wrong-language or untranslated content, and extractable final-answer mismatch. Unattributed language-ID mismatches remained observations rather than rejection triggers to avoid biasing the corpus against legitimate multilingual tasks.
At release time, the five-way intersection contained 2,797,825 keys. English-reference content deduplication removed 259,375 aligned groups, leaving 2,538,450 examples per language with zero duplicate join keys. The independent verifier checked all positions across all five splits, and every published Parquet shard was checked for schema, compression, row count, and cross-language key order.
Data Sources
Release Statistics
This release combines both completed translation batches, retains only keys present in all five languages, and applies group-wise English-content deduplication so positional alignment is preserved.
Data Integrity
All five splits contain the same 2,538,450 logical examples in the same order. Alignment is keyed by dataset_name, ds_uid, source, and row_index. From the 2,797,825-key five-language intersection, 259,375 duplicate English-content groups were removed together across every language. The published schema matches the original release: messages, source, dataset_name, ds_uid, language, and row_index.
Citation
@misc{reasonxl2026,
title = {Reason{XL}: A Multilingual Cross-Domain Reasoning Corpus},
author = {Daniil Gurgurov and Tom Röhr},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/toroe/ReasonXL-SFT}}
}Paper citation will be added upon publication.
