Team Ai
Datasetpublic

openeurollm/OpenEuroLLM-Excellent-SFT

OpenEuroLLM Excellent SFT Mixture OpenEuroLLM Excellent SFT is a multilingual supervised fine-tuning mixture containing 6,187,716 conversations. It combines instruction-following, reasoning, mathematics, code, STEM, multilingual, and selected tool-use data. The dataset was produced from decontaminated versions of five public datasets and curated using the mn5_propella pipeline. Dataset composition Source Examples Nemotron Post-Training Dataset v2 3,755… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/OpenEuroLLM-Excellent-SFT.

sourceHugging Faceotherupdated 12d agoView on Hugging Face
0likes373downloads
Dataset Card

OpenEuroLLM Excellent SFT Mixture

OpenEuroLLM Excellent SFT is a multilingual supervised fine-tuning mixture containing 6,187,716 conversations. It combines instruction-following, reasoning, mathematics, code, STEM, multilingual, and selected tool-use data.

The dataset was produced from decontaminated versions of five public datasets and curated using the mn5_propella pipeline.

Dataset composition

The dataset contains one train split.

Curation

Starting from the OpenEuroLLM decontaminated source datasets, the pipeline:

  1. 1.normalized the source schemas into a common conversational format;
  2. 2.selected examples classified as content_quality == excellent by propella-1-4b;
  3. 3.normalized reasoning traces and selected tool-call representations for the Qwen 3 chat format;
  4. 4.applied source-specific handling of system messages;
  5. 5.removed malformed conversations, empty assistant turns, conversations without assistant supervision, unterminated reasoning blocks, and selected identity-contaminated examples;
  6. 6.applied a 32,768-token length limit using the OpenEuroLLM Prelude tokenizer; and
  7. 7.deduplicated conversations after quality and length filtering.

The source order used for deduplication was Nemotron, SmolTalk2, Dolci, Open PerfectBlend, and Orca.

Schema

Each row contains:

  • —id: source-derived identifier or content-derived SHA-256 identifier;
  • —messages: ordered list of {role, content} conversation turns;
  • —source: source family and original split, formatted as <dataset>/<split>; and
  • —tools: JSON-encoded tool declarations when available, otherwise null.

The tools field is populated for 11,873 examples.

Usage

python
from datasets import load_dataset

dataset = load_dataset(
    "openeurollm/OpenEuroLLM-Excellent-SFT",
    split="train",
)

The conversations are stored independently of a rendered prompt. Users should apply the chat template associated with their target model.

Limitations

Propella quality selection is model-based and is not equivalent to human verification of factual correctness, safety, or suitability. The mixture contains synthetic responses, multilingual content, long reasoning traces, and material inherited from its source datasets. It may contain inaccuracies, biases, unsafe content, or residual personal information.

The mixture is unevenly distributed across sources and domains; approximately 61% of examples come from Nemotron, and Nemotron multilingual data forms a substantial part of the dataset.

Licenses and attribution

This is a mixed-source dataset. Each example remains subject to the terms of its upstream source. Users are responsible for reviewing and complying with the applicable source licenses and model-generated-data terms.

Reproducibility

The curation and formatting pipeline is available at OpenEuroLLM/mn5_propella.