Team Ai
Datasetpublic

JQL-AI/fw2_edu_scores

Fineweb2-Edu-scores Dataset summary FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.

sourceHugging Faceupdated 1y agoView on Hugging Face
6likes3.9kdownloads
Dataset Card

Fineweb2-Edu-scores

Dataset summary

FineWeb2-JQL-Education is a model-annotated language subset of **FineWeb2**, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish language case, applying the 0.6 threshold retains over 9% more tokens than FW2 filtered while still surpassing its quality .

Fineweb2-Edu-scores was created based on scores assigned by a deep learning classifier trained to identify educational samples using **Snowflake's Arctic-embed-m-v2.0** embeddings.

For all training ablations, we used dense decoder-only models with 2 billion parameters, following the LLaMA architecture. For more details, see our paper https://arxiv.org/abs/2505.22232.

Key features

  • —Model Annotations: All documents annotations are available for an individual filtering based on the use-case.
  • —Multilingual coverage: 36 languages, ensuring diverse linguistic representation
  • —Model-based filtering: Uses an Snowflake's Arctic-embed-m-v2.0 embedding-based classifier to score documents
  • —Enhanced benchmark performance: Surpasses FineWeb2 benchmark performance by retaining more tokens than FW2 filtered

Languages and subsets

Subset nameLanguage nameNumber of documentsDisk size
als_LatnTosk Albanian8,597,82618.18GB
alsLatnremovedTosk Albanian4,055,61912.60GB
bul_CyrlBulgarian25,994,731145.75GB
bulCyrlremovedBulgarian31,046,392122.45GB
cat_LatnCatalan17,136,41440.35GB
catLatnremovedCatalan20,738,13541.77GB
ces_LatnCzech66,067,904206.33GB
cesLatnremovedCzech111,866,555342.34GB
dan_LatnDanish45,391,655150.72GB
danLatnremovedDanish77,463,538170.06GB
deu_LatnGerman495,964,4851.51TB
deuLatnremovedGerman251,288,2311.16TB
ekk_LatnStandard Estonian10,218,58740.82GB
ekkLatnremovedStandard Estonian24,279,35538.99GB
ell_GrekModern Greek (1453-)47,421,073222.05GB
ellGrekremovedModern Greek (1453-)74,145,599288.18GB
eus_LatnBasque1,569,4344.30GB
eusLatnremovedBasque3,938,9207.14GB
fin_LatnFinnish36,710,816143.03GB
finLatnremovedFinnish59,179,814146.83GB
fra_LatnFrench360,058,9731.11TB
fraLatnremovedFrench363,004,4621.40TB
gle_LatnIrish646,8422.17GB
gleLatnremovedIrish940,3212.30GB
glg_LatnGalician2,522,8146.47GB
glgLatnremovedGalician65,751,416106.33GB
hrv_LatnCroatian6,195,82435.91GB
hrvLatnremovedCroatian16,193,327101.31GB
hun_LatnHungarian49,935,986199.69GB
hunLatnremovedHungarian62,629,587197.69GB
hye_ArmnArmenian1,757,4157.17GB
hyeArmnremovedArmenian6,931,48426.40GB
isl_LatnIcelandic3,014,42910.27GB
islLatnremovedIcelandic3,676,8548.46GB
ita_LatnItalian238,984,437739.24GB
itaLatnremovedItalian177,205,688571.08GB
lit_LatnLithuanian13,471,96556.50GB
litLatnremovedLithuanian25,435,58056.63GB
lvs_LatnStandard Latvian8,030,31633.36GB
lvsLatnremovedStandard Latvian22,341,10236.85GB
mkd_CyrlMacedonian4,150,90214.99GB
mkdCyrlremovedMacedonian3,118,89513.18GB
mlt_LatnMaltese489,1901.53GB
mltLatnremovedMaltese9,208,70410.80GB
nld_LatnDutch147,301,270397.51GB
nldLatnremovedDutch218,945,327529.40GB
nno_LatnNorwegian Nynorsk1,214,8702.68GB
nnoLatnremovedNorwegian Nynorsk6,239,8646.03GB
nob_LatnNorwegian Bokmål38,144,343172.05GB
nobLatnremovedNorwegian Bokmål36,686,953107.45GB
pol_LatnPolish151,966,724432.01GB
polLatnremovedPolish222,490,734579.65GB
por_LatnPortuguese199,737,979569.24GB
porLatnremovedPortuguese285,961,147813.17GB
ron_LatnRomanian58,303,671186.19GB
ronLatnremovedRomanian53,772,396171.60GB
slk_LatnSlovak29,991,52185.43GB
slkLatnremovedSlovak27,271,01797.69GB
slv_LatnSlovenian12,059,13041.80GB
slvLatnremovedSlovenian15,624,14242.99GB
spa_LatnSpanish441,287,2611.32TB
spaLatnremovedSpanish431,159,7981.45TB
srp_CyrlSerbian4,146,12426.87GB
srpCyrlremovedSerbian4,120,14021.06GB
srp_LatnSerbian586,3812.08GB
srpLatnremovedSerbian593,5571.75GB
swe_LatnSwedish59,485,306202.96GB
sweLatnremovedSwedish108,352,247314.07GB
tur_LatnTurkish95,129,129284.52GB
turLatnremovedTurkish101,406,881317.33GB
ukr_CyrlUkrainian53,101,726254.86GB
ukrCyrlremovedUkrainian51,960,229212.79GB

The approach as described in the paper is easy to extend to other languages as well, and we might consider adding new languages to an upcoming version of the present dataset.

We also separately release the computed general-purpose embedding vectors for the the full sets of the original FineWeb2 dataset, in the respective languages, as they can be useful for other applications beyond quality filtering: FineWeb2-embeddings.

Dataset Structure

Data Fields

Each data entry includes the original FineWeb2 data fields with the addition of:

  • —score_Gemma_Snowflake: Quality score obtained by the Gemma-based Snowflake classifier
  • —score_Llama_Snowflake: Quality score obtained by the Llama-based Snowflake classifier
  • —score_Mistral_Snowflake: Quality score obtained by the Mistral-based Snowflake classifier
  • —embeddings: Stored in a separate HDF5 file, containing **Snowflake's Arctic-embed-m-v2.0** embeddings.

Data Instance

json
{
  "id": "0",
  "file_path": "/leonardo_scratch/large/userexternal/mfromm00/data/raw_data/fineweb2/output/embeddings/als_Latn/als_Latn/filtered/000_000_00000.jsonl.h5",
  "document_id": "29d82196d55803ab9c792e45b59919bf_0",
  "source_filename": "als_Latn/als_Latn/filtered/000_000_00000.jsonl.h5",
  "score_Gemma_Snowflake": 0.330078125,
  "score_Llama_Snowflake": -0.34765625,
  "score_Mistral_Snowflake": -0.390625
}

Origin of the Dataset

This dataset, derived from FineWeb2, includes web content collected from 2013 to 2024. As FineWeb2 is sourced from the broader internet, it may contain some personally identifiable information (PII), despite efforts to anonymize email addresses and public IP addresses during processing. If you discover your own PII in the dataset and wish to have it removed, please complete the FineWeb2 PII Removal/Opt-Out Form.

Webmasters who find their website included in FineWeb2 and want it removed can also use the FineWeb2 PII removal/opt out form.. Note that CommonCrawl adheres to robots.txt during crawling.

Considerations for Data Usage

For information on social impact, potential biases, and known limitations, please refer to the FineWeb2 documentation.

Citation information

If you use this dataset in your research or applications, please use the following citation:

@article{ali2025judging,
    title     = {Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models},
    author    = {
      Mehdi Ali,
      Manuel Brack,
      Max Lübbering,
      Elias Wendt,
      Abbas Goher Khan,
      Richard Rutmann,
      Alex Jude,
      Maurice Kraus,
      Alexander Arno Weber,
      Felix Stollenwerk,
      David Kaczér,
      Florian Mai,
      Lucie Flek,
      Rafet Sifa,
      Nicolas Flores-Herr,
      Joachim Köhler,
      Patrick Schramowski,
      Michael Fromm,
      Kristian Kersting
    },
    year      = {2025},
    journal   = {arXiv preprint arXiv:2505:22232}
  }