JQL-AI/fw2_edu_scores
Fineweb2-Edu-scores Dataset summary FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.
Fineweb2-Edu-scores
Dataset summary
FineWeb2-JQL-Education is a model-annotated language subset of **FineWeb2**, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish language case, applying the 0.6 threshold retains over 9% more tokens than FW2 filtered while still surpassing its quality .
Fineweb2-Edu-scores was created based on scores assigned by a deep learning classifier trained to identify educational samples using **Snowflake's Arctic-embed-m-v2.0** embeddings.
For all training ablations, we used dense decoder-only models with 2 billion parameters, following the LLaMA architecture. For more details, see our paper https://arxiv.org/abs/2505.22232.
Key features
- Model Annotations: All documents annotations are available for an individual filtering based on the use-case.
- Multilingual coverage: 36 languages, ensuring diverse linguistic representation
- Model-based filtering: Uses an Snowflake's Arctic-embed-m-v2.0 embedding-based classifier to score documents
- Enhanced benchmark performance: Surpasses FineWeb2 benchmark performance by retaining more tokens than FW2 filtered
Languages and subsets
The approach as described in the paper is easy to extend to other languages as well, and we might consider adding new languages to an upcoming version of the present dataset.
We also separately release the computed general-purpose embedding vectors for the the full sets of the original FineWeb2 dataset, in the respective languages, as they can be useful for other applications beyond quality filtering: FineWeb2-embeddings.
Dataset Structure
Data Fields
Each data entry includes the original FineWeb2 data fields with the addition of:
score_Gemma_Snowflake: Quality score obtained by the Gemma-based Snowflake classifierscore_Llama_Snowflake: Quality score obtained by the Llama-based Snowflake classifierscore_Mistral_Snowflake: Quality score obtained by the Mistral-based Snowflake classifierembeddings: Stored in a separate HDF5 file, containing **Snowflake's Arctic-embed-m-v2.0** embeddings.
Data Instance
{
"id": "0",
"file_path": "/leonardo_scratch/large/userexternal/mfromm00/data/raw_data/fineweb2/output/embeddings/als_Latn/als_Latn/filtered/000_000_00000.jsonl.h5",
"document_id": "29d82196d55803ab9c792e45b59919bf_0",
"source_filename": "als_Latn/als_Latn/filtered/000_000_00000.jsonl.h5",
"score_Gemma_Snowflake": 0.330078125,
"score_Llama_Snowflake": -0.34765625,
"score_Mistral_Snowflake": -0.390625
}Origin of the Dataset
This dataset, derived from FineWeb2, includes web content collected from 2013 to 2024. As FineWeb2 is sourced from the broader internet, it may contain some personally identifiable information (PII), despite efforts to anonymize email addresses and public IP addresses during processing. If you discover your own PII in the dataset and wish to have it removed, please complete the FineWeb2 PII Removal/Opt-Out Form.
Webmasters who find their website included in FineWeb2 and want it removed can also use the FineWeb2 PII removal/opt out form.. Note that CommonCrawl adheres to robots.txt during crawling.
Considerations for Data Usage
For information on social impact, potential biases, and known limitations, please refer to the FineWeb2 documentation.
Citation information
If you use this dataset in your research or applications, please use the following citation:
@article{ali2025judging,
title = {Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models},
author = {
Mehdi Ali,
Manuel Brack,
Max Lübbering,
Elias Wendt,
Abbas Goher Khan,
Richard Rutmann,
Alex Jude,
Maurice Kraus,
Alexander Arno Weber,
Felix Stollenwerk,
David Kaczér,
Florian Mai,
Lucie Flek,
Rafet Sifa,
Nicolas Flores-Herr,
Joachim Köhler,
Patrick Schramowski,
Michael Fromm,
Kristian Kersting
},
year = {2025},
journal = {arXiv preprint arXiv:2505:22232}
}
