ksolovev/fine-news
FineNews FineNews is a multilingual news-text dataset for language-model research. It filters and deduplicates a 2021–2025 snapshot of the INFINI-NEWS Corpus, which extracts articles from Common Crawl CC-News. At a glance Measure FineNews Input articles 852,824,802 Infini-News rows from 2021–2025 Output 392,627,654 physical rows Files 294,509 Parquet files Folders 60 publication months (2021-01 to 2025-12), then language Language folders 129… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news.
FineNews
FineNews is a multilingual news-text dataset for language-model research. It filters and deduplicates a 2021–2025 snapshot of the INFINI-NEWS Corpus, which extracts articles from Common Crawl CC-News.
At a glance
The output row count is not a unique-article count. FineNews repeats some retained articles to weight near-duplicate clusters for training. Repeated copies keep the same id; deduplicate by id if you need unique articles.
Schema
The three top-level columns are:
id is a source-file and row location, not the Infini-News warc_record_id. metadata.date and metadata.year_month refer to the extracted publication date. They are not verified first-availability timestamps. metadata.warc_filename identifies the source crawl file.
Quick start
Download one Parquet file:
from huggingface_hub import hf_hub_download
import pyarrow.parquet as pq
path = hf_hub_download(
repo_id="ksolovev/fine-news",
filename="2025-12/en/000_00000.parquet",
repo_type="dataset",
)
table = pq.read_table(path)Dataset creation
Source data
The INFINI-NEWS Corpus provides article text extracted from CC-News WARC files. FineNews uses an earlier 2021–2025 snapshot of that corpus. The source corpus card describes its WARC provenance, extraction, and language metadata.
Processing pipeline
The pipeline applies global URL deduplication, language-aware quality filtering, exact deduplication within publication months, and MinHash near-deduplication within publication-month/language groups. It then repeats selected rows to weight near-duplicate clusters for training.
Quality thresholds come from FineWeb2, with a 1.15× relaxation of repetition thresholds for news. The pipeline uses a 100-word minimum and 15,000-word maximum. It uses GlotLID to fill missing language labels.
The output count is 53.96% below the input row count, including repeated rows. Japanese retention was 39.6%, below the pipeline's 40% target.
Limitations
- Publication-month folders do not establish when article text became available. In one checked 22,137-row file, 941 extracted publication dates fell after their WARC file dates. That single-file finding is not a corpus-wide rate.
- News coverage, language labeling, and filtering vary by language, publisher, and time. Retention alone does not establish quality or representativeness.
- Article text can contain personal information, errors, and harmful content.
- Parquet files can differ in the inferred type of an all-null nested metadata field. A bulk reader may need to normalize schemas.
Terms of use
FineNews is publicly downloadable. Article text remains subject to its publishers' copyright; no dataset-wide text license is assigned. Public access does not grant rights to the underlying article text. Users must assess their intended use and applicable law. See the source corpus terms for the Infini-News release.
Citation
If you use FineNews, cite this dataset page and the INFINI-NEWS Corpus.
