Team Ai
Datasetpublic

surogate/fineweb2-ro-bert

FineWeb2-Ro-BERT FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here. Key Features Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders. Usage You… See the full description on the dataset page: https://huggingface.co/datasets/surogate/fineweb2-ro-bert.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes459downloads
Dataset Card

FineWeb2-Ro-BERT

FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here.

Key Features

  • —Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders.

Usage

You can load this dataset using the Hugging Face datasets library:

python
from datasets import load_dataset

dataset = load_dataset("OpenLLM-Ro/fineweb2-ro-bert", split="train")