datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conll2003-generative
Dataset Card for CoNLL-2003 with NER Workflow Enhancements
This dataset is a modified version of the CoNLL-2003 dataset, enhanced to support an LLM-based Named Entity Recognition (NER) workflow. Two new columns, sentence and entities, have been added. You can find the code used to generate this version together with the data files.
Named Entities
As in the original CoNLL-2003 task, this dataset focuses on four types of named entities:
Persons (PER)
Locations (LOC)… See the full description on the dataset page: https://huggingface.co/datasets/areias/conll2003-generative.fineweb-edu-ar
FineWeb-Edu-Ar
FineWeb-Edu-Ar is a machine-translated Arabic version of the FineWeb-Edu dataset designed to support the development of Arabic small language models (SLMs).
Dataset Details:
Languages: Arabic, English (paired)
Size: 202 billion tokens
License: CC-BY-NC-4.0
Source: Machine-translated from the deduplicated version of Hugging Face’s FineWeb-Edu dataset
Translation model: facebook/nllb-200-distilled-600M
Application:
FineWeb-Edu-Ar is suitable for pre-training… See the full description on the dataset page: https://huggingface.co/datasets/kaust-generative-ai/fineweb-edu-ar.Galician-Generative-Eval-Prompts
Galician Generative Evaluation Prompts
Dataset Summary
This dataset contains a stratified sample of 60 input texts (prompts) for continuation and their respective original references, selected to cover a variety of linguistic registers and domains. It is designed to be executed within the Galician Rapid Human Evaluation Framework (galician-rapid-human-gen-eval) for LLMs in Galician.
Its primary purpose is to act as an input for the evaluation of model's open-ended… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Galician-Generative-Eval-Prompts.generative-rlhf-training-data
Generative RLHF-V Online Interaction Training Data
This repository contains the exact 4,053-example multimodal preference mixture
used by our online interaction training runs. The train split preserves the
original training order through training_index.
Composition
Source
Examples
PKU-Alignment/align-anything
2,048
PKU-Alignment/BeaverTails-V
2,005
Total
4,053
The Align-Anything subset was sampled from
text-image-to-text/new/train_40k.parquet… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyZou/generative-rlhf-training-data.
