CorefUD
Datasets
All datasets matching “CorefUD”corefud-1-4
corefud-1-4
Dataset Summary
This repository provides a standardized and reformatted version of the
original corefud-1-4 coreference resolution dataset.
The purpose of this formatting is to provide a unified document structure
across multiple coreference datasets in order to simplify:
cross-dataset comparison,
multilingual experimentation,
benchmarking of coreference resolution systems,
interoperability between NLP pipelines,
and reproducible evaluation settings.… See the full description on the dataset page: https://huggingface.co/datasets/lattice-nlp/corefud-1-4.CorefUD_1.2
[!NOTE]
Dataset origin: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-5478
Description
CorefUD is a collection of previously existing datasets annotated with coreference, which we converted into a common annotation scheme. In total, CorefUD in its current version 1.2 consists of 25 datasets for 16 languages. The datasets are enriched with automatic morphological and syntactic annotations that are fully compliant with the standards of the Universal Dependencies… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/CorefUD_1.2.SummIt-CorefUD
Dataset Card for Summ-it++ CorefUD
Dataset Summary
Summ-it++ CorefUD is a standardized, harmonized version of a 10% subset of the Summ-it++ corpus (the second largest coreference corpus for Portuguese). Adapted for evaluation in Universal Dependencies / CorefUD format, it contains two distinct variants designed for benchmarking entity coreference resolution models and LLMs in Portuguese:
Summ-it++ Original Harmonizado: A format-harmonized version of the original… See the full description on the dataset page: https://huggingface.co/datasets/manu-pac/SummIt-CorefUD.
