Team Ai
Datasetpublic

HealthDataHub/PARHAF-response_to_treatment-annotated

Dataset Card for PARHAF-response_to_treatment-annotated Reporting Issues & Contributing If you encounter any errors or inconsistencies in this dataset, please report them in the discussion section of the "Community" tab on Hugging Face. For more substantial contributions or collaboration opportunities, feel free to contact us directly. Dataset Summary PARHAF-response_to_treatment-annotated is a subpart of the PARHAF corpus, an open… See the full description on the dataset page: https://huggingface.co/datasets/HealthDataHub/PARHAF-response_to_treatment-annotated.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
0likes869downloads
Dataset Card

Dataset Card for PARHAF-responsetotreatment-annotated

<div align="center"> <p> <a href="https://huggingface.co/spaces/HealthDataHub/PARTAGES" style="display:inline;"><img src="img/PARTAGES BASELINE_RVB.png" alt="Logo" style="height:20px;vertical-align:middle; margin-right:8px; display:inline; margin-bottom:0.2em; margin-top:0.2em;" title="PARTAGES project"; /></a> <a href="https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf" style="display:inline;" title="Etalab 2.0 license"><img src="img/Logo-licence-ouverte2.svg" style="height:20px; display:inline; margin-bottom:0.2em; margin-top:0.2em;"/></a> <a href="https://creativecommons.org/licenses/by/4.0/deed.en" style="display:inline;"><img src="https://mirrors.creativecommons.org/presskit/buttons/88x31/png/by.png" style="height:20px; display:inline; margin-bottom:0.2em; margin-top:0.2em;" title="CC BY 4.0 license"/></a> </p> </div>

Reporting Issues & Contributing

If you encounter any errors or inconsistencies in this dataset, please report them in the discussion section of the "Community" tab on Hugging Face.

For more substantial contributions or collaboration opportunities, feel free to contact us directly.

Dataset Description

  • —Points of Contact: GILLES Flavien, KHALIL Youness

Dataset Summary

PARHAF-responsetotreatment-annotated is a subpart of the PARHAF corpus, an open French corpus of human-authored clinical reports of fictional patients.

It was created to support the development and evaluation of clinical NLP systems for the extraction and classification of treatment response levels from oncology consultation reports using RECIST-based categories..

This dataset contains training data only. The test set will remain under embargo to enable future evaluations under controlled conditions, limiting the risk of LLM contamination through prior data exposure. Please contact us for access to the test data.

This training dataset is divided into a train split (80%) and a dev split (20%) to facilitate experimental design and reproducibility across teams. Teams are free to use the full training set or define a different split configuration.

Each patient record was:

  • —written by a senior medical resident
  • —reviewed by another senior medical resident, from the same specialty
  • —annotated by a specialist of the use case
  • —curated by another specialist of the use case
Data statistics

DATASET SUMMARY

IndicatorValue
Complete Dataset
Number of JSON files108
Total annotations211
Average document length2573 characters
80% Threshold (annotations)168

Complete dataset distribution by type

TypeAnnotations% of total
DocumentMetadata10851.18%
Span10348.82%

Complete dataset distribution by populated field

Type.FieldOccurrences% of total annot.
DocumentMetadata.Nomenclature10851.18%
Span.Justification10348.82%

Train / Dev Split Results

IndicatorTRAINDEV
Number of files8523
Number of annotations16843
Percentage of dataset79.62%20.38%
Avg. document length (chars)25892515

Train / Dev distribution by type

TypeTRAINDEV% train of type
DocumentMetadata852378.70%
Span832080.58%

Train / Dev distribution by populated field

Type.FieldTRAINDEV% train of field
DocumentMetadata.Nomenclature852378.70%
Span.Justification832080.58%

Train / Dev distribution by value

Type.Field.ValueTRAINDEV% train
Span.Justification.Justification832080.58%
DocumentMetadata.Nomenclature.ReponsePartielle33684.62%
DocumentMetadata.Nomenclature.MaladieProgressive17385.00%
DocumentMetadata.Nomenclature.MaladieStable15575.00%
DocumentMetadata.Nomenclature.NonApplicable9660.00%
DocumentMetadata.Nomenclature.ReponseComplete8372.73%
DocumentMetadata.Nomenclature.NonDetermine30100.00%
Data Origin

The clinical reports are extracted from the PARHAF corpus. Please refer to PARHAF documentation for more information about this corpus.

Languages

  • —fr_FR

Dataset Structure

We distribute both a Hugging Face dataset and a standalone version of the corpus. The standalone dataset consists of a JSON file per patient report, in UIMA CAS JSON format. This format constitutes the canonical version of the corpus. The Hugging Face dataset (Parquet/Arrow) is a derived representation generated automatically from the JSON files.

Both formats therefore contain identical information and differ only in storage layout.

One dataset instance corresponds to one report.

Hugging Face dataset

This snippet shows how to extract and iterate over medical report information per patient using the datasets library.

python
import pandas as pd
from datasets import load_dataset

dfs = {cfg: load_dataset("HealthDataHub/PARHAF-response_to_treatment-annotated", cfg, split="train").to_pandas()
       for cfg in ["document_metadata", "spans"]}

for patient_raw in dfs["document_metadata"].itertuples():
    report_id = patient_raw.report
    text = patient_raw.full_text
    annot_type = patient_raw.annotation_type
    nomenclature = patient_raw.attribute_Nomenclature
    report_spans = dfs["spans"][dfs["spans"]["report"] == report_id]
    ...

Data Fields

PathTypeDescriptionPossible values
document_metadata
reportstringIdentifiant unique du rapport
full_textstringTexte intégral du rapport
spans
reportstringIdentifiant du rapport
span_idintegerIdentifiant de l'annotation
span_typestringType de l'entité annotéeSpan
beginintegerOffset de début
endintegerOffset de fin
span_textstringTexte de l'entité
attribute_JustificationstringJustificationJustification

Data Splits

Only the training set is released here. The remaining portion of the corpus will be temporarily embargoed to enable future evaluations under controlled conditions, thereby limiting the risk of large language model contamination through prior exposure to the data. You can evaluate your system on the test set through the CodaBench platform.

Annotation Guidelines

You can find the detailed annotation protocol here: annotation_guidelines.pdf

Licensing Information

This dataset is released under licenses:

  • —CC BY 4.0
  • —Etalab 2.0

Citation Information

[More Information Needed]