cleanlab
Datasets
All datasets matching “cleanlab”FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/FinQA-hallucination-detection.stanford-politeness
Stanford Politeness Dataset
This dataset contains politeness classification data based on the Stanford Politeness Corpus for active learning and fine-tuning tasks.
Dataset Description
The dataset is organized into two main directories:
Active Learning
X_labeled_full.csv - Labeled examples
X_unlabeled.csv - Unlabeled examples for active learning
extra_annotations.npy - Additional annotation data
test.csv - Test set
Fine-tuning
train.csv - Training… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/stanford-politeness.insurance-claims-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here:
https://github.com/cleanlab/structured-output-benchmark/
fire-financial-ner-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here:
https://github.com/cleanlab/structured-output-benchmark/
pii-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here:
https://github.com/cleanlab/structured-output-benchmark/
cifar-10-subset
CIFAR-10 Subset
This dataset contains a subset of the CIFAR-10 image classification dataset.
Dataset Description
A curated subset of the CIFAR-10 dataset useful for demonstrating image classification and data quality techniques without requiring the full dataset.
Usage
# Download the dataset
wget https://huggingface.co/datasets/Cleanlab/cifar-10-subset/resolve/main/CIFAR-10-subset.zip
unzip CIFAR-10-subset.zip
from huggingface_hub import hf_hub_download
#… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/cifar-10-subset.
