datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lilm2-data-curationrukiga-yoruba-short-text-proverbs
Dataset Card for Rukiga-Yoruba Short Text and Proverbs Dataset
Dataset summary
This dataset pairs short original Rukiga (cgg) and Yoruba (yor) text —
everyday sentences, greetings, a market/introduction dialogue, and proverbs —
with English translations. It was created for a university data-curation
course assignment (Practical II: End-to-End Data Curation), by two student
authors each writing content in a language they speak natively. 78 entries
total: 63 Rukiga… See the full description on the dataset page: https://huggingface.co/datasets/rukiga-yoruba-datacuration/rukiga-yoruba-short-text-proverbs.patho-ssl-data-curation
Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology
Abstract Vision foundation models (FMs) are accelerating the devel- opment of digital pathology algorithms and transforming biomedical research. These models learn, in a self-supervised manner, to represent histological features in highly heterogeneous tiles extracted from whole-slide images (WSIs) of real-world patient samples. The performance of these FMs is significantly influenced by the size… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/patho-ssl-data-curation.Data-Curation-for-Visual-AI-Module-5-VisDrone
Dataset Card for Voxel51/VisDrone2019-DET
This is a FiftyOne dataset with 8629 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("dgural/Data-Curation-for-Visual-AI-Module-5-VisDrone")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/dgural/Data-Curation-for-Visual-AI-Module-5-VisDrone.Data-Curation-for-Visual-AI-Module-4-VisDrone
Dataset Card for 2024.10.06.22.04.02
This is a FiftyOne dataset with 8629 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("dgural/Data-Curation-for-Visual-AI-Module-4-VisDrone")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/dgural/Data-Curation-for-Visual-AI-Module-4-VisDrone.eagle-data-curation
Merged Dataset Standard Filtered
This folder contains the final training-ready dataset produced by the current standard filtering pipeline.
Files
merged_dataset.filtered.standard.back.jsonl: final filtered dataset, schema-consistent with the raw input
Filtering Strategy
The current pipeline uses the standard strategy defined in:
/home/dhz/eagle-data-curation/configs/process-open-perfectblend.standard.yaml
Applied operators and parameters:… See the full description on the dataset page: https://huggingface.co/datasets/aaa23123/eagle-data-curation.spike_sorting_curation_datadatacuration-verl6000Q_A3_curation_data_P3datacuration-verl-curatedatacuration-verl-filterdatacuration-pool-sample
