curation
mlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_metamath-ggufmlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_unnatural_instructions-ggufmlfoundations-dev_-_oh-dcft-v1.1-no-curation-ggufmlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_slim_orca-ggufmlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_evol_instruct-ggufmlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_airoboros-ggufoh-dcft-v1.3_no-curation_gpt-4o-mini_scale_2x-GGUFmlfoundations-dev_-_oh-dcft-v1-no-curation-gguf
Datasets
All datasets matching “curation”lilm2-data-curationlerobot-curation-metrics-figures
LeRobot curation metrics: blog figures and plugin screenshots
The ten figures from A Practical Guide to LeRobot Dataset Quality Metrics (a Hugging Face article) and its longer companion post, You Can't Watch 10,000 Episodes. Measure These Instead. Both are guides to the metrics worth computing on a LeRobot v3 dataset before you train a policy or fine-tune a VLA.
All signals are synthetic and illustrative. Every number printed on a figure is computed from its synthetic signal… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/lerobot-curation-metrics-figures.MNIST-Curation
Curation of the famous MNIST Dataset
The curation was done using qualitative analysis of the dataset, following visualization techniques like PCA and UMAP and score-based categorization of the samples using metrics like hardness, mistakenness, or uniqueness.
The code of the curation can be found on GitHub:👉 https://github.com/Conscht/MNIST_Curation_Repo/tree/main
This curated version of MNIST introduces an additional IDK (“I Don’t Know”) label for digits that are ambiguous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/Consscht/MNIST-Curation.oh-dcft-v1.3_no-curation_gpt-4o-mini_scale_4xqwen35-2b-tool-use-qwen36-27b-curation-candidates
Full candidate collections: 2B tool use + 27B data curation
This public Dataset contains two complete, unredacted, exact-40 candidate collections:
Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and
233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL,
Spider, and TravelPlanner.
Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021
targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.patho-ssl-data-curation
Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology
Abstract Vision foundation models (FMs) are accelerating the devel- opment of digital pathology algorithms and transforming biomedical research. These models learn, in a self-supervised manner, to represent histological features in highly heterogeneous tiles extracted from whole-slide images (WSIs) of real-world patient samples. The performance of these FMs is significantly influenced by the size… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/patho-ssl-data-curation.
