datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vineyard-pruning
Dataset Card for Vineyard Dataset for Pruning
This is a FiftyOne dataset with 536 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/vineyard-pruning")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/vineyard-pruning.vineyard_pruning_segmentation
Vineyard Pruning Detection
A segmentation dataset for determining where to prune vineyards. The dataset contains 536 images with with pixel-level mask annotations across 3 categories: trunk, shoot, and pruned shoot.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{pacioni2025vineyard,
title={Vineyard dataset for automatic pruning based on main parts localization},
author={Pacioni, Elia and… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/vineyard_pruning_segmentation.SFT-for-Pruningfairness-pruning-pairs-en
Fairness Pruning Prompt Pairs — English
Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis.
This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs.
Dataset Summary
Each record contains a pair of prompts that are identical except for a single… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-en.fairness-pruning-pairs-es
Fairness Pruning Prompt Pairs — Spanish
Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis, with a focus on Spanish-language bias patterns.
This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs. It is the Spanish companion to the English dataset, enabling… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-es.bge-retrieval-data-ivf-pruning-438Kbge-retrieval-data-ivf-cluster-pruning-438Kbge-retrieval-data-ivf-passage-pruning-438Kbge-retrieval-data-ivf-query-pruning-fixed-438Kbge-retrieval-data-ivf-pruning-200Kbge-retrieval-data-ivf-cluster-pruning-200KCanopyLedger-Pruning-Observations
CanopyLedger Pruning Observations
This dataset card describes normalized pruning observations assembled from approved arborist field records.
Observation lineage register
Material class
Decision
Record source
License code
reuse_weight
authorized_on
pruning observation
cleared
Neighborhood Arbor Log
ODC-BY-1.0
9
2025-01-14
pruning observation
cleared
Crown Care Notebook
CC-BY-4.0
9
2025-04-06
pruning observation
review
Seasonal Crew Ledger
CC0-1.0
12… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/CanopyLedger-Pruning-Observations.bge-retrieval-data-ivf-query-pruning-fixed-200Kfr-wiki-popular-200-tokenizer-pruning
French Wikipedia corpus for tokenizer pruning
Prepared by Daniil Koblov. Contains 200 complete plain-text article extracts:
180 training articles and 20 held-out articles, split with Python's random seed 1337.
Candidates come from the 2025 monthly French Wikipedia top-1000 pageview lists.
Articles must appear in at least three months. Ranking uses month recurrence,
then the sum of reciprocal monthly ranks. The first 200 qualifying articles are
shuffled and split. Non-article… See the full description on the dataset page: https://huggingface.co/datasets/dakoblov/fr-wiki-popular-200-tokenizer-pruning.bge-retrieval-data-ivf-passage-pruning-100Kstablebridge-pruning-eval
Stablebridge Pruning Evaluation Dataset
Evaluation dataset for the Stablebridge context pruner/highlighter model, measuring sentence-level pruning quality on US stablecoin regulatory documents.
Dataset Structure
File
Records
Description
queries.jsonl
93
Regulatory queries (JSONL with _id and text fields)
corpus.jsonl
38
US stablecoin regulatory documents (full text)
qrels/test.tsv
2,704
Query-document relevance judgments
pruning_labels/test.jsonl
10,006… See the full description on the dataset page: https://huggingface.co/datasets/sugiv/stablebridge-pruning-eval.bge-retrieval-data-ivf-pruning-100Kbge-retrieval-data-ivf-query-pruning-fixed-100Ktruthy-dpo-v0.1-pruning-test-resultNot a serious or useable dataset, just used this https://github.com/Pleias/marginalia library and edited this notebook to check the quality of the selections in the jondurbin/truthy-dpo-v0.1 dataset
reasoning-pruning-pt-math-25-gemma4-r2pruning_calibrationreasoning-pruning-pt-bbh-logical-25-gemma4-r2bge-retrieval-data-ivf-passage-pruning-200Kbge-retrieval-data-ivf-pruning-50Kbge-retrieval-data-ivf-passage-pruning-fixed-100Kamortized-neuron-pruning-effects
Amortized Neuron-Ablation Effects (for pruning)
Exact mean-ablation effects of individual MLP neurons in transformer LMs, paired
with cheap forward/backward signals, for the task of amortized causal
intervention-effect prediction at neuron granularity (pruning).
The intervention unit is a single MLP neuron (blocks.{l}.mlp.hook_post channel).
For every (prompt, neuron) we record the exact effect of replacing that neuron's
last-token activation with its dataset mean, m(x… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/amortized-neuron-pruning-effects.Dataset-Pruninglang-pruning-uk-uber-text-2pruning-datasetcircuit-pruning-scores
