Team Ai
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oopere /fairness-pruning-pairs-en Fairness Pruning Prompt Pairs — English Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis. This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs. Dataset Summary Each record contains a pair of prompts that are identical except for a single… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-en.texttext-classificationn<1K1 likes145 downloads2mo agoHugging Face02oopere /fairness-pruning-pairs-es Fairness Pruning Prompt Pairs — Spanish Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis, with a focus on Spanish-language bias patterns. This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs. It is the Spanish companion to the English dataset, enabling… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-es.texttext-classificationn<1K1 likes109 downloads2mo agoHugging Face03sugiv /stablebridge-pruning-eval Stablebridge Pruning Evaluation Dataset Evaluation dataset for the Stablebridge context pruner/highlighter model, measuring sentence-level pruning quality on US stablecoin regulatory documents. Dataset Structure File Records Description queries.jsonl 93 Regulatory queries (JSONL with _id and text fields) corpus.jsonl 38 US stablecoin regulatory documents (full text) qrels/test.tsv 2,704 Query-document relevance judgments pruning_labels/test.jsonl 10,006… See the full description on the dataset page: https://huggingface.co/datasets/sugiv/stablebridge-pruning-eval.tabulartext-classification10K<n<100K0 likes42 downloads7mo agoHugging Face04dakoblov /fr-wiki-popular-200-tokenizer-pruning French Wikipedia corpus for tokenizer pruning Prepared by Daniil Koblov. Contains 200 complete plain-text article extracts: 180 training articles and 20 held-out articles, split with Python's random seed 1337. Candidates come from the 2025 monthly French Wikipedia top-1000 pageview lists. Articles must appear in at least three months. Ranking uses month recurrence, then the sum of reciprocal monthly ranks. The first 200 qualifying articles are shuffled and split. Non-article… See the full description on the dataset page: https://huggingface.co/datasets/dakoblov/fr-wiki-popular-200-tokenizer-pruning.tabularn<1K0 likes39 downloads7d agoHugging Face05iliakabanov /token_pruningtextn<1K0 likes26 downloads2d agoHugging Face06sugiv /stablebridge-regulatory-pruning-evaltabular10K<n<100K0 likes10 downloads7mo agoHugging Face07ZachSun /visual_pruning_kltext10K<n<100K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.