datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Prompt_Injection_PIDSnigerian-pidgin-eval
Wazobia Labs — Nigerian Pidgin Gold-Standard Evaluation Set
Version: v0.3 — May 2026
Entries: 253 dual-verified gold-standard entries
Builder: Wazobia Labs
License: CC-BY-4.0 — commercial use permitted with attribution
Contact: wazobialabs@gmail.com
What This Is
Every AI lab building for Nigerian Pidgin has the same unsolved problem: they cannot evaluate whether their model actually works.
No gold-standard benchmark exists. No culturally verified test set. No sarcasm… See the full description on the dataset page: https://huggingface.co/datasets/WAZOBIALABS/nigerian-pidgin-eval.piqa_yoruba_pidgin
Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin
Dataset Summary
This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures.
It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios.
The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.Blind_Spots_of_Frontier_Models
Qwen3-0.6B-Base — Blind Spots Dataset
Model Tested
Qwen/Qwen3-0.6B-Base
Type: Causal Language Model (base / pretraining only — not instruction-tuned)
Parameters: 0.6B (0.44B non-embedding)
Released: April–May 2025 by Alibaba Cloud's Qwen Team
Context Length: 32,768 tokens
How the Model Was Loaded
The model was loaded in a Google Colab T4 GPU notebook using HuggingFace transformers >= 4.51.0(required because the qwen3 architecture key was added in… See the full description on the dataset page: https://huggingface.co/datasets/Pidoxy/Blind_Spots_of_Frontier_Models.Eng-PidginBioData
Eng-PidginBioData: English–Nigerian Pidgin Biology Translation Dataset
This dataset is archived on Zenodo with DOI:
https://doi.org/10.5281/zenodo.18888857
Dataset Summary
Eng-PidginBioData is a domain-specific parallel corpus for English ↔ Nigerian Pidgin machine translation focused on biological and scientific texts. The dataset contains 2,300 sentence pairs extracted from open-source biological research papers and manually translated into Nigerian Pidgin.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/coderGit/Eng-PidginBioData.naija-pidgin-health-qa-rivers-2026PID_Simulation_Resultsgclm-pidgin-text-corpusPidginPIdemo
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[train]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/michaelpenaariet/PIdemo.
