Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MuhammadAnas1657 /Prompt_Injection_PIDStext100K<n<1M1 likes56 downloads1mo agoHugging Face02WAZOBIALABS /nigerian-pidgin-eval Wazobia Labs — Nigerian Pidgin Gold-Standard Evaluation Set Version: v0.3 — May 2026 Entries: 253 dual-verified gold-standard entries Builder: Wazobia Labs License: CC-BY-4.0 — commercial use permitted with attribution Contact: wazobialabs@gmail.com What This Is Every AI lab building for Nigerian Pidgin has the same unsolved problem: they cannot evaluate whether their model actually works. No gold-standard benchmark exists. No culturally verified test set. No sarcasm… See the full description on the dataset page: https://huggingface.co/datasets/WAZOBIALABS/nigerian-pidgin-eval.textn<1K1 likes38 downloads5mo agoHugging Face03taresco /piqa_yoruba_pidgin Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin Dataset Summary This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures. It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios. The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.textquestion-answeringn<1K2 likes25 downloads11mo agoHugging Face04Pidoxy /Blind_Spots_of_Frontier_Models Qwen3-0.6B-Base — Blind Spots Dataset Model Tested Qwen/Qwen3-0.6B-Base Type: Causal Language Model (base / pretraining only — not instruction-tuned) Parameters: 0.6B (0.44B non-embedding) Released: April–May 2025 by Alibaba Cloud's Qwen Team Context Length: 32,768 tokens How the Model Was Loaded The model was loaded in a Google Colab T4 GPU notebook using HuggingFace transformers >= 4.51.0(required because the qwen3 architecture key was added in… See the full description on the dataset page: https://huggingface.co/datasets/Pidoxy/Blind_Spots_of_Frontier_Models.texttext-generationn<1K0 likes20 downloads7mo agoHugging Face05coderGit /Eng-PidginBioData Eng-PidginBioData: English–Nigerian Pidgin Biology Translation Dataset This dataset is archived on Zenodo with DOI: https://doi.org/10.5281/zenodo.18888857 Dataset Summary Eng-PidginBioData is a domain-specific parallel corpus for English ↔ Nigerian Pidgin machine translation focused on biological and scientific texts. The dataset contains 2,300 sentence pairs extracted from open-source biological research papers and manually translated into Nigerian Pidgin. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/coderGit/Eng-PidginBioData.texttranslation1K<n<10K0 likes17 downloads7mo agoHugging Face06Charley890 /naija-pidgin-health-qa-rivers-2026textquestion-answeringn<1K0 likes16 downloads6mo agoHugging Face07japaneil /PID_Simulation_Resultstabular100M<n<1B0 likes12 downloads1y agoHugging Face08Guavacoderepo /gclm-pidgin-text-corpustext100K<n<1M1 likes7 downloads1y agoHugging Face09Dikko /Pidgintexttable-question-answering1K<n<10K0 likes5 downloads1y agoHugging Face10michaelpenaariet /PIdemogated Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [train] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/michaelpenaariet/PIdemo.textn<1K0 likes3 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.