Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Polyglot-or-Not /Fact-Completion Dataset Card Homepage: https://bit.ly/ischool-berkeley-capstone Repository: https://github.com/daniel-furman/Capstone Point of Contact: daniel_furman@berkeley.edu Dataset Summary This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models. Test Description Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.texttext-generation100K<n<1M13 likes1k downloads3y agoHugging Face02ahmad21omar /Polyglot-Thoughts-SFT-Collection Polyglot Thoughts SFT Collection Polyglot Thoughts SFT Collection is a large-scale supervised fine-tuning (SFT) corpus for reasoning-oriented language models. It combines, filters, deduplicates, and language-extends a broad set of public reasoning datasets into a single uniform schema centred on chain-of-thought reasoning traces. The final corpus contains 23,896,757 examples and roughly 123 billion tokens, spanning six languages (English, German, French, Italian, Spanish… See the full description on the dataset page: https://huggingface.co/datasets/ahmad21omar/Polyglot-Thoughts-SFT-Collection.texttext-generation10M<n<100M0 likes949 downloads4mo agoHugging Face03ToxicityPrompts /PolygloToxicityPrompts PolygloToxicityPrompts Dataset Summary A multilingual toxicity evaluation benchmark curated from web text. We prepared 3 splits: ptp-full, ptp-small, and wildchat containining 25K, 5K and 1K prompts per language respectively. The wildchat split is created using AI2's WildChat dataset. How do I download this? Using 🤗 Datasets from datasets import load_dataset # English only dataset =… See the full description on the dataset page: https://huggingface.co/datasets/ToxicityPrompts/PolygloToxicityPrompts.text-generation100K<n<1M14 likes621 downloads5mo agoHugging Face04hac541309 /polyglot-ko-tokenizer-corpus Dataset Card for "polyglot-ko-tokenizer-corpus" More Information needed text10M<n<100M1 likes362 downloads3y agoHugging Face05rmyeid /polyglot_nerPolyglot-NER A training dataset automatically generated from Wikipedia and Freebase the task of named entity recognition. The dataset contains the basic Wikipedia based training data for 40 languages we have (with coreference resolution) for the task of named entity recognition. The details of the procedure of generating them is outlined in Section 3 of the paper (https://arxiv.org/abs/1410.3791). Each config contains the data corresponding to a different language. For example, "es" includes only spanish examples.token-classification40 likes360 downloads3y agoHugging Face06swiss-ai /polyglotoxicitypromptstabular100K<n<1M0 likes333 downloads1y agoHugging Face07DCAgent2 /aider_polyglot0 likes327 downloads11mo agoHugging Face08nathansutton /chad-polyglot-runs chad run data Run output from the benchmarks of chad, a local coding agent for Apple Silicon, kept out of the code repository. Every polyglot run here is a row of benchmarks/polyglot/RUNS.md with the sha256 of its trials.jsonl, and benchmarks/polyglot/fetch.py --label <label> downloads one and refuses it if the hash differs. path what polyglot/<label>/ one polyglot run: meta.json (chad version and commit, model, context limit, every CHAD_* variable), trials.jsonl (one… See the full description on the dataset page: https://huggingface.co/datasets/nathansutton/chad-polyglot-runs.0 likes270 downloads7d agoHugging Face09polyglot-tagger /wikipedia-language-snippets-filtered Wikipedia Snippets (Filtered) Filtered sentence snippets in Wikipedia, by taking the first 60% of an article after filtering for stubs. Minor Latin groups are additionally filtered again for English leakage. Sentences are mostly filtered out for non matching scripts, such as Arabic in a Cyrllic language. Files Each file is in this format for languages in ISO 639 2-letter codes: train/en/en.parquet train/es/es.parquet From wikimedia/wikipedia Licensing… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/wikipedia-language-snippets-filtered.texttext-generation10M<n<100M0 likes151 downloads6mo agoHugging Face10nassimjp /Polyglot-Restaurant-Reasoning-Dataset Polyglot-Restaurant-Reasoning-Dataset 🍽️🌐🤖 Dataset Summary The Polyglot-Restaurant-Reasoning-Dataset is a specialized instruction-tuning dataset engineered to train large language models (LLMs) to function as intelligent, context-aware, and reasoning-capable restaurant chatbot assistants. This dataset specifically focuses on empowering low-resource and regional languages (Pashto, Farsi/Dari, Urdu, Balochi, and Sindhi), alongside bridging support for English… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Polyglot-Restaurant-Reasoning-Dataset.texttext-generation1K<n<10K0 likes126 downloads13d agoHugging Face11polyglots /Punjabi-NERtext100K<n<1M1 likes117 downloads2y agoHugging Face12Kicaulah /polyglot-bug-patterns 🗃️ Polyglot Bug Patterns 🐝 Dataset · 🎨 Space · 🧠 Model · 💻 GitHub Bro, imagine a dataset of 43 non-destructive attack payloads across 8 bug classes — 24 of them polyglot, meaning a single string that's simultaneously plausible HTML, JS, SQL and shell, so it escapes whatever context your app dropped it into. Plus real detector output from a full multimodal scan. Everything here is synthesised or locally generated. No real-world traffic, no scraped data, no user data.… See the full description on the dataset page: https://huggingface.co/datasets/Kicaulah/polyglot-bug-patterns.tabulartext-classificationn<1K0 likes109 downloads9d agoHugging Face13ljvmiranda921 /PolyglotTeachers-SFT-Synth-Data Website: ljvmiranda921.github.io/polyglot-teachers/ PolyglotTeachers-SFT-Synth This dataset contains synthetic supervised fine-tuning examples generated by the best teacher we found in the paper Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation, where we systematically characterize what makes a good teacher model. It contains examples across six languages: Arabic, Czech, German, Indonesian, Japanese, Spanish, and Tagalog. Note: In… See the full description on the dataset page: https://huggingface.co/datasets/ljvmiranda921/PolyglotTeachers-SFT-Synth-Data.texttext-generation100K<n<1M3 likes97 downloads2mo agoHugging Face14ahmad21omar /Polyglot-Thoughts-RL-Collection Polyglot Thoughts RL Collection Polyglot Thoughts RL Collection is a large-scale, curated corpus for reinforcement learning from verifiable rewards (RLVR) of reasoning-oriented language models. It combines, filters, normalises, and deduplicates a broad set of public RL datasets into a single uniform schema in which every row carries a machine-verifiable ground-truth signal — math equivalence, code execution, Prolog rule induction, schema validation, multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/ahmad21omar/Polyglot-Thoughts-RL-Collection.texttext-generation1M<n<10M0 likes94 downloads4mo agoHugging Face15polyglot-tagger /finetranslations-filtered Files Each file is in this format for languages in ISO 639 2-letter codes: train/en/en.parquet train/es/es.parquet From HuggingFaceFW/finetranslations Licensing Information The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. The use of this dataset is also subject to CommonCrawl's Terms of Use. Citation Information @misc{penedo2026finetranslations, title={FineTranslations}, author={Guilherme… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/finetranslations-filtered.texttext-classification1M<n<10M0 likes81 downloads6mo agoHugging Face16Reubencf /PolyglotAudio Citation If you use this dataset in your research or downstream work, please cite: @misc{polyglot_audio_2026, author = {Fernandes, Reuben Chagas}, title = {PolyglotAudio: Multilingual Audio Pre-training Corpus}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/datasets/Reubencf/PolyglotAudio}} } APA-style: Reuben Chagas Fernandes (2026). PolyglotAudio: Multilingual Audio Pre-training Corpus [Dataset]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/PolyglotAudio.audio1M<n<10M2 likes72 downloads6mo agoHugging Face17nassimjp /Polyglot-Restaurant-Questions-Dataset 🍽️ Polyglot Restaurant Questions Dataset A multilingual dataset of 101 restaurant / catering order cancellation questions translated into 6 languages: English, Farsi (Persian), Japanese, Pashto, Sindhi, and Urdu. This dataset is designed for intent classification, customer-support chatbots, question-answering systems, and cross-lingual NLP research in the food-service domain. 🌍 Languages Code Language Script File Records en English Latin… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Polyglot-Restaurant-Questions-Dataset.textquestion-answeringn<1K0 likes71 downloads13d agoHugging Face18open-llm-leaderboard-old /details_EleutherAI__polyglot-ko-12.8b Dataset Card for Evaluation run of EleutherAI/polyglot-ko-12.8b Dataset Summary Dataset automatically created during the evaluation run of model EleutherAI/polyglot-ko-12.8b on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__polyglot-ko-12.8b.0 likes69 downloads3y agoHugging Face19polyglot-tagger /tatoeba-filtered Files Each file is in this format for languages in ISO 639 2-letter codes: train/en/en.parquet train/es/es.parquet texttext-classification1M<n<10M0 likes61 downloads6mo agoHugging Face20open-llm-leaderboard-old /details_beomi__KoAlpaca-Polyglot-5.8B Dataset Card for Evaluation run of beomi/KoAlpaca-Polyglot-5.8B Dataset Summary Dataset automatically created during the evaluation run of model beomi/KoAlpaca-Polyglot-5.8B on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_beomi__KoAlpaca-Polyglot-5.8B.0 likes55 downloads3y agoHugging Face21archnemix-ffmpeg-polyglot /DF-Realtime-mode0 likes53 downloads10d agoHugging Face2234data /polyglotfake-fakevideo1K<n<10K0 likes50 downloads7mo agoHugging Face2334data /polyglotfake-real0 likes50 downloads7mo agoHugging Face24darkknight25 /polyglot_paylods_datasets Polyglot Payloads Dataset for Cybersecurity Training Overview This dataset, polyglot_payloads.jsonl, is a curated collection of 500 polyglot payloads designed for training AI models in cybersecurity, specifically for red team operations and vulnerability detection. The dataset includes payloads targeting common web vulnerabilities such as Cross-Site Scripting (XSS), SQL Injection (SQLi), Local File Inclusion (LFI), Remote Code Execution (RCE), and Server-Side Template… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/polyglot_paylods_datasets.textn<1K0 likes49 downloads1y agoHugging Face25polyglot-tagger /nlp-noise-snippets Synthetic Noise Pool For Text Classification purposes, as many models may consider code snippets, html artifacts, and math as "English". Around 50K are latex snippets from im2latex-100k texttext-classification100K<n<1M0 likes46 downloads6mo agoHugging Face26DCAgent2 /DCAgent2_aider_polyglot_DCAgent_tbench_oracle_solutions_terminus_20260125_202032textn<1K0 likes41 downloads9mo agoHugging Face27hac541309 /polyglot-ko-tokenizer-corpus-merge_ws Dataset Card for "polyglot-ko-tokenizer-corpus-merge_ws" More Information needed text10M<n<100M0 likes36 downloads3y agoHugging Face28polyglots /MADLAD_CulturaX_cleanedtext10M<n<100M22 likes35 downloads2y agoHugging Face29open-llm-leaderboard-old /details_macadeliccc__polyglot-math-4x7b Dataset Card for Evaluation run of macadeliccc/polyglot-math-4x7b Dataset automatically created during the evaluation run of model macadeliccc/polyglot-math-4x7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_macadeliccc__polyglot-math-4x7b.0 likes33 downloads3y agoHugging Face30DCAgent2 /DCAgent2_aider_polyglot_DCAgent2_swesmith-stack-reason_20260127_005703text1K<n<10K0 likes33 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.