datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi_Language_Audio2TextThis dataset by Mozilla Common Voice (https://commonvoice.mozilla.org/en/datasets) is crafted by Udyan Sachdev
Voice datasets play a pivotal role in training and evaluating speech-to-text models, influencing advancements in natural language processing. This dataset outlines the creation of a comprehensive text dataset from 40,571 MP3 audio files sourced from the Common Voice project. The dataset aims to serve as a benchmark for training and evaluating speech-to-text models in English, French… See the full description on the dataset page: https://huggingface.co/datasets/UdyanSachdev/Multi_Language_Audio2Text.toxicity_multilanguage_datasetmulti-open
African Languages Lab Multi-Open
multi-open is the open-source multilingual subset released by the
African Languages Lab. It contains English-target
parallel text for 31 African languages.
Project website: https://the-african-languages-lab.github.io/
The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African
NLPIssaka et al., ACL 2026.
The paper presents All Lab's broader collaborative program: systematic and quality-controlled
data infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/multi-open.synthetic-multilanguage-conversationtask1577_amazon_reviews_multi_japanese_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.mini-multilanguageNSFW_Multilanguage_Chat_DatasetThanks to utsavm/NSFW_Chat_Dataset, I translate it so it can be more useful.
🚨 18+ Only! NSFW & Spicy Content Ahead 🚨
Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you.
📜 What’s Inside?
This dataset features two columns:
input → Boyfriend’s dialogue (aka what YOU say 😉)… See the full description on the dataset page: https://huggingface.co/datasets/Raphael172/NSFW_Multilanguage_Chat_Dataset.code-gen-multi-languageamazon-reviews-multi-all-languagesmultilanguagetask1576_amazon_reviews_multi_english_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1576_amazon_reviews_multi_english_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1576_amazon_reviews_multi_english_language_classification.Zeroshot-multilanguages-2.0
Dataset Card for "Zeroshot-multilanguages-2.0"
More Information needed
LanguageQA
Dataset Card for SAKURA-LanguageQA
This dataset contains the audio and the single/multi-hop questions/answers of the language track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., the language spoken in the speech) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/LanguageQA.task1574_amazon_reviews_multi_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1574_amazon_reviews_multi_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1574_amazon_reviews_multi_language_identification.translation-multilanguage-v2flan_combined_task1574_amazon_reviews_multi_language_identificationZeroshot-multilanguages-2.1alpaca_multi_language_train_reasoning_simple
このデータセットは、有名なALPACAデータセットの一部を使ったデータセットです。
##日本語と英語(プラス中国語)の同じ情報が記載されています。
「Reasoning」というプロンプトを使えるようにしています。
Reasoningを使用することにより、予測精度を上げられるようにしました。
Reasoningの内容は、回答に使用すべき言語をしているだけです。
必要に応じて、ユーザーが変更してみてください。
「Category」という属性で、レコードが分類されています。
open_qa, closed_qa, classification
brainstorm, creative
translation, question_to_question
「Reasoning」を推論するためのレコードが追加されています。
詳しい情報はこちらのブログを参考にしてください。
multi_language_train0625
経緯
このデータセットはQEUプロジェクトのBONSAI2の学習のために開発されました。
特長
alpacaデータセットの一部がベースですが、大幅に変更されています。
言語: 英語、日本語、中国語
reasoningという情報が入っています。使わなくともかまいません。
open_qa, closed_qa, classification, evaluation, question to question
参考サイト
QEUR23_ CHRLTM14 : 閑話休題~Predibaseでfinetuneを使ってみる(SOLAR LLM)
flan_combined_task1576_amazon_reviews_multi_english_language_classificationarticle-20-multilanguage-bindings
Every hash chain startup is wrong about multi-language bindings for cryptographic ledgers.
Multi-Language Bindings for Cryptographic Ledgers: PyO3, NAPI-RS, and CGo Interoperability
The Problem
The AIOSS (AI Open Signed Storage) cryptographic ledger format must be accessible from multiple programming languages to serve diverse deployment environments, including Python for data science workflows, Go for infrastructure tooling, and JavaScript/Node.js for web-based… See the full description on the dataset page: https://huggingface.co/datasets/Anticloud/article-20-multilanguage-bindings.article-20-multilanguage-bindings
Every hash chain startup is wrong about multi-language bindings for cryptographic ledgers.
Multi-Language Bindings for Cryptographic Ledgers: PyO3, NAPI-RS, and CGo Interoperability
The Problem
The AIOSS (AI Open Signed Storage) cryptographic ledger format must be accessible from multiple programming languages to serve diverse deployment environments, including Python for data science workflows, Go for infrastructure tooling, and JavaScript/Node.js for web-based… See the full description on the dataset page: https://huggingface.co/datasets/kleinnner/article-20-multilanguage-bindings.supergoose_flan_combined_task1574_amazon_reviews_multi_language_identificationCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données supergoose/flan_combined_task1574_amazon_reviews_multi_language_identification.
Lots-of-LoRAs_task1574_amazon_reviews_multi_language_identificationCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données Lots-of-LoRAs/task1574_amazon_reviews_multi_language_identification.
Lots-of-LoRAs_task1576_amazon_reviews_multi_english_language_classificationCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données Lots-of-LoRAs/task1576_amazon_reviews_multi_english_language_classification.
multi-language-messages-01CodeTransOceanのMultilingualTransデータセットのsplit trainをopenAI messages形式に調整。
Lots-of-LoRAs_task1577_amazon_reviews_multi_japanese_language_classificationCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.
all-lab-text-multisupergoose_flan_combined_task1576_amazon_reviews_multi_english_language_classificationCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données supergoose/flan_combined_task1576_amazon_reviews_multi_english_language_classification.
Zeroshot-multilanguages
