Team Ai
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /ubuntu_irc Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 329,115 6.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc.texttext-generation100K<n<1M0 likes787 downloads1y agoHugging Face02common-pile /ubuntu_irc_filtered Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 234,982 5.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.texttext-generation100K<n<1M2 likes484 downloads1y agoHugging Face03Tamazight-NLP /IRCAM-CORPUS Dataset Card for IRCAM Corpus A text corpus containing texts written in various Tamazight dialects of Morocco published by IRCAM (Institut Royal de la Culture Amazighe). Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/IRCAM-CORPUS.texttext-generationn<1K1 likes332 downloads3y agoHugging Face04AdaptKey /ustax-irc-qa-36k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v1. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-36k.textquestion-answering10K<n<100K0 likes31 downloads7mo agoHugging Face05abdelhaqueidali /IRCAM-Website-Datatexttranslationn<1K0 likes22 downloads4mo agoHugging Face06AdaptKey /ustax-irc-qa-89k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v2. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-89k.textquestion-answering10K<n<100K0 likes21 downloads7mo agoHugging Face07david-ar /synthetic-irc-data Synthetic IRC Conversation Dataset Dataset Description This dataset contains 1,500 synthetic IRC-style conversations featuring multiple participants, including an AI character named Em. The conversations were generated to replicate authentic IRC chat dynamics with natural flow, interruptions, and varied engagement levels. Dataset Summary Total conversations: 1,500 Total size: ~10MB Format: JSONL with IRC-style formatting Language: English License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/synthetic-irc-data.texttext-generation1K<n<10K2 likes13 downloads1y agoHugging Face08abdelhaqueidali /ircam-dglai-dataset Dataset Card for DGLAi-Augmented An augmented, quad-lingual electronic parallel dataset based on the Dictionnaire Général de la Langue Amazighe (DGLAi). The original standard source fields (Amazigh, Arabic, French) are curated by IRCAM, supplemented with automated English translations to maximize its utility for modern machine translation, cross-lingual NLP, and LLM applications. Dataset Details Dataset Description The original Dictionnaire Général… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/ircam-dglai-dataset.tabulartranslation10K<n<100K1 likes8 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.