Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /ubuntu_irc Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 329,115 6.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc.texttext-generation100K<n<1M0 likes787 downloads1y agoHugging Face02common-pile /ubuntu_irc_filtered Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 234,982 5.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.texttext-generation100K<n<1M2 likes484 downloads1y agoHugging Face03referencesource /ibc-irc-code-edition-adoption-by-state Which edition of the International Building Code, Residential Code and 15 other I-Codes the ICC's adoption chart of January 2024 records for each US state Canonical, always-current version: https://referencesource.org/ibc-irc-code-edition-adoption-by-state/ Machine-readable: https://referencesource.org/ibc-irc-code-edition-adoption-by-state/data.json — this mirror is a point-in-time copy. Last verified: 2026-09-30 Stale after: 2027-03-29 (past this date, prefer the canonical… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ibc-irc-code-edition-adoption-by-state.textn<1K0 likes79 downloads3d agoHugging Face04open-llm-leaderboard /FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-detailsgated Dataset Card for Evaluation run of FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs Dataset automatically created during the evaluation run of model FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face05AdaptKey /ustax-irc-qa-36k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v1. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-36k.textquestion-answering10K<n<100K0 likes31 downloads7mo agoHugging Face06abdelhaqueidali /IRCAM-Website-Datatexttranslationn<1K0 likes22 downloads4mo agoHugging Face07AdaptKey /ustax-irc-qa-89k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v2. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-89k.textquestion-answering10K<n<100K0 likes21 downloads7mo agoHugging Face08suolyer /pile_ubuntu-irctextn<1K0 likes18 downloads4y agoHugging Face09david-ar /synthetic-irc-data Synthetic IRC Conversation Dataset Dataset Description This dataset contains 1,500 synthetic IRC-style conversations featuring multiple participants, including an AI character named Em. The conversations were generated to replicate authentic IRC chat dynamics with natural flow, interruptions, and varied engagement levels. Dataset Summary Total conversations: 1,500 Total size: ~10MB Format: JSONL with IRC-style formatting Language: English License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/synthetic-irc-data.texttext-generation1K<n<10K2 likes13 downloads1y agoHugging Face10ktgiahieu /IR_combined_filteredtextn<1K0 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.