datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ubuntu_irc
Ubuntu IRC
Description
Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain.
We downloaded all chats from all channels up until March of 2025.
We consider all messages for given channel on a given day as a single document.
We removed system messages as well as those from known bots.
Dataset Statistics
Documents
UTF-8 GB
329,115
6.3
License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc.ubuntu_irc_filtered
Ubuntu IRC
Description
Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots.
Dataset Statistics
Documents
UTF-8 GB
234,982
5.3
License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.ibc-irc-code-edition-adoption-by-state
Which edition of the International Building Code, Residential Code and 15 other I-Codes the ICC's adoption chart of January 2024 records for each US state
Canonical, always-current version: https://referencesource.org/ibc-irc-code-edition-adoption-by-state/
Machine-readable: https://referencesource.org/ibc-irc-code-edition-adoption-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-09-30
Stale after: 2027-03-29 (past this date, prefer the canonical… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ibc-irc-code-edition-adoption-by-state.FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-details
Dataset Card for Evaluation run of FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs
Dataset automatically created during the evaluation run of model FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-details.ustax-irc-qa-36k
US Federal Tax Law QA Dataset (IRC — 36K pairs)
Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC),
used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v1.
Generation Pipeline
IRC full text stored in a Qdrant vector store (chunked at ~512 tokens)
An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk
Generated pairs are deduplicated and split into train/validation
Statistics
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-36k.IRCAM-Website-Dataustax-irc-qa-89k
US Federal Tax Law QA Dataset (IRC — 36K pairs)
Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC),
used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v2.
Generation Pipeline
IRC full text stored in a Qdrant vector store (chunked at ~512 tokens)
An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk
Generated pairs are deduplicated and split into train/validation
Statistics
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-89k.pile_ubuntu-ircsynthetic-irc-data
Synthetic IRC Conversation Dataset
Dataset Description
This dataset contains 1,500 synthetic IRC-style conversations featuring multiple participants, including an AI character named Em. The conversations were generated to replicate authentic IRC chat dynamics with natural flow, interruptions, and varied engagement levels.
Dataset Summary
Total conversations: 1,500
Total size: ~10MB
Format: JSONL with IRC-style formatting
Language: English
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/synthetic-irc-data.IR_combined_filtered
