datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ubuntu_irc
Ubuntu IRC
Description
Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain.
We downloaded all chats from all channels up until March of 2025.
We consider all messages for given channel on a given day as a single document.
We removed system messages as well as those from known bots.
Dataset Statistics
Documents
UTF-8 GB
329,115
6.3
License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc.ubuntu_irc_filtered
Ubuntu IRC
Description
Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots.
Dataset Statistics
Documents
UTF-8 GB
234,982
5.3
License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.IRCAM-CORPUS
Dataset Card for IRCAM Corpus
A text corpus containing texts written in various Tamazight dialects of Morocco published by IRCAM (Institut Royal de la Culture Amazighe).
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/IRCAM-CORPUS.irc_disentangle
Dataset Card for IRC Disentanglement
Dataset Summary
Disentangling conversations mixed together in a single stream of messages is a difficult task, made harder by the lack of large manually annotated datasets. This new dataset of 77,563 messages manually annotated with reply-structure graphs that both disentangle conversations and define internal conversation structure. The dataset is 16 times larger than all previously released datasets combined, the first to include… See the full description on the dataset page: https://huggingface.co/datasets/jkkummerfeld/irc_disentangle.stocks-IRCON-1D-candlesstocks-IRCTC-1D-candlespile-ubuntu_irc-broken
⚠️ Warning: This dataset will probably make you run out of memory if you try loading it. Don't do it.
Dataset Creation Process
These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered.
Citations
If you use this dataset, please cite the original Pile papers:
@article{gao2020pile,
title={The Pile: An 800GB dataset of diverse… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-ubuntu_irc-broken.IRCB-26
IRCB '26 Datasets
This repository hosts the datasets for the Inter-Regional Communication Benchmark (IRCB): a standardized benchmark suite for evaluating inter-regional communication in multi-regional models.
Code and benchmark implementation:
https://github.com/team-ircb/IRCB-26
The datasets are organized by task and observation model:
delayed_memory/
gaussian/{train,val,test}.h5
poisson/{train,val,test}.h5
relay_decision/
gaussian/{train,val,test}.h5
poisson/{train… See the full description on the dataset page: https://huggingface.co/datasets/ircb/IRCB-26.ubuntu-irc-days
Ubuntu IRC channel-days
Five years of two Ubuntu IRC channels, one row per channel-day, used as the
showcase corpus for the MADS Data Analysis and Visualisation course.
What one row is
One row is one channel on one day. The date is a column; the individual
message times sit inside the text, one message per line:
column
type
meaning
created
datetime64[ms]
the day
channel
string
#ubuntu-uk or #ubuntu-nl
text
string
every message that day, [HH:MM]… See the full description on the dataset page: https://huggingface.co/datasets/pttrn-io/ubuntu-irc-days.ibc-irc-code-edition-adoption-by-state
Which edition of the International Building Code, Residential Code and 15 other I-Codes the ICC's adoption chart of January 2024 records for each US state
Canonical, always-current version: https://referencesource.org/ibc-irc-code-edition-adoption-by-state/
Machine-readable: https://referencesource.org/ibc-irc-code-edition-adoption-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-09-30
Stale after: 2027-03-29 (past this date, prefer the canonical… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ibc-irc-code-edition-adoption-by-state.T1x-IRC-8K
Dataset Card: T1x-IRC-8K
Repository: huggingface.co/datasets/mathboylinlin/T1x-IRC-8K
Status: repository created; data upload pending until manuscript
acceptance.
Summary
T1x-IRC-8K is a dataset of intrinsic reaction coordinate (IRC) trajectories
used to train and evaluate the MARC-TS transition-state prediction pipeline.
reactions 8,209
IRC geometries 1,088,725
Locked split (to be released with the dataset)
train 7… See the full description on the dataset page: https://huggingface.co/datasets/mathboylinlin/T1x-IRC-8K.FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-details
Dataset Card for Evaluation run of FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs
Dataset automatically created during the evaluation run of model FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-details.ustax-irc-qa-36k
US Federal Tax Law QA Dataset (IRC — 36K pairs)
Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC),
used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v1.
Generation Pipeline
IRC full text stored in a Qdrant vector store (chunked at ~512 tokens)
An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk
Generated pairs are deduplicated and split into train/validation
Statistics
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-36k.nq
Data Description
Homepage: https://github.com/KID-22/Cocktail
Repository: https://github.com/KID-22/Cocktail
Paper: [Needs More Information]
Dataset Summary
All the 16 benchmarked datasets in Cocktail are listed in the following table.
Dataset
Raw Website
Cocktail Website
Cocktail-Name
md5 for Processed Data
Domain
Relevancy
# Test Query
# Corpus
MS MARCO
Homepage
Homepage
msmarco
985926f3e906fadf0dc6249f23ed850f
Misc.
Binary
6,979
542,203
DL19
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/nq.ircam-dglai
IRCAM DGLAI — Amazigh Dictionary Dataset
18,420 structured entries from the IRCAM Dictionnaire Général
de la Langue Amazighe Informatisé (DGLAI).
Fields
Field
Description
id
IRCAM entry ID
tifinagh
Word in Tifinagh script
latin
Word in Latin romanization (ALA-LC / IRCAM standard)
grammar
Grammar type (nom masculin, verbe, adjectif...)
meaning_fr
French meaning
meaning_ar
Arabic meaning
construct_state
Alternate forms when available
lang… See the full description on the dataset page: https://huggingface.co/datasets/hbouqssi/ircam-dglai.scidocs
Data Description
Homepage: https://github.com/KID-22/Cocktail
Repository: https://github.com/KID-22/Cocktail
Paper: [Needs More Information]
Dataset Summary
All the 16 benchmarked datasets in Cocktail are listed in the following table.
Dataset
Raw Website
Cocktail Website
Cocktail-Name
md5 for Processed Data
Domain
Relevancy
# Test Query
# Corpus
MS MARCO
Homepage
Homepage
msmarco
985926f3e906fadf0dc6249f23ed850f
Misc.
Binary
6,979
542,203
DL19
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/scidocs.IRCAM-Website-Dataustax-irc-qa-89k
US Federal Tax Law QA Dataset (IRC — 36K pairs)
Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC),
used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v2.
Generation Pipeline
IRC full text stored in a Qdrant vector store (chunked at ~512 tokens)
An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk
Generated pairs are deduplicated and split into train/validation
Statistics
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-89k.ubuntu_irc_datespile_ubuntu-ircIRCInterpres RAG Collection (IRC) contains the Medieval Latin dictionaries like Medieval Latin Word Vocabulary, William Whitaker's Words, Lexicon Abbreviaturarum, and the English to Latin Dictionary. Together with parallel copara include grosenthal/latin_english_parallel and magistermilitum/tridis_translate_la_en the vector database for RAG system of Interpres is formed.
ubuntu_irc_kummerfeld_ft_20_window_last
Dataset Card for "ubuntu_irc_kummerfeld_ft_20_window_last"
More Information needed
dl19
Data Description
Homepage: https://github.com/KID-22/Cocktail
Repository: https://github.com/KID-22/Cocktail
Paper: [Needs More Information]
Dataset Summary
All the 16 benchmarked datasets in Cocktail are listed in the following table.
Dataset
Raw Website
Cocktail Website
Cocktail-Name
md5 for Processed Data
Domain
Relevancy
# Test Query
# Corpus
MS MARCO
Homepage
Homepage
msmarco
985926f3e906fadf0dc6249f23ed850f
Misc.
Binary
6,979
542,203
DL19
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/dl19.fiqa
Data Description
Homepage: https://github.com/KID-22/Cocktail
Repository: https://github.com/KID-22/Cocktail
Paper: [Needs More Information]
Dataset Summary
All the 16 benchmarked datasets in Cocktail are listed in the following table.
Dataset
Raw Website
Cocktail Website
Cocktail-Name
md5 for Processed Data
Domain
Relevancy
# Test Query
# Corpus
MS MARCO
Homepage
Homepage
msmarco
985926f3e906fadf0dc6249f23ed850f
Misc.
Binary
6,979
542,203
DL19
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/fiqa.dda-annotatable-ubuntu-ircfever
Data Description
Homepage: https://github.com/KID-22/Cocktail
Repository: https://github.com/KID-22/Cocktail
Paper: [Needs More Information]
Dataset Summary
All the 16 benchmarked datasets in Cocktail are listed in the following table.
Dataset
Raw Website
Cocktail Website
Cocktail-Name
md5 for Processed Data
Domain
Relevancy
# Test Query
# Corpus
MS MARCO
Homepage
Homepage
msmarco
985926f3e906fadf0dc6249f23ed850f
Misc.
Binary
6,979
542,203
DL19
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/fever.trec-covid
Data Description
Homepage: https://github.com/KID-22/Cocktail
Repository: https://github.com/KID-22/Cocktail
Paper: [Needs More Information]
Dataset Summary
All the 16 benchmarked datasets in Cocktail are listed in the following table.
Dataset
Raw Website
Cocktail Website
Cocktail-Name
md5 for Processed Data
Domain
Relevancy
# Test Query
# Corpus
MS MARCO
Homepage
Homepage
msmarco
985926f3e906fadf0dc6249f23ed850f
Misc.
Binary
6,979
542,203
DL19
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/trec-covid.scaling_mia_the_pile_00_Ubuntu_IRCsynthetic-irc-data
Synthetic IRC Conversation Dataset
Dataset Description
This dataset contains 1,500 synthetic IRC-style conversations featuring multiple participants, including an AI character named Em. The conversations were generated to replicate authentic IRC chat dynamics with natural flow, interruptions, and varied engagement levels.
Dataset Summary
Total conversations: 1,500
Total size: ~10MB
Format: JSONL with IRC-style formatting
Language: English
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/synthetic-irc-data.employeurs-ircantec
Employeurs IRCANTEC
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Employeurs IRCANTEC qui est disponible à l'adresse https://www.data.gouv.fr/datasets/585bac6088ee3878fa3f4e5d
Description
Institution de Retraite Complémentaire des Agents Non Titulaires de l’Etat et des Collectivités publiques.
Le régime de l’Ircantec s’applique d’une part aux salariés des employeurs relevant du champ d’application, d’autre… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/employeurs-ircantec.
