irc
Datasets
All datasets matching “irc”ubuntu_irc
Ubuntu IRC
Description
Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain.
We downloaded all chats from all channels up until March of 2025.
We consider all messages for given channel on a given day as a single document.
We removed system messages as well as those from known bots.
Dataset Statistics
Documents
UTF-8 GB
329,115
6.3
License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc.ubuntu_irc_filtered
Ubuntu IRC
Description
Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots.
Dataset Statistics
Documents
UTF-8 GB
234,982
5.3
License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.IRCAM-CORPUS
Dataset Card for IRCAM Corpus
A text corpus containing texts written in various Tamazight dialects of Morocco published by IRCAM (Institut Royal de la Culture Amazighe).
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/IRCAM-CORPUS.irc_disentangle
Dataset Card for IRC Disentanglement
Dataset Summary
Disentangling conversations mixed together in a single stream of messages is a difficult task, made harder by the lack of large manually annotated datasets. This new dataset of 77,563 messages manually annotated with reply-structure graphs that both disentangle conversations and define internal conversation structure. The dataset is 16 times larger than all previously released datasets combined, the first to include… See the full description on the dataset page: https://huggingface.co/datasets/jkkummerfeld/irc_disentangle.stocks-IRCON-1D-candlesstocks-IRCTC-1D-candles
