Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /ubuntu_irc Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 329,115 6.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc.texttext-generation100K<n<1M0 likes787 downloads1y agoHugging Face02common-pile /ubuntu_irc_filtered Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 234,982 5.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.texttext-generation100K<n<1M2 likes484 downloads1y agoHugging Face03Tamazight-NLP /IRCAM-CORPUS Dataset Card for IRCAM Corpus A text corpus containing texts written in various Tamazight dialects of Morocco published by IRCAM (Institut Royal de la Culture Amazighe). Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/IRCAM-CORPUS.texttext-generationn<1K1 likes332 downloads3y agoHugging Face04jkkummerfeld /irc_disentangle Dataset Card for IRC Disentanglement Dataset Summary Disentangling conversations mixed together in a single stream of messages is a difficult task, made harder by the lack of large manually annotated datasets. This new dataset of 77,563 messages manually annotated with reply-structure graphs that both disentangle conversations and define internal conversation structure. The dataset is 16 times larger than all previously released datasets combined, the first to include… See the full description on the dataset page: https://huggingface.co/datasets/jkkummerfeld/irc_disentangle.texttoken-classification100K<n<1M7 likes315 downloads2y agoHugging Face05average-developer /stocks-IRCON-1D-candlesn<1K0 likes235 downloads21h agoHugging Face06average-developer /stocks-IRCTC-1D-candlesn<1K0 likes234 downloads21h agoHugging Face07timaeus /pile-ubuntu_irc-broken ⚠️ Warning: This dataset will probably make you run out of memory if you try loading it. Don't do it. Dataset Creation Process These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered. Citations If you use this dataset, please cite the original Pile papers: @article{gao2020pile, title={The Pile: An 800GB dataset of diverse… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-ubuntu_irc-broken.text10K<n<100K0 likes214 downloads1y agoHugging Face08ircb /IRCB-26 IRCB '26 Datasets This repository hosts the datasets for the Inter-Regional Communication Benchmark (IRCB): a standardized benchmark suite for evaluating inter-regional communication in multi-regional models. Code and benchmark implementation: https://github.com/team-ircb/IRCB-26 The datasets are organized by task and observation model: delayed_memory/ gaussian/{train,val,test}.h5 poisson/{train,val,test}.h5 relay_decision/ gaussian/{train,val,test}.h5 poisson/{train… See the full description on the dataset page: https://huggingface.co/datasets/ircb/IRCB-26.0 likes204 downloads14d agoHugging Face09pttrn-io /ubuntu-irc-days Ubuntu IRC channel-days Five years of two Ubuntu IRC channels, one row per channel-day, used as the showcase corpus for the MADS Data Analysis and Visualisation course. What one row is One row is one channel on one day. The date is a column; the individual message times sit inside the text, one message per line: column type meaning created datetime64[ms] the day channel string #ubuntu-uk or #ubuntu-nl text string every message that day, [HH:MM]… See the full description on the dataset page: https://huggingface.co/datasets/pttrn-io/ubuntu-irc-days.texttext-classification1K<n<10K0 likes155 downloads2mo agoHugging Face10referencesource /ibc-irc-code-edition-adoption-by-state Which edition of the International Building Code, Residential Code and 15 other I-Codes the ICC's adoption chart of January 2024 records for each US state Canonical, always-current version: https://referencesource.org/ibc-irc-code-edition-adoption-by-state/ Machine-readable: https://referencesource.org/ibc-irc-code-edition-adoption-by-state/data.json — this mirror is a point-in-time copy. Last verified: 2026-09-30 Stale after: 2027-03-29 (past this date, prefer the canonical… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ibc-irc-code-edition-adoption-by-state.textn<1K0 likes79 downloads3d agoHugging Face11mathboylinlin /T1x-IRC-8K Dataset Card: T1x-IRC-8K Repository: huggingface.co/datasets/mathboylinlin/T1x-IRC-8K Status: repository created; data upload pending until manuscript acceptance. Summary T1x-IRC-8K is a dataset of intrinsic reaction coordinate (IRC) trajectories used to train and evaluate the MARC-TS transition-state prediction pipeline. reactions 8,209 IRC geometries 1,088,725 Locked split (to be released with the dataset) train 7… See the full description on the dataset page: https://huggingface.co/datasets/mathboylinlin/T1x-IRC-8K.0 likes38 downloads21d agoHugging Face12open-llm-leaderboard /FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-detailsgated Dataset Card for Evaluation run of FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs Dataset automatically created during the evaluation run of model FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face13AdaptKey /ustax-irc-qa-36k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v1. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-36k.textquestion-answering10K<n<100K0 likes31 downloads7mo agoHugging Face14IR-Cocktail /nq Data Description Homepage: https://github.com/KID-22/Cocktail Repository: https://github.com/KID-22/Cocktail Paper: [Needs More Information] Dataset Summary All the 16 benchmarked datasets in Cocktail are listed in the following table. Dataset Raw Website Cocktail Website Cocktail-Name md5 for Processed Data Domain Relevancy # Test Query # Corpus MS MARCO Homepage Homepage msmarco 985926f3e906fadf0dc6249f23ed850f Misc. Binary 6,979 542,203 DL19 Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/nq.0 likes27 downloads2y agoHugging Face15hbouqssi /ircam-dglai IRCAM DGLAI — Amazigh Dictionary Dataset 18,420 structured entries from the IRCAM Dictionnaire Général de la Langue Amazighe Informatisé (DGLAI). Fields Field Description id IRCAM entry ID tifinagh Word in Tifinagh script latin Word in Latin romanization (ALA-LC / IRCAM standard) grammar Grammar type (nom masculin, verbe, adjectif...) meaning_fr French meaning meaning_ar Arabic meaning construct_state Alternate forms when available lang… See the full description on the dataset page: https://huggingface.co/datasets/hbouqssi/ircam-dglai.1 likes27 downloads3mo agoHugging Face16IR-Cocktail /scidocs Data Description Homepage: https://github.com/KID-22/Cocktail Repository: https://github.com/KID-22/Cocktail Paper: [Needs More Information] Dataset Summary All the 16 benchmarked datasets in Cocktail are listed in the following table. Dataset Raw Website Cocktail Website Cocktail-Name md5 for Processed Data Domain Relevancy # Test Query # Corpus MS MARCO Homepage Homepage msmarco 985926f3e906fadf0dc6249f23ed850f Misc. Binary 6,979 542,203 DL19 Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/scidocs.0 likes23 downloads2y agoHugging Face17abdelhaqueidali /IRCAM-Website-Datatexttranslationn<1K0 likes22 downloads4mo agoHugging Face18AdaptKey /ustax-irc-qa-89k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v2. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-89k.textquestion-answering10K<n<100K0 likes21 downloads7mo agoHugging Face19nkandpa2 /ubuntu_irc_dates100K<n<1M0 likes20 downloads1y agoHugging Face20suolyer /pile_ubuntu-irctextn<1K0 likes18 downloads4y agoHugging Face21wnkh /IRCInterpres RAG Collection (IRC) contains the Medieval Latin dictionaries like Medieval Latin Word Vocabulary, William Whitaker's Words, Lexicon Abbreviaturarum, and the English to Latin Dictionary. Together with parallel copara include grosenthal/latin_english_parallel and magistermilitum/tridis_translate_la_en the vector database for RAG system of Interpres is formed. text10K<n<100K0 likes18 downloads6mo agoHugging Face22timyangyazhou /ubuntu_irc_kummerfeld_ft_20_window_last Dataset Card for "ubuntu_irc_kummerfeld_ft_20_window_last" More Information needed text10K<n<100K0 likes17 downloads3y agoHugging Face23IR-Cocktail /dl19 Data Description Homepage: https://github.com/KID-22/Cocktail Repository: https://github.com/KID-22/Cocktail Paper: [Needs More Information] Dataset Summary All the 16 benchmarked datasets in Cocktail are listed in the following table. Dataset Raw Website Cocktail Website Cocktail-Name md5 for Processed Data Domain Relevancy # Test Query # Corpus MS MARCO Homepage Homepage msmarco 985926f3e906fadf0dc6249f23ed850f Misc. Binary 6,979 542,203 DL19 Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/dl19.0 likes17 downloads2y agoHugging Face24IR-Cocktail /fiqa Data Description Homepage: https://github.com/KID-22/Cocktail Repository: https://github.com/KID-22/Cocktail Paper: [Needs More Information] Dataset Summary All the 16 benchmarked datasets in Cocktail are listed in the following table. Dataset Raw Website Cocktail Website Cocktail-Name md5 for Processed Data Domain Relevancy # Test Query # Corpus MS MARCO Homepage Homepage msmarco 985926f3e906fadf0dc6249f23ed850f Misc. Binary 6,979 542,203 DL19 Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/fiqa.0 likes15 downloads2y agoHugging Face25Brendan /dda-annotatable-ubuntu-irctabular100K<n<1M1 likes15 downloads2y agoHugging Face26IR-Cocktail /fever Data Description Homepage: https://github.com/KID-22/Cocktail Repository: https://github.com/KID-22/Cocktail Paper: [Needs More Information] Dataset Summary All the 16 benchmarked datasets in Cocktail are listed in the following table. Dataset Raw Website Cocktail Website Cocktail-Name md5 for Processed Data Domain Relevancy # Test Query # Corpus MS MARCO Homepage Homepage msmarco 985926f3e906fadf0dc6249f23ed850f Misc. Binary 6,979 542,203 DL19 Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/fever.0 likes14 downloads2y agoHugging Face27IR-Cocktail /trec-covid Data Description Homepage: https://github.com/KID-22/Cocktail Repository: https://github.com/KID-22/Cocktail Paper: [Needs More Information] Dataset Summary All the 16 benchmarked datasets in Cocktail are listed in the following table. Dataset Raw Website Cocktail Website Cocktail-Name md5 for Processed Data Domain Relevancy # Test Query # Corpus MS MARCO Homepage Homepage msmarco 985926f3e906fadf0dc6249f23ed850f Misc. Binary 6,979 542,203 DL19 Homepage… See the full description on the dataset page: https://huggingface.co/datasets/IR-Cocktail/trec-covid.0 likes13 downloads2y agoHugging Face28parameterlab /scaling_mia_the_pile_00_Ubuntu_IRCtextn<1K0 likes13 downloads2y agoHugging Face29david-ar /synthetic-irc-data Synthetic IRC Conversation Dataset Dataset Description This dataset contains 1,500 synthetic IRC-style conversations featuring multiple participants, including an AI character named Em. The conversations were generated to replicate authentic IRC chat dynamics with natural flow, interruptions, and varied engagement levels. Dataset Summary Total conversations: 1,500 Total size: ~10MB Format: JSONL with IRC-style formatting Language: English License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/synthetic-irc-data.texttext-generation1K<n<10K2 likes13 downloads1y agoHugging Face30french-open-data /employeurs-ircantec Employeurs IRCANTEC [!NOTE] Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Employeurs IRCANTEC qui est disponible à l'adresse https://www.data.gouv.fr/datasets/585bac6088ee3878fa3f4e5d Description Institution de Retraite Complémentaire des Agents Non Titulaires de l’Etat et des Collectivités publiques. Le régime de l’Ircantec s’applique d’une part aux salariés des employeurs relevant du champ d’application, d’autre… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/employeurs-ircantec.0 likes13 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.