Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ken-Z /Latin-Audio Dataset Summary Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training. Alignment and curation: Kaiyuan Zhao Language: Latin (Classical) Uses This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.audiotext-to-speech10K<n<100K9 likes2.9k downloads2mo agoHugging Face02PleIAs /Latin-PD 🇲🇪 Latin Public Domain Books (Latin) 🇲🇪 Latin-Public Domain or Latin-PD is a large collection aiming to aggregate all Latin monographies and periodicals in the public domain. As of June 2024, it is the largest Latin open corpus. Dataset summary The collection contains 16,521,454,086 words (159,070 titles) recovered from multiple sources, including the Internet Archive and various European national libraries and cultural heritage institutions (BDH, BNF). Each parquet… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Latin-PD.tabular100K<n<1M8 likes2.1k downloads2y agoHugging Face03Fece228 /latin-literature-dataset-170MThis is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin. The dataset is split in two parts: preprocessed with basic cltk tools, ready for work, and raw text data. It must be noted, however, that the latter contains text in Greek, Hebrew, and other languages, with references and contractions text100M<n<1B11 likes766 downloads4y agoHugging Face04LatinNLP /latin-summarizer-dataset ✨ LatinSummarizer Dataset ✨ Note: If Dataset Viewer is not available, see samples of dataset for samples from the dataset.The LatinSummarizer Dataset is a comprehensive collection of Latin texts designed to support natural language processing research for a low-resource language. It provides parallel data for various tasks, including translation (Latin-to-English) and summarization (extractive and abstractive). This dataset was created for a… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/latin-summarizer-dataset.summarization0 likes585 downloads1y agoHugging Face05TartarusXXX /uyghur-cv-latinaudio100K<n<1M1 likes508 downloads9mo agoHugging Face06LatinNLP /LatinSummarizer LatinSummarizer Dataset Structure aligned_en_la_data_raw.csv aligned_en_la_data_cleaned.csv aligned_en_la_data_cleaned_with_stanza.csv concat_aligned_data.csv concat_cleaned.csv latin_wikipedia_cleaned.csv latin_wikipedia_raw.csv latin-literature-dataset-170M_raw_cleaned.csv latin-literature-dataset-170M_raw_cleaned_chunked.csv Elsa_aligned/ README.md Details aligned_en_la_data_raw.csv This dataset contains aligned Latin (la) - English (en)… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/LatinSummarizer.texttranslation1M<n<10M0 likes439 downloads2y agoHugging Face07TomasGuija /LatinFontsSVGs SVG Font Dataset Overview We present a curated dataset of vector font glyphs stored as SVG files, designed for research in generative modeling, structured vector graphics and typography synthesis. The dataset was created for the development and evaluation of our paper: DesigNet: Learning to Draw Vector Graphics as Designers Do Related Resources Paper (arXiv) : https://arxiv.org/abs/2604.06494 Code: https://github.com/TomasGuija/DesigNet… See the full description on the dataset page: https://huggingface.co/datasets/TomasGuija/LatinFontsSVGs.tabular1M<n<10M0 likes396 downloads5mo agoHugging Face08KhalfounMehdi /arabic-latin-invoices-synthetic Synthetic Arabic/Latin Invoices — label-first multimodal dataset A large, perfectly-labeled synthetic invoice dataset for training multimodal invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin layouts, and all three numeral glyph systems. Every sample is generated label-first: a structured invoice record is sampled, then rendered to pixels via a headless browser, and bounding boxes are read back from the same DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.imageimage-to-text1K<n<10K1 likes177 downloads4mo agoHugging Face09grosenthal /latin_english_translation Dataset Card for "latin_english_parallel" 101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation. For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences. Each sample is annotated with the index and file (and therefore author/work) that the sample is from. If you find errors… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_translation.texttranslation100K<n<1M14 likes158 downloads3y agoHugging Face10julian-schelb /latin-classical-intertextuality-corpus Latin Classical Authors Corpus This dataset contains processed texts from classical Latin authors, serving as a retrieval corpus for intertextuality research. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Cicero and classical Latin literature. Related Datasets This corpus is part of the Latin Jerome Intertextuality collection: Queries: latin-classical-intertextuality-queries - The whole works of… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-corpus.texttext-retrieval10K<n<100K4 likes137 downloads1mo agoHugging Face11julian-schelb /latin-classical-intertextuality-labels Latin Jerome Intertextuality Labels This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.tabulartext-retrieval1K<n<10K2 likes130 downloads1mo agoHugging Face12reesjon9 /Latin-Audio Dataset Summary Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training. Alignment and curation: Kaiyuan Zhao Language: Latin (Classical) Uses This dataset is built for training and evaluating speech processing models for… See the full description on the dataset page: https://huggingface.co/datasets/reesjon9/Latin-Audio.audiotext-to-speech1K<n<10K0 likes125 downloads10mo agoHugging Face13itserr /WP8-Latin-Embeddings-Indices0 likes115 downloads2y agoHugging Face14ApyHTML19 /Medication_Boxes_Arabe_Latin Medication Boxes — Arabic / Latin This dataset contains photos of medication boxes, with packaging text in Arabic and Latin script (French and others). It is annotated for COCO instance segmentation and was built for MediSeG, which segments each box so the text on it can be read afterwards. One class: 1 = medicine_box (0 = background) Format: COCO JSON (polygons + bbox) Images: original resolution, never resized Structure train/ 540 images val/… See the full description on the dataset page: https://huggingface.co/datasets/ApyHTML19/Medication_Boxes_Arabe_Latin.imageimage-segmentationn<1K0 likes96 downloads9d agoHugging Face15julian-schelb /latin-classical-intertextuality-queries Latin Classical Intertextuality Queries This dataset contains query texts used for finding intertextual relationships with classical Latin authors. It comprises the whole works of Hieronymus (Jerome) and Lactantius, which are searched against a corpus of classical Latin literature. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Lactantius and classical Latin literature. Related Datasets This queries… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-queries.texttext-retrieval10K<n<100K2 likes95 downloads1mo agoHugging Face16grosenthal /latin_english_parallel Dataset Card for "latin_english_parallel" 101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation. For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences. Additionally, the English translations were both 1. copyrighted and 2. outdated. As such, we decided to modernize and… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_parallel.texttranslation100K<n<1M11 likes90 downloads3y agoHugging Face17wnkh /Tridis-latinimage100K<n<1M0 likes90 downloads10mo agoHugging Face18AncientLanguages /Latin-CC-170M Latin-CC-170M Reupload of the Corpus Corporum as Parquet, originally from Kaggle, and reuploaded as CSV by Fece228/latin-literature-dataset-170M. License These works are public domain. Original README This is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin up to 19th century, which includes: Classical Latin: works of Caesar, Cicero and many more Medieval Latin: a… See the full description on the dataset page: https://huggingface.co/datasets/AncientLanguages/Latin-CC-170M.tabular1K<n<10K1 likes81 downloads10mo agoHugging Face19wandb /deita-10k-v0-sft-latinSame as HuggingFaceH4/deita-10k-v0-sft but without non-latin text. text10K<n<100K1 likes79 downloads3y agoHugging Face20Blakus /Latinoamerican_Spanish_Voice_DatasetCompiled, and curated from the Crowdsourced high-quality speech datasets made by Google, available at https://openslr.org/resources.php. This dataset consists of a wavs folder with the audios plus a .txt file with the path to the audio and the speaker's transcription. Example: wavs/vem_05223_00896110924.wav|Los corazones de pollo son una delicia. wavs/vem_04310_01196944169.wav|Es un plato muy nutritivo. wavs/vem_02484_00854567505.wav|En este momento estoy enviando a sus mails unos links de… See the full description on the dataset page: https://huggingface.co/datasets/Blakus/Latinoamerican_Spanish_Voice_Dataset.text-to-speech1 likes74 downloads2y agoHugging Face21wnkh /medieval-latinimage10K<n<100K0 likes69 downloads10mo agoHugging Face22professorf /latin-vulgate latin-vulgate For machine learning translation of English-Latin and Latin-English — the complete Latin Vulgate Bible from the 4th Century AD and the complete English Douay-Rheims Bible from 1582. Citation If you use this data set, please support me by citing the repository. See APA Style. APA Citation: Flor, Nick. (2023). Latin Vulgate. GitHub. https://huggingface.co/datasets/professorf/latin-vulgate MLA Citation: Flor, Nick. Latin Vulgate. 2023, GitHub… See the full description on the dataset page: https://huggingface.co/datasets/professorf/latin-vulgate.2 likes66 downloads2y agoHugging Face23COMHIS /eacl26-detect-latin Dataset of the EACL 2026 paper Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark (arXiv:2510.19585) The benchmark contains 724 annotated pages from eighteenth-century British books. 594 pages contain Latin and 130 pages do not. Each example pairs the original OCR text, LLM-corrected OCR text, page-level metadata, Latin text segments, image references, and manual annotations for Latin regions and text spans. The original dataset record is… See the full description on the dataset page: https://huggingface.co/datasets/COMHIS/eacl26-detect-latin.imagetext-classificationn<1K0 likes50 downloads22h agoHugging Face24daidalos-project /latin_treebanks_ud_testNOTE: This template for datasheets for ancient language data is based on the proposal by Gebru et al. 2021 https://arxiv.org/abs/1803.09010. The majority of questions is taken from there, but in a slightly rearanged order. However, as some questions are not relevant for historical data, they have been left out. Questions which are of relevance for Humanities scholars researching ancient languages have been added. The question have been answered to the best of the knowledge of the Daidalos… See the full description on the dataset page: https://huggingface.co/datasets/daidalos-project/latin_treebanks_ud_test.texttoken-classification1K<n<10K0 likes49 downloads2mo agoHugging Face25linhgiang /classical-latin-greek-corpus Classical Latin & Greek Heritage Corpus (Perseus Digital Library) Total Records: 414,569 Format: Apache Parquet (Zstandard) File Size: 48.96 MB Source: Open Heritage Corpus 2.0 / Perseus Digital Library texttext-retrieval100K<n<1M0 likes49 downloads10d agoHugging Face26abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes48 downloads3mo agoHugging Face27zinaro /kurdish-latin-wikipedia-sentences Kurdish Latin Wikipedia Sentences Dataset This dataset consists of 78,004 Kurdish sentences extracted from Wikipedia. All sentences are written in Latin script and consist of 12 to 18 words. The dataset has been carefully cleaned to remove numbers, dates, or non-textual elements. Dataset Highlights Source: Wikipedia (Kurdish content) Script: Kurdish Latin Content: Pure textual sentences (no numbers, dates, or special characters) Sentence Length: 12 to 18 words per… See the full description on the dataset page: https://huggingface.co/datasets/zinaro/kurdish-latin-wikipedia-sentences.text10K<n<100K0 likes47 downloads2y agoHugging Face28ramonpzg /latin_musicaudion<1K0 likes46 downloads3y agoHugging Face29Dddixyy /latin_italian_parallel Italian-Latin Parallel Corpus (30,000 Sentences) This dataset provides approximately 30,000 parallel sentences between Italian and Latin. It is designed for tasks such as machine translation and cross-linguistic research. 🌐 The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited. Key Features Size: ~30,000 translation pairs. Languages: Italian (it), Latin (la). Source… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin_italian_parallel.texttranslation10K<n<100K3 likes46 downloads1y agoHugging Face30PoetryMTEB /LatinEpicIntertextualityRetrieval Latin Epic Intertextuality Retrieval BEIR-style Retrieval benchmark for Latin epic intertextuality under PoetryMTEB. Given a passage from Valerius Flaccus, Argonautica Book 1, retrieve the corresponding verse line(s) in Vergil (Aeneid), Lucan (Bellum Civile), Ovid (Metamorphoses), or Statius (Thebaid) that traditional scholarship identifies as parallels. Gold parallels: Burns et al., NAACL-HLT 2021 (paper; repo), 945 curated pairs. Dataset Card Item… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/LatinEpicIntertextualityRetrieval.tabulartext-retrieval10K<n<100K0 likes46 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.