datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.Latin-PD
🇲🇪 Latin Public Domain Books (Latin) 🇲🇪
Latin-Public Domain or Latin-PD is a large collection aiming to aggregate all Latin monographies and periodicals in the public domain. As of June 2024, it is the largest Latin open corpus.
Dataset summary
The collection contains 16,521,454,086 words (159,070 titles) recovered from multiple sources, including the Internet Archive and various European national libraries and cultural heritage institutions (BDH, BNF). Each parquet… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Latin-PD.latin-literature-dataset-170MThis is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin. The dataset is split in two parts: preprocessed with basic cltk tools, ready for work, and raw text data. It must be noted, however, that the latter contains text in Greek, Hebrew, and other languages, with references and contractions
latin-summarizer-dataset
✨ LatinSummarizer Dataset ✨
Note: If Dataset Viewer is not available, see samples of dataset for samples from the dataset.The LatinSummarizer Dataset is a comprehensive collection of Latin texts designed to support natural language processing research for a low-resource language. It provides parallel data for various tasks, including translation (Latin-to-English) and summarization (extractive and abstractive).
This dataset was created for a… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/latin-summarizer-dataset.uyghur-cv-latinLatinSummarizer
LatinSummarizer Dataset
Structure
aligned_en_la_data_raw.csv
aligned_en_la_data_cleaned.csv
aligned_en_la_data_cleaned_with_stanza.csv
concat_aligned_data.csv
concat_cleaned.csv
latin_wikipedia_cleaned.csv
latin_wikipedia_raw.csv
latin-literature-dataset-170M_raw_cleaned.csv
latin-literature-dataset-170M_raw_cleaned_chunked.csv
Elsa_aligned/
README.md
Details
aligned_en_la_data_raw.csv
This dataset contains aligned Latin (la) - English (en)… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/LatinSummarizer.LatinFontsSVGs
SVG Font Dataset
Overview
We present a curated dataset of vector font glyphs stored as SVG files, designed for research in generative modeling, structured vector graphics and typography synthesis.
The dataset was created for the development and evaluation of our paper:
DesigNet: Learning to Draw Vector Graphics as Designers Do
Related Resources
Paper (arXiv) : https://arxiv.org/abs/2604.06494
Code: https://github.com/TomasGuija/DesigNet… See the full description on the dataset page: https://huggingface.co/datasets/TomasGuija/LatinFontsSVGs.arabic-latin-invoices-synthetic
Synthetic Arabic/Latin Invoices — label-first multimodal dataset
A large, perfectly-labeled synthetic invoice dataset for training multimodal
invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin
layouts, and all three numeral glyph systems.
Every sample is generated label-first: a structured invoice record is sampled, then
rendered to pixels via a headless browser, and bounding boxes are read back from the same
DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.latin_english_translation
Dataset Card for "latin_english_parallel"
101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation.
For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences.
Each sample is annotated with the index and file (and therefore author/work) that the sample is from. If you find errors… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_translation.latin-classical-intertextuality-corpus
Latin Classical Authors Corpus
This dataset contains processed texts from classical Latin authors, serving as a retrieval corpus for intertextuality research. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Cicero and classical Latin literature.
Related Datasets
This corpus is part of the Latin Jerome Intertextuality collection:
Queries: latin-classical-intertextuality-queries - The whole works of… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-corpus.latin-classical-intertextuality-labels
Latin Jerome Intertextuality Labels
This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models for… See the full description on the dataset page: https://huggingface.co/datasets/reesjon9/Latin-Audio.WP8-Latin-Embeddings-IndicesMedication_Boxes_Arabe_Latin
Medication Boxes — Arabic / Latin
This dataset contains photos of medication boxes, with packaging text in Arabic and Latin script (French and others). It is annotated for COCO instance segmentation and was built for MediSeG, which segments each box so the text on it can be read afterwards.
One class: 1 = medicine_box (0 = background)
Format: COCO JSON (polygons + bbox)
Images: original resolution, never resized
Structure
train/ 540 images
val/… See the full description on the dataset page: https://huggingface.co/datasets/ApyHTML19/Medication_Boxes_Arabe_Latin.latin-classical-intertextuality-queries
Latin Classical Intertextuality Queries
This dataset contains query texts used for finding intertextual relationships with classical Latin authors. It comprises the whole works of Hieronymus (Jerome) and Lactantius, which are searched against a corpus of classical Latin literature. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Lactantius and classical Latin literature.
Related Datasets
This queries… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-queries.latin_english_parallel
Dataset Card for "latin_english_parallel"
101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation.
For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences. Additionally, the English translations were both 1. copyrighted and 2. outdated. As such, we decided to modernize and… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_parallel.Tridis-latinLatin-CC-170M
Latin-CC-170M
Reupload of the Corpus Corporum as Parquet, originally from
Kaggle,
and reuploaded as CSV by
Fece228/latin-literature-dataset-170M.
License
These works are public domain.
Original README
This is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin up to 19th century, which includes:
Classical Latin: works of Caesar, Cicero and many more
Medieval Latin: a… See the full description on the dataset page: https://huggingface.co/datasets/AncientLanguages/Latin-CC-170M.deita-10k-v0-sft-latinSame as HuggingFaceH4/deita-10k-v0-sft but without non-latin text.
Latinoamerican_Spanish_Voice_DatasetCompiled, and curated from the Crowdsourced high-quality speech datasets made by Google, available at https://openslr.org/resources.php.
This dataset consists of a wavs folder with the audios plus a .txt file with the path to the audio and the speaker's transcription.
Example:
wavs/vem_05223_00896110924.wav|Los corazones de pollo son una delicia.
wavs/vem_04310_01196944169.wav|Es un plato muy nutritivo.
wavs/vem_02484_00854567505.wav|En este momento estoy enviando a sus mails unos links de… See the full description on the dataset page: https://huggingface.co/datasets/Blakus/Latinoamerican_Spanish_Voice_Dataset.medieval-latinlatin-vulgate
latin-vulgate
For machine learning translation of English-Latin and Latin-English — the complete Latin Vulgate Bible from the 4th Century AD and the complete English Douay-Rheims Bible from 1582.
Citation
If you use this data set, please support me by citing the repository. See APA Style.
APA Citation:
Flor, Nick. (2023). Latin Vulgate. GitHub. https://huggingface.co/datasets/professorf/latin-vulgate
MLA Citation:
Flor, Nick. Latin Vulgate. 2023, GitHub… See the full description on the dataset page: https://huggingface.co/datasets/professorf/latin-vulgate.eacl26-detect-latin
Dataset of the EACL 2026 paper Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark (arXiv:2510.19585)
The benchmark contains 724 annotated pages from eighteenth-century British books. 594 pages contain Latin and 130 pages do not. Each example pairs the original OCR text, LLM-corrected OCR text, page-level metadata, Latin text segments, image references, and manual annotations for Latin regions and text spans.
The original dataset record is… See the full description on the dataset page: https://huggingface.co/datasets/COMHIS/eacl26-detect-latin.latin_treebanks_ud_testNOTE: This template for datasheets for ancient language data is based on the proposal by Gebru et al. 2021 https://arxiv.org/abs/1803.09010. The majority of questions is taken from there, but in a slightly rearanged order. However, as some questions are not relevant for historical data, they have been left out. Questions which are of relevance for Humanities scholars researching ancient languages have been added. The question have been answered to the best of the knowledge of the Daidalos… See the full description on the dataset page: https://huggingface.co/datasets/daidalos-project/latin_treebanks_ud_test.classical-latin-greek-corpus
Classical Latin & Greek Heritage Corpus (Perseus Digital Library)
Total Records: 414,569
Format: Apache Parquet (Zstandard)
File Size: 48.96 MB
Source: Open Heritage Corpus 2.0 / Perseus Digital Library
Kabyle-Latin-to-Tifinagh-Parallel-Corpus
Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.kurdish-latin-wikipedia-sentences
Kurdish Latin Wikipedia Sentences Dataset
This dataset consists of 78,004 Kurdish sentences extracted from Wikipedia. All sentences are written in Latin script and consist of 12 to 18 words. The dataset has been carefully cleaned to remove numbers, dates, or non-textual elements.
Dataset Highlights
Source: Wikipedia (Kurdish content)
Script: Kurdish Latin
Content: Pure textual sentences (no numbers, dates, or special characters)
Sentence Length: 12 to 18 words per… See the full description on the dataset page: https://huggingface.co/datasets/zinaro/kurdish-latin-wikipedia-sentences.latin_musiclatin_italian_parallel
Italian-Latin Parallel Corpus (30,000 Sentences)
This dataset provides approximately 30,000 parallel sentences between Italian and Latin. It is designed for tasks such as machine translation and cross-linguistic research.
🌐 The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited.
Key Features
Size: ~30,000 translation pairs.
Languages: Italian (it), Latin (la).
Source… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin_italian_parallel.LatinEpicIntertextualityRetrieval
Latin Epic Intertextuality Retrieval
BEIR-style Retrieval benchmark for Latin epic intertextuality under PoetryMTEB.
Given a passage from Valerius Flaccus, Argonautica Book 1, retrieve the corresponding verse line(s) in Vergil (Aeneid), Lucan (Bellum Civile), Ovid (Metamorphoses), or Statius (Thebaid) that traditional scholarship identifies as parallels.
Gold parallels: Burns et al., NAACL-HLT 2021 (paper; repo), 945 curated pairs.
Dataset Card
Item… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/LatinEpicIntertextualityRetrieval.
