datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fontsfont_crops_v4font_crops_v2fontfondant-cc-25m
Dataset Card for Fondant Creative Commons 25 million (fondant-cc-25m)
Changelog
Release
Description
v0.1
Release of the Fondant-cc-25m dataset
Dataset Summary
Fondant-cc-25m contains 25 million image URLs with their respective Creative Commons
license information collected from the Common Crawl web corpus.
The dataset was created using Fondant, an open source framework that aims to simplify and speed up
large-scale data processing by making… See the full description on the dataset page: https://huggingface.co/datasets/fondant-ai/fondant-cc-25m.font-square-v2
Accessing the font-square-v2 Dataset on Hugging Face
The font-square-v2 dataset is hosted on Hugging Face at blowing-up-groundhogs/font-square-v2. It is stored in WebDataset format, with tar files organized as follows:
tars/train/: Contains {000..499}.tar shards for the main training split.
tars/fine_tune/: Contains {000..049}.tar shards for fine-tuning.
Each tar file contains multiple samples, where each sample includes:
An RGB image (.rgb.png)
A black-and-white image (.bw.png)… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-v2.font-square-pretrain-20M
📚 Citation
If you use this dataset in your research, please cite these papers:
@article{pippi2023evaluating,
title={Evaluating Synthetic Pre-Training for Handwriting Processing Tasks},
author={Pippi, Vittorio and Cascianelli, Silvia and Baraldi, Lorenzo and Cucchiara, Rita},
journal={Pattern Recognition Letters},
year={2023},
publisher={Elsevier}
}
@InProceedings{pippi2025zeroshot,
author = {Pippi, Vittorio and Quattrini, Fabio and Cascianelli, Silvia and Tonioni… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-pretrain-20M.my-fontsgithub-code-fontend-lang
github-code fontend code
Dwonload
方式一
huggingface-cli download --resume-download LiXiang12/github-code-fontend-lang --include "*/*.zip" --repo-type dataset --local-dir github_code
方式二
进入Files and versions/data直接下载zip文件
数据统计
fontfont-square-v2-pairs-vaeeng-fonts128-1mASCII_Alphabet_Dataset_571_Fonts
Dataset Description
This dataset provides programmatically generated ASCII representations of the English alphabet rendered using 571 fonts from the PyFiglet library. Each letter (A–Z) is available in multiple typographic styles, resulting in a structured and high-variability dataset suitable for research, experimentation, and creative applications.
The dataset was created to support tasks involving text-based pattern recognition, synthetic data generation, typography analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/beta3/ASCII_Alphabet_Dataset_571_Fonts.heb-fonts-2font-hh104-hh1multi_lingal_fonts
Dataset Statistics
#
Language
Code
Font Count
1
English
en
5201
2
Hindi
hi
117
3
Kannada
kn
78
4
Tamil
ta
36
5
Telugu
te
70
6
Marathi
mr
117
7
Punjabi
pa
71
8
Bengali
bn
32
9
Odia
or
68
10
Malayalam
ml
72
11
Gujarati
gu
51
12
Sanskrit
sa
117
13
Japanese
ja
40
14
Korean
ko
40
15
Chinese
zh
9
16
German
de
47
17
French
fr
118
18
Italian
it
157
19
Russian
ru
62
20
Arabic
ar
26
21
Spanish
es
174
22
Thai
th
144
Total Fonts: 6… See the full description on the dataset page: https://huggingface.co/datasets/v1v1d/multi_lingal_fonts.FontTransfer
NomGenie: Font Diffusion for Sino-Nom Language
NomGenie is a specialized image-to-image dataset designed for font generation and style transfer within the Sino-Nom (Hán-Nôm) script system. This dataset facilitates the training of deep learning models—particularly Diffusion Models and GANs—to preserve the historical and structural integrity of Vietnamese Nom characters while applying diverse typographic styles.
Dataset Description
The dataset consists of paired images:… See the full description on the dataset page: https://huggingface.co/datasets/dzungpham/FontTransfer.analyse_fonciere_dataViet-Font-mc4-Textfont_crops_v5early_printed_books_font_detection_loaded
Dataset Card for "early_printed_books_font_detection_loaded"
More Information needed
chinese_fonts_common_128x128
Dataset Card for "chinese_fonts_common_128x128"
More Information needed
galdrastafir-fonts
Galdrastafir Font Recognition Dataset
Synthetic font-in-the-wild images generated with ControlNet + diffusion for training
font recognition models. Each image has per-layer metadata (most with bounding boxes)
for every font rendered into the scene. Fonts appear as signs, sculptures, paintings,
fashion, screens, graffiti, and more — rendered in 3D with diverse transformations.
Columns
column
type
description
image_id
str
unique image identifier (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Nafnlaus/galdrastafir-fonts.E3-miu-GNN
Neo Mixed-Granularity Atomistic Dataset
Neo is the dataset collection for
E3-miu-GNN, an E(3)-equivariant
graph neural network developed by Fona Group. This repository hosts the
training data. GitHub hosts the model, GUI, training and inference workflows,
dataset preparation code, tests, API documentation, and phonon tools.
Project status
Datasets: canonical Tiny through Large and composite SE, Plus, and Max are
materialized with the 2026-07-25 non-OMat24 stress… See the full description on the dataset page: https://huggingface.co/datasets/FonaTech/E3-miu-GNN.eng-fonts64-augBMA_Fondue_Images
Bangkok Metropolitan Urban Issue Image (Traffy Fondue Issue Buckets)
BMA_Fondue_Image Dataset is an original raw dataset without frame labels.
heb-fonts-3font-hh104-hh1-hh2eng-fonts32-augearly_printed_books_font_detection
Early Printed Books Font Detection
Photographs of 35,623 pages from books printed between the mid-15th and the end of the 18th century, each labelled by experts with the font group or groups used on the page. This is a mirror of Dataset of Pages from Early Printed Books with Multiple Font Groups by Mathias Seuret, Saskia Limbach, Nikolaus Weichselbaumer, Andreas Maier and Vincent Christlein, deposited on Zenodo in August 2019 and described in their HIP'19 paper.
The page images… See the full description on the dataset page: https://huggingface.co/datasets/biglam/early_printed_books_font_detection.datacomp-small-clip
Production-ready
data processing made
easy
and
shareable
Explore the Fondant docs »
Dataset Card for fondant-ai/datacomp-small-clip
This is a dataset containing image urls and their CLIP embeddings, based on the datacomp_small dataset, and processed with fondant.
Dataset Details
Dataset Description
Large (image) datasets are often unwieldy to use due to their… See the full description on the dataset page: https://huggingface.co/datasets/fondant-ai/datacomp-small-clip.svg-fonts
Dataset Card for svg-fonts
Dataset Description
This dataset contains SVG code examples for training and evaluating SVG models for image vectorization.
Dataset Structure
Features
The dataset contains the following fields:
Field Name
Description
Filename
Unique ID for each SVG
Svg
SVG code
Usage
from datasets import load_dataset
dataset = load_dataset("starvector/svg-fonts")… See the full description on the dataset page: https://huggingface.co/datasets/starvector/svg-fonts.
