Dheerraa/TinyLibrary
TinyLibrary Paper: TinyLibrary: Do Age-Graded Curricula Help Small Vision-Language Models?, accepted at the BabyLM 2026 Workshop at EMNLP 2026. Code: github.com/dhevarghese/TinyLibrary Dataset summary TinyLibrary contains synthetic captioning, visual question-answering, and reasoning conversations generated from illustrated pages in the English-language portion of the International Children's Digital Library (ICDL). It was created for studying age-based curricula… See the full description on the dataset page: https://huggingface.co/datasets/Dheerraa/TinyLibrary.
TinyLibrary
Paper: TinyLibrary: Do Age-Graded Curricula Help Small Vision-Language Models?, accepted at the BabyLM 2026 Workshop at EMNLP 2026.
Code: github.com/dhevarghese/TinyLibrary
Dataset summary
TinyLibrary contains synthetic captioning, visual question-answering, and reasoning conversations generated from illustrated pages in the English-language portion of the International Children's Digital Library (ICDL). It was created for studying age-based curricula in small vision-language models.
The paper-version dataset contains 128,823 conversations totaling approximately 21.7M words, referencing 43,184 distinct page images from 46,721 filtered pages. A conversation may contain multiple question-answer pairs.
Release contents
The Hugging Face release contains the synthetic annotations, per-sample difficulty scores, sample mappings, and reconstructed curriculum orders used in the study.
The release does not contain ICDL books or page images, original OCR text, intermediate page descriptions, or trained model checkpoints. Image filenames are references only, so multimodal training requires independently obtained source images.
The code repository contains the filtering and generation code, prompts, scoring and curriculum-construction scripts, training and evaluation utilities, and analysis code.
Data construction
- Mistral-7B-Instruct cleans the original book OCR and estimates a book-level reader-age band. Long books are classified in chunks and assigned a band by majority vote.
- Page filtering selects illustrated pages using OCR text coverage and visual edge-content criteria.
- CogVLM2 generates a detailed description of each selected page image.
- Llama-3-8B-Instruct converts these descriptions into LLaVA-style captioning, VQA, and reasoning conversations. Generation prompts vary vocabulary, sentence structure, and reasoning complexity according to the inferred age band.
The annotations are synthetic rather than the books' original prose or responses collected from children. Age bands are model-derived estimates, not publisher ratings or validated measures of educational suitability.
Annotation files and counts
Files follow {task}_{age_band}.json, where task is caption, vqa, or reasoning. Each file contains a JSON array of conversation records.
Hugging Face configurations expose each task/age file separately, alongside difficulty_scores and sample_mapping. The all split refers to the complete file rather than the experimental training split. Experimental train/validation membership is recorded in sample_mapping.jsonl.
File checksums are provided in annotations.sha256.
The suffix 12 denotes the 12+ age band. These are the coarse categories used in the study rather than exact ages.
Record schema
Task and age band are encoded in the filename rather than stored as separate fields in each record.
The training manifest uses the composite identifier:
{task}_{age_band}_{id}Use this identifier, rather than the page ID alone, when joining annotations with difficulty scores or sample mappings.
Difficulty scores
difficulty_scores.jsonl contains readability and teacher-model scores copied from the verified scored manifest.
The initial user prompt is masked when computing NLL. In multi-turn conversations, later user turns and assistant responses are scored. The teacher models receive text only.
Lower NLL indicates greater predictability under a particular teacher model and should not be interpreted directly as reading difficulty for a child. Scores produced with different teacher tokenizers should not be treated as a common absolute scale.
Curriculum orders and split
The study compares developmental age order, reverse age order, readability, teacher NLL, and random ordering, together with random-block and alternative-teacher controls. Both staged and competence-based schedules are included.
The experimental split contains 126,247 training samples and 2,576 validation samples, using a fixed 2% sample-level split with seed 42. It is not book-disjoint or page-disjoint.
The release contains 63 unique curriculum sequences covering 72 training configurations. Nine text-only controls reuse existing ten-epoch sequences; this reuse is recorded through applies_to in curriculum_orders/metadata.json.
Staged and random sequences include every training sample once per pass. Competence-based sequences sample with replacement and can therefore repeat some records while omitting others.
Each .npy file is a one-dimensional array of zero-based unsigned 32-bit indices into the training subset defined by sample_mapping.jsonl.
Sample mapping
For example, the first sample in a curriculum sequence can be resolved as follows:
import json
from pathlib import Path
import numpy as np
root = Path(".")
with (root / "sample_mapping.jsonl").open(encoding="utf-8") as handle:
mapping = [json.loads(line) for line in handle]
train = sorted(
(row for row in mapping if row["split"] == "train"),
key=lambda row: row["split_index"],
)
sequence = np.load(
root / "curriculum_orders/ep10/developmental_staged_seed0.npy",
allow_pickle=False,
mmap_mode="r",
)
sample = train[int(sequence[0])]
with (root / sample["annotation_file"]).open(encoding="utf-8") as handle:
annotation = json.load(handle)[sample["annotation_index"]]
print(sample["sample_id"], annotation["conversations"])Reconstruction note
The released curriculum orders are reconstructed sampler-input sequences rather than saved training traces. They reproduce the original split, random seeds, sampling, and tie-breaking implementation.
The sequences describe sample selection before dataloader sharding, batching, sequence packing, and training stopping. They therefore do not specify the exact packed-batch order or final sequence prefix consumed by a checkpoint.
For competence-based sampling, epochs determines the number of draws rather than guaranteeing complete corpus passes or identical realized token exposure. The eligible prefix grows per draw with competence_c0=0.1. Random ordering ignores the schedule argument, so its public schedule metadata is null.
Reconstruction provenance, environment versions, validation results, and checksums are provided in export_metadata.json and curriculum_orders/metadata.json.
Intended use and limitations
TinyLibrary is intended for research on training-data order, synthetic multimodal annotations, and data-efficient language and vision-language learning. It is not a human-annotated grounding benchmark or a validated educational resource for children.
License and source material
TinyLibrary is released for research use subject to the terms described below.
The release contains synthetic annotations, difficulty scores, sample mappings, and curriculum-order metadata. It does not include ICDL books, page images, or original OCR text.
To the extent that rights are held by the dataset authors, permission is granted to use, reproduce, and modify the released data for non-commercial research and evaluation. This permission does not grant rights to the underlying ICDL books or illustrations, which remain subject to their original copyright and licensing terms.
Some annotations were generated using Meta Llama 3. Use of these annotations must comply with the applicable Meta Llama 3 Community License. In particular, that license restricts the use of Llama 3 outputs to improve other large language models. Users are responsible for ensuring that their intended use complies with these upstream terms.
The accompanying source code is separately licensed under the MIT License.
No warranty is provided regarding the availability of rights for uses beyond those described above. Users are responsible for determining whether their intended use complies with applicable copyright, licensing, and other legal requirements.
Citation
@inproceedings{varghese2026tinylibrary,
title = {TinyLibrary: Do Age-Graded Curricula Help Small Vision-Language Models?},
author = {Dheeraj Varghese},
booktitle = {BabyLM 2026 Workshop at EMNLP 2026},
year = {2026},
url = {https://openreview.net/forum?id=3S9UR6Xizi}
}