Team Ai
Datasetpublic

Dheerraa/TinyLibrary

TinyLibrary Paper: TinyLibrary: Do Age-Graded Curricula Help Small Vision-Language Models?, accepted at the BabyLM 2026 Workshop at EMNLP 2026. Code: github.com/dhevarghese/TinyLibrary Dataset summary TinyLibrary contains synthetic captioning, visual question-answering, and reasoning conversations generated from illustrated pages in the English-language portion of the International Children's Digital Library (ICDL). It was created for studying age-based curricula… See the full description on the dataset page: https://huggingface.co/datasets/Dheerraa/TinyLibrary.

sourceHugging Faceotherupdated 25d agoView on Hugging Face
0likes212downloads
Dataset Card

TinyLibrary

Paper: TinyLibrary: Do Age-Graded Curricula Help Small Vision-Language Models?, accepted at the BabyLM 2026 Workshop at EMNLP 2026.

Code: github.com/dhevarghese/TinyLibrary

Dataset summary

TinyLibrary contains synthetic captioning, visual question-answering, and reasoning conversations generated from illustrated pages in the English-language portion of the International Children's Digital Library (ICDL). It was created for studying age-based curricula in small vision-language models.

The paper-version dataset contains 128,823 conversations totaling approximately 21.7M words, referencing 43,184 distinct page images from 46,721 filtered pages. A conversation may contain multiple question-answer pairs.

Release contents

The Hugging Face release contains the synthetic annotations, per-sample difficulty scores, sample mappings, and reconstructed curriculum orders used in the study.

FileContents
{task}_{age_band}.jsonTwelve captioning, VQA, and reasoning annotation files.
difficulty_scores.jsonlReadability and teacher-model scores for 128,823 samples.
sample_mapping.jsonlSample IDs, annotation locations, split membership, and index mappings.
curriculum_orders/ep5/*.npy, curriculum_orders/ep10/*.npyReconstructed training-sample sequences.
curriculum_orders/metadata.jsonMetadata describing each curriculum sequence.
export_metadata.jsonExport provenance and validation information.
annotations.sha256, artifacts.sha256Checksums for annotations and exported artifacts.

The release does not contain ICDL books or page images, original OCR text, intermediate page descriptions, or trained model checkpoints. Image filenames are references only, so multimodal training requires independently obtained source images.

The code repository contains the filtering and generation code, prompts, scoring and curriculum-construction scripts, training and evaluation utilities, and analysis code.

Data construction

  1. 1.Mistral-7B-Instruct cleans the original book OCR and estimates a book-level reader-age band. Long books are classified in chunks and assigned a band by majority vote.
  2. 2.Page filtering selects illustrated pages using OCR text coverage and visual edge-content criteria.
  3. 3.CogVLM2 generates a detailed description of each selected page image.
  4. 4.Llama-3-8B-Instruct converts these descriptions into LLaVA-style captioning, VQA, and reasoning conversations. Generation prompts vary vocabulary, sentence structure, and reasoning complexity according to the inferred age band.

The annotations are synthetic rather than the books' original prose or responses collected from children. Age bands are model-derived estimates, not publisher ratings or validated measures of educational suitability.

Annotation files and counts

Files follow {task}_{age_band}.json, where task is caption, vqa, or reasoning. Each file contains a JSON array of conversation records.

Hugging Face configurations expose each task/age file separately, alongside difficulty_scores and sample_mapping. The all split refers to the complete file rather than the experimental training split. Experimental train/validation membership is recorded in sample_mapping.jsonl.

Age bandFilename suffixCaptioningVQAReasoningTotal
3–53_511,40211,61111,61034,623
6–86_86,1106,1406,13918,389
9–129_125,1315,1825,17515,488
12+1219,91820,24320,16260,323
Total42,56143,17643,086128,823

File checksums are provided in annotations.sha256.

The suffix 12 denotes the 12+ age band. These are the coarse categories used in the study rather than exact ages.

Record schema

FieldMeaning
idSource-page identifier. It can recur across annotation files.
imageFilename of the corresponding page image, which is not included.
conversationsOrdered conversation messages.
conversations[].fromhuman for a prompt or gpt for a response. These are format labels, not evidence of human authorship.
conversations[].valueMessage text. A prompt can contain an <image> placeholder.

Task and age band are encoded in the filename rather than stored as separate fields in each record.

The training manifest uses the composite identifier:

text
{task}_{age_band}_{id}

Use this identifier, rather than the page ID alone, when joining annotations with difficulty scores or sample mappings.

Difficulty scores

difficulty_scores.jsonl contains readability and teacher-model scores copied from the verified scored manifest.

FieldDefinition
sample_idComposite sample ID matching the annotation record and sample mapping.
fk_gradeFlesch–Kincaid grade of the concatenated prompt and response text after removing the image placeholder.
n_wordsWhitespace-delimited word count of the prompt-and-response text.
nllMean per-token NLL under SmolLM2-135M.
nll_qwenMean per-token NLL under Qwen3-0.6B.
nll_gemmaMean per-token NLL under Gemma-3-270M.
n_tokensNon-padding token count after the primary teacher scorer's truncation, including the initial prompt.

The initial user prompt is masked when computing NLL. In multi-turn conversations, later user turns and assistant responses are scored. The teacher models receive text only.

Lower NLL indicates greater predictability under a particular teacher model and should not be interpreted directly as reading difficulty for a child. Scores produced with different teacher tokenizers should not be treated as a common absolute scale.

Curriculum orders and split

The study compares developmental age order, reverse age order, readability, teacher NLL, and random ordering, together with random-block and alternative-teacher controls. Both staged and competence-based schedules are included.

The experimental split contains 126,247 training samples and 2,576 validation samples, using a fixed 2% sample-level split with seed 42. It is not book-disjoint or page-disjoint.

The release contains 63 unique curriculum sequences covering 72 training configurations. Nine text-only controls reuse existing ten-epoch sequences; this reuse is recorded through applies_to in curriculum_orders/metadata.json.

Staged and random sequences include every training sample once per pass. Competence-based sequences sample with replacement and can therefore repeat some records while omitting others.

Each .npy file is a one-dimensional array of zero-based unsigned 32-bit indices into the training subset defined by sample_mapping.jsonl.

Sample mapping

FieldMeaning
sample_idUnique composite sample identifier.
manifest_indexRow in the original complete manifest.
annotation_fileJSON file containing the conversation.
annotation_indexPosition within that file's JSON array.
source_idOriginal page ID from the annotation record.
imageReferenced page-image filename, not an included image.
task, age_bandAnnotation type and inferred age category.
splittrain or validation.
split_indexPosition within that split, preserving original manifest order.

For example, the first sample in a curriculum sequence can be resolved as follows:

python
import json
from pathlib import Path

import numpy as np

root = Path(".")

with (root / "sample_mapping.jsonl").open(encoding="utf-8") as handle:
    mapping = [json.loads(line) for line in handle]

train = sorted(
    (row for row in mapping if row["split"] == "train"),
    key=lambda row: row["split_index"],
)

sequence = np.load(
    root / "curriculum_orders/ep10/developmental_staged_seed0.npy",
    allow_pickle=False,
    mmap_mode="r",
)

sample = train[int(sequence[0])]

with (root / sample["annotation_file"]).open(encoding="utf-8") as handle:
    annotation = json.load(handle)[sample["annotation_index"]]

print(sample["sample_id"], annotation["conversations"])

Reconstruction note

The released curriculum orders are reconstructed sampler-input sequences rather than saved training traces. They reproduce the original split, random seeds, sampling, and tie-breaking implementation.

The sequences describe sample selection before dataloader sharding, batching, sequence packing, and training stopping. They therefore do not specify the exact packed-batch order or final sequence prefix consumed by a checkpoint.

For competence-based sampling, epochs determines the number of draws rather than guaranteeing complete corpus passes or identical realized token exposure. The eligible prefix grows per draw with competence_c0=0.1. Random ordering ignores the schedule argument, so its public schedule metadata is null.

Reconstruction provenance, environment versions, validation results, and checksums are provided in export_metadata.json and curriculum_orders/metadata.json.

Intended use and limitations

TinyLibrary is intended for research on training-data order, synthetic multimodal annotations, and data-efficient language and vision-language learning. It is not a human-annotated grounding benchmark or a validated educational resource for children.

License and source material

TinyLibrary is released for research use subject to the terms described below.

The release contains synthetic annotations, difficulty scores, sample mappings, and curriculum-order metadata. It does not include ICDL books, page images, or original OCR text.

To the extent that rights are held by the dataset authors, permission is granted to use, reproduce, and modify the released data for non-commercial research and evaluation. This permission does not grant rights to the underlying ICDL books or illustrations, which remain subject to their original copyright and licensing terms.

Some annotations were generated using Meta Llama 3. Use of these annotations must comply with the applicable Meta Llama 3 Community License. In particular, that license restricts the use of Llama 3 outputs to improve other large language models. Users are responsible for ensuring that their intended use complies with these upstream terms.

The accompanying source code is separately licensed under the MIT License.

No warranty is provided regarding the availability of rights for uses beyond those described above. Users are responsible for determining whether their intended use complies with applicable copyright, licensing, and other legal requirements.

Citation

bibtex
@inproceedings{varghese2026tinylibrary,
  title     = {TinyLibrary: Do Age-Graded Curricula Help Small Vision-Language Models?},
  author    = {Dheeraj Varghese},
  booktitle = {BabyLM 2026 Workshop at EMNLP 2026},
  year      = {2026},
  url       = {https://openreview.net/forum?id=3S9UR6Xizi}
}