Team Ai
Datasetpublic

Praha-Labs/DaTikZ-V4-Instruct-100K

DaTikZ-V4 Instruction 100K A cleaned instruction-tuning dataset for text-to-TikZ generation, derived from the first 100,000 rows of nllg/DaTikZ-V4. The original dataset provides rendered diagram images and corresponding TikZ source code. This derived dataset adds natural-language, imperative user instructions that describe how to recreate each diagram. These instructions are intended as model inputs, with the original TikZ code serving as the supervised target.… See the full description on the dataset page: https://huggingface.co/datasets/Praha-Labs/DaTikZ-V4-Instruct-100K.

sourceHugging Faceapache-2.0updated 15d agoView on Hugging Face
0likes215downloads
Dataset Card

DaTikZ-V4 Instruction 100K

A cleaned instruction-tuning dataset for text-to-TikZ generation, derived from the first 100,000 rows of `nllg/DaTikZ-V4`.

The original dataset provides rendered diagram images and corresponding TikZ source code. This derived dataset adds natural-language, imperative user instructions that describe how to recreate each diagram. These instructions are intended as model inputs, with the original TikZ code serving as the supervised target.

Dataset construction

For every source row, both the rendered image and TikZ source code were provided to nvidia/Qwen3.8-27B-NVFP4.

The image supplied the visual composition and appearance. The TikZ source was used as authoritative supporting information for visible details that a vision model might overlook, including:

  • —exact labels and mathematical notation;
  • —colors and line styles;
  • —arrow directions and endpoints;
  • —node relationships and layout;
  • —plot annotations and legends;
  • —repeated or visually subtle elements.

The generated output was required to be a natural-language, imperative user request such as “Create…”, “Draw…”, or “Design…”. It was explicitly not intended to be a passive image caption, code explanation, or reference to the source image or TikZ code.

Source and provenance

  • —Source dataset: nllg/DaTikZ-V4
  • —Source revision: 33734c83608211682be11001a1618856fc1979dd
  • —Source split: train
  • —Source range: rows 0–99999
  • —Source license: Apache-2.0
  • —Caption model: nvidia/Qwen3.8-27B-NVFP4
  • —Model revision: 482ca0f3832238542f8f5295dde86b5f22711d80
  • —Prompt version: caption-v2
  • —Prompt SHA-256: 3cd9267b40dc9331851d48a8efe796baeff4f258d57b24090d2a8c73aa2cd0e6
  • —Generation run ID: d97cc0c4f0f810d0
  • —Maximum instruction length: 300 words
  • —Initial generation ceiling: 512 tokens
  • —Truncation retry ceiling: 768 tokens
  • —Context length: 32,768 tokens

Processing results

ResultRows
Source rows frozen100,000
Accepted training examples98,450
Rejected examples1,550
Input too long670
Invalid instruction810
Truncated at retry ceiling70

The final production audit passed with zero violations. Every source row reached a terminal accepted or rejected state.

Default training schema

The default configuration loads shards/*.parquet and contains accepted training examples.

FieldDescription
idStable row identifier
source_row_indexOriginal row index in the frozen source slice
file_idOriginal source file identifier
png_imageRendered PNG image stored as bytes
tikz_codeOriginal TikZ source code and supervised target
instructionGenerated imperative natural-language request
source_datasetSource dataset identifier
source_revisionPinned source revision
caption_modelModel used to generate the instruction
caption_model_revisionPinned model revision
prompt_versionPrompt specification version
image_sha256Image-content checksum
tikz_sha256TikZ-source checksum

Additional configurations and files

  • —default: 98,450 accepted training examples.
  • —rejected: 1,550 excluded rows with explicit rejection reasons.
  • —attempts: generation-attempt history for reproducibility and diagnostics.
  • —run-metadata.json: complete pinned generation identity and provenance.
  • —checksums.json: physical and logical checksums for exported shards.
  • —stats.json: processing and token statistics.
  • —validation-report.json: final validation report.
  • —export.meta.json: export identity and row counts.

Intended use

The primary intended use is supervised fine-tuning of models that map natural-language diagram instructions to TikZ source code:

  • —input: instruction
  • —target: tikz_code

The image is retained for multimodal experiments, verification, evaluation and future image-conditioned training.

Important differences from the source dataset

This is not an unmodified mirror of DaTikZ-V4. It:

  1. 1.freezes only source rows 0–99,999 at a pinned revision;
  2. 2.adds model-generated instruction-style requests;
  3. 3.preserves stable source identifiers and content hashes;
  4. 4.rejects rows that fail context, truncation or instruction validation;
  5. 5.exports deterministic Parquet shards sorted by source index;
  6. 6.includes checksums, provenance, rejection records and attempt history.

Limitations

  • —Instructions are model-generated and have not been individually human-reviewed.
  • —The first 100,000 source rows are used rather than a random sample.
  • —Some generated instructions may contain factual omissions or stylistic artifacts.
  • —Acceptance indicates structural and policy validation, not guaranteed semantic perfection.
  • —The source dataset may contain artifacts inherited from scientific documents and other upstream sources.
  • —Users should perform task-specific quality evaluation before training production models.

License and attribution

The source dataset is distributed under the Apache License 2.0. This derived dataset retains that license and provides attribution to `nllg/DaTikZ-V4`.

When using this dataset, please also review and cite the original DaTikZ dataset and its associated work.