SmartWhatt/thai_spatial_reasoning
Thai Spatial Reasoning 1.0.0 Thai spatial captions and question-answer pairs for continual pretraining of a Thai foundation VLM. Synthetic means the Thai text and the spatial annotations are generated or processed; the primary images are real photographs. Images are not covered by one blanket licence, so the card does not point at one file: every image has its own licence, creator, and attribution in rights/attribution.csv and rights/ledger.parquet. The dataset-level other… See the full description on the dataset page: https://huggingface.co/datasets/SmartWhatt/thai_spatial_reasoning.
Thai Spatial Reasoning 1.0.0
Thai spatial captions and question-answer pairs for continual pretraining of a Thai foundation VLM. Synthetic means the Thai text and the spatial annotations are generated or processed; the primary images are real photographs.
Images are not covered by one blanket licence, so the card does not point at one file: every image has its own licence, creator, and attribution in `rights/attribution.csv` and `rights/ledger.parquet`. The dataset-level other licence applies to the annotations, the generated text, and the rendered images. (The Hub rejects a relative license_link; the ledger inside the repository is the licence statement.)
Contents
data/*.parquet is the canonical table: one row per image, with an Image column, one caption, a list of four question-answer structures, the relation list, and the rights fields. cpt/examples-*.parquet is the text-only index that expands one image row into five logical examples, so image bytes are stored once.
Method
- Real images come from a source with object boxes. Rights are checked per image against a licence allowlist; a verification date and an evidence URL are recorded for each one. Images with person labels wait for a human privacy decision and do not ship without it.
- Splits are assigned per duplicate group and per rendered scene family before any text exists. No group crosses a split boundary.
- Facts come from conservative image-frame geometry over boxes (a real separation on one axis and real alignment on the other, no clipped axis, no occluded reference object) and from source annotations for containment. Contact, depth, and distance are never inferred from 2-D boxes; for the renders they are measured from known 3-D state.
- Thai text is assembled from those facts with deterministic templates, so each sentence traces to fact ids, and the factual meaning does not depend on a model.
- An object whose class repeats in an image is named by what it looks like before any position is used ("คนที่ใส่เสื้อสีเขียว", not "คนที่อยู่ทางซ้าย"). A pinned vision model returns one value per field — colour, garment, pose, material — from a closed list; it writes no Thai, and an attribute can neither create nor remove a spatial fact. Renders are never sent to it, because their colours are known from the scene state.
- Every relation and every sentence is checked against its source fact, and a Thai-speaking review sample of 0 examples is scored before release.
Evaluation and quality
The Thai-speaking review gate is still open for this build. A sample of examples are queued in reports/audit_queue.csv for a native-speaker review of the relation, the language, and the faithfulness. This build ships before those verdicts exist, and it will be re-packaged with them; treat the relation labels as machine-verified only.
Validation report: `reports/validation.json`. Audit report: `reports/audit.json`. Splits: `reports/split_manifest.json`. Release gates checked: rightscomplete=pass, privacyreviewcomplete=pass, splitintegrity=pass, logicalexamplesperimage=pass, artifactintegrity=pass, humanaudit=fail, lexiconreview=fail, declaredscale=fail, releasesize=pass.
Real-image acceptance in this build: {'real': {'accepted': 5778, 'images': 20823, 'rate': 0.2775}, 'rendered': {'accepted': 28986, 'images': 40000, 'rate': 0.7247}}. Relations available: {'above': 33571, 'behind': 177397, 'below': 33571, 'far': 4975, 'infrontof': 177397, 'largerthan': 104893, 'leftof': 46479, 'near': 81090, 'nottouching': 199465, 'occludes': 25414, 'rightof': 46479, 'smaller_than': 104893, 'touching': 9364}.
Intended use
Continual pretraining and evaluation of Thai vision-language models on spatial language: left/right, above/below, in front of/behind, near/far, contact, size comparison, and containment, in image, object, and world reference frames.
Limitations
- Image-frame left/right and above/below are reliable only where the boxes are clearly separated and aligned; the dataset does not claim 3-D layout from photographs.
- Renderer output is flat-shaded and synthetic: it is a controlled probe for relation contrasts, not a substitute for photographs.
- Relation balance reflects what the source annotations support. Under-represented relations are supplemented by renders, not invented for photographs.
- Thai text is template generated.
- Descriptions such as
เสื้อสีเขียวcome from a pinned vision model, not from the source annotation. A spot check of six objects found five correct; an attribute is only used when it tells one object apart from the others in the same image, so an ambiguous description is dropped rather than applied to the wrong object. Treat the attribute text as model-generated.
Version and citation
Version 1.0.0, built by run release-v13 with config thai_spatial_reasoning_v1.yaml (config hash dff750b131deec59, code ad8324a6dcd67a2a0780b05eb0f139108f3a19a3). See the project's CITATIONS.md for the methodology references.
Removal and contact
To request removal of an image or to report a rights problem, open an issue in the project repository or contact the dataset maintainers (smartwhatt/thaispatialreasoning). Each image row carries source_url, image_license, and rights_evidence_uri, which is enough to identify and remove the exact image in a future revision.
Token budget
Measured with the Qwen/Qwen3-VL-2B-Instruct image processor over every one of the 34,764 images in this release, so these are the tokens a Qwen3-VL model actually sees, not an estimate. Text is counted with the processor tokenizer over captions, questions and answers.
Total: 26,764,589 tokens — 18,970,783 visual plus 7,793,806 text, an average of 545 visual and 224 text tokens per image.
Real images cost far more than renders: 1,778 visual tokens per image on average against 300 for a render, because a 1240x1754 photograph is not a 640x480 scene. The training splits are render-heavy, so most of the training budget today is the fixed 300-token renders, and the real photographs are where the per-image weight sits.
The continuous-pretraining index is a separate figure: 243,348 rows carrying 8,175,985 text tokens, and about 140,800,645 tokens if every row trains with its image attached. Quote that number when a training recipe asks for a slice size.
