deb0naire/Bodhisetu
Bodhisetu A multimodal cultural-heritage dataset from India, collected via the SMRITI field-data-collection platform. Overview Bodhisetu documents 3,000+ heritage entities across 6 Indian languages (Tamil, Kannada, Assamese, Maithili, Bodo, Nepali) and 14 taluks spanning 5 states. Each entity is captured as one or more images and short-form videos, accompanied by field descriptions. This release contains 5,753 resources (4,461 images + 1,292 videos, ~91 GB)… See the full description on the dataset page: https://huggingface.co/datasets/deb0naire/Bodhisetu.
Bodhisetu
A multimodal cultural-heritage dataset from India, collected via the SMRITI field-data-collection platform.
Overview
Bodhisetu documents 3,000+ heritage entities across 6 Indian languages (Tamil, Kannada, Assamese, Maithili, Bodo, Nepali) and 14 taluks spanning 5 states. Each entity is captured as one or more images and short-form videos, accompanied by field descriptions.
This release contains 5,753 resources (4,461 images + 1,292 videos, ~91 GB), stratified over (taluk, media_type) to match the full-corpus proportions from the SMRITI pilot.
Two views of the same data
- `raw` — fields as submitted by collectors in the field (descriptions in the speaker's voice, user-provided tags, native-language names).
- `enriched` — AI-curated versions produced by the SMRITI enrichment pipeline (refined descriptions grounded against web + Wikipedia, structured taxonomy tags, normalised entity names).
Both configs reference the same media files. Joining the two parquets on entity_uid gives the raw↔enriched comparison for each entity.
Sampling
The 5,753 resources in this release were drawn from the Smriti pilot via proportion-matched stratified sampling on (taluk × media_type).
- Reference distribution. For each of the 14 Phase-1 taluks and each media type (IMAGE, VIDEO), we counted resources in the raw Smriti collection (29,575 resources over 28 cells). Per-cell shares form the target distribution.
- Target pool. The curated (enriched) layer contributes 17,989 eligible resources across the same 28 cells. CONSENT PDFs are excluded.
- Strict proportion matching. Corpus size is maximised while preserving proportions exactly:
K_max = min over cells of (curated_available_c / raw_share_c). This run was constrained byDhemaji × VIDEO(178 curated videos, 3.05% raw share) →K_max = 178 / 0.0305 ≈ 5,836. - Domain diversification. Within each
(taluk, media_type)cell, resources are sampled across domains with square-root-proportional allocation and a minimum-one-per-domain floor, so small categories (Traditions & Customs, Governance, People & Personalities) are not starved by large categories (Religion, Buildings). - Blob availability filter. 82 resources that failed face-anonymisation (mostly files mis-typed as IMAGE in the source DB but containing H.264 streams) were dropped. Final release: 5,753 resources.
- Max per-cell drift from the raw reference: 0.01%. Media split matches the raw corpus exactly at 76.6% images / 23.4% videos.
Seed: 20260507 (deterministic, reproducible).
Loading
from datasets import load_dataset
ds_enriched = load_dataset("51Hypers/Bodhisetu", "enriched", split="train")
ds_raw = load_dataset("51Hypers/Bodhisetu", "raw", split="train")Schema (21 columns, identical in both configs)
Privacy
All faces in images and videos have been automatically blurred using SCRFD face detection with feathered elliptical Gaussian blur, prior to release. Voice memos are released as transcriptions only (no raw audio). EXIF metadata has been stripped.
License
Released under CC BY 4.0.
Citation
@misc{bodhisetu2026,
title = {Bodhisetu: A Multimodal Multilingual Cultural-Heritage Dataset from India},
year = {2026},
note = {Available at https://huggingface.co/datasets/51Hypers/Bodhisetu},
}