Team Ai
Datasetpublic

deb0naire/Bodhisetu

Bodhisetu A multimodal cultural-heritage dataset from India, collected via the SMRITI field-data-collection platform. Overview Bodhisetu documents 3,000+ heritage entities across 6 Indian languages (Tamil, Kannada, Assamese, Maithili, Bodo, Nepali) and 14 taluks spanning 5 states. Each entity is captured as one or more images and short-form videos, accompanied by field descriptions. This release contains 5,753 resources (4,461 images + 1,292 videos, ~91 GB)… See the full description on the dataset page: https://huggingface.co/datasets/deb0naire/Bodhisetu.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes408downloads
Dataset Card

Bodhisetu

A multimodal cultural-heritage dataset from India, collected via the SMRITI field-data-collection platform.

Overview

Bodhisetu documents 3,000+ heritage entities across 6 Indian languages (Tamil, Kannada, Assamese, Maithili, Bodo, Nepali) and 14 taluks spanning 5 states. Each entity is captured as one or more images and short-form videos, accompanied by field descriptions.

This release contains 5,753 resources (4,461 images + 1,292 videos, ~91 GB), stratified over (taluk, media_type) to match the full-corpus proportions from the SMRITI pilot.

Two views of the same data

  • —`raw` — fields as submitted by collectors in the field (descriptions in the speaker's voice, user-provided tags, native-language names).
  • —`enriched` — AI-curated versions produced by the SMRITI enrichment pipeline (refined descriptions grounded against web + Wikipedia, structured taxonomy tags, normalised entity names).

Both configs reference the same media files. Joining the two parquets on entity_uid gives the raw↔enriched comparison for each entity.

Sampling

The 5,753 resources in this release were drawn from the Smriti pilot via proportion-matched stratified sampling on (taluk × media_type).

  1. 1.Reference distribution. For each of the 14 Phase-1 taluks and each media type (IMAGE, VIDEO), we counted resources in the raw Smriti collection (29,575 resources over 28 cells). Per-cell shares form the target distribution.
  2. 2.Target pool. The curated (enriched) layer contributes 17,989 eligible resources across the same 28 cells. CONSENT PDFs are excluded.
  3. 3.Strict proportion matching. Corpus size is maximised while preserving proportions exactly: K_max = min over cells of (curated_available_c / raw_share_c). This run was constrained by Dhemaji × VIDEO (178 curated videos, 3.05% raw share) → K_max = 178 / 0.0305 ≈ 5,836.
  4. 4.Domain diversification. Within each (taluk, media_type) cell, resources are sampled across domains with square-root-proportional allocation and a minimum-one-per-domain floor, so small categories (Traditions & Customs, Governance, People & Personalities) are not starved by large categories (Religion, Buildings).
  5. 5.Blob availability filter. 82 resources that failed face-anonymisation (mostly files mis-typed as IMAGE in the source DB but containing H.264 streams) were dropped. Final release: 5,753 resources.
  6. 6.Max per-cell drift from the raw reference: 0.01%. Media split matches the raw corpus exactly at 76.6% images / 23.4% videos.

Seed: 20260507 (deterministic, reproducible).

Loading

python
from datasets import load_dataset

ds_enriched = load_dataset("51Hypers/Bodhisetu", "enriched", split="train")
ds_raw      = load_dataset("51Hypers/Bodhisetu", "raw",      split="train")

Schema (21 columns, identical in both configs)

ColumnTypeNotes
entity_uidstringsynthetic 16-char hex; same across configs
namestringentity name
languagestringISO: TAM, KAN, ASM, MAI, BRX, NEP
talukstringadministrative sub-division
statestringIndian state
domainstringtop-level heritage category (16 domains)
subdomainstringfine-grained category
location_textstringfree-text location
gps_lat, gps_lngfloatoptional
keywords_tagsstring (JSON list)keyword tags
significance_statusstringenum: LOCALSIGNIFICANCE, REGIONALSIGNIFICANCE, …
frequency_seasonalitystringenum
time_period_erastringenum
colloquial_namestring
resource_idstringoriginal resource ID
media_pathstringrelative path under blobs/
media_typestringIMAGE or VIDEO
descriptionstringcaption / description of this resource
transcriptstringvoice-memo transcript, if present
transcript_languagestring

Privacy

All faces in images and videos have been automatically blurred using SCRFD face detection with feathered elliptical Gaussian blur, prior to release. Voice memos are released as transcriptions only (no raw audio). EXIF metadata has been stripped.

License

Released under CC BY 4.0.

Citation

bibtex
@misc{bodhisetu2026,
  title  = {Bodhisetu: A Multimodal Multilingual Cultural-Heritage Dataset from India},
  year   = {2026},
  note   = {Available at https://huggingface.co/datasets/51Hypers/Bodhisetu},
}