hellosindh/indus-script-models
0
1---2language:3- und4license: cc-by-4.05tags:6- indus-script7- ancient-scripts8- archaeology9- nlp10- text-generation11- sequence-modeling12- grammar-analysis13- undeciphered-script14library_name: transformers15pipeline_tag: text-generation16---17 18# Indus Script Models19 20Trained models for validating, predicting, and generating sequences in the undeciphered21Indus Valley Script (2600–1900 BCE). Built on 3,310 real archaeological inscriptions.22 23---24 25## Quick Start (3 steps)26 27```bash28# Step 1 — Clone the repo29git clone https://huggingface.co/hellosindh/indus-script-models30cd indus-script-models31 32# Step 2 — Install dependencies33pip install torch transformers34 35# Step 3 — Run the demo36python inference.py --task demo37```38 39---40 41## What you can do42 43### 1. Validate a sequence44Is this inscription grammatically valid?45 46```bash47python inference.py --task validate --sequence "T638 T177 T420 T122"48```49 50Output:51```52Sequence : T638 T177 T420 T12253BERT : 0.965054N-gram : 0.893055ELECTRA : 0.941056Ensemble : 0.941057Verdict : VALID (>=85%)58```59 60### 2. Predict a masked sign61What sign most likely fills the missing position?62 63```bash64python inference.py --task predict --sequence "T638 [MASK] T420 T122"65```66 67Output:68```69Position 1 predictions:70 T177 18.3%71 T243 12.1%72 T653 9.4%73 T684 7.2%74 T650 5.8%75```76 77### 3. Generate new sequences78 79```bash80# Generate 10 sequences (default threshold 85%)81python inference.py --task generate --count 1082 83# More variety, less strict84python inference.py --task generate --count 20 --threshold 0.7885 86# High quality only87python inference.py --task generate --count 5 --threshold 0.9288```89 90### 4. Score any sequence91 92```bash93python inference.py --task score --sequence "T604 T123 T609"94```95 96---97 98## Generating more diverse or longer sequences99 100Open `inference.py` and find the `task_generate` function. Change the temperature list:101 102**More random — forces rare signs to appear:**103```python104# Change this line:105temps = [0.85, 0.90, 1.00, 1.10]106# To:107temps = [1.10, 1.20, 1.30, 1.40]108```109 110**Longer sequences:**111Find the `generate()` method inside `load_nanogpt()` and change `max_len`:112```python113# Default (avg 7 signs):114def generate(self, temperature=0.85, top_k=40, max_len=15):115 116# For longer sequences:117def generate(self, temperature=0.85, top_k=40, max_len=25):118 119# For shorter sequences:120def generate(self, temperature=0.85, top_k=40, max_len=6):121```122 123---124 125## Pros and cons of tuning126 127| Setting | Effect | Good for | Watch out for |128|---|---|---|---|129| Temperature 0.7–0.8 | Very focused, repeats common signs | High quality outputs | Low diversity |130| Temperature 0.9–1.0 | Balanced — default | General use | Nothing |131| Temperature 1.1–1.3 | More variety, rare signs appear | Exploring vocabulary | Some unusual sequences |132| Temperature above 1.4 | Very random | Stress testing | Most sequences fail quality gate |133| Threshold 0.85 | Strict — default | Publication quality | Slower generation |134| Threshold 0.75 | Relaxed | Larger datasets | Lower average quality |135| Threshold 0.92 | Very strict | Highest confidence only | Very few sequences pass |136| max_len 6 | Short sequences | Matching real length distribution | Misses complex patterns |137| max_len 20+ | Long sequences | Complex grammar patterns | Not representative of real seals |138 139---140 141## Displaying Indus glyphs142 143Sequences use sign IDs like T638, T177. To see actual glyphs:144 1451. Search for **indus-brahmi-font** and download it1462. The `glyphs` field in output shows the rendered glyph characters1473. Open `data/id_to_glyph.json` to see the full sign to character mapping1484. If want to see mapping with T, open `data/indus_tokenizer/indus_id_map.json`149 150Without the font installed, glyphs show as boxes or question marks.151The sign IDs (T638, T177 etc.) always work regardless of font.152 153---154 155## Repo structure156 157```158indus-script-models/159├── inference.py run this for all tasks160├── indus_ngram.py required by ngram_model.pkl — do not move161├── README.md162├── models/163│ ├── nanogpt_indus.pt NanoGPT generator (153K params, PPL 13.3)164│ ├── ngram_model.pkl N-gram RTL model (88.2% pairwise accuracy)165│ ├── mlm/ TinyBERT masked language model (val loss 2.06)166│ ├── cls/ TinyBERT classifier (89.0% test accuracy)167│ ├── electra/ ELECTRA discriminator (95.1% token accuracy)168│ └── deberta/ DeBERTa discriminator (87.1% test accuracy)169└── data/170 ├── id_to_glyph.json 641 sign ID to glyph character mappings171 └── indus_tokenizer/ custom tokenizer for Indus Script172```173 174---175 176## How the pipeline works177 178**Stage 1 — Train on 3,310 real inscriptions:**179 180Four models trained independently, each learning a different aspect of grammar:181 182- **TinyBERT MLM** — learns which sign can fill a masked position in a sequence183- **TinyBERT Classifier** — learns to tell valid sequences from corrupted ones184- **N-gram RTL** — learns right-to-left transition probabilities between signs185- **ELECTRA** — learns token-level discrimination between real and fake signs186- **NanoGPT** — learns to generate new sequences from scratch187 188**Stage 2 — Generate and filter:**189 190NanoGPT generates candidate sequences in RTL order, then flips them to LTR.191Each candidate is scored by three models: BERT (50%) + N-gram (25%) + ELECTRA (25%).192Only sequences scoring 85% or higher are kept as valid synthetic sequences.193Sequences that exactly match real inscriptions are separated as seal reproductions.194Result: 5,000 novel sequences with 752 exact seal matches as validation evidence.195 196**Stage 3 — Retrain on combined data:**197 198The 5,000 synthetic sequences were combined with 3,310 real sequences (8,310 total).199All models were retrained on the larger dataset. Results improved significantly:200 201| Model | Before | After |202|---|---|---|203| TinyBERT accuracy | 78.4% | 89.0% |204| NanoGPT perplexity | 32.5 | 13.3 |205| DeBERTa accuracy | 80.5% | 87.1% |206 207The final 5,000 sequences in the dataset were generated with these retrained models.208 209---210 211## Key findings212 213- **RTL reading confirmed** — right-to-left has 12% stronger grammatical structure than LTR214- **Grammar proven** — entropy chain H1 to H2 to H3 = 6.03 to 3.41 to 2.39 bits (language-like decay)215- **Zipf law confirmed** — R squared = 0.968, language-like token distribution216- **752 seal reproductions** — model independently reproduced real archaeological inscriptions217- **Sign roles discovered:**218 - PREFIX signs at reading end: T638, T604, T406, T496219 - SUFFIX signs at reading start: T123, T122, T701, T741220 - CORE signs in the middle: T101, T268, T177, T243221 222---223 224## Known limitations225 226**DeBERTa calibration issue:**227DeBERTa scores near-zero for all sequences due to confidence calibration failure.228It is logged in output but excluded from the quality gate.229BERT, N-gram, and ELECTRA handle all scoring.230 231**Vocabulary coverage:**232Only about 26% of the 641 known Indus signs appear reliably in generated sequences.233475 signs appear 10 times or fewer in the real corpus — too rare for the model to learn.234This is a property of the archaeological record, not a model bug.235No synthetic corpus can reliably generate signs that barely exist in the training data.236 237**Short sequences:**238The model rarely generates length-2 sequences even though they are common in real inscriptions.239If you need shorter outputs, set `max_len=4` in the generate function.240 241---242 243## Dataset244 245The 5,000 synthetic sequences with full scores and sign index are available at:246 247[hellosindh/indus-script-synthetic](https://huggingface.co/datasets/hellosindh/indus-script-synthetic)248 