Team Ai
Apppublic

Piyush23890/Sign_Language_Decoder

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes
README.md380 linesDownload Raw Back to root
1---2license: mit3title: sign decoder4sdk: streamlit5emoji: πŸ‘€6colorFrom: yellow7colorTo: yellow8python_version: "3.10"9---10# 🀟 SignBridge β€” Indian Sign Language Smart Communication System11 12> Real-time ISL gesture β†’ Text β†’ Speech, with live English ↔ Hindi translation.  13> Single-click desktop app β€” no Python required for end users.14 15![System Active](https://img.shields.io/badge/status-active-10b981?style=flat-square)16![Python](https://img.shields.io/badge/Python-3.10%2B-3b82f6?style=flat-square&logo=python)17![Flask](https://img.shields.io/badge/Flask-3.0-white?style=flat-square&logo=flask)18![MediaPipe](https://img.shields.io/badge/MediaPipe-0.10-orange?style=flat-square)19![License](https://img.shields.io/badge/license-MIT-green?style=flat-square)20 21---22 23## Table of Contents241. [What is SignBridge?](#what-is-signbridge)252. [Features](#features)263. [System Architecture](#system-architecture)274. [Project Structure](#project-structure)285. [Quick Start (End Users)](#quick-start-end-users)296. [Developer Setup](#developer-setup)307. [Data Collection](#data-collection)318. [Training Models](#training-models)329. [Running the App](#running-the-app)3310. [Building the .exe](#building-the-exe)3411. [Controls & Keyboard Shortcuts](#controls--keyboard-shortcuts)3512. [Tech Stack](#tech-stack)3613. [Troubleshooting](#troubleshooting)37 38---39 40## What is SignBridge?41 42SignBridge converts **Indian Sign Language (ISL)** hand gestures captured via a standard webcam into readable English text and spoken audio β€” in real time.43 44It bridges the communication gap between India's ~6.3 million hearing-impaired ISL users and the general public, requiring no specialist hardware beyond a laptop camera.45 46---47 48## Features49 50| Feature | Detail |51|---|---|52| **Static sign recognition** | A–Z alphabet via Random Forest (126 MediaPipe landmark features) |53| **Dynamic word recognition** | "Hello", "Thank You" via LSTM β†’ ONNX (30-frame sequences) |54| **Motion-based switching** | Automatically picks static or dynamic mode β€” no buttons needed |55| **Smart sentence builder** | Auto-spacing, backspace, clear β€” builds natural sentences |56| **Live translation** | English ↔ Hindi via `deep-translator` (Google) |57| **Text-to-Speech** | Browser Web Speech API (no server round-trip) |58| **Speech-to-Text** | Browser Web Speech API β†’ optional Hindi translation |59| **WebSocket UI** | Real-time updates via Flask-SocketIO (no page refreshes) |60| **Single-click .exe** | PyInstaller bundle for Windows β€” no Python needed |61| **Git LFS ready** | Large model files tracked correctly |62 63---64 65## System Architecture66 67```68Webcam Frame (OpenCV)69        β”‚70        β–Ό71MediaPipe Hands72  21 landmarks Γ— 3 coords Γ— 2 hands = 126 features73        β”‚74        β–Ό75  Motion Score  =  β€–keypoints_t βˆ’ keypoints_{t-1}β€–76        β”‚77   β”Œβ”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”78   β”‚ motion < 0.10           β”‚ motion β‰₯ 0.1079   β–Ό                         β–Ό80Random Forest            ONNX LSTM81(static A–Z)         (dynamic words)82   β”‚                         β”‚83   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜84              β–Ό85     SentenceBuilder86  (auto-space Β· backspace Β· clear)87              β”‚88              β–Ό89    Deep-Translator  (EN ↔ HI)90              β”‚91        β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”92        β–Ό           β–Ό93  WebSocket      Video Feed94  (SocketIO)     (MJPEG stream)95        β”‚96        β–Ό97   Browser UI98  (HTML + JS)99        β”‚100   Web Speech API101  (TTS + STT)102```103 104---105 106## Project Structure107 108```109SignBridge/110β”‚111β”œβ”€β”€ app.py                      ← Main Flask application (run this)112β”œβ”€β”€ sentence_builder.py         ← SentenceBuilder class (also importable)113β”‚114β”œβ”€β”€ templates/115β”‚   └── index.html              ← UI template (Jinja2)116β”‚117β”œβ”€β”€ static/118β”‚   β”œβ”€β”€ style.css               ← Dark glassmorphism styles119β”‚   └── script.js               ← WebSocket + STT/TTS + clipboard JS120β”‚121β”œβ”€β”€ dataset/                    ← Static landmark CSVs (created by collection)122β”‚   β”œβ”€β”€ A/data.csv123β”‚   β”œβ”€β”€ B/data.csv124β”‚   └── ...Z/data.csv125β”‚126β”œβ”€β”€ dynamic_dataset/            ← Dynamic .npy sequences127β”‚   β”œβ”€β”€ hello/128β”‚   β”‚   β”œβ”€β”€ 0.npy ... N.npy129β”‚   └── thank_you/130β”‚       β”œβ”€β”€ 0.npy ... N.npy131β”‚132β”œβ”€β”€ isl_alphabet_model.pkl      ← Trained static RF model  (Git LFS)133β”œβ”€β”€ dynamic_sign_model.h5       ← Trained LSTM model       (Git LFS)134β”œβ”€β”€ dynamic_sign_model.onnx     ← ONNX version for runtime (Git LFS)135β”‚136β”œβ”€β”€ hand_landmarks_dataset.py   ← Collect static data137β”œβ”€β”€ collect_dynamic_data.py     ← Collect dynamic data138β”œβ”€β”€ merge_dataset.py            ← Merge per-letter CSVs β†’ final_dataset.csv139β”œβ”€β”€ train_model.py              ← Train Random Forest140β”œβ”€β”€ train_dynamic_model.py      ← Train LSTM141β”œβ”€β”€ convert_to_onnx.py          ← Convert .h5 β†’ .onnx142β”œβ”€β”€ run_setup_wizard.py         ← All-in-one first-time setup143β”œβ”€β”€ live_predict.py             ← Standalone webcam prediction demo144β”‚145β”œβ”€β”€ test_sentence_builder.py    ← Unit tests146β”œβ”€β”€ requirements.txt147β”œβ”€β”€ SignBridge.spec             ← PyInstaller build spec148β”œβ”€β”€ .gitignore149└── .gitattributes              ← Git LFS config150```151 152---153 154## Quick Start (End Users)155 1561. Download `SignBridge.exe` from the `dist/` folder (or the GitHub Release).1572. Double-click `SignBridge.exe`.1583. Your browser opens automatically at `http://127.0.0.1:5000`.1594. Show your hand to the webcam and start signing βœ‹.160 161> **No Python, no installation required.**162 163---164 165## Developer Setup166 167### 1 β€” Clone the repository168 169```bash170git lfs install171git clone https://github.com/HetviPandav123/sign-language-smart-communication.git172cd sign-language-smart-communication173git lfs pull        # download model files tracked via LFS174```175 176### 2 β€” Create a virtual environment177 178```bash179python -m venv venv180 181# Windows182venv\Scripts\activate183 184# Linux / macOS185source venv/bin/activate186```187 188### 3 β€” Install dependencies189 190```bash191pip install -r requirements.txt192```193 194> **GPU users:** Replace `tensorflow` with `tensorflow-gpu` in requirements.txt  195> **CPU-only machines:** Use `tensorflow-cpu` instead196 197---198 199## Data Collection200 201### Option A β€” Automated wizard (recommended for first-time setup)202 203```bash204python run_setup_wizard.py205```206 207The wizard guides you through:208- Recording 50 samples per letter (A–Z) with on-screen prompts209- Recording 30 gesture sequences each for "Hello" and "Thank You"210- Training both models automatically after collection211 212### Option B β€” Manual collection213 214**Static signs (A–Z):**215```bash216# Collect 200 samples for letter A217python hand_landmarks_dataset.py --sign A --samples 200218 219# Repeat for B through Z220python hand_landmarks_dataset.py --sign B --samples 200221# ...222```223 224**Dynamic gestures:**225```bash226python collect_dynamic_data.py --action hello     --samples 200227python collect_dynamic_data.py --action thank_you --samples 200228```229 230> **Tips for good data:**231> - Use consistent lighting (avoid backlighting)232> - Vary hand distance (30–80 cm from camera)233> - Slightly vary the angle between samples for robustness234 235---236 237## Training Models238 239### Train static Random Forest model240 241```bash242python train_model.py243```244 245Output: `isl_alphabet_model.pkl` + `label_map.pkl`246 247### Train dynamic LSTM model248 249```bash250python train_dynamic_model.py251```252 253Output: `dynamic_sign_model.h5`254 255### Convert LSTM β†’ ONNX (required for app.py)256 257```bash258python convert_to_onnx.py259```260 261Output: `dynamic_sign_model.onnx`262 263### Merge static CSVs (optional β€” for inspection)264 265```bash266python merge_dataset.py267```268 269Output: `final_dataset.csv`270 271---272 273## Running the App274 275```bash276python app.py277```278 279The browser opens automatically at `http://127.0.0.1:5000`.280 281> If the browser doesn't open, navigate there manually.282 283---284 285## Building the .exe286 287Requires both models to be trained and present first.288 289```bash290pip install pyinstaller291pyinstaller SignBridge.spec292```293 294The executable is written to `dist/SignBridge.exe`.295 296> The spec file excludes TensorFlow from the bundle (it is mocked at runtime)  297> which keeps the .exe size manageable. Only `onnxruntime` is bundled for inference.298 299---300 301## Controls & Keyboard Shortcuts302 303### In-app buttons304 305| Button | Action |306|---|---|307| πŸ”Š Speak | Read the sentence aloud (Web Speech API) |308| ⌫ Backspace | Delete last letter or word |309| βœ– Clear | Reset the entire sentence |310| 🎀 Tap to Speak | Toggle Speech-to-Text |311| πŸ“‹ Copy | Copy sentence / STT text to clipboard |312| Language selector | Switch display between English and Hindi |313 314### Keyboard shortcuts315 316| Key | Action |317|---|---|318| `Enter` | Speak the current sentence |319| `Backspace` | Delete last token |320| `C` | Clear sentence |321 322---323 324## Tech Stack325 326| Category | Technology | Purpose |327|---|---|---|328| Language | Python 3.10+ | Core backend |329| Computer Vision | OpenCV 4.x | Webcam capture, MJPEG streaming |330| Hand Tracking | MediaPipe Hands | 21-point landmark detection per hand |331| Static ML | Scikit-learn RandomForest | Letter classification A–Z |332| Dynamic DL | TensorFlow/Keras LSTM | Word-level gesture sequence recognition |333| Inference | ONNX Runtime | Fast, TF-free inference in production |334| Web Backend | Flask + Flask-SocketIO | HTTP routes + WebSocket push updates |335| Frontend | HTML5 / CSS3 / Vanilla JS | UI, camera feed display |336| TTS / STT | Web Speech API (browser) | No server latency |337| Translation | deep-translator (Google) | EN ↔ HI live translation |338| Packaging | PyInstaller | Single .exe for Windows distribution |339| Large Files | Git LFS | .pkl / .h5 / .onnx / .exe version control |340 341---342 343## Troubleshooting344 345### "Paging file too small" on Windows346Caused by Flask debug mode + memory-mapped files.  347**Fix:** `app.py` already sets `debug=False`. If you still see this, restart your PC to clear the paging file.348 349### Webcam not detected350```bash351# Check which index works (try 0, 1, 2)352python -c "import cv2; cap=cv2.VideoCapture(0); print(cap.isOpened())"353```354Change `cv2.VideoCapture(0)` in `app.py` to the correct index.355 356### MediaPipe import error in .exe357The `SignBridge.spec` already includes `collect_all('mediapipe')` and `sys._MEIPASS` path resolution. Rebuild with the provided spec β€” do not use `pyinstaller app.py` directly.358 359### Static sign flickering360Each letter is locked after `STATIC_FRAMES=5` stable frames and won't repeat until the hand moves away. Increase `STATIC_FRAMES` in `app.py` for stricter locking.361 362### Poor recognition accuracy363- Collect more samples per sign (200+ recommended)364- Vary lighting conditions during collection365- Ensure both hands are visible for two-handed signs366- Re-train with the new data367 368### Translation not working369Requires an internet connection. `deep-translator` uses Google Translate API. If offline, the raw English sentence is shown as fallback.370 371---372 373## Author374 375**Hetvi Pandav**  376BE – Artificial Intelligence & Machine Learning377 378---379 380⭐ If SignBridge helped you, star the repo!