blackcow63/spreadsheetlabeller
0
SpreadsheetGraph — Labeling Tool & Training Pipeline
Hierarchical knowledge graph construction from spreadsheets for context-preserving retrieval.
Quick Start (Labeling Tool)
1. Google credentials
See SETUP.md for full GCP setup including Drive folders and the metadata sheet.
Minimal local setup:
- Go to Google Cloud Console
- Create a project (or use existing), enable Google Sheets API and Google Drive API
- Create a Service Account → download JSON key
- Save it as
credentials/service_account.json - Share the Drive root folder and metadata sheet with the service account email
2. Install & run
pip install -r requirements.txt
python app.pyOpen http://localhost:8000 in your browser.
3. Labeling workflow
- Enter your name in the top-right input
- Click cells to select (Shift+click for range, Ctrl/Cmd+click to toggle)
- Press keyboard shortcuts to label:
- V = Value (numeric data point)
- D = Attribute (categorical/text data)
- C = Column Header at active level (yellow shades); Shift+C = one level up
- R = Row Header at active level (orange shades); Shift+R = one level up (group header shortcut)
- A = Aggregation (formula/computed cell)
- M = Metadata / Title (blue)
- E = Empty (structural spacer)
- X = Clear label
- T + digit 0–9 = set active table (T0 = sheet-level metadata, T1+ = data tables)
- L + digit 1–9 = set active header level (L1 = most specific, L2+ = broader)
- Press Enter or click SAVE & NEXT to save and load the next sheet
Labels are stored with table IDs and header levels (e.g. col_header_2, row_header_1). Labeled data is saved as JSON in data/labeled/ and uploaded to Google Drive.
See LABELING_GUIDE.md for the full labeling guide with examples (available in English and Polish).
Deployment
Railway / Hugging Face Spaces
- Push to GitHub
- Connect to Railway or Hugging Face Spaces
- Set environment variables (see SETUP.md for the full list):
GOOGLE_SERVICE_ACCOUNT_JSON— service account key JSONDRIVE_ROOT_FOLDER_ID,METADATA_SHEET_ID, etc.APP_PASSWORD(optional gate)- Deploy
Local
cp .env.example .env # fill in your IDs
python app.pyTraining Pipeline
# 1. Extract structural features from labeled JSON → parquet
python features.py
# 2. Train LightGBM baseline (cell role classification)
python train_baseline.py
# 3. Train GNN model (cell classification + edge prediction)
python train_gnn.py
# 4. Generate text chunks from labeled data
python chunk_builder.py
# 5. Run evaluation (graph chunks vs. naive flat extraction)
python evaluate.pyData flow
xlsx on Drive
→ xlsx_client (parse cells, formulas, styles)
→ labeling UI (human assigns labels, tables, levels)
→ labeled JSON (cells with label, table, formatting)
→ features.py (structural features → parquet)
→ train_gnn.py / train_baseline.py (models)
→ predict.py (GNN inference → prefill UI labels)
→ chunk_builder.py (labeled JSON → RAG chunks)