Team Ai
Apppublic

blackcow63/spreadsheetlabeller

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

SpreadsheetGraph — Labeling Tool & Training Pipeline

Hierarchical knowledge graph construction from spreadsheets for context-preserving retrieval.

Quick Start (Labeling Tool)

1. Google credentials

See SETUP.md for full GCP setup including Drive folders and the metadata sheet.

Minimal local setup:

  1. 1.Go to Google Cloud Console
  2. 2.Create a project (or use existing), enable Google Sheets API and Google Drive API
  3. 3.Create a Service Account → download JSON key
  4. 4.Save it as credentials/service_account.json
  5. 5.Share the Drive root folder and metadata sheet with the service account email

2. Install & run

bash
pip install -r requirements.txt
python app.py

Open http://localhost:8000 in your browser.

3. Labeling workflow

  • —Enter your name in the top-right input
  • —Click cells to select (Shift+click for range, Ctrl/Cmd+click to toggle)
  • —Press keyboard shortcuts to label:
  • —V = Value (numeric data point)
  • —D = Attribute (categorical/text data)
  • —C = Column Header at active level (yellow shades); Shift+C = one level up
  • —R = Row Header at active level (orange shades); Shift+R = one level up (group header shortcut)
  • —A = Aggregation (formula/computed cell)
  • —M = Metadata / Title (blue)
  • —E = Empty (structural spacer)
  • —X = Clear label
  • —T + digit 0–9 = set active table (T0 = sheet-level metadata, T1+ = data tables)
  • —L + digit 1–9 = set active header level (L1 = most specific, L2+ = broader)
  • —Press Enter or click SAVE & NEXT to save and load the next sheet

Labels are stored with table IDs and header levels (e.g. col_header_2, row_header_1). Labeled data is saved as JSON in data/labeled/ and uploaded to Google Drive.

See LABELING_GUIDE.md for the full labeling guide with examples (available in English and Polish).

Deployment

Railway / Hugging Face Spaces

  1. 1.Push to GitHub
  2. 2.Connect to Railway or Hugging Face Spaces
  3. 3.Set environment variables (see SETUP.md for the full list):
  4. 4.GOOGLE_SERVICE_ACCOUNT_JSON — service account key JSON
  5. 5.DRIVE_ROOT_FOLDER_ID, METADATA_SHEET_ID, etc.
  6. 6.APP_PASSWORD (optional gate)
  7. 7.Deploy

Local

bash
cp .env.example .env  # fill in your IDs
python app.py

Training Pipeline

bash
# 1. Extract structural features from labeled JSON → parquet
python features.py

# 2. Train LightGBM baseline (cell role classification)
python train_baseline.py

# 3. Train GNN model (cell classification + edge prediction)
python train_gnn.py

# 4. Generate text chunks from labeled data
python chunk_builder.py

# 5. Run evaluation (graph chunks vs. naive flat extraction)
python evaluate.py

Data flow

xlsx on Drive
  → xlsx_client (parse cells, formulas, styles)
  → labeling UI (human assigns labels, tables, levels)
  → labeled JSON (cells with label, table, formatting)
  → features.py (structural features → parquet)
  → train_gnn.py / train_baseline.py (models)
  → predict.py (GNN inference → prefill UI labels)
  → chunk_builder.py (labeled JSON → RAG chunks)