vinaylodhi1712/text2sql-api
0
NL2SQL — Natural Language to SQL (Multi-Agent)
Lightweight NL2SQL assistant that converts plain-English questions into validated SQLite SQL and executes them against a sample e‑commerce dataset or user-uploaded CSV/XLSX files.
Features
- Domain routing (customers / orders / products)
- Query vagueness detection and automatic suggestion
- LLM-driven SQL generation with schema-aware prompts
- SQL validation and fuzzy literal correction
- Upload CSV/XLSX support with temporary in-memory query execution
Quick start (backend)
Windows (PowerShell)
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
uvicorn api:app --reload --port 8000Linux / macOS
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn api:app --reload --port 8000Frontend (React + Vite)
cd frontend
npm install
npm run dev # development server (http://localhost:5173)
# or production build
npm run build
npm run previewEnvironment configuration
- Copy
frontend/.env.exampletofrontend/.env. - Set
VITE_API_URLto your backend host if not usinghttp://localhost:8000. - For the backend, set
GROQ_API_KEYif you want LLM-generated suggested queries.
Repository layout
api.py— FastAPI endpoints and upload/session handlinggraph.py— State graph setup for the agent pipelineagents/— Individual agent modules (router_agent.py,sql_generator.py,sql_executor.py, etc.)session_manager.py— temporary session storage and CSV/XLSX helperstext2sql.db— sample SQLite database used by the demofrontend/— React + Tailwind UI
System architecture (workflow)
flowchart TD
User[User] --> UI[Frontend: React / Streamlit]
UI --> API[FastAPI]
API --> Graph[StateGraph pipeline]
Graph --> Router[Router Agent]
Router --> Filter[Filter Check Agent]
Filter --> SQLGen[SQL Generator Agent]
SQLGen --> Validator[Query Validator Agent]
Validator --> Fuzzy[Fuzzy Matcher Agent]
Fuzzy --> Exec[SQL Executor Agent]
Exec --> DB[SQLite text2sql.db / in-memory upload DB]
DB --> Exec
Exec --> UIResearch papers & related work
- "Seq2SQL" — Zhong et al., 2017 — early sequence-to-SQL model ideas.
- "SQLNet" — Xu et al., 2017 — sketch-based SQL generation improving over Seq2Seq.
- "Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task" — Yu et al., 2018 — dataset that many modern text-to-SQL systems use.
- "IRNet" — Guo et al., 2019 — schema linking and bridging natural language to SQL.
- "PICARD" — Scholak et al., 2021 — constrained decoding for improved SQL generation (works with LLMs).
Notes & next steps
- Consider replacing the LLM prompt pipeline with a fine-tuned seq2seq model for deterministic SQL generation when reproducibility matters.
- Add unit tests for agent outputs to catch regressions in prompt changes.
- For production, move session storage from in-memory
SessionManagerto Redis and secure file uploads.
