BabyRider/hitl-prompt-optimization-api
🌍 HITL Prompt Optimization API
Detect ambiguity in natural language and generate an ambiguity-preserving prompt (so a model answers each valid interpretation instead of collapsing to one).
Quick links: API docs · Run locally · Test · Deploy on Spaces
⚡ TL;DR (Docker)
docker build -t hitl-prompt-optimization-api .
docker run --rm -p 7860:7860 hitl-prompt-optimization-apiThen:
- Swagger UI: http://localhost:7860/docs
- Analyze:
POST http://localhost:7860/analyze/
This repo exposes a small HTTP API that takes free-form text and returns:
- Detected ambiguity categories (A/B/C/D)
- Character spans where ambiguity likely occurs
- A rewritten “final prompt” that preserves ambiguity by asking for multiple interpretations rather than collapsing to one
It’s designed for HITL (human-in-the-loop) workflows where you want an assistant to explore interpretations instead of prematurely disambiguating.
Hugging Face Spaces config reference
🧰 Core technologies
- API framework: FastAPI + Uvicorn
- ML runtime: PyTorch
- NLP models: Hugging Face Transformers (BERT tokenizer/model family)
- Models included in repo:
models/ambiguity_classifier/model.pt(multi-label ambiguity classifier)models/span_detector/span_detector.pt(span/sequence tagger)- Heuristics: rule-based math ambiguity span extraction
🗂️ Project structure (high level)
ml_backend/api/main.py: FastAPI app and HTTP routespipeline/narrative_processor.py: sentence splitting + per-sentence inference + prompt constructionmodels/ambiguity_classifier/: predicts ambiguity class(es) for a sentencemodels/span_detector/: predicts word-level spans and maps them back to character offsetsagents/resolver/: class-specific prompt “strategies” that preserve ambiguitypipeline/math_span_heuristics.py: rule-based extraction for arithmetic/operator ambiguity
🧠 Methods used (what the system is doing)
1) Sentence segmentation
pipeline/narrative_processor.py uses a simple regex splitter (./!/?) to break the input into sentences.
2) Ambiguity classification (A/B/C/D)
models/ambiguity_classifier/infer_classifier.py loads a fine-tuned classifier and predicts one or more ambiguity classes.
Classes used in the codebase:
- A: implicit / arithmetic ambiguity (missing or unclear operations)
- B: structural / attachment ambiguity
- C: referential ambiguity (e.g., pronouns with multiple antecedents)
- D: symbol–meaning ambiguity (multiple semantic readings)
The classifier is multi-label (more than one class can be predicted for the same sentence).
3) Span detection
models/span_detector/infer_span_detector.py runs a token/sequence tagger that predicts BIO-style labels per word (e.g., B-A, I-B, O). It then maps word spans back to character offsets (start, end) for easier UI highlighting.
4) Heuristic math spans (augmentation)
If the classifier predicts math/implicit ambiguity (A, and in the narrative processor also B), pipeline/math_span_heuristics.py adds extra spans for operator words (e.g., times, divided) and simple patterns (e.g., four times two).
5) “Preserve ambiguity” prompt construction
The output prompt is intentionally written to elicit multiple interpretations.
- API path uses
pipeline/narrative_processor.build_narrative_resolver_prompt() - There is also an alternate “resolver agent” implementation in
agents/resolver/used bypipeline/run_pipeline.py
🔌 API
Endpoints
GET /→ health checkPOST /analyze/→ analyze text and return ambiguity-aware output
Request/response schema
Request:
{ "text": "..." }Response:
{
"original_text": "...",
"sentences": ["..."],
"ambiguities": [
{
"sentence_index": 0,
"sentence": "...",
"classes": ["A", "C"],
"spans": [{ "text": "...", "start": 0, "end": 10, "class": "A" }]
}
],
"final_prompt": "..."
}FastAPI docs when running locally:
GET /docs(Swagger UI)GET /openapi.json
📦 Dependencies (packages + versions)
Direct dependencies are listed in ml_backend/requirements.txt.
- Pinned:
torch==2.10.0+cpu(installed from the PyTorch CPU wheel index)- Unpinned (installed at latest compatible versions unless you pin them):
fastapi,uvicorn,transformers,tqdm,regex,pydantic,slowapi
To record the exact versions you have installed locally, run:
pip list --format=freezeIf you want fully reproducible builds, consider generating and checking in a lock file (e.g., requirements.lock) from that output.
🛠️ Build and run locally
You can run this either via Docker (recommended) or directly with Python.
Option A: Docker (recommended)
docker build -t hitl-prompt-optimization-api .
docker run --rm -p 7860:7860 hitl-prompt-optimization-apiThen open:
- http://localhost:7860/ (health)
- http://localhost:7860/docs (Swagger)
Option B: Python (venv)
From the repo root:
python -m venv .venv
./.venv/Scripts/activate
pip install -r ml_backend/requirements.txt
uvicorn ml_backend.api.main:app --host 0.0.0.0 --port 7860✅ Test locally (URLs + examples)
Health check:
curl http://localhost:7860/Analyze endpoint:
PowerShell:
curl http://localhost:7860/analyze/ -Method POST -ContentType 'application/json' -Body '{"text":"After Leah informed Morgan that the server was overloaded, they initiated a restart."}'POSIX shells (bash/zsh):
curl -X POST http://localhost:7860/analyze/ \
-H "Content-Type: application/json" \
-d "{\"text\":\"After Leah informed Morgan that the server was overloaded, they initiated a restart.\"}"Narrative test runner
There’s also a simple script that prints out the generated ambiguity-preserving prompt:
python run_narrative_tests.py🚀 Publish / deploy on Hugging Face Spaces (Docker)
This repo is already configured for Spaces with sdk: docker in the README front-matter and a Dockerfile that runs Uvicorn on port 7860.
Steps:
- Create a new Space on Hugging Face
- SDK: Docker
- Push this repo to the Space (either via Git or by connecting a GitHub repo)
- Spaces will build the Docker image and run it automatically
- Once running, your Space will expose the same endpoints:
/health/docsSwagger/analyze/POST
If you change dependencies, Spaces will rebuild automatically on the next push.
🧊 Notes on CUDA packages in Spaces builds
If your Space logs show nvidia-*-cu12 packages being installed even though you didn't list them, it's usually because pip install torch on Linux resolves to a CUDA-enabled PyTorch wheel by default.
This repo's Docker build installs CPU-only PyTorch using the PyTorch CPU wheel index in Dockerfile. If you deploy on a GPU Space and want CUDA-enabled PyTorch, remove that CPU-only install step and install torch normally.
