mokshak/vera-rubric-decision-engine
Vera Rubric Decision Engine
Submission bot for the magicpin Vera AI Challenge. It exposes the required HTTP API, stores judge-pushed context in memory, and composes grounded merchant/customer actions from the JSON context it receives.
The submitted live bot is deterministic and does not require a paid API key or any runtime LLM call. OpenRouter/LLM tooling is used only for offline calibration and local judging experiments, never as a dependency for /v1/tick.
How It Works
- The judge pushes category, merchant, customer, and trigger JSON into
POST /v1/context. POST /v1/tickranks active triggers by expected rubric score and returns up to 20 message actions.POST /v1/replyhandles auto-replies, commitment, off-topic replies, STOP/hostility, and ended conversations.POST /v1/teardownis an extra helper for clean local/HF reruns; it is not required by the challenge judge.
Challenge-required endpoints:
GET /v1/healthzGET /v1/metadataPOST /v1/contextPOST /v1/tickPOST /v1/reply
Additional debug/test endpoint:
POST /v1/teardown
Compatibility aliases without the /v1 prefix are also exposed (/healthz, /metadata, /context, /tick, /reply) because the public challenge page mentions both shorthand endpoint names and the /v1/... submission contract. The submitted base URL remains the same.
Scoring Strategy
The composer is a rubric-optimized decision engine:
- Extracts evidence from merchant, category, trigger, and customer context.
- Flattens nested context into a MerchantCAR so every downstream decision sees stable field names.
- Uses JITAI-style severity/receptivity/intervention-fit scores to decide whether the moment is worth acting on.
- Uses Fogg B=MAP scores to separate motivation, CTA ability, and prompt timing.
- Applies prospect-theory plus Cialdini framing: loss recovery, gain momentum, scarcity, social proof, authority, reciprocity, commitment, or liking.
- Structures copy with deterministic PAS/BAB/scarcity frameworks so the body has a fact, a cost/opportunity hook, and one concrete reply action.
- Converts soft CTAs into commitment CTAs such as
Reply YES and I'll prepare... now. - Suppresses weak merchant-only trigger moments when the context has too little why-now evidence, rather than sending a generic fallback.
- Uses a pharmacy-safe body contract for merchant-facing pharmacy copy: stock, refill, batch, expiry, checklist, audit, or compliance anchors only.
- Builds deterministic Tree-of-Thought frame diagnostics and Best-of-N variants for each trigger.
- Runs a deterministic constitutional audit/repair pass for generic copy, multiple CTAs, corporate tone, weak facts, and repeated weak action types.
- Uses empirical category action priors from in-memory reply outcomes to reduce cold-start action mistakes without random live exploration.
- Scores each candidate across decision quality, specificity, category fit, merchant fit, and engagement compulsion.
- Sends only the highest-scoring valid action.
- Uses category playbooks for dentists, salons, restaurants, gyms, and pharmacies.
- Validates output for hallucinated numbers,
Noneleaks, repeatedDr. Dr., weak generic copy, repeated bodies, missing CTA shape, and unsafe customer outreach.
Every sent message should include why now, a real merchant/category/customer fact, one CTA, category-appropriate voice, and a low-friction next action.
Run Locally
python -m pip install -r requirements-dev.txt
python dataset/generate_dataset.py --seed-dir dataset --out expanded
uvicorn app.main:app --host 0.0.0.0 --port 8080In another terminal:
pytest -q
python -m compileall app bot.py scripts tests
python scripts/generate_submission.py
python scripts/lint_submission.py
python scripts/score_proxy.py 43
python scripts/geval_calibrate.pyscripts/geval_calibrate.py skips cleanly unless OPENROUTER_API_KEY is set. The official judge_simulator.py also needs a scorer LLM key. The hosted bot itself does not need one.
For LLM judging with a floor-case table, use the wrapper instead of editing the official simulator:
set OPENROUTER_API_KEYS=key1,key2
set BOT_URL=https://mokshak-vera-rubric-decision-engine.hf.space
set JUDGE_SCENARIO=all
python scripts/run_judge_with_trigger_log.pyIt prints the lowest-scoring trigger IDs and can write JSON when JUDGE_TRIGGER_LOG_JSON is set.
Optional Local LLM Copy Polish
The submitted HF deployment does not use runtime LLM polish. Set these only for local experiments:
OPENAI_API_KEYOPENAI_MODEL, defaultgpt-4o-mini
The model receives only the deterministic plan and evidence. It must return structured JSON with body, cta, send_as, suppression_key, and rationale. The bot falls back to deterministic copy if the call fails, times out, changes protected fields, or invents numbers. Do not enable this in the submitted Space; deterministic runtime is the reliability tradeoff.
Optional OpenRouter Calibration
Set these only for offline quality checks:
OPENROUTER_API_KEYOPENROUTER_MODEL, defaultopenrouter/autoGEVAL_LIMIT, default10
This script critiques generated submission.jsonl rows against the five official dimensions with a Prometheus-style reference bank and reports low-scoring cases. It is intentionally not called by /v1/tick.
Free Deployment
Primary target: Hugging Face Docker Space.
Current hosted base URL:
https://mokshak-vera-rubric-decision-engine.hf.spaceStart command:
uvicorn app.main:app --host 0.0.0.0 --port $PORTHealth check path:
/v1/healthzUse one process/worker only. State is in memory, so multiple workers would split conversations and suppression keys.
Backup targets:
- Koyeb Free Instance
- Render Free Web Service using
render.yaml
Free hosts can sleep. Before submission, hit /v1/healthz, then keep it warm:
python scripts/keep_warm.py https://your-bot.example --interval 900Required Submission Details
Set these deployment environment variables:
CONTACT_EMAIL:mokshagnak004@gmail.comTEAM_NAME: optional display nameTEAM_MEMBER: optional member nameSUBMITTED_AT: optional ISO timestamp
You do not need OpenAI, Groq, Gemini, OpenRouter, Redis, Postgres, or any paid key for the live deterministic bot.
Tradeoffs
In-memory state is simple, private, and fast, but the service must run as a single worker and should not restart during judging. Free LLM APIs are useful for local experiments, but relying on free quota during live judging is risky, so this bot treats LLM usage as optional polish only.
What More Context Would Help
Quality would improve further with verified customer consent mappings, real slot inventory, item-level order or dispense history, offer eligibility rules, customer segment aggregates, and locality-level peer benchmarks. When those facts are absent, the bot avoids inventing them.
