Team Ai
Apppublic

mokshak/vera-rubric-decision-engine

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes
App README

Vera Rubric Decision Engine

Submission bot for the magicpin Vera AI Challenge. It exposes the required HTTP API, stores judge-pushed context in memory, and composes grounded merchant/customer actions from the JSON context it receives.

The submitted live bot is deterministic and does not require a paid API key or any runtime LLM call. OpenRouter/LLM tooling is used only for offline calibration and local judging experiments, never as a dependency for /v1/tick.

How It Works

  1. 1.The judge pushes category, merchant, customer, and trigger JSON into POST /v1/context.
  2. 2.POST /v1/tick ranks active triggers by expected rubric score and returns up to 20 message actions.
  3. 3.POST /v1/reply handles auto-replies, commitment, off-topic replies, STOP/hostility, and ended conversations.
  4. 4.POST /v1/teardown is an extra helper for clean local/HF reruns; it is not required by the challenge judge.

Challenge-required endpoints:

  • —GET /v1/healthz
  • —GET /v1/metadata
  • —POST /v1/context
  • —POST /v1/tick
  • —POST /v1/reply

Additional debug/test endpoint:

  • —POST /v1/teardown

Compatibility aliases without the /v1 prefix are also exposed (/healthz, /metadata, /context, /tick, /reply) because the public challenge page mentions both shorthand endpoint names and the /v1/... submission contract. The submitted base URL remains the same.

Scoring Strategy

The composer is a rubric-optimized decision engine:

  • —Extracts evidence from merchant, category, trigger, and customer context.
  • —Flattens nested context into a MerchantCAR so every downstream decision sees stable field names.
  • —Uses JITAI-style severity/receptivity/intervention-fit scores to decide whether the moment is worth acting on.
  • —Uses Fogg B=MAP scores to separate motivation, CTA ability, and prompt timing.
  • —Applies prospect-theory plus Cialdini framing: loss recovery, gain momentum, scarcity, social proof, authority, reciprocity, commitment, or liking.
  • —Structures copy with deterministic PAS/BAB/scarcity frameworks so the body has a fact, a cost/opportunity hook, and one concrete reply action.
  • —Converts soft CTAs into commitment CTAs such as Reply YES and I'll prepare... now.
  • —Suppresses weak merchant-only trigger moments when the context has too little why-now evidence, rather than sending a generic fallback.
  • —Uses a pharmacy-safe body contract for merchant-facing pharmacy copy: stock, refill, batch, expiry, checklist, audit, or compliance anchors only.
  • —Builds deterministic Tree-of-Thought frame diagnostics and Best-of-N variants for each trigger.
  • —Runs a deterministic constitutional audit/repair pass for generic copy, multiple CTAs, corporate tone, weak facts, and repeated weak action types.
  • —Uses empirical category action priors from in-memory reply outcomes to reduce cold-start action mistakes without random live exploration.
  • —Scores each candidate across decision quality, specificity, category fit, merchant fit, and engagement compulsion.
  • —Sends only the highest-scoring valid action.
  • —Uses category playbooks for dentists, salons, restaurants, gyms, and pharmacies.
  • —Validates output for hallucinated numbers, None leaks, repeated Dr. Dr., weak generic copy, repeated bodies, missing CTA shape, and unsafe customer outreach.

Every sent message should include why now, a real merchant/category/customer fact, one CTA, category-appropriate voice, and a low-friction next action.

Run Locally

bash
python -m pip install -r requirements-dev.txt
python dataset/generate_dataset.py --seed-dir dataset --out expanded
uvicorn app.main:app --host 0.0.0.0 --port 8080

In another terminal:

bash
pytest -q
python -m compileall app bot.py scripts tests
python scripts/generate_submission.py
python scripts/lint_submission.py
python scripts/score_proxy.py 43
python scripts/geval_calibrate.py

scripts/geval_calibrate.py skips cleanly unless OPENROUTER_API_KEY is set. The official judge_simulator.py also needs a scorer LLM key. The hosted bot itself does not need one.

For LLM judging with a floor-case table, use the wrapper instead of editing the official simulator:

bash
set OPENROUTER_API_KEYS=key1,key2
set BOT_URL=https://mokshak-vera-rubric-decision-engine.hf.space
set JUDGE_SCENARIO=all
python scripts/run_judge_with_trigger_log.py

It prints the lowest-scoring trigger IDs and can write JSON when JUDGE_TRIGGER_LOG_JSON is set.

Optional Local LLM Copy Polish

The submitted HF deployment does not use runtime LLM polish. Set these only for local experiments:

  • —OPENAI_API_KEY
  • —OPENAI_MODEL, default gpt-4o-mini

The model receives only the deterministic plan and evidence. It must return structured JSON with body, cta, send_as, suppression_key, and rationale. The bot falls back to deterministic copy if the call fails, times out, changes protected fields, or invents numbers. Do not enable this in the submitted Space; deterministic runtime is the reliability tradeoff.

Optional OpenRouter Calibration

Set these only for offline quality checks:

  • —OPENROUTER_API_KEY
  • —OPENROUTER_MODEL, default openrouter/auto
  • —GEVAL_LIMIT, default 10

This script critiques generated submission.jsonl rows against the five official dimensions with a Prometheus-style reference bank and reports low-scoring cases. It is intentionally not called by /v1/tick.

Free Deployment

Primary target: Hugging Face Docker Space.

Current hosted base URL:

text
https://mokshak-vera-rubric-decision-engine.hf.space

Start command:

bash
uvicorn app.main:app --host 0.0.0.0 --port $PORT

Health check path:

text
/v1/healthz

Use one process/worker only. State is in memory, so multiple workers would split conversations and suppression keys.

Backup targets:

  • —Koyeb Free Instance
  • —Render Free Web Service using render.yaml

Free hosts can sleep. Before submission, hit /v1/healthz, then keep it warm:

bash
python scripts/keep_warm.py https://your-bot.example --interval 900

Required Submission Details

Set these deployment environment variables:

  • —CONTACT_EMAIL: mokshagnak004@gmail.com
  • —TEAM_NAME: optional display name
  • —TEAM_MEMBER: optional member name
  • —SUBMITTED_AT: optional ISO timestamp

You do not need OpenAI, Groq, Gemini, OpenRouter, Redis, Postgres, or any paid key for the live deterministic bot.

Tradeoffs

In-memory state is simple, private, and fast, but the service must run as a single worker and should not restart during judging. Free LLM APIs are useful for local experiments, but relying on free quota during live judging is risky, so this bot treats LLM usage as optional polish only.

What More Context Would Help

Quality would improve further with verified customer consent mappings, real slot inventory, item-level order or dispense history, offer eligibility rules, customer segment aggregates, and locality-level peer benchmarks. When those facts are absent, the bot avoids inventing them.