Team Ai
Apppublic

satwik2217/Startup_Operations

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

πŸš€ Startup Operations: Strategic AI Benchmark

Startup Operations is a sophisticated Reinforcement Learning (RL) environment designed to test an AI agent's ability to navigate high-stakes, long-term strategic trade-offs. Unlike simple linear simulations, this environment models the non-linear realities of tech startups, including Technical Debt, Market Volatility, and Diminishing Returns.


🧠 What Makes This Different

Traditional benchmarks often evaluate "what a model knows." Startup Operations evaluates how a model decides when faced with:

  • β€”Irreversible Consequences: Excessive marketing spend today can lead to a "Quality Death Spiral" that is impossible to recover from in later steps.
  • β€”Hidden Costs (Technical Debt): Choosing growth (Gear 3) has an invisible tax on product stability. The agent must proactively "refactor" via R&D before the system breaks.
  • β€”Environmental Chaos: Success isn't just about following a script; it's about maintaining a "margin of safety" to survive stochastic market shocks.

πŸ’Ž The Three Stones of Complexity

To qualify for research-grade evaluation, this environment implements three advanced "Stones" of logic:

  1. 1.Technical Debt & Quality Decay: Aggressive growth strategies (Gear 3) cause "Product Quality" to decay. If stability drops below 0.7, user churn spikes exponentially.
  2. 2.Stochastic Market Events: The simulation is non-deterministic. A 4% probability trigger introduces "Black Swan" eventsβ€”Funding Crunches (2x burn) or Tax Holidays (0.5x burn)β€”testing agent robustness.
  3. 3.Market Saturation (Expertise Stone): Acquisition follows a logarithmic curve. As the user base grows, the cost-per-acquisition increases, preventing "infinite growth" loops and forcing efficiency.

🎯 Task Progression

Task IDNameDifficultyChallenge
ramen-profitableSurvivalEasySurvive 30 days. Focus on lean operations.
growth-blitzAcquisitionMediumScale to 1,000 users. Requires managing Quality vs. Spend.
series-a-valuationMarket LeaderHard90-day maximization. High exposure to Black Swan events.

πŸ“Š Observation & Action Spaces

Action Space (The 4 Gears)

  • β€”1: Save Cash: Minimal burn, zero growth.
  • β€”2: Balanced: Moderate growth, moderate R&D.
  • β€”3: Aggressive Marketing: High growth, triggers Quality Decay.
  • β€”4: Product Development: Zero growth, triggers Quality Recovery.

Observation Space

The agent perceives a rich state-vector including:

  • β€”Financials: Cash Balance, Burn Rate, Daily Revenue.
  • β€”Growth: Active User Base.
  • β€”Stability: Product Quality (0.0 to 1.5).
  • β€”Sentiment: Market Status (e.g., Stable, Funding Crunch, Viral Success).

πŸ§ͺ Reward Function Logic

The reward is a multi-objective weighted score calculated at each step, clamped between 0.01 and 0.99 for Meta spec compliance.

$$R = (Survival \times 0.2) + (Growth \times 0.4) + (Health \times 0.2) + (Quality \times 0.2)$$

Note: If Product Quality drops below 0.7, the final reward is multiplied by the quality score as a "Technical Debt Penalty," discouraging unsustainable strategies.


βš™οΈ Technical Specifications

ComponentSpecification
RuntimeFastAPI / Uvicorn
ExecutionSynchronous, Turn-based
StatefulnessPersistent Business State (across episode)
StochasticityProbability-based Market Events (4% trigger)
ComplianceOpenEnv v1 Protocol

🐳 Deployment & Usage

Building Locally

bash
docker build -t startup-operations:latest -f server/Dockerfile .