satwik2217/Startup_Operations
π Startup Operations: Strategic AI Benchmark
Startup Operations is a sophisticated Reinforcement Learning (RL) environment designed to test an AI agent's ability to navigate high-stakes, long-term strategic trade-offs. Unlike simple linear simulations, this environment models the non-linear realities of tech startups, including Technical Debt, Market Volatility, and Diminishing Returns.
π§ What Makes This Different
Traditional benchmarks often evaluate "what a model knows." Startup Operations evaluates how a model decides when faced with:
- Irreversible Consequences: Excessive marketing spend today can lead to a "Quality Death Spiral" that is impossible to recover from in later steps.
- Hidden Costs (Technical Debt): Choosing growth (Gear 3) has an invisible tax on product stability. The agent must proactively "refactor" via R&D before the system breaks.
- Environmental Chaos: Success isn't just about following a script; it's about maintaining a "margin of safety" to survive stochastic market shocks.
π The Three Stones of Complexity
To qualify for research-grade evaluation, this environment implements three advanced "Stones" of logic:
- Technical Debt & Quality Decay: Aggressive growth strategies (Gear 3) cause "Product Quality" to decay. If stability drops below 0.7, user churn spikes exponentially.
- Stochastic Market Events: The simulation is non-deterministic. A 4% probability trigger introduces "Black Swan" eventsβFunding Crunches (2x burn) or Tax Holidays (0.5x burn)βtesting agent robustness.
- Market Saturation (Expertise Stone): Acquisition follows a logarithmic curve. As the user base grows, the cost-per-acquisition increases, preventing "infinite growth" loops and forcing efficiency.
π― Task Progression
π Observation & Action Spaces
Action Space (The 4 Gears)
- 1: Save Cash: Minimal burn, zero growth.
- 2: Balanced: Moderate growth, moderate R&D.
- 3: Aggressive Marketing: High growth, triggers Quality Decay.
- 4: Product Development: Zero growth, triggers Quality Recovery.
Observation Space
The agent perceives a rich state-vector including:
- Financials: Cash Balance, Burn Rate, Daily Revenue.
- Growth: Active User Base.
- Stability: Product Quality (0.0 to 1.5).
- Sentiment: Market Status (e.g., Stable, Funding Crunch, Viral Success).
π§ͺ Reward Function Logic
The reward is a multi-objective weighted score calculated at each step, clamped between 0.01 and 0.99 for Meta spec compliance.
$$R = (Survival \times 0.2) + (Growth \times 0.4) + (Health \times 0.2) + (Quality \times 0.2)$$
Note: If Product Quality drops below 0.7, the final reward is multiplied by the quality score as a "Technical Debt Penalty," discouraging unsustainable strategies.
βοΈ Technical Specifications
π³ Deployment & Usage
Building Locally
docker build -t startup-operations:latest -f server/Dockerfile .