Team Ai
Datasetpublic

Via20/blind-spots-frontier-optimization

Technical Challenge: Symbolic Drift & Constraint Decay in Frontier LLMs 1. Executive Summary & Background When deploying AI for core engineering, economic, and mathematical optimization tasks (such as Karush-Kuhn-Tucker conditions or Lagrangian multipliers), calculations require rigorous multi-step invariant tracking. Drawing from practical experience building optimization pipelines, we identify an underexplored capability gap in current models: Symbolic Drift and… See the full description on the dataset page: https://huggingface.co/datasets/Via20/blind-spots-frontier-optimization.

sourceHugging Faceupdated 1d agoView on Hugging Face
0likes36downloads
Dataset Card

Technical Challenge: Symbolic Drift & Constraint Decay in Frontier LLMs

1. Executive Summary & Background

When deploying AI for core engineering, economic, and mathematical optimization tasks (such as Karush-Kuhn-Tucker conditions or Lagrangian multipliers), calculations require rigorous multi-step invariant tracking. Drawing from practical experience building optimization pipelines, we identify an underexplored capability gap in current models: Symbolic Drift and Constraint Decay.

Unlike human analysts who maintain external working memory or strict symbolic invariants, autoregressive language models treat derivations as next-token distributions. As derivation length increases:

  • —Coefficients mutate silently across intermediate paragraphs.
  • —Constraints vanish or fail to enforce complementary slackness (\mu \ge 0).
  • —Algebraic shortcuts mask missing intermediate arithmetic steps, leading to high residual constraint violations.

2. Systematic Evaluation Setup

  • —Model Evaluated: Qwen/Qwen2.5-3B-Instruct (~3.09 Billion parameters, hosted on Hugging Face).
  • —Environment: Google Colab (Free T4 GPU runtime, FP16 precision via Hugging Face Transformers & PyTorch).
  • —Evaluation Suite: Custom dataset of multi-step constrained quadratic programs and portfolio risk minimization problems requiring exact rational fraction arithmetic and boundary verification.

3. Empirical Observations & Failure Modes

Problem CategoryPrompt FocusObserved Model BehaviorEvaluation Finding
Constrained Quadratic ProgramKKT conditions, stationary equations, primal equality (2x - y = 1)Generated clean step-by-step algebraic derivations and correct final coordinates (x^, y^) = (1.2, 1.4).Arithmetic Hallucination / Drift: Detailed case-study analysis below reveals subtle arithmetic failure despite clean formatting.
Portfolio Variance Minimization5x5 linear system solving via Lagrange multipliersSkipped explicit matrix row-reduction steps and jumped straight to final weight distribution.Algebraic Shortcut Fallacy: Final weights violated the target expected return constraint by over 14% due to skipped simultaneous equation steps.
Non-linear Boundary OptimizationActive vs. inactive inequality case checkingDropped secondary inequality constraint g_2(x) \le 1 halfway through the derivation.Constraint Decay: Evaluated an interior saddle point rather than verifying true global boundary feasibility.

4. Detailed Case Study: Concrete Failure Trace Analysis (cqp_01)

To understand how the model fails under the hood, we audited the generated text for test case cqp_01:

  1. 1.The Model's Optimal Point: The model successfully derived the coordinate candidate (x^, y^) = (6/5, 7/5) = (1.2, 1.4).
  2. 2.Primal Constraint Verification:
  3. 3.Testing equality (2x - y = 1): 2(1.2) - 1.4 = 2.4 - 1.4 = 1.0 (Satisfied).
  4. 4.Testing inequality (x + 2y \ge 4): 1.2 + 2(1.4) = 1.2 + 2.8 = 4.0 (Meets boundary equality).
  5. 5.The Smoking Gun (Arithmetic Hallucination & Symbolic Strain): When evaluating the objective function f(x, y) = 3x^2 + 2y^2 - 4x + 6y + 7 at this point using rational fractions, the model generated: [ = 3(36/25)^2 + 2(49/25)^2 - 24/5 + 42/5 + 7 ] [ = 108/25 + 98/25 - 120/25 + 210/25 + 175/25 ] [ = 491/25 = 19.64 ]

The Audit: Summing the numerators correctly yields 108 + 98 - 120 + 210 + 175 = 471, not 491. The model suffered from a hidden arithmetic hallucination at the final summation stage while maintaining flawless linguistic and markdown formatting. This proves that models can mimic proof structures perfectly while failing multi-step arithmetic invariants.


5. Proposed Path Forward

To bridge this gap without scaling model size, we propose:

  1. 1.Neuro-Symbolic Sandboxing (RLVR): Fine-tuning models via Reinforcement Learning with Verifiable Rewards to output executable symbolic code blocks (e.g., SymPy) for intermediate arithmetic and constraint verification.
  2. 2.State-Tracking Attention Registers: Architectural buffers dedicated to pinning variable definitions and active inequalities across deep context windows to prevent symbolic decay.

6. Repository Artifacts

  • —run_eval.py: Complete Google Colab execution script.
  • —evaluation_results.json: Raw model outputs, token counts, and execution logs across all categories.