Team Ai
Apppublic

code-shrish/itr-detection

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

ITR Fraud Detection Environment ๐Ÿฆ๐Ÿ”

An OpenEnv-compliant environment where AI agents learn to detect fraud in Indian Income Tax Returns (ITR). Built for the Meta PyTorch ร— Hugging Face OpenEnv Hackathon.

๐ŸŽฏ What It Does

This environment simulates the work of a tax auditor at the Indian Income Tax Department. An AI agent:

  1. 1.Reviews an ITR filing (income, deductions, TDS, high-value transactions)
  2. 2.Investigates specific fields for irregularities
  3. 3.Cross-references data points to find discrepancies
  4. 4.Requests supporting documents (Form 16, rent receipts, bank statements)
  5. 5.Flags anomalies with reasoning and severity
  6. 6.Renders a final verdict: legitimate, suspicious, or fraudulent

This is a real-world task โ€” not a game. Tax auditors perform these exact steps daily to detect billions in tax fraud.


๐Ÿ›๏ธ Project Relevance & The AI Auditor

Why This Idea Matters

India's tax collection system is transitioning to a "Faceless Assessment" model. Manually auditing millions of ITR filings is physically impossible and prone to human error or bias. Our environment provides a training ground for AI Auditors that can:

  • โ€”Scalability: Process thousands of returns per minute.
  • โ€”Consistency: Apply the same audit logic to every citizen fairly.
  • โ€”Complexity: Identify multi-layered "layering" techniques used in money laundering and tax evasion.

Role of Large Language Models (LLMs)

Modern LLMs (like GPT-4o, Gemini 1.5 Pro) are uniquely suited for this task because:

  • โ€”Tax Law Training: These models have read thousands of pages of tax codes, case law, and IT Department circulars (Section 80C, 80D, HRA rules, etc.).
  • โ€”Pattern Recognition: They are excellent at spotting inconsistencies between text-heavy descriptions (standard business books) and numerical data.
  • โ€”Reasoning: Unlike traditional static rule-engines, an AI Agent can decide which document to ask for next based on previous findings, mimicking the curiosity of a human auditor.
  • โ€”Explainability: They don't just flag fraud; they explain why it's fraud in natural language, making them perfect assistants for human review.

๐Ÿš€ Quick Start

Installation

bash
pip install -e .

Run the Baseline Agent (No API Key Needed)

bash
python inference.py --local

Run Inference Script Required by Hackathon

bash
export HF_TOKEN="your-huggingface-or-openai-key-here"
export API_BASE_URL="https://router.huggingface.co/v1" # (Optional: defaults to OpenAI)
export MODEL_NAME="gpt-4o-mini" # (Optional: defaults to gpt-4o-mini)
python inference.py

Start the Server

bash
python -m uvicorn server.app:app --host 0.0.0.0 --port 8000

Use the Client

python
from client import ITRFraudEnv
from models import ITRAction, ActionType, VerdictType

with ITRFraudEnv(base_url="http://localhost:8000") as client:
    # Reset for easy task
    obs = client.reset(task_id="easy")
    
    # Investigate salary
    result = client.step(ITRAction(
        action_type=ActionType.INVESTIGATE_FIELD,
        field_name="income.salary"
    ))
    
    # Render verdict
    result = client.step(ITRAction(
        action_type=ActionType.RENDER_VERDICT,
        verdict=VerdictType.FRAUDULENT,
        confidence=0.9,
        explanation="Income mismatch with TDS records"
    ))

๐Ÿ“‹ Tasks

TaskDifficultyFraud PatternsMax Steps
Obvious Red FlagsEasyIncome mismatch, 80C over-limit10
The Subtle CheatMediumPhantom HRA, penny stock LTCG, tax computation error15
The Elaborate SchemeHardShell companies, income shifting, cash deposits, fake donations25

Each task has a programmatic grader that scores performance from 0.0 to 1.0 based on:

  • โ€”Anomaly detection accuracy
  • โ€”Verdict correctness
  • โ€”Investigation completeness
  • โ€”Efficiency (fewer steps = better)

๐ŸŽฎ Action Space

ActionDescriptionKey Parameters
investigate_fieldDrill into a specific fieldfield_name
cross_referenceCompare two data pointsfield_a, field_b
request_documentRequest supporting docsdocument_type
flag_anomalyFlag suspicious findinganomaly_field, anomaly_reason, anomaly_severity
render_verdictFinal decisionverdict, confidence, explanation

Available Document Types

  • โ€”form_16 โ€” Employer TDS certificate
  • โ€”bank_statement โ€” Bank account records
  • โ€”rent_receipts โ€” For HRA verification
  • โ€”investment_proofs โ€” 80C/80D proofs
  • โ€”capital_gains_statement โ€” Stock trading records
  • โ€”business_books โ€” Business accounting
  • โ€”related_party_records โ€” Related entity transactions

๐Ÿ“Š Observation Space

Each observation contains:

  • โ€”ITR Summary: Full tax return data (income, deductions, TDS, transactions, previous years)
  • โ€”Last Action Result: What happened from the last action
  • โ€”Investigation Results: Findings from investigations
  • โ€”Document Results: Requested documents and any discrepancies
  • โ€”Flagged Anomalies: List of anomalies the agent has flagged
  • โ€”Step Info: Current step number, max steps, task description

๐Ÿ† Reward Function

ComponentValueDescription
Correct anomaly flag+0.10Flagging a real irregularity
Useful investigation+0.05Investigating a relevant field
Useful cross-reference+0.08Finding a real discrepancy
Useful document+0.06Document reveals discrepancies
Correct verdict+0.30Right final call
Efficiency bonusup to +0.05Fewer steps = more reward
False flag-0.05Flagging a non-issue
Redundant action-0.02Repeating an investigation
Timeout-0.10Running out of steps

๐Ÿณ Docker Deployment

bash
# Build
docker build -t itr-fraud-env -f server/Dockerfile .

# Run
docker run -p 8000:8000 itr-fraud-env

# Test
curl http://localhost:8000/health

๐Ÿ“ Project Structure

โ”œโ”€โ”€ models.py              # Pydantic Action/Observation/State models
โ”œโ”€โ”€ client.py              # HTTP client (ITRFraudEnv)
โ”œโ”€โ”€ inference.py           # Baseline inference script
โ”œโ”€โ”€ openenv.yaml           # OpenEnv manifest
โ”œโ”€โ”€ pyproject.toml         # Package configuration
โ”œโ”€โ”€ tasks/
โ”‚   โ”œโ”€โ”€ task_easy.py       # Easy: Obvious red flags
โ”‚   โ”œโ”€โ”€ task_medium.py     # Medium: Subtle inconsistencies
โ”‚   โ””โ”€โ”€ task_hard.py       # Hard: Multi-layered fraud
โ”œโ”€โ”€ data/
โ”‚   โ””โ”€โ”€ itr_generator.py   # Synthetic ITR data generator
โ””โ”€โ”€ server/
    โ”œโ”€โ”€ itr_environment.py # Core environment logic
    โ”œโ”€โ”€ app.py             # FastAPI server
    โ”œโ”€โ”€ requirements.txt   # Server dependencies
    โ””โ”€โ”€ Dockerfile         # Container image

๐Ÿ”ง OpenEnv API

EndpointMethodDescription
/resetPOSTStart new episode ({"task_id": "easy"})
/stepPOSTExecute action
/stateGETGet episode state
/healthGETHealth check
/infoGETEnvironment metadata

๐Ÿ“œ License

MIT