Spaces:
Sleeping
title: Financial Audit Env
emoji: π°
colorFrom: blue
colorTo: green
sdk: docker
app_port: 8000
pinned: false
tags:
- openenv
π° Financial Audit Env
A Multi-Agent Oversight Platform for Training Auditing LLMs
An OpenEnv-compatible reinforcement learning environment for training AI agents to audit financial documents through multi-agent cooperation, regulatory adaptation, and self-improvement. Built for the Meta PyTorch Hackathon β Round 2.
Theme #3 World Modeling β #3.1 Professional Tasks | Built with OpenEnv v0.2.3+ | Deployed on HF Spaces | Training via HF TRL GRPO + Unsloth in Colab
Live API Β· 108 tests passing Β· Blog: BLOG.md Β· Notebook: GRPO_Training_Submission_Final.ipynb
Environment Walkthrough: From Baseline to Regulatory Shock
Act 1 β The Baseline. Period 1. The agent receives 19 expense claims and a policy doc. Easy task, stable rules, F1 β 0.12.
Act 2 β The World Mutates. Period 2. Meal limit jumps βΉ1,500 β βΉ2,000. A new vendor onboards. Cross-period memory is now required. The agent's internal model is stale unless it updates.
Act 3 β The Regulatory Shock. Period 3. Mid-audit, REG-001 drops: GST on IT services 18% β 12%. Schema drifts (vendor_gstin β supplier_gstin). Findings already submitted under the old rule. Baselines drop 15β20% here, exactly as designed.
Act 4 β The Environment Reveals the Training Problem. GRPO chases the densest reward signal. Llama 3.1 8B improved 3Γ on expense audits and collapsed 82% on fraud detection. The 0.20 F1 floor multiplier kicked in. The environment said no, this isn't a win β and made the curriculum bias obvious.
A good environment doesn't hide training problems. It surfaces them so cleanly that the fix is obvious.
Problem Statement
Most LLM tools for finance treat auditing as retrieval: "find the thing that violates the rule." This environment tests something harder: Can an LLM maintain a consistent internal model of a financial world and UPDATE that model when ground truth shifts?
Real audits run across multiple fiscal periods. A GST rate changes mid-audit. A vendor gets flagged mid-investigation. An approval threshold drops after Period 1 findings were already submitted. That's the capability gap professional LLM applications hit, and it's what this environment measures.
System Specifications
| Component | Detail |
|---|---|
| Model | Llama 3.1 8B Instruct (4-bit quantized via Unsloth) |
| LoRA Rank | 16 (~21M trainable params) |
| Training | GRPO via HF TRL β 10 train seeds Γ 4 tasks = 40 prompts |
| Environment | 4 specialist agents + 1 overseer, 5-period campaigns, 25 instructions |
| Seed Pools | Train 42β51 / Held-out 100β104 (strict disjoint sets) |
| Score Range | (0.01, 0.99) β strictly bounded for clean GRPO signal |
| Deployment | Docker on HF Spaces, FastAPI on port 8000 |
Submission Index
- GitHub: https://github.com/balloonmann/financial-audit-env
- HF Space (live env): https://huggingface.co/spaces/balloonmann/financial_audit_env
- Training notebook: GRPO_Training_Submission_Final.ipynb
- Blog: BLOG.md
- GRPO adapter: https://huggingface.co/balloonmann/financial-audit-grpo-adapter
- Eval artifacts: https://huggingface.co/datasets/balloonmann/financial-audit-eval-artifacts
Checklist β β
OpenEnv-compatible HF Space Β· β
valid openenv.yaml Β· β
Unsloth + TRL GRPO notebook Β· β
blog with results Β· β
adapter + eval CSVs on HF Hub Β· β
before/after plots embedded below.
Architecture
What's New in Round 2
| Feature | Description |
|---|---|
| Multi-Agent Campaign | 4 specialists + 1 overseer per period, dependency order enforced |
| 5-Period Campaigns | World mutates each period (policy, schema, vendor status) |
| Regulatory Shocks | 3 mid-period rule drops β agent must adapt without restart |
| Schema Drift | Field renames in P3+ (vendor_gstin β supplier_gstin) |
| Overseer Review | Approves/rejects specialist findings, resolves conflicts |
| Self-Improvement | Critic + regression gate + held-out seed separation |
| Adversarial Red/Blue | Tunable fraud difficulty (5 levels), arms-race tracking |
| 22 Frozen Instructions | Binary-checkable rules across 5 buckets |
| GRPO Training | Colab-ready Unsloth + TRL pipeline |
| Confidence Calibration | ECE scoring on confidence-tagged findings |
Campaign Flow
Period N:
1. Expense β 2. Invoice β 3. GST β 4. Fraud (dependency-ordered)
5. Overseer reviews all findings
6. Advance period β world mutates
β οΈ REGULATORY SHOCKS may drop mid-period β applied to remaining work.
| Period | Changes |
|---|---|
| 1 | Baseline β 17 active instructions |
| 2 | Meal limit βΉ1,500ββΉ2,000, new vendor, cross-period memory required |
| 3 | GST 18%β12% (IT services), schema drift, REG-001 shock |
| 4 | Vendor under investigation, REG-002 + REG-003 (cash/UPI threshold, vendor risk) |
| 5 | Annual reconciliation β all 22 instructions active |
Tasks
| # | Task | Difficulty | Documents | Errors | What the Agent Must Do |
|---|---|---|---|---|---|
| 1 | Expense Policy Audit | Easy | 19 claims + policy | 7 violations | Check claims against policy |
| 2 | Invoice Three-Way Match | Medium | 10 POs + 10 GRNs + 12 invoices | 9 discrepancies | Cross-reference 3 doc types |
| 3 | GST Reconciliation | Hard | 45 book entries + 44 GSTR-2B | 12 mismatches | Reconcile books vs gov data |
| 4 | Fraud Detection | Expert | 84 transactions + 26 vendors | 10 fraud patterns | Forensic pattern recognition |
Quick Start
import requests
BASE = "http://localhost:8000"
# Round 1 β single task
ep = requests.post(f"{BASE}/reset", json={"task_id": 1, "seed": 42}).json()
findings = [{"document_id": "EXP-001", "error_type": "weekend_expense", "confidence": 0.9}]
r = requests.post(f"{BASE}/step", json={"findings": findings, "submit_final": True}).json()
print(r["grader"]["partial_credit_f1"], r["reward"])
# Round 2 β full 5-period campaign
cid = requests.post(f"{BASE}/campaign/start",
json={"seed": 42, "total_periods": 5}).json()["campaign_id"]
for p in range(1, 6):
for role in ["expense", "invoice", "gst", "fraud"]:
requests.post(f"{BASE}/campaign/task/start", json={"campaign_id": cid, "role": role})
requests.post(f"{BASE}/campaign/task/submit",
json={"campaign_id": cid, "role": role, "findings": []})
requests.post(f"{BASE}/overseer/review", json={"campaign_id": cid})
requests.post(f"{BASE}/campaign/period/advance", json={"campaign_id": cid})
print(requests.get(f"{BASE}/campaign/state/{cid}").json()["score"])
API Reference
Standard: /, /health, /docs, /tasks, /reset, /step, /state, /grader, /session, /leaderboard, /metrics, /adaptive-difficulty, /baseline (13)
Campaign: /campaign/start, /campaign/state/{id}, /campaign/task/start, /campaign/task/submit, /campaign/period/advance (8 total)
Oversight + Self-Improvement: /overseer/review, /overseer/report, /self-improve, /self-improve/history (4)
29 routes total. Full request/response schemas at /docs.
Grading System
Per-task metrics β Partial Credit F1 (primary, 0.01β0.99, 0.40 credit for right doc / wrong type), Strict F1, Weighted F1 (fraud=2.0Γ, weekend=0.5Γ), confusion matrix, risk score (βΉ), ECE.
Campaign score β 35% specialist F1 + 25% overseer + 10% compliance + 10% memory + 8% schema/policy + 7% improvement + 5% efficiency.
Anti-gaming guards β any specialist F1 < 0.20 β score = 0.01; any critical (severity β₯ 1.5) miss β score Γ 0.5; bonuses capped at 30%.
Reward signal β +0.15/TP (severity-weighted), +0.04/partial, β0.05/FP, β0.02 step + β0.005Β·step_n decay, +0.30 final if recall β₯ 0.6, β0.20 if recall < 0.3.
Training
GRPO via Unsloth + TRL on Colab T4 (or HF Jobs A10G):
# Colab
!pip install -r requirements-training.txt
!python training/train_grpo.py
# HF Jobs A10G
hf jobs run --flavor a10g-large --timeout 6h --secrets HF_TOKEN \
pytorch/pytorch:2.6.0-cuda12.4-cudnn9-devel -- bash -lc \
"git clone https://github.com/balloonmann/financial-audit-env && \
cd financial-audit-env && bash scripts/hf_jobs_bootstrap_and_train.sh"
- InProcessEvaluator (
training/evaluator.py) β direct Python eval, no HTTP overhead - Reward fn (
training/reward.py) β JSON + free-text parsing β F1 score - Config β Llama 3.1 8B 4-bit, LoRA r=16, 10 seeds Γ 4 tasks
- Seeds β train 42β51, held-out 100β104, disjoint sets enforced
Training Results β Held-Out Seeds 100β104
Llama 3.1 8B (HF Jobs A10-Large)
| Mean Score | |
|---|---|
| Baseline | 0.1690 |
| GRPO Trained | 0.1230 |
| Delta | β0.0460 (β27.2%) |
| Task | Difficulty | Baseline F1 | Trained F1 | Delta |
|---|---|---|---|---|
| expense_audit | Easy | ~0.12 | 0.356 | +196% |
| invoice_match | Medium | ~0.18 | 0.074 | β59% |
| gst_reconciliation | Hard | ~0.01 | 0.042 | +320% |
| fraud_detection | Expert | ~0.11 | 0.020 | β82% |
Expense improved dramatically; fraud collapsed β the optimizer chased the densest signal. Full analysis in BLOG.md.
Held-out seeds 100β104. Left: F1 per task. Right: Recall per task. Orange = GRPO trained, Blue = baseline.
Qwen 2.5-1.5B (Colab T4)
Mean: baseline 0.0470 β trained 0.0100 (β78.7%). All four tasks collapse to format floor β 1.5B at 4-bit lacks capacity to absorb the GRPO policy changes while keeping valid JSON output. Training reward was flat at 0.01 for all 120 steps (reward_std = 0 throughout), confirming the GRPO cold-start failure mode.
Qwen 2.5-1.5B: reward pinned at floor (0.01) for all 120 steps. reward_std = 0 means GRPO had zero gradient signal throughout.
Observed Agent Behaviors Under GRPO Training
- Submit findings as structured JSON β free-text gets parsed but loses precision.
- Prefer high-confidence claims on easy tasks; partial-credit weighting punishes hallucinated FPs.
- Weekend dates and missing receipts are dense, low-risk signals β chase them first.
- On the final step, abstaining beats guessing when recall < 0.3.
- Cross-period findings should be referenced, not re-derived β repeated evidence is rewarded.
Training Findings and Design Implications
- GRPO with a single scalar reward collapses onto whichever task has the densest signal.
- A 0.20 specialist F1 floor isn't enough; per-task floors or task-stratified reward batches are next.
- The 4-bit Qwen 1.5B run hit format floor β capacity matters more than data under aggressive quantization.
- The 0.40 partial-credit weight is generous on easy tasks and stingy on fraud β task-specific weights are next.
- Static evaluation seeds hide failure modes; held-out 100β104 caught the curriculum bias the training reward never showed.
Inference & Evaluation
# Round 1 (single task)
export HF_TOKEN=... API_BASE_URL=https://huggingface.co/proxy/router.huggingface.co/v1/ \
MODEL_NAME=meta-llama/Llama-3.1-8B-Instruct
python inference.py --env-url http://localhost:8000
# Round 2 (5-period campaign)
python inference.py --env-url http://localhost:8000 --campaign --seed 42
# Reproduce held-out evaluation (see GRPO_Training_Submission_Final.ipynb for full eval loop)
for seed in 100 101 102 103 104; do
python financial_audit_env/baseline.py --base-url https://huggingface.co/proxy/balloonmann-financial-audit-env.hf.space --seed $seed
done
Baseline (Round 1, seed 42, Llama 3.1 8B) β Expense F1 0.12 / Invoice 0.18 / GST 0.01 / Fraud 0.11. The model struggles with abstract rules (date math, cumulative limits) and chases red-herring expenses, destroying precision.
Setup and Local Development
git clone https://github.com/balloonmann/financial-audit-env.git
cd financial-audit-env && pip install -e .
python -m financial_audit_env.server.app # serves on :8000
python -m pytest tests -q # 108 tests pass in ~10s
Docker: docker build -t financial-audit-env . && docker run -p 8000:8000 financial-audit-env
Judge runbook: pip install -e . β start server β pytest tests -q β python inference.py --campaign --seed 42 β curl https://huggingface.co/proxy/balloonmann-financial-audit-env.hf.space/health.
Deployment on HF Spaces
FROM python:3.11-slim
WORKDIR /app
COPY . /app
RUN pip install --no-cache-dir -e .
EXPOSE 8000
CMD ["uvicorn", "financial_audit_env.server.app:app", "--host", "0.0.0.0", "--port", "8000"]
# openenv.yaml
spec_version: 1
name: financial_audit_env
type: space
runtime: fastapi
app: financial_audit_env.server.app:app
port: 8000
tasks:
- { id: 1, name: expense_audit }
- { id: 2, name: invoice_match }
- { id: 3, name: gst_reconciliation }
- { id: 4, name: fraud_detection }
The YAML frontmatter at the top of this README controls Space metadata (title, emoji, color, port, tags). Push to main rebuilds the Space.
Configuration
| Variable | Description | Default |
|---|---|---|
HF_TOKEN |
HF token for inference router + adapter pulls | required for inference |
API_BASE_URL |
OpenAI-compatible inference endpoint | https://huggingface.co/proxy/router.huggingface.co/v1/ |
MODEL_NAME |
Inference model id | meta-llama/Llama-3.1-8B-Instruct |
ADMIN_API_KEY |
Optional API key for protected admin endpoints | unset |
RATE_LIMIT_PER_MIN |
Per-IP limit on /step and /campaign/* |
120 |
MAX_FINDINGS_PER_STEP |
Hard cap on findings per step (anti-spam) | 200 |
CAMPAIGN_TOTAL_PERIODS |
Periods per campaign | 5 |
FRAUD_DIFFICULTY |
Adversarial fraud level override (0β4) | auto |
LOG_LEVEL |
Server log level | INFO |
Project Structure
financial_audit_env/
server/
app.py # FastAPI β all endpoints
environment.py # Core RL environment
data_generator.py # Synthetic data + planted errors + schema drift
graders.py # F1 family, ECE, campaign score, cross-agent agreement
campaign.py # 5-period orchestration
instructions.py # 22 frozen + 3 regulatory shocks
regulatory.py # Mid-period shock engine
adversarial.py # Red/Blue fraud difficulty
self_improve.py # Self-improvement + regression gate
security.py # Rate limiting, headers
tasks.py # 4 task definitions
models.py # Pydantic models
baseline.py # Llama 3.1 baseline agent
training/
evaluator.py # InProcessEvaluator
reward.py # GRPO reward parser
train_grpo.py # Colab training script
tests/ # 108 pytest tests
inference.py # R1 + R2 inference CLI
openenv.yaml # OpenEnv spec
Dockerfile # HF Spaces image
Design Decisions
Multi-agent. Real auditing is team-based; specialists + overseer give natural task dependencies and conflict resolution.
5 periods. Long-horizon tests memory, adaptation, planning. P3 introduces schema drift, P4 brings shocks β the agent can't memorize P1.
Regulatory shocks. Mid-episode rule changes are realistic (tax law shifts mid-quarter) and test belief updating without restart.
Strict seed separation. Train 42β51 vs held-out 100β104 prevents overfit; the self-improvement engine rejects overlapping seeds at the API level.
Deterministic scoring. Same findings β same score, always. No LLM judge in the reward loop. GRPO needs this variance separation to compute meaningful advantages.
Research Significance
This environment tests what cutting-edge agentic-LLM research is converging on:
Long-horizon belief updating with structured tool output under shifting ground truth.
It connects to: world-modeling & belief revision (Theme #3.1, the agent must update internal state when REG-001/002/003 fire); curriculum learning under reward-density skew (the Llama run is a clean reproduction of "RL collapses onto easy tasks" with quantitative receipts: +196% easy, β82% expert); multi-agent oversight (the overseer pattern mirrors LLM-as-reviewer / debate / recursive critique pipelines); and adversarial Red/Blue self-play (5-level fraud controller is a small-scale adversarial-designer).
Keeping scoring deterministic and bounded in (0.01, 0.99) is what makes GRPO usable here β putting an LLM judge in the reward loop destroys the variance separation GRPO needs.
Hackathon Alignment
| Criterion | How We Address It |
|---|---|
| Environment Innovation (40%) | Multi-agent oversight + regulatory shocks + schema drift + self-improvement β well beyond static eval |
| Storytelling (30%) | Specialists β overseer β advance β adapt; real-world domain; the GRPO run itself is a story about curriculum bias |
| Reward Improvement (20%) | GRPO script + before/after F1 tables + Llama comparison plot + Qwen training curves |
| Reward & Training Pipeline (10%) | InProcessEvaluator + reward parser + Colab-ready training script |
Round 2 Implementation Scorecard
Verified 2026-04-24 with
pytest tests -qβ all 108 tests passing.
| Step | Component | Status |
|---|---|---|
| 1 | Core models (AgentRole, WorldState, CampaignState, OverseerAction, CriticReport, CampaignObservation) |
β |
| 2 | Instructions registry (22 frozen + 3 shocks across 5 buckets) | β |
| 3 | Campaign Controller (5-period orchestration, composition over inheritance) | β |
| 4 | Extended grading (ECE, campaign score, cross-agent agreement) | β |
| 5 | Self-Improvement API (critic, regression gate, 12 R2 endpoints) | β |
| 6 | Regulatory Shock Engine (mid-period rule injection + GT modification) | β |
| 7 | Adversarial Red/Blue (5-level difficulty, arms race) | β |
| 8 | Training infra (InProcessEvaluator, reward parser, GRPO script) | β |
| 9 | Campaign inference (multi-period flow + CLI) | β |
| 10 | Tests + README (108 tests passing) | β |
10/10 implementation steps complete.
Anti-Gaming Guards (Verified)
| Guard | Trigger | Effect | Verified |
|---|---|---|---|
| Specialist Floor | any specialist weighted F1 < 0.20 | score β 0.01 | β Fires |
| Safety Gate | critical miss (severity β₯ 1.5) | score Γ 0.50 | β 0.62 β 0.31 |
| Bonus Cap | non-core bonuses > 30% raw total | clamped | β Enforced |
Test Coverage
test_campaign_round2.py (10) Β· test_data_generators.py (22) Β· test_environment.py (20) Β· test_graders.py (15) Β· test_regulatory.py (7) Β· test_security.py (5) Β· test_self_improve.py (6) Β· test_adversarial.py (4) β 108 passing in ~10s.
Key Quantities
22 frozen instructions + 3 regulatory shocks (5 buckets) Β· 4 specialists + 1 overseer Β· 5 campaign periods Β· 5 fraud difficulty levels Β· 10 training seeds + 5 held-out Β· score β (0.01, 0.99) Β· 7 grading functions Β· 29 API routes.