# Stress harness

Drives the **real** `CmsClient` over signed HTTP against a real CMS. No stubs —
the seam between agent and CMS is where three defects have now been found, and
none of them were visible to unit tests on either side.

**Run against staging. Never against production.**

Credentials are read from `C:\taxpilot-ai-staging\shared\.env`.

```
python scripts/stress/harness.py register   <n> [tag]   # registration throughput
python scripts/stress/harness.py lifecycle  <n> [tag]   # register + 3 transitions each
python scripts/stress/harness.py duplicates <n>         # same message id n times
python scripts/stress/harness.py audit                  # what the CMS thinks is unfinished

python scripts/stress/concurrent_claims.py              # 4 workers, one outbox
python scripts/stress/network_interruption.py           # CMS unreachable, then back
```

## What these measure, and what they do not

`register` and `lifecycle` **do not run OCR**. They exercise registration, state
transitions, the queue, idempotency, rate limiting and memory. A figure from them
is a workflow-scalability figure and must never be reported as end-to-end.

## Reading the numbers

Throughput is bound by `config('agent.rate_limit')` — 120 calls per minute per
agent key by default. Registration costs one call per document; a full lifecycle
costs four. A run slower than that ceiling means something else is wrong; a run
at the ceiling means the limit is the constraint, not the machine.

## Defects these have found

1. Response envelope mismatch — the client read `["data"]`, intake returned the
   record at the top level. First live call failed.
2. Rate limits treated as refusals — 69 of 100 documents stalled in a burst.
3. Outbox rows read rather than claimed — two workers, one row, one document
   registered twice.
