# The command queue

How the CMS asks its agent to do something, without ever calling it.

---

## The problem

A customer presses **Connect WhatsApp** in their own dashboard. The thing that
must act is the agent — which runs on other infrastructure and, by ADR-0005, is
the *client* in this relationship. The CMS has no outbound client, no credential
for the agent, and no idea where it lives.

The obvious answer is to give it all three. That would mean a public inbound
endpoint on every agent, a stored credential in every CMS, and two directions of
traffic to secure instead of one.

## The answer

The CMS writes down what it wants. The agent — already polling this installation
for proposal decisions — collects it on its next pass, does the work, and reports
back through the same signed channel.

```mermaid
sequenceDiagram
    participant C as Customer
    participant CMS as TaxPilot CMS
    participant AI as TaxPilot AI
    participant P as Evolution

    C->>CMS: Connect WhatsApp
    CMS->>CMS: queue whatsapp.connect (pending)
    Note over CMS,AI: the CMS never calls out
    AI->>CMS: POST commands/claim
    CMS-->>AI: whatsapp.connect (now in_progress)
    AI->>P: create session
    P-->>AI: QR code
    AI->>CMS: POST commands/{id}/result
    CMS-->>C: display the QR
    C->>P: scan with the phone
```

Nothing new is exposed on either side.

---

## Lifecycle

```
pending ──claim──> in_progress ──> completed
   │                    │      └──> failed        attempts exhausted
   │                    └─release─> pending       retry, with backoff
   ├──sweep──> expired                            nobody collected it
   └──cancel──> cancelled                         a person changed their mind
```

Every transition is written to `ai_command_events`, append-only. The status
column says where a command *is*; the trail says how it got there — and only the
second answers "why did connecting fail three times on Tuesday", which is the
question actually asked when a customer complains.

---

## Two kinds of duplicate, two different guards

**The same request arriving twice** — a browser retry, a double click. Guarded by
`idempotency_key`, unique: the second call returns the command already created.

**Two different requests for the same thing at once.** Guarded by a unique index
on `active_key`, which holds `"{agent}:{type}"` while a command is live and
becomes NULL the moment it reaches a terminal state. MySQL permits any number of
NULLs in a unique index, so finished commands stop competing while live ones
cannot duplicate.

The second **cannot** be an application check. Two requests a millisecond apart
both read "nothing active" and both insert. The database decides, or nothing
does.

It matters more than it sounds: two QR codes in flight means scanning the second
silently invalidates the first, and the customer watches their screen stop
working for no reason they can see.

---

## Surviving a restart

**Nothing is lost, because nothing is local.** Every command's state lives in the
CMS. A process killed mid-flight leaves a claim behind rather than losing work.

Two sweeps run before any claim is handed out — deliberately on the poll rather
than on a schedule, so a restarted agent recovers its own interrupted work on its
very next pass with no cron involved:

| | |
| --- | --- |
| **Expired** | past `expires_at` and never collected → `expired`. An expired command must never be handed out and then swept a second later |
| **Stale claim** | `in_progress` for longer than 5 minutes → back to `pending`. This is how a command survives the agent being killed |

The consequence for handlers: **they must be safe to run twice.** A command whose
result never reached the CMS will come back.

---

## Retries

The agent reports *what happened* and *whether trying again could plausibly
work*. Everything else is the CMS's decision, where it is durable and auditable.

`retryable` is the agent's call because only it knows the difference:

- *Evolution was unreachable* → worth another attempt
- *This provider cannot be linked by QR* → never will be, and three attempts
  only delay telling the customer something true

Backoff is exponential — 30s, 60s, 120s. A provider that is down stays down for
a while, and retrying every poll turns one outage into a stream of identical
failures nobody learns anything from.

---

## Expiry

| Command | Window | Why |
| --- | --- | --- |
| `whatsapp.connect` | 2 minutes | It answers somebody stood at a screen. Two minutes later they have gone |
| `whatsapp.status` | 5 minutes | Background work |
| `whatsapp.disconnect` | 5 minutes | Background work |

---

## The endpoints

Both on the CMS, both called by the agent, both requiring the `commands.process`
grant — which, per ADR-0003, a newly issued agent does not hold.

```
POST  /api/agent/v1/commands/claim           → 200 with a command, or 204
POST  /api/agent/v1/commands/{uuid}/result   → completed | failed
```

Claiming and reporting are **separate calls**, because the work between them
happens on another machine and can take seconds. A single request that claimed,
waited and returned would hold a connection open across a QR generation and lose
the command entirely if it dropped.

The identifier is a UUID, never the row id — that leaks how many commands an
installation has ever issued.

An agent can only see and report on its own work; another agent's command is a
404, not a 403, so the existence of other deployments is not confirmed either.

---

## The QR code in transit

The customer's browser has to draw it, and the browser talks to the CMS, so the
code travels: agent → CMS result → screen.

It is a credential the whole way. Scanning one links a device and grants read
access to every conversation in the account. So:

- it lives 90 seconds by construction
- `LinkRequest` masks it in `repr()` and `str()` — a dataclass repr is how
  secrets reach a log file
- it appears in exactly one place, the command result, and nowhere else
- the CMS must clear it once the link completes or the window passes

---

## What this does not change

- **ADR-0005** — the Agent API is still the only channel, and the agent still
  initiates every conversation
- **ADR-0001** — one deployment, one installation, one WhatsApp account. A
  handler that selected a provider from a command payload would be the first
  step toward multi-tenancy, so handlers bind their provider at registration
- **ADR-0008** — commands carry operational work, never client data. Nothing
  here files anything
