# Clinic CMS — Enterprise Audit

**Date:** 4 August 2026
**Scope:** full codebase, database, security posture, deployment, and commercial readiness
**Method:** direct inspection, `composer audit` / `npm audit`, the 803-test suite run today (2,486 assertions, green), and live-install verification against clinic.tmsoagency.com. Findings are marked **[verified]** when demonstrated by a test, a command, or a reproduced behaviour today, and **[observed]** when based on code reading alone.

---

## 1. Executive Summary

This is an unusually well-engineered codebase for its stage. The module architecture is real and *enforced* — an architecture test fails the build when core imports a module or a module reaches past another's published contracts — and the domain modelling consistently makes the safe thing the default: append-only ledgers, snapshot-on-issue records, database-level guards (unsigned quantities, unique idempotency keys) backing every application-level check. Static analysis (PHPStan, no baseline suppressions), Pint, and 803 passing tests gate every change. Zero known dependency vulnerabilities in either composer or npm trees **[verified]**.

The risks are almost entirely **operational rather than architectural**, and they cluster around one fact: **development, testing and production currently share a single server and a single MySQL instance.** The live database has been damaged twice by this arrangement (an emptied `migrations` table after a restore; an emptied `sequences` table discovered today, which blocked all patient registration with a 500). The release signing key sits on the same web server, which converts a single server compromise into a supply-chain compromise of every future clinic. These are solvable with process and hosting changes, not rewrites.

Secondary theme: several modules are **built but not yet operational** — messaging runs on the log driver with no inbound webhook, CDSS rules ship inactive pending clinical review, the formulary is empty, and the pharmacy module is roughly 40% complete. None of these is broken; they are honest staging, but the gap between "code complete" and "commercially live" should be tracked deliberately.

---

## 2. System Health Scores

| Area | Score | One-line justification |
|---|---|---|
| Architecture | **9/10** | Enforced boundaries, published contracts, slots, DTOs; deduction for the half-built pharmacy module live on the production install |
| Security | **7.5/10** | Application-level security is strong; operational security (key custody, shared prod/dev host, missing STOP webhook) drags it down |
| Database | **8.5/10** | Ledgers, snapshots, constraints and indexes are exemplary; restore fragility is proven, not theoretical |
| Performance | **7/10** | Fine at target scale (small clinics); no config/route cache on this host, no monitoring, no load evidence beyond tests |
| Code quality | **9/10** | Consistent, gated, and documented with intent-level comments; minor debt in the dual UI stacks |

---

## 3. Critical Issues

### C1. Production, development and testing share one server and one MySQL instance — **[verified, twice bitten]**

- **How it happened:** the product was developed on the machine that then became its first production host.
- **Business impact:** live patient data has already been damaged twice by development activity (restored DB with empty `migrations` table; empty `sequences` table found today — patient registration returned 500s until repaired). A stray `migrate:fresh` from a test session is one keystroke from destroying the live clinic.
- **Technical impact:** `config:cache` is forbidden on this box (tests would then hit the live DB), and route caching breaks module routes — so production runs permanently uncached, and two safety valves normal Laravel deployments rely on are unavailable.
- **Short-term:** treat `sequences:reconcile --dry-run` and a `migrations`-table check as mandatory post-restore steps (command exists, tested today); keep `.env.testing` guards in place; never run destructive artisan commands without checking `APP_ENV`.
- **Long-term:** move the live clinic to its own cPanel host — the product is literally designed for that deployment. This one change also unlocks config/route caching in production and removes findings C1, and halves C2's blast radius.
- **Prevention:** production hosts never carry a `.env.testing`, a test database, or developer tooling.

### C2. Release signing key — **RESOLVED before this audit; verified today**

- **Correction:** the audit draft carried a stale note. Fresh inspection (2026-08-04) shows **no production private key exists anywhere on this box**: `C:\laragon\release-keys` holds only the committed test fixtures, and the licence server's `UPDATE_SIGNING_KEY` is unset. The first keypair — generated on this machine on 3 Aug — was **securely destroyed rather than shipped**, and its SPKI SHA-256 fingerprint is blacklisted in `ReleaseKeyProvenanceTest` so it can never be reintroduced even re-wrapped.
- **Current posture (fail-safe):** `config/updates.php` ships `public_key => ''`, and `PackageSignatureTest` proves the updater installs nothing unsigned. A build that cannot verify is strictly safer than one carrying a key an attacker on this box could have copied.
- **Remaining action (at first release, not before):** run `tools/generate-release-key.php` on a machine that serves nothing (standalone, no network; procedure in `docs/UPDATES.md`); embed only the public half in the clinic build; set `UPDATE_SIGNING_KEY` only on the signing machine.
- **Prevention:** already in place — the fingerprint blacklist plus the provenance test.

### C3. Backups are local-only, with no restore verification and no alerting — **[known, recorded]**

- **Why it matters:** the daily 02:30 dump lands on the same disk as the database it protects. A disk failure, ransomware event, or the server compromise from C2 destroys both. Nobody is alerted if the dump silently stops.
- **Short-term:** copy the nightly dump off-box (any object storage or even a second machine); add a scheduled check that alerts when the newest dump is older than 26 hours.
- **Long-term:** automated restore-verification (restore into a scratch DB, count rows, compare against expectations) — the Updates module already contains a tested dump/restore implementation to build on.
- **Prevention:** the monitoring itself.

### C4. Inbound messaging does not exist: STOP is not honoured, deliveries are never confirmed — **[verified today; latent until credentials are entered]**

- **Severity note:** High *today* (log driver, nothing reaches patients); **Critical the day real credentials are entered**, because a patient replying STOP would be silently ignored — a compliance failure under both Meta's platform policy and PTA rules.
- **How it happened:** the outbound half was built end-to-end (consent, quiet hours, dedupe, retry, 36 tests); the webhook half was scoped but not yet built (task #123 carries the full spec including signature verification — `X-Twilio-Signature` HMAC-SHA1, `X-Hub-Signature-256` HMAC-SHA256 — without which the public webhook would let anyone opt patients in or out).
- **Fix:** build task #123 (webhooks + STOP + delivery receipts + test-send button) **before** credentials are entered, and make the settings screen refuse to go live without the webhook URL confirmed.
- **Prevention:** the test-send button, so the first real send is never the test.

---

## 4. High Priority Issues

| # | Finding | Impact | Fix |
|---|---|---|---|
| H1 | **Pharmacy module ~40% complete but enabled on live.** Schema (4 migrations), enums, models, FEFO allocator and quantity parser are done and gated; `DispensingService`, all screens, dashboard and 14 reports are not. Menu entries are dormant behind `Route::has()` so nothing user-facing is broken **[verified]** | Incomplete surface; the FEFO concurrency path has no tests yet | Complete tasks #120–#122; do not announce the module until the seven listed invariants are tested |
| H2 | **No CI pipeline.** The suite runs manually, on the shared box, against one shared `clinic_test` DB (concurrent runs corrupt each other — recorded incident) | A change can reach live untested; two agents/testers collide | Short: discipline + the existing "one runner at a time" rule. Long: any CI (GitHub Actions with MySQL service) — the suite is already CI-shaped |
| H3 | **No error monitoring or alerting.** `LOG_LEVEL=warning`, log files only; today's registration-blocking 500 was discovered by the user hitting it | Failures are invisible until a human trips over them | Short: a scheduled task alerting on ERROR lines. Long: Sentry/Flare-class capture, per-install |
| H4 | **`MAIL_MAILER=log` on the live install** — password resets and any mail currently deliver to a log file until the clinic configures SMTP in Settings **[verified]** | Staff who request a reset receive nothing, with no visible error | Correct as a shipped default; add a Settings-screen banner ("email is not configured — resets will not deliver") mirroring the messaging outbox's honest "log only" warning |
| H5 | **Audit rows lose the actor on console-initiated actions** — `user_id = NULL` on today's archive because the audit logger reads the request context, not the actor passed to the service **[verified today]** | "Who did this" unanswerable for CLI/scheduled actions in a system whose audit trail is otherwise chained and tamper-evident | Let `AuditLogger` accept an explicit actor and fall back to request context; backfill is impossible, so fix soon |
| H6 | **`QUEUE_CONNECTION=database` with no worker.** Currently zero `ShouldQueue` classes exist **[verified]**, so it is latent — but the first future queued job (or a package's) rots silently forever | Silent non-delivery of whatever queues first | Set `sync` for cPanel deployments, or drain via scheduler (`queue:work --stop-when-empty` each minute); add an architecture test forbidding `ShouldQueue` while no worker strategy exists |
| H7 | **The scheduler only began running today.** `Clinic-Scheduler` was created this session; licence heartbeat and update checks had therefore never fired since deployment, and nothing noticed | The licence server believed nothing, and no one knew | Fixed locally; add server-side monitoring on the licence server ("installation silent > 48h" alert) so a dead scheduler is seen from the other end |

---

## 5. Medium Priority

- **M1 — Restore runbook.** `sequences:reconcile` exists and is tested, but the restore procedure that caused both incidents is undocumented. Write it: restore → check `migrations` count → `sequences:reconcile --dry-run` → apply.
- **M2 — Test data in the live register** (the archived Faker patient, oddly-archived patient 1). Hygiene symptom of C1. Purge deliberately or leave archived, but record which.
- **M3 — Major-version updates pending** (Laravel 13, spatie/permission 8, Pest 4, Intervention 4). No advisories today **[verified]**, so schedule as a maintenance window, not an emergency.
- **M4 — Two UI stacks.** Bootstrap+jQuery (with yajra/datatables — which *is* used, via `QueryObject`/`data-table` **[verified]**) and the `.px-*` premium tier coexist; four screens (Invoices, Payments, Users, Roles) remain unmigrated. Converge on `.px-*` screen by screen; retire datatables when its last table goes.
- **M5 — `RecommendedQuantity` has no test file.** Two real dosage-math bugs were found by probing today ("twice daily" halved every course; "5ml" recommended 105 units). The 14 probe cases must become a dataset test.
- **M6 — Data-subject erasure is deferred by design** (archive comments call it "a separate, deliberate process") but no such process exists. Acceptable for the market; document the policy so a clinic can answer a request.
- **M7 — `docs/ARCHITECTURE.md` ends stale** (pre-dates ~6 modules).
- **M8 — Suite runtime 13.5 min**, dominated by two real-mysqldump tests (26s + 81s) and per-file cold boots (~60–100s first test). Fine today; group slow tests for a future CI matrix.
- **M9 — Insurance fields** are referenced in product/design briefs but have no schema. Decide in or out.

## 6. Low Priority

- Appointments "payment" column awaits a Billing read-port (noted earlier this project).
- Portal quarantine: pending (never-reviewed) uploads accumulate indefinitely; add an age-based nudge on the review screen.
- `laravel.log` (test channel) reached 7.9 MB on the shared box — symptom of C1, cosmetic otherwise.
- Formulary import is CLI-only (`formulary:import`); a screen can wait until the data exists.

---

## 7. What Is Right (and must not be "improved" away)

Verified this session, mostly by tests that would fail if regressed:

- **Module system:** contracts/events/DTOs/enums are the only legal cross-module surface; slots (`encounter.sidebar`, `patient.clinical`, `patient.administrative`, settings tabs) let modules contribute UI without the host knowing them; every module is removable.
- **Patient files:** magic-byte sniffing (browser MIME ignored), image re-encode (strips EXIF, kills polyglots — a PHP script named `report.jpg` claiming `image/jpeg` is refused **[verified]**), ULID storage names, private disk, attachment-only streaming with `nosniff` and a deny-all CSP, downloads logged to the record-access log.
- **Audit:** hash-chained, tamper-evident (`detects tampering that bypasses the model` passes), secrets redacted, model-level refusal of update/delete.
- **Money:** integer minor units end-to-end; allocation sums exactly; currency mixing refused.
- **Records that patients hold copies of** (prescriptions, protocols, dispensings, messages, upload decisions) snapshot their content; editing a template never rewrites history.
- **Concurrency:** booking double-book prevention holds at the DB when the app check is bypassed **[verified]**; sequences are gapless and transactional; messaging dedupe and dispensing idempotency are unique-index-backed; FEFO allocator refuses to run unlocked.
- **API:** `/api/v1`, the exact requested envelope, allow-listed resources (never model dumps), registered error codes (architecture-tested), same policies and services as the web, uniform 404 for cross-branch vs nonexistent.
- **Auth:** 2FA, passkeys, password history/policy, lockout+unlock, seat limits, last-super-admin protection, no public registration, throttled uniform-response password reset.
- **Licensing/Updates:** offline grace ladder, signed packages (wrong key/product/tampered bytes all refused **[verified]**), pre-update backup with byte-for-byte restore test, business-hours guard, env/patient-data never overwritten.
- **OWASP top-10 pass:** no injection findings (Eloquent/bound queries throughout), XSS guarded (escaped Blade; release notes explicitly escaped **[test]**), CSRF on all forms, mass-assignment strict (`$fillable` minimal; status fields deliberately unfillable), SSRF surface none (outbound calls are to fixed provider endpoints), secrets write-only in UI, `APP_DEBUG=false`, `SESSION_SECURE_COOKIE=true` **[verified]**.

---

## 8. Missing Features (commercial readiness)

| Feature | State |
|---|---|
| Pharmacy dispensing (service, queue, screens, reports) | ~40%, specced in tasks #120–122 |
| Inbound messaging: STOP, delivery receipts, test send | Specced (task #123), not built |
| Formulary content | Importer built; no data yet (user-sourced) |
| CDSS rules live | 6 shipped inactive by design; needs clinical sign-off |
| Off-site backup + alerting | Missing (C3) |
| Error monitoring | Missing (H3) |
| CI | Missing (H2) |
| Patient merge (duplicates are *detected* and paused, not mergeable) | Missing |
| Reports centre (cross-module; pharmacy's 14 reports pending) | Partial |
| Insurance data | Undecided (M9) |

---

## 9. Recommended Roadmap

**Phase 1 — Critical fixes (before any new clinic installs):**
C2 move the signing key off-box → C3 off-site backups + staleness alert → C1 split prod from dev (one cPanel migration; enables config/route cache) → H5 audit actor fix → M1 restore runbook.

**Phase 2 — Product stabilisation:**
Finish pharmacy (#120–122, with the seven FEFO invariants tested) → messaging inbound (#123) → M5 quantity-parser tests → H4 mail-unconfigured banner → H6 queue decision + guard test → H2 minimal CI.

**Phase 3 — Commercial SaaS readiness:**
H3 error monitoring per install → H7 licence-server-side silence alerting → M4 finish `.px-*` migration, retire datatables/jQuery → installer polish (the 9-step provisioning in TENANCY.md as a guided flow) → M3 framework major upgrades in a maintenance window.

**Phase 4 — Advanced:**
Patient merge → reports centre → insurance (if in) → OTP self-service portal atop the link infrastructure → per-patient language for messaging templates.

---

*Prepared from direct inspection and this session's verified test/live evidence. Where a finding rests on session memory rather than a fresh probe it is marked as such; nothing here is assumed from convention.*
