Chimti Workers & Automation Plan (OTLU-admin style)
Date: 2026-08-01. Researched against chimti-api as-is (routes, Prisma schema,
modules). Everything listed below is backed by an existing table, field or
function — nothing invented.
Current state (verified)
- Only ONE background worker exists: `startEmailSyncWorker` (IMAP sync,
in-process setInterval, `EMAIL_SYNC_INTERVAL_SECONDS`, default 60s).
- Everything else runs INLINE in request handlers: communication sends fire
during the order-status PATCH; held messages need a MANUAL retry endpoint;
demo reset is a manual endpoint; processing-queue escalation is a manual POST.
- Redis is provisioned in docker-compose (port 6380) but completely unused —
no queue library in package.json.
- No cleanup anywhere: AuthSession / LoginOtpChallenge / LoginAuthAttempt /
AuditLog / WhatsAppWebhookEvent rows grow forever.
Infra recommendation — two steps
**W0 (now, zero new infra):** add `src/worker.js` in chimti-api — same codebase,
same image, run as a SECOND process (Coolify: duplicate app with start command
`node src/worker.js`, or PM2 second entry). Inside: a tiny job registry —
`{ name, intervalSeconds, envFlag, run() }` — with a Postgres advisory lock per
job (`pg_try_advisory_lock`) so multiple instances never double-run. Every job
ships OFF by default behind an env flag (`WORKER_OUTBOX=1` etc.). Move the
existing email IMAP sync into this registry too.
**W1 (when volume grows):** switch the dispatcher-style jobs to BullMQ on the
already-provisioned Redis. The registry API stays the same; only the scheduler
changes. Do NOT start here — queue infra before queue need is pure overhead.
A `worker_heartbeats` table (name, lastRunAt, lastOkAt, lastError) gives the
admin app a Workers health card later (System → Settings).
Worker catalog — 23 workers by category
A. Communication (biggest load + reliability win)
| # | Worker | Cadence | What it does (backing) |
|---|--------|---------|------------------------|
| 1 | **Outbox dispatcher** | 30s | Deliver `CommunicationOutboxMessage` QUEUED + retryable FAILED via provider; log `CommunicationDeliveryAttempt`. Moves sending OUT of the order-PATCH request path (today inline). |
| 2 | **Held-message releaser** | 5 min + on top-up | `HELD_INSUFFICIENT_CREDITS` → re-attempt when wallet ≥ creditCost (today: manual `/communication-outbox/:id/retry`). |
| 3 | **Delivery-status reconciler** | 2 min | Drain `WhatsAppWebhookEvent` rows → conversation/message status updates (SENT→DELIVERED→READ), idempotent. |
| 4 | **Credit low-balance alerter** | daily | `CommunicationCreditWallet` below threshold → alert brand owner + Chimti team; prevents silent HELD pile-ups. |
| 5 | **Email IMAP sync** | 60s | EXISTS — relocate into the worker process so API pods stay request-only. |
B. Revenue / SaaS lifecycle
| # | Worker | Cadence | What it does |
|---|--------|---------|--------------|
| 6 | **Trial & subscription lifecycle** | daily | `BrandSubscription`: TRIAL past `endsAt` → EXPIRED; endsAt in 14/7/3 days → reminder + founder digest line; ACTIVE past endsAt → PAUSED. |
| 7 | **Invoice overdue + dunning** | daily | `Invoice.dueAt` passed & not PAID → OVERDUE + fire existing `PAYMENT_REMINDER` communication rule. |
| 8 | **Founder daily digest** | 8 am | Yesterday's sales/collections/new leads/risk flags → WhatsApp/email to founder. Uses billing + orders reads that already exist. |
| 9 | **Portfolio nightly snapshot** | nightly | Roll up MRR/ARR/collection per brand into a snapshot table → founder deck loads instantly + gets history/trend. |
C. Operations SLA
| # | Worker | Cadence | What it does |
|---|--------|---------|--------------|
| 10 | **Promise-breach watcher** | 15 min | Orders with `promisedAt` passed & not READY/DELIVERED → flag + notify store manager (internal chat / WhatsApp). |
| 11 | **Shipment schedule nudger** | hourly | `Shipment.scheduledAt` within 60 min & no rider → alert; overdue in-transit legs escalate. |
| 12 | **Processing-queue auto-escalation** | 15 min | Items past `etaMins` → call the SAME logic as the existing manual escalate endpoint (`escalatedAt`). |
| 13 | **Inventory reorder alerter** | daily | `InventoryItem.stock <= reorderPoint` → purchase list to owner. Fields already exist. |
D. CRM / Growth
| # | Worker | Cadence | What it does |
|---|--------|---------|--------------|
| 14 | **Lead follow-up nudger** | hourly | `LeadRecord.followUpAt` passed & stage open → ping Sales/BD; schema already has the index `(stage, followUpAt)`. |
| 15 | **Post-delivery NPS ping** | hourly | DELIVERED + N hours → feedback message via the communication engine; fills Customer Experience feedback buckets. |
| 16 | **Churn-risk tagger** | weekly | `lastOrderAt` older than threshold → recompute health/segment tags shown in the Customers register. |
| 17 | **Loyalty renewal reminder** | daily | `LoyaltySubscription` nearing end → renewal message; expire lapsed ones. |
| 18 | **Ad campaign budget guard** | daily | `AdCampaign` past end date / budget → auto-pause + owner notification. |
E. Platform hygiene
| # | Worker | Cadence | What it does |
|---|--------|---------|--------------|
| 19 | **Demo workspace auto-reset** | nightly | `resetDemoWorkspace()` ALREADY EXISTS as a function — just schedule it (skip if a demo is mid-session). |
| 20 | **Auth sweeper** | daily | Delete expired `AuthSession`, `LoginOtpChallenge`, stale `LoginAuthAttempt`. |
| 21 | **Audit/webhook retention** | weekly | Archive `AuditLog` + processed `WhatsAppWebhookEvent` older than 180d. |
| 22 | **Photo/storage compactor** | weekly | Large `OrderItemPhoto` data-URLs → Drive storage module; keeps Postgres lean. |
| 23 | **AI account status refresher** | hourly | `refreshAiSettingStatus()` exists — keep Genie's Codex card current without a user visit. |
Rollout order (value ÷ effort)
1. **Wave 1 — Comms reliability:** #1, #2, #3 (+#5 relocation). Direct customer-visible reliability; removes inline sending load.
2. **Wave 2 — Revenue:** #6, #7, #8. Money never sleeps on a manual click.
3. **Wave 3 — Ops SLA:** #10, #12, #13, #11.
4. **Wave 4 — Growth:** #14, #15, #16, #17.
5. **Wave 5 — Hygiene:** #19, #20, #21, #22, #23, #9, #18.
Every wave = one drop in chimti-api (`src/worker.js` + jobs + env flags) with
zero schema changes except the heartbeat table (Wave 1).
Implementation status — BUILT (2026-08-01)
All five waves are implemented in chimti-api as **21 runnable jobs** (25 unit
tests green, run `npm run test:workers`):
- `src/worker.js` — runner: advisory locks, staggered start, heartbeats,
graceful SIGTERM drain. Exits cleanly unless `WORKERS_ENABLED=1`.
- `src/workers/` — support.js + wave1-comms / wave2-revenue / wave3-ops /
wave4-growth / wave5-hygiene. Delivery reuses the communication engine's
`retryCommunicationMessage` / `enqueueOrderCustomerUpdate`; demo reset reuses
`resetDemoWorkspace()`.
- Two catalog items intentionally reduced (schema honesty):
**#18 ads guard** is report-only (AdCampaign has no end/budget dates) and
**#23 AI status refresher** is deferred (refresh helper lives un-exported
inside routes/ai.js — export it first). Both noted above.
Deploy (with the Phase-5 cutover or after)
Coolify: duplicate the chimti-api app → same repo/image, start command
`npm run worker`, no public domain. Env = API env **plus**:
WORKERS_ENABLED=1
WORKER_ALERT_EMAIL=me@jagjeetsinghsethi.com # optional; alerts otherwise log-only
WORKER_DIGEST_HOUR=8
WORKER_SNAPSHOT_HOUR=2
WORKER_DEMO_RESET_HOUR=4
# any job off: WORKER_OUTBOX=0, WORKER_NPS_PING=0, …
# any cadence: WORKER_OUTBOX_INTERVAL_SECONDS=15, …
First boot creates `worker_heartbeats` + `worker_snapshots` itself (no
migration). Check health: `SELECT * FROM worker_heartbeats ORDER BY name;`