Docs / architecture/workers-automation-plan.md

Chimti Workers & Automation Plan (OTLU-admin style)

Date: 2026-08-01. Researched against chimti-api as-is (routes, Prisma schema,

modules). Everything listed below is backed by an existing table, field or

function — nothing invented.

Current state (verified)

in-process setInterval, `EMAIL_SYNC_INTERVAL_SECONDS`, default 60s).

during the order-status PATCH; held messages need a MANUAL retry endpoint;

demo reset is a manual endpoint; processing-queue escalation is a manual POST.

no queue library in package.json.

AuditLog / WhatsAppWebhookEvent rows grow forever.

Infra recommendation — two steps

**W0 (now, zero new infra):** add `src/worker.js` in chimti-api — same codebase,

same image, run as a SECOND process (Coolify: duplicate app with start command

`node src/worker.js`, or PM2 second entry). Inside: a tiny job registry —

`{ name, intervalSeconds, envFlag, run() }` — with a Postgres advisory lock per

job (`pg_try_advisory_lock`) so multiple instances never double-run. Every job

ships OFF by default behind an env flag (`WORKER_OUTBOX=1` etc.). Move the

existing email IMAP sync into this registry too.

**W1 (when volume grows):** switch the dispatcher-style jobs to BullMQ on the

already-provisioned Redis. The registry API stays the same; only the scheduler

changes. Do NOT start here — queue infra before queue need is pure overhead.

A `worker_heartbeats` table (name, lastRunAt, lastOkAt, lastError) gives the

admin app a Workers health card later (System → Settings).

Worker catalog — 23 workers by category

A. Communication (biggest load + reliability win)

| # | Worker | Cadence | What it does (backing) |

|---|--------|---------|------------------------|

| 1 | **Outbox dispatcher** | 30s | Deliver `CommunicationOutboxMessage` QUEUED + retryable FAILED via provider; log `CommunicationDeliveryAttempt`. Moves sending OUT of the order-PATCH request path (today inline). |

| 2 | **Held-message releaser** | 5 min + on top-up | `HELD_INSUFFICIENT_CREDITS` → re-attempt when wallet ≥ creditCost (today: manual `/communication-outbox/:id/retry`). |

| 3 | **Delivery-status reconciler** | 2 min | Drain `WhatsAppWebhookEvent` rows → conversation/message status updates (SENT→DELIVERED→READ), idempotent. |

| 4 | **Credit low-balance alerter** | daily | `CommunicationCreditWallet` below threshold → alert brand owner + Chimti team; prevents silent HELD pile-ups. |

| 5 | **Email IMAP sync** | 60s | EXISTS — relocate into the worker process so API pods stay request-only. |

B. Revenue / SaaS lifecycle

| # | Worker | Cadence | What it does |

|---|--------|---------|--------------|

| 6 | **Trial & subscription lifecycle** | daily | `BrandSubscription`: TRIAL past `endsAt` → EXPIRED; endsAt in 14/7/3 days → reminder + founder digest line; ACTIVE past endsAt → PAUSED. |

| 7 | **Invoice overdue + dunning** | daily | `Invoice.dueAt` passed & not PAID → OVERDUE + fire existing `PAYMENT_REMINDER` communication rule. |

| 8 | **Founder daily digest** | 8 am | Yesterday's sales/collections/new leads/risk flags → WhatsApp/email to founder. Uses billing + orders reads that already exist. |

| 9 | **Portfolio nightly snapshot** | nightly | Roll up MRR/ARR/collection per brand into a snapshot table → founder deck loads instantly + gets history/trend. |

C. Operations SLA

| # | Worker | Cadence | What it does |

|---|--------|---------|--------------|

| 10 | **Promise-breach watcher** | 15 min | Orders with `promisedAt` passed & not READY/DELIVERED → flag + notify store manager (internal chat / WhatsApp). |

| 11 | **Shipment schedule nudger** | hourly | `Shipment.scheduledAt` within 60 min & no rider → alert; overdue in-transit legs escalate. |

| 12 | **Processing-queue auto-escalation** | 15 min | Items past `etaMins` → call the SAME logic as the existing manual escalate endpoint (`escalatedAt`). |

| 13 | **Inventory reorder alerter** | daily | `InventoryItem.stock <= reorderPoint` → purchase list to owner. Fields already exist. |

D. CRM / Growth

| # | Worker | Cadence | What it does |

|---|--------|---------|--------------|

| 14 | **Lead follow-up nudger** | hourly | `LeadRecord.followUpAt` passed & stage open → ping Sales/BD; schema already has the index `(stage, followUpAt)`. |

| 15 | **Post-delivery NPS ping** | hourly | DELIVERED + N hours → feedback message via the communication engine; fills Customer Experience feedback buckets. |

| 16 | **Churn-risk tagger** | weekly | `lastOrderAt` older than threshold → recompute health/segment tags shown in the Customers register. |

| 17 | **Loyalty renewal reminder** | daily | `LoyaltySubscription` nearing end → renewal message; expire lapsed ones. |

| 18 | **Ad campaign budget guard** | daily | `AdCampaign` past end date / budget → auto-pause + owner notification. |

E. Platform hygiene

| # | Worker | Cadence | What it does |

|---|--------|---------|--------------|

| 19 | **Demo workspace auto-reset** | nightly | `resetDemoWorkspace()` ALREADY EXISTS as a function — just schedule it (skip if a demo is mid-session). |

| 20 | **Auth sweeper** | daily | Delete expired `AuthSession`, `LoginOtpChallenge`, stale `LoginAuthAttempt`. |

| 21 | **Audit/webhook retention** | weekly | Archive `AuditLog` + processed `WhatsAppWebhookEvent` older than 180d. |

| 22 | **Photo/storage compactor** | weekly | Large `OrderItemPhoto` data-URLs → Drive storage module; keeps Postgres lean. |

| 23 | **AI account status refresher** | hourly | `refreshAiSettingStatus()` exists — keep Genie's Codex card current without a user visit. |

Rollout order (value ÷ effort)

1. **Wave 1 — Comms reliability:** #1, #2, #3 (+#5 relocation). Direct customer-visible reliability; removes inline sending load.

2. **Wave 2 — Revenue:** #6, #7, #8. Money never sleeps on a manual click.

3. **Wave 3 — Ops SLA:** #10, #12, #13, #11.

4. **Wave 4 — Growth:** #14, #15, #16, #17.

5. **Wave 5 — Hygiene:** #19, #20, #21, #22, #23, #9, #18.

Every wave = one drop in chimti-api (`src/worker.js` + jobs + env flags) with

zero schema changes except the heartbeat table (Wave 1).

Implementation status — BUILT (2026-08-01)

All five waves are implemented in chimti-api as **21 runnable jobs** (25 unit

tests green, run `npm run test:workers`):

graceful SIGTERM drain. Exits cleanly unless `WORKERS_ENABLED=1`.

wave4-growth / wave5-hygiene. Delivery reuses the communication engine's

`retryCommunicationMessage` / `enqueueOrderCustomerUpdate`; demo reset reuses

`resetDemoWorkspace()`.

**#18 ads guard** is report-only (AdCampaign has no end/budget dates) and

**#23 AI status refresher** is deferred (refresh helper lives un-exported

inside routes/ai.js — export it first). Both noted above.

Deploy (with the Phase-5 cutover or after)

Coolify: duplicate the chimti-api app → same repo/image, start command

`npm run worker`, no public domain. Env = API env **plus**:


WORKERS_ENABLED=1

WORKER_ALERT_EMAIL=me@jagjeetsinghsethi.com   # optional; alerts otherwise log-only

WORKER_DIGEST_HOUR=8

WORKER_SNAPSHOT_HOUR=2

WORKER_DEMO_RESET_HOUR=4

# any job off: WORKER_OUTBOX=0, WORKER_NPS_PING=0, …

# any cadence: WORKER_OUTBOX_INTERVAL_SECONDS=15, …

First boot creates `worker_heartbeats` + `worker_snapshots` itself (no

migration). Check health: `SELECT * FROM worker_heartbeats ORDER BY name;`