Skip to content
← Sheet 02 / Projects

Professional

From Polling to Events

Replacing the CRM's polling loops with events published from the database, a single authenticated event stream for the browser, and a durable queue for every production email.

Result

Notifications arrive live instead of every 30 seconds, and emails go out within seconds instead of on a 15-minute sweep.

01

Problem

The CRM found out about changes by asking over and over. The notification badge polled every 30 seconds, conversations every 10, live calls every 5 and customer-portal activity every 10, while notification emails went out on a 15-minute sweep. Screens were stale, requests were wasted, and emails often arrived long after the thing they were about.

02

Start with an inventory

Before changing anything, I traced every polling loop in the codebase (client, worker, shared packages and database) and sorted them into three groups: loops events should replace, loops that should stay as recovery for upstream services whose webhooks can go missing, and timers that never fetch anything at all. That inventory became the plan, and it stopped this turning into an accidental rewrite.

03

Architecture

  • Changes are published from the same database transaction that makes them, so an event can never describe something that didn't commit.
  • Postgres LISTEN/NOTIFY wakes every API instance and the worker the moment a change commits. No separate message broker: Postgres was already the transaction boundary.
  • Each signed-in user gets one authenticated Server-Sent Events stream. It only says what changed; the browser refetches with its own permissions, so nothing private travels over the stream, and a reconnect simply refetches, so nothing is missed.
  • The listener is a dedicated, self-healing connection that never lets an error escape, and the API shuts down gracefully inside the container's stop window.

04

Getting email right

Every production email (notifications, portal sign-in links and member invitations) now goes through a durable queue and a delivery runner in the worker. Workers take a time-limited lease on each delivery, and a hand-off is refused once that lease has expired, so a slow worker can't send something another has already picked up. Content, preferences and eligibility are rechecked before every attempt. If a crash leaves it unclear whether a message left, the delivery is recorded as uncertain rather than sent twice.

05

A bug found along the way

The old email job claimed work and sent it inside one database transaction that only committed afterwards. A crash after a successful send could roll back the claim, despite comments claiming at-most-once delivery. The new runner replaces that lifecycle instead of copying the promise.

06

Rollout

The new path shipped behind a feature switch, with the old sweep kept as a rollback lever, plus summary logging and a warning when the delivery backlog grows. After that, I retired the legacy email path and moved messaging and live calls onto the same event stream.

07

What stays on a timer

Some polling is the right answer. Reconciliation with upstream providers still runs on a schedule, because no amount of internal eventing can replace a webhook that never arrives. Scheduled reminders still fire at their due time.

08

Stack

PostgreSQL (Supabase) with pgTAP tests, Hono, node-postgres, Nodemailer, React with TanStack Query, Vitest and Playwright.