Files
familysync/.planning/phases/09-faster-write-back/09-CONTEXT.md
T

8.5 KiB
Raw Blame History

Phase 9: Faster Write-Back - Context

Gathered: 2026-06-12 Status: Ready for planning

## Phase Boundary

Make a created/edited/deleted event reach Fastmail in ~1-2s instead of waiting up to ~15s for the next interval tick, by signalling the outbox drain on enqueue (event-driven) — with every existing outbox durability guarantee intact.

In scope: the in-process signal mechanism, its wiring into the existing drain path, and a trailing re-drain guarantee under bursts.

Out of scope (locked by roadmap — do NOT reopen):

  • Redis for the drain (Redis stays only for list SSE). Single-process, single Node module instance by design.
  • node-cron (silently skipped ticks in the long-lived process — setInterval only). See [[node-cron-skips-in-long-running-process]].
  • Any change to the route handlers' optimistic-202 / no-inline-CalDAV contract.
  • Any change to outbox durability semantics (fresh-etag-before-PUT, 412 conflict flow, per-uid exactly-once, edit-as-move create-before-delete, dead-letter/backoff).
  • Multi-process / multi-replica row-claim (the isDraining declaration comment documents the future UPDATE ... status='processing' path; not this phase).
## Implementation Decisions

Signal mechanism (locked from ROADMAP.md — restated so the planner doesn't re-derive)

  • D-01: New artifact is a single zero-dependency in-process EventEmitter at apps/api/src/lib/outboxTrigger.ts. No new npm dependency.
  • D-02: The signal funnels through the existing isDraining-guarded runOutboxDrain() path via a drainRequested flag. The trigger must NEVER call runOutboxDrain() directly in a way that bypasses the isDraining guard or escapes the error-caught wrapper (Pitfall 5 — no double-drain).
  • D-03: The signal is published after the enqueue transaction commits — never between the two inserts of an edit-as-move (Pitfall 6). For a move, publish once after both CREATE and DELETE rows are committed, so create-before-delete ordering is never raced by the signal.
  • D-04: Route handlers still return optimistic 202 immediately and make no inline CalDAV call; the signal is fire-and-forget.

Burst behavior — trailing re-drain (the main decision this discussion locked)

  • D-05: Guarantee the last edit in a burst. When a signal arrives while a drain is already in flight, set drainRequested and re-run the drain exactly once after the current one finishes (drain-again-if-requested loop), so the LAST edit in a rapid burst still lands in ~1-2s rather than waiting for the next 15s tick.
  • D-06: No debounce/coalesce window and no rejection of mid-drain signals. Rationale: two-user household — burst volume is tiny; correctness/snappiness of the trailing edit outweighs Fastmail request-rate politeness. (A debounce knob was considered and rejected as premature.)
  • D-07: Implementation must preserve exactly-once per uid when the signal and the 15s fallback overlap — the trailing re-drain reuses the same pending AND next_attempt_at <= NOW() selection and isDraining guard, so no duplicate CalDAV PUT for the same row (Success Criterion 4).

Fallback interval

  • D-08: Keep the 15s setInterval fallback exactly as-is. It still runs for startup catch-up and transient-error recovery (Success Criterion 5). Do not lengthen it — idle DB polling cost is negligible for this deployment and the criterion names 15s explicitly.

Verification

  • D-09: Prove Success Criterion 1 with an automated integration test on the trigger wiring (Vitest): assert that enqueuing a row publishes the signal and that runOutboxDrain is invoked promptly (stub/mock the CalDAV dispatch so no live Fastmail is needed). Also cover the trailing-re-drain case (signal during in-flight drain → exactly one follow-up drain) and the no-double-PUT overlap case.
  • D-10: No operator-stopwatch human checkpoint is required for this phase. Rationale: the dev bypass user (id 1) has no CalDAV credential/calendars, so live event-create 422s in the dev stack — see [[dev-data-user1-no-calendars]]; an automated wiring test is the durable, CI-guarded evidence. (Real wall-clock-to-Fastmail latency is bounded by the CalDAV round-trip regardless and is implicitly exercised in production use.)

Claude's Discretion

  • Exact shape of the drainRequested/re-drain loop (where the flag lives, whether the loop sits inside runOutboxDrain's finally or in the scheduler wrapper) — planner/researcher decides, provided D-02/D-05/D-07 hold.
  • Whether the EventEmitter is a module singleton vs. a tiny custom signal object — either is fine; zero-dep is the only hard constraint.
  • Test file placement follows the existing convention (API tests live under apps/api/tests/, never src/) — see [[api-integration-test-db]].

<canonical_refs>

Canonical References

Downstream agents MUST read these before planning or implementing.

Phase scope & guarantees

  • .planning/ROADMAP.md §"Phase 9: Faster Write-Back" — goal, the 5 Success Criteria, and the 3 owned pitfalls (no double-drain, create-before-delete under concurrent enqueues, hard constraints).

Existing implementation to extend (read before changing)

  • apps/api/src/broker/outboxWorker.tsrunOutboxDrain() (the guarded drain), isDraining flag + its multi-process caveat comment, startOutboxWorker() (the 15s setInterval), edit-as-move ordering + durable CR-04 sibling-status gate, 412/hardfail/backoff/dead-letter handling. This is the path the signal must funnel into.
  • apps/api/src/routes/events.ts — the three enqueue sites (create ~L300, edit/edit-as-move transaction ~L389432, delete ~L511) where the post-commit signal must be published, and the sync-status endpoint the PWA polls.
  • apps/api/src/index.ts §L136141 — background-worker startup wiring (startOutboxWorker beside startBrokerPoller/startReminderScheduler); the trigger must be initialised consistently with this pattern, only under isMainModule().

No external specs/ADRs beyond ROADMAP.md — the durability decisions (D-04…D-12, CR-04/05, WR-01/04/06) are documented inline in outboxWorker.ts and events.ts.

</canonical_refs>

<code_context>

Existing Code Insights

Reusable Assets

  • runOutboxDrain() + isDraining guard (outboxWorker.ts): the signal reuses this verbatim — no parallel drain path. The trailing re-drain (D-05) is a loop around it, not a second drainer.
  • startOutboxWorker() setInterval pattern: the fallback stays; the new trigger sits alongside it and shares the same guarded entrypoint.
  • Existing post-write triggerTargetedResync + RESYNC_TIMEOUT_MS flow already makes status='done' mean "local cache reflects the write" — unaffected by this phase; the signal only changes when the drain starts.

Established Patterns

  • setInterval-only background workers (poller, reminderScheduler, outboxWorker) started exclusively under isMainModule() in index.ts — keeps timers out of the test process. The trigger follows the same gating.
  • Per-uid exactly-once via the pending AND next_attempt_at <= NOW() selection + isDraining; the signal path must not introduce a second selection that could double-dispatch.

Integration Points

  • Enqueue → signal: publish from events.ts after each db.insert/db.transaction commits (create, delete, and the move pair).
  • Signal → drain: outboxTrigger listener sets drainRequested / invokes the guarded drain in outboxWorker.ts.
  • Startup: initialise the trigger subscription in index.ts next to startOutboxWorker().

</code_context>

## Specific Ideas

The 1-2s target is "signal on enqueue, don't wait for the interval" — not a hard real-time budget. Actual latency is bounded by the single CalDAV round-trip to Fastmail, which is unchanged; this phase only removes the up-to-15s queueing delay before that round-trip begins.

## Deferred Ideas
  • Multi-process / multi-replica outbox drain — durable DB row-claim (UPDATE calendar_outbox SET status='processing' WHERE id=? AND status='pending') instead of the in-memory isDraining guard. Documented in the isDraining declaration comment; out of scope while the deployment is single-process. Belongs to a future scaling phase, not v1.1.
  • Debounce/coalesce knob for burst write-back — considered and rejected for now (D-06). Revisit only if Fastmail request-rate ever becomes a concern with more members (see [[project-nmember-expansion]]).

None of the above are in scope for Phase 9.


Phase: 9-Faster-Write-Back Context gathered: 2026-06-12