Skip to content
Selected work
Durable execution · Distributed systems · 2026

Trellis

Reliability is not retry logic sprinkled over a request handler. It is a durable execution boundary that survives process loss, and a compensation path for the work that already succeeded.

Builder · Workflow design, signal handling, compensation, and queue isolation

~2s

end to end, against a 15s budget

300s

injected hang every step must survive

2

isolated task queues, separate worker processes

01

Problem

Every business step in the brief calls a function that randomly throws or sleeps for 300 seconds. A normal request handler either blocks on it, loses the work when the process dies, or leaves a payment charged against an order that never shipped.

02

Why it matters

The interesting part is not retrying. It is what the system owes the customer when step three fails after step two took their money, and whether a human approval can be part of a workflow without a thread waiting on it.

03

Architecture

8 stages · select to inspect

The manual-review gate is a durable timer racing a human approve_order signal — approve and it proceeds, let the timer win and the order cancels. Shipping is a child workflow on its own task queue in its own worker process, so it can be scaled or deployed independently; on exhausted retries it signals the parent, which restarts it up to a bound before compensating.
04

Technical challenges

01

Human approval without a waiting thread

A durable timer raced against an inbound signal replaces the usual blocking call, so an order can park at manual review for as long as the business allows without holding any process open.

02

Compensation after partial success

Cancelling before shipment is honored, but if payment already went through the terminal state records the refund compensation rather than silently dropping the charge.

03

Failure isolation across queues

Two workers polling two queues means shipping failures cannot starve order processing, and the child physically runs in a different process.

05

Tradeoffs

Child workflow over an inline activity

Costs an extra queue and worker, buys independent scaling and a failure boundary that shows up in the event history.

Bounded restarts before compensating

Unbounded retry on a genuinely dead carrier turns a recoverable order into an indefinite one; the bound forces a decision.

Postgres for application state, Temporal for execution state

Keeps the workflow engine authoritative about progress without making it the system of record for orders.

06

Experiments

  1. 01Ran the full lifecycle repeatedly against injected random failures and 300-second hangs to confirm the time budget held.
  2. 02Exercised cancel, address-update, approve, and dispatch-failed signals against workflows parked at different steps.
  3. 03Forced carrier dispatch to exhaust its retries to verify the parent restart bound and the compensation path.
07

Results

Completes in about two seconds against a 15-second budget despite a random throw-or-hang on every business step.

Live status query reporting the current step, flags, per-activity retry counts, and the last error.

Whole stack reproducible with one docker compose command — Temporal server, Postgres, two workers, and the API.

08

Lessons learned

  • A durable timer racing a signal removes a whole class of state machine that would otherwise need its own storage.
  • Queue isolation is only real when the child runs in a different process, not just under a different name.
  • Event history is the observability story; reading it top to bottom explains the run better than any log line.
09

Future work

  • Add idempotency keys on the payment activity so a duplicated start cannot double-charge.
  • Model the refund as a first-class compensating workflow rather than a terminal-state record.
  • Load-test queue isolation under a sustained shipping backlog.