Trellis
Reliability is not retry logic sprinkled over a request handler. It is a durable execution boundary that survives process loss, and a compensation path for the work that already succeeded.
Builder · Workflow design, signal handling, compensation, and queue isolation
+2 more
~2s
end to end, against a 15s budget
300s
injected hang every step must survive
2
isolated task queues, separate worker processes
Problem
Every business step in the brief calls a function that randomly throws or sleeps for 300 seconds. A normal request handler either blocks on it, loses the work when the process dies, or leaves a payment charged against an order that never shipped.
Why it matters
The interesting part is not retrying. It is what the system owes the customer when step three fails after step two took their money, and whether a human approval can be part of a workflow without a thread waiting on it.
Architecture
8 stages · select to inspect
Technical challenges
Human approval without a waiting thread
A durable timer raced against an inbound signal replaces the usual blocking call, so an order can park at manual review for as long as the business allows without holding any process open.
Compensation after partial success
Cancelling before shipment is honored, but if payment already went through the terminal state records the refund compensation rather than silently dropping the charge.
Failure isolation across queues
Two workers polling two queues means shipping failures cannot starve order processing, and the child physically runs in a different process.
Tradeoffs
Child workflow over an inline activity
Costs an extra queue and worker, buys independent scaling and a failure boundary that shows up in the event history.
Bounded restarts before compensating
Unbounded retry on a genuinely dead carrier turns a recoverable order into an indefinite one; the bound forces a decision.
Postgres for application state, Temporal for execution state
Keeps the workflow engine authoritative about progress without making it the system of record for orders.
Experiments
- 01Ran the full lifecycle repeatedly against injected random failures and 300-second hangs to confirm the time budget held.
- 02Exercised cancel, address-update, approve, and dispatch-failed signals against workflows parked at different steps.
- 03Forced carrier dispatch to exhaust its retries to verify the parent restart bound and the compensation path.
Results
Completes in about two seconds against a 15-second budget despite a random throw-or-hang on every business step.
Live status query reporting the current step, flags, per-activity retry counts, and the last error.
Whole stack reproducible with one docker compose command — Temporal server, Postgres, two workers, and the API.
Lessons learned
- A durable timer racing a signal removes a whole class of state machine that would otherwise need its own storage.
- Queue isolation is only real when the child runs in a different process, not just under a different name.
- Event history is the observability story; reading it top to bottom explains the run better than any log line.
Future work
- Add idempotency keys on the payment activity so a duplicated start cannot double-charge.
- Model the refund as a first-class compensating workflow rather than a terminal-state record.
- Load-test queue isolation under a sustained shipping backlog.