Last month, teams on Tessel ran 52,000 replays before publishing a change. One in eleven showed a difference the author had not expected, and in 600 cases the diff was serious enough that the change never shipped. Replay takes a draft workflow, runs it against the last seven days of real events, and shows you exactly which outputs would change. Here is how it works.
What replay does
When you press replay on a draft, Tessel collects every trigger event the live workflow received in the past seven days. It runs the draft against each one in a sandbox, then compares the draft's output for every block with what the live version produced at the time. You get a diff, not a test report: which runs changed, which block changed them, and how.
Where the events come from
Replay does not need a test fixture, because the traces already hold the inputs. Every run stores the input and output of each block, so the trigger event and every intermediate value are already on disk in the workspace's region. Replay reads them from there and never moves data across regions.
For a busy workflow, seven days can mean hundreds of thousands of events. We run them in parallel on isolated workers, and a replay of 50,000 events usually finishes in under four minutes.
Reading the diff
The summary groups changed runs by block and by field, so one logic change does not show up as 300 separate surprises.
Click any group to open one of the affected runs, with the live trace and the replayed trace side by side.
Side effects stay in the sandbox
A replay must never post a journal entry, send a Slack message or charge a card. The sandbox treats connector calls in three ways:
Writes are intercepted, recorded in the replay trace and answered with the response the live run received.
Reads return the values stored in the original trace, so a replay is repeatable: run it twice and you get the same diff.
Approvals are skipped and marked as pending, so nobody is asked in Slack to approve a run that never happened.
The rounding change
In July, a team changed how a logic block rounded currency amounts, from rounding each line to rounding the total. It looked harmless and passed review. Replay showed 311 of 48,213 runs flipping from matched to variance, all of them invoices with more than ten lines. The author kept per-line rounding, and nothing reached production.
The best time to find a bug is last week, before you ship it.
Thirty days on Team
Seven days catches most problems, but not the ones tied to month-end or a monthly billing cycle. We are extending the replay window to 30 days on the Team plan, using the same traces that are already kept for that long. Enterprise workspaces will be able to choose any window inside their one-year retention.
Replay is included on every paid plan, and the pricing page lists the trace retention for each one.
SM
Sofia Marques
Tracing lead
Sofia leads the work on replay and the run inspector inside the tracing team. Before Tessel she built the event store for a European airline, replaying booking events to find out why one seat had been sold twice.