In September, Tessel executed 2.1 billion runs across 12 regions. Every one of them left a trace with the input and output of each block, its timing, the decisions an agent made and who approved what. On the Team plan we keep that record for 30 days, and you can open any of it in under 200 ms. This post explains how the pipeline works, what we got wrong the first time, and what it costs us.
What a trace holds
A run is one trigger firing one workflow from start to finish. A trace is the record of that run, split into spans, one span per block. A typical finance workflow passes through 14 blocks, so it produces 14 spans plus a header. The median trace is 11 KB compressed. The largest we stored last month was 38 MB, a NetSuite export that someone pushed through a single logic block.
Each span carries the block version, start and end time, the typed input and output, and for agents the tool calls, spend and approver. The typing matters. Because every block declares its schema, we can store payloads as columns instead of JSON blobs, and that is where most of the savings below come from.
Writing without slowing the run
Our first rule was that tracing can never add latency to a block. The p95 per block is 184 ms, and the trace write is not allowed to be part of that number. Each runtime worker appends spans to a local ring buffer and returns. A sidecar drains the buffer every 250 ms and ships batches to a regional ingest queue.
If the queue is unavailable, the sidecar spills to local disk and retries with backoff. In the last 12 months we have lost zero spans. We came close once.
Fig. 1 · Worker ring buffer, sidecar, regional queue, columnar store
The Frankfurt disk incident
In March, a misconfigured volume in eu-central filled up during a queue outage, and spill files stopped writing for 41 seconds. The ring buffers held every span until the queue recovered, so nothing was lost. It was still closer than we liked. We now alert when spill disk reaches 50 percent, and each worker's buffer is sized for 90 seconds of peak traffic.
Typed payloads as columns
Spans land in a columnar store partitioned by workspace, region and hour. Data never leaves the region a workspace is pinned to, so we run 12 independent clusters instead of one large one. That made residency simple and capacity planning a little harder.
Storing typed payloads as columns changed three things for us:
Compression went from 4.1x with JSON to 11.6x, because a column of invoice amounts compresses far better than thousands of nested objects.
Filtered queries, such as every run where
customer_idmatches one value, read a single column instead of every payload.Schema changes are explicit. A new block version with a new field gets a new column, and older traces keep their original shape.
Thirty days of traces for the whole platform comes to about 1.9 PB after compression. Enterprise workspaces keep a full year, and their plan covers the storage.
Opening a trace in 200 ms
Most people open a trace from a Slack link or a failed-run alert. That path is predictable, so we optimise for it. The header of every run sits in a small hot index keyed by run id. Opening a trace reads the header, then fetches that run's spans from a single partition. The p95 from click to first paint in the trace view is 162 ms.
If a trace takes longer to open than the bug took to happen, nobody will open it.
Search across many runs is slower, and the product says so. A query over 30 days of a busy workflow can take several seconds. We stream partial results as each hour partition returns, so you see the first matches long before the last one arrives.
What comes next
Two projects are in progress. Replay already uses traces to run a draft workflow against last week's real events, and we are extending that window to 30 days on Team. We are also opening trace export, so you can stream spans into your own Snowflake or S3 bucket and keep them for as long as your auditors ask.
If this is the kind of problem you enjoy, we are hiring a Staff Engineer for tracing, in Rotterdam or remote within the EU.
MW
Maarten de Wit
Engineering lead, Tracing
Maarten leads the team that owns traces, replay and the run inspector. Before Tessel he worked on payment ledgers at a Dutch bank, which is where he learned to keep every record.