/

How we trace two billion runs a month

Engineering

Published

Read time

9

min

How we trace two billion runs a month

Inside the pipeline that stores inputs, outputs and timings for every block, and opens a 30-day trace in under 200 ms.

MW

Maarten de Wit

Engineering lead, Tracing

On this page

Share

In September, Tessel executed 2.1 billion runs across 12 regions. Every one of them left a trace with the input and output of each block, its timing, the decisions an agent made and who approved what. On the Team plan we keep that record for 30 days, and you can open any of it in under 200 ms. This post explains how the pipeline works, what we got wrong the first time, and what it costs us.

What a trace holds

A run is one trigger firing one workflow from start to finish. A trace is the record of that run, split into spans, one span per block. A typical finance workflow passes through 14 blocks, so it produces 14 spans plus a header. The median trace is 11 KB compressed. The largest we stored last month was 38 MB, a NetSuite export that someone pushed through a single logic block.

Each span carries the block version, start and end time, the typed input and output, and for agents the tool calls, spend and approver. The typing matters. Because every block declares its schema, we can store payloads as columns instead of JSON blobs, and that is where most of the savings below come from.

type Span = {
  runId: string;       // ULID, sortable by time
  block: string;       // "reconcile-invoices@v14"
  startedAt: number;   // epoch microseconds
  durationMs: number;
  input: Columns;      // typed by the block schema
  output: Columns;
  agent?: { spendUsd: number; tools: string[]; approver?: string };
};
type Span = {
  runId: string;       // ULID, sortable by time
  block: string;       // "reconcile-invoices@v14"
  startedAt: number;   // epoch microseconds
  durationMs: number;
  input: Columns;      // typed by the block schema
  output: Columns;
  agent?: { spendUsd: number; tools: string[]; approver?: string };
};
type Span = {
  runId: string;       // ULID, sortable by time
  block: string;       // "reconcile-invoices@v14"
  startedAt: number;   // epoch microseconds
  durationMs: number;
  input: Columns;      // typed by the block schema
  output: Columns;
  agent?: { spendUsd: number; tools: string[]; approver?: string };
};

Writing without slowing the run

Our first rule was that tracing can never add latency to a block. The p95 per block is 184 ms, and the trace write is not allowed to be part of that number. Each runtime worker appends spans to a local ring buffer and returns. A sidecar drains the buffer every 250 ms and ships batches to a regional ingest queue.

If the queue is unavailable, the sidecar spills to local disk and retries with backoff. In the last 12 months we have lost zero spans. We came close once.

Fig. 1 · Worker ring buffer, sidecar, regional queue, columnar store

The Frankfurt disk incident

In March, a misconfigured volume in eu-central filled up during a queue outage, and spill files stopped writing for 41 seconds. The ring buffers held every span until the queue recovered, so nothing was lost. It was still closer than we liked. We now alert when spill disk reaches 50 percent, and each worker's buffer is sized for 90 seconds of peak traffic.

Typed payloads as columns

Spans land in a columnar store partitioned by workspace, region and hour. Data never leaves the region a workspace is pinned to, so we run 12 independent clusters instead of one large one. That made residency simple and capacity planning a little harder.

Storing typed payloads as columns changed three things for us:

  • Compression went from 4.1x with JSON to 11.6x, because a column of invoice amounts compresses far better than thousands of nested objects.

  • Filtered queries, such as every run where customer_id matches one value, read a single column instead of every payload.

  • Schema changes are explicit. A new block version with a new field gets a new column, and older traces keep their original shape.

Thirty days of traces for the whole platform comes to about 1.9 PB after compression. Enterprise workspaces keep a full year, and their plan covers the storage.

Opening a trace in 200 ms

Most people open a trace from a Slack link or a failed-run alert. That path is predictable, so we optimise for it. The header of every run sits in a small hot index keyed by run id. Opening a trace reads the header, then fetches that run's spans from a single partition. The p95 from click to first paint in the trace view is 162 ms.

If a trace takes longer to open than the bug took to happen, nobody will open it.

Search across many runs is slower, and the product says so. A query over 30 days of a busy workflow can take several seconds. We stream partial results as each hour partition returns, so you see the first matches long before the last one arrives.

What comes next

Two projects are in progress. Replay already uses traces to run a draft workflow against last week's real events, and we are extending that window to 30 days on Team. We are also opening trace export, so you can stream spans into your own Snowflake or S3 bucket and keep them for as long as your auditors ask.

If this is the kind of problem you enjoy, we are hiring a Staff Engineer for tracing, in Rotterdam or remote within the EU.

MW

Maarten de Wit

Engineering lead, Tracing

Maarten leads the team that owns traces, replay and the run inspector. Before Tessel he worked on payment ledgers at a Dutch bank, which is where he learned to keep every record.

Keep reading

More from the blog. Picked for this post.

Engineering

5

min read

Replaying last week before you ship

How replay runs a draft workflow against seven days of real events and shows the diff before anything reaches production.

Engineering

6

min read

Why every block compiles to TypeScript

Readable code you can review, test and take with you. Why we chose TypeScript as the output of every logic block.

Product

7

min read

Agents need budgets, not just prompts

A prompt says what an agent should do. A spend cap, a tool list and a named approver say what it may never do.

Newsletter

New posts, monthly. Plain text, no tracking.

One email a month · Unsubscribe in one click

One email a month. Only what shipped.

A1

Product

C1

Layouts

D1

Company

E1

Template

F1

All systems operational
fra41 ms
iad38 ms
sin52 ms

C2

Elsewhere

XLinkedInGitHubYouTube

E2

Move across the wordmark168 blocks · iso 30°

One email a month. Only what shipped.

A1

Product

C1

Layouts

D1

Company

E1

Template

F1

All systems operational
fra41 ms
iad38 ms
sin52 ms

C2

Elsewhere

XLinkedInGitHubYouTube

E2

Move across the wordmark168 blocks · iso 30°

One email a month. Only what shipped.

A1

Product

C1

Layouts

D1

Company

E1

Template

F1

All systems operational
fra41 ms
iad38 ms
sin52 ms

C2

Elsewhere

XLinkedInGitHubYouTube

E2

Move across the wordmark168 blocks · iso 30°

Create a free website with Framer, the website builder loved by startups, designers and agencies.