Skip to content
Apsan Works

Email Outreach at Scale: Why the State Machine Is the Product

Outreach at volume is a state machine problem. Tracking each lead's position, enforcing send limits, and routing replies correctly are the real work.

5 min read

A three-step email sequence is easy to build. Send something on day one, something else on day four, a follow-up on day ten. At one lead, that’s a scheduler and a template renderer.

At a thousand leads, all in different states, receiving different variants, with rate limits to respect and replies to detect, the same problem is a state machine with real engineering behind it.

What the state actually is

Each lead in a running sequence needs its own durable record: which step it is on, when it last received something, which variant it received, whether it has replied, whether it has bounced, and when it is eligible for the next action.

A single shared sequence object does not hold this. The sequence defines the steps and the timing. Each lead’s record describes where that specific lead is within those steps, and what is safe to happen next for that lead, independent of every other lead running through the same sequence at the same time.

That separation is the first thing most implementations get wrong. A sequence is a template, and a lead record is an instance. The engineering lives in the instances.

The state has to survive restarts

State that lives in memory is fine for a demo. In production, a process restart, a deployment, or a crash wipes it. Every lead loses its position. Some go back to step one, some never move again. Neither is acceptable.

This is the same principle behind durable pipelines: each stage’s output must exist independently of the process that produced it. The same applies here. Lead state belongs in a database from the first line of production code: which step, which timestamp, which outcome. Not in a map in memory that will disappear.

A database record also makes the system inspectable. When something goes wrong for a specific lead, reading one row tells you exactly what happened and when. Debugging from memory state or inferred logs is a different problem entirely.

Concurrency means per-lead scheduling

A sequence with a thousand active leads does not run in a loop. It runs as a queue of scheduled jobs, one per lead per pending step, each firing at the right time for that lead.

The alternative, a cron job that iterates through all leads and decides what each one should receive, has a ceiling. It is harder to parallelize, slower at scale, and harder to reason about when rate limits need to be applied per account rather than per run.

Per-lead scheduling means: when a lead reaches step two, a job for step three is scheduled at the right offset. That job runs, does its work, and schedules step four. The sequence advances one lead at a time, driven by the job queue rather than a periodic sweep.

Send limits are constraints the scheduler enforces

Burning a sending account is expensive. Most email infrastructure providers enforce their own rate limits, but those are hard stops. The soft limits, daily send caps per account, weekday-only delivery, restricted sending hours, cooldown after a transient failure, are business rules the scheduler needs to enforce before they become infrastructure errors.

This means the scheduler cannot blindly fire a job when its timestamp arrives. Before sending, it checks: is the account under its daily cap? Is this a permitted sending hour? Has enough time passed since the last failure for this lead? If any check fails, the job is requeued for the next eligible window, not abandoned.

None of this is complex logic. It is bookkeeping, and the bookkeeping has to be correct every time because the cost of a mistake is not a test failure; it is a blacklisted sending account.

Reply detection is a state transition

When a lead replies, the sequence stops for that lead. That sounds obvious. The implementation is specific: the reply detection has to run on an interval, compare incoming mail against active leads, and write a state transition to each matching lead record before the next scheduled send fires.

A race condition here sends the next step to someone who already replied. That is not a test case to avoid; it is a likely failure mode in any implementation that checks for replies lazily.

The reply transition also determines what happens next. Some sequences move a replied lead into a different state for manual follow-up. Some trigger a different sequence. Some just stop. The state machine needs to represent those paths, and the transition logic needs to be correct before the sequence goes live with real leads.

What this looks like in practice

The platform we built for internal outreach use, described in the Email Automator case study, stores all lead state in SQLite with an explicit stage marker per lead, per step, per send attempt. The job queue drives sequence progression. Send limits are enforced as pre-send checks. Reply detection runs on a polling interval and writes a final state to the lead record before the next step can fire.

The result is a sequence engine that can run thousands of concurrent leads, restart cleanly, and let you answer the question “what happened to this specific lead and why” by reading one record. That inspectability is part of the product.

When this level of design is worth it

A small sequence running a few dozen leads at low frequency does not need all of this. A simple scheduler and an in-memory state object is fine, as long as everyone understands it will need to be rewritten before it runs at real volume.

The threshold is roughly: when a sequence runs more leads than you can manually audit on failure, or when the sending account has real value that a mistake could damage, the architecture described here pays for itself quickly. Below that, the simpler version is fine and the rewrite cost is known.

We built and operate this kind of system ourselves. See how we approach workflow automation, or read about when outreach automation is worth custom code.

Tell us what is slowing you down

A short conversation is usually enough to tell whether this is a build, an automation, or something you should not do at all. We will tell you which.