Durable Pipelines: Why One Retry Policy Is Not Enough
When a pipeline chains slow stages, retrying from the start discards prior work. Stage boundaries need to be checkpoints, not just function calls.
When a pipeline chains five slow stages, retrying from the start on a failure at stage five throws away four stages of work. That cost is negligible when stages are fast. When any stage involves GPU generation, large document analysis, or third-party API calls that take minutes each, it becomes real.
Adding retry logic is only half of it. The more important question is where the retry boundary sits.
The problem with treating a pipeline as one job
The straightforward implementation runs every stage in sequence inside a single job, with one retry policy at the top. If stage five of seven fails, the retry re-runs from stage one. That works when stages are cheap. It falls apart when any single stage is expensive enough that re-running it on every downstream failure is unacceptable.
Retrying the pipeline conflates two different things: retrying the stage that failed, and re-running prior stages that already succeeded. The first is necessary. The second is waste.
A pipeline that loses the work of four successful stages to a failure in the fifth has a checkpoint problem, not a retry problem.
Stage boundaries as durable checkpoints
The fix is one principle applied consistently: each stage’s output lands in durable storage before the next stage begins.
Durable storage means it survives process crashes, worker restarts, and transient third-party outages. For media pipelines this is object storage: each generated clip, assembled file, or processed artefact gets written to a blob store before the next stage starts. For data pipelines it is a database row or file: extracted content, classified records, enriched data all persisted before the next stage consumes them. The handoff between stages is a storage read, not a function argument.
Once that is true, retrying a failed stage reads its input from the previous stage’s output and writes its own output to durable storage. Nothing prior is at risk.
State separate from artefacts
Keeping job state in the same place as the artefacts makes both harder to work with. A job record with a status, a stage marker, and an error field costs nothing to maintain and makes every failure traceable without opening object storage and guessing what each file represents.
The practical split: metadata (what the pipeline is doing, which stage it reached, what happened) goes into a relational row or document. Artefacts (video files, generated images, extracted content) go into object storage or a blob store. Inspecting a run means reading one small record. Replaying a stage from a checkpoint means reading one artefact. Reasoning about the pipeline never requires moving large files.
Per-stage retry policies
Not every stage has the same failure mode or the same safe retry behaviour.
A generation stage that calls an external model API is a good candidate for exponential backoff with a cap, since most failures are transient rate limits or brief service interruptions. An assembly stage that stitches together files already in object storage is usually safe to retry immediately and repeatedly. A publish stage that posts to an external platform or sends an email is not idempotent: a naive retry sends the email twice. That stage needs a guard that checks whether the action already completed before attempting it again.
A single retry policy at the pipeline level cannot express these differences. Per-stage queues with their own retry configuration can.
What the architecture looks like in practice
Each stage is an independently queued task. Stage one completes, writes its output to storage, and enqueues stage two. Stage two reads from storage, does its work, writes its own output, and enqueues stage three. A failure at any point leaves all prior output intact. The failed stage gets re-queued with its own policy and re-runs from its durable input.
The pipeline’s overall state lives in a job record: stage reached, timestamp, error if any. Progress is visible without reading artefacts. Resuming from a failed stage is re-queuing that stage alone, not restarting the pipeline.
This is the architecture behind the video pipeline we built for AwesomeGene, where clip generation, assembly, narration, subtitle timing, and publish are each separately queued tasks. A transient API error at clip generation retries that stage with backoff. A failure at final publish retries with an idempotency check. Nothing prior reruns.
The same checkpoint discipline applies to shorter data pipelines. The six-stage document processing pipeline described in automating document processing is built on the same logic: a failed extraction step does not discard the normalisation work that preceded it.
When this level of design is not worth it
Short pipelines with fast, cheap, idempotent stages do not need this. If the entire pipeline takes seconds and is free to rerun, re-running from the start on failure costs nothing and per-stage queues are pure overhead.
The threshold is roughly: when any single stage takes more than a minute, costs real money to rerun, or calls a third-party service with rate limits, the pipeline is worth building with checkpoint architecture. When none of those are true, sequential execution is fine.
We built the AwesomeGene pipeline with this architecture. See the case study, or read about how we approach workflow automation.