Semogram Docs
Data PipelinesRun and operate

Checkpoint recovery

Resume eligible interrupted runs without blindly repeating external effects

Checkpoint recovery continues an eligible failed or canceled pipeline run using its original execution snapshot. It retains the run ID and finished node checkpoints. It is different from launching a new run, which reads the newly selected version and may reread changing sources.

What you need

You need a Semogram account as a workspace owner or administrator, access to the pipeline, and a failed or canceled run with its original snapshot and usable checkpoints. Its earlier Workflow must be stopped before recovery. Run artifacts that have expired cannot support recovery. Older executions with incomplete checkpoint coverage can be rejected.

Inspect external destinations before approval. A connector can commit a write and lose its response before the node checkpoint is saved. Failure or cancellation does not establish that nothing was written.

What is preserved

StateRecovery behavior
Finished nodeIts durable output is restored instead of repeating the node
Partial or skipped checkpointRetained as recorded; recovery is not an instruction to replay every partial result
Unfinished ordinary node claimReviewed and cleared so the node can be attempted again
Materializer with saved progressProgress is preserved for its supported continuation path
Source cursorOnly durably committed progress is retained; inspect connector-specific semantics
External write without a saved completionOutcome must be checked before permitting another attempt
Original version and snapshotRetained; editing today's draft does not modify this run

Files needed for checkpoint restoration must remain available and pass integrity checks. A live external dataset reference is not a historical copy of the source: inspect whether its contents changed since the original attempt.

Recover in the UI

  1. Open the pipeline's Operations page and the affected run's Run recovery page.
  2. Read Durable node checkpoints. Identify finished nodes and unfinished claims.
  3. Confirm the earlier workflow has stopped. Inspect destination rows, object paths or materialized records for unfinished writes.
  4. When requested, check External outcomes reviewed only after establishing that the unfinished operations can safely run again.
  5. Select Resume from checkpoints. Read the acceptance message and follow the same run ID to its final state.
  6. Inspect Recovery history, restored node outputs and the actual destination result.

Recovery is currently a signed-in UI workflow. There is no public API-key recovery action or exposed MCP resume tool. An assistant can help investigate and prepare a review, but should not claim that it resumed the run through an unavailable tool.

Assistant investigation prompt
Investigate this failed order-refresh run without launching another one. Report its original version, completed nodes, unfinished node errors and any available output or connector receipts. Identify which external destination checks I must perform before using Run recovery. Do not infer that a failed write had no effect.

Example: destination acknowledgement was lost

The source returned orders A-100 and A-101. The destination accepted A-100, then the worker stopped before recording completion. The run now reports an unfinished egress node.

Read the destination using the same order IDs. If A-100 exists, determine whether the installed write capability safely upserts it or would append a duplicate. A checkpoint acknowledgement does not change that connector behavior. If replay is unsafe, repair or reconcile the destination through its supported operation before approval. If you cannot establish the outcome, keep the run stopped and retain the evidence.

After safe recovery, compare both IDs and their values, not just the Completed label. Record the provider receipt or reconciliation evidence alongside the investigation.

When to launch a new run

Use a new run for a deliberately changed version or a new business invocation. If recovery is refused because the snapshot/checkpoint artifacts are unavailable, first inspect the old effects, then decide whether a new run can safely execute. It will not automatically inherit the old checkpoint protection.

Stopping a queued/running run is available in Run recovery. It preserves history and finished effects; it does not roll back writes. Read back the stopped state before preparing recovery.