Semogram Docs
Data PipelinesRun and operate

Incremental reads and checkpoints

Understand load modes, cursors, retry boundaries and replay before scheduling

An ingress node can request Full load or Incremental. The connector's implemented contract determines whether it can honor incremental extraction, pagination and checkpoints. A mode field does not add incremental support to a connector that lacks it.

Choose a mode

ModeUseVerify
FullSmall snapshots and the first fixtureIntended dataset bounds and whether the destination repeats rows
IncrementalA connector with supported cursor/watermark behaviorInitial cursor, ordering, checkpoint advancement and late/changed records

Source checkpoints track extraction progress. They are not an automatic transaction spanning every downstream write. Named intermediate outputs and run/node checkpoints support execution recovery; they are distinct from the source's logical cursor and the destination's mutation receipt.

Configure and test

In Studio, select the ingress endpoint, inspect its loaded contract and set Load mode. Use the supported flow controls for batch size, retry and concurrency. For programmatic graph authoring, the ingress section can include:

Ingress definition fragment
{
  "dataEndpointId": "<ENDPOINT_UUID>",
  "mode": "incremental",
  "strategy": "parser",
  "output": "orders",
  "flow": {
    "checkpoint": {
      "enabled": true,
      "resetPolicy": "manual"
    },
    "batch": {
      "size": 100
    },
    "retry": {
      "maxAttempts": 3,
      "backoff": "exponential"
    }
  },
  "concurrency": 1
}

This fragment belongs under a full node's ingress property; it is not a standalone API request. Connector-specific supported settings still govern execution.

  1. Begin with two records and a known cursor/watermark.
  2. Execute once and inspect record IDs, cursor/checkpoint and output.
  3. Add one new record, run again and check what the connector returns.
  4. Change an earlier record and test whether updates/late arrivals are included.
  5. Simulate a recoverable failure in a controlled test and inspect replay effects before scheduling.

Replay and resets

A reset can reread old records. Establish whether the destination appends, replaces or upserts and whether the connector has durable deduplication/receipts for that path. Do not infer exactly-once behavior from checkpointing or API idempotency alone.

Changing the endpoint target, source order, checkpoint or write mode changes replay meaning. Record the old version/run state, inspect existing output and use the supported reset/recovery controls. Never reset a production cursor merely to make a test appear successful.