Semogram Docs
Data Pipelines

Overview

Build, version and run project workflows that read endpoints, transform records, materialize facts and write results

A data pipeline describes work Semogram should perform on data: where to read it, how to process it and where to put the result. For example, read orders from a database, select the useful fields and write a reporting table. Other pipelines turn records into ontology facts, preserve evidence assertions, enrich rows with an external tool or run a forecaster.

The pipeline is the reusable definition. A run is one execution of a saved version of that definition. Saving a draft or a version does not read records or write output.

Arrows show data movement during a real run. The source and destination are workspace resources; the pipeline processing belongs to a project.Source endpointRead external dataProcess recordsTransform, map or enrichDestination endpointWrite or materializeVerified resultInspect rows and evidence
Saving the graph configures this flow; a run performs the work

What belongs where

ResourceScopeWhat it supplies
Plugin installationWorkspaceAn installed capability and protected connection configuration
Data endpointWorkspaceA named target and direction: read, write or store ontology facts
PipelineProjectNodes, their configuration and the connections between them
Pipeline versionPipelineA saved graph revision
RunProject and pipelineExecution state, step results and a snapshot of the selected version
SchedulePipelineTiming, time zone and overlap/missed-time behavior

Two projects can reference the same workspace endpoint. They still own separate pipelines and run histories. A node selects an endpoint or installed capability; database passwords belong on the installation, not in the graph or an assistant prompt.

The parts of a pipeline

A node performs a task. A port names a node's input or output and declares its data shape. An edge connects an output port to an input port and establishes a dependency. Intermediate outputs are data handles/pools passed to downstream nodes; drawing an edge does not create an external database table.

TaskNode or combinationResult
Read external dataIngressRecords from a workspace source endpoint
Select, clean, join or classifyTransformAn installed transform's output
Inspect or stage an intermediate resultAssetA staged data artifact; optional configured persistence
Enrich each row with a read-only remote toolMCP toolOriginal rows plus a result column
Build business entities and relationshipsOntology mapping → Ontology materializerMapped facts persisted to a supported ontology store
Preserve claims and evidenceAssertion mapping → Assertion materializerAssertions persisted with their evidence links and conflict policy
Make a predictionPredictionA prediction from an existing forecaster and its evidence query
Write records to another systemEgressData committed through a destination endpoint

A Prediction node queues separate non-blocking work; check its prediction job as well as the pipeline run. A mapping creates a fact set; materialization persists it. Query and Skill nodes provide context/guidance and currently have no default standalone execution step. Every node kind and its limits are covered in the node reference.

Build, check, save, run

The first four stages prepare the definition. Running executes a saved snapshot. Verification checks real output, which configuration validation cannot prove.1. Describe and reviewAssistant proposes a graph2. Edit the draftInspect nodes and bindings3. ValidateConfiguration preflight only4. Save a versionRecord the graph revision5. RunExecute a saved snapshot6. VerifyCheck actual outputs
Drafts, versions and runs are separate records; validate before saving and inspect results after execution
  1. Describe the task and its expected result in Pipeline Studio. The assistant proposes a graph; review the endpoints, transformations and output before accepting it.
  2. Inspect node fields, ports and dependencies. Use the manual controls where you need precise configuration.
  3. Validate the draft. Configuration preflight checks graph structure and executable bindings. It does not execute reads, models or external writes.
  4. Save a version and inspect which version is active.
  5. Run a small fixture and check the actual result. A queued response means accepted, not completed.
  6. Once the output is correct, configure a schedule if recurring execution is needed.

You need a Semogram account with workspace membership, a project in that workspace and permission to create and run pipelines. Reading or writing also requires access to the selected endpoints and external systems. Creating a pipeline does not grant those permissions.

Build toward your result

GoalGuide
Understand the complete flow with two known rowsFirst pipeline
Project and rename fieldsTransform records
Combine primary and reference recordsLook up reference data
Write result recordsWrite records
Persist business entitiesBuild ontology facts
Persist claims with evidenceBuild assertions
Classify text with AIClassify records
Add information from a remote MCP serverEnrich with MCP
Include an existing forecasterRun predictions
Use a capability your team builtCustom plugin pipeline

Each example states its prerequisites, connection settings, fixture, graph and verification steps. Programmatic examples separate HTTP API requests from MCP tool calls; an API tab appears only for an implemented public action.

Operate the pipeline

Use Authoring for authoring and version history, Run and operate for validation, runs, schedules, monitoring and checkpoint recovery, and Reference for the complete graph document and public interfaces.

Execution guarantees depend on the selected capability. In particular, append is not automatically deduplicated, partial writes can survive a failed run, and cancellation does not roll back completed external effects. Inspect the run and the destination before repeating a write.