Overview
Build, version and run project workflows that read endpoints, transform records, materialize facts and write results
A data pipeline describes work Semogram should perform on data: where to read it, how to process it and where to put the result. For example, read orders from a database, select the useful fields and write a reporting table. Other pipelines turn records into ontology facts, preserve evidence assertions, enrich rows with an external tool or run a forecaster.
The pipeline is the reusable definition. A run is one execution of a saved version of that definition. Saving a draft or a version does not read records or write output.
What belongs where
| Resource | Scope | What it supplies |
|---|---|---|
| Plugin installation | Workspace | An installed capability and protected connection configuration |
| Data endpoint | Workspace | A named target and direction: read, write or store ontology facts |
| Pipeline | Project | Nodes, their configuration and the connections between them |
| Pipeline version | Pipeline | A saved graph revision |
| Run | Project and pipeline | Execution state, step results and a snapshot of the selected version |
| Schedule | Pipeline | Timing, time zone and overlap/missed-time behavior |
Two projects can reference the same workspace endpoint. They still own separate pipelines and run histories. A node selects an endpoint or installed capability; database passwords belong on the installation, not in the graph or an assistant prompt.
The parts of a pipeline
A node performs a task. A port names a node's input or output and declares its data shape. An edge connects an output port to an input port and establishes a dependency. Intermediate outputs are data handles/pools passed to downstream nodes; drawing an edge does not create an external database table.
| Task | Node or combination | Result |
|---|---|---|
| Read external data | Ingress | Records from a workspace source endpoint |
| Select, clean, join or classify | Transform | An installed transform's output |
| Inspect or stage an intermediate result | Asset | A staged data artifact; optional configured persistence |
| Enrich each row with a read-only remote tool | MCP tool | Original rows plus a result column |
| Build business entities and relationships | Ontology mapping → Ontology materializer | Mapped facts persisted to a supported ontology store |
| Preserve claims and evidence | Assertion mapping → Assertion materializer | Assertions persisted with their evidence links and conflict policy |
| Make a prediction | Prediction | A prediction from an existing forecaster and its evidence query |
| Write records to another system | Egress | Data committed through a destination endpoint |
A Prediction node queues separate non-blocking work; check its prediction job as well as the pipeline run. A mapping creates a fact set; materialization persists it. Query and Skill nodes provide context/guidance and currently have no default standalone execution step. Every node kind and its limits are covered in the node reference.
Build, check, save, run
- Describe the task and its expected result in Pipeline Studio. The assistant proposes a graph; review the endpoints, transformations and output before accepting it.
- Inspect node fields, ports and dependencies. Use the manual controls where you need precise configuration.
- Validate the draft. Configuration preflight checks graph structure and executable bindings. It does not execute reads, models or external writes.
- Save a version and inspect which version is active.
- Run a small fixture and check the actual result. A queued response means accepted, not completed.
- Once the output is correct, configure a schedule if recurring execution is needed.
You need a Semogram account with workspace membership, a project in that workspace and permission to create and run pipelines. Reading or writing also requires access to the selected endpoints and external systems. Creating a pipeline does not grant those permissions.
Build toward your result
| Goal | Guide |
|---|---|
| Understand the complete flow with two known rows | First pipeline |
| Project and rename fields | Transform records |
| Combine primary and reference records | Look up reference data |
| Write result records | Write records |
| Persist business entities | Build ontology facts |
| Persist claims with evidence | Build assertions |
| Classify text with AI | Classify records |
| Add information from a remote MCP server | Enrich with MCP |
| Include an existing forecaster | Run predictions |
| Use a capability your team built | Custom plugin pipeline |
Each example states its prerequisites, connection settings, fixture, graph and verification steps. Programmatic examples separate HTTP API requests from MCP tool calls; an API tab appears only for an implemented public action.
Operate the pipeline
Use Authoring for authoring and version history, Run and operate for validation, runs, schedules, monitoring and checkpoint recovery, and Reference for the complete graph document and public interfaces.
Execution guarantees depend on the selected capability. In particular, append is not automatically deduplicated, partial writes can survive a failed run, and cancellation does not roll back completed external effects. Inspect the run and the destination before repeating a write.