Semogram Docs
Data Endpoints

Data Endpoints

Understand endpoint roles, data direction, plugin connections and how projects use workspace data

A data endpoint is a saved configuration that tells Semogram which data to access, through which installed capability, and for what purpose. You give it a name, select its role and connection, and specify a target such as a database table, object path, API request or ontology store.

For example, instead of putting database credentials and table settings into every pipeline, you configure one endpoint for your orders table. Pipelines reference that endpoint when they need those orders. The records remain in the source until an operation reads them; saving the endpoint does not ingest them.

An endpoint is not necessarily a web URL. A database endpoint can be identified by a table selector; a storage endpoint by a path. The HTTP Dataset plugin is one case where the target includes a request URL.

The three endpoint roles and data direction

The endpoint creation form offers Ingress, Egress and Materialization. These describe what Semogram does with the selected data. The stored role names are source, destination and ontology_store respectively.

Role in the UIData directionWhat it doesInstalled capability
Ingress (source)External data → SemogramReads records for a pipeline or supported read operationRead
Egress (destination)Semogram → external targetWrites output using the connector's supported write behaviorWrite
Materialization (ontology_store)Ontology facts → configured fact storeStores structured ontology facts for supported query operationsFact materializer
Data direction for the three endpoint rolesIngress reads external records into Semogram. Egress writes Semogram output to an external target. Materialization writes ontology facts to a configured fact store, where supported ontology queries can read them back. Arrows show data movement during execution, not automatic synchronization.Ingress · readExternal dataTable, file or APISource endpointSemogramPipeline / supported readEgress · writeSemogramPipeline outputDestination endpointExternal targetTable or object pathMaterialization · store factsSemogramOntology factsOntology-store endpointFact storePostgres / Apache Jena
Arrows show data direction when an operation runs; saving an endpoint does not move data

Ingress: bring records into Semogram

An ingress endpoint selects data to read. Examples include a Postgres orders table, MongoDB customer collection, CSV files under an S3 path, or shipment records returned by an API.

The endpoint specifies the source; a pipeline or supported read operation performs the read. A source node can use full or incremental execution only where the connector implements the required behavior. Reading records does not grant permission to change them.

Egress: send output to a target

An egress endpoint selects where a supported writer sends data. For example, a pipeline can write cleaned orders into a separate Postgres table or export output files under an S3 path.

The installed writer determines the available formats and write modes. Append adds records, upsert updates/inserts using keys where supported, and overwrite can replace existing target data. Those modes are not interchangeable and are not available on every connector. Verify the target, permissions and applicable write policy before execution.

A connection with both read and write capabilities can support separate source and destination endpoints. One saved endpoint binds to one installed capability; selecting a source does not make it a destination automatically.

Materialization: store ontology facts

A materialization endpoint selects a fact store, such as Postgres relational ontology storage or an Apache Jena named graph. A project ontology binding uses it to materialize entities, relationships and facts through the supported runtime.

Materialization writes facts into the configured store. Supported ontology or SPARQL queries can subsequently read them. A store endpoint is therefore different from a raw source table: its target configures fact storage, not an arbitrary collection of business rows. Available query behavior depends on the store implementation.

What about the other role names in the API?

The shared endpoint schema also contains event_source, event_sink, api, file_store, search_index, graph_store and other. The current manual creation flow exposes the three roles above and maps them to read, write and fact-materializer capabilities. The broader schema does not establish a working creation/execution flow for every additional role. Use operations the installed capability and runtime actually support.

Plugin, installation, capability and endpoint

These are separate resources with different jobs:

ResourceWhat it containsPostgres example
Plugin packageCode and contracts implementing supported operationsThe Postgres connector package
Plugin installationConfiguration for a particular connectionHost, database, user, protected password and TLS settings
Installed capabilityOne operation supplied by that installationRead, write or ontology fact materialization
Data endpointA chosen name, role, capability reference and concrete targetA source named orders targeting a particular table

The installation answers “How do we connect?” The endpoint answers “Which data do we use, and how?”

For a Postgres source, the installation contains database connection settings; the endpoint target contains the schema and table. For S3, the installation selects bucket/region; the target selects a path. For HTTP Dataset, the target describes the request and response extraction. News/Web places its source URL on the installation and uses an articles target on the endpoint. Target settings follow each plugin's contract rather than a universal table/path form.

Credentials stay on protected installation settings or supported secret references. Projects reference the endpoint instead of copying credentials into every pipeline.

Role, contract and target are different

SettingQuestion it answersExample
RoleWhat is this endpoint used for?Ingress/source
Contract kindWhich access pattern does the capability implement?source_stream or source_query
TargetWhat exact data should that operation access?Schema public, table orders

A source_stream contract selects source records through the connector's stream interface. The word “stream” does not by itself mean a live event feed. A source_query contract supplies a connector-specific query. A write contract supplies destination settings. An ontology_fact_store contract supplies fact-store settings.

The chosen contract must exist on the installed capability, and the target must satisfy that contract's required fields and types. Choosing a contract name cannot add an operation the plugin does not implement.

Endpoints belong to the workspace

Plugin installations and endpoints are workspace resources. Projects in that workspace can reference them, subject to access rules. Two projects reading the same table can reuse one source endpoint rather than creating two connections or copying configuration.

An endpoint references a specific installed capability ID. It is not bound merely to a package name or to a project. The selected installation must belong to the same workspace. A project still provides context for its pipelines and ontology bindings.

The diagram below shows configuration references, not the direction of a data transfer. Both project pipelines use the same saved endpoint, which selects a table through the installed Postgres read capability.

Configuration references: one endpoint shared by two projectsThe Postgres database contains public.orders. A workspace Postgres installation holds connection settings and provides a read capability. The orders data endpoint selects that table through the capability. Project A's pipeline and Project B's pipeline both reference the same endpoint.WorkspacePostgres databasepublic.ordersActual table and recordsPostgres installationConnection + credentialsInstalled read capabilityData endpointordersTable: public.ordersProject AOrders pipelineProject BRevenue pipelinereference
Arrows show configuration references; both projects use the same workspace endpoint

Changing a shared endpoint's target or connection can affect every consumer. Inspect its pipelines, queries and ontology bindings before replacing the capability, renaming it or deleting it. Workspace scope does not grant everyone every read/write operation.

Example 1: read an orders table

Suppose your test database contains this table:

order_idtotal
1120.00
275.50

The table is named orders and belongs to the database schema named public. Its database-qualified name is therefore public.orders. Neither name is supplied by Semogram; use an existing table or create a test fixture in your database.

You choose orders as the endpoint name and sales as its grouping namespace in Semogram. These labels need not match the database schema. For clarity, this example deliberately keeps the endpoint name short.

  1. Open Plugins → Explore in the workspace, choose Postgres's read capability and install it with your database host, port, database, user, password and TLS settings.
  2. Check that connection using a database user permitted to read the table.
  3. Open Data Endpoints → New data endpoint. Describe this source to the assistant, or choose Edit manually.
  4. Choose Ingress, the installed Postgres read capability and source_stream. Set name orders, namespace sales, and this target:

Open workspace Data Endpoints → New data endpoint → Edit manually, choose the direction and installed capability described in this example, then fill the target fields. Enter values in the labeled controls rather than pasting the whole JSON object.

UI fieldExample value
Schemapublic
Tableorders

Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Review the endpoint name, direction, capability and selected target before saving.

In the platform assistant or your connected MCP assistant, ask:

Assistant prompt
Create the endpoint described on this page using these settings:
name: <ENDPOINT_NAME_FROM_THIS_EXAMPLE>
namespace: <ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>
role: source
contractKind: source_stream
target / schema: public
target / table: orders
pluginCapabilityInstallationId: <INSTALLED_CAPABILITY_UUID>
Use the actual installed capability and the endpoint name/namespace selected in this example. Show the proposed direction, connection and target before saving. Keep credentials on the installation.

Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.

Use a workspace API key with endpoints:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.

HTTP API request
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints" \
  --header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
  --header "Idempotency-Key: <UNIQUE_KEY_FOR_THIS_ENDPOINT>" \
  --header "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
  "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
  "role": "source",
  "contractKind": "source_stream",
  "target": {
    "schema": "public",
    "table": "orders"
  },
  "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>"
}
JSON

Call source_create with the arguments below through an authenticated workspace MCP connection. Replace the name/namespace placeholders with the labels chosen in this example and use the actual installed capability UUID. Set the role/contract to the direction described here; the workspace is resolved from the connection. This configures an endpoint and does not execute a read or write.

MCP request
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "source_create",
    "arguments": {
      "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
      "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
      "role": "source",
      "contractKind": "source_stream",
      "target": {
        "schema": "public",
        "table": "orders"
      },
      "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>",
      "idempotencyKey": "<UNIQUE_KEY_FOR_THIS_ENDPOINT>"
    }
  }
}
  1. Add a useful description and set the key hint to order_id after checking uniqueness and nulls. Save the endpoint.
  2. Read a small supported preview, or run a small saved pipeline with a source node referencing this endpoint. Compare IDs 1 and 2 and their totals with the database.

The endpoint now identifies the table for consumers. It has not created the table or started recurring ingestion. The Postgres setup example includes complete installation settings and SQL for a small fixture if you want to perform this exercise.

Example 2: write cleaned orders elsewhere

Suppose you want the pipeline's cleaned results in a separate table named clean_orders in the database schema analytics. Prepare that test destination and the required write privileges before execution.

Install/select Postgres's write capability. Create another endpoint named clean_orders with namespace sales, Egress role and write contract. Its target can be:

Open workspace Data Endpoints → New data endpoint → Edit manually, choose the direction and installed capability described in this example, then fill the target fields. Enter values in the labeled controls rather than pasting the whole JSON object.

UI fieldExample value
Schemaanalytics
Tableclean_orders
Write modeappend

Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Review the endpoint name, direction, capability and selected target before saving.

In the platform assistant or your connected MCP assistant, ask:

Assistant prompt
Create the endpoint described on this page using these settings:
name: <ENDPOINT_NAME_FROM_THIS_EXAMPLE>
namespace: <ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>
role: destination
contractKind: write
target / schema: analytics
target / table: clean_orders
target / writeMode: append
pluginCapabilityInstallationId: <INSTALLED_CAPABILITY_UUID>
Use the actual installed capability and the endpoint name/namespace selected in this example. Show the proposed direction, connection and target before saving. Keep credentials on the installation.

Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.

Use a workspace API key with endpoints:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.

HTTP API request
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints" \
  --header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
  --header "Idempotency-Key: <UNIQUE_KEY_FOR_THIS_ENDPOINT>" \
  --header "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
  "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
  "role": "destination",
  "contractKind": "write",
  "target": {
    "schema": "analytics",
    "table": "clean_orders",
    "writeMode": "append"
  },
  "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>"
}
JSON

Call source_create with the arguments below through an authenticated workspace MCP connection. Replace the name/namespace placeholders with the labels chosen in this example and use the actual installed capability UUID. Set the role/contract to the direction described here; the workspace is resolved from the connection. This configures an endpoint and does not execute a read or write.

MCP request
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "source_create",
    "arguments": {
      "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
      "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
      "role": "destination",
      "contractKind": "write",
      "target": {
        "schema": "analytics",
        "table": "clean_orders",
        "writeMode": "append"
      },
      "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>",
      "idempotencyKey": "<UNIQUE_KEY_FOR_THIS_ENDPOINT>"
    }
  }
}

A pipeline can read through the orders source, transform the records, and write through the clean_orders destination. The endpoints identify input and output; the pipeline defines the processing between them.

For the first test, use two known rows and a dedicated target. Run the saved pipeline and inspect the receiving table and run output. Append can produce duplicates when repeated unless the implemented flow provides the appropriate idempotency. Use an upsert mode only when supported and with verified keys.

Example 3: materialize facts into an ontology store

Suppose a project models customers and their orders as ontology entities and relationships. You want to store the resulting facts in Postgres rather than treating them as a raw orders source.

Install/select Postgres's ontology fact-store capability and prepare a database schema with the required materialization privileges. Create an endpoint named facts, namespace ontology, Materialization role and ontology_fact_store contract:

Open workspace Data Endpoints → New data endpoint → Edit manually, choose the direction and installed capability described in this example, then fill the target fields. Enter values in the labeled controls rather than pasting the whole JSON object.

UI fieldExample value
Typepostgres_ontology_fact_store
Schemaontology
Table prefixmain

Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Review the endpoint name, direction, capability and selected target before saving.

In the platform assistant or your connected MCP assistant, ask:

Assistant prompt
Create the endpoint described on this page using these settings:
name: <ENDPOINT_NAME_FROM_THIS_EXAMPLE>
namespace: <ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>
role: ontology_store
contractKind: ontology_fact_store
target / type: postgres_ontology_fact_store
target / schema: ontology
target / tablePrefix: main
pluginCapabilityInstallationId: <INSTALLED_CAPABILITY_UUID>
Use the actual installed capability and the endpoint name/namespace selected in this example. Show the proposed direction, connection and target before saving. Keep credentials on the installation.

Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.

Use a workspace API key with endpoints:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.

HTTP API request
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints" \
  --header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
  --header "Idempotency-Key: <UNIQUE_KEY_FOR_THIS_ENDPOINT>" \
  --header "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
  "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
  "role": "ontology_store",
  "contractKind": "ontology_fact_store",
  "target": {
    "type": "postgres_ontology_fact_store",
    "schema": "ontology",
    "tablePrefix": "main"
  },
  "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>"
}
JSON

Call source_create with the arguments below through an authenticated workspace MCP connection. Replace the name/namespace placeholders with the labels chosen in this example and use the actual installed capability UUID. Set the role/contract to the direction described here; the workspace is resolved from the connection. This configures an endpoint and does not execute a read or write.

MCP request
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "source_create",
    "arguments": {
      "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
      "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
      "role": "ontology_store",
      "contractKind": "ontology_fact_store",
      "target": {
        "type": "postgres_ontology_fact_store",
        "schema": "ontology",
        "tablePrefix": "main"
      },
      "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>",
      "idempotencyKey": "<UNIQUE_KEY_FOR_THIS_ENDPOINT>"
    }
  }
}

Here ontology is the example database schema and main is the storage prefix used by the materializer. The endpoint name facts is your Semogram label. Bind this store in the project's ontology flow, materialize a known entity/fact and query it through the supported ontology path. Inspect its identifier and source evidence.

For Apache Jena, the installation configures the Fuseki dataset and the endpoint target identifies the named graph and entity IRI prefix. Both are store configurations; they are not interchangeable with a generic table-read target.

What else an endpoint records

SettingWhy it is usefulWhat it does not do
Name and namespaceIdentify and group endpointsCreate external tables or database schemas
Description and owner teamExplain the data and who maintains accessGrant permissions
Column and primary-key hintsDescribe expected fields and identityCreate constraints or guarantee uniqueness
FreshnessState expected currencySchedule reads
Trust rankRecord the assessed reliability of the sourceAutomatically verify every record
Accepts/emitsDescribe formats used by the operationMake unsupported formats work

Use discovered schema and actual records to check metadata. Update affected mappings and consumers when an upstream field or type changes.

From configuration to running data

StageWhat happens
InstallConfigure the plugin capability and connection
CreateSave the endpoint's role, contract, target and metadata
Validate/checkValidate selectors and run supported connection probes
Discover/previewInspect supported schema and a bounded sample
ExecuteRun a read, pipeline, write or materialization operation
InspectCompare records, outputs, receipts and evidence
ScheduleConfigure recurring pipeline execution after verification

These steps establish different things. A successful connection check does not prove that every table is readable or writable. A valid target does not prove that its external object exists. A preview does not complete ingestion, and setting freshness does not create a schedule.

Uploaded files follow a managed storage flow: upload and preview the bytes, create the endpoint and pipeline, then execute that pipeline. Selecting a new file version is separate from processing it.

Permissions and plugin-specific operations

You need a Semogram account with workspace access and permission for the intended operations, plus appropriate external-system access. Configuring an endpoint does not grant additional credentials or permissions.

Generic destination writes and durable source mutations are different paths. The current governed source-write, snapshot/recovery and maintenance operations require Iceberg-specific support, along with the applicable workspace settings, endpoint policy and actor grants. Installing another connector's writer does not add those guarantees.

Choose your setup

Data you want to exposeGuide
PostgreSQL table/queryPostgres
MongoDB collectionMongoDB
Objects in your bucketS3 Bucket
Record-oriented HTTP API/datasetHTTP Dataset
Feeds, articles or web sourcesNews/Web
Transactional catalog tableIceberg
Ontology fact storagePostgres store or Apache Jena
File on your computerDataset upload
Your own connector implementationCustom source or custom destination

Transforms, skills and secret providers support processing/configuration but do not expose standalone data endpoints simply because they are plugins. External MCP servers supply tools through a separate integration flow.

For exact fields and operation behavior, use the endpoint reference. For failed checks or unexpected records, use connection and troubleshooting.