Semogram Docs
Data EndpointsSetup guides

HTTP Dataset

Shipment API records endpoint setup and verification

What this endpoint does

An HTTP Dataset endpoint describes a request to an API or downloadable dataset and how to extract records from its response. It can read supported JSON, CSV and ZIP datasets, follow configured pagination and apply time windows for recurring pipeline reads.

Like all data endpoints, it belongs to a workspace and can be referenced by projects in that workspace, subject to permissions. Saving it configures access; it does not start ingestion.

Directions and supported operations

This is an Ingress/read endpoint: HTTP service → Semogram. Select the installed read capability and source_stream contract. JSON, CSV and supported ZIP responses are mapped into records; pagination, field coercion and time windows follow the target contract.

A configured HTTP POST can fetch records from a read API. It does not turn this connector into an Egress writer. This package does not implement a generic destination or ontology fact store. To send records to a service, install a separate custom write capability whose runtime implements that service's write API. Source API credentials and rate limits remain applicable to every read.

Plugin used

This endpoint uses the HTTP Dataset read capability. Installation steps for this example are included below. The HTTP Dataset installation guide provides optional further detail. Connection details and credentials stay on the installation; the endpoint selects a particular target through that installed capability.

What you need

  • A Semogram account with workspace access and permission to manage endpoints
  • An existing matching installation, or the connection details to create one using the steps below
  • Have the service URL, HTTP method, authentication requirements, a sample response and the API’s pagination rules. Store authentication on the installation and permit the service domain where an allowlist is configured

Example setup

This fictional shipment API returns records in data and the next-page cursor in nextCursor. Each record contains shipment_id. https://api.example.com/shipments is a placeholder, not a working service. We choose shipments as the endpoint name and logistics as its namespace. Replace the URL, field paths and pagination settings with those from your API.

Install and configure the plugin

  1. Open Plugins in this workspace and select Explore.
  2. Find HTTP Dataset, inspect its publisher/version and select its read capability.
  3. Name the installation and fill its connection settings using your actual external-system details.
  4. Save and run the supported connection check. Fix any reported error before creating the endpoint.

Example installation configuration:

Open workspace Plugins → Explore, choose the matching capability and fill its installation settings. Enter values in the labeled controls rather than pasting the whole JSON object.

UI fieldExample value
Auth → Typebearer
Auth → TokenYOUR_API_TOKEN
Allowed domainsAdd YOUR_API_HOST
Max records100
Timeout ms30000

Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Enter credentials in the protected fields and review the selected installation before saving.

In the platform assistant or your connected MCP assistant, ask:

Assistant prompt
Install the plugin described on this page in this workspace. Discover its catalog entry, select the matching capability and propose the installation using the connection settings shown here. Ask me to enter credentials in protected installation fields. Show the selected plugin/version, capability and non-secret settings before saving.

Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.

Use plugin_catalog_list / plugin_catalog_get to obtain the discovery ID and matching capability class (reads, writes or factStores). Call plugin_installation_create with the arguments below through an authenticated MCP connection. The workspace comes from that connection. Enter credentials through an authorized protected configuration path; do not send real secrets as conversational prompt text.

MCP request
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "plugin_installation_create",
    "arguments": {
      "capabilityClass": "<MATCHING_CAPABILITY_CLASS>",
      "discoveryId": "<DISCOVERY_ID_FROM_CATALOG>",
      "name": "<INSTALLATION_NAME>",
      "config": {
        "auth": {
          "type": "bearer",
          "token": "YOUR_API_TOKEN"
        },
        "allowedDomains": [
          "YOUR_API_HOST"
        ],
        "maxRecords": 100,
        "timeoutMs": 30000
      },
      "idempotencyKey": "<UNIQUE_KEY_FOR_THIS_INSTALLATION>"
    }
  }
}

This operation has no standalone public /api/v1 plugin-installation/catalog route in the current implementation. Use the UI or MCP methods shown here.

This operation has no standalone public /api/v1 plugin-installation/catalog route in the current implementation. Use the UI or MCP methods shown here.

Replace every YOUR_… value. Enter credentials only in protected installation configuration. The configuration above is separate from the endpoint target below; selecting a target does not create or authenticate the connection.

Identify the API response

Use your API's own documented URL and confirm that its response matches your configured paths. The fictional response used in this example is:

Example API response
{"data":[{"shipment_id":"shipment_1","status":"delivered"}],"nextCursor":"page_2"}

The next request sends cursor=page_2. The final response must indicate there is no next page using the connector-supported cursor behavior. If your API differs, change recordsPath, responsePath and the request parameter. For a public API, omit authentication; for another auth method use the installed contract's supported fields. This page does not provision an API service.

Use the assistant

Open Data Endpoints → New data endpoint in the workspace. Describe your actual target and choose a name for the endpoint. For the example above, you could ask:

Create a source endpoint named shipments in namespace logistics.
Use our installed HTTP Dataset read capability with contract source_stream
and set these target fields:
Request / Url: https://api.example.com/shipments
Request / Method: GET
Response / Format: json
Response / Records path: data
Pagination / Type: cursor
Pagination / Parameter: cursor
Pagination / Response path: nextCursor
Pagination / Max pages: 10
Primary keys: shipment_id
Review the proposed connection and target before saving.

Replace example values with your own. Give the assistant the installed connection reference; keep credentials in protected installation settings.

Configure manually

Choose Edit manually, select the matching installed capability and configure:

FieldValue
Nameshipments
Namespacelogistics
Rolesource
Contractsource_stream

Open workspace Data Endpoints → New data endpoint → Edit manually, choose the direction and installed capability described in this example, then fill the target fields. Enter values in the labeled controls rather than pasting the whole JSON object.

UI fieldExample value
Request → Urlhttps://api.example.com/shipments
Request → MethodGET
Response → Formatjson
Response → Records pathdata
Pagination → Typecursor
Pagination → Parametercursor
Pagination → Response pathnextCursor
Pagination → Max pages10
Primary keysAdd shipment_id

Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Review the endpoint name, direction, capability and selected target before saving.

In the platform assistant or your connected MCP assistant, ask:

Assistant prompt
Create the endpoint described on this page using these settings:
name: <ENDPOINT_NAME_FROM_THIS_EXAMPLE>
namespace: <ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>
role: source
contractKind: source_stream
target / request / url: https://api.example.com/shipments
target / request / method: GET
target / response / format: json
target / response / recordsPath: data
target / pagination / type: cursor
target / pagination / parameter: cursor
target / pagination / responsePath: nextCursor
target / pagination / maxPages: 10
target / primaryKeys: shipment_id
pluginCapabilityInstallationId: <INSTALLED_CAPABILITY_UUID>
Use the actual installed capability and the endpoint name/namespace selected in this example. Show the proposed direction, connection and target before saving. Keep credentials on the installation.

Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.

Use a workspace API key with endpoints:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.

HTTP API request
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints" \
  --header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
  --header "Idempotency-Key: <UNIQUE_KEY_FOR_THIS_ENDPOINT>" \
  --header "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
  "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
  "role": "source",
  "contractKind": "source_stream",
  "target": {
    "request": {
      "url": "https://api.example.com/shipments",
      "method": "GET"
    },
    "response": {
      "format": "json",
      "recordsPath": "data"
    },
    "pagination": {
      "type": "cursor",
      "parameter": "cursor",
      "responsePath": "nextCursor",
      "maxPages": 10
    },
    "primaryKeys": [
      "shipment_id"
    ]
  },
  "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>"
}
JSON

Call source_create with the arguments below through an authenticated workspace MCP connection. Replace the name/namespace placeholders with the labels chosen in this example and use the actual installed capability UUID. Set the role/contract to the direction described here; the workspace is resolved from the connection. This configures an endpoint and does not execute a read or write.

MCP request
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "source_create",
    "arguments": {
      "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
      "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
      "role": "source",
      "contractKind": "source_stream",
      "target": {
        "request": {
          "url": "https://api.example.com/shipments",
          "method": "GET"
        },
        "response": {
          "format": "json",
          "recordsPath": "data"
        },
        "pagination": {
          "type": "cursor",
          "parameter": "cursor",
          "responsePath": "nextCursor",
          "maxPages": 10
        },
        "primaryKeys": [
          "shipment_id"
        ]
      },
      "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>",
      "idempotencyKey": "<UNIQUE_KEY_FOR_THIS_ENDPOINT>"
    }
  }
}

Set the primary-key hint to shipment_id after checking uniqueness and nulls. Add a useful description, owner and expected freshness. Save in the workspace. Inspect the installed version’s contract before adding optional settings.

Verify the endpoint

  1. Open the saved endpoint and run its supported connection check and schema inspection.
  2. Read a small sample through the supported preview. If the connector has no preview, select a project, open Pipeline Studio and create a pipeline with a source node referencing this endpoint.
  3. Configure a bounded test input using the connector's supported settings, validate the pipeline, save a version and run it.
  4. Inspect returned records or the completed run's output. Compare expected keys and values with the example fixture before scheduling anything.

Compare shipment IDs with the API and follow at least two pages. Confirm final/empty cursors terminate the read instead of repeating pages.

Check connectivity, target validity and actual data separately. Saving does not import records or schedule execution.

Manage the endpoint

Select the matching format for CSV or ZIP. Query parameter values are strings. Configure recurring windows using HTTP windows and pagination; since_last also needs incremental pipeline execution.

Keep credentials on the installation. Review consumers before replacing capabilities, changing targets or deleting endpoints.

FAQ

Why does verification fail?

Check HTTP method, authentication, domain restrictions, records path and cursor names. A 200 response containing an error envelope is not record verification.

Can another project use it?

Yes, within the same workspace and subject to permissions. The endpoint remains workspace-scoped.

What comes after verification?

Use sources in a bounded pipeline, destinations in a supported write flow, and stores in ontology bindings. Inspect real output before scheduling recurring work.