Semogram Docs
PluginsSetup guides

HTTP Dataset

Install HTTP Dataset, configure its capabilities and verify a bounded operation.

Use @craven/http-dataset to read record-oriented JSON, CSV, and ZIP datasets from HTTP APIs.

Install

Before you start

You need a Semogram account with access to the workspace and permission to manage plugin installations. You also need access to the external system; installing a plugin does not grant credentials or provision it.

Identify the provider URL, authentication, response format and location of its records. Determine its pagination scheme and rate limits. Keep provider credentials in installation configuration, separate from the endpoint’s request and parsing rules.

Configure the installation

  1. Open Plugins in your workspace and select Explore.
  2. Find HTTP Dataset, inspect its publisher, version and capabilities, and choose the matching capability.
  3. Give the installation a useful name and fill in its connection configuration.
  4. Save it and run the connection check where the capability provides one. Read the returned error before proceeding.

This is an illustrative configuration. Replace every example value; never paste real credentials into an assistant prompt or public document.

Open workspace Plugins → Explore, choose the matching capability and fill its installation settings. Enter values in the labeled controls rather than pasting the whole JSON object.

UI fieldExample value
Timeout ms30000
Rate limit ms250
Max records1000
Allowed domainsAdd api.example.com

Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Enter credentials in the protected fields and review the selected installation before saving.

In the platform assistant or your connected MCP assistant, ask:

Assistant prompt
Install the plugin described on this page in this workspace. Discover its catalog entry, select the matching capability and propose the installation using the connection settings shown here. Ask me to enter credentials in protected installation fields. Show the selected plugin/version, capability and non-secret settings before saving.

Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.

Use plugin_catalog_list / plugin_catalog_get to obtain the discovery ID and matching capability class (reads, writes or factStores). Call plugin_installation_create with the arguments below through an authenticated MCP connection. The workspace comes from that connection. Enter credentials through an authorized protected configuration path; do not send real secrets as conversational prompt text.

MCP request
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "plugin_installation_create",
    "arguments": {
      "capabilityClass": "<MATCHING_CAPABILITY_CLASS>",
      "discoveryId": "<DISCOVERY_ID_FROM_CATALOG>",
      "name": "<INSTALLATION_NAME>",
      "config": {
        "timeoutMs": 30000,
        "rateLimitMs": 250,
        "maxRecords": 1000,
        "allowedDomains": [
          "api.example.com"
        ]
      },
      "idempotencyKey": "<UNIQUE_KEY_FOR_THIS_INSTALLATION>"
    }
  }
}

This operation has no standalone public /api/v1 plugin-installation/catalog route in the current implementation. Use the UI or MCP methods shown here.

Configuration fields

FieldTypeRequiredDetails
timeoutMsintegerNoDefault: 30000.
rateLimitMsintegerNoDefault: 0.
maxRecordsintegerNoDefault: 100000.
maxResponseBytesintegerNoDefault: 262144000.
allowedDomainsarrayNoItems: string.
userAgentstringNoDefault: "Semogram-HttpDataset/0.0.1".
maxRedirectsintegerNoDefault: 5.
authobjectNoProtected credentials shared by endpoints using this installation. Nested fields: type, token, name, value, username, password.

View the complete published contracts for nested settings and endpoint selectors. Inspect the installed version’s contract before configuring optional fields; catalog availability and versions can vary by deployment.

Make your first read

Follow the data endpoint scenario for target setup and verification.

Open Data Endpoints, choose the installed capability and configure a concrete target. For this plugin, an illustrative target is:

Open workspace Data Endpoints → New data endpoint → Edit manually, choose the direction and installed capability described in this example, then fill the target fields. Enter values in the labeled controls rather than pasting the whole JSON object.

UI fieldExample value
Request → Urlhttps://api.example.com/orders
Request → MethodGET
Request → Query{"limit": "100"} (object editor)
Response → Formatjson
Response → Records pathdata.orders
Pagination → Typenone
Primary keysAdd id
Streamorders

Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Review the endpoint name, direction, capability and selected target before saving.

In the platform assistant or your connected MCP assistant, ask:

Assistant prompt
Create the endpoint described on this page using these settings:
name: <ENDPOINT_NAME_FROM_THIS_EXAMPLE>
namespace: <ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>
role: source
contractKind: source_stream
target / request / url: https://api.example.com/orders
target / request / method: GET
target / request / query / limit: 100
target / response / format: json
target / response / recordsPath: data.orders
target / pagination / type: none
target / primaryKeys: id
target / stream: orders
pluginCapabilityInstallationId: <INSTALLED_CAPABILITY_UUID>
Use the actual installed capability and the endpoint name/namespace selected in this example. Show the proposed direction, connection and target before saving. Keep credentials on the installation.

Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.

Use a workspace API key with endpoints:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.

HTTP API request
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints" \
  --header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
  --header "Idempotency-Key: <UNIQUE_KEY_FOR_THIS_ENDPOINT>" \
  --header "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
  "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
  "role": "source",
  "contractKind": "source_stream",
  "target": {
    "request": {
      "url": "https://api.example.com/orders",
      "method": "GET",
      "query": {
        "limit": "100"
      }
    },
    "response": {
      "format": "json",
      "recordsPath": "data.orders"
    },
    "pagination": {
      "type": "none"
    },
    "primaryKeys": [
      "id"
    ],
    "stream": "orders"
  },
  "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>"
}
JSON

Call source_create with the arguments below through an authenticated workspace MCP connection. Replace the name/namespace placeholders with the labels chosen in this example and use the actual installed capability UUID. Set the role/contract to the direction described here; the workspace is resolved from the connection. This configures an endpoint and does not execute a read or write.

MCP request
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "source_create",
    "arguments": {
      "name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
      "namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
      "role": "source",
      "contractKind": "source_stream",
      "target": {
        "request": {
          "url": "https://api.example.com/orders",
          "method": "GET",
          "query": {
            "limit": "100"
          }
        },
        "response": {
          "format": "json",
          "recordsPath": "data.orders"
        },
        "pagination": {
          "type": "none"
        },
        "primaryKeys": [
          "id"
        ],
        "stream": "orders"
      },
      "pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>",
      "idempotencyKey": "<UNIQUE_KEY_FOR_THIS_ENDPOINT>"
    }
  }
}

Fetch a bounded sample, compare row counts with the provider response and inspect the mapped primary keys. Test a second page when pagination is configured; a first-page success is not proof that cursor or link traversal works.

Installation is not ingestion. Use the saved endpoint in a supported query or a small pipeline, validate the pipeline, execute it and inspect its run. Add a schedule only after the bounded run succeeds.

Manage the connection

Formats are json, csv, zip_json and zip_csv. Configure recordsPath or zipEntry as needed. Pagination supports none, page, offset, cursor and link. Query parameter values are strings. A since_last or rolling time window requires the provider’s actual date parameter names and formats.

Change connection settings on the installation and target selectors on the endpoint. Recheck the connection after credential rotation. Review consuming endpoints and pipelines before replacing a capability or removing its installation.

For assistant-driven setup, use catalog discovery, installation creation and installation checks. Use real IDs returned by discovery, not the example name as an ID.

FAQ

Why does the check or first run fail?

Check recordsPath, ZIP entry, response format, authentication and pagination parameters. Empty fields keeps provider fields; a configured mapping projects and optionally coerces them. Inspect domain restrictions, redirect limits and response-size limits when requests fail.

Does installing this plugin start a sync?

No. A check verifies a supported connection probe. Queries and pipeline runs perform reads; recurring ingestion needs a schedule. Inspect run status, records and evidence to confirm actual work.

Can an assistant use every operation after installation?

Only operations supported by the capability and permitted for the caller. The MCP tool reference marks plugin and connector requirements; declaring a write capability does not grant every source-write or maintenance operation.

What should I verify before using production data?

Check a known small input, the expected result, permissions, supported read/write behavior and failure handling. These guides describe the implemented contracts; they are not a claim that your external system has already passed a live connection test.