HTTP windows and pagination
Configure bounded API reads and recurring time windows
What this endpoint can do
An HTTP Dataset endpoint can request API records in pages and limit each recurring read to a configured time window. It parses responses using the target's format and record path.
Plugin used
This uses the HTTP Dataset read capability. Connection/authentication settings belong on its installation. HTTP Dataset installation provides further details.
What you need
Have an installed HTTP Dataset capability, a reachable API, its request/response documentation and known test records. Window parameter names must be accepted by that API. For since_last, use an incremental pipeline; setting a window alone does not schedule execution.
Example setup
For an unauthenticated public API, open Plugins → Explore, select HTTP Dataset's read capability, install it without auth and run its supported check. For an authenticated API, configure protected auth on the installation first.
Open workspace Data Endpoints → New data endpoint. Use the assistant or Edit manually, select source role, that installed capability and source_stream. Choose an endpoint name such as earthquakes and namespace public_data; these are your Semogram labels.
The example below requests the USGS earthquake API. Its response array is features, its record identifier is id, and it accepts starttime/endtime date filters. These are fields of this example API, not universal HTTP Dataset fields. Set window on the target and omit the same time parameters from request.query.
Window fields
| Field | Meaning | Default |
|---|---|---|
startParameter | Parameter that receives the window start | Required |
endParameter | Parameter that receives the window end | None |
location | query, or body for POST requests | query |
mode | since_last resumes where the last successful run ended; rolling reads the last lookback every run | since_last |
lookback | ISO 8601 duration such as P30D or PT6H: the first since_last window, the most any window reaches back, and the rolling length | Required |
overlap | Duration re-read before the last window end, for records that arrive late | None |
format | iso, date (YYYY-MM-DD), epoch_ms, or epoch_s | iso |
Open workspace Data Endpoints → New data endpoint → Edit manually, choose the direction and installed capability described in this example, then fill the target fields. Enter values in the labeled controls rather than pasting the whole JSON object.
| UI field | Example value |
|---|---|
| Request → Url | https://earthquake.usgs.gov/fdsnws/event/1/query |
| Request → Query → Format | geojson |
| Request → Query → Minmagnitude | 4.5 |
| Request → Query → Orderby | time-asc |
| Request → Query → Limit | 20000 |
| Response → Format | json |
| Response → Records path | features |
| Window → Start parameter | starttime |
| Window → End parameter | endtime |
| Window → Mode | since_last |
| Window → Lookback | P30D |
| Window → Overlap | PT1H |
| Primary keys | Add id |
Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Review the endpoint name, direction, capability and selected target before saving.
In the platform assistant or your connected MCP assistant, ask:
Create the endpoint described on this page using these settings:
name: <ENDPOINT_NAME_FROM_THIS_EXAMPLE>
namespace: <ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>
role: source
contractKind: source_stream
target / request / url: https://earthquake.usgs.gov/fdsnws/event/1/query
target / request / query / format: geojson
target / request / query / minmagnitude: 4.5
target / request / query / orderby: time-asc
target / request / query / limit: 20000
target / response / format: json
target / response / recordsPath: features
target / window / startParameter: starttime
target / window / endParameter: endtime
target / window / mode: since_last
target / window / lookback: P30D
target / window / overlap: PT1H
target / primaryKeys: id
pluginCapabilityInstallationId: <INSTALLED_CAPABILITY_UUID>
Use the actual installed capability and the endpoint name/namespace selected in this example. Show the proposed direction, connection and target before saving. Keep credentials on the installation.Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.
Use a workspace API key with endpoints:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints" \
--header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
--header "Idempotency-Key: <UNIQUE_KEY_FOR_THIS_ENDPOINT>" \
--header "Content-Type: application/json" \
--data-binary @- <<'JSON'
{
"name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
"namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
"role": "source",
"contractKind": "source_stream",
"target": {
"request": {
"url": "https://earthquake.usgs.gov/fdsnws/event/1/query",
"query": {
"format": "geojson",
"minmagnitude": "4.5",
"orderby": "time-asc",
"limit": "20000"
}
},
"response": {
"format": "json",
"recordsPath": "features"
},
"window": {
"startParameter": "starttime",
"endParameter": "endtime",
"mode": "since_last",
"lookback": "P30D",
"overlap": "PT1H"
},
"primaryKeys": [
"id"
]
},
"pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>"
}
JSONCall source_create with the arguments below through an authenticated workspace MCP connection. Replace the name/namespace placeholders with the labels chosen in this example and use the actual installed capability UUID. Set the role/contract to the direction described here; the workspace is resolved from the connection. This configures an endpoint and does not execute a read or write.
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "source_create",
"arguments": {
"name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
"namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
"role": "source",
"contractKind": "source_stream",
"target": {
"request": {
"url": "https://earthquake.usgs.gov/fdsnws/event/1/query",
"query": {
"format": "geojson",
"minmagnitude": "4.5",
"orderby": "time-asc",
"limit": "20000"
}
},
"response": {
"format": "json",
"recordsPath": "features"
},
"window": {
"startParameter": "starttime",
"endParameter": "endtime",
"mode": "since_last",
"lookback": "P30D",
"overlap": "PT1H"
},
"primaryKeys": [
"id"
]
},
"pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>",
"idempotencyKey": "<UNIQUE_KEY_FOR_THIS_ENDPOINT>"
}
}
}The window ends at the run's sync time. since_last needs the pipeline's
source node in incremental mode: its position advances only when the node
succeeds, so a failed run is read again next time. Pair an overlap with an
upsert keyed on the primary key so re-read records do not duplicate
Pagination
Pagination lives in target.pagination for HTTP Dataset. Supported types are none, page, offset, cursor and link. Match parameter names and response paths to the API; a type declaration does not discover its protocol.
| Field | Purpose |
|---|---|
location | Query parameters or POST body |
parameter | Page, offset or cursor request parameter |
responsePath | Response path supplying the next cursor/link where required |
start | Initial page/offset where applicable |
pageSizeParameter, pageSize | API page-size parameter and value |
limitParameter, limit | API limit parameter and value |
maxPages | Bound on pages; published default is 100 |
Start with two small pages and compare keys across the boundary. Confirm termination, retry behavior and empty responses. A page bound can deliberately truncate results; inspect logs and counts rather than treating every bounded read as a full export.
See the HTTP Dataset scenario and published contract