Iceberg
Read snapshots and perform governed writes to transactional Iceberg tables
What this endpoint does
An Iceberg endpoint selects a registered transactional table in an Iceberg catalog. It supports table reads and, through supported operations, historical snapshots, governed source writes and table maintenance. Those operations have separate access and policy requirements.
Like all data endpoints, it belongs to a workspace and can be referenced by projects in that workspace, subject to permissions. Saving it configures access; it does not start ingestion.
Directions and supported operations
| Operation | Direction | Behavior |
|---|---|---|
| Table read | Iceberg → Semogram | Read a registered table, including supported pinned historical snapshots |
| Governed source mutation | Semogram → the selected Iceberg table | Append, upsert, update or delete through the public source-write path when policy permits it |
| Maintenance | Semogram → catalog/warehouse | Compact files, expire eligible snapshots and inspect orphan files under separate permissions |
This integration uses a source endpoint and a transactional table adapter for governed mutations. It is not the generic destination-writer configuration used by Postgres, MongoDB or S3. Selecting Ingress does not authorize mutation: source-write settings, an endpoint-bound policy and an actor grant are separate requirements.
The adapter supports single-table run atomicity, durable operation receipts and snapshot conflict checks. Identity comes from the approved policy; duplicate and schema rules must match the advertised capabilities. The lower-level table adapter also supports replacement; the public source-write request does not expose a replace operation. Tables use unpartitioned Parquet in the supported PostgreSQL catalog/S3 warehouse configuration. Compaction and retention do not imply arbitrary partition evolution or automatic physical orphan deletion.
Plugin used
This endpoint uses the Iceberg read capability. Installation steps for this example are included below. The Iceberg installation guide provides optional further detail. Connection details and credentials stay on the installation; the endpoint selects a particular target through that installed capability.
What you need
- A Semogram account with workspace access and permission to manage endpoints
- An existing matching installation, or the connection details to create one using the steps below
- The installation can access the catalog and object warehouse, and the table is registered in that catalog. Know its namespace and table name and have a small known sample
Read example setup
This example assumes your catalog already contains a table named orders in namespace raw. The target namespace is an array because that is the Iceberg contract. We choose orders as the Semogram endpoint name and warehouse as its grouping namespace. The endpoint namespace does not change the catalog namespace. order_id is an example table column.
Install and configure the plugin
- Open Plugins in this workspace and select Explore.
- Find Iceberg, inspect its publisher/version and select its read capability.
- Name the installation and fill its connection settings using your actual external-system details.
- Save and run the supported connection check. Fix any reported error before creating the endpoint.
Example installation configuration:
Open workspace Plugins → Explore, choose the matching capability and fill its installation settings. Enter values in the labeled controls rather than pasting the whole JSON object.
| UI field | Example value |
|---|---|
| Catalog id | operations |
| Namespace | Add raw |
| Catalog uri | postgresql+psycopg2://USER:PASSWORD@YOUR_CATALOG_HOST/catalog?sslmode=verify-full |
| Warehouse | s3://YOUR_BUCKET/tables |
| S3endpoint | https://YOUR_STORAGE_HOST |
| S3region | YOUR_REGION |
| S3access key id | YOUR_ACCESS_KEY |
| S3secret access key | YOUR_SECRET_KEY |
Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Enter credentials in the protected fields and review the selected installation before saving.
In the platform assistant or your connected MCP assistant, ask:
Install the plugin described on this page in this workspace. Discover its catalog entry, select the matching capability and propose the installation using the connection settings shown here. Ask me to enter credentials in protected installation fields. Show the selected plugin/version, capability and non-secret settings before saving.Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.
Use plugin_catalog_list / plugin_catalog_get to obtain the discovery ID and matching capability class (reads, writes or factStores). Call plugin_installation_create with the arguments below through an authenticated MCP connection. The workspace comes from that connection. Enter credentials through an authorized protected configuration path; do not send real secrets as conversational prompt text.
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "plugin_installation_create",
"arguments": {
"capabilityClass": "<MATCHING_CAPABILITY_CLASS>",
"discoveryId": "<DISCOVERY_ID_FROM_CATALOG>",
"name": "<INSTALLATION_NAME>",
"config": {
"catalogId": "operations",
"namespace": [
"raw"
],
"catalogUri": "postgresql+psycopg2://USER:PASSWORD@YOUR_CATALOG_HOST/catalog?sslmode=verify-full",
"warehouse": "s3://YOUR_BUCKET/tables",
"s3Endpoint": "https://YOUR_STORAGE_HOST",
"s3Region": "YOUR_REGION",
"s3AccessKeyId": "YOUR_ACCESS_KEY",
"s3SecretAccessKey": "YOUR_SECRET_KEY"
},
"idempotencyKey": "<UNIQUE_KEY_FOR_THIS_INSTALLATION>"
}
}
}This operation has no standalone public /api/v1 plugin-installation/catalog route in the current implementation. Use the UI or MCP methods shown here.
This operation has no standalone public /api/v1 plugin-installation/catalog route in the current implementation. Use the UI or MCP methods shown here.
Replace every YOUR_… value. Enter credentials only in protected installation configuration. The configuration above is separate from the endpoint target below; selecting a target does not create or authenticate the connection.
Prepare catalog access
This runtime uses a PostgreSQL PyIceberg SQL catalog and an S3-compatible warehouse. Remote catalog connections require sslmode=verify-full; storage uses HTTPS. Catalog/namespace segments use lowercase letters, digits and underscores and start with a letter.
Have your catalog administrator register a small test table in namespace raw named orders, with known order_id values. Object files in a bucket without catalog registration are insufficient. This setup uses your existing catalog/warehouse; it does not provision those external services.
Use the assistant
Open Data Endpoints → New data endpoint in the workspace. Describe your actual target and choose a name for the endpoint. For the example above, you could ask:
Create a source endpoint named orders in namespace warehouse.
Use our installed Iceberg read capability with contract source_stream
and set these target fields:
Namespace: raw
Table: orders
Review the proposed connection and target before saving.Replace example values with your own. Give the assistant the installed connection reference; keep credentials in protected installation settings.
Configure manually
Choose Edit manually, select the matching installed capability and configure:
| Field | Value |
|---|---|
| Name | orders |
| Namespace | warehouse |
| Role | source |
| Contract | source_stream |
Open workspace Data Endpoints → New data endpoint → Edit manually, choose the direction and installed capability described in this example, then fill the target fields. Enter values in the labeled controls rather than pasting the whole JSON object.
| UI field | Example value |
|---|---|
| Namespace | Add raw |
| Table | orders |
Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Review the endpoint name, direction, capability and selected target before saving.
In the platform assistant or your connected MCP assistant, ask:
Create the endpoint described on this page using these settings:
name: <ENDPOINT_NAME_FROM_THIS_EXAMPLE>
namespace: <ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>
role: source
contractKind: source_stream
target / namespace: raw
target / table: orders
pluginCapabilityInstallationId: <INSTALLED_CAPABILITY_UUID>
Use the actual installed capability and the endpoint name/namespace selected in this example. Show the proposed direction, connection and target before saving. Keep credentials on the installation.Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.
Use a workspace API key with endpoints:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints" \
--header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
--header "Idempotency-Key: <UNIQUE_KEY_FOR_THIS_ENDPOINT>" \
--header "Content-Type: application/json" \
--data-binary @- <<'JSON'
{
"name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
"namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
"role": "source",
"contractKind": "source_stream",
"target": {
"namespace": [
"raw"
],
"table": "orders"
},
"pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>"
}
JSONCall source_create with the arguments below through an authenticated workspace MCP connection. Replace the name/namespace placeholders with the labels chosen in this example and use the actual installed capability UUID. Set the role/contract to the direction described here; the workspace is resolved from the connection. This configures an endpoint and does not execute a read or write.
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "source_create",
"arguments": {
"name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
"namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
"role": "source",
"contractKind": "source_stream",
"target": {
"namespace": [
"raw"
],
"table": "orders"
},
"pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>",
"idempotencyKey": "<UNIQUE_KEY_FOR_THIS_ENDPOINT>"
}
}
}Set the primary-key hint to order_id after checking uniqueness and nulls. Add a useful description, owner and expected freshness. Save in the workspace. Inspect the installed version’s contract before adding optional settings.
Verify the endpoint
- Open the saved endpoint and run its supported connection check and schema inspection.
- Read a small sample through the supported preview. If the connector has no preview, select a project, open Pipeline Studio and create a pipeline with a source node referencing this endpoint.
- Configure a bounded test input using the connector's supported settings, validate the pipeline, save a version and run it.
- Inspect returned records or the completed run's output. Compare expected keys and values with the example fixture before scheduling anything.
Compare a bounded read with known rows and record its snapshot ID. Historical reads require an available snapshot and supported snapshot/as-of operation.
Check connectivity, target validity and actual data separately. Saving does not import records or schedule execution.
Write example: append one row to the same test table
This writes back to raw.orders through the source endpoint created above. It does not create an Egress endpoint. Use a dedicated test table with columns order_id and total, compatible with the row below, and confirm ID 3 is absent before starting.
- Configure the workspace's source-write setting for the intended governed write path. Bind an approved write policy to this endpoint allowing append and matching the adapter's run atomicity, schema and concurrency capabilities. Configure identity keys such as
order_idonly when the table and policy support them. - Grant the calling actor write access to this specific endpoint and operation. For the MCP path, use an API key with
data:writeand the endpoint's required grant; an OAuth read connection alone does not authorize this mutation. External catalog and warehouse credentials also need write access. - Read the test endpoint immediately before mutation. Keep the returned snapshot
idandversion, and the saved endpoint UUID. These values come from the read result and endpoint, not from the example table name. - Through an authenticated, initialized MCP connection, call
source_write_executewith these arguments, replacing all placeholders:
Ask your connected assistant to read this test endpoint, inspect its current snapshot and approved append policy/grant, then propose appending order ID 3 with total 42.00. Review the proposed mutation before execution. Keep the operation's idempotency key when inspecting or recovering an uncertain result.
Use a workspace API key with data:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints/<ENDPOINT_UUID>/source-writes" \
--header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
--header "Idempotency-Key: orders-test-append-3-unique-operation" \
--header "Content-Type: application/json" \
--data-binary @- <<'JSON'
{
"contractVersion": 1,
"expectedSourceSnapshot": {
"id": "<SNAPSHOT_ID_FROM_READ>",
"version": "<SNAPSHOT_VERSION_FROM_READ>"
},
"operation": "append",
"rows": [
{
"order_id": 3,
"total": 42.0
}
]
}
JSONThese are the arguments for source_write_execute, not endpoint form fields. Send them through an authenticated, initialized MCP connection using tools/call.
{
"sourceId": "<ENDPOINT_UUID>",
"request": {
"contractVersion": 1,
"expectedSourceSnapshot": {
"id": "<SNAPSHOT_ID_FROM_READ>",
"version": "<SNAPSHOT_VERSION_FROM_READ>"
},
"operation": "append",
"rows": [{"order_id":3,"total":42.00}]
},
"idempotencyKey": "orders-test-append-3-unique-operation"
}- Inspect the returned operation state/receipt. Read the endpoint again and compare ID
3, total42.00, affected counts and resulting snapshot with the receipt. A queued or failed response is not proof of a completed write. - For an uncertain result, inspect/recover the original operation and reuse the same idempotency key only for the exact same request. A new key means new work. If the source changed, obtain a fresh snapshot and review the change before issuing a new mutation; do not remove the concurrency check.
Upsert sends rows and uses the policy's identity. Update sends where predicates and set; delete sends where. Those requests require their own allowed operation and grant. An endpoint read, an append grant and an endpoint primary-key hint do not authorize all mutations.
Snapshot history, compaction and expiration are separate operations. Retained snapshots can support historical reads; expired snapshots cannot be assumed available. Orphan discovery reports candidates and does not by itself delete storage objects.
Manage the endpoint
Durable source writes and maintenance require Iceberg-specific runtime support. Bind the approved policy and actor grant before mutation, then inspect receipts and snapshots. See write access.
Keep credentials on the installation. Review consumers before replacing capabilities, changing targets or deleting endpoints.
FAQ
Why does verification fail?
Check namespace, catalog registration, TLS and warehouse permissions. An object path alone is not a registered Iceberg table.
Can another project use it?
Yes, within the same workspace and subject to permissions. The endpoint remains workspace-scoped.
What comes after verification?
Use sources in a bounded pipeline, destinations in a supported write flow, and stores in ontology bindings. Inspect real output before scheduling recurring work.