S3 Bucket
Read Parquet files and write Parquet output through S3 endpoint capabilities
What this endpoint does
An S3 endpoint selects an object path in a bucket configured on an installation. A source endpoint reads supported data files; a separate destination endpoint writes output through an installed write capability. The endpoint does not create a bucket.
Like all data endpoints, it belongs to a workspace and can be referenced by projects in that workspace, subject to permissions. Saving it configures access; it does not start ingestion.
Directions and supported operations
| Direction | Capability and contract | Behavior |
|---|---|---|
| Ingress: storage → Semogram | read, source_stream | Read Parquet record files selected by an object path |
| Egress: Semogram → storage | write, write | Upload pipeline output as Parquet objects under a target prefix |
Create separate source and destination endpoints, each selecting its corresponding capability. The bucket and credentials belong to the installation; the path belongs to the endpoint. Read access needs listing and object reads; writing needs object uploads and the permissions required by the storage service's multipart-upload flow.
The destination is object append, not row upsert, row delete or table replacement. Atomicity is per object, not the whole run. Policy writes reject existing object keys. A retry that creates a new key can still duplicate logical records. The destination target uses path; do not add a writeMode field that its contract does not declare.
Plugin used
This plugin provides separate read and write capabilities. Select read for Ingress or write for Egress; the example below installs read first and then explains how to install write. Installation steps for this example are included below. The S3 Bucket installation guide provides optional further detail. Connection details and credentials stay on the installation; the endpoint selects a particular target through that installed capability.
What you need
- A Semogram account with workspace access and permission to manage endpoints
- An existing matching installation, or the connection details to create one using the steps below
- The installation selects your bucket and region, and its credentials can list/read the intended objects. Know the object path and file formats, and prepare a small set of known rows
Read example setup
This example assumes the configured bucket already contains supported order-export files under orders/. That path comes from your storage layout. We choose order_exports as the endpoint name and files as its namespace. order_id is an illustrative column in those files, not a field supplied by S3.
Install and configure the plugin
- Open Plugins in this workspace and select Explore.
- Find S3 Bucket, inspect its publisher/version and select its read capability.
- Name the installation and fill its connection settings using your actual external-system details.
- Save and run the supported connection check. Fix any reported error before creating the endpoint.
Example installation configuration:
Open workspace Plugins → Explore, choose the matching capability and fill its installation settings. Enter values in the labeled controls rather than pasting the whole JSON object.
| UI field | Example value |
|---|---|
| Bucket | YOUR_BUCKET |
| Region | YOUR_REGION |
| Credentials → Access key id | YOUR_ACCESS_KEY |
| Credentials → Secret access key | YOUR_SECRET_KEY |
Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Enter credentials in the protected fields and review the selected installation before saving.
In the platform assistant or your connected MCP assistant, ask:
Install the plugin described on this page in this workspace. Discover its catalog entry, select the matching capability and propose the installation using the connection settings shown here. Ask me to enter credentials in protected installation fields. Show the selected plugin/version, capability and non-secret settings before saving.Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.
Use plugin_catalog_list / plugin_catalog_get to obtain the discovery ID and matching capability class (reads, writes or factStores). Call plugin_installation_create with the arguments below through an authenticated MCP connection. The workspace comes from that connection. Enter credentials through an authorized protected configuration path; do not send real secrets as conversational prompt text.
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "plugin_installation_create",
"arguments": {
"capabilityClass": "<MATCHING_CAPABILITY_CLASS>",
"discoveryId": "<DISCOVERY_ID_FROM_CATALOG>",
"name": "<INSTALLATION_NAME>",
"config": {
"bucket": "YOUR_BUCKET",
"region": "YOUR_REGION",
"credentials": {
"accessKeyId": "YOUR_ACCESS_KEY",
"secretAccessKey": "YOUR_SECRET_KEY"
}
},
"idempotencyKey": "<UNIQUE_KEY_FOR_THIS_INSTALLATION>"
}
}
}This operation has no standalone public /api/v1 plugin-installation/catalog route in the current implementation. Use the UI or MCP methods shown here.
This operation has no standalone public /api/v1 plugin-installation/catalog route in the current implementation. Use the UI or MCP methods shown here.
Replace every YOUR_… value. Enter credentials only in protected installation configuration. The configuration above is separate from the endpoint target below; selecting a target does not create or authenticate the connection.
Prepare example data
The current reader selects Parquet objects, not CSV or arbitrary bucket files. Create a two-row fixture locally with Python and PyArrow:
python -m pip install pyarrowimport pyarrow as pa
import pyarrow.parquet as pq
pq.write_table(pa.table({
"order_id": [1, 2],
"total": [120.00, 75.50]
}), "orders.parquet")Upload orders.parquet with your storage console to orders/orders.parquet in the configured test bucket. Allow the read installation to list this prefix and fetch this object. Use target path orders without a trailing slash: the reader matches the exact key or keys beginning with orders/. A full S3 URI is not the path selector. The first bounded read should contain the two fixture rows.
For compatible storage, set the installation's HTTPS endpoint and path-style option as required by your provider.
Use the assistant
Open Data Endpoints → New data endpoint in the workspace. Describe your actual target and choose a name for the endpoint. For the example above, you could ask:
Create a source endpoint named order_exports in namespace files.
Use our installed S3 Bucket read capability with contract source_stream
and set these target fields:
Path: orders
Review the proposed connection and target before saving.Replace example values with your own. Give the assistant the installed connection reference; keep credentials in protected installation settings.
Configure manually
Choose Edit manually, select the matching installed capability and configure:
| Field | Value |
|---|---|
| Name | order_exports |
| Namespace | files |
| Role | source |
| Contract | source_stream |
Open workspace Data Endpoints → New data endpoint → Edit manually, choose the direction and installed capability described in this example, then fill the target fields. Enter values in the labeled controls rather than pasting the whole JSON object.
| UI field | Example value |
|---|---|
| Path | orders |
Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Review the endpoint name, direction, capability and selected target before saving.
In the platform assistant or your connected MCP assistant, ask:
Create the endpoint described on this page using these settings:
name: <ENDPOINT_NAME_FROM_THIS_EXAMPLE>
namespace: <ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>
role: source
contractKind: source_stream
target / path: orders
pluginCapabilityInstallationId: <INSTALLED_CAPABILITY_UUID>
Use the actual installed capability and the endpoint name/namespace selected in this example. Show the proposed direction, connection and target before saving. Keep credentials on the installation.Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.
Use a workspace API key with endpoints:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints" \
--header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
--header "Idempotency-Key: <UNIQUE_KEY_FOR_THIS_ENDPOINT>" \
--header "Content-Type: application/json" \
--data-binary @- <<'JSON'
{
"name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
"namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
"role": "source",
"contractKind": "source_stream",
"target": {
"path": "orders"
},
"pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>"
}
JSONCall source_create with the arguments below through an authenticated workspace MCP connection. Replace the name/namespace placeholders with the labels chosen in this example and use the actual installed capability UUID. Set the role/contract to the direction described here; the workspace is resolved from the connection. This configures an endpoint and does not execute a read or write.
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "source_create",
"arguments": {
"name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
"namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
"role": "source",
"contractKind": "source_stream",
"target": {
"path": "orders"
},
"pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>",
"idempotencyKey": "<UNIQUE_KEY_FOR_THIS_ENDPOINT>"
}
}
}Set the primary-key hint to order_id after checking uniqueness and nulls. Add a useful description, owner and expected freshness. Save in the workspace. Inspect the installed version’s contract before adding optional settings.
Verify the endpoint
- Open the saved endpoint and run its supported connection check and schema inspection.
- Read a small sample through the supported preview. If the connector has no preview, select a project, open Pipeline Studio and create a pipeline with a source node referencing this endpoint.
- Configure a bounded test input using the connector's supported settings, validate the pipeline, save a version and run it.
- Inspect returned records or the completed run's output. Compare expected keys and values with the example fixture before scheduling anything.
Compare selected object names, supported file formats, rows and keys with storage. Confirm what the path selects; not all files under a prefix necessarily share a schema.
Check connectivity, target validity and actual data separately. Saving does not import records or schedule execution.
Write example: export two orders as Parquet objects
Use the two-row orders/orders.parquet fixture and order_exports source created above. The destination will be named verified_order_files, namespace files, and will write to a separate bucket prefix verified-orders/. It does not modify the input Parquet file.
- In workspace Plugins → Explore, install S3 Bucket → write using the bucket, region and protected credentials configuration above. For compatible storage, include the required HTTPS endpoint and path-style setting. Give this installation upload permissions on the dedicated output prefix, including the multipart-upload permissions required by your provider.
- Open Data Endpoints → New data endpoint → Edit manually. Set name
verified_order_files, namespacefiles, Egress (roledestination), the installed write capability and contractwrite. - Save this target; the writer combines the prefix with each runtime output file name:
Open workspace Data Endpoints → New data endpoint → Edit manually, choose the direction and installed capability described in this example, then fill the target fields. Enter values in the labeled controls rather than pasting the whole JSON object.
| UI field | Example value |
|---|---|
| Path | verified-orders/ |
Nested labels above identify the containing group. Lists use the form’s list controls; open-ended objects use its object editor. Labels and available options follow the installed version’s contract. Review the endpoint name, direction, capability and selected target before saving.
In the platform assistant or your connected MCP assistant, ask:
Create the endpoint described on this page using these settings:
name: <ENDPOINT_NAME_FROM_THIS_EXAMPLE>
namespace: <ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>
role: destination
contractKind: write
target / path: verified-orders/
pluginCapabilityInstallationId: <INSTALLED_CAPABILITY_UUID>
Use the actual installed capability and the endpoint name/namespace selected in this example. Show the proposed direction, connection and target before saving. Keep credentials on the installation.Replace placeholders with real accessible resources. The assistant prepares the operation; inspect its proposed inputs and result.
Use a workspace API key with endpoints:write. Set SEMOGRAM_API_KEY in your shell; replace resource placeholders with real IDs. This is an HTTP resource request, not an MCP JSON-RPC message.
curl --request POST "https://platform.semogram.com/api/v1/data-endpoints" \
--header "Authorization: Bearer ${SEMOGRAM_API_KEY}" \
--header "Idempotency-Key: <UNIQUE_KEY_FOR_THIS_ENDPOINT>" \
--header "Content-Type: application/json" \
--data-binary @- <<'JSON'
{
"name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
"namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
"role": "destination",
"contractKind": "write",
"target": {
"path": "verified-orders/"
},
"pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>"
}
JSONCall source_create with the arguments below through an authenticated workspace MCP connection. Replace the name/namespace placeholders with the labels chosen in this example and use the actual installed capability UUID. Set the role/contract to the direction described here; the workspace is resolved from the connection. This configures an endpoint and does not execute a read or write.
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "source_create",
"arguments": {
"name": "<ENDPOINT_NAME_FROM_THIS_EXAMPLE>",
"namespace": "<ENDPOINT_NAMESPACE_FROM_THIS_EXAMPLE>",
"role": "destination",
"contractKind": "write",
"target": {
"path": "verified-orders/"
},
"pluginCapabilityInstallationId": "<INSTALLED_CAPABILITY_UUID>",
"idempotencyKey": "<UNIQUE_KEY_FOR_THIS_ENDPOINT>"
}
}
}- In a workspace project, open Pipeline Studio. Connect an ingress node referencing
order_exportsto an egress node referencingverified_order_files. Use only the fixture file. Review the target, supported object-append policy and caller permissions, validate, save a version and run once. - Inspect the run and list
verified-orders/using your storage console. Download the newly produced Parquet file(s) and inspect them with a Parquet reader. Across the output files, expect two records with order IDs1and2and totals matching the fixture. File count can vary with batching; verify record count rather than assuming one file per run. - Confirm the original
orders/orders.parquetobject is unchanged. Record the output keys and run ID so you can distinguish a retry from a new export.
A failed run can leave already committed objects. Policy writes reject an existing key; new keys can still duplicate the same logical rows. Do not blindly rerun append after partial failure. No destination field adds row-level upsert or rollback to this object writer.
Manage the endpoint
For output, create a separate destination endpoint on the installed write capability with write contract and a dedicated target such as {"path":"verified-orders/"}. Test its files before production use.
Keep credentials on the installation. Review consumers before replacing capabilities, changing targets or deleting endpoints.
FAQ
Why does verification fail?
Check bucket/region, path semantics and object permissions. Do not assume a full S3 URI is interchangeable with the published path selector.
Can another project use it?
Yes, within the same workspace and subject to permissions. The endpoint remains workspace-scoped.
What comes after verification?
Use sources in a bounded pipeline, destinations in a supported write flow, and stores in ontology bindings. Inspect real output before scheduling recurring work.