Dataset upload
Create an endpoint from a file and verify its ingestion
What this endpoint does
An upload endpoint exposes a file preserved in Semogram's managed storage. A pipeline can read the selected file version and turn it into usable records. Uploading or selecting a version is separate from running ingestion.
Storage used
The upload flow manages its storage binding. You do not need to install your own S3 plugin for this example. For files already in your bucket, use an S3 source endpoint instead; the S3 installation guide explains that optional alternative.
What you need
You need a Semogram account with workspace access and permission to upload data. Prepare CSV, JSON, NDJSON, Parquet or ZIP. Maximum upload is 256 MiB; JSON and ZIP processing is limited to 32 MiB. Use a small known fixture first.
Example setup
Save the following as orders.csv on your computer. order_id and total are column names in this example file, not platform-provided fields.
order_id,total
1,120.00
2,75.50Use this file for the steps below. The endpoint name is a label you choose for it in Semogram. Have a project available for the pipeline created by the upload flow.
Upload and preview
- Open the workspace's Data Endpoints → Upload dataset.
- Select the file and let upload finish; interrupted uploads can resume.
- Review detected columns, sample rows and parsing errors. Check dates, quoted CSV values, nulls and identifiers.
- Choose Create endpoint and pipeline and provide the requested project context for the pipeline.
- Inspect the resulting endpoint and pipeline in Studio. The endpoint is workspace-scoped; the pipeline belongs to its project.
Original bytes are preserved privately. An upload endpoint uses the platform-managed file path; do not invent a connector target or replace its internal upload references manually.
Verify the endpoint
Run the created pipeline and compare output with the original fixture. Previewing bytes is not the same as completing ingestion. Check run state, row counts, types and a known record.
Manage the endpoint
Upload a new file version to an existing upload endpoint and choose Use this version when ready. Inspect the selected version and rerun the consuming pipeline; choosing a version is separate from ingestion. Review consumers before deleting an endpoint.
FAQ
Do I need to install S3 first?
The upload flow manages its storage binding. This differs from an endpoint reading your own S3 installation.
Why does upload succeed but parsing fail?
Check supported format, encoding, ZIP contents and the lower JSON/ZIP processing limit. Inspect the returned error before retrying.
Does a new version update existing pipeline output?
Run the pipeline to process the selected version, then inspect the new output and evidence