Classify records
Use a bundled AI transform to classify a small text fixture and review its rationale
This example classifies two issue descriptions into billing or delivery. It uses Postgres read and the agentic classification transform. Unlike field projection, the labels are model outputs that require review; this fixture demonstrates the flow, not a measured accuracy claim.
What you need
You need a Semogram account with workspace membership, a project in that workspace and permission to create and run pipelines. Reading or writing also requires access to the selected endpoints and external systems. Creating a pipeline does not grant those permissions.
Use a reachable test PostgreSQL database and its SQL client. You need installation/endpoint management access as well as pipeline read/write/execute access. The connection must be reachable from the execution runtime, not only from your laptop.
Prepare the input
CREATE TABLE public.pipeline_demo_issues (issue_id integer PRIMARY KEY, description text NOT NULL);
INSERT INTO public.pipeline_demo_issues VALUES
(1, 'My invoice charged the order twice'),
(2, 'The parcel has not arrived');Install Postgres read in workspace Plugins → Explore using these fields:
| Installation field | Value |
|---|---|
| Host | Your reachable database host |
| Port | 5432, or your configured port |
| Database | The test database containing the fixture |
| User | A database identity with the required access |
| Password | Enter in the protected installation field |
| Ssl | Enable when the database requires TLS |
Check the connection. Create an Ingress endpoint demo_issues in namespace pipelines, Stream contract, target Schema public, Table pipeline_demo_issues. Inspect the two known rows.
Install the classifier
In workspace Plugins → Explore, find Craven Transforms (@craven/transforms), inspect the required operation and install/enable it. Its bundle installation configuration is empty; the operation’s business parameters belong on the transform node. Use the capability installation ID, not the bundle installation ID, for capabilityRefId. The Transforms guide provides optional background.
Choose classification and ensure the deployment’s model runtime is configured. Bundle installation does not configure model credentials. If model execution is unavailable, validation/runtime must report that limitation; another deterministic transform is not an equivalent classifier.
Configure the pipeline
Create Ingress → Transform, with ingress Full load/parser and output issues. Connect its output to classifier input in.
Read demo_issues and classify description into billing or delivery using the installed classification transform. Define billing as invoice, payment or charge issues; delivery as arrival, shipment or parcel issues. Store category, category_confidence and category_reason. Show the configuration before running and keep the result for human review.Select the installed classification capability. Source fields: description. Add labels billing (“Invoice, payment or charge issues”) and delivery (“Arrival, shipment or parcel issues”). Set Output field category, Confidence field category_confidence, Rationale field category_reason, Multi label disabled. Bind input issues, output classified_issues.
This is the node’s transform section inside a complete pipeline document, not a standalone API/MCP action. Resource placeholders must be replaced with the installed IDs.
{
"capabilityRefId": "<CLASSIFICATION_CAPABILITY_UUID>",
"function": "classification",
"input": "issues",
"output": "classified_issues",
"params": {
"sourceFields": [
"description"
],
"labels": [
{
"label": "billing",
"description": "Invoice, payment or charge issues"
},
{
"label": "delivery",
"description": "Arrival, shipment or parcel issues"
}
],
"multiLabel": false,
"outputField": "category",
"confidenceField": "category_confidence",
"rationaleField": "category_reason"
}
}Validate, save and run
Validate the draft and resolve errors. Preflight checks configuration only; it does not produce records or run a model. Save version, inspect the active saved graph and Run once. Open the run and compare the named step's output with the expected result below. Keep the run ID and inspect errors/counts, not only the assistant's response. Programmatic launch uses HTTP POST /api/v1/projects/<PROJECT_ID>/pipelines/<PIPELINE_ID>/execute with pipelines:execute and an Idempotency-Key, or MCP pipeline_execute with pipelineId, projectId and idempotencyKey.
If the Studio result only exposes an artifact handle, open the corresponding output/artifact inspection rather than treating that handle as a row preview. To persist rows externally, add an Egress endpoint and test its write semantics; the examples below do not silently create a destination.
Review the result
Expect issue 1 to be a billing candidate and issue 2 a delivery candidate. Inspect actual labels, confidence and rationale; do not hard-code that expectation as proof. Compare evidence linkage to each source row, and test ambiguous text and empty descriptions. Confidence is the reported assessment, not demonstrated accuracy. Establish quality on a representative labeled dataset before operational use.