Semogram Docs
PluginsTransforms reference

entity-resolution

Configure entity-resolution parameters, input/output bindings and lineage.

Use AI to propose provisional entity clusters; reconcile source identities through the persistent registry before treating them as canonical.

Install

Install the Transforms bundle, then select the entity-resolution capability on a pipeline transform node. Runtime entrypoint: transforms/entity-resolution.ts. The installed capability must be enabled and pass preflight.

Configure

Execution: AI-based. Declared runtime support: supported.

Parameters

FieldTypeRequiredDetails
candidateKeysarrayYesSemantic comparison hints. Differences in these fields do not prevent candidates from being compared. Items: string.
entityTypestringNoHuman name for the entity being resolved, such as patient or customer. Default: "entity".
instructionsstringNoOptional match/merge guidance for ambiguous records.
entityIdFieldstringNoField for a provisional cluster label, never a persistent canonical identity. Default: "resolvedEntityId".
confidenceFieldstringNoField that stores resolution confidence for each row. Default: "resolutionConfidence".
rationaleFieldstringNoField that stores resolution rationale for each row. Default: "resolutionRationale".
canonicalFieldstringNoCompatibility field; always false until a separate persistent registry decision. This transform only emits provisional clusters. Default: "isCanonical".

View the complete contract for nested parameter shapes, defaults, named bindings and lineage.

Bindings

  • Input in: Input (required)
  • Output out: Output

Verify the result

Create a small input with known values that exercise entity-resolution and supply the required parameters above. Validate, save and run the pipeline. Inspect the output records and compare row counts and changed fields with the input. Check the run’s errors and evidence before accepting a conversational summary.

Review the confidence and rationale fields where supplied. Include ambiguous examples and inspect the result; model output is not an observed business fact.

Lineage and changes

Lineage mode: row_preserving. This capability declares one output row per input row in the same order. Verify that downstream evidence still refers to the correct source row.

Changing parameters changes the computation. Review the saved pipeline revision before rerunning it. Publishing a new plugin version and editing node parameters are different changes.

FAQ

Why does validation fail?

Check required parameters, the installed capability version, runtime support and the named input/output bindings. A missing nested field may be visible only in the full contract.

Is this a standalone MCP tool?

It is a plugin capability used by a pipeline, not a new top-level MCP tool with this operation name. An assistant can inspect, validate and execute the pipeline through the MCP reference.

Are its IDs canonical?

No. They are provisional cluster labels. Use persistent identity reconciliation and review before treating a cluster as a canonical business entity.