Record an outcome
Evaluate a completed prediction using observed events, values and evidence
An evaluation records an observation for a completed typed prediction. It is a new immutable version containing the prediction reference, policy snapshot, outcome, observation time, evidence assertion IDs, evaluator and metrics. It does not alter the original probability, prompt or source evidence.
Establish the equipment event
For Equipment failure risk, occurrence means an unplanned mechanical failure causing at least one hour downtime within the run's window. Preventive maintenance is excluded. Keep equipment identity, incident time, downtime duration and the distinction between observation and a model's interpretation clear.
You need a Semogram account with project access and permission to record an evaluation. The prediction must be Completed and carry a forecast kind/output and horizon end. Use actual incident/monitoring evidence; do not label the synthetic setup rows as proof of a future failure.
Choose the outcome
| Outcome | Meaning |
|---|---|
| occurred | The defined event occurred in the prediction window |
| not_occurred | The closed window was observed sufficiently to establish non-occurrence |
| unknown | The observation does not establish either result |
| ambiguous | Available evidence leaves multiple plausible interpretations |
| unresolvable | A review establishes that the outcome cannot be resolved |
Pending means no settled evaluation yet. Missing evidence after seven days is not automatically not_occurred. A query/configuration failure is not proof of unresolvability.
When recording is allowed
A probability event can be recorded as occurred after invocation but before the horizon ends: its early occurrence settles the event. Other states and other forecast kinds require observedAt at or after horizonEndsAt. observedAt cannot precede invocation for an early probability occurrence or be in the future.
If you reference assertions, each ID must belong to the same workspace/project and have been created no later than observedAt. This check does not itself establish that every assertion proves the equipment event; inspect its content, provenance and review history.
Record and read back
Open the completed prediction's Outcome panel. Choose the permitted Outcome, actual Observed at time and Notes, then Record evaluation. Before a probability window closes, the UI offers only occurrence. Its simple panel does not expose numeric/categorical observedValue or evidence assertion-ID input; use MCP for those richer evaluations.
Ask a connected assistant to prepare an evaluation for the actual prediction and incident: defined failure occurred at the documented time, with the real project assertion IDs and your explanation. Review the event/window/evidence and time before prediction_evaluation_create. Ask it to list the saved evaluations afterward. Do not ask it to invent an observation from the forecast itself.
Replace prediction UUID, observation time, evidence IDs and key. The fixed time below is illustrative and valid only when it describes an actual observation allowed by the target run's window.
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "prediction_evaluation_create",
"arguments": {
"predictionId": "<COMPLETED_PREDICTION_UUID>",
"outcomeState": "occurred",
"observedAt": "2026-10-06T14:00:00Z",
"evidenceAssertionIds": [
"<ACTUAL_PROJECT_ASSERTION_UUID>"
],
"evaluatorKind": "human",
"notes": "P-101 had an unplanned mechanical failure with two hours downtime during this prediction window.",
"idempotencyKey": "<UNIQUE_EVALUATION_KEY>"
}
}
}Use prediction_evaluation_list to inspect the immutable versions, and prediction_get to inspect the latest outcome state. Recording updates the prediction's outcome, independently of its completed execution state.
Numeric and category observations
For numeric downtime, use evaluationPolicy.forecastPath $.value and supply observedValue as a number in hours. For a categorical output, use forecastPath $.category and supply the observed label. A human review may require additional operational rules to derive those values; a binary occurred/not_occurred flag alone does not supply them.
Correct an evaluation
Retain the older evaluation and record a new one with the corrected observation and notes. New evaluation versions preserve the history. Recording an outcome or a correction does not automatically retrain a model or rerun the earlier prediction.