Operating outcome resolution
Track due windows, investigate pending evaluations and retain corrections
Outcome operations establish what actually happened after a prediction. They do not rerun the forecast. Each evaluation is immutable, and a later correction is another evaluation version rather than an edit to the original model result.
You need a Semogram account with project access and write permission to record evaluations. Use a Completed typed prediction with its real event definition, invocation and horizon end. Reference actual observations in the same project.
Equipment review queue
For each Equipment failure risk prediction, identify the window and check incident/monitoring evidence for P-101. The event is an unplanned mechanical failure causing at least one hour downtime. Preventive visits do not count.
| Condition | Decision |
|---|---|
| Defined probability event observed inside the window | Occurred may be recorded before window end |
| Closed window with adequate monitoring and no qualifying event | Record not_occurred with the supporting explanation |
| Window still open, no event | Keep Pending |
| Missing observations | Do not infer non-occurrence |
| Conflicting evidence | Review; use ambiguous when it is the supported conclusion |
| Resolution attempt failed technically | Keep Pending and investigate the job/configuration |
Non-occurrence and other outcome states require observedAt at or after horizonEndsAt. Observed time cannot be in the future. Assertion references must exist in the same workspace/project and have been created no later than the observation time.
Manual review
In the prediction's Outcome panel, choose the permitted outcome, actual observation time and notes. Read back the evaluation. The simple form does not expose observedValue or evidence assertion IDs; use prediction_evaluation_create for those richer fields, then prediction_evaluation_list to inspect history.
Review this completed P-101 prediction against the actual incident and monitoring records. State the event definition and exact window, identify supporting project assertion IDs, and explain whether occurred, not_occurred or an unresolved state is justified. Prepare an evaluation for my review; do not use the forecast itself as outcome evidence.The record an outcome guide gives complete field and MCP examples. There is no public API-key evaluation action.
Query resolution
The forecaster can pin a published outcome-query release in its evaluation policy. This is separate from the evidence query used to forecast. The authorized outcome job considers due Pending work, executes that pinned query and requires one explicit verdict with real assertion evidence IDs and a complete result.
| Failure | What to inspect |
|---|---|
| No/multiple verdict rows | Subject filtering and one-verdict query contract |
| Result truncated | Bounds, limit and completeness diagnostics |
| Missing/foreign assertion IDs | Evidence exists in this prediction's project |
| Unavailable release | Pinned release and query permissions |
| Job never runs | Authorized deployment job and credentials |
| Invalid verdict | Mapping paths and accepted outcome states |
Read outcome notes and evaluations on the actual prediction. A failed attempt leaves Pending with diagnostics, not Unresolvable. A policy edit on today's forecaster does not rewrite an older prediction's pinned policy. Use manual evaluation when a justified observation is available and the old automatic policy cannot resolve it.
The current query tick records event verdicts and evidence but does not forward a numeric/category observedValue for scoring. Use explicit richer evaluations for those kinds. See query-based resolution for the query contract.
Corrections and coverage
If an incident was misclassified, preserve the old evaluation and record a new corrected observation with notes and evidence. Read the latest state and complete evaluation history. The corrected outcome does not retrain the model or rewrite its original evidence/output.
Periodically distinguish predictions awaiting an open window from closed windows lacking evidence, technical resolution failures and completed evaluations. Report those groups separately. A score based only on resolved binary outcomes does not establish that every prediction was evaluable, and overlapping equipment windows need an explicit cohort interpretation.