Semogram Docs
Forecasting and predictionsRun and operate

Operating outcome resolution

Track due windows, investigate pending evaluations and retain corrections

Outcome operations establish what actually happened after a prediction. They do not rerun the forecast. Each evaluation is immutable, and a later correction is another evaluation version rather than an edit to the original model result.

You need a Semogram account with project access and write permission to record evaluations. Use a Completed typed prediction with its real event definition, invocation and horizon end. Reference actual observations in the same project.

Equipment review queue

For each Equipment failure risk prediction, identify the window and check incident/monitoring evidence for P-101. The event is an unplanned mechanical failure causing at least one hour downtime. Preventive visits do not count.

ConditionDecision
Defined probability event observed inside the windowOccurred may be recorded before window end
Closed window with adequate monitoring and no qualifying eventRecord not_occurred with the supporting explanation
Window still open, no eventKeep Pending
Missing observationsDo not infer non-occurrence
Conflicting evidenceReview; use ambiguous when it is the supported conclusion
Resolution attempt failed technicallyKeep Pending and investigate the job/configuration

Non-occurrence and other outcome states require observedAt at or after horizonEndsAt. Observed time cannot be in the future. Assertion references must exist in the same workspace/project and have been created no later than the observation time.

Manual review

In the prediction's Outcome panel, choose the permitted outcome, actual observation time and notes. Read back the evaluation. The simple form does not expose observedValue or evidence assertion IDs; use prediction_evaluation_create for those richer fields, then prediction_evaluation_list to inspect history.

Outcome review prompt
Review this completed P-101 prediction against the actual incident and monitoring records. State the event definition and exact window, identify supporting project assertion IDs, and explain whether occurred, not_occurred or an unresolved state is justified. Prepare an evaluation for my review; do not use the forecast itself as outcome evidence.

The record an outcome guide gives complete field and MCP examples. There is no public API-key evaluation action.

Query resolution

The forecaster can pin a published outcome-query release in its evaluation policy. This is separate from the evidence query used to forecast. The authorized outcome job considers due Pending work, executes that pinned query and requires one explicit verdict with real assertion evidence IDs and a complete result.

FailureWhat to inspect
No/multiple verdict rowsSubject filtering and one-verdict query contract
Result truncatedBounds, limit and completeness diagnostics
Missing/foreign assertion IDsEvidence exists in this prediction's project
Unavailable releasePinned release and query permissions
Job never runsAuthorized deployment job and credentials
Invalid verdictMapping paths and accepted outcome states

Read outcome notes and evaluations on the actual prediction. A failed attempt leaves Pending with diagnostics, not Unresolvable. A policy edit on today's forecaster does not rewrite an older prediction's pinned policy. Use manual evaluation when a justified observation is available and the old automatic policy cannot resolve it.

The current query tick records event verdicts and evidence but does not forward a numeric/category observedValue for scoring. Use explicit richer evaluations for those kinds. See query-based resolution for the query contract.

Corrections and coverage

If an incident was misclassified, preserve the old evaluation and record a new corrected observation with notes and evidence. Read the latest state and complete evaluation history. The corrected outcome does not retrain the model or rewrite its original evidence/output.

Periodically distinguish predictions awaiting an open window from closed windows lacking evidence, technical resolution failures and completed evaluations. Report those groups separately. A score based only on resolved binary outcomes does not establish that every prediction was evaluable, and overlapping equipment windows need an explicit cohort interpretation.