Monitoring predictions
Separate execution health, evidence quality and outcome coverage
Monitor the prediction record, the evidence it used and its eventual evaluation separately. Completed means output was validated and saved; it does not mean the prediction was correct. Pending is an outcome state, not proof that model execution is still running.
You need a Semogram account with project read access. Inspect the actual equipment prediction ID returned by a run or watchlist invocation.
Inspect one prediction
Open the prediction record. Read execution status, subject, horizon end, forecaster version, output, evidence/provenance and available usage. Inspect Outcome separately and follow evaluation history when present.
Read this P-101 prediction and its evaluations. Report execution state separately from outcome state, its pinned versions, actual horizon end, evidence completeness and saved output. Report usage/cost only when available. If it is overdue and Pending, report the diagnostic notes instead of guessing that no failure happened.Call prediction_get with predictionId. Call prediction_evaluation_list with the same predictionId. Use the initialized project connection; on the workspace connection also supply projectId. Inspect the returned data, completeness and warnings. These read tools do not resolve outcomes.
No public API-key prediction-read route is exposed. Use the supported UI or MCP interface.
Read the signals
| Signal | Interpretation | Investigation |
|---|---|---|
| Queued | Accepted but not completed | Delivery/workflow identity; retain prediction ID |
| Running | Execution underway | Evidence query, model stage, external calls and budgets |
| Failed | Execution did not save a successful result | Actual error and any external-operation receipts |
| Cancelled | Invocation stopped | Prior side effects can remain |
| Completed + Pending | Output saved; outcome unsettled | Is the window closed? Is the policy manual or query? |
| Completed + evaluated | An immutable observation was recorded | Review event definition, evidence and evaluator |
| Evidence hasMore/diagnostics | Evidence may be incomplete | Query bounds and saved evidence snapshot |
Evidence retention is bounded; a summary does not establish that all relevant source records were considered. Inspect query bounds and diagnostics. Invocation cutoff metadata does not automatically make every connector apply historical filtering.
Scheduled equipment check
For a daily P-101 watchlist, read lastRunAt and the expected nextRunAt, then locate the actual new prediction. Compare its pinned forecaster/evidence versions with the intended setup. Inspect the seven-day end date and verify the output against its schema.
After the window closes, inspect evaluations. Missing incident records cannot establish not_occurred unless adequate monitoring supports that conclusion. Overlapping daily windows are related observations; do not present them as independent equipment failures.
Usage and quality
Inspect model/external-call counts, token usage, duration and available cost state. Missing estimates are unavailable, not zero cost. Budget failures need configuration/evidence review, not just a higher cap.
Use evaluation metrics for observed quality, with the cohort and version limitations explained in Metrics. Completion rate and resolved-outcome coverage are operational checks; neither is forecast accuracy. No automatic alerts or retraining are implied by the record views.