Semogram Docs
Forecasting and predictionsCheck what happened

Scores and calibration

Measure predictions against observations without conflating execution and accuracy

Evaluation metrics compare a saved forecast with an observed result under the stored policy. They measure the predictions you evaluated, not every subject in your business. Unknown, ambiguous and unresolvable observations produce no numeric score.

Probability: Brier score

For probability p and actual occurrence y (1 for occurred, 0 for not_occurred), Brier score is (p − y)². Lower is better, with 0 a perfect probability for that observation.

Illustrative failure forecastObservationBrier score
0.7occurred0.09
0.7not_occurred0.49
0.2not_occurred0.04

One surprising event does not prove that a probabilistic forecast was invalid. Evaluate a representative set of predictions made before their outcomes, under the same event definition and horizon.

Calibration

The forecaster evaluation report groups probabilities into ten bins. Each bin includes bounds, count, mean forecast probability and observed occurrence frequency. If many predictions near 0.7 experience the defined event about 70% of the time, that bin is calibrated. A small or empty bin does not establish reliability; its averages may be null.

Read the report in the equipment forecaster's score/evaluation UI or call forecaster_evaluation_report with forecasterId. The response includes sampleSize, meanBrierScore and calibration. No eligible evaluated predictions produces sample size 0 and null mean score, not perfect performance.

Report population limits

The current aggregate covers completed probability predictions for the forecaster and selects the latest eligible occurred/not_occurred evaluation for each. It does not expose filters for version, horizon or period. If a newer evaluation is Unknown, an older eligible binary evaluation can still enter that report. Inspect evaluation history and use an explicit analysis for corrected cohorts rather than treating the report as a latest-status audit.

Versions with different prompts, evidence or windows can share a report. Compare compatible cohorts deliberately, include unresolved/unevaluated counts and avoid presenting selective successful observations as the overall track record.

Other kinds

KindRequired observed valueMetric
NumericNumber in the same units as forecastPathAbsolute error and squared error
CategoricalExact observed value/labelEquality accuracy, 1 or 0
ScenarioDocumented outcome and notesNo built-in numeric scenario score

For numeric/categorical predictions, configure evaluationPolicy.forecastPath to the actual output field. The default fallback $.prediction does not match the simple form's value/category fields. Use prediction_evaluation_create to supply observedValue and inspect the resulting metrics; the simple UI outcome panel does not expose that field.

Build a trustworthy study

Keep each forecast's invocation time, evidence cutoff, event/window definition, version and outcome observations. Exclude post-outcome evidence from historical inputs. Compare against an agreed baseline with the same observation coverage. Document missing outcomes and observation bias. No automatic backtest/training endpoint is implied by these metric functions, and recording evaluations does not automatically retrain the model.