Decision quality methodology
What we measure—and what the numbers cannot prove.
Decize follows a decision from its original frame to its eventual outcome. This page defines the measures built from that record, the rules that keep hindsight out, and the limits a responsible reader should keep beside every score.
Measurement contract
The record comes before the score.
A number is useful only if the inputs cannot be rewritten after the outcome and the denominator is visible. Decize applies three rules before interpretation begins.
01
Record the expectation
The decision, rationale, owner, confidence, forecast, and review date are captured before the result is known.
02
Wait for a real outcome
Unresolved rows do not count as misses, and elapsed time alone does not turn an open question into an answer.
03
Show the denominator
Rates remain blank below the display floor and forecast summaries carry their sample size and early-sample status.
Metric definitions
One formula, one denominator, one honest reading.
These are the definitions used by the product today. Product thresholds are identified as product thresholds; they are not dressed up as universal statistical laws.
- Outcome closure
Calculation
mature decisions with a recorded outcome ÷ all mature decisionsWhen it appears
Hidden until at least 5 decisions are mature.
Responsible reading
Whether the team is completing the learning loop. It does not say whether the decisions were good.
- On-time execution
Calculation
elapsed deadlines implemented on or before deadline ÷ all elapsed deadlinesWhen it appears
Hidden until at least 5 deadlines have elapsed.
Responsible reading
Whether commitments were completed when promised. A timely decision can still produce a poor outcome.
- Owner coverage
Calculation
decisions with an owner ÷ all decisions in the selected reporting rangeWhen it appears
Hidden until the range contains at least 5 decisions.
Responsible reading
Whether accountability is explicit. It measures presence of an owner, not that person’s performance.
- Brier score
Calculation
mean of (forecast probability − observed result)²When it appears
After 5 resolved forecasts; labelled an early sample below 20.
Responsible reading
Lower is better. Zero means every probability matched the result perfectly; the score alone does not explain why.
- Skill versus base rate
Calculation
1 − (Brier score ÷ Brier score from always predicting the sample base rate)When it appears
With the Brier score; left blank when every result is identical and the baseline is zero.
Responsible reading
Positive is better than the sample’s base-rate forecast, zero is parity, and negative is worse.
- Reliability bins
Calculation
mean predicted probability compared with observed frequency inside four probability bandsWhen it appears
Each band is withheld until it contains at least 5 resolved forecasts.
Responsible reading
Points near the diagonal are better calibrated. Empty bands mean insufficient data, not perfect calibration.
The five-record minimum protects against brittle percentages and very small slices. It is deliberately conservative product behavior, not a claim that n=5 makes an estimate precise. Forecast results below 20 are separately labelled early sample.
Integrity rules
The outcome is not allowed to rewrite the forecast.
These controls make the measures auditable inside a workspace and harder to flatter after the fact.
- 1
The decision-time frame locks
Once a decision is marked made, the question, rationale, assumptions, forecast, expected impact, and reversibility cannot be silently rewritten to fit the result.
- 2
Unresolved forecasts do not score
A probability enters forecast metrics only after its result has been recorded. A forecast that has not reached or received its resolution is excluded rather than treated as wrong.
- 3
AI readings do not become human judgment
Model-inferred confidence is excluded from the confidence-versus-outcome view until a person confirms it. Generated summaries remain proposals until accepted.
- 4
Small groups stay blank
Rates and probability bands use a five-record privacy and display floor. Five is not presented as statistical certainty; scores below 20 resolved forecasts remain explicitly early-sample results.
- 5
Organization learning is the unit
Decize reports shared workflow and outcome measures. It does not publish leaderboards or individual decision-quality scores.
Interpretation limits
Evidence about a workflow, not a verdict on a team.
Decize measures recorded behavior. The results are descriptive, not causal: a higher closure rate does not prove that the product caused better decisions, and a lower Brier score does not identify which process produced it.
Small samples remain uncertain even after the display floor. The mix and difficulty of decisions can change across periods, teams, and organizations, so comparisons need context. Decize does not convert these measures into employee rankings or professional advice.
The outcome-closure chart displays up to 12 mature weekly cohorts and waits for 10 mature weeks before presenting a trend. That ten-week rule is a Decize interpretation guardrail, not a universal threshold for statistical significance.
Primary sources
The ideas behind the product, linked at the source.
External references explain the established methods and terminology. They do not endorse Decize, validate its implementation, or replace the product limits above.
- 1Verification of Forecasts Expressed in Terms of Probability
Glenn W. Brier · Monthly Weather Review · 1950
The original paper behind the mean-squared probability score Decize reports.
- 2Sampling Uncertainty and Confidence Intervals for the Brier Score
Bradley, Schwartz and Hashino · Weather and Forecasting · 2008
Why a score calculated from a small set of forecasts should be treated as an estimate, not a verdict.
- 32015 Letter to Shareholders
Amazon · published 2016
The primary source for the one-way-door and two-way-door decision terminology used in Decize.
- 4Process or Product Monitoring and Control
NIST/SEMATECH e-Handbook of Statistical Methods
Background for reading measurements over time without treating a single point as a trend.
Inspect the product
Read the definitions beside the evidence.
The documentation explains what ships, the public ledger shows real decisions from building Decize, and the contact page reaches the team responsible for both.