Decision quality methodology

What we measure—and what the numbers cannot prove.

Decize follows a decision from its original frame to its eventual outcome. This page defines the measures built from that record, the rules that keep hindsight out, and the limits a responsible reader should keep beside every score.

Measurement contract

The record comes before the score.

A number is useful only if the inputs cannot be rewritten after the outcome and the denominator is visible. Decize applies three rules before interpretation begins.

  1. 01

    Record the expectation

    The decision, rationale, owner, confidence, forecast, and review date are captured before the result is known.

  2. 02

    Wait for a real outcome

    Unresolved rows do not count as misses, and elapsed time alone does not turn an open question into an answer.

  3. 03

    Show the denominator

    Rates remain blank below the display floor and forecast summaries carry their sample size and early-sample status.

Metric definitions

One formula, one denominator, one honest reading.

These are the definitions used by the product today. Product thresholds are identified as product thresholds; they are not dressed up as universal statistical laws.

Outcome closure

Calculation

mature decisions with a recorded outcome ÷ all mature decisions

When it appears

Hidden until at least 5 decisions are mature.

Responsible reading

Whether the team is completing the learning loop. It does not say whether the decisions were good.

On-time execution

Calculation

elapsed deadlines implemented on or before deadline ÷ all elapsed deadlines

When it appears

Hidden until at least 5 deadlines have elapsed.

Responsible reading

Whether commitments were completed when promised. A timely decision can still produce a poor outcome.

Owner coverage

Calculation

decisions with an owner ÷ all decisions in the selected reporting range

When it appears

Hidden until the range contains at least 5 decisions.

Responsible reading

Whether accountability is explicit. It measures presence of an owner, not that person’s performance.

Brier score

Calculation

mean of (forecast probability − observed result)²

When it appears

After 5 resolved forecasts; labelled an early sample below 20.

Responsible reading

Lower is better. Zero means every probability matched the result perfectly; the score alone does not explain why.

Skill versus base rate

Calculation

1 − (Brier score ÷ Brier score from always predicting the sample base rate)

When it appears

With the Brier score; left blank when every result is identical and the baseline is zero.

Responsible reading

Positive is better than the sample’s base-rate forecast, zero is parity, and negative is worse.

Reliability bins

Calculation

mean predicted probability compared with observed frequency inside four probability bands

When it appears

Each band is withheld until it contains at least 5 resolved forecasts.

Responsible reading

Points near the diagonal are better calibrated. Empty bands mean insufficient data, not perfect calibration.

The five-record minimum protects against brittle percentages and very small slices. It is deliberately conservative product behavior, not a claim that n=5 makes an estimate precise. Forecast results below 20 are separately labelled early sample.

Integrity rules

The outcome is not allowed to rewrite the forecast.

These controls make the measures auditable inside a workspace and harder to flatter after the fact.

  1. 1

    The decision-time frame locks

    Once a decision is marked made, the question, rationale, assumptions, forecast, expected impact, and reversibility cannot be silently rewritten to fit the result.

  2. 2

    Unresolved forecasts do not score

    A probability enters forecast metrics only after its result has been recorded. A forecast that has not reached or received its resolution is excluded rather than treated as wrong.

  3. 3

    AI readings do not become human judgment

    Model-inferred confidence is excluded from the confidence-versus-outcome view until a person confirms it. Generated summaries remain proposals until accepted.

  4. 4

    Small groups stay blank

    Rates and probability bands use a five-record privacy and display floor. Five is not presented as statistical certainty; scores below 20 resolved forecasts remain explicitly early-sample results.

  5. 5

    Organization learning is the unit

    Decize reports shared workflow and outcome measures. It does not publish leaderboards or individual decision-quality scores.

Interpretation limits

Evidence about a workflow, not a verdict on a team.

Decize measures recorded behavior. The results are descriptive, not causal: a higher closure rate does not prove that the product caused better decisions, and a lower Brier score does not identify which process produced it.

Small samples remain uncertain even after the display floor. The mix and difficulty of decisions can change across periods, teams, and organizations, so comparisons need context. Decize does not convert these measures into employee rankings or professional advice.

The outcome-closure chart displays up to 12 mature weekly cohorts and waits for 10 mature weeks before presenting a trend. That ten-week rule is a Decize interpretation guardrail, not a universal threshold for statistical significance.

Primary sources

The ideas behind the product, linked at the source.

External references explain the established methods and terminology. They do not endorse Decize, validate its implementation, or replace the product limits above.

  1. 1
    Verification of Forecasts Expressed in Terms of Probability

    Glenn W. Brier · Monthly Weather Review · 1950

    The original paper behind the mean-squared probability score Decize reports.

  2. 2
    Sampling Uncertainty and Confidence Intervals for the Brier Score

    Bradley, Schwartz and Hashino · Weather and Forecasting · 2008

    Why a score calculated from a small set of forecasts should be treated as an estimate, not a verdict.

  3. 3
    2015 Letter to Shareholders

    Amazon · published 2016

    The primary source for the one-way-door and two-way-door decision terminology used in Decize.

  4. 4
    Process or Product Monitoring and Control

    NIST/SEMATECH e-Handbook of Statistical Methods

    Background for reading measurements over time without treating a single point as a trend.

Inspect the product

Read the definitions beside the evidence.

The documentation explains what ships, the public ledger shows real decisions from building Decize, and the contact page reaches the team responsible for both.