A number is only as good as the evidence behind it. The PFV toolkit runs on estimates — deliberately, because estimates are cheap and they tell you where to look. But estimates cannot carry an investment case, and pretending otherwise is how methodologies lose the room. This page names the three rungs of confidence — observed, hypothesized, verified — and what each one can honestly support.
Per-step queue time, processing time, and rework rates entered from memory and team judgment. The worksheet locates the waiting. Fast, free, and directionally useful — people are usually right about which steps hurt, and often wrong about how much. Nothing about the cause is known yet.
Honestly supports: deciding where to look in your system-of-record data; ranking which process to examine first; a shared map the team can argue with.
Cannot support: naming a cause, claiming a bottleneck, budget reallocation, or any number a CFO will interrogate. A large queue locates a delay; it does not explain it.
You work the dominant waits through the six mechanisms — capacity, batching, dependency, WIP overload, rework, policy — and name the one the pattern suggests. Citing concrete indicators you actually see raises this to evidence-supported, which is still self-reported: the toolkit cannot read your systems. “Unknown — needs evidence” is a legitimate verdict here.
Honestly supports: choosing which candidate experiment to run first; a conservative, clearly-labelled scenario in the calculator; a conversation with the step’s owner about what they actually see.
Cannot support: presenting the mechanism upstairs as a proven cause. A hypothesis is a question you have not yet answered.
The same map and the same mechanism, re-fed from status-transition timestamps your systems already record: CRM field and stage history (e.g. Salesforce), ERP workflow logs, ticketing and ITSM audit tables. The logs settle both questions at once — how long the step really waited, and what it was really waiting on. One export and a spreadsheet gets you here.
Honestly supports: internal prioritization decisions; SLA redesign; process-owner commitments; before/after comparison when you change the process; the strongest figures the toolkit will model. Where logs and estimates agree within roughly ±30%, your team's judgment is calibrated — that is worth knowing in itself.
Cannot support: certainty, or proof that a fix worked. Verifying the cause is not the same as demonstrating the remedy — and the cost models layered on top (COPQ multipliers, cost-of-delay rates, industry cost benchmarks) are still industry averages, not your organization's numbers.
Once the cause is verified and you have acted on it, per-step flow metrics can be extracted continuously, with your own cost inputs replacing the industry averages: your loaded labor rates, your delay costs, your overhead. Lead time, queue time, and PCE become tracked operational metrics rather than one-time findings.
This is a continuous-improvement and maintenance practice, not a higher grade of confidence in the diagnosis. It is how you re-measure after an intervention — and re-measuring is what earns Improved: the intervention produced measurable operational gains. It is also how you notice when a process that was fixed quietly drifts back.
Still cannot support: certainty. Even measured systems carry surge behavior and model assumptions — which is why honest figures at every level carry ranges, not points.
The ladder is the method: the map tells you where to look, the diagnosis tells you what to ask, and the logs tell you what is true. Every figure in this toolkit is labeled with the rung it stands on — and instrumentation, if you go on to adopt it, tells you whether it stayed true.