Design a baseline yourself
What you will learn Not a chapter to read but one to do. Take one of your projects, write the measurement plan on a single page, and test whether it actually works.
Why now
As measuring and reporting results showed, past cycle time has no retroactive measurement. Postpone this exercise a week and that much of the baseline is gone for good.
The last line is not optional. Filled in by managers alone, every number becomes an estimate.
Step 1 — Nail down where the work starts and ends (10 min)
The most common mistake. Too wide a scope and you cannot tell later what improved.
| Item | Yours |
|---|---|
| Where does it start? | |
| Where does it end? | |
| What is one unit? | (one enquiry / one invoice / one report) |
| How many per month? |
Step 2 — Find numbers you already count (10 min)
Do not invent new ones. Use figures already produced by a system or a ledger. Starting to measure something new changes behaviour, which contaminates the very baseline you are taking.
| Metric | Value today | Where you checked | Verified |
|---|---|---|---|
| Cycle time per item | □ | ||
| Error / rework rate | □ | ||
| Items per month | □ |
Three rows is enough. Three tracked for six months beats ten attempted and none completed.
Step 3 — Window and owner (10 min)
| Item | Yours |
|---|---|
| Measurement window | (long enough to average out weekday and seasonal variance — usually 2–4 weeks) |
| Start date | |
| Who pulls the same numbers after deployment | (a name) |
| Reporting cadence | (monthly / quarterly) |
The third row is the whole table. If a consultant measures before and nobody measures after, there is nothing to compare against. Check now whether that person will still be around after launch.
Step 4 — Test your own plan (10 min)
Run the plan through four questions.
Test 1 — Can it be pulled the same way in three weeks?
From what you wrote under "where you checked", someone else must be able to produce the same number. If not, that is a memory, not a measurement.
Test 2 — Does this number move when things improve?
It is common to pick a metric that does not move even when cycle time drops. Write down now which direction, and by how much, this number should move if the AI works. If you cannot, the metric is wrong.
Test 3 — Did you leave room for the hidden costs?
Without these, "40 min → 12 min" becomes a lie later.
Test 4 — Did you write where the saved time goes?
Ticking the last box is honest. But left as is, the financial effect records as zero. Note now that this is a change management problem, not a measurement one.
What a finished one looks like
Check yourself
1. Why pick metrics you already count rather than new ones?
Answer
Starting to measure something new changes behaviour and contaminates the baseline. A figure already produced by a system is also the only kind you can pull the same way after deployment.
2. Why test "could someone else pull this in three weeks?"
Answer
If they cannot, it is a memory rather than a measurement. A baseline is only worth anything compared by the same method to the post-deployment value; if the method is not reproducible, there is no comparison.
3. Why measure rework rate and review time from the start?
Answer
Without them the improvement is overstated. Cycle time falling from 40 to 12 minutes means little at a 40% rework rate — the real figure is 12 minutes plus the fixing. Added later, there is no baseline to compare against.
You are ready to keep the numbers; next, keep the evidence → Audit trail and observability