WCapsuleM8

Measure

On-Time Delivery (OTD)

The percentage of deliveries that arrived when they were promised. The most commonly reported operational measure and the most commonly manipulated, usually without anyone intending to.

Also called: Delivery performance · OTIF (on time in full) · Schedule adherence · DIFOT

Formula

OTD = deliveries on time ÷ total deliveries × 100

Unit
Percentage
Direction
Higher is better
How often
Weekly for operational control, monthly for reporting and trend.
Normally owned by
Production Planner, Despatch Coordinator, Inside Sales / Sales Coordinator

What it is

On-time delivery measures how often the company delivered when it said it would. It is the customer's primary experience of everything the chain does — a customer never sees the schedule, the capacity plan or the material shortage, only whether the goods arrived on the day.

The measure is simple. What makes it treacherous is that it depends on three definitions that are rarely written down: on time against which date, measured at which event, and counted at what level of granularity. Change any of the three and the same month's performance can be reported anywhere from 71% to 98%, all honestly.

The stricter variant, OTIF, requires the delivery to be both on time and complete. It is a harder measure and a more useful one, because a delivery that arrives on the day with 90% of the quantity has not solved the customer's problem.

What each term means

Deliveries on time
Deliveries arriving on or before the promised date. Early delivery counts as on time only if the customer accepts it — many do not.
Total deliveries
All deliveries due in the period, counted consistently — by order line, by consignment, or by order, but not switching between them.
Promised date
The date originally confirmed to the customer, not a date revised later to match what happened.
In full
For OTIF, the delivery must also be complete. A short delivery fails even if it arrives on the day.

One month, four defensible answers

Worked through with real numbers.

Inputs

Order lines due in the month
218
Lines despatched on or before the original promise date
156
Lines despatched on or before a revised date
191
Lines that also arrived complete, against original promise
142
Lines arriving at the customer on time (allowing 2 days transit)
149

Calculation

  1. Against the original promise, measured at despatch:
  2. 156 ÷ 218 × 100 = 71.6%
  3. Against a revised promise, measured at despatch:
  4. 191 ÷ 218 × 100 = 87.6%
  5. Against the original promise, measured at customer receipt:
  6. 149 ÷ 218 × 100 = 68.3%
  7. OTIF — original promise, at receipt, complete:
  8. 142 ÷ 218 × 100 = 65.1%

Result
Between 65.1% and 87.6%, depending entirely on the definition used.

How to read it
All four numbers are arithmetically correct and one of them is honest. The 87.6% is the one most likely to be reported, because rescheduling a slipping order updates the date in the system automatically and the original is overwritten. The customer's own supplier scorecard will show something close to 68% — measured at their gate, against the date they were first given. A 20-point gap between what a supplier believes and what the customer measures is common and is almost always caused by this definition, not by disagreement about the facts.

How to decide what good looks like

We do not publish benchmark figures we cannot source, because the ones in circulation compare businesses using incompatible definitions. This is the method instead.

Fix the definition before fixing the number. Write down which date, which event and which granularity, and do not change any of the three afterwards — most reported improvements in OTD are definition changes.

Measure against the original promise date. If it can be edited when an order slips, the measure reports how good the company is at rescheduling, not at delivering.

Establish your own baseline over three to six months on that fixed definition. It will usually be lower than whatever was previously reported, and that drop is information rather than a deterioration.

Where a customer maintains a supplier scorecard, take their definition as the one that matters commercially. Their number is the one that affects whether you keep the business, however you prefer to measure internally.

Set the target from the gap between your baseline and their measurement, and improve the causes rather than the reporting. Late causes fall into a small number of buckets — material, capacity, quality rework, subcontract queue, and dates that were never achievable — and each has a different owner.

Published industry averages are not useful here, because they aggregate companies using incompatible definitions. A supplier reporting 98% and one reporting 82% may be performing identically.

Where it misleads

  • PitfallMeasuring against a revised date

    Why it happens: When an order slips, the date in the system is updated so the schedule stays current. The original is overwritten and the order records as on time.

    What to do instead: Store the original promise date in a field that cannot be edited after release, and measure against it. Record changes as revisions with reasons.

  • PitfallMeasuring at despatch rather than at delivery

    Why it happens: Despatch is the event the supplier controls and can see.

    What to do instead: Measure both, and know which one the customer uses. Transit time is part of the promise even when it is a carrier's responsibility.

  • PitfallCounting early deliveries as on time

    Why it happens: Early feels better than late, so it is scored as a success.

    What to do instead: Check what the customer wants. Many run scorecards with a delivery window and score early deliveries as failures, because early arrivals consume their storage and cash.

  • PitfallCounting a partial delivery as on time

    Why it happens: Something arrived on the day, and the order line shows movement.

    What to do instead: Use OTIF where the customer needs the full quantity. A delivery of 388 against an order for 400 did not solve their problem.

  • PitfallSwitching granularity between periods

    Why it happens: Order-level counting flatters the number relative to line-level counting, and reporting tools default differently.

    What to do instead: Fix the level of counting and state it on every report. Line level is usually the honest choice.

  • PitfallReporting the number without the causes

    Why it happens: A single percentage fits on a dashboard and provokes no follow-up questions.

    What to do instead: Report the failures by cause every month. The percentage tells you there is a problem; only the causes tell anyone what to do on Monday.

How it gets gamed

Rarely dishonestly. Mostly these are things a reasonable person does when a number becomes a target.

  • Editing the promise date when an order is going to be late.
  • Quoting generous lead times so that hitting them is easy — perfect OTD with an uncompetitive lead time.
  • Shipping partial quantities on the due date to stop the clock.
  • Counting at order level rather than line level so one on-time line rescues a whole order.
  • Excluding orders held for customer-caused reasons without stating how many were excluded.
  • Shipping early into a customer's window they did not want, and scoring it as success.

Where it fits

Common questions

What is the difference between OTD and OTIF?

OTD asks only whether the delivery arrived on time. OTIF requires it to be on time and complete. OTIF is the harder measure and the more honest one, because a delivery arriving on the promised day with 90% of the quantity has not met the customer's need. Most customer scorecards use OTIF or something close to it.

Why is my customer's delivery score so much lower than mine?

Nearly always a definition difference rather than a factual one. They measure arrival at their gate against the date they were first given, counted per line, requiring full quantity. Suppliers commonly measure despatch against the current system date, per order. A twenty-point gap between the two is unremarkable, and the customer's version is the one that decides whether you keep the business.

Should the promise date ever be changed?

The commitment can and sometimes must change — telling the customer early is right. But the original should never be overwritten, because that is the thing performance is measured against. Keep the original, add a revision with a reason, and make the number of revisions visible; it is often more revealing than the OTD figure itself.

Is early delivery on time?

Only if the customer says so. Many large buyers operate a delivery window and score anything outside it as a failure in either direction, because early arrivals consume their warehouse space and pay for stock they do not need yet. Check the scorecard rules before assuming early is safe.

What is a good on-time delivery percentage?

Unanswerable without the definition, which is why published averages should be ignored. A supplier reporting 98% measured at despatch against revised dates may be performing worse than one reporting 82% measured at receipt against original promises. Fix your definition, establish your baseline, and compare only to yourself and to the customer's scorecard.

Where do late deliveries actually come from?

In most make-to-order businesses, a small number of recurring causes: material not available when planned, capacity lost at the constraint, rework after a quality failure, subcontract queue times, and dates that were never achievable when they were promised. The last one is the largest and the least often recorded, because it is created at quotation and only discovered at delivery.

Tools that calculate this

They do the arithmetic and show the workings. If you only want the method, everything you need is on this page.

Reviewed 2026-08-15. The formulas behind every CapsuleM8 tool are published in the methods reference.