Supplier Scorecard
Score suppliers from individual deliveries rather than opinion: on-time-in-full, defects, documentation, responsiveness and price, weighted into one comparable number and printed as a supplier review pack. Runs entirely in your browser. Nothing is uploaded.
Version 1.0.0 · Updated Aug 5, 2026
Overview
How to use Supplier Scorecard
The complete in-tool guidance, reproduced here so you can read it before you download.
What this tool does
CM8-78 builds a supplier scorecard out of individual deliveries. You record what actually happened each time something arrived — was it on time, was it all there, was it accepted, was the paperwork right, how did they behave, what did it cost — and the tool turns those events into one weighted number per supplier, a league table, and a short list of who needs a conversation.
The point of doing it this way is that the score cannot be argued with in the way an opinion can. "They are hopeless" is a view. "Eleven of your last fourteen deliveries were late, and four of those were also short" is a fact, and it is the sentence that changes behaviour.
Everything runs inside this single file. There is no account, no upload and no network request, so supplier names, prices and performance never leave the computer you are using. That matters: a scorecard is commercially sensitive to your supplier and, in aggregate, to you.
What counts as one record
One record is one delivery, or one order line if you receive orders in parts and want each part judged separately. Be consistent: if you split some orders into lines and not others, the supplier whose orders you split will have more records and their bad ones will weigh less.
Only record deliveries you actually assessed. A register that contains every problem delivery and only the memorable good ones will produce a score that is far too low, and the supplier will be right to say so. If you cannot assess everything, assess a consistent sample — every delivery in one week a month, say — and say that on the report.
Supplier names are matched without regard to capitals, so "Kestrel Components" and "kestrel components" group together, but a typo or a trading name that differs from the one on the invoice will create a second supplier with a shorter, weaker record. Check the league table for near-duplicates before you circulate anything.
How a delivery is scored
Each delivery produces four component scores, each out of 100, which are then weighted into one figure.
Delivery component = (on-time score + in-full score) ÷ 2 on-time score = 100 if it arrived on or before the agreed date, otherwise 100 − (10 × days late), floored at 0 in-full score = 100 if the quantity was correct, otherwise 0
Timing and completeness are half each, so arriving punctually with the wrong quantity scores 50, not 100. Ten points per day late means a delivery ten or more days late scores zero for timing, however good the excuse. Arriving early does not earn more than 100 — early deliveries cost you storage, and a supplier who ships a week early to hit a target has not done you a favour.
Quality component = (acceptance score + defect score + documentation score) ÷ 3 acceptance score = 100 if accepted without rework, concession or return, otherwise 0 defect score = 100 − (20 × defects found), floored at 0 documentation score = 100 if the paperwork was complete and correct, otherwise 0
Documentation is a full third of quality, which surprises people. It is there because missing certificates, wrong batch records and absent delivery notes generate real work at your end and, in regulated supply chains, can stop you using the goods at all. If paperwork genuinely does not matter for what you buy, tick the box every time and it stops affecting anything.
Responsiveness component = (responsiveness score − 1) × 25 Price component = (price score − 1) × 25
The two judgement scores run 1 to 5 and are mapped so that 1 scores 0 and 5 scores 100. A 1 means the component is worth nothing, not one fifth of something — that is the honest reading of "hard to reach, problems go unanswered". The options carry descriptions rather than bare numbers so that two different buyers score the same behaviour roughly the same way.
Weighted score = (wdelivery × delivery + wquality × quality + wresponse × responsiveness + wprice × price) ÷ (wdelivery + wquality + wresponse + wprice)
Supplier score = mean of the weighted scores of that supplier's assessed deliveries
The supplier score is a plain mean of deliveries. A large delivery and a small one count the same. That is a deliberate choice — it measures reliability, not spend — but it means a supplier who is perfect on small orders and poor on the big one will score better than the money says. The value column is there to keep that visible.
Setting the weights
The four weights decide what "good" means to you, and they should add up to 100. The tool divides by whatever they actually add up to, so a mistake will not break the score, but the headline tile tells you when the sum is not 100 because a set of weights that does not add up is usually a set nobody has agreed.
The defaults — delivery 40, quality 30, responsiveness 15, price 15 — suit a business where a stock-out costs more than a few per cent on the unit price. Move them to match reality:
- Just-in-time manufacturing — push delivery up, often to 50 or more. A late part stops a line.
- Regulated or safety-critical goods — push quality up. A concession is not a small thing.
- Commodity buying with several interchangeable sources — price can justify 25 or 30, because switching is cheap and quality differences are small.
- Single-source or hard-to-replace suppliers — weight price low. You are not going to leave over price, so scoring it heavily just makes the number move without changing any decision.
Agree the weights before you look at the scores, and write them into the report notes. Weights chosen after seeing the results are not a measurement, they are a justification.
On time in full, and why it is harsh
On-time-in-full rate = deliveries that were on time AND in full AND accepted ÷ deliveries assessed × 100
OTIF is deliberately harsher than any of its parts, and it is normal for it to sit well below the on-time figure. That is the point. A delivery that is on time but two pallets short is not a good delivery — somebody still has to chase the balance, re-plan the work and receive it twice. A delivery that is on time and complete but rejected on quality is worse still.
Because all three conditions must hold, OTIF falls faster than any single measure. A supplier at 90 per cent on each of the three, if the failures land on different deliveries, will show roughly 73 per cent OTIF. When you see a large gap between the on-time column and the OTIF column, the problem is not the haulier — it is what is on the vehicle.
Quote OTIF with its basis. "58 per cent OTIF across 12 assessed deliveries in four months" is a statement somebody can check. "Our OTIF is 58 per cent" is not.
The minimum-deliveries rule
The most common way a scorecard misleads is that a new supplier with two flawless deliveries appears at the top of the table and gets more work, while the supplier who has delivered a hundred times and slipped four times sits below them. Two deliveries is not a track record; it is two coin tosses.
So the tool will not rank a supplier until they reach the minimum number of assessed deliveries, set on the Settings tab and defaulting to three. Below that they still appear in the league table, with their percentages shown, but the score column reads "—" and their standing says how many deliveries they still need. They are kept out of the two ranking charts entirely, and they can never appear in the below-threshold review list.
Three is a floor, not a target. If you receive weekly from most suppliers, set it to eight or twelve; the scores become far more stable. Raising the minimum does not hide anything — the deliveries are all still in the register, the defect chart still counts them, and the monthly trend still includes them.
Defects and the defect rate
Defect rate = total defects found ÷ deliveries assessed
The rate is defects per delivery, not defects per unit, because this tool does not record how many units were in each delivery. That is an honest limitation and it matters: a supplier sending pallet loads and a supplier sending single items are not comparable on this measure. Compare like with like, or record the same kind of delivery for both.
Each record also carries a "defects per delivery" figure in the CSV export. For a single delivery it is simply the defect count, because one record is one delivery — it becomes a rate only when averaged across a supplier, which is exactly what the chart and the tiles do.
Zero is a normal answer. A register in which nobody ever records a defect usually means nobody is inspecting, not that nothing is wrong.
Reading the league table
The league table gives each supplier their delivery count, the four percentage measures, their mean responsiveness and price scores and the overall weighted score, ranked. Scored suppliers come first, in score order; suppliers still under the minimum follow, marked as not yet scored.
The total row at the foot is calculated across every delivery in the filter, not by averaging the rows above it. Those two are different numbers whenever suppliers have different delivery counts, and the delivery-level figure is the one that describes what actually happened to you.
Read the columns together. High on-time with low in-full is a picking or capacity problem. High quality with low responsiveness is a supplier who makes good product and is painful to deal with — often worth keeping and worth telling. Low price scores across every supplier in a category usually says more about your specification than about the market.
What is dragging a score down
Any scored supplier below your acceptable score appears in the second table with the components that cost them the most points. The arithmetic is simple:
Points lost on a component = (component weight ÷ total weight) × (100 − component score)
So a supplier averaging 64 on quality, where quality carries 30 of 100 points, loses 10.7 points — and if responsiveness at 33 out of 100 with a weight of 15 loses another 10.0, you now know the meeting is about defects and about answering the phone, not about price. The two largest losses are named; the rest are in the numbers.
The value column shows what that supplier was paid across the deliveries in the current filter. It is not part of the score. It is there so the review can be proportionate: a supplier scoring 58 on 2 per cent of your spend is a different problem from one scoring 72 on a third of it.
Using it in a supplier review
Set the date filter to the period you are reviewing, check that the weights and the acceptable score are the ones you agreed, fill in the report header, and print. The report carries the headline figures, all four charts, both tables and the full delivery register, so the supplier can see every record behind their number rather than just the number.
Send the register in advance. Disputed records are common and usually legitimate — an agreed date that moved, a shortage they told you about, a rejection later overturned. Correct them before the meeting rather than arguing in it. Then talk about the two components in the drag column and agree what changes by when.
What this tool cannot tell you
- Whether the agreed date was reasonable. Lateness is measured against the date you entered. If your own ordering is late, the supplier carries a score for your planning.
- Whether you assessed fairly. The tool cannot see the deliveries you did not record.
- Total cost. The price score is a judgement of competitiveness, not a calculation. Rework, expediting, stock cover held because a supplier is unreliable and the cost of a stop-out are all real and none of them are in here.
- Risk. Financial stability, single-site dependency, sub-tier exposure and compliance are not performance measures and are not scored. A supplier can score 95 and still be the biggest risk you carry.
- Cause. A falling score tells you something changed. It does not say whether the supplier lost a key person, took on too much work, or is reacting to how you buy.
Printing and sharing
Print Report produces a report from whatever the current filter shows: header, the six headline figures, the four charts, both tables, the full register and your closing notes. Print to PDF to circulate it.
The scope line under the title states the filter in force, and the report prints the number of records included. Clear the filters before issuing anything described as the full period, and never send a report filtered to one supplier's bad quarter without saying so on it.
Saving your work
Deliveries, settings and the report header are written to this browser's local storage as you type, and the toolbar shows the time of the last save. That storage belongs to one browser on one computer: another browser, a private window, a second machine or a clean-up tool that clears site data will not have it.
Treat Export .json as the real save — one file containing everything, which Import .json restores anywhere. Export CSV gives you the register for spreadsheet work, including the derived OTIF flag, weighted score and defect figures, and covers every filtered record rather than only those drawn on screen. Reset asks twice, then erases everything this tool has stored. There is no undo.
Accuracy & disclaimer
This tool calculates from what you enter, using the weights you choose. It cannot verify a delivery date, a defect count or a judgement score, and every figure it produces depends on all three being recorded honestly and consistently across suppliers.
Scores are an internal management aid. They are not a contractual measurement unless your contract says the method is, they are not a quality-management certification, and they are not advice on whether to keep, replace or challenge a supplier. Before acting on a score, check the underlying records and the terms you actually agreed.