A supplier scorecard is worth building for one reason only: it decides who gets the next order. Every other use — the annual review document, the procurement dashboard, the evidence folder for an audit — is decoration. The test of whether yours works is simple. If a supplier’s score changes and nothing about your next purchase order changes, you do not have a rating system, you have a spreadsheet. That is why the design has to start from allocation and work backwards to the metrics, because a score that is attached to volume gets taken seriously by the supplier and a score that is attached to nothing gets negotiated.
This guide sets out the whole mechanism: the five dimensions worth measuring and how to weight them, the specific metrics inside each one and the thresholds that make them meaningful, where the data has to come from if the system is to survive past the second quarter, why most scorecards die from having too many metrics, how to convert a weighted score into an allocation split, the agenda for a quarterly review that produces decisions rather than discussion, how to handle a falling score without turning the relationship into a disciplinary process, and the failure modes to design around. Production reference for this guide — QUANZHOU JUNYUAN BAGS, custom waterproof bags since 2014, 4,950 m² SGS-verified facility, MOQ 500 pieces per style, sampling in 6–10 working days, bulk in 35–50 days, FOB Xiamen.



The scorecard exists to move volume, not to rank people
Most vendor rating systems are built as measurement exercises and then quietly abandoned, because measurement without consequence is administrative cost. A working supplier scorecard is a decision rule: given two approved suppliers both capable of making the product, it says which one gets this order and in what proportion. That single property does the work. It tells the supplier exactly what behaviour is rewarded, it removes the argument from the allocation conversation, and it gives you a defensible reason for a decision you would otherwise make on gut feel. A vendor rating system that produces a number and no decision is worse than no system, because it creates the appearance of rigour while the real decision is still being made informally.
The consequence has to be real and it has to run in both directions. A high score earns a larger share, earlier visibility of the range plan, and first refusal on new styles. A low score earns a smaller share and a defined remediation period — not a penalty clause, not a public ranking, and not a threat. Suppliers respond to allocation because allocation is revenue; they respond to scorecards only insofar as the scorecard changes allocation.
It also has to be transparent. Share the metric definitions, the weights and the thresholds with the supplier before the first scoring period, and share the score itself afterwards. A rating computed privately and revealed only when it is bad reads as a weapon, and a supplier that expects to be judged unfairly will optimise for the appearance of performance rather than the substance. Our supplier audit checklist covers the qualification side of the same relationship; the scorecard is what happens after qualification, every quarter, for as long as you buy.
Five dimensions and the weights that make them behave
Resist the temptation to measure everything. Five dimensions cover everything that actually decides whether a supplier is worth more volume, and each one needs a weight that reflects how fast it costs you money when it fails. Quality failures are the most expensive and the slowest to detect. Delivery failures are visible immediately and damage your own promises. Responsiveness determines how quickly the other four get fixed. Cost matters but is the easiest to measure elsewhere. Compliance is low-frequency and catastrophic rather than routine.
| Dimension | Suggested weight | Why that weight | Primary metric | Data source |
|---|---|---|---|---|
| Quality | 35 | A defect reaches your customer and costs returns, reviews and warranty | Defects per 1,000 units at AQL inspection | Inspection reports, one per shipment |
| Delivery | 25 | A late order breaks your launch, your listing or your own commitment | On time in full, against the confirmed date | PO date versus receipt date |
| Responsiveness | 15 | Everything else improves or decays at the speed of communication | Acknowledgement and quote turnaround | Email or portal timestamps |
| Cost | 15 | Important, but already visible in every quotation | Normalised landed competitiveness | Your own quote comparison sheet |
| Compliance | 10 | Low frequency, but a failure blocks a market entirely | Audit and test report currency | Certificate expiry dates |
Two rules about weights. Keep them stable for at least four quarters, because a weight that moves is a signal nobody can act on — the supplier cannot tell whether its score changed because it changed or because you did. And adjust them only for a structural reason you can state in one sentence, such as a new market requiring tighter compliance evidence, or a shift to a channel where late delivery costs more than it used to.
Scoring itself should be simple enough to compute on one page: each metric scores 0 to 5 against a published threshold, the dimension score is the mean of its metrics, and the total is the weighted sum converted to a percentage. Anything more elaborate than that will not be maintained, and a system that is not maintained is worse than none because people keep quoting last year’s number as if it were current.
Quality: measure what the inspection already produces
Quality is the highest-weighted dimension and the one most often measured badly, typically as a vague satisfaction judgement. It does not need to be. Every pre-shipment inspection already produces the data: the number of units inspected, the defect count by classification, and the pass or fail outcome. Capture three numbers per shipment and the metric computes itself.
| Metric | Definition | Score 5 | Score 3 | Score 0 |
|---|---|---|---|---|
| Defects per 1,000 units | Major and minor defects found at AQL inspection, per thousand units inspected | Under 8 | 8 to 25 | Over 40 or any critical defect |
| First-pass inspection pass rate | Share of shipments passing inspection without rework | Over 95% | 85% to 95% | Under 75% |
| Customer-facing defect rate | Returns or warranty claims attributed to manufacturing, per thousand sold | Under 5 | 5 to 15 | Over 25 |
| Repeat defect rate | Share of defects that appeared in a previous shipment of the same style | Zero | One repeat | Two or more repeats |
The fourth row is the one worth adding deliberately. A defect that appears once is a manufacturing event; the same defect appearing twice is a process that has not been fixed, and it tells you something about corrective action discipline that no single inspection result will. Track it by root cause rather than by symptom — "weld leak at the base seam" rather than "leak" — or you will never see the repetition.
Set the sampling basis once and never change it per shipment. Our AQL sampling guide sets out the standard levels and how to read them, and consistency matters more than strictness here: a defect rate computed on a different sample size each quarter measures your sampling plan rather than the supplier. If you also run statistical process control on a style, the process capability data from our SPC overview makes a better leading indicator than any inspection result. The underlying framing — prevention, appraisal and failure cost — is set out clearly by the American Society for Quality, and it is worth agreeing internally before you set thresholds, because it determines whether a defect caught late counts as a quality problem or a cost problem.
Delivery: score against the date that was confirmed, not the date that was hoped
On time in full is the only delivery metric that matters, and it is routinely corrupted by measuring against the wrong date. If the supplier quoted 35 days, you asked for 28, and the supplier agreed to 28, then 28 is the confirmed date and 33 days is late. If nobody confirmed anything, the score is meaningless and the fix is a booking process, not a metric. The discipline of confirming a date in writing before the order is placed is what makes delivery measurable at all.
Measure in full as well as on time, because partial shipment is the most common way a delivery problem is disguised. A shipment that arrives on the confirmed date with 80 percent of the quantity is not 80 percent on time — it is late, since you cannot sell 80 percent of a range. Score OTIF as a binary per purchase order line and count partials as failures.
- Confirm the delivery date in writing at order placement, and record it as the baseline for that line.
- Score on time in full at line level, not at order level, so a mixed order does not hide a failed line.
- Record the reason code for every miss — material delay, capacity, quality rework, documentation, or your own late approval.
- Exclude misses caused by your own late approval from the score, but track them separately, because they are your problem and they distort the picture if ignored.
- Track the production-side calendar too: our lead time guide shows which parts of the cycle are genuinely controllable.
One nuance that improves the metric considerably: separate predictability from speed. A supplier that consistently delivers in 42 days and says so is often more useful than one that promises 30 and delivers anywhere between 28 and 45, because you can plan around a slow number and you cannot plan around a variable one. If you want to capture this, score the standard deviation of delivery performance against the confirmed date, not just the mean.
Responsiveness: the dimension nobody measures and everybody complains about
Responsiveness feels subjective and is in fact the easiest dimension to instrument, because every interaction leaves a timestamp. It also predicts the other four: a supplier that acknowledges an enquiry in four hours and quotes in three days is a supplier whose quality problem will be told to you early, whose delivery slip will be flagged in advance rather than discovered, and whose compliance documents will be current. Communication latency is a leading indicator of everything else.
- Enquiry acknowledgement time, measured from your email or portal message to a substantive reply — not an automatic receipt. Target under one business day.
- Quotation turnaround, measured from complete technical pack to a priced quotation. Target under five working days for an existing style, ten for a new one.
- Sample dispatch time against the committed sampling window, and whether the commitment was met without chasing.
- Escalation response: how long a confirmed problem takes to reach a proposed corrective action with a date.
- Documentation response: how long certificates, test reports or declarations take to produce on request.
Record these from your own systems, not from the supplier’s account of itself. If you use a shared portal, the timestamps are free; if you use email, a simple log maintained by whoever places the orders takes minutes a week. The important part is that the number exists without anyone having to form an opinion, because opinion-based scores are the ones suppliers argue with and the ones that quietly stop being updated.
Cost: competitiveness, not the lowest number
Cost is weighted at fifteen percent precisely because it is the dimension you already manage continuously through quotation. Scoring it on the raw unit price is a mistake for two reasons: it punishes a supplier for taking on your difficult, low-volume or highly customised styles, and it rewards a supplier for quoting narrowly and adding the rest later. Score competitiveness on a normalised basis instead — same Incoterm, same quantity, same specification, same validity window, converted to landed cost.
Then score three behaviours rather than one number. Price stability: how often a confirmed price changes before shipment, and whether the supplier honours a quoted price on a reorder. Cost transparency: whether a requested cost breakdown is provided in a usable form, line by line. And cost initiative: whether the supplier brings unsolicited cost reduction proposals with evidence attached, which our companion guide on value analysis and value engineering describes in full. A supplier that brings two validated cost ideas a year is worth more than one that is two percent cheaper and never volunteers anything.
This is also where a normalisation method is not optional. Different suppliers quote on different bases, and the cheapest quote on the page is frequently the narrowest in scope. Compare like for like first — the same exercise our quotation comparison guide walks through — then score what remains.
Compliance: a currency date, not a certificate folder
Compliance is weighted low because it fails rarely, and scored pass or fail because when it fails the consequence is not a bad quarter but a blocked market. The metric is not whether the supplier has ever been audited; it is whether each required document is current today. That turns compliance into a small set of expiry dates, which is a far better thing to manage than a folder of PDFs.
| Item | What to record | Failure consequence |
|---|---|---|
| Social compliance audit | Standard, date, and expiry of the current report | Retailer programmes may delist the product |
| Quality management certification | Certificate number and surveillance audit date | Some channels require it as a gate |
| Restricted substances testing | Test report date per material, and scope | Market access; recall exposure |
| Material declarations | Current for the construction actually being produced | A changed material with an old declaration is a real risk |
| Origin and classification evidence | Documents supporting duty treatment | Unexpected duty bills and clearance delay |
The trap in this dimension is drift: the audit is current, but a material was substituted eight months ago and the restricted substances report still describes the old one. Score the document against the current bill of materials, not against the last time anyone checked. Our pages on social compliance audits and restricted substances testing explain what each document actually evidences and how long it stays meaningful. For the management-system side, the ISO standards catalogue lists the certifications most suppliers cite, and the two fields worth capturing are the certificate number and the date of the last surveillance audit rather than the issue date.
Data sources: if a person has to type it, it will stop
This is the practical reason scorecards die, and it is worth designing against explicitly. A metric that requires someone to remember, judge or type something will be maintained for two quarters and then quietly abandoned, after which the scorecard keeps circulating with stale numbers and nobody notices. Every metric therefore needs a source that produces the number as a by-product of work that was happening anyway.
| Metric | Automatic source | Cost to collect | Risk if manual |
|---|---|---|---|
| Defect rate | Inspection report issued for every shipment | Zero; the report exists regardless | Judgement replaces counting |
| On time in full | Purchase order date and goods receipt date in your system | Zero if both are recorded | Negotiated after the fact |
| Acknowledgement time | Email or portal timestamps | Near zero with a shared portal | Subjective and arguable |
| Quote turnaround | Date the technical pack was sent versus date received | One field per RFQ | Forgotten on urgent RFQs |
| Certificate currency | Expiry date in a tracked register | One update per certificate | Discovered at the worst moment |
| Cost initiative count | Log of proposals received with evidence | Low | Never recorded at all |
Two structural habits make this durable. Put the collection into a process that already exists — the inspection report, the goods receipt, the RFQ record — rather than creating a new form. And assign one named owner for the whole scorecard, not one owner per metric. A metric with a shared owner has no owner, and that is usually the first thing to go quiet.
A useful sanity test: ask the owner to produce last quarter’s score in under ten minutes without asking anyone for anything. If they cannot, the data sources are wrong, and no amount of metric design will save it.
The most common failure is too many metrics, and nobody reads them
The failure mode that kills scorecards is not disagreement about weights or bad data. It is length. A scorecard built by committee accumulates metrics because each stakeholder adds the two things they care about, and eighteen months later it has forty-one rows, runs to six pages, and is read by nobody except during an audit. At that point it has negative value: it consumes maintenance effort and produces no decision.
- Cap it at twelve to fifteen metrics across five dimensions. If a new metric is added, an existing one has to be removed.
- Keep it to one page. If it does not fit on a single screen or a single sheet, it will not be discussed in a meeting, and a document that is not discussed does not drive decisions.
- Every metric must have a defined threshold and a defined action at each score. A metric without an action attached is decoration.
- Report the trend, not just the level. Three quarters of direction tells you whether a problem is being fixed; one quarter tells you very little.
- Show the supplier the same one page. If the document is too long to share, it is too long.
There is a second failure worth naming because it is subtler: metrics that measure what is easy rather than what matters. Inspection pass rate is easy to collect and is a reasonable metric, but customer-facing defect rate is what actually costs you money, and the two diverge whenever inspection is catching the wrong things. If you can only maintain five metrics, pick the five closest to your own cost, not the five easiest to count.
Turning a score into an allocation
This is the step most scorecards skip, and it is the only step that changes behaviour. The rule should be published, mechanical and boring: a score band maps to a share of the next period’s volume, and the mapping does not move because someone has a preference. Boring is the goal — the moment allocation becomes a discussion, the scorecard stops being the reason for it.
| Band | Weighted score | Allocation share | Additional consequence | Review trigger |
|---|---|---|---|---|
| A | 90 to 100 | 60% to 70% | First refusal on new styles; earlier visibility of the range plan | Annual, unless something changes |
| B | 75 to 89 | 25% to 35% | Eligible for new styles; one improvement target set | Quarterly |
| C | 60 to 74 | 0% to 10% | Defined remediation plan with dates; no new styles | Monthly until it moves |
| D | Below 60 | 0% | No orders while remediation is open | Re-qualification required |
Three design points make this work in practice. Keep a floor of real volume with your second supplier even when its score is lower, because continuity is worth more than the optimisation — our guide to second source qualification explains why a dormant second source is not a second source. Never allocate one hundred percent to a single supplier on the strength of a score alone, because the scorecard measures performance, not solvency, fire, or a regional power restriction. And state the allocation before the scoring period starts, so the supplier knows the rules of the game it is playing.
Allocation also has to account for capability, which the scorecard deliberately does not measure. A supplier that scores 88 but cannot weld the construction you need does not get the order regardless of its band. Treat score as the tiebreaker among qualified suppliers, never as a substitute for qualification.
The quarterly review agenda that produces decisions
A review without an agenda becomes a status update, and a status update is what suppliers learn to survive. Sixty minutes with a fixed structure produces more than ninety minutes without one, and the structure should end with decisions recorded rather than actions discussed. This is the agenda that works.
- Five minutes: confirm the score. Numbers only, no discussion. Any dispute is taken offline against the published definition, not argued in the meeting.
- Fifteen minutes: quality. Every defect category with a root cause and a corrective action with a date. Check whether last quarter’s actions closed.
- Ten minutes: delivery. Every miss with its reason code, and whether the confirmed date was realistic in the first place.
- Ten minutes: forward plan. Share the next two quarters of likely volume by style, so capacity can be planned rather than hoped for.
- Ten minutes: cost and compliance. Any price movement and its driver, any document nearing expiry, any cost proposal with evidence.
- Ten minutes: agree and record. One improvement target each, with a measurable definition and a date, plus the allocation for next quarter.
Two things make this agenda effective. Sharing the forward plan is the highest-value item on it, because it converts the supplier from an order-taker into a planner and is the single biggest lever on delivery performance — capacity booked early is capacity delivered. And recording the allocation decision in the same document as the score closes the loop: the supplier can see, in one place, what the score was and what it bought.
Hold the review on the same cadence regardless of performance. Cancelling a review because things are going well teaches the supplier that the process only runs when there is a problem, which makes the next review feel like a disciplinary meeting rather than a planning one.
When a score falls: remediation, not punishment
A falling score is information, and the instinct to convert it into a penalty is usually counterproductive. Penalties create an incentive to manage the metric rather than the cause — to ship the easy styles, to negotiate the inspection level, to argue about whether a defect is major or minor. Remediation creates an incentive to fix the process, provided it is structured properly and the volume consequence is already known.
A workable remediation has four parts: a specific measurable problem rather than "quality is down"; a root cause the supplier states in its own words; a corrective action with a date and an owner on the supplier side; and a verification method that will show whether it worked, usually the next two shipments at the same or tighter inspection. Anything less specific will produce a polite letter and no change.
Equally important is checking your own contribution before concluding. Late approvals, changed specifications after quotation, unclear technical packs and unrealistic confirmed dates cause a measurable share of delivery and quality failures, and a supplier that is being scored on your own variability will disengage from the process entirely. Track your own reason codes honestly — if a third of delivery misses are yours, that is the highest-return item in the whole review. Our capacity planning guide covers the seasonality that makes unrealistic dates most likely.
Designing out the ways the system gets gamed
Every metric can be gamed, and the interesting question is not whether a supplier will try but which games your design invites. Inspection pass rate invites negotiating the AQL level. Defect rate invites classifying defects down. On time in full invites the supplier refusing to confirm a date, or confirming a very long one. Quote turnaround invites a fast, incomplete quotation. Each has a structural fix that is better than policing.
- Fix the sampling level and the defect classification in writing at the start of the relationship, and change it only by agreement.
- Count any shipment not confirmed against a date as a delivery failure, so refusing to commit is not a way to avoid scoring.
- Score quotation completeness alongside quotation speed, so a fast partial quote does not beat a slower complete one.
- Audit the classification periodically: have someone re-read a sample of inspection reports and check the defect categories independently.
- Keep at least one metric the supplier cannot influence directly, such as customer-facing defect rate, so the overall score cannot be managed purely from the supplier side.
The deeper protection is a relationship in which the scorecard is genuinely two-sided. Give the supplier a channel to score you — on specification clarity, approval speed, payment punctuality and forecast accuracy — and act on it visibly. A supplier that can tell you your technical packs are incomplete and see them improve next quarter will treat the scorecard as a shared tool rather than an instrument of control, and the data will get better as a result.
If you want to see what a supplier’s own process documentation looks like before you score it at all, review how a custom bag programme runs from first sample through bulk production and ask any candidate supplier for the same level of detail. Minimum order quantity is 500 pieces per style, sampling takes 6–10 working days, bulk production runs 35–50 days, and quotations are issued FOB Xiamen. A supplier that can describe its own process in specific, dated terms is usually a supplier whose scorecard will be easy to populate.
Frequently Asked Questions
Q1. What should a supplier scorecard measure?
Five dimensions: quality, delivery, responsiveness, cost and compliance. Suggested weights are 35, 25, 15, 15 and 10. Each dimension holds two to three metrics with published thresholds so the score can be computed without judgement.
Q2. How many metrics should a scorecard have?
Twelve to fifteen at most, on one page. Scorecards die from length: once the document runs past a single sheet it stops being discussed in meetings, and a document that is not discussed does not drive decisions.
Q3. What is a good on time in full target for a bag supplier?
Ninety-five percent against the confirmed date is a reasonable target. The important detail is scoring against a date confirmed in writing at order placement, and counting partial shipments as failures.
Q4. How do I measure quality without relying on opinion?
Use the inspection report you already receive: defects per thousand units inspected, first-pass pass rate, customer-facing defect rate, and whether any defect category has repeated from a previous shipment.
Q5. Should the scorecard be shared with the supplier?
Yes. Share the metric definitions, weights and thresholds before the first scoring period and share the score afterwards. A rating computed privately and revealed only when it is bad reads as a weapon and gets managed rather than improved.
Q6. How does a score turn into order allocation?
Publish a fixed band-to-share mapping, for example: 90 to 100 earns 60 to 70 percent, 75 to 89 earns 25 to 35 percent, 60 to 74 earns up to 10 percent with remediation, below 60 earns none. Keep the mapping mechanical so allocation stops being a discussion.
Q7. Should I ever allocate one hundred percent to my best supplier?
No. The scorecard measures performance, not solvency, fire or a regional power restriction. Keep a real volume floor with a qualified second source even when its score is lower, because continuity is worth more than the optimisation.
Q8. What is the most common reason scorecards get abandoned?
Manual data collection. A metric requiring someone to remember, judge or type something survives about two quarters. Design every metric to be a by-product of work that already happens, such as inspection reports and goods receipts.
Q9. How often should I review supplier performance?
Quarterly, on the same cadence regardless of performance. Cancelling reviews when things are going well teaches the supplier that the process only runs when there is a problem, which turns the next review into a disciplinary meeting.
Q10. What should a quarterly supplier review agenda contain?
Confirm the score in five minutes, quality root causes in fifteen, delivery misses in ten, forward volume plan in ten, cost and compliance in ten, and agreed recorded decisions in ten. Sharing the forward plan is usually the highest-value item.
Q11. How should I handle a falling score?
With remediation rather than penalty: a specific measurable problem, a root cause stated by the supplier, a corrective action with a date and owner, and a verification method over the next two shipments.
Q12. Can a supplier game the scorecard?
Yes. Inspection pass rate invites negotiating the AQL; on time in full invites refusing to confirm a date. Fix sampling levels in writing, count unconfirmed shipments as failures, and keep one metric the supplier cannot influence, such as customer-facing defect rate.
Q13. How should cost be scored?
On normalised landed competitiveness rather than raw unit price, plus three behaviours: price stability on reorders, transparency when a cost breakdown is requested, and whether cost reduction proposals arrive with evidence attached.
Q14. What compliance data belongs on a scorecard?
Expiry dates, not certificate folders: social compliance audit, quality management certification, restricted substances testing per material, material declarations, and origin evidence. Score each against the bill of materials currently being produced.
Q15. Should the supplier also score me?
Yes. Give a channel for feedback on specification clarity, approval speed, payment punctuality and forecast accuracy, and act on it visibly. A third of delivery misses are typically caused by the buyer, and that is the highest-return item available.
Q16. How long should I keep weights stable?
At least four quarters. A weight that moves is a signal the supplier cannot act on, because it cannot tell whether its score changed because it changed or because you did.
Q17. What is the difference between supplier qualification and a scorecard?
Qualification is a one-time gate: can this supplier make the product acceptably? The scorecard is the ongoing mechanism that decides, every quarter, how much of your volume that qualified supplier receives.
People Also Ask
What is a supplier scorecard?
A weighted rating across quality, delivery, responsiveness, cost and compliance that maps directly to how much of your next order each supplier receives.
How many metrics should a supplier scorecard have?
Twelve to fifteen, on one page. Longer scorecards stop being read and stop driving decisions.
How do you weight supplier performance?
Quality 35, delivery 25, responsiveness 15, cost 15, compliance 10 is a workable starting split, held stable for at least four quarters.
What is on time in full?
The share of purchase order lines delivered complete by the date confirmed in writing at order placement. Partial shipments count as failures.
How do I stop a scorecard being gamed?
Fix sampling and defect classification in writing, count unconfirmed dates as late, score quotation completeness as well as speed, and keep one customer-facing metric.
Should supplier scores affect order volume?
Yes. Allocation is the only consequence that changes supplier behaviour, and it should follow a published band-to-share rule.