Home/Blog /Procurement

Building a vendor evaluation scorecard that survives scrutiny

Weighted criteria, anchored scales, and the four failure modes that quietly turn a scorecard into a rubber stamp for a decision already made.

A vendor evaluation scorecard is a weighted model that turns a software selection into a comparable number per vendor. Done properly it makes a decision defensible, repeatable and fast. Done in the usual way it produces a number that justifies whichever vendor the loudest person in the room already preferred — and does so with a spreadsheet's air of objectivity.

The difference between the two is almost entirely about sequence.

Set the weights before you see a vendor

This is the rule that everything else depends on. Choose your criteria and assign their weights before any demo, any pricing, any conversation with a sales team. Write them down and circulate them.

The reason is straightforward: once you have seen a product you like, you will unconsciously weight the criteria it happens to win. This is not a character flaw; it is how everyone's judgement works. Fixing the weights in advance is the only reliable defence, and it converts the scorecard from a justification device into an evaluation.

A useful discipline: have each stakeholder allocate 100 points across the criteria independently, then average. Where allocations differ sharply you have found a genuine disagreement about what the organisation is buying, and it is far cheaper to resolve that in week one than in month six.

Separate gates from scored criteria

Two kinds of requirement get muddled constantly, and muddling them is how organisations buy software that cannot legally be used.

  • Must-haves (gates). Binary. A vendor either meets them or is eliminated. Data residency in a specific jurisdiction, a required certification, a hard integration, a compliance regime. These are never scored and never traded off against price — a vendor that fails a gate is out even if it wins everywhere else.
  • Scored criteria. Matters of degree, where more is better and trade-offs are legitimate. Usability, roadmap, support responsiveness, implementation effort, total cost.

Putting a must-have into the scored section is a common and expensive mistake. A high enough score elsewhere will always outweigh it, and the model will cheerfully recommend a vendor you cannot actually deploy.

How many criteria should you use?

Between eight and fifteen scored criteria for most software selections.

Below eight, the model is too coarse to separate vendors that are genuinely close. Above fifteen, individual weights shrink to a few percent each, no single score can move the total, and the exercise produces the appearance of rigour while actually flattening the real differences into noise. If you find yourself with thirty criteria, group them: three to five weighted categories, each containing a handful of sub-criteria that roll up.

CategoryTypical weightExample sub-criteria
Functional fit30–40%Core workflow coverage, configurability, reporting
Total cost of ownership15–25%Licence, implementation, integration, internal admin effort
Technical fit15–20%API quality, deployment model, data export, performance
Vendor viability10–15%Financial stability, customer base, roadmap credibility
Service and support10–15%SLA, onboarding, escalation path, regional coverage

These are starting points, not a template to copy unchanged. If your genuine constraint is that implementation must finish inside a quarter, then implementation effort deserves a far larger weight than the table suggests — and saying so out loud is exactly the point of the exercise.

Anchor the scale

An unanchored 1–5 scale means five different things to five evaluators. Anchor every point to an observable fact, per criterion, before scoring begins.

For "API quality", for example:

  • 5 — Documented REST and webhook APIs covering every object we need, with a sandbox and versioning policy.
  • 4 — Full API coverage, documented, no sandbox.
  • 3 — API covers our main objects; some gaps require export files.
  • 2 — Limited API; core workflows need manual export.
  • 1 — No API, or API available only on a higher tier we are not buying.

Writing anchors is tedious and it is where most of the value is. It converts "I felt like a 4" into a claim someone can check, and it makes scores comparable across evaluators who never spoke to each other.

Who should score, and how?

Three to five evaluators, drawn from the groups who will live with the decision: an owner from the using team, someone from engineering or IT, someone from security or procurement. More than five and scheduling collapses; fewer than three and one person's blind spot becomes the organisation's.

Score independently first. Collect scores before any group discussion. If the group scores together, the most senior voice sets an anchor for everyone else within the first two criteria, and the remaining "independent" scores are anything but.

Then discuss only where scores diverge by two points or more. Convergent scores need no meeting. Divergent ones almost always reveal that two people were scoring different things — which is a definition problem, not a vendor problem, and it is fixable in ten minutes.

Run a sensitivity check before you announce anything

Take your final model and shift each weight by ±10%, one at a time. Does the winner change?

  • No. The decision is robust. Say so in the write-up — it is a strong claim and it pre-empts the obvious challenge.
  • Yes, on one weight. Your decision genuinely rests on that weight. That is fine, but you must now defend the weight explicitly rather than the conclusion, and the write-up should say so plainly.
  • Yes, on several. The vendors are effectively tied on your criteria. Stop scoring and decide on something the model does not capture — a reference call, a paid pilot, or the exit cost if you are wrong.

A scorecard's job is not to remove judgement from a decision. It is to make the judgement visible, so that the thing being argued about is a weight everyone can see rather than a preference nobody will admit to.

Four failure modes to watch for

  1. Weights set after the demos. The scorecard is now a justification. Everything downstream is theatre.
  2. Price folded in as a scored criterion at a high weight. Price then swamps fit, and you buy the cheapest tool rather than the right one. Better: score fit and quality, then compare three-year total cost of ownership across the shortlist as a separate step.
  3. Scoring from the vendor's own materials. Every vendor scores 5 on their own datasheet. Score from a scripted demo of your workflow, a trial, and reference calls.
  4. No record of why. Six months later somebody will ask why you picked this vendor. Keep the weighted model, the anchors and the divergence notes — that package answers the question in one minute instead of one afternoon.

What the scorecard is not

It is not an oracle. A two-point gap between the top two vendors is inside the noise of any scoring exercise, and treating it as decisive is a misreading of your own model's precision. When the top two are close, the honest conclusion is "either is acceptable" — and the decision should then turn on something the scorecard cannot measure, such as which vendor's team you would rather call at 2am.

It is also not a substitute for understanding the market. Analyst grids like the Magic Quadrant, Wave and MarketScape are good sources for the criteria worth scoring — a considered view of what matters in a category — even when their rankings cannot know enough about you to be useful.

Frequently asked questions

How many criteria should a vendor scorecard have?

Between eight and fifteen scored criteria for most software selections. Fewer than eight and the model is too coarse to separate close vendors; more than fifteen and the weights become so small that individual scores stop moving the total, which produces the illusion of rigour while actually flattening real differences. Group related items into a small number of weighted categories rather than adding more line items.

Who should score the vendors?

Three to five evaluators drawn from the groups who will live with the decision — typically an owner from the using team, someone from engineering or IT, and someone from security or procurement. Have them score independently before any group discussion, then discuss only the criteria where scores diverge by two points or more. Scoring as a group first lets the most senior voice in the room set the anchor for everyone else.

Keep reading

Related articles

Vendor evaluation

Magic Quadrant vs. Forrester Wave vs. IDC MarketScape: what each grid actually measures

Three analyst firms, three two-dimensional grids, three different questions. Knowing which is which decides whether your citation strengthens an argument or quietly undermines it.

12 min read
Decision frameworks

Build vs. buy: a decision framework that accounts for the second year

Most build-versus-buy analyses compare a build estimate against a first-year licence fee. That comparison is wrong in both directions, and predictably so.

10 min read