Strategic Sourcing | | | 5 min read

Your Supplier Scorecard Probably Is Not Doing Anything

Your Supplier Scorecard Probably Is Not Doing Anything: Strategic Sourcing | Sourcing Tomorrow

Key takeaways

  • Pooling 740 measured effects from 38 studies, the unconditional effect of performance-based contracting was 0.008 at p = .303, meaning no reliable advantage over simply buying the service.
  • Of those 740 effects, 21 percent were negative and significant. Roughly one application in five made performance worse rather than better.
  • Once contract design was controlled for, the mean effect rose to 0.119 and became significant at the 1 percent level, so design carried the result rather than measurement itself.
  • Contracts built on process or output targets performed worse than outcome-focused contracts, with a negative coefficient significant at the 5 percent level.
  • Incentive pay directs where an agent spends attention, so heavy metrics on countable duties pull effort away from duties nobody is counting.
740
Measured effects pooled across 38 studies and 10 service areas
Bel, Espaillat and Esteve, Journal of Public Administration Research and Theory, 2026, 36(2), 250-266
21%
Those measured effects that were negative and statistically significant
Bel, Espaillat and Esteve, Journal of Public Administration Research and Theory, 2026, 36(2), 250-266
0.008
Unconditional effect of adopting performance-based contracting, not significant at p = .303
Bel, Espaillat and Esteve, Journal of Public Administration Research and Theory, 2026, 36(2), 250-266

Supplier scorecards are close to standard practice, and they rest on an assumption that rarely gets tested: that measuring performance improves it. You measure, they improve.

The evidence does not support that, and the failure is not that people picked the wrong metrics. It is more structural than that.

What happens when you pool the studies

A meta-regression published in April synthesized 740 measured effects from 38 studies across 10 service areas, examining whether performance-based contracting actually delivers (Bel, Espaillat and Esteve, Journal of Public Administration Research and Theory, 2026).

The headline result is uncomfortable. The unconditional effect of adopting performance-based contracting was positive but not statistically significant, a coefficient of 0.008 at p = .303. On its own, tying payment to measured performance did not reliably beat simply buying the service.

The distribution matters more than the average. Of those 740 effects, 31 percent were positive and significant, 48 percent were not significant, and 21 percent were negative and significant. Roughly one measured application in five made performance worse, not better.

What separated the ones that worked

Once the analysis controlled for how contracts were designed, the mean effect rose to 0.119 and became significant at the 1 percent level. Design was doing the work, not the existence of measurement.

Which dimensions you decide to count is a category decision before it is a measurement one, the same way the specifier decides where price actually changes minds. The clearest single moderator was what got measured. Contracts built on process or output targets carried a negative coefficient, significant at the 5 percent level and stable across robustness checks. Outcome-focused contracts outperformed them. The authors also identify residual control, collaboration and shared goals as central to success.

Why process metrics actively distort behavior

There is a well-established explanation for that result, and it predates the scorecard software by decades.

The multitask analysis of incentives established that when an agent has several duties, incentive pay does more than motivate effort. In the authors' words, it "serves to direct the allocation of the agents' attention among their various duties" (Holmstrom and Milgrom, Journal of Law, Economics, and Organization, 1991).

The consequence is the part that should worry anyone running a scorecard. When some duties are easy to measure and others are not, strong incentives pull effort toward the measurable ones and away from the rest. You do not get more total effort. You get the same effort, redirected toward whatever you decided to count.

What that looks like on a supplier scorecard

On-time delivery is easy to measure. Whether the supplier warned you early about a problem developing three tiers down is not. Invoice accuracy is easy to count. Whether their engineer stayed late to solve something that was your fault is not.

Weight the first heavily enough and you will get excellent on-time delivery, achieved partly by padding lead times, holding buffer inventory you pay for indirectly, and shipping partial orders that technically arrive on schedule. The metric improves. The relationship gets worse. Nothing was gamed in any sense the supplier would recognize as dishonest.

What to do differently

Measure fewer things, and measure outcomes

The meta-regression is direct on this: process and output measures were associated with worse results than outcome measures. An outcome is the thing you actually wanted. Not "responded within four hours" but "the line did not stop." Not "attended the quarterly review" but "the defect rate fell."

Separate what you can measure from what you cannot

The multitask work has a specific design implication: allocate tasks so that easily measured work is separated from work that is hard to measure. If one supplier owns both a countable duty and an uncountable one, heavy metrics on the first will quietly starve the second.

Stop renegotiating the metric

Frequent revision of what counts is corrosive. Each change tells the supplier the target moves when they hit it, which is the fastest way to teach them to manage the measurement rather than the work.

Accept that some of it is unmeasurable and manage it directly

Residual control, collaboration and shared goals came out as central to success. That is the same ground where a negotiation is decided before the concessions start. None of those live on a scorecard. They live in the relationship, in what you commit to, and in whether the supplier believes the next conversation will be fair.

An honest limit on the evidence

The meta-regression covers public outsourcing rather than firm-to-firm purchasing, and much of the wider evidence on metric gaming comes from labor economics and healthcare. Whether the effect sizes transfer cleanly to a commercial supplier relationship is not established.

What does transfer is the mechanism. Attention follows measurement wherever an agent has more duties than the principal can count, and that is every supplier relationship any of us manage.

The short version

A scorecard is not a neutral instrument that reveals performance. It is an incentive that redirects effort, and the evidence says it does that badly about as often as it does it well. Measure the outcome you actually want, keep the list short, leave the target alone once it is set, and handle the unmeasurable parts by managing the relationship rather than by adding another column.

For more on category strategy and supplier management, see our articles. To discuss a supplier performance program, get in touch.

References

  1. Bel, G., Espaillat, P., and Esteve, M. (2026). The performance of performance-based contracting in public outsourcing: a meta-regression analysis. Journal of Public Administration Research and Theory, 36(2), 250 to 266. doi.org/10.1093/jopart/muaf037
  2. Holmstrom, B., and Milgrom, P. (1991). Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design. Journal of Law, Economics, and Organization, 7 (Special Issue), 24 to 52. doi.org/10.1093/jleo/7.special_issue.24 (open copy: sfu.ca)

Disclosure: This article is published by SourcingTomorrow and reflects our analysis and commentary on procurement and sourcing practice. It is for informational purposes only and does not constitute legal, financial, or operational advice, and should not be relied upon for purchasing, contracting or vendor selection decisions. Consult qualified advisors for guidance on specific situations.

You do not get more total effort. You get the same effort, redirected toward whatever you decided to count.

Claudio Tartaglia

What 740 measured effects say about performance-based contracting

Measure Result Reading
Positive and significant effects31%It works less than a third of the time
Non-significant effects48%Most often it changes nothing measurable
Negative and significant effects21%One in five applications made performance worse
Unconditional effect of adoption0.008, p = .303Adopting it alone is not the intervention
Effect once contract design is controlled0.119, significant at 1%Design carries the result
Process or output measures versus outcome measures-0.038, significant at 5%Measuring activity underperforms measuring results

Bel, Espaillat and Esteve, The performance of performance-based contracting in public outsourcing: a meta-regression analysis, Journal of Public Administration Research and Theory, 2026, 36(2), 250-266. Publication bias tests were non-significant, Egger p = .132 and Begg p = .153.

Frequently Asked Questions

Do supplier scorecards actually improve supplier performance?
Less reliably than most teams assume. A meta-regression pooling 740 measured effects from 38 studies found the unconditional effect of performance-based contracting was 0.008 at p = .303, statistically indistinguishable from no effect. Only 31 percent of measured effects were positive and significant, while 21 percent were negative and significant. Measurement by itself is not the intervention; how the contract is designed is.
Why would measuring supplier performance make things worse?
Because incentives redirect attention rather than simply increasing effort. The multitask analysis of incentive contracts established that incentive pay serves to direct how an agent allocates attention across their duties. When some duties are easy to measure and others are not, strong incentives on the countable ones pull effort away from the rest. Delivery dates improve while early warning about an upstream problem quietly stops arriving.
Should I measure process or outcomes?
Outcomes, according to the strongest available evidence. In the meta-regression, contracts built on process or output targets carried a negative coefficient significant at the 5 percent level and stable across robustness checks, while outcome-focused contracts performed better. In practice that means measuring whether the line stopped rather than whether the supplier responded within four hours, since the response time is a proxy and the stoppage is the thing you care about.
How many metrics should a supplier scorecard carry?
Fewer than most carry. Every metric you add moves supplier attention toward that dimension and away from everything unmeasured, so a long scorecard spreads attention thin while creating an impression of thoroughness. Keep the list short enough that each item represents an outcome you would genuinely trade other things to get, and manage the remaining dimensions through the relationship rather than by adding columns.
Does this research apply to commercial supplier relationships?
Partly, and the limit is worth stating. The meta-regression covers public outsourcing rather than firm-to-firm purchasing, and much of the wider evidence on metric distortion comes from labor economics and healthcare. The effect sizes may not transfer cleanly. The mechanism does transfer, because attention follows measurement wherever an agent has more duties than the buyer can count, which describes every supplier relationship.

Join the Discussion

Get procurement insights, sourcing strategies, and supply chain intelligence delivered to your inbox.

Subscribe to the Newsletter