Act III · The Objections

Industrialist Paper No. 32

The Score Will Be Gamed

By Andrew Kornuta · July 27, 2026 · 6 min read

Trust has to scale or the network does not work. A buyer in a line-down situation cannot personally re-qualify every unfamiliar shop, so the system has to carry trust on their behalf, compressed into a signal that routing can act on. I have called that signal a trust score, and the objection arrives immediately and correctly: any number that controls who gets work will be attacked by the people who want the work. The moment a score decides routing, it stops being a neutral measurement and becomes a prize.

Claim: a manufacturing trust score survives adversarial pressure only when it is computed from verified events bound to artifacts — delivery records, NCR closures, cert packets whose scope and expiration someone actually checked — rather than from self-claims, stars, or paperwork accepted at face value. If capability can be asserted and certificates accepted without verification, the score will be gamed, and the drift appears as high self-reported ratings sitting next to rising quality escapes and forged or expired certifications in the award record.

This is Goodhart's Law, and the precise version is the argument. Charles Goodhart wrote in 1975 that "any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." The popular paraphrase — when a measure becomes a target it ceases to be a good measure — belongs to Marilyn Strathern rather than Goodhart, and that distinction matters. The failure is not that measurement is bad. It is that measurement under incentive pressure decays unless somebody actively defends it.

The objection at its strongest

The critic's best case is that reputation systems get gamed everywhere, and manufacturing offers richer targets than most. The economic pull is enormous. Harvard research by the economist Michael Luca found that one additional Yelp star raises a restaurant's revenue by five to nine percent, and where a signal is worth that much, an industry grows up to manufacture it. Amazon sued the operators of more than ten thousand Facebook groups that brokered fake reviews, one of them with over forty thousand members. In 2024 the Federal Trade Commission finalized a rule, 16 CFR Part 465, banning fake and AI-generated reviews outright, effective October 21, 2024, with penalties up to $51,744 per violation, and Chair Lina Khan noted that fake reviews "pollute the marketplace and divert business away from honest competitors."

Manufacturing's version is worse than a fake star, because the thing being faked here is a certificate of competence. Peer-reviewed work on faking ISO 9001 documented one certification body that opened a China office and within weeks submitted roughly 700 certificates that its U.S. office recognized as fraudulent, and it shut the office down. The aviation supply chain then lived the consequence: a director of a parts distributor called AOG Technics forged Authorised Release Certificates for some sixty thousand engine parts, inventing fake quality managers to give the paperwork credibility, before a UK court sentenced him to four years and eight months. A cert packet is only trust if someone checked it.

What the gameable version looks like

The bad version trusts what suppliers say about themselves. It lets a shop list processes, tolerances, and certifications it cannot support, then renders that self-portrait as a clean profile and a high rating. It accepts a certificate as a scanned image instead of as a claim with a scope and an expiration to verify. It leans on human moderation that cannot possibly scale to a national RFQ channel, so spammers and optimistic bidders outrun the reviewers. Worst of the four, it treats a star average as a score, which rewards volume and age over demonstrated fit, so the shop that quotes everything and ships adequately outranks the specialist who declines poor-fit work honestly. Each of these is a door, and adversaries are patient about doors.

The governed design

The score I'd build is computed from events, not adjectives. It weights what a supplier did — acknowledged an RFQ, hit a commit date, closed an NCR, shipped a complete cert packet, earned a repeat award — over what a supplier claimed. Every input binds to an artifact, whether a delivery record, a receiving result, a CMM report, or a corrective-action closure, so that the score points back to evidence a buyer could audit. Capability assertions require proof before they unlock better RFQs. Certificates get verified for scope and expiration rather than accepted as pictures.

Spam and no-response have to carry consequences while an honest decline does not, or the network trains its suppliers to quote everything in sight. The score decays with time, so a good year three years ago cannot coast. And it carries an appeal path, because a wrong attribution or a buyer-side failure ought to be correctable rather than silently punishing a supplier who did nothing wrong. One enforcement boundary holds all of it up, and it is load-bearing: only verified events write to the score.

How you would know

Measure the share of a supplier's score that comes from verified events versus self-reported claims, because a gameable system leans on the latter. Track the certificate-verification rate, the fraction of cert packets checked for scope and expiration rather than accepted on sight. Watch for the tell of gaming, which is high self-ratings sitting right next to rising NCRs and receiving rejections. Monitor decline discipline, so honest no-fit declines stay distinguishable from ghosting. If appeal volume spikes, the attribution logic is wrong; if score volatility spikes with no matching change in outcomes, somebody is pushing on the controller.

Implications

If the score is a popularity number, the critic is right, and the network routes the country's work to whoever is best at looking good. If the score is a controller fed only by verified events, gaming gets expensive, because faking the signal now means faking delivery, closure, and certification in front of buyers who can check. A national coordination layer inherits only the trust it can defend, and a trust layer that can be faked is worse than none at all, because it launders bad actors into confident awards. The practical failure mode is gaming.

Next I'll put the same machinery under a different kind of pressure, this time not from a fraudster but from the platform's own business model: if trust cannot be faked into the score, can it simply be bought?

Questions to Ask

  1. What share of this supplier's score comes from verified events versus self-reported claims?
  2. Is each capability claim backed by an artifact, or by the supplier's own description?
  3. Are certificates verified for scope and expiration, or accepted as images?
  4. Does an honest decline cost the same as ghosting, or less?
  5. What happens to the score when a job's outcome is later attributed to the wrong party?
  6. Which single metric would tell us the score is being gamed before the bad part ships?