On the podcast The Age of AI, hosted by the technology journalist Lara Lewington with Ben Clark, the guest was Raj Bharat Patel, VP of AI Transformation at Holistic AI. Asked plainly whether AI is fundamentally biased, he answered that it likely is, because the training data is a record of human history and human history carries bias inside it. The interesting part came next, when he explained what a governance team is supposed to do about that. You cannot drive bias to zero without wrecking the model, he argued, so the job is to set a tolerance, a defined level of disparity the organization is willing to accept, and to hold the system to it.
He gave a warning with it. When Google pushed its Gemini image tool too hard toward correction, it produced historically absurd pictures and destroyed the trust it was trying to build, which is what happens when a team chases maximal fairness instead of a chosen level of it. The number matters more than the direction, and someone has to pick it.
Run that idea across the vendor market and an uncomfortable pattern appears. If fairness is a tolerance the organization chooses, then a fairness claim only means something when a specific person picked the number, wrote down why, and can show the live model is being held to it. A dashboard full of fairness metrics with no chosen threshold behind it is Fairness Theater, the appearance of a governed system sitting on a decision no one ever made. Two questions sort the entire market, and almost no vendor passes both of them for bias.
Key Terms
Fairness Theater: A bias threshold that looks governed but was never actually chosen, owned, or justified by anyone.
Bias Tolerance: The specific level of disparity between groups an organization decides it will accept for a given use case.
Measurement Half: The ability to measure a fairness metric against live production outputs and act when it moves past the line.
Ownership Half: A named owner, a written rationale, an approval trail, and a review date attached to that specific tolerance.
Disparate Impact: A gap in outcomes across protected groups, often measured as an impact ratio, with four-fifths a long-standing reference line.
Fairness Drift: The movement of a model's fairness profile after deployment as the data flowing through it changes.
Accountable Owner: The named person who signs off on the tolerance and answers for it when a regulator or a court asks.
The Two-Part Test
A governed bias tolerance has to satisfy two conditions, and they tend to live in different tools. The first is the measurement half, which means the fairness metric is checked against real production outputs and something happens when it drifts past the accepted line. Monitoring and observability platforms are strong here, because watching a running model is what they were built to do.
The second is the ownership half, which means the tolerance is a specific number that a named person chose, with a written reason, an approval record, and a date it will be reviewed. Governance platforms and the model risk management discipline in banking are strong here, because assigning accountability and keeping a defensible record is their whole tradition.
A number that gets measured but was never chosen is a readout without a decision behind it. A number that gets chosen but never measured is a policy with no system holding it up. Governing a bias tolerance requires the decision and the measurement wired together, and because the two halves usually sit in separate products, the wiring is left to the buyer. That gap is where most programs end up with neither a chosen number nor proof that any number holds.
Who Delivers Which Half
The table below places the platforms alphabetically against the two questions, whether they measure bias against live outputs and whether they govern the tolerance with an owner, a rationale, and a sign-off. The deep dives that follow group the same vendors by the half they actually deliver.
Platform | Measures bias vs. live outputs | Governs the tolerance (owner, rationale, sign-off) | Where it lands |
|---|---|---|---|
Arthur AI | Yes | Emerging | Measurement half |
Credo AI | By integration | Yes | Ownership half |
Fiddler AI | Yes | Left to buyer | Measurement half |
Holistic AI | Yes | Yes, newer | Attempts both |
Warden AI | Yes (HR) | Certifies, not owns | Measurement half |
GRC incumbents | No | Yes | Neither half for bias |
The measurement half · metrics live, ownership thin
Arthur, the model-monitoring platform from Arthur AI, measures fairness against live outputs by comparing results across demographic subgroups, even when the protected attribute is not a model input, and it lets teams set custom fairness thresholds that fire an alert when a model drifts past them. It has started adding the ownership layer, with accountable owners and real-time intervention described in its own governance guidance, though independent roundups still rate its evidence and compliance documentation as thinner than its monitoring. The measurement is real and the ownership record is catching up.
Fiddler, the observability platform from Fiddler AI, tracks fairness metrics on production traffic and notifies the model, engineering, or risk teams when a metric falls outside an accepted threshold. The threshold and the alert are strong, and the decision about what the number should be, who signs it, and why, is left to the customer to make and record elsewhere.
Warden AI, which runs continuous bias auditing aimed at hiring tools, monitors fairness over time and issues its Warden Assured certification that buyers now use to clear procurement. It measures and certifies well, and the choosing and owning of the underlying tolerance still sits with the employer, where the law also puts the liability.
The ownership half · governs the tolerance, measures by integration
Credo AI, through its governance platform and AI registry, is built around the ownership half. It assigns governance tasks to named stakeholders, keeps review dates, approval logs, and stakeholder sign-offs, and, in its own words, aligns policies, risk thresholds, and responsibilities as a single record, an approach the UK government's assurance catalogue describes as tracking accountability across the AI lifecycle. Credo reaches the measurement half by connecting to outside observability tools rather than measuring bias itself, so its strength is making sure a number is chosen, owned, and evidenced rather than producing the number.
Attempts both · the most complete single-vendor story, with real gaps
Holistic AI, whose governance platform supplied the thesis for this piece through its own executive, is the closest thing to a single vendor covering both halves. It measures bias natively through a testing suite and the JobFair line of hiring-bias research, and it offers configurable risk thresholds, approval workflows, and enforcement that intervenes when a threshold is crossed. The gaps are worth stating plainly. Independent bias-tool roundups still file Holistic nearer the framework and pre-deployment side alongside Credo, which means its runtime bias-enforcement claim is newer than its testing pedigree, and a buyer should ask it to show a bias number that is owned, justified, signed, and held against live outputs rather than accept that chain from a product tour. Being the most complete option in a thin field is a real credit and a low bar at the same time.
Neither half for bias · owns the paper, cannot measure the model
The established governance, risk, and compliance suites, among them IBM OpenPages, MetricStream, OneTrust, and ServiceNow, supply the ownership machinery in abundance, with owners, approval workflows, and policy records, and they were placed at the documentation layer in our compliance-theater analysis for a reason. They do not natively measure bias against live model outputs, so an owned threshold in one of these systems governs a number that nothing is checking, which is ownership without measurement.
The Law Wants a Chosen Number, and Most Organizations Never Set One
The frameworks that shape enterprise AI deliberately refuse to pick the number for you. The NIST AI Risk Management Framework leaves each organization to define its own risk tolerance, and practitioners who implement it report that far fewer companies actually define what acceptable looks like than talk about managing risk, which is the empty threshold seen from the framework side. When the tolerance is left undefined, teams make quiet assumptions and governance drifts into whatever the data team happened to default to.
The law is starting to demand that the number be chosen and shown. New York City's Local Law 144 has required annual independent bias audits of automated hiring and promotion tools since 2023, with impact ratios published on the employer's own website. Colorado passed the first comprehensive state AI law in 2024, revised it in 2026, and it takes effect in 2027, requiring companies that deploy AI for consequential decisions to use reasonable care against algorithmic discrimination. The European Union's AI Act classifies employment and hiring systems as high risk and attaches testing, documentation, and risk-management duties, with those high-risk obligations now deferred to late 2027.
The gap between the mandate and the practice is the whole argument, and it is measurable. A December 2025 audit by the New York State Comptroller, reported in industry coverage, found that the city agency had received only two complaints in two years, surveyed thirty-two companies and identified one instance of non-compliance, while the Comptroller's own reviewers found at least seventeen. Even where a chosen, published bias number is the law, most of the organizations that owe one have not been made to produce it. A workable model of what a governed number looks like already exists in finance, where a board-approved risk appetite sets a specific figure, names the person who signs it, records the trigger, and reports to the audit committee on a schedule.
Six Things Your Program Should Produce Without Scrambling
The practical test is short, and it works on any organization that has AI making decisions about people. For every model that shapes a hiring, lending, housing, insurance, or similar decision, the team should be able to produce six things on request without a scramble.
The fairness metric being used, such as an impact ratio across the relevant protected groups.
The exact tolerance, the specific number the organization has decided it will accept.
The named owner who signed off on that number and answers for it.
The date the tolerance was set and the date it is due for review.
The documented reason it was set where it was, tying the number to the use case and the law.
The live measurement showing the running model is inside that number today.
Every blank in that list is a piece of theater. The counting exercise that follows is just as cheap, since it costs nothing but honesty: count the models that make consequential decisions, then count how many can fill in all six items, and treat the difference as exposure that exists regardless of how polished the dashboard looks. When the same discipline is turned on a vendor, the question becomes concrete, so ask each platform to show where a bias tolerance is set, who signs it, where the rationale is stored, and how the running model is measured against that exact figure. A strong answer walks the whole chain, and a weak answer shows a configurable threshold field and a chart, which is the difference this piece is about.
Sources
"The Age of AI with Lara Lewington and Ben Clark," episode featuring Raj Bharat Patel of Holistic AI. podcasts.apple.com
Credo AI, "Evidence" and platform documentation. credo.ai · UK Government assurance catalogue profile: gov.uk
Fiddler AI, "How to Track Fairness and Bias in Predictive and Generative AI." fiddler.ai
Arthur AI, "AI Governance Framework Guide." arthur.ai · Independent comparison: infomineo.com
Holistic AI, AI Governance Platform. holisticai.com · Independent bias-tool roundup: trysight.ai
Warden AI, "AI Fairness Testing" and "Bias Audit Legal Requirements." warden-ai.com
New York City Department of Consumer and Worker Protection, "Automated Employment Decision Tools" (Local Law 144). nyc.gov
New York State Comptroller Local Law 144 enforcement audit, as reported in "Algorithmic Fairness Audits: A Compliance Guide for 2026." risktemplate.com
Warden AI, "Are Bias Audits Required by Law" (Colorado and EU AI Act timing). warden-ai.com
Logicalis, "AI Risk Management Framework Compliance and Risk Appetite" (NIST leaves tolerance undefined). us.logicalis.com
CFO Connect, "CFO AI Governance Framework" (board-approved risk appetite example). cfoconnect.eu
GetAIGovernance, "AI Governance Platforms That Cannot See Your Models Are Selling You Compliance Theater." getaigovernance.net
Our Take
The AI Governance Take
Raj Bharat Patel is right on the core point, that zero bias is the wrong target and a chosen, held tolerance is the right one, and the direction of the law is proving him correct. The bar that argument sets is higher than most of the market clears, because a tolerance that nobody chose or owns is theater no matter how precise the metric behind it, and the industry has become very good at selling the metric, the threshold field, and the alert while leaving the choosing and the owning to the buyer.
Read alongside our earlier analysis of compliance theater, this forms one grid with two axes. The first axis asks whether you can see the system in production at all, and the second asks whether you can hold that system to a number a named person actually chose. A platform can pass one axis and fail the other, and a governed AI program needs both.
Honesty requires naming what stays hard even when the will is there. Fairness metrics conflict with one another mathematically, so a model that looks fair on one measure can look unfair on another, and no team can satisfy all definitions at once. Pushing a bias metric harder trades against accuracy, the tradeoff Gemini illustrated in public, and there is still no industry-standard threshold to copy. Many organizations do not collect the protected-attribute data that measurement depends on, so even a well-meaning team can struggle to produce the six items above. None of that excuses leaving the number blank, and all of it is a reason to name an owner who has to work through it.
Buyers who want to close the gap can compare platforms by which half they deliver in the AI Governance category at GetAIGovernance.net, where the tools that measure fairness against live behavior sit alongside the ones that make an accountable owner choose and defend the number, so a team can assemble both halves on purpose rather than discover it has neither.