A dispenser that did not see every hand.
In 2017, a short video circulated widely online. An automatic soap dispenser responded to a lighter-skinned hand but failed to activate for a darker-skinned hand. When a white paper towel was placed over the same hand, the dispenser worked. The device was quickly labelled the “racist” soap dispenser.
The label was provocative, but the more useful question is not whether the device held an intention. It is why a product designed for the public could work reliably for one user and fail for another. The dispenser appears to have relied on an optical or infrared sensor rather than artificial intelligence. Its behaviour may have reflected sensor sensitivity, calibration, material choices or limited testing across skin tones.
That distinction matters. Unequal technology outcomes do not begin and end with machine-learning models. They can originate in hardware, design assumptions, procurement decisions and test conditions. AI can then amplify the same pattern by applying it faster, more broadly and with greater authority.
Bias is a system problem
Bias is often described as a data-quality issue: improve the training data and the problem disappears. Better data is essential, but the diagnosis is incomplete. NIST treats AI as a socio-technical system and distinguishes among systemic, computational and human-cognitive sources of bias. In practice, these sources interact throughout the lifecycle.
Where bias enters |
What can go wrong |
Executive question |
Problem definition |
The system optimizes the wrong objective or ignores an important harm. |
Are we solving the right problem, and for whom? |
Data and representation |
Some populations, contexts or edge cases are missing or underrepresented. |
Who is represented, and who is absent? |
Labels and proxies |
A measurable variable is substituted for the outcome that actually matters. |
Does the target measure the intended concept? |
Model and thresholds |
Performance or decision thresholds differ materially across groups. |
How does performance vary beyond the average? |
Deployment |
A model is used in a context that differs from its testing environment. |
Is this use consistent with the evidence? |
Human and organizational controls |
People over-trust outputs, lack recourse or cannot identify incidents. |
Who can challenge, override and stop the system? |
Three real-world lessons
1. Representation changes performance
The 2018 Gender Shades study evaluated three commercial gender-classification systems across gender and skin-type groups. Darker-skinned women were the most frequently misclassified group, with error rates reported as high as 34.7%, compared with a maximum error rate of 0.8% for lighter-skinned men. The study also found that two widely used benchmark datasets were overwhelmingly composed of lighter-skinned subjects.
The lesson is not that every facial system will perform the same way. NIST testing has shown that demographic effects vary significantly by algorithm, application and image conditions. The lesson is that aggregate performance can conceal material differences. Testing must reflect the people, environments and consequences of the intended use.
2. A neutral-looking proxy can reproduce inequality
A 2019 Science study examined an algorithm used to identify patients who should receive additional care-management support. The system used health-care spending as a proxy for health need. Because spending patterns also reflected unequal access to care, Black patients with the same risk score were, on average, considerably sicker than White patients. Replacing cost with a more direct measure of health need substantially increased the share of Black patients identified for additional help.
The model did what it was asked to do. The failure was in the objective. This is one of the most important executive lessons in AI governance: a convenient metric can encode historical inequality while appearing objective and mathematically precise.
3. Deployment controls matter as much as model accuracy
In 2023, the U.S. Federal Trade Commission announced a proposed order prohibiting Rite Aid from using facial-recognition technology for surveillance purposes for five years. The FTC alleged that the retailer failed to implement reasonable safeguards and that false matches led to consumer harm, with Black, Asian, Latino and women consumers especially likely to be affected.
The case was not only about an algorithm. It involved image quality, watchlist practices, testing, employee response, transparency and governance. Even a technically capable model can become unsafe when the surrounding operating process is weak.
Accuracy is not the same as fairness
A single performance score is attractive because it is easy to compare and report. It can also create false confidence. A model that is 95% accurate overall may still fail disproportionately for a smaller group, perform poorly under certain lighting or language conditions, or create more serious consequences when it is wrong for one population than another.
Fairness is also context-dependent. Different definitions – equal error rates, equal access, equal treatment or equal outcomes – can conflict. There is rarely one mathematical test that resolves the issue. Organizations need to define the relevant harms, affected stakeholders and acceptable trade-offs before selecting metrics.
The question is not simply, “Is the model accurate?” It is, “Accurate for whom, under what conditions, and with what consequence when it is wrong?”
What business leaders should expect
Define the decision and the harm
Document the intended use, the people affected, the decisions influenced and the outcomes that must be prevented. |
Challenge data, labels and proxies
Understand data provenance, representation, exclusions and whether labels measure the business or social outcome that actually matters. |
Test performance by relevant groups
Evaluate error rates, thresholds and outcomes across meaningful demographic, geographic, linguistic and operational segments. |
Design recourse and human oversight
Ensure consequential decisions can be reviewed, explained, challenged and corrected by people with clear authority. |
Govern vendors and third parties
Require evidence of testing, documentation, incident handling, model limitations and material changes – not only contractual assurances. |
Monitor real-world outcomes
Track drift, complaints, overrides, incidents and downstream outcomes after deployment. Approval is the start of governance, not the end. |
Executive perspective
Bias cannot be eliminated by a final fairness test or a diversity statement. It has to be managed as a lifecycle and operating-model discipline. That means involving business owners, risk teams, legal and compliance functions, technical specialists and people who understand how the system will affect real users.
The goal is not to demand perfect neutrality from every system. It is to make assumptions visible, measure performance honestly, provide meaningful recourse and prevent automation from scaling harms that the organization has not understood. Organizations that do this well will be better positioned to earn trust, satisfy emerging regulatory expectations and deploy AI at scale with fewer surprises.
Key takeaways
- The soap-dispenser example is not primarily a story about malicious intent; it is a story about incomplete design and testing.
- Harmful bias can enter through objectives, sensors, data, labels, proxies, models, human decisions and deployment processes.
- Overall accuracy can hide unequal subgroup performance and unequal consequences.
- Fairness requires context-specific definitions, stakeholder input and explicit trade-off decisions.
- Responsible AI governance must continue after deployment through monitoring, recourse, incident management and accountability.
Primary sources and further reading
- NIST SP 1270 – Towards a Standard for Identifying and Managing Bias in Artificial Intelligence
- NIST – There Is More to AI Bias Than Biased Data
- Buolamwini and Gebru – Gender Shades
- NIST – Face Recognition Vendor Test: Demographic Effects
- Obermeyer et al. – Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations
- U.S. FTC – Rite Aid Facial Recognition Enforcement
- NIST – AI Risk Management Framework
- Ren and Heacock – Sensitivity of Infrared Sensor Faucets on Different Skin Colours