Bias in AI: Lessons from the “Racist” Soap Dispenser

How design assumptions, proxy choices and uneven testing can turn technical decisions into unequal outcomes
A white and yellow liquid dispenser placed on a marble countertop, showcasing its sleek design and functional purpose.

Executive summary

  • A widely shared soap-dispenser video became a memorable example of technology working differently for different people. The device was probably not AI, which makes the lesson more important: harmful bias can enter before a model is ever trained.
  • Bias is rarely caused by one bad dataset or one developer. It can arise through problem definition, sensors, historical data, labels, proxy variables, model design, testing, human decisions and the environment in which a system is deployed.
  • Overall accuracy is not enough. A system can perform well on average while producing materially worse outcomes for particular groups or at the intersections of several characteristics.
  • Responsible AI requires lifecycle governance: clear accountability, representative evaluation, human recourse, ongoing monitoring and the willingness to stop or redesign a system when harms cannot be managed.

A dispenser that did not see every hand.

In 2017, a short video circulated widely online. An automatic soap dispenser responded to a lighter-skinned hand but failed to activate for a darker-skinned hand. When a white paper towel was placed over the same hand, the dispenser worked. The device was quickly labelled the “racist” soap dispenser.

The label was provocative, but the more useful question is not whether the device held an intention. It is why a product designed for the public could work reliably for one user and fail for another. The dispenser appears to have relied on an optical or infrared sensor rather than artificial intelligence. Its behaviour may have reflected sensor sensitivity, calibration, material choices or limited testing across skin tones.

That distinction matters. Unequal technology outcomes do not begin and end with machine-learning models. They can originate in hardware, design assumptions, procurement decisions and test conditions. AI can then amplify the same pattern by applying it faster, more broadly and with greater authority.

Bias is a system problem

Bias is often described as a data-quality issue: improve the training data and the problem disappears. Better data is essential, but the diagnosis is incomplete. NIST treats AI as a socio-technical system and distinguishes among systemic, computational and human-cognitive sources of bias. In practice, these sources interact throughout the lifecycle.

Where bias enters
What can go wrong
Executive question
Problem definition
The system optimizes the wrong objective or ignores an important harm.
Are we solving the right problem, and for whom?
Data and representation
Some populations, contexts or edge cases are missing or underrepresented.
Who is represented, and who is absent?
Labels and proxies
A measurable variable is substituted for the outcome that actually matters.
Does the target measure the intended concept?
Model and thresholds
Performance or decision thresholds differ materially across groups.
How does performance vary beyond the average?
Deployment
A model is used in a context that differs from its testing environment.
Is this use consistent with the evidence?
Human and organizational controls
People over-trust outputs, lack recourse or cannot identify incidents.
Who can challenge, override and stop the system?

Three real-world lessons

1. Representation changes performance

The 2018 Gender Shades study evaluated three commercial gender-classification systems across gender and skin-type groups. Darker-skinned women were the most frequently misclassified group, with error rates reported as high as 34.7%, compared with a maximum error rate of 0.8% for lighter-skinned men. The study also found that two widely used benchmark datasets were overwhelmingly composed of lighter-skinned subjects.

The lesson is not that every facial system will perform the same way. NIST testing has shown that demographic effects vary significantly by algorithm, application and image conditions. The lesson is that aggregate performance can conceal material differences. Testing must reflect the people, environments and consequences of the intended use.

2. A neutral-looking proxy can reproduce inequality

A 2019 Science study examined an algorithm used to identify patients who should receive additional care-management support. The system used health-care spending as a proxy for health need. Because spending patterns also reflected unequal access to care, Black patients with the same risk score were, on average, considerably sicker than White patients. Replacing cost with a more direct measure of health need substantially increased the share of Black patients identified for additional help.

The model did what it was asked to do. The failure was in the objective. This is one of the most important executive lessons in AI governance: a convenient metric can encode historical inequality while appearing objective and mathematically precise.

3. Deployment controls matter as much as model accuracy

In 2023, the U.S. Federal Trade Commission announced a proposed order prohibiting Rite Aid from using facial-recognition technology for surveillance purposes for five years. The FTC alleged that the retailer failed to implement reasonable safeguards and that false matches led to consumer harm, with Black, Asian, Latino and women consumers especially likely to be affected.

The case was not only about an algorithm. It involved image quality, watchlist practices, testing, employee response, transparency and governance. Even a technically capable model can become unsafe when the surrounding operating process is weak.

Accuracy is not the same as fairness

A single performance score is attractive because it is easy to compare and report. It can also create false confidence. A model that is 95% accurate overall may still fail disproportionately for a smaller group, perform poorly under certain lighting or language conditions, or create more serious consequences when it is wrong for one population than another.

Fairness is also context-dependent. Different definitions – equal error rates, equal access, equal treatment or equal outcomes – can conflict. There is rarely one mathematical test that resolves the issue. Organizations need to define the relevant harms, affected stakeholders and acceptable trade-offs before selecting metrics.

The question is not simply, “Is the model accurate?” It is, “Accurate for whom, under what conditions, and with what consequence when it is wrong?”

What business leaders should expect

Define the decision and the harm

Document the intended use, the people affected, the decisions influenced and the outcomes that must be prevented.

Challenge data, labels and proxies

Understand data provenance, representation, exclusions and whether labels measure the business or social outcome that actually matters.

Test performance by relevant groups

Evaluate error rates, thresholds and outcomes across meaningful demographic, geographic, linguistic and operational segments.

Design recourse and human oversight

Ensure consequential decisions can be reviewed, explained, challenged and corrected by people with clear authority.

Govern vendors and third parties

Require evidence of testing, documentation, incident handling, model limitations and material changes – not only contractual assurances.

Monitor real-world outcomes

Track drift, complaints, overrides, incidents and downstream outcomes after deployment. Approval is the start of governance, not the end.

Executive perspective

Bias cannot be eliminated by a final fairness test or a diversity statement. It has to be managed as a lifecycle and operating-model discipline. That means involving business owners, risk teams, legal and compliance functions, technical specialists and people who understand how the system will affect real users.

The goal is not to demand perfect neutrality from every system. It is to make assumptions visible, measure performance honestly, provide meaningful recourse and prevent automation from scaling harms that the organization has not understood. Organizations that do this well will be better positioned to earn trust, satisfy emerging regulatory expectations and deploy AI at scale with fewer surprises.

Key takeaways

  • The soap-dispenser example is not primarily a story about malicious intent; it is a story about incomplete design and testing.
  • Harmful bias can enter through objectives, sensors, data, labels, proxies, models, human decisions and deployment processes.
  • Overall accuracy can hide unequal subgroup performance and unequal consequences.
  • Fairness requires context-specific definitions, stakeholder input and explicit trade-off decisions.
  • Responsible AI governance must continue after deployment through monitoring, recourse, incident management and accountability.

Primary sources and further reading

Prefer the concise version?

Get the two-page executive brief

A concise, visual overview of where AI bias enters, what real-world cases reveal, and the controls business leaders should expect.
By submitting this form, you agree to receive the requested content and occasional DataFuel insights. You can unsubscribe at any time.