This is the second paper in a five-part sequence about artificial intelligence. Can Software Cheat? separates software behavior from consciousness and responsibility. This paper examines perspective and judgment. What Are We Actually Afraid Artificial Intelligence Will Do? separates present harms, plausible risks, and speculative catastrophes by mechanism and evidence. Compounding Intelligence explains how people and organizations can build durable advantage with these tools. The AI Industrial System then moves outward to the physical, financial, and institutional system that makes those gains possible.
TL;DR
An artificial-intelligence answer can be wrong, unfair, interpretive, or incomplete, and those are different failures. Calling them all bias hides the cause and often points to the wrong remedy. Trust should depend on the kind of claim, the sources and assumptions behind it, the people affected, and the consequence of error.
A disagreement between trusted sources
A reader recently told me about an apparent disagreement between two sources she trusted.
She had asked BibleQuestions.com about premillennialism and postmillennialism. Its answer conflicted with what she understood John MacArthur’s New American Standard Bible (NASB) Study Bible to say. Which source was biased?
Because I do not have her exact question or the site’s exact response, I cannot reconstruct the disagreement. That’s part of the problem: a slight wording change can change what evidence a system retrieves, what assumptions it makes, and what answer it gives. I can evaluate the sources’ stated positions, but not the answer that appeared on her screen.
What I can verify is that neither source comes from nowhere. BibleQuestions.com says it uses artificial intelligence, video, and text to help people understand the Bible from a “grace-centered perspective.”[1] MacArthur has publicly and emphatically argued for premillennialism, grounding it in a particular approach to biblical interpretation and Israel's future.[2]
The disagreement does not establish which conclusion is correct or prove that one source malfunctioned; it could reflect different interpretive frameworks, source selections, or understandings of the question.
We use the word bias for all of those things, and it's becoming too imprecise to help us.
Disagreement is not proof of bias, nor is agreement proof of truth.
Five different problems hiding inside one word
When someone says an artificial-intelligence answer is biased, I want to know what kind of claim they're making. I propose a five-part diagnostic framework to separate problems that are commonly bundled together.
The first is factual error. The answer says that an event occurred when it did not, attributes a quotation to the wrong person, calculates a number incorrectly, cites a source that does not exist, or presents an obsolete fact as current. This is the most straightforward category, although deciding what counts as a fact can still require judgment.
The second is statistical or representational bias. The data used to build, tune, test, or retrieve information for a system do not adequately represent the people, conditions, language, or situations the system is applied to. A medical system tested primarily on one population may perform worse on another. A speech-recognition system may understand some accents better than others. A hiring model trained on historical decisions may reproduce patterns that should not be preserved.
Image generation provides a visible example. In one study of Stable Diffusion, neutral occupation prompts amplified racial and gender disparities: software developers appeared almost exclusively as pale, stereotypically masculine faces, while housekeepers appeared with darker skin tones and stereotypically feminine features.[8]
The third is harmful systemic or allocative bias. A system participates in a process that unfairly distributes opportunities, burdens, attention, or risk. Consider a hypothetical loan system that denies credit more often to similarly qualified applicants from a protected group because historical decisions, proxy variables, or the surrounding lending rules carry an existing disparity into new decisions. The problem may involve the software, the data, the institutional rules, the people using it, or all of the above. A technically accurate prediction can still be used in an unjust process. Appropriate remedies include testing outcomes across relevant groups, examining data and proxy variables, giving applicants specific reasons for adverse decisions, and providing meaningful review or recourse.[6][7]
The fourth concerns a disciplinary or interpretive framework, sometimes including value judgments. A theologian reads a text through one tradition. An economist evaluates a policy through a particular theory of incentives. A historian decides which events are causally important. A physician weighs benefits and harms using a clinical standard. A lawyer distinguishes controlling authority from persuasive authority. These are ways of deciding what evidence means, not merely collections of facts.
The fifth is missing context or undisclosed source selection. The answer may be reasonable within one set of assumptions but misleading because the user does not know what those assumptions are. A theological assistant might answer entirely from one denomination’s sources without identifying that tradition or acknowledging credible alternatives. A medical assistant might answer for the average patient while omitting a condition that makes this patient an exception. The words may be accurate within the hidden frame and still mislead the reader.
These categories can overlap. A narrow dataset can create both statistical error and disparate harm. A hidden interpretive framework can cause an answer to sound more settled than the evidence warrants. A factual mistake can be repeated so widely that it becomes institutional practice.
The remedies differ accordingly. A false date calls for correction, an unrepresentative dataset for better data and testing, and discriminatory allocation may require changing the surrounding institution. An interpretive dispute calls for visible assumptions, sources, and credible alternatives, while missing context calls for better questions and more honest disclosure.
Calling all five biases can make a conversation sound morally serious while leaving the actual problem untouched.
There is no answer from nowhere
Can Software Cheat? drew a central boundary: an artificial-intelligence system does not possess a human worldview. Software can produce language that sounds certain, compassionate, ideological, evasive, or angry without experiencing certainty, compassion, ideology, embarrassment, or anger.
Even so, an answer can reflect a point of view.
The apparent paradox disappears when we stop looking for a mind inside the software and examine the system around it. Every answer is shaped by some combination of the question, training data, system instructions, safety policies, source corpus, retrieval method, ranking choices, product design, and conversation history. People chose or influenced every part of that environment. The output may not express one coherent philosophy, but it is not independent of human choices.
This is not unique to artificial intelligence. A newspaper’s front page, a museum exhibit, a school curriculum, a search result, a court opinion, and a medical guideline all reflect decisions about relevance, evidence, authority, and presentation. The best examples make those decisions disciplined and inspectable. The worst conceal advocacy behind a claim of neutrality.
The National Institute of Standards and Technology (NIST) provides a related, narrower risk-management taxonomy. It identifies systemic, statistical or computational, and human sources of bias, and treats artificial intelligence as part of a larger human and institutional system rather than an isolated algorithm.[3] The agency also warns that bias cannot be assessed meaningfully without a task and context. A system is not biased in the abstract; it is biased relative to a use, a population, a measure, or a consequence.
This matters because “remove the bias” is often not an executable instruction. Which bias? Measured against what baseline? For which people? In which application? At what cost to accuracy, pluralism, privacy, or another form of fairness?
My proposed standard is artificial intelligence whose relevant sources, assumptions, context, uncertainty, and alternatives are visible enough for people to judge, rather than an impossible perspective-free system.
When a viewpoint becomes a trust problem
Recognizing that interpretation is unavoidable does not make every answer equally good. “That is just a point of view” can become a convenient excuse for error, propaganda, or discrimination.
A viewpoint becomes a serious trust problem when the system presents a contested judgment as fact, conceals a source or instruction that materially shapes the answer, or excludes credible alternatives without explanation.
The problem becomes more consequential when the system claims neutrality while advancing a particular interest, expresses more certainty than the evidence supports, or is applied outside the population and conditions in which it was tested. It is especially serious when errors fall unevenly across groups, or an affected person cannot inspect, question, or appeal the result.
Harmful bias does not require intent. A system can discriminate without intending to; an unrepresentative dataset can cause harm without a developer intending it; and an institution can narrow the answers people receive without announcing a preferred ideology.
This is where the language of perspective must not erase the language of responsibility. The software has no moral intention, but the people and institutions that select, deploy, and rely on it remain responsible for reasonable precautions, for foreseeable consequences, and for responding when unexpected harms emerge.
The same principle, with different standards
Transparency is not one universal checklist. What a useful answer must disclose depends on the type of question and the stakes.
Theology
The reader’s eschatology question is interpretive. The underlying texts exist, but capable readers disagree about how to interpret prophecy, how biblical covenants relate to Israel and the church, and whether the millennium described in Revelation should be understood literally or otherwise.
An artificial-intelligence answer should identify the principal framework it is using, name the sources or tradition that shape it, distinguish quotation from interpretation, acknowledge credible alternatives, and state uncertainty where the source tradition itself is divided. It should not present an interpretive conclusion as though it were a laboratory measurement.
That does not mean every interpretation is equally valid. Evidence still matters. Texts can be misquoted. Traditions can be described inaccurately. Arguments can be inconsistent. Some interpretations may be better supported than others. Disclosure allows the reader to evaluate the reasoning instead of mistaking an invisible framework for a neutral answer.
It also preserves a necessary boundary. Artificial intelligence can help retrieve passages, compare commentaries, identify assumptions, and formulate questions. It is not a theological authority. The judgment belongs to the people and communities doing the interpreting.
Medicine
Medical advice requires a much stronger standard because an error can cause immediate physical harm. Here it is not enough to say that a recommendation reflects one clinical viewpoint.
For consequential clinical decision-support software, the answer should disclose the intended patient population, the relevant inputs, the quality and representativeness of the supporting data, the basis for the recommendation, important limitations, known and unknown patient-specific factors, and the current sources it relies on. A qualified professional should be able to review the basis independently rather than accept the output because it sounds authoritative.
That is substantially the logic in the United States Food and Drug Administration’s (FDA) current guidance for certain clinical decision-support software. The agency emphasizes enabling health care professionals to independently review the basis for recommendations, including inputs, algorithms, validation results, limitations, data representativeness, and patient-specific information.[4] It also warns about automation bias, the tendency to rely too heavily on an automated suggestion. The point is not to replace clinical judgment but to augment it: the software supplies analysis, while the professional evaluates its basis and remains responsible for the decision.
Education
Education falls between those examples. Some questions have well-established answers. Others are genuine scholarly disputes. Still others concern civic values, literature, history, or public policy, where selection and framing are part of the lesson.
An educational system should distinguish consensus from controversy, identify the curricular or disciplinary frame, represent credible alternatives in proportion to their evidentiary standing, and help students trace claims back to sources. It should not manufacture false balance by treating every objection as equally credible. Nor should it hide real disagreement for the sake of a simpler answer.
Most importantly, the system should help students reason, not merely hand them conclusions. In practice, that means teachers decide when to use the tool, students can inspect the sources and assumptions behind its answers, and both can challenge or reject its output. The United States Department of Education’s 2023 report makes the same division of responsibility: keep humans in the loop, keep teachers in charge of major instructional decisions, and let educators, not software, set educational goals.[5]
Useful disclosure should expose the important structure when it matters without making every answer longer.
What useful disclosure can and cannot do
Useful disclosure has limits, and four practical complications deserve attention.
First, transparency cannot reveal how a large model arrived at every output; even developers cannot always trace a sentence to a specific training example or provide a complete causal explanation of a model’s internal activity.
A perfect mechanistic explanation is not the only useful disclosure. The people or organizations that build and operate a system can identify its intended use, source set, retrieval boundaries, governing instructions, evaluation results, known limitations, and confidence conditions. A product can distinguish a quotation from a generated summary, show which documents supported an answer, identify disputed questions and the main credible frameworks, and tell the user when current information or professional review is required.
Second, disclosure must not manufacture false balance: present credible alternatives in proportion to the evidence supporting them.
Third, ordinary users cannot audit every answer, nor should they be expected to. I do not personally inspect every component in an aircraft before boarding it or reproduce every clinical study before accepting treatment. Trustworthy institutions reduce the verification burden through standards, testing, professional duties, audits, monitoring, and accountability.
Artificial intelligence needs the same division of labor. Those who build and operate AI systems must test them and disclose relevant limits. Deploying organizations must choose appropriate systems and controls. Professionals must preserve their duty of care. Regulators and independent evaluators must examine consequential uses. Users must apply skepticism proportionate to the stakes.
That division of labor does not make individual judgment optional. Artificial intelligence makes answers faster and easier to obtain, while consequential use often demands more human cognition: framing the problem, noticing omitted context, judging source quality, and knowing when to seek independent review. Critical thinking becomes more valuable as polished answers become cheaper.
Fourth, context must clarify discrimination, not reclassify it as a viewpoint. Disclosure is useful only if it helps people identify and correct the underlying harm.
Context helps us name the failure without making it disappear.
Calibrated trust
Two tempting shortcuts are blind trust and categorical distrust.
Blind trust accepts a fluent answer because it is convenient, confident, or agreeable. Categorical distrust rejects an answer because software produced it or because the system once made a mistake. Neither position effectively evaluates the actual claim.
Calibrated trust is harder. It asks enough questions to match confidence to evidence and consequence:
What exactly is the claim?
Is it factual, predictive, interpretive, normative, or a recommendation?
Which sources and time period support it?
What assumptions or framework shape the answer?
What relevant personal, institutional, or historical context may be missing?
Are there credible alternatives, and does the answer represent their evidentiary weight fairly?
What happens if the answer is wrong, and what level of independent verification does that consequence justify?
Those questions are not unique to artificial intelligence. They are the questions I should ask of a consultant, news story, study, sermon, investment thesis, medical recommendation, or my own confident memory.
Artificial intelligence makes the discipline more important because it can synthesize an answer quickly and remove many of the cues we once used to recognize a source’s point of view. A book has an author. A newspaper has a masthead. A church states its beliefs. A study identifies its methods and limitations. A generated answer may arrive without any comparable label.
The absence of visible authorship, methods, and institutional identity can feel like neutrality, but it often lacks context.
What would change my mind?
I would revise the five-part framework if evidence showed that the distinctions consistently confused users, concealed important forms of harm, or failed to improve decisions. I would become less confident in disclosure as a remedy if well-designed studies showed that source, assumption, uncertainty, and alternative-view information did not help people calibrate trust, or if it predictably overwhelmed them and worsened outcomes.
I would also revise this position if a better practical framework emerged, one that separated technical performance, distributive harm, interpretation, and context more cleanly.
Criticism becomes useful when it is precise enough to act on. That precision should sharpen criticism, not protect artificial intelligence from it.
Can Software Cheat? argued that capable software is not a moral agent. This paper adds an equally important qualification: the absence of a mind inside the machine does not make its output neutral. The answer still reflects human choices about sources, objectives, access, instructions, testing, and use.
Once those choices are visible, we can ask the next question more intelligently: What, specifically, are we afraid artificial intelligence will do, by what mechanism, and under whose control?
Questions, corrections, or disagreements are welcome. You can reach me directly at dave@aworkingmodel.com.
Sources
BibleQuestions.com, “About Us” and home page, accessed September 19, 2026.
John MacArthur, “Why Every Calvinist Should Be a Premillennialist, Part 1”, Grace to You, March 25, 2007.
National Institute of Standards and Technology, Towards a Standard for Identifying and Managing Bias in Artificial Intelligence, NIST Special Publication 1270, March 2022.
U.S. Food and Drug Administration, Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff, revised January 2026.
U.S. Department of Education, Office of Educational Technology, Artificial Intelligence and the Future of Teaching and Learning: Insights and Recommendations, May 2023.
Consumer Financial Protection Bureau, Consumer Financial Protection Circular 2022-03, “Adverse Action Notification Requirements in Connection With Credit Decisions Based on Complex Algorithms.”
National Institute of Standards and Technology, AI Risk Management Framework Playbook.
Federico Bianchi et al., “Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale”, FAccT ’23, DOI 10.1145/3593013.3594095.
