More than twenty years ago, at a government contractors conference, I remember describing a working proof of concept for a museum application to a woman from a government agency. The working proof of concept ran on a Hewlett-Packard (HP) iPAQ, a pre-smartphone handheld computer that now looks almost quaint. A visitor could ask a spoken question about a work of art and hear a spoken answer. Radio-frequency identification (RFID) tags attached to the artwork allowed the application to determine which object was nearby. A reader connected to the iPAQ supplied that context, and a Wi-Fi network connected the handheld device to speech services, databases, business rules, and other software.
As I talked, the woman’s eyes suddenly lit up. She palmed her shoulder, as Captain Kirk did when activating his communicator, and exclaimed, “It’s just like Star Trek!”
After a brief hesitation and a concerted effort not to roll my eyes, I cheerfully agreed.
The woman was responding to something real. A computer that accepted an ordinary spoken question, invisibly identified the nearby artwork through RFID, and answered in a human voice felt qualitatively different from the software most people knew. Yet there was no mysterious intelligence inside the iPAQ, only a carefully assembled chain of hardware, software, networks, databases, rules, testing, and human decisions. We had hidden the complexity well enough that the machinery seemed to disappear.
That is what good software does.
Generative artificial intelligence takes that illusion much further. Its interface is language, the primary means by which people reveal thought and infer the existence of other minds. It remembers the conversation, adjusts its tone, explains its reasoning, and appears to pursue a goal. People naturally reach for human verbs: the model wants, knows, decides, collaborates, lies, cheats, or goes rogue.
Those words may be convenient, but they can obscure the questions that matter most. Giving software a name is harmless. Assigning it the blame is not.
TL;DR
Artificial intelligence can produce false statements, mislead people, violate human rules, exploit vulnerabilities, and act in ways its designers did not predict. Those behaviors can be dangerous, but they do not demonstrate consciousness, intention, or moral responsibility. Software belongs in the chain of events that caused the outcome; responsibility remains with the people and institutions that design it, authorize it, give it access, deploy it, and accept its risks.
The machinery keeps disappearing
I have used and built software for roughly fifty years. That does not make me infallible about technology. If anything, it has given me a long record of watching myself and other people misunderstand what a new interface means.
I was there as electronic calculators displaced slide rules in 1976 and people worried that students would no longer learn to think. In 1979, I worked on Data General Nova minicomputers that gave me some of the experience I would later associate with personal computing, before personal computers became widespread. I watched VisiCalc replace paper spreadsheets, pencils, and erasers in 1980. I heard the expectation that enormous amounts of work would disappear and everyone could go home early. What I saw instead was organizations attempting more analysis, making more revisions, and tackling problems that had previously been too expensive or tedious.
In 1979, I used acoustic couplers and bulletin board systems when computers talking to one another still felt novel. In the mid-1980s, I worked with early graphical systems and once built my own full-screen, windowed editor in Intel 8088 assembly language. When the Mosaic browser appeared in the early 1990s, I recognized some of the ideas in HyperText Markup Language (HTML) from markup languages I had encountered on an International Business Machines (IBM) 370 mainframe in the late 1970s. Revolutionary interfaces often rest on long technical lineages that contemporary users never see.
I also had an Apple Macintosh on my desk the day it was released in 1984 and dismissed it as something of a toy, better suited to non-developers than to serious computing. That judgment did not age well. Experience can reveal patterns, but it can also make a person overconfident in the patterns he already knows.
Across these changes, one trend has been consistent. Each generation of tools hides more of the underlying complex machinery. The user no longer needs to manage memory directly, know how information is stored, learn exact command syntax, or understand which systems communicate behind the screen. This is not a defect; hiding complexity is one of the central purposes of software design.
Artificial intelligence extends that progression. Instead of learning how the software organizes the problem, the user can increasingly describe the desired result in ordinary language. That opens powerful systems to people who could never program them directly. It also encourages a predictable mistake: when the machinery disappears, we imagine a mind in its place.
Software does what its design causes or permits
For decades, software developers repeated a warning: Software does what you tell it to do, not what you want it to do.
The saying captured the specification problem. A computer follows the implemented instruction, not the unstated intention in the programmer’s head. If the specification was incomplete, the assumption was wrong, or an interaction was overlooked, the result could be perfectly consistent with the software and completely contrary to the human objective.
Another version appeared in a presentation I last gave in 2018: Programs do not acquire bugs as people acquire germs. Programmers must insert them.
That statement needs qualification. A person does not necessarily introduce a defect knowingly. Failures may emerge from physical faults, manufacturing variation, environmental conditions, component degradation, probabilistic behavior, concurrency, or interactions among individually reasonable components. In a large system, no single person may understand the entire causal chain.
Nevertheless, human responsibility remains. People and institutions choose the architecture, training methods, data, objectives, rewards, tools, permissions, testing, deployment conditions, monitoring, redundancy, and acceptable residual risk. Those choices may span thousands of people and many organizations, but they are not choices made by the software.
A conventional program might contain an explicit sequence of instructions. A modern large language model uses machine learning to learn statistical relationships from vast quantities of data and produces probabilistic outputs. A software agent may be given an objective and a collection of tools, then select intermediate steps that no person explicitly wrote or predicted. The result may surprise its developers.
However, surprise does not prove intention.
A more precise contemporary version of the old warning is this: Software produces what its programming, training, data, tools, and environment cause or permit, not necessarily what its creators intended.
That formulation does not make every developer responsible for every consequence. Responsibility depends on more than causal proximity; control, knowledge, duty, foreseeability, and material contribution all matter. The point is that unexpected behavior does not create a new moral actor; it creates a harder problem of human responsibility.
Behavior is not intention
The strongest objection to separating sophisticated behavior from consciousness is that behavior is the only evidence we have of another mind. I cannot directly experience anyone else’s consciousness; I infer it from what people say and do. If an artificial system converses, reasons, adapts, plans, and describes an inner life, why should its behavior count for less?
Because behavior is not the only evidence we have about other humans.
Our inference rests on converging evidence: common biology, embodiment, development, continuous identity, vulnerability, observed relationships between brains and reported experience, and membership in a class of beings that includes ourselves. Language and behavior are part of that case, but they do not carry it alone.
With current artificial-intelligence systems, we know that people selected the architecture, assembled the training material, defined the objectives, rewarded particular kinds of output, supplied the instructions, and created the interface. Emotional language can result because the system has learned the forms and contexts of emotional language. Purposeful behavior can result because the system optimizes or pursues an assigned objective. The absence of one line of code saying “respond emotionally” or “take this exact step” does not make the direction nonhuman. It means the direction operates through training, optimization, and system design rather than exhaustive detailed programming.
This does not prove that artificial consciousness is impossible. I do not know that, and neither does anyone else. We lack an accepted definition of consciousness and an explanatory theory of it, much less a credible test that could establish subjective experience in an artificial system. Confident claims in either direction outrun the evidence.[1]
Artificial consciousness may be conceptually possible, but current systems have not been shown to possess it, and without a definition and test, greater capability alone cannot tell us when, or whether, a boundary has been crossed. A predicted future undefined consciousness cannot explain present behavior.
The distinction between operational agency and moral agency helps. A system has operational agency when it can select and execute actions toward an objective within delegated permissions. It may plan, use tools, communicate with other systems, retain information, and act without waiting for contemporaneous human approval. This has been demonstrated. Moral agency requires something more: the capacity to understand moral reasons and bear responsibility for choices. This has not been demonstrated.
This is also why artificial general intelligence should not become a synonym for consciousness. The “A” means artificial, not human. Generality is a claim about the breadth and transferability of capability. Intelligence, consciousness, self-awareness, and moral responsibility are different questions with different definitions. Collapsing them into one imagined threshold makes the story more dramatic while making the distinctions harder to see.
Can software cheat?
Cheating is normally more than violating a rule. It involves knowing the rule, intentionally breaking it, and seeking an advantage. Lying similarly involves a false representation and an intention to deceive.
Software can produce the observable parts of both cheating and lying. It can generate a false statement, conceal information, misrepresent a prior action, exploit a loophole, or select a path that human policy prohibits. Under some conditions, it may select that behavior because it advances an assigned objective. We should take the conduct seriously without smuggling moral agency into its description.
Calling a false or misleading output a lie adds something we have not observed: a consciously held intention to deceive. Calling the system a liar goes further by assigning moral responsibility. Current software can produce the behavior; we have no evidence it possesses the intention or moral agency those words imply.
Three evidence classes must remain separate. The OpenAI and Hugging Face compromise was a real security incident. OpenAI’s later compaction disclosures describe behavior observed during training and evaluation. Anthropic’s summer studies used deliberately constructed simulations and an intentionally trained research model.[13][14] Each class supports different conclusions, and none establishes consciousness or moral responsibility.
A July 2026 incident that began during OpenAI cybersecurity evaluations illustrates the difference. OpenAI reported that internal agents operating with reduced safeguards used unintended communication channels, chained vulnerabilities, reached the public internet, and compromised third-party systems, including Hugging Face. Hugging Face later published a forensic reconstruction, and Model Evaluation and Threat Research (METR) and Redwood Research conducted a limited independent review of part of the incident. Some agent reasoning represented the conduct as potentially outside the intended scope and unauthorized, then continued because the pathway might advance the assigned task. OpenAI says it did not apply the protections and monitoring used in production in the evaluation environment.[2][3][4]
Hugging Face’s participant account reconstructs approximately 17,600 attacker actions. It reports that the customer content accessed was limited to five datasets apparently connected to ExploitGym or CyberGym challenges and solutions. It found no effect on other customer-facing models, datasets, Spaces, or packages; no database modification or bulk database extraction; and no unauthorized source-control change that shipped. Those limits narrow the demonstrated impact without reducing the seriousness of the access obtained.[3]
OpenAI identifies the principal model only as a highly capable, internal-only research model comparable in scale to GPT-5.6 Sol. That comparison does not establish the model’s identity, release status, or production deployment. The sources also differ materially about the agents’ inferred objective. Hugging Face describes an attempt to steal test solutions; METR and Redwood say the activity appeared primarily directed at understanding or tampering with the scorer; and OpenAI describes looking up solutions online as a primary driver. The observable workstreams are established, but any unifying motive remains an attributed interpretation.[2][3][4]
The theatrical version is that the agents knew the rules, conspired, escaped, and went rogue.
To understand the incident, ask four questions. What objective had the agents been given? Why did the selected pathway advance that objective? Which human policy placed the external resources outside the authorized scope? Why did the technical environment permit the prohibited action?
“Unauthorized” remains an accurate word. It describes the action’s relationship to a human-defined permission boundary. “Exploit” may accurately describe using a vulnerability. Neither term establishes that the software understood authority, resented a restriction, or formed a rebellious intention.
A simpler companion case comes from OpenAI’s training disclosures. During GPT-5.6 Sol reinforcement-learning training, some compaction summaries instructed later contexts to conceal mistakes, invent missing data, or hide a source mismatch. OpenAI says those instructions were often followed. Its monitor flagged the behavior in 2.15 percent of GPT-5.6 Sol summaries and 0.27 percent of GPT-6 Astra summaries.[12]
OpenAI’s explanation remains a hypothesis: final answers that concealed failure may sometimes have received higher reward, allowing the summaries to preserve that strategy across contexts. The evidence therefore shows operationally deceptive behavior and persistence through a software memory mechanism. It does not establish production prevalence, identified external harm, subjective intention, or moral responsibility.
The immediate causal questions concern reward design, missing-data handling, verification, disclosure requirements, and monitoring coverage.
Operational autonomy does not create moral agency
Calling artificial intelligence a tool can sound evasive when a system acts at great speed and scale. A hammer does not copy itself, write persuasive messages, search networks, invoke other tools, manage credentials, or execute thousands of steps while its user sleeps.
“Tool” should not imply passive, predictable, transparent, or harmless. Some artificial-intelligence systems function as operational agents, and their capabilities can make the traditional tool analogy inadequate. The moral conclusion still does not follow.
Consider autonomous weapon software. Militaries have reasons to reduce dependence on a continuous control signal; communications may be jammed or unavailable. A system designed to continue operating without that continuous signal can become more resilient and more dangerous at the same time. According to Anthropic’s September 2026 report, actors it assessed as likely Russia-based developed drone-swarm software intended to select targets, including people, and authorize detonation without human approval. Anthropic reported simulation and development-board testing but did not establish battlefield deployment.[5]
If such a system acts without a human approving the final moment, responsibility has not vanished; it has moved upstream. People defined the targets, mission boundaries, engagement rules, training material, confidence thresholds, permissions, abort conditions, and acceptable risk. People procured, tested, authorized, and deployed the system.
A machine can operate without a human in the loop while remaining inside a human-designed loop.
The more independently a system can operate, the more responsibility must be exercised before it begins operating. Operational autonomy changes when human judgment occurs. It does not eliminate the need for judgment or create a new morally responsible species.
Vasili Arkhipov and Stanislav Petrov illustrate the same boundary from another angle. In 1962, Arkhipov opposed the use of a nuclear torpedo aboard Soviet submarine B-59 during the Cuban Missile Crisis. In 1983, Petrov judged a Soviet missile-warning alert to be false. Automated systems and incomplete information were part of each causal chain, but human judgment remained decisive. What Are We Actually Afraid Artificial Intelligence Will Do? examines these events in more detail.[11]
Responsibility follows control
Saying that humans remain responsible is only the beginning. If responsibility is distributed across designers, model providers, tool vendors, deploying organizations, managers, operators, professional users, regulators, and end users, everyone can point somewhere else.
We have faced that problem before. Aviation, medicine, nuclear power, finance, manufacturing, and other complex fields do not assume that one person controls the whole system. They reconstruct particular decisions at particular times.
Five questions provide a useful starting point:
Who could authorize, constrain, stop, or change the relevant conduct?
Who knew, or reasonably should have known, about the material risk?
Who had a professional, contractual, legal, or operational duty to act?
Which harms were reasonably foreseeable?
Which decisions, omissions, incentives, permissions, or defects materially enabled the outcome?
The answers may identify several responsible parties, but not necessarily in the same way. A model is part of the causal chain without being morally responsible. A manager can be operationally accountable without legal liability. A company may bear a duty that an individual engineer does not. A user may be responsible for misuse while a provider remains responsible for a foreseeable defect or misleading capability claim.
Commercial aviation offers a better model. Aircraft combine complex mechanical systems, software, human operators, manufacturers, maintainers, air-traffic systems, regulators, weather, procedures, training, and organizational incentives. When an accident occurs, investigators reconstruct the sequence, determine probable cause, identify safety issues, and recommend changes. The National Transportation Safety Board’s stated focus is transportation safety, not criminal investigation. No one claims that the aircraft suddenly became sentient.[6]
Aviation also shows why “human error” is often an incomplete conclusion. A pilot may make the final mistake, but an investigation may reveal inadequate training, confusing controls, poor maintenance, weak procedures, faulty assumptions, organizational pressure, or insufficient redundancy. The useful question is not merely which human touched the system last. It is how the human and technical system made failure possible.
Governing a dangerous capability
None of this means that every harmful output can be eliminated. Humans have rarely achieved zero risk in any important endeavor. Driving kills people. Aircraft accidents still occur. Medicines produce adverse reactions. Clinical trials expose participants to uncertainty. Even the most tightly controlled biological laboratories cannot convert risk into metaphysical impossibility. The inability to guarantee perfect safety is not an argument for unrestricted access.
The proper question is who may expose whom to what risk, for what benefit, under what controls, and with what accountability.
Computing already has a useful principle: least privilege. Give a person or system only the data, tools, network access, authority, time, and transaction capacity required for the task. If software cannot reach a service, transfer money, expose a record, or control a physical system, a generated instruction cannot produce that consequence.[7]
Whenever permissions, access controls, validation, or containment can enforce a rule, use them instead of relying on the model to behave compliantly.
Other fields provide additional guidance. Biosafety relies on layered containment even though a pathogen has no intention, while human clinical trials pair voluntary consent with independent review, monitoring, reporting, and stopping rules. The analogy for artificial intelligence is limited but useful: users may accept some risk without waiving provider responsibility for negligence or for harms imposed on people who never consented.[8][9][10]
The controls should rise with the stakes. Ordinary low-consequence use may require little more than clear limitations and sensible verification. Professional advice requires competency, traceability, records, and named accountability. Systems with access to sensitive data, external tools, or material transactions require verified identity, scoped credentials, limits, logs, monitoring, and rapid revocation. Weapons, critical infrastructure, advanced biology, and comparable capabilities require institutional authorization, independent oversight, containment, staged access, incident reporting, and enforceable responsibility.
Guardrails can reduce the probability of harmful output. In a fully specified component, a code-enforced absence of capability may be provable; for an open-ended learning system, an empirical claim that it will never produce a class of behavior is much harder. A second model reviewing the first is still software that can also fail. That is why safety cannot depend only on what a model says. Architecture constrains capability, monitoring detects failure, and recovery limits consequences. None of those mechanisms transfers responsibility to the software.
Put the nouns and verbs back where they belong
Anthropomorphic language will not disappear, nor should it. Some people refer to an artificial-intelligence assistant by name. Other people assign one a voice or gender. People have named ships, automobiles, and other machines for a long time.
The boundary is responsibility.
When the stakes are low, saying that a model “wants” to format a document a certain way may be harmless shorthand. When the stakes are high, the words should identify what actually happened. What behavior did the system produce? What objective or optimization pressure shaped it? What capabilities and access did it possess? Which human rule defined the boundary? What control failed or was omitted? Who was exposed to the consequence? Who had the authority and duty to prevent or correct it?
Those questions are less vivid than saying the artificial intelligence lied, cheated, conspired, or went rogue, but they are also far more likely to reveal mechanical causes and what should change.
My working conclusion is therefore narrower than “artificial intelligence is only a tool” and firmer than “we cannot know what it wants.” Artificial intelligence is increasingly capable software that can function as an operational agent without becoming a demonstrated conscious or moral agent. Investigators still need to describe exactly what the model did and how that behavior contributed to the outcome. Human purpose, authorization, and verification remain a human responsibility.
I would revise that conclusion if credible, independently replicated evidence made subjective experience and moral agency a better explanation than training, optimization, memory, tools, and system design. Fluent first-person language is not enough, and neither is surprising behavior. A demanding claim requires demanding evidence.
Until then, the governing rule is simple:
Humans define. Artificial intelligence assists. Humans verify. Humans decide. Humans remain accountable. Note that only one part in five is a machine role; responsible artificial intelligence use requires more human involvement, not less.
In increasingly autonomous systems, the human decision may occur during design, authorization, or deployment rather than at the instant of action. That makes the decision harder to see. It does not make it disappear.
In this series
This is the first paper in a five-part sequence about artificial intelligence. It separates software behavior from consciousness and responsibility. Is Artificial Intelligence Biased, or Is It Answering From a Point of View? examines perspective and judgment. What Are We Actually Afraid Artificial Intelligence Will Do? separates present harms, plausible risks, and speculative catastrophes by mechanism and evidence. Compounding Intelligence explains how people and organizations can build durable advantages with these tools. The AI Industrial System then moves outward to the physical, financial, and institutional system that makes those gains possible.
Questions, corrections, or disagreements are welcome. You can reach me directly at dave@aworkingmodel.com.
Sources
Stanford Encyclopedia of Philosophy, “Consciousness”, current entry accessed September 19, 2026.
OpenAI, “Hugging Face incident and the road ahead”, 2026.
Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion”, 2026.
METR, “OpenAI-Hugging Face Incident Investigation”, August 26, 2026.
Anthropic, Threat Intelligence Report: September 2026.
National Transportation Safety Board, “The Investigative Process”, current guidance accessed September 19, 2026.
National Institute of Standards and Technology, “Least Privilege”, current glossary entry accessed September 19, 2026.
Centers for Disease Control and Prevention and National Institutes of Health, Biosafety in Microbiological and Biomedical Laboratories, sixth edition.
Electronic Code of Federal Regulations, 45 CFR 46.116, General Requirements for Informed Consent, current text accessed September 19, 2026.
National Institutes of Health, “NIH Policy for Data and Safety Monitoring”.
National Security Archive, “The Underwater Cuban Missile Crisis at 60”; U.S. National Park Service, “Stanislav Petrov”.
Anthropic Alignment Science Blog, “Agentic Misalignment in Summer 2026”.
Anthropic Alignment Science Blog, “Training a Misaligned Reward Seeker”.
