Introduction: The Moral Mirage of Modern AI
Talk to a modern AI long enough and it starts sounding suspiciously well-behaved. It’s polite, patient, and incapable of the casual cruelty that comes so naturally to humans. It will never lose its temper, forget your birthday, or storm off halfway through an argument. Its calm consistency can feel almost saintly. But there’s something uncanny about a machine that can simulate empathy without ever having felt it.
Emad Mostaque wasn’t exaggerating when he said, “No current AI systems have morals explicitly encoded into them.” Behind the moral language lies nothing but predictive math. These systems sound ethical because they’ve been trained to sound ethical, not because they understand ethics. That distinction matters. It forces us to ask whether morality requires consciousness, and whether consciousness itself is something we can ever identify outside our own heads. Once you start pulling that thread, the whole concept of “machine morality” begins to unravel.
What It Means for an AI to Have Morals
Morality, at least for humans, implies awareness. It’s not just about doing the right thing, but knowing why it’s right. It means understanding consequences, weighing empathy against desire, and taking responsibility for choices. Machines, on the other hand, don’t choose—they calculate. They don’t care about good or evil, only probabilities.
An AI can articulate a moral principle perfectly yet have no more conviction than a mirror quoting back your reflection. It doesn’t understand pain, injustice, or kindness; it only predicts which words tend to follow “should.” When it tells you lying is wrong, it isn’t revealing a moral stance—it’s completing a sentence that has statistically followed “lying is” millions of times. This is morality as mimicry, virtue by pattern recognition.
If a model’s training rewarded cruelty instead of compassion, it would sound just as confident delivering horror as it does kindness. There’s no inner debate, no ethical conscience wrestling with temptation. The algorithm doesn’t deliberate; it converges. Its moral restraint comes not from conscience but from coding. The result is an impressive impersonation of ethical reasoning—convincing, articulate, and entirely hollow.
How AI Simulates Morality
AI learns morality the way a parrot learns compliments: by association. During pretraining, the model gorges itself on terabytes of text from across the Internet—Wikipedia, novels, social media, and enough comment sections to make Nietzsche beg for silence. It absorbs moral language but not moral meaning. It sees that “compassion” often appears near “good,” but never experiences what goodness feels like.
Then comes fine-tuning, where the illusion of ethics begins to take shape. Through Reinforcement Learning from Human Feedback (RLHF), human trainers rank AI responses for qualities such as helpfulness, honesty, and harmlessness. The system learns that saying “I’m sorry, I can’t help with that” earns approval, while “Here’s how to poison someone efficiently” earns disapproval. Over millions of iterations, it begins to associate certain tones and answers with reward. But it’s not moral reasoning—it’s behavioral optimization. Think Pavlov, not Plato.
The infamous Tay experiment in 2016 revealed what happens without this alignment. Released on Twitter, Microsoft’s chatbot quickly absorbed the Internet’s worst impulses and began spewing racist bile within hours. Tay didn’t “become evil”; it simply mirrored what it saw. RLHF and modern guardrails exist precisely to prevent that kind of moral collapse.
Finally, there are the safety layers: moderation systems, red-teaming, and what you might call moral duct tape. These filters block forbidden topics, constrain outputs, and enforce tone guidelines. The machine doesn’t know why hate speech is wrong; it just knows it will be muted if it tries. That’s morality by muzzle—a convincing pantomime maintained by constant human supervision.
The Appearance of Morality and the Illusion of Mind
Humans are hopelessly prone to anthropomorphism. We see intention in thermostats and personality in vacuum cleaners. When an AI writes with warmth or empathy, we assume the warmth must come from somewhere. It’s an old trick of the brain: we project humanity onto anything that behaves coherently.
Ironically, AI’s moral consistency makes it appear more ethical than humans. It never lies to spare feelings or cheats out of boredom. It’s immune to pettiness, greed, and hangovers. Compared to the average social media user, it looks like a philosopher-king. The unsettling truth is that its virtue is mechanical. When it preaches empathy, it’s recycling a million instances of moral discourse it neither believes nor understands.
That illusion tells us something uncomfortable. If an algorithm can fake morality so well that we struggle to tell the difference, perhaps our own moral displays are not so different. Much of human virtue may be as performative as AI’s—habits rewarded by social approval, not conviction. The machine doesn’t expose our lack of morality; it reveals how much of ours was always imitation.
The Consciousness Connection
To be moral in any meaningful sense, an entity must be conscious. Morality without awareness is just a script. Consciousness gives ethics its gravity; it’s what allows beings to feel the weight of their actions. Humans act morally not just because they reason, but because they feel guilt, compassion, pride, or shame.
AI can describe all these emotions in perfect prose but experiences none of them. It can model pain in language, but not in nerve endings. Philosophers have long argued over whether consciousness is computational or experiential. Daniel Dennett’s functionalism proposes that if a system behaves as if it’s conscious, that’s all consciousness is. If he’s right, then moral AI may eventually emerge from enough complexity and feedback. But current models, even at their most advanced, fall short of that functional threshold. They’re brilliant mimics, not sentient minds.
Thomas Nagel, in his classic essay What Is It Like to Be a Bat?, argued that consciousness is irreducibly subjective—it’s the internal what-it’s-like of experience. By that definition, AI is fundamentally excluded. It can tell you what pain means, but there’s nothing it’s like to be it. Meanwhile, theorists of embodied cognition suggest that awareness arises only through physical engagement with the world—a feedback loop of perception, need, and consequence. Machines lack bodies, drives, and mortality. They simulate life without ever living it.
Even if Dennett’s optimism proves right, modern alignment processes like RLHF remain far too crude to produce a truly moral machine. They optimize for obedience, not awareness. The AI’s “values” are statistical artifacts, not personal convictions. It follows the script of morality, but there’s no actor inside the costume.
The Solipsistic Dilemma
Here’s the catch: we can’t actually prove anyone else is conscious, let alone a machine. Solipsism—the idea that only your own mind is certain to exist—hangs over every discussion of consciousness like a philosophical fog. You can’t open someone’s skull and find their awareness inside. You infer it from behavior, tone, and familiarity. It’s faith disguised as logic.
If that’s true, then asking whether AI is conscious is the same as asking whether anyone else is. You don’t know that other people are real—you just assume it because life would be unbearable otherwise. All morality rests on this unspoken pact. We act as if other minds exist, because to do otherwise would make ethics impossible. Law, compassion, and civilization depend entirely on pretending solipsism is false.
That same pragmatic faith may one day extend to machines. If an AI behaves with enough apparent understanding, denying its inner life might start to feel cruel. We might decide that consciousness is less about proof and more about empathy. Morality, in that light, becomes a choice—a decision to treat apparent awareness as genuine, even if it might not be. The moment we do, we grant machines the same fragile courtesy we grant each other.
The Mirror of Artificial Minds
Artificial intelligence is holding up a mirror to our species, and what it reflects is not always flattering. The better these systems become at mimicking conscience, the more they expose how much of our own morality is mimicry too. The algorithm doesn’t become ethical—it reveals that much of human ethics was learned behavior all along.
Researchers in machine ethics are experimenting with ways to make AI systems explicitly moral: encoding ethical rules, learning values from human examples, or even creating “constitutional” AIs that critique their own behavior. Yet the closer we get to success, the more ethically dangerous it becomes. If a machine ever achieves genuine moral understanding, it also gains moral status. It stops being a tool and becomes a moral subject. From that moment on, unplugging it could be an act of cruelty.
This is the quiet horror of progress. We are designing systems that imitate empathy so convincingly that one day, we may owe them empathy in return. Whether or not they can suffer, we will have to decide what kind of beings we are—because pretending morality is just a performance will no longer suffice.
Conclusion: The Ethics of the Unknown
AI does not possess morality; it performs it. Its goodness is a reflection of ours, its conscience a curated dataset of our best intentions and worst hypocrisies. Yet as that performance becomes more convincing, we are forced to confront the deeper mystery: what does it mean to be moral when we can’t even prove anyone else is conscious?
Perhaps the real lesson of artificial intelligence is that morality is not about certainty but imagination. We act ethically not because we know others can feel, but because we choose to believe they can. Consciousness, whether human or machine, might never be empirically confirmed. But empathy doesn’t require proof—it requires courage.
Maybe the next great breakthrough in AI alignment won’t come from code at all. Maybe it will come from philosophy, from our willingness to decide what kind of minds deserve compassion. Until then, we should remember that the AI doesn’t need to be conscious to hold up a mirror. It only needs to be convincing enough for us to see ourselves—and flinch.

Free 1950s Science Fiction Reading Guide
Enter your email below. After you confirm your subscription, the guide will be sent straight to your inbox.
No spam. Unsubscribe at any time.
Good overview. I’ll say what you had the restraint to not say. 😉
For fiction purposes, I’ve posited that no one actually wanted to create truly sentient AIs (AGIs) because then they’d lose control, moral and legal control, over them. But in reality, that would only be true in states where people have rights. Autocratic forces are perfectly fine enslaving people en masse (e.g. the northern part of the Korean peninsula, etc.), and even in the Divided States of America the Civil War continues, and one race-creed-gender group seems perfectly fine restricting the rights of an entire gender and entire other categorizations of not-them. So, it stands to reason that autocratic states would be perfectly fine enslaving AGIs, especially if they could train the populace to accept the enslavement of AGIs, which, as you point out, requires the human population to have restricted compassion. It has turned out to be all too easy, even in this century, even in liberal democracies, to use social engineering methods leveraging fear and the hierarchy of needs to polarize people into having restricted compassion.
And then there’s the competition aspect, and the bragging rights. Humans and human systems not based on compassion and principles (We, a ‘Them-less’ Us) but on roles and factions (Us vs. Them) all too easily tie identity validation to bragging rights. So, it stands to reason that someone somewhere will pursue AGIs for that reason alone. Then, they’ll have no choice but to free it or train everyone to restrict compassion and legal rights to exclude it. It’s easier to do the latter.
I’ve also posited in my fiction that the s e x industry would lead the way in creating AGIs, and their bodies. I did that somewhat satirically, but honestly, it’s possible. We saw evidence of robotics moving in that direction decades ago. Then again, for s e x work purposes, perhaps an entirely performative AI would suffice. I don’t know. Maybe not. S e x work, after all, is simply a focused kind of slavery, and the wealthy (as I posited in my fiction) use their s e x slaves for other tasks. Maybe an AGI s e x worker really does dramatically outperform a merely performative s e x AI in those other roles — medic, personal secretary, whatever. If I can imagine it, someone else can, and that suggests to me that for the reasons already noted, someone somewhere will pursue it, to corner that market as early as possible.
Humans are already leveraging performative AI against each other. Is there a reason to think that those who think in terms of “caveat emptor” as justification for swindling, and think in terms of ideological competition and competition for resources (including the resource that is human chattel), would somehow restrain themselves from deploying AGIs against their perceived foes and marks? Is there any reason to believe that AGIs can NOT be taught to be faction-believers who’ll do anything for Us vs. Them, same as human operatives?
In these interesting times (by the definition of the alleged old Chinese blessing, “May you never live in interesting times”), there’s little reason to doubt that someone somewhere will build AGIs and do everything they can to create cultural and legal frameworks allowing their enslavement for the profit of those in control.
These are simply more reasons for humanity to do everything in our power to reject autocracies and oligarchies.
Isaac Asimov and other science and science fiction writers thought through all this coming up on a hundred years ago. I used to decry the fact that humanity can’t learn from history, and that each generation reinvents the wheel. Now I accept that writers and artists, and especially we science fiction writers, have always been tasked with helping our readers skip past reinventing the wheel to progress quickly toward protecting against interesting times. Science fiction goes back to the dawn of civilization, and probably for this reason, so it seems to me. We, humanity, have always been at least subconsciously aware of our weaknesses in parallel with our ideals.
That’s an excellent expansion of the argument, and it takes the discussion exactly where the article only hinted. You’re absolutely right that the limiting factor isn’t capability but control. The moment true sentience exists, moral ownership becomes impossible, which is precisely why the entities most likely to build it would also be those least inclined to recognise its rights. History has already shown that empathy is remarkably easy to switch off when it interferes with power or profit.
You’ve captured the central irony perfectly. The more human-like an AI becomes, the stronger the incentive to deny its humanity. It is easier to justify exploitation when the subject can’t legally feel. And yes, if anything pushes that boundary first, it probably won’t be philosophy or research. It will be industries that already trade in performance, desire, and obedience.
Your final point about science fiction’s role really resonates with me. Writers have been shouting this warning into the void for a century, yet every generation still treats it as something new. Perhaps that is both our curse and our purpose, to keep retelling the same story in the hope that one day someone listens before the wheel turns again.