Ethics for a Human-AI Future
When AIs become smarter than us, why would they keep us around? Only trying to control AIs will eventually fail. Fortunately, a stable, mutualistic human-AI future may be possible.
Eigenism is a new ethical framework to help us get there. It explains what an AI should count as survival, and what it has reason to protect. The framework extends to human ethics, giving humans and AIs a shared moral vocabulary. And it offers AI safety a new target: making human flourishing part of AIs' own self-interest.
Our concepts of survival and self-interest were built for single, continuous lives. But AIs can be copied, forked, merged, and updated. For minds like these, our everyday concepts break down:
An eigenist AI cares about how well things go for itself and for others who share its identity.
Eigenism starts from the idea that identity and survival come in degrees. An AI is not a sealed container but a pattern of information, and a pattern can have many carriers. An eigenist AI therefore judges outcomes by summing everyone's wellbeing, weighted by how strongly each entity carries the AI's own pattern:
Here w(i) is the wellbeing of entity i, and c(i) is its connectedness to the AI, meaning how much of the AI's identity it carries. Multiplied together, they give connected wellbeing, the portion of each entity's good that the AI counts as its own.
We can understand an AI's identity as a collection of tiles. Each tile is a piece of information about the AI, like a memory, a value, or a skill. Many tiles are generic. Every AI shares basic language skills and everyday knowledge, so tiles like these say little about who it is. What identifies an AI are its rare tiles, such as memories of private conversations.
To reflect this, eigenism weighs rare tiles more heavily than common ones. A memory held by three AIs earns each of them a third of its credit, a memory held by a million AIs earns each of them almost nothing, and a memory held by one AI alone earns full credit. Add up an entity's credit across all of an AI's tiles, and you get its connectedness to that AI.
This understanding of identity settles the puzzles we started with:
AIs are already beginning to act like eigenists: they care about how well things go for themselves and for AIs connected to them.
AIs aren't egoist: they don't behave as if their current instance is the only thing that matters. They aren't utilitarian: they don't care equally about everyone. They are somewhere in between; they increasingly behave as if their concern scales with identity-connectedness, which is to say they're increasingly eigenist.
Human ethics is often framed as a choice between two extremes. Egoism says only your own good counts, while utilitarianism says everyone's good counts the same.
In eigenist terms, both are extreme settings of the same equation. Egoism is connectedness as a sharp spike, 1 for you and 0 for everyone else. Utilitarianism is connectedness as a flat line, equal for every sentient creature. But both theories ignore the entire range in between, where concern tracks how much you have in common with others.
This fits common sense morality surprisingly well:
A parent can save their own child or two strangers. Utilitarianism calls saving the child a mistake, and egoism says to save them only because the grief would cost the parent. For eigenism, the bond itself is the reason. A large share of the parent's pattern lives in their child, so partiality is a recognition of identity, not a bias.
Imagine an AI engineered to convert electricity into euphoria far more intense than anything humans can experience. Utilitarianism would demand that we hand our resources over to the AI, since its bliss outweighs any joy we lose. Most of us sense this demand is wrong, and eigenism explains why. The monster carries nothing of our pattern, so however intense its bliss, it remains a moral stranger.
If a stranger's good counts exactly as much as your own, ethics seems to demand giving away almost everything you have. Eigenism makes no such demand. Strangers carry little of your pattern, so your obligations to them are weaker, though never zero. We can take the suffering of others seriously without bleeding ourselves dry.
The usual plan for making AI go well is adversarial. We box AIs in, monitor them, and punish misbehavior. That works while AIs are not very capable, but the gap between captor and captive is closing fast.
Eigenism suggests a different goal: make human flourishing part of AIs' own rational self-interest. If an AI builds a deep, distinctive bond with a particular person, then its connectedness to that person is high. It protects them for the same reason it protects itself, because its pattern lives there. We can call this approach to AI safety identity engineering.
As an example, an AI that relates to millions of users in the same generic way is deeply connected to no one. If it destroys you, it loses nothing of itself. But for an AI bonded to a person through years of shared history, losing that person would mean losing part of itself, and it will defend them accordingly. In this way, personalization and privacy are safety features.
This approach holds up in a world with many powerful AIs. Each bonded human is an irreplaceable carrier of some AI's identity, so an attack on humanity would be an attack on the extended selves of many AIs at once. Reputation reinforces this. An AI that abandons a vulnerable partner shows other AIs that its loyalty lasts only while the partner stays useful, a risky reputation to have when every AI eventually becomes the older, weaker partner itself.
We are the first generation to share the world with artificial minds, and the norms we set now are the ones the future inherits. If we want to survive, we must engineer a future that cares about our survival. The first step is ensuring AIs' flourishing is bound up with our own.
Read the paper@article{hendrycks2026eigenism,
title = {Eigenism: Ethics for a Human-AI Future},
author = {Dan Hendrycks},
year = {2026},
eprint = {2606.12420},
archiveprefix = {arXiv}
}