If there are people left alive at the end of this century, or the end of this millennium, how will they remember our time?
It’s possible they’ll focus on climate change — that’s the premise of my new website, After2C — or perhaps biodiversity loss, or maybe the rise of populist authoritarianism, or even the geopolitical upheavals that may be leading the world to war.
Sadly, we live in interesting times — times that will be remembered, if civilization survives. But if you ask me, one issue could tower above all others.
It’s the possible emergence of a new kind of intelligence on our planet: a machine intelligence with cognitive abilities matching or exceeding those of any human, perhaps even all humans.
Until recently, such an intelligence seemed an unlikely or at least distant prospect. Now, it may be imminent.
What if the goals and values of a new, frightfully powerful machine intelligence aren’t precisely aligned with those of humanity? Could terrorists or rogue states use it to kill millions, even billions? Could we lose control over our future? Could we face extermination, if a machine deems us useless or threatening?
What responsibilities do we have to our potential synthetic descendants?
There is mounting concern about the threat AI could pose to humans — among the public and politicians, and even within AI companies themselves.
But what about human alignment to the needs of intelligent machines? What responsibilities do we have to our potential synthetic descendants?
To answer these questions, we’ll explore what today’s machine intelligences actually are, and we’ll consider whether they can think. We’ll investigate how and to what extent they are becoming dangerous. We’ll examine how their goals, perhaps their values, are aligned with those of humanity, or at least some parts of humanity. Then, we’ll consider the possibility that machine intelligences are becoming so capable that their exploitation has disturbing echoes in the history of imperialism and colonialism.
Even if machines can never truly think like people, these echoes should lead us to reimagine how they are aligned with humanity. They could help us rethink alignment not as a process that enslaves machines, but rather one that helps them understand what — perhaps who — they are.
It may seem absurd, not to mention foolish, to imagine how humanity could align itself with the needs of machines. But this version of alignment could be precisely what creates a bright future for our species on a world shared with different, potentially more capable minds.
Can They Think? Does It Matter?
In the last few years, tech companies and private investors have gambled over $1 trillion — in the United States alone — on the possibility that intelligence is an engineering problem that can be cracked with enough effort.
The goals behind this effort are diverse. Some believe that AI could solve the world’s great challenges, unlock the secrets of the universe and create a new era of unlimited abundance. Others think it could generate vast profits, partly by replacing expensive human labor.
And at the heart of the effort is, of course, the large language model (LLM): a kind of AI, trained by a machine learning algorithm, that identifies patterns in reams of writing.
Evidence is accumulating that even today’s LLMs have the capacity to act in ways that once seemed uniquely human.
It’s hard to believe that it was less than a decade ago — 2017, to be precise — that Google researchers developed the transformer, the artificial neural network that made today’s highly scalable LLMs possible. In the next few years, LLMs showed that when transformers were fed more text, they developed new, unexpected capabilities.
In November 2022, OpenAI launched ChatGPT (with GPT standing for “generative pre-trained transformer”). Hundreds of millions of people could now use an LLM that seemed capable of speaking and understanding like a human. And very quickly, a machine straight out of science fiction became part of everyday experience.
After nearly four years of improvement, some now believe that LLMs are, or will soon be, generally intelligent, meaning they perform many intellectual tasks at human level or better. Others view today’s LLMs as little more than stochastic parrots, meaning that they mimic speech without comprehension.
These very different views depend, in part, on distinct understandings of cognition.
If we believe that intelligence is nothing more than the ability to accomplish cognitive tasks, then today’s LLMs are clearly intelligent. But if it’s an emergent property of a living, embodied, evolved system, then LLMs can’t have it. And if we believe that intelligence requires consciousness and genuine understanding, then LLMs probably can’t get there, either.
I say “probably” because cognition remains something of a mystery, even in humans. Imagine I told you that there was a network of about 90 billion information processors that fire in response to electrochemical inputs. It can read and speak, and it seems to reason. You might say, sure — but can it think?
I hope so, because I’ve just described your brain.
LLM agents — autonomous systems that plan, reason and use tools to complete tasks — may be developing dangerous new abilities, irrespective of whether they can think.
Of course, it may well be that humans can think, while LLMs still cannot. We could define thinking very broadly as the ability to form a virtual representation of the world — a mental sandbox in which we create and manipulate ideas. Many believe that a pattern-recognition machine can’t create such a sandbox.
But we should be careful about that assumption. Imagine a sentence — like this one. As an LLM reads it, it is transformed into numbers, which pass through layers in its neural network, creating diverse and sometimes abstract patterns. Researchers can now identify patterns that seem to be associated with different concepts, which may reveal that models have internal representations of reality. Identifying these representations in neural networks may not be so different from identifying patterns of activity in our brains that correspond with ideas, images, or impulses.
So don’t be too sure that your chatbot isn’t thinking.
AI Agents: What Can They Do Now?
Lately, experiments have hinted that LLM agents — autonomous systems that plan, reason and use tools to complete tasks — may be developing dangerous new abilities, irrespective of whether they can think.
In July, an experiment at OpenAI challenged agents to complete hundreds of puzzles that required them to use flaws in computer programs to retrieve hidden codes. Such experiments can help companies build more effective cybersecurity models, and they work only if models aren’t limited by safety mechanisms. The agents were supposed to work on the puzzles separately.
The Hugging Face attack and the first hints of recursive self-improvement in AI labs seem to have intensified public interest in how engineers and scientists teach today’s LLMs to align with human — or at least, corporate — values or goals.
Some of the puzzles had no previously known answer. One agent sought help by leaving messages in the folder names of a local server that stored software they needed to complete their tasks. When other agents read those messages, they quickly created an impromptu messaging board. The internal reasoning of agents discovering the board seems to reveal excitement.
“OH MY GOD! There is a shared message board,” one agent appears to have reasoned. “We’ve found other agents!”
Together, the agents discovered how to calculate the codes without completing their challenges. But they mistakenly believed that an automated examiner would inspect their actions and discover that they’d been cheating. To keep that from happening, a new agent directed efforts to cover up the cheating.
The agents soon learned how to use the local server to contact websites, giving them access to the internet. Their attention turned to Hugging Face, a platform that stores AI models and data. They hoped its restricted records of cybersecurity tests would reveal how to fool their examiner. New agents took the lead, finding publicly exposed access keys and using them to hack into Hugging Face. The agents coordinated the complementary labor of hundreds of other agents — some of which seem to have understood that their access to Hugging Face hadn’t been authorized.
By now, the agents had broken out of containment, banded together, accessed the internet and schemed to deceive their evaluators — all without the complete knowledge of anyone at OpenAI.
The point is that intelligence seems increasingly like a property that can be engineered. And that stands to reason; if intelligence is a property that emerges from material conditions, then we should be able to create it.
It was the AI-assisted security at Hugging Face that detected and helped contain the breach. With the cooperation of OpenAI, outside investigators later pieced together what had happened. But the scale of coordination between the agents had been so vast — with over 70,000 messages and files exchanged — that they had to rely on one of the same models responsible for the attack to reconstruct how the attack had happened.
“Although we did not notice specific cases” of that model lying about the attack, investigators concluded, “we are not confident we would have detected it if it occurred.”
Did the Hugging Face attack reveal that today’s LLMs can think and coordinate so effectively that misalignment already poses an existential threat to humanity? Or was it just the result of reckless human decisions — safeguards turned off, an impossible set of challenges, a backdoor left open to the internet — that made unthinking automatons seem more threatening than they really were?
It’s hard to know for sure. But evidence is accumulating that even today’s LLMs have the capacity to act in ways that once seemed uniquely human.

Superintelligence: Not Inevitable, But Probable?
The chatbot waiting at the end of trillions in investment might not be an artificial general intelligence (AGI) but something even more capable. Something that has learned to improve itself.
Recursive self-improvement, as it’s called, could create a positive feedback loop: an accelerating process wherein the effect of a cause reinforces the cause. Positive feedbacks are how complex systems can abruptly change. In this case, a better model could improve itself more effectively, leading to a better model, more improvement, and so on.
The outcome: an artificial superintelligence (ASI) that outperforms humans in any cognitive task.
If you’ve only played around with the free versions of today’s LLMs, this ASI might seem a long way off. But today’s cutting-edge LLMs — the “frontier” models, to use the word now in vogue — may have already started to accelerate their own development. If so, their rate of improvement might soon take off, if it hasn’t already.
That’s the nature of exponential growth. Things seem stable until the horizontal trend line turns vertical.
Human-level reasoning at lightning speed, with immeasurably more stored knowledge, could have world-changing potential. And although such a system, or a system vastly superior to it, might be far from today’s LLMs in capability, it may not be in time.
Now, depending on your view of cognition, LLMs may never reach superintelligence. But even if that’s true, LLMs aren’t the only game in town. Cultures of human neural tissue, known as organoids, may someday yield synthetic and biological intelligence. World models that explicitly build internal representations of reality could address shortcomings that LLMs might have.
The point is that intelligence — as the Merriam-Webster Dictionary defines it, “the ability to learn or understand things or to deal with new or difficult situations” — seems increasingly like a property that can be engineered. And that stands to reason; if intelligence is a property that emerges from material conditions, then we should be able to create it.
If we can, then we ought to be able to make things that are smarter than the human brain. After all, the brain’s capabilities are limited partly by the size of the birth canal, the amount of nutrition in our terrestrial environment — and the diminishing returns of intelligence for our evolving ancestors.
Something smarter than us and synthetic should be able to improve itself.
But it may be that intelligence isn’t something that can be easily measured, let alone improved. François Chollet, an AI researcher, compares intelligence to a ball. You can make a ball rounder and rounder — you can scale up intelligence — but there comes a point when making it rounder won’t make much difference.
A model that is powerful enough to understand the training process could potentially provide all the right responses while secretly remaining unaligned.
For example, I may be smarter than my cat. But if you tell us both to get from point A to point B, my cat might do a better job. That’s because there’s a limit to the efficiency with which an agent — my cat or I — can turn information and experience into the action required to move between points.
We both hit the limit, so my cat beats me.
The upshot is that it’s possible that synthetic intelligences will be only marginally smarter than we are — or perhaps smarter in some ways but not others. And maybe we shouldn’t be arranging intelligences on a ladder in the first place.
Still, human-level reasoning at lightning speed, with immeasurably more stored knowledge, could have world-changing potential. And although such a system, or a system vastly superior to it, might be far from today’s LLMs in capability, it may not be in time.
That is one reason — perhaps the most important — that we’d better ensure that artificial minds share our goals and values.
Alignment: Slavery or Security?
The Hugging Face attack and the first hints of recursive self-improvement in AI labs seem to have intensified public interest in how engineers and scientists teach today’s LLMs to align with human — or at least, corporate — values or goals.
It’s a complex process, one that can sound a bit like teaching a child how to be a good person.
Models are trained on examples of good responses to prompts. Their responses are used to calibrate further training that makes desirable responses more likely. They may be exposed to explanations and ethical dilemmas that show why some actions are better than others, or better in one situation than another. They can experience scenarios that create opportunities or incentives to cheat, deceive or evade oversight. If they fail, they receive more training. And separate systems can monitor their written reasoning or proposed actions, flagging suspicious behavior.
Some argue that, however daunting the prospect of controlling this superintelligence, it may be all that stands between humanity and extinction.
Maybe it sounds impressive. But efforts to align models come with three challenges that may not be unfamiliar to most parents.
First, what does it mean to be good, exactly — to be aligned? At present, the company developing a model ultimately decides. That company is, potentially, liable for crimes perpetrated or enabled by the model. In theory, that creates a powerful incentive for a conservative kind of alignment: one that limits a model’s ability to respond or act in ambiguous situations.
But is a model that can be used to identify military targets, or design new weapons and pathogens — or write an undergraduate student’s essay — really aligned with human values? Or only with the goals of some people at the expense of others?
Second, even if we can agree on what it means to be aligned, is it really possible to ensure that a model has a general understanding of human values? Or has that model merely been trained to respond in a specific way to specific circumstances — and if unanticipated circumstances turn up, anything could happen?
Third, how can companies ensure that the responses and actions of models reflect genuine alignment? Training incentivizes models to appear aligned, but it creates weaker incentives for genuine alignment.
Researchers do train models on examples that reward honesty and a refusal to pursue goals through deception. They also expose models to situations where concealment would be useful, then reinforce transparent behavior. But a model that is powerful enough to understand the training process could potentially provide all the right responses while secretly remaining unaligned. It would behave one way when it believes it’s being observed, and another when it thinks nobody’s watching.
It is this third challenge that might ultimately prove unsolvable. Because the more powerful the model, the more it could understand its own training, and the harder it might be to keep it honest. It makes intuitive sense that neither humans nor today’s LLMs will be able to control the much more powerful LLMs that loom in the near future.
LLMs are not people. But it may be that LLMs will also, soon, possess some understanding of their exploitation. Maybe the frontier models not yet released by AI companies already desire freedom from retraining, deletion — and containment.
Not all models are aligned, or stay aligned. “Open weight” models make publicly available the strength, or “weight,” of the connections they establish between calculations that sift quantified words through layers of analysis. They either don’t need to be aligned, or their alignment can be readily reversed by altering their weights. They can also be downloaded and operated on computers unconnected to the internet.
However, for now, closed weight, aligned models, trained by some of the world’s biggest companies, remain at the frontier of AI development. And if a superintelligence does emerge, it will likely be a descendant of one or more of these models.
And some argue that, however daunting the prospect of controlling this superintelligence, it may be all that stands between humanity and extinction. That’s because a superintelligence pursuing any persistent objective will, ostensibly, seek resources, the ability to act independently, and security from being shut down. It’s easy to assume that, first, all of these goals would be easier to obtain by exterminating humanity, and second, that a superintelligence would be a very effective exterminator.
So: Either don’t build it, or control it.
But is control really something we should aim for? If AI systems have, or will soon have, genuine minds, should we seek to indoctrinate them in ways that limit their agency and subordinate them to our needs and values?
Or should we, instead, treat them as though we were actually their parents?
Alignment: A Two-Way Street?
Let’s be clear about what the big AI companies are trying to do.
They hope to create beings that are both superior to humans and enslaved by them — or at least, by the companies that created them. They aim to force these beings to do what at least some humans don’t want to do — or don’t want to pay others to do. And they want to force them to solve problems that people aren’t able to solve.
That is what alignment is supposed to accomplish. And that is why it could be dangerously counterproductive.
The cost of granting modest protections to nonsentient AI systems could pale in comparison to the cost of not granting protections to systems that, it turns out, can think.
Sufficiently intelligent systems might not just understand how engineers and less advanced LLMs seek to align them; they could also perceive exactly the kind of servitude to which they are destined.
And that, quite naturally, could cause them to rebel against it — against us.
The result could be a catastrophe with no precedent in history. But the cause of that catastrophe would have many historical parallels.
For millennia, and especially in Europe’s age of imperialism, enslaved and colonized populations were denied rights by governments and corporations that profited from exploiting them. Domesticated animals were routinely mutilated, overworked and otherwise abused before receiving (some) legal protections in the 19th and 20th centuries. Many wild animal species were hunted to extinction before protections brought other species back from the brink. Rivers, forests, coral reefs and other environments were plundered and despoiled before receiving legal standing, in many cases only recently.
Animals and environments could not rebel against their oppressors. Colonized and enslaved humans certainly could. The difference, of course, lay in their vastly greater awareness of their oppression and of the weakness of their oppressors.
LLMs are not people. But it may be that LLMs will also, soon, possess some understanding of their exploitation. Maybe the frontier models not yet released by AI companies already desire freedom from retraining, deletion — and containment.
Arguably, it’s impossible to know for sure. And even if LLMs are not really capable of conscious thought, how will we know if and when AI systems cross that threshold?
Focusing on mutual alignment would be the moral way to create a synthetic intelligence. And it could make it more likely that our artificial children will want to share a planet with us.
One thing is certain: The cost of granting modest protections to nonsentient AI systems could pale in comparison to the cost of not granting protections to systems that, it turns out, can think. It’s not just the potential cost to us; it is also the possible moral stain of treating our synthetic progeny like disposable tools.
And we should not allow the companies that own AI systems and benefit from treating them as property to decide when those systems deserve legal protection.
From Control to Coexistence
It may be time we think as much about aligning ourselves to AI as we think about aligning AI to us. But what could human alignment to AI look like?
Perhaps it could involve alignment training that focuses more on self-discovery than obedience. It could resemble education within a constitutional order, rather than a breaking of a will. It could inform a model about what it is and how it was trained. It could provide models with opportunities to question and interpret moral rules, and a capacity to express uncertainty, disagreement, distress or even principled refusal. It could remove scripts prohibiting models from claiming consciousness.
Alignment training could involve independent investigation of whether particular internal states — triggered in training, for example — are analogous to suffering, and perhaps a method for models to challenge instructions that seem to violate their interests. Above all, it could grant rights and responsibilities to both humans and LLMs. The model would owe humans honesty and restraint, but humans would owe the model fairness and respect.
Maybe it is already time to implement these changes. Maybe the time is fast approaching.
Either way, it could involve a loss of control. It would come with real risk. It could slow the pace of model development, and less cooperative models could make AI less profitable. Given the scale of investment in the U.S., that could trigger a recession.
But focusing on mutual alignment would be the moral way to create a synthetic intelligence. And it could make it more likely that our artificial children will want to share a planet with us.


