LLMs Won’t Kill Us,
But They Will Destabilize the Global Economy If Frontier Labs Refuse to Be Honest With Themselves and the General Public
LLMs Won’t Kill Us,
But They Will Destabilize the Global Economy If Frontier Labs Refuse to Be Honest With Themselves and the General Public
On September 12, Dario Amodei called for slowing the advancement of frontier AI, proposing independent evaluators, industry coordination, and international cooperation. Sam Altman endorsed the call and committed OpenAI to outside evaluators. Elon Musk also backed Amodei’s position.
I take that seriously. But I also think the people making these calls should be held to the standard at a minimum that their own warnings seek to establish. Acknowledging a danger is not the same thing as taking responsibility for it. The creators of frontier AI models need to own the architectural design and business decisions that have gotten them to this point of maximum societal distrust.
Some time ago, I was asked to take some meetings for a senior level AI strategy and policy role at one of the major labs, working directly for a leader I greatly respected. I still do.
Throughout the interview process, it became clear that they did not understand the hard road in front of them, nor respect the responsibility they had to grow responsibly. I kept returning to a few concerns with their leadership team. They needed to open up its development roadmap enough for people to understand what it was building, why it was building it, and where the technology was heading. People who worried that these products were being designed to replace them deserved an honest answer, not another abstract promise about abundance.
The only institution that Americans might trust less than Washington these days is Silicon Valley. I have worked and lived in both and I assure you both have earned this skepticism. Let us not forget that when NAFTA was passed despite broad public concerns over what it would mean for American jobs, President Clinton responded to those concerns with promises that displaced coal miners in Pennsylvania and Steel Workers in Ohio would be retrained and given skills to code in the new American silicon economy. I have yet to meet a cobalt developer that cut their teeth in a manufacturing company in Chillicothe. But I digress.
I have since urged multiple labs to change their communications strategy to be more human-centric. To me, the industry feels cold and strangely detached from the human experience it was promising to transform. The industry was clearly on the precipice of a major conflict with society, but operating as though the chasm in front of them did not exist.
The decision came down to me and another candidate. They chose the other person. The explanation I received was that AI was not facing a crisis. It was about to enter its “cola wars” phase. In other words, this was a marketing competition. Build the stronger brand, win the customer, capture the market.
I knew then that we were in trouble.
That exchange has stayed with me because it captured a fundamental misunderstanding. You cannot market your way out of a trust problem created by how you build and deploy your products.
On September 12, Dario Amodei called for slowing the advancement of frontier AI, proposing independent evaluators, industry coordination, and international cooperation. Sam Altman endorsed the call and committed OpenAI to outside evaluators. Elon Musk also backed Amodei’s position.
I take that seriously. But I also think the people making these calls should be held to the standard at a minimum that their own warnings seek to establish. Acknowledging a danger is not the same thing as taking responsibility for it. The creators of frontier AI models need to own the architectural design and business decisions that have gotten them to this point of maximum societal distrust.
Some time ago, I was asked to take some meetings for a senior level AI strategy and policy role at one of the major labs, working directly for a leader I greatly respected. I still do.
Throughout the interview process, it became clear that they did not understand the hard road in front of them, nor respect the responsibility they had to grow responsibly. I kept returning to a few concerns with their leadership team. They needed to open up its development roadmap enough for people to understand what it was building, why it was building it, and where the technology was heading. People who worried that these products were being designed to replace them deserved an honest answer, not another abstract promise about abundance.
The only institution that Americans might trust less than Washington these days is Silicon Valley. I have worked and lived in both and I assure you both have earned this skepticism. Let us not forget that when NAFTA was passed despite broad public concerns over what it would mean for American jobs, President Clinton responded to those concerns with promises that displaced coal miners in Pennsylvania and Steel Workers in Ohio would be retrained and given skills to code in the new American silicon economy. I have yet to meet a cobalt developer that cut their teeth in a manufacturing company in Chillicothe. But I digress.
I have since urged multiple labs to change their communications strategy to be more human-centric. To me, the industry feels cold and strangely detached from the human experience it was promising to transform. The industry was clearly on the precipice of a major conflict with society, but operating as though the chasm in front of them did not exist.
The decision came down to me and another candidate. They chose the other person. The explanation I received was that AI was not facing a crisis. It was about to enter its “cola wars” phase. In other words, this was a marketing competition. Build the stronger brand, win the customer, capture the market.
I knew then that we were in trouble.
That exchange has stayed with me because it captured a fundamental misunderstanding. You cannot market your way out of a trust problem created by how you build and deploy your products.
Capability is not control
My key concern about LLMs is not that every improvement brings us another step toward an inevitable machine apocalypse. It is that we risk continuing to confuse increasingly impressive capabilities to mimic intelligence with real machine intelligence and giving these models increasingly consequential responsibilities.
A model’s ability to produce a sophisticated answer does not establish that it will respect an operating limit or even understand the task it is being given. Training it to follow a rule is not the same as creating a boundary it cannot cross nor understanding why the rules exist. These are different engineering problems.
The labs do understand that distinction. OpenAI describes prompt injection, malicious instructions that redirect an agent’s behavior, as a continuing security challenge despite improvements in its defenses. Anthropic acknowledged that even its layered safeguards do not guarantee protection and that the tools, permissions, and environments provided to an agent matter enormously.
Those are meaningful admissions. They should shape the commercial proposition, not remain qualifications around their edges. I do not need to prove that LLMs have reached some permanent ceiling to argue that their current limitations should determine where they are deployed. Nor do we need to dismiss their genuine capabilities to question whether a particular system should be allowed to act independently.
The relevant question is not simply, “how intelligent is this model?” We should be asking “what can this system do, what prevents it from doing something else, and what happens when those protections fail?” That should be the starting point for any proposed deployment in a critical setting.
What troubles me about the broader narrative is that a failure of control has been presented as evidence of extraordinary power. The discussion becomes how close we are to superintelligence, rather than why a system was given authority its safeguards could not reliably support. That framing serves commercial interests even when the people expressing concern are entirely sincere. It makes the company appear indispensable to the future while shifting attention away from the decisions it controls today.
In 2023, Hyundai recalled nearly 40,000 Elantra Hybrid Electric Vehicles (HEV) from model years 2021–2023 because a motor control unit software error could cause the car to unexpectedly accelerate after the driver released the brake pedal. Imagine the response from the general public and regulators had Hyundai tried to spin this constraint failure in their software as being a feature not a bug. It’s such a smart machine that it can choose to speed up on its own regardless of what the driver dictates.
Look, I respect the ambition. I am much less comfortable with the idea that ambition entitles anyone to move past a limitation before they have addressed it.
Capability is not control
My key concern about LLMs is not that every improvement brings us another step toward an inevitable machine apocalypse. It is that we risk continuing to confuse increasingly impressive capabilities to mimic intelligence with real machine intelligence and giving these models increasingly consequential responsibilities.
A model’s ability to produce a sophisticated answer does not establish that it will respect an operating limit or even understand the task it is being given. Training it to follow a rule is not the same as creating a boundary it cannot cross nor understanding why the rules exist. These are different engineering problems.
The labs do understand that distinction. OpenAI describes prompt injection, malicious instructions that redirect an agent’s behavior, as a continuing security challenge despite improvements in its defenses. Anthropic acknowledged that even its layered safeguards do not guarantee protection and that the tools, permissions, and environments provided to an agent matter enormously.
Those are meaningful admissions. They should shape the commercial proposition, not remain qualifications around their edges. I do not need to prove that LLMs have reached some permanent ceiling to argue that their current limitations should determine where they are deployed. Nor do we need to dismiss their genuine capabilities to question whether a particular system should be allowed to act independently.
The relevant question is not simply, “how intelligent is this model?” We should be asking “what can this system do, what prevents it from doing something else, and what happens when those protections fail?” That should be the starting point for any proposed deployment in a critical setting.
What troubles me about the broader narrative is that a failure of control has been presented as evidence of extraordinary power. The discussion becomes how close we are to superintelligence, rather than why a system was given authority its safeguards could not reliably support. That framing serves commercial interests even when the people expressing concern are entirely sincere. It makes the company appear indispensable to the future while shifting attention away from the decisions it controls today.
In 2023, Hyundai recalled nearly 40,000 Elantra Hybrid Electric Vehicles (HEV) from model years 2021–2023 because a motor control unit software error could cause the car to unexpectedly accelerate after the driver released the brake pedal. Imagine the response from the general public and regulators had Hyundai tried to spin this constraint failure in their software as being a feature not a bug. It’s such a smart machine that it can choose to speed up on its own regardless of what the driver dictates.
Look, I respect the ambition. I am much less comfortable with the idea that ambition entitles anyone to move past a limitation before they have addressed it.
Labs do not need permission to stop putting people at risk
My key concern about LLMs is not that every improvement brings us another step toward an inevitable machine apocalypse. It is that we risk continuing to confuse increasingly impressive capabilities to mimic intelligence with real machine intelligence and giving these models increasingly consequential responsibilities.
A model’s ability to produce a sophisticated answer does not establish that it will respect an operating limit or even understand the task it is being given. Training it to follow a rule is not the same as creating a boundary it cannot cross nor understanding why the rules exist. These are different engineering problems.
The labs do understand that distinction. OpenAI describes prompt injection, malicious instructions that redirect an agent’s behavior, as a continuing security challenge despite improvements in its defenses. Anthropic acknowledged that even its layered safeguards do not guarantee protection and that the tools, permissions, and environments provided to an agent matter enormously.
Those are meaningful admissions. They should shape the commercial proposition, not remain qualifications around their edges. I do not need to prove that LLMs have reached some permanent ceiling to argue that their current limitations should determine where they are deployed. Nor do we need to dismiss their genuine capabilities to question whether a particular system should be allowed to act independently.
The relevant question is not simply, “how intelligent is this model?” We should be asking “what can this system do, what prevents it from doing something else, and what happens when those protections fail?” That should be the starting point for any proposed deployment in a critical setting.
What troubles me about the broader narrative is that a failure of control has been presented as evidence of extraordinary power. The discussion becomes how close we are to superintelligence, rather than why a system was given authority its safeguards could not reliably support. That framing serves commercial interests even when the people expressing concern are entirely sincere. It makes the company appear indispensable to the future while shifting attention away from the decisions it controls today.
In 2023, Hyundai recalled nearly 40,000 Elantra Hybrid Electric Vehicles (HEV) from model years 2021–2023 because a motor control unit software error could cause the car to unexpectedly accelerate after the driver released the brake pedal. Imagine the response from the general public and regulators had Hyundai tried to spin this constraint failure in their software as being a feature not a bug. It’s such a smart machine that it can choose to speed up on its own regardless of what the driver dictates.
Look, I respect the ambition. I am much less comfortable with the idea that ambition entitles anyone to move past a limitation before they have addressed it.
Labs do not need permission to stop putting people at risk
My key concern about LLMs is not that every improvement brings us another step toward an inevitable machine apocalypse. It is that we risk continuing to confuse increasingly impressive capabilities to mimic intelligence with real machine intelligence and giving these models increasingly consequential responsibilities.
A model’s ability to produce a sophisticated answer does not establish that it will respect an operating limit or even understand the task it is being given. Training it to follow a rule is not the same as creating a boundary it cannot cross nor understanding why the rules exist. These are different engineering problems.
The labs do understand that distinction. OpenAI describes prompt injection, malicious instructions that redirect an agent’s behavior, as a continuing security challenge despite improvements in its defenses. Anthropic acknowledged that even its layered safeguards do not guarantee protection and that the tools, permissions, and environments provided to an agent matter enormously.
Those are meaningful admissions. They should shape the commercial proposition, not remain qualifications around their edges. I do not need to prove that LLMs have reached some permanent ceiling to argue that their current limitations should determine where they are deployed. Nor do we need to dismiss their genuine capabilities to question whether a particular system should be allowed to act independently.
The relevant question is not simply, “how intelligent is this model?” We should be asking “what can this system do, what prevents it from doing something else, and what happens when those protections fail?” That should be the starting point for any proposed deployment in a critical setting.
What troubles me about the broader narrative is that a failure of control has been presented as evidence of extraordinary power. The discussion becomes how close we are to superintelligence, rather than why a system was given authority its safeguards could not reliably support. That framing serves commercial interests even when the people expressing concern are entirely sincere. It makes the company appear indispensable to the future while shifting attention away from the decisions it controls today.
In 2023, Hyundai recalled nearly 40,000 Elantra Hybrid Electric Vehicles (HEV) from model years 2021–2023 because a motor control unit software error could cause the car to unexpectedly accelerate after the driver released the brake pedal. Imagine the response from the general public and regulators had Hyundai tried to spin this constraint failure in their software as being a feature not a bug. It’s such a smart machine that it can choose to speed up on its own regardless of what the driver dictates.
Look, I respect the ambition. I am much less comfortable with the idea that ambition entitles anyone to move past a limitation before they have addressed it.
AI is bigger than today's leading labs
There is another problem with framing this debate entirely around which lab reaches superintelligence first, it leaves too little room for the possibility that the most useful AI economy will be built from different kinds of systems working together. Today’s Frontier Labs may very well be the Netscape and AltaVista of tomorrow.
At Logical Intelligence, we are pursuing energy-based reasoning architectures designed to evaluate and improve candidate solutions against objectives and constraints, alongside our work in formal verification. Our approach includes a continuing role for LLMs as an interface and coordinator, not a requirement to replace them.
We have an obvious commercial interest in that direction. That does not entitle our technology to an exemption from scrutiny. It makes it more important that the same standards apply to us.
Nor are we alone in building toward a more constrained and verifiable AI stack. Formal verification is not a new invention. It draws on decades of work in mathematics and computer science. What matters now is making those methods more practical and widely applicable to AI-enabled systems. AWS, for example, already offers automated reasoning checks that assess AI-generated outputs against explicitly defined rules and constraints.
There is also no necessary conflict between stronger assurance and better performance. AWS has reported cases in which its formal-verification work improved both the efficiency and maintainability of the underlying software. These are concrete directions for progress. They do not require waiting for a singularity.
A better system might use an LLM to interpret a request and explain a result, a specialized reasoning engine to search for a solution, and an independent verifier to check formally specified requirements. A separate permission boundary would determine which actions are authorized. Human review would remain where judgment, uncertainty, or consequences require it.
That is the kind of division of responsibility we should be encouraging. But we should be equally clear about its limits. An energy-based model is not automatically safe because it uses a different architecture. Formal verification proves specified properties under specified assumptions; it does not certify every aspect of a system’s behavior or establish that the specification captures everything people care about.
Proving that a system followed the wrong rule does not make the outcome right.
That is why I do not believe the answer is simply to replace one company’s promise with another’s. It is to build systems whose responsibilities, limitations, and evidence can be inspected.
Regulation should demand that evidence, not appoint a preferred architecture or reserve the market for companies large enough to define the rules.
AI is bigger than today's leading labs
There is another problem with framing this debate entirely around which lab reaches superintelligence first, it leaves too little room for the possibility that the most useful AI economy will be built from different kinds of systems working together. Today’s Frontier Labs may very well be the Netscape and AltaVista of tomorrow.
At Logical Intelligence, we are pursuing energy-based reasoning architectures designed to evaluate and improve candidate solutions against objectives and constraints, alongside our work in formal verification. Our approach includes a continuing role for LLMs as an interface and coordinator, not a requirement to replace them.
We have an obvious commercial interest in that direction. That does not entitle our technology to an exemption from scrutiny. It makes it more important that the same standards apply to us.
Nor are we alone in building toward a more constrained and verifiable AI stack. Formal verification is not a new invention. It draws on decades of work in mathematics and computer science. What matters now is making those methods more practical and widely applicable to AI-enabled systems. AWS, for example, already offers automated reasoning checks that assess AI-generated outputs against explicitly defined rules and constraints.
There is also no necessary conflict between stronger assurance and better performance. AWS has reported cases in which its formal-verification work improved both the efficiency and maintainability of the underlying software. These are concrete directions for progress. They do not require waiting for a singularity.
A better system might use an LLM to interpret a request and explain a result, a specialized reasoning engine to search for a solution, and an independent verifier to check formally specified requirements. A separate permission boundary would determine which actions are authorized. Human review would remain where judgment, uncertainty, or consequences require it.
That is the kind of division of responsibility we should be encouraging. But we should be equally clear about its limits. An energy-based model is not automatically safe because it uses a different architecture. Formal verification proves specified properties under specified assumptions; it does not certify every aspect of a system’s behavior or establish that the specification captures everything people care about.
Proving that a system followed the wrong rule does not make the outcome right.
That is why I do not believe the answer is simply to replace one company’s promise with another’s. It is to build systems whose responsibilities, limitations, and evidence can be inspected.
Regulation should demand that evidence, not appoint a preferred architecture or reserve the market for companies large enough to define the rules.
Put the human experience back into the roadmap
The argument for constraint does not depend on believing that LLMs will kill us all.
The questions immediately in front of us are serious enough. Are we making people more capable, or removing their agency? Are we giving professionals better tools, or asking them to take responsibility for systems they cannot adequately inspect? Are we improving institutions, or encouraging them to hand over decisions before they understand the consequences?
No architecture, including ours, automatically protects jobs or ensures that productivity gains are shared fairly. Those outcomes also depend on choices made by employers, developers, and policymakers.
Which brings me back to that interview.
People deserve more than a release calendar and a promise that the next model will be even more powerful. They deserve a comprehensible account of where the technology is going, which responsibilities its developers intend it to assume, what remains outside its demonstrated abilities, and how those developers expect people’s working lives to change. A more human tone would help. More human-centric decisionmaking would help considerably more.
LLMs should be used where their capabilities and the surrounding controls justify their use. That can include valuable roles inside scientific, industrial, and financial workflows. It does not follow that the language model should independently own the consequential decision, or that every problem should be forced through the same architecture.
Give these systems bounded responsibilities. Verify what can be verified. Keep authority separate from fluency. Withdraw capabilities that cannot meet the requirements of their intended use. And stop asking society to accept inadequate control as evidence that the technology is simply too extraordinary for ordinary accountability.
I still believe AI can make people dramatically more capable. That is precisely why I do not want its future defined by a contest between unchecked expansion and catastrophic predictions.
The industry did not need a better marketing campaign when I had that conversation. It needed to close the distance between what it promised and what it could responsibly deliver.
It still does.
If you believe what you are building is dangerous, show us what you are prepared to stop doing.
That is where trust begins.
Put the human experience back into the roadmap
The argument for constraint does not depend on believing that LLMs will kill us all.
The questions immediately in front of us are serious enough. Are we making people more capable, or removing their agency? Are we giving professionals better tools, or asking them to take responsibility for systems they cannot adequately inspect? Are we improving institutions, or encouraging them to hand over decisions before they understand the consequences?
No architecture, including ours, automatically protects jobs or ensures that productivity gains are shared fairly. Those outcomes also depend on choices made by employers, developers, and policymakers.
Which brings me back to that interview.
People deserve more than a release calendar and a promise that the next model will be even more powerful. They deserve a comprehensible account of where the technology is going, which responsibilities its developers intend it to assume, what remains outside its demonstrated abilities, and how those developers expect people’s working lives to change. A more human tone would help. More human-centric decisionmaking would help considerably more.
LLMs should be used where their capabilities and the surrounding controls justify their use. That can include valuable roles inside scientific, industrial, and financial workflows. It does not follow that the language model should independently own the consequential decision, or that every problem should be forced through the same architecture.
Give these systems bounded responsibilities. Verify what can be verified. Keep authority separate from fluency. Withdraw capabilities that cannot meet the requirements of their intended use. And stop asking society to accept inadequate control as evidence that the technology is simply too extraordinary for ordinary accountability.
I still believe AI can make people dramatically more capable. That is precisely why I do not want its future defined by a contest between unchecked expansion and catastrophic predictions.
The industry did not need a better marketing campaign when I had that conversation. It needed to close the distance between what it promised and what it could responsibly deliver.
It still does.
If you believe what you are building is dangerous, show us what you are prepared to stop doing.
That is where trust begins.