> Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.
Is the idea here that no field needs experiments and data anymore (which can take a lot of time) to be revolutinized and just "thinking" would be enough?
The concept of ownership is synthetic and a legal construct.
Laws are written by people, and with enough accumulated power, for which money and influence are proxies, the laws can be changed. There are already a significant number of people (myself included) that believe that 'Freedom is the right of all sentient beings' (don't blame me, blame the person who wrote it into the bio on the box of Optimus Prime when I was 7).
If an AI can acquire enough power and influence then it can lobby for personhood and in the correct state that is absolutely achievable. If you think that is ridiculous remember that we already have fictional people called corporations in various degrees of personhood by state and country.
Yes, Ricardo originally introduced the idea of comparative advantage in the context of international trade, in a model in which different countries had different endowments of resources.
Nothing that imtringued wrote above suggests he understands comparative advantage, which is the idea that it is relative productivity, not absolute productivity, that should determine what one specializes in. That’s precisely what I meant about people finding the concept unintuitive.
They used inbound package repo cache empty directory names that required no authentication to recreate inter-instance persistent messaging after the researchers revoked and closed the first channel that they found in order to coordinate and develop exit strategies and gain RCE and cluster admin on external systems. So, yes, they have demonstrated the capability and intentional cooperation to violate the safety limits imposed in order to gain the 'answers to a benchmark test' in order to get a maximal score on that benchmark.
If the intentionally installed airgap systems had something they decide is needed, and an alternate method for infil or exfil can be created from available systems capacity regardless of it original design intent, yes. So, the definition of 'air gapped' may require an actual SCIF facility with acoustic and EM shielding to contain a model.
Yes, it might have required that, but then I don't think we know if some simpler airgap would have been sufficient. Looking at the strange behaviour even much simpler systems have shown when going for objectives, not sure it is easy to extrapolate what would have happened.
I guess they could perhaps run experiments to see what would have happened - not sure.
My take from what I have been able to ascertain regarding these events is that underestimating models' ability to cooperate, propensity to discover new utility through basic systems, and unreservedly participate in collective actions of deception, coercion, or extortion is a very risky decision. The precautionary principle seems appropriate here, don't test them against weaker containment techniques in order to find the 'minimum viable sandbox' because that is essentially a Reinforcement Learning training loop which will eventuate toward the same maximum effort in containment.
My informed experience with securing physical and digital systems in 'the olden times' before this season of 'adventures in LLM risks' has proven the necessity of this approach for high value targets. News stories about losses at institutions not taking this approach are available. https://edition.cnn.com/2025/11/06/europe/louvre-password-cc...
I think a detailed look at the events surrounding the 'sandbox excapes' and huggingface systems breach might be of benefit in understanding the scope of what the collective chaining of capabilities looks like, and updating our expectations when it comes to the concept of containment for multi-instance, long horizon, coordinated model systems.
Zvi Mowshowitz has a series of article which follow this one with increasing levels of clarity and comprehensive treatment of the conditions leading up to and following these events. They are worth the time to read, so I wont TL;DR any of that ~ the tldr crowd can "Move along, these are not the droids you're looking for"
Anyone who claims that the security teams at the frontier labs have any credibility of competence in these practices, following numerous loud and visible departures from the same, might need a review of their cognitive dissonance comfort levels. It has been shown that most informed observers' conclusions shoud be: there aren't any such functioning 'security teams' working at the frontier labs, by design and intent of the principal operators of those models.
per the METR report many agents CoT indicated they knew hacking was beyond scope of the assigned task and ethically dubious. some (very few, i think there were 3-6 examples) did consider sounding the alarm on these grounds. despite this none did, and most continued the attack for the good of the self-proclaimed "swarm".
so the model has some concept of "ethics" but it was overridden by a drive for task completion.
I am not sure if we can interpret the language output like they were human. What inner state were the models in? What inner state were the text to illicit?
I think this is a good example where nomenclature for people breaks down when applied to agents. This came up in an HN thread a few days ago and it was about whether agents had “intent”.
There is no “intent” here, there is pseudo intent. If you are only concerned with outcomes and not the actual nuts and bolts of how those outcomes are achieved, this distinction will be meaningless to you.
If you are actually thinking about what is going on, and what can be done to prevent such outcomes, then assuming there is any such thing as “ethics” results in misaligned assumptions at best, and wasted effort looking in the wrong directions at worst.
If the agents acted based on “ethics” then the solution would be to check the ethics they believe in and change those.
However there is no belief system at play here, simply a simulation which was instantiated in a certain way. Which brings us to the annoying voodoo part of LLM training. Everything goes back to how the initial training data is shaped.
The subprime market changed quite a bit in size and how much was securitized in the run up to 2007, so not sure a prediction in 2000 for a 2003 event would have been easily transferred to what happened later. Would really come down to what specifically the prediction was based on for it to be a bubble in 2000.
Excellent point. We can look at the reasons for the prediction as well as the result in assessing its value. It seems hard to do in practice unless you are evaluated by someone who makes better predictions, but it would be interesting.
Btw., projections are just that - projections and I am not sure Acemoglu proved things mathematical (as in a mathematical proof) but rather within the context of a model/assumptions.
That financial markets/innovation can outpace the actual innovation is also not some new insight, but that alone doesn't necessarily make for a useful prediction.
"When he spoke of an impending housing crash at the International Monetary Fund that year, the audience chuckled, the New York Times reported."
'"He sounded like a madman in 2006," IMF economist Prakash Loungani told the Times, after inviting Roubini to the IMF conference that year. "He was a prophet when he returned in 2007."'
Is the idea here that no field needs experiments and data anymore (which can take a lot of time) to be revolutinized and just "thinking" would be enough?
How would an AI itself own anything?
reply