"Stealing millennium problems" would be more complete if not accurate description. And that 99.99% of the market is not interested in solving millennium problems is the other fact.
These exist. Very elaborate systems in many cases. Otherwise Hermes would probably have 30,000. Like a lot of agent tools they merge a ton of slop though so I wouldn’t consider them leaders.
There is no need to ask. If you want to test SOTA models today, there are obviously only two: GPT-6 Astra and Fable 5.1.
The models listed in the paper are from early 2025 and are no longer relevant, much less on the frontier.
That Claude version is no longer available today, Gemini 2.5 Pro will be shutdown next month, and the OpenAI models are only available via the API today.
Fortunately, a fellow commenter was so kind and did it with Astra. Didn't do that well either [0]. I'm sure GPT-7 will be super mega ASI regardless (since GPT-6 Astra already claimed AGI in the minds of Jen-Hsun, et al.)...
I'll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. "Just" having been trained on every rule, strategy guide and likely most games of chess on the world wide web isn't even enough for an LLM to play that game reliably. Yet the same model could code a competitive chess engine, just like a model struggling to count can write advanced maths papers. Fascinating tools, but tools nonetheless.
Maybe watch some HuskIRL videos to temper your expectations. Sure, frontier models providers may alter their harnesses to better target chess, but that’s lipstick on a pig imo. The models themselves are not, in isolation, capable of solving general tasks. We haven’t modeled intelligence sufficiently. We’re in a local minimum and throwing billions of dollars at a gamble that that local minimum can facilitate the concentration of wealth even further and fully realize the American dream of eliminating the middle class.
Really? It was being sold as a total replacement for jobs like software engineering and being an attorney, but its looking a lot more that its just going to be a tool those professions use and doesn't actually seem to be taking jobs away.
I find it amusing that you're describing a huge misallocation of capital and a society enabling such, and that is the optimisitic scenario (in my mind anyway).
I very much agree that the next models will be better, heck, I still suck at hobbyist training and could probably coax t5 to do better in Chess specifically, just need to get loads of data from Stockfish.
Thing is, given what GPT-6 Astra was trained on and what models of a similar class can do (including developing a competitive chess engine), it is often paradoxical and somewhat surprising how little these models have gained in actually capability that is in the training data, but not RLHFd to hell, so to speak. Tracking the state of pieces, I suspect given similar in Sudoku [0], is what these models struggle with in game settings, whilst tracking the state of code changes can be reliable over 250k tokens. Essentially, for the latter they were trained in the specific manner that lead them to abstract the capability, but that doesn't track to the former, which is a massive difference between LLMs data focused training and human learning.
So yeah, GPT-7 or any upcoming/present LLM could do massively better in Chess than GPT-6 Astra, but not because the approach was emergent out of pure data. Rather, it requires a very specific training data type and stack for a model to gain capabilities that track a specific task long enough to adhere to the rules of a game such as chess.
I'm wondering if instructing it to track the board state in a file would make a significant difference then.
It reminds me of the ARC-AGI-3 issue where not dropping the thinking tokens between turns or something like that + a new context compaction method increased the performance dramatically. However, I think that is not applicable here.
So what is the supposed leap?
One agent per option to change, evaluating the board state that there move would create, by having a army evaluate the remaining piece options and average over that? Wee-Free-Man as a hierarchical army ?
Pet-LLMs trained on one thing?
Honestly, for intelligence I don't know and I doubt anyone can claim to know. Maybe JEPA, there is potential concerning some shortcomings inherent to LLMs but it has its own, maybe scaling up the electron microscope stuff Google just did (though the connections are inferred), maybe future implementations of autoregressive and diffusion LLMs can at some point address its issues after all, maybe something else entirely.
All I know is, AGI, as in actual intelligence, is quite a massive accomplishment to claim and we shouldn't loose sight of that fact, especially as "not being intelligent" does not make these models any less impressive, fascinating to work on or useful in many tasks. Personally, the only thing I am fairly convinced on is that if we were to find a way to create actual intelligence, it likely wouldn't start out as useful as todays LLMs are and may thus be dismissed early. But again, pure speculation on that front.
If for leap you just mean more utility from LLMs as they are, then I'll pretty confidently put my money on higher quality, not more, training data for a wide range of verifiable tasks. What makes maths, coding, etc. comparatively easy to make gains in (though less verifiable tasks can also make similar as seen with the writing in Kimi K2).
I suggest you don't participate if you cannot discuss calmly the opinion that the EU should be more inviting towards the UK, even if you consider the opinion absurd.
How can EU be the one expected to be inviting when they didn't kick the UK out in the first place? The UK is the one that left on its own. The EU doesn't have to do anything, I think they are quite open to the UK or anyone else in the continent joining. I just find it absurd a British person is complaining without the UK first doing the homework and showing joining EU is something tenable from their side via an internal vote about joining.
If you want a disclaimer, I am neither european nor british if that matters. From outside the situation seems very clear to me.
I could go into exhaustive details, but the EU made absolutely no case to the British people why they should be part of the EU, and that case is still unmade. Brexit was an EU political issue happening inside the EU yet it and the British people were treated with ignorance and condescension, like it was a wholly British minor political issue and the very thought of leaving was laughable (openly laughable because Nigel Farage got laughed at in the EU parliament when he put it to them that he was going to make Britain leave, which he succeeded in doing). That case for the benefits of joining has still never been made, rather assumed as self-evident. There has been no publicity at all towards the British people that they could rejoin and it would be beneficial, which is the step the EU would reasonably take yet refuses to do. Hence the British people similarly have no interest in begging to be part of this organisation, the mechanics and political organisation of which is very poorly understood.
Whats wrong in laughing at farage? You left, if you eant to rejoin then you have to make the case and prove you are ready for it. The EU can't read your brain, why should it invite when you left. It should not even invite unlesss you did a poll and proved you want to join.
I agree that he’s a laughable character but they were laughing at the very idea he might do what he was saying he was going to do, rather than taking him seriously. If they’d have taken it seriously, the UK might still be part of the EU. Shows how open and liberal their thinking is to any ideas that don’t fit their paradigm.
The far right wing stupidity needs to run its course. There's nothing outside forces could have done to prevent it, attempts by EU to steer the Brexit issue would only have led to accusations of "EU interference" and casued a stronger Brexit. A few millions dead and total destruction of QoL would generally be the usual for far right wing to run its course like in the last world wars. Immunity only develops after a dose of the disease. I am not saying I like this, it's just how it is.
These parties dislike the EU anyways. So what would be the point in trying to woo them to a cause they don't believe in? Some lessons are best learned from europeans again returning to killing each other en masse. Unfortunately.
The EU didn't participate in the Brexit campaign because that would have rightfully been seen as undue meddling in internal politics of the UK.
In any case I don't understand what you want us to do? Should the EU grovel before the UK and beg it to return? That won't happen. You guys voted to leave. You can come back if you want but don't expect any red carpets.
It's telling that ten years on it's mainly you Brits still kvetching about this.
My main point in this thread was about Ursula proposing that the EU is democratic while she benefits from it not being democratic. I’m more just stirring the pot rather than kvetching. The more distance we get from Brexit it becomes easier to see the defects inherent within the EU project.
I really don't understand this heisenberg nonsense at all. If he wants to join he should say (and prove) that, if he doesn't want to join then directly say that.
Vaguely telling the agent what the issue is and what behavior I expect solves the issue with a fraction of the effort.
Some claim that the tech debt only keeps increasing and that the result will be unmaintainable. This is not my experience, and I don't think it is theirs either. These claims are often entirely speculative.
I, and I think most experienced developers, can recognize the type of code that incurs a maintenance cost down the line; that will make adding new code take longer. And AI writes such code "relatively" frequently. I love having the AI to write code, but I find it extremely important to review it - to make sure that it's correct, understandable, and not going to be a problem later.
I find it unnecessary for most non-critical code, such as client applications.
I doubt that any supposed future extra effort for the AI to add new code is remotely comparable to the upfront effort of you reviewing the code manually.
I know that this is the case today for native mobile apps, and I speak from hundreds of hours of experience over the last four months on such a project where I stopped reviewing the code.
We are already here today, and this balance is only going to further shift to the point where it is obvious that the hands-on approach is no longer competitive.
Everything about what you're said strikes me as sounding like "I don't bother wearing a seatbelt, because my experience is that I don't get in accidents" .. and also "I don't write automated tests, because I already hand tested my code and it works".
And neither one of those statements is very convincing to me.
And what you said strikes me as speculation not based on actual experience in using AI in this way, with a healthy dose of condescension added.
Anyway, I think we shared our viewpoints, and neither of us is going to change their mind until either my project fails spectacularly, or you change your approach in the future to use AI more autonomously.
I've had bugs the agents can't fix or figure out. Sometimes those involve third-party, proprietary, broken code (read: Windows APIs). Sometimes they just involve complex deployment situation on the client code (I work on desktop apps) where the agent can't figure out what's wrong/makes wrong assumptions/goes nowhere. Sometimes the agent is just very dumb and tunnels vision on the wrong fix.
FWIW, I've also had bugs the agent fixed that I probably never would've figured out without LLMs - LLMs are definitely useful! But I need to keep understanding how the code works so I can take over the reigns when the LLM fails.
I had edited my comment at the same moment you added yours. It made an 11 Pro Max I installed it behave almost like a brand new phone as far as snappiness. This blew my mind a bit, given what happened with iOS 26. Mind you, I only have 10 minutes of testing so far.
I'm not a lawyer but I don't think Sam Altman 'knowingly accessed' anything.
Are you sure that is applicable here?
And for the first count with 'knowingly accessed', he would need to have accessed classified national-defense or atomic-energy information, otherwise we are back to 'intentionally accessed'.
The first count is "or any restricted data", not classified material. A technological restriction, is enough.
"Knowingly accessed" has never meant you personally. Operators of a botnet don't know directly what they access. They know that the autonomous software is built to access restricted things.
> or any restricted data, as defined in paragraph y. of section 11 of the Atomic
Energy Act of 1954, with the intent or reason to believe that such information so obtained is to be used
to the injury of the United States, or to the advantage of any foreign nation
I understand it was applied in the case of leaking classified CIA material to WikiLeaks, where a former CIA software engineer was sentenced to 40 years in prison.
The service works just fine and it is a decent way to get the latest news.
For example, rumors about latest AI developments are available there, but not here on HN.
X also works well for media, which on HN you can only see after navigating to an external website.
Many comment sections are rough, but usually not worth visiting apart from the top comments anyway.
I don't like that Ads are hard to distinguish from regular posts.
I believe the main reason for not using X is a political boycott. Many disagree politically, or don't entertain certain delusions about Musk that are popular with progressives.
I'd be more interested in your concrete alternative recommendation on where to get the rumors/insider news about AI rather than your opinion on how I should use social media.
There isn't much TBH. All newswires are even propaganda by selection mostly. I'd say the newswires that still sort of pretend to be news and usually are only biased through selection: AP and Reuters.
Its relevant because those same issues occur with large models. Also, the "crappy small model", was a model trained for safety tasks, and outperformed the frontier lab safety models.
Open models have not yet caught up with February's Mythos checkpoint.
Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.
reply