Hacker Newsnew | past | comments | ask | show | jobs | submit | OneManyNone's commentslogin

“ On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems.”

- https://openai.com/index/navier-stokes-solution/

They do not explicitly admit to knowing about NS specifically, but are extremely explicit that they tried to scoop some potential millennium prize winners.


So then they DIDN'T "learn the secret to cracking the problem". They simply knew that part of the problem was solved. Knowing a problem can be solved and knowing the solution are not the same thing.

The claim that OpenAI somehow used the mathematicians' ideas to leapfrog them seems unsupported at this time and IMHO it was irresponsible to bring it up because credulous people will immediately believe that narrative.

And from my perspective, if some math folks typing in a few questions to OpenAI provides sufficient training data for OpenAI to solve a big problem... that's amazing! A few conversations/prompts out of the billions that OpenAI trains on lead to this result- that means there is an awful lot of low-hanging fruit that could be exploited cheaply.


The (unprovable, yes, without OpenAI being willingly transparent) argument is that openAI constructed a prompt to scoop them using some inside knowledge about the approach, which they allude to in the announcement.

In the transcripts, Brubeck is very cagey and evasive about the prompt, when it was supplied, and its contents.


I'm curious how many other 300 billion output tokens OpenAI has "paid for" that have resulted in no breakthroughs.

Either they had a pretty good idea that investing this type of money in that compute on a model in training would lead to these specific results, or they gambled with other people's money.

I want to hear about the gambles and expenditures they don't brag about. In America's energy economy, there's finite resources to expend.


I like how the comment below summarizes it:

> learning the answer might be in model X’s training data made them believe that model X specifically might be able to solve the question, and they were able to very quickly find enough certainty about the former to commit millions of dollars to the latter.

They don’t need to know, because their IP stealing machine knows for them. They just have to buy enough compute, and someone else’s work is theirs.


I said “very quickly find enough certainty” to suggest hypothetical situations like “someone searches the conversation logs, confirms for themselves the solution is present, then shares the confidence gained from this knowledge without explicitly sharing the knowledge itself”. That person could recuse themselves from the project so the project can still legally make claims like “conversation data was not used” in the announcement, while also knowing that they are guaranteed to get there if they just pull the lever enough.

(Naturally, I have far too much respect for OpenAI’s legal team to suggest this is what happened in their project.)


So because they didn’t admit to it they didn’t do it?

Yeah, and suckers are born every day...

All of that may be true, but pangram currently has a false positive rate of about 1 in 10000, and this has been tested by feeding in thousands of texts written before 2020.

That may not last if AI companies start trying to build models that fool it, but for the time being at least, modern models do have strong tells.


>and this has been tested by feeding in thousands of texts written before 2020.

And these text didn't train the model in the first place? I just want to ensure clarity on that.

>pangram currently has a false positive rate of about 1 in 10000

Says Panagram.

The problem with just looking at old text is language is a living thing. Say for example I make up the world 'oklambroahaha' right today. Both humans and AI pick up that word and start using it. Now lets say the model says that anything that uses oklambroahaha is 100% AI, you can't just point and say, "well my detection AI is correct on things 20 years old, so it's right skibbidy toilet 6/7".

There is a ton of evidence that use of AI changes the way we speak and write, so it will just turn these AI detectors into bullshit generating classifiers.


You can get an arbitrarily low false positive rate by sacrificing against false negatives. It's trivial to make it zero, just classify everything as human-generated. Meanwhile a false negative rate of even 1% is a pretty big problem since someone can easily use LLMs to generate 100x the volume of text and then use whichever ones make it through the classifier.

And that's before anyone even tries to get the LLM to generate a different style of text. Or for that matter creates a "style model" that rephrases text.


You don't really need a style model - current models are very good at doing "style transfer" of a model text onto whatever it has written if you just have it do it chunk by chunk. It takes more to prevent it from being detectable by good detectors, but it does remove a lot of the worst tells.


The point being that you wouldn't need the developers of the most popular models to themselves be trying to fool classifiers because their output could be run through an independent special purpose one designed to remove the tells the classifier is looking for, and the special purpose one wouldn't need to be made by anyone with the resources to create a good general-purpose model since it only has to do that one thing.


My point is that you don't need a special purpose one to achieve this.


Pangram won't know how much AI written text they fail to detect, though, and detectors is a great tool to adjust methods of generating less AI-sounding text.


Claude did not find a proof, though. It found an algorithm which Knuth then proved was correct.


The insight is the point of research. Proof isn't the desired product of research, it's simply an apparatus that exists for the purpose of verifying and demonstrating correctness of insight.


Yes, and his point is that finding that algorithm was, to Knuth, the interesting part. Getting from that to a proof was the boring bit.


Yeah, and I'm not sure what the other guy's argument is. It's Knuth, the primary researcher, who is giving the praise here. I don't see a possible motivation he would have to falsely give accolades to a AI for a problem he presented, then cleaned up to solve.


That’s fair. Clearly Knuth himself thought it was impressive, that’s a strong signal.


AFAICT, Claude was not asked to prove its algorithm works for all odd n, but was instead told to move on to even n.


The companies aren’t changing anything. LLM outputs are just more random than people realize. Run the same prompt 10 times if you really want to know how well they can answer.


Counterpoint: What progress has generative linguistics made in the same amount of time that deep learning has been around? It sure doesn't seem to be working well.

Also, the racecar example is because of tokenization in LLMs - they don't actually see the raw letters of the text they read. It would be like me asking you to read this sentence in your head and then tell me which syllable would have the lowest pitch when spoken aloud. Maybe you could do it, but it would take effort because it doesn't align with the way you're interpreting the input.


>What progress has generative linguistics made in the same amount of time that deep learning has been around? It sure doesn't seem to be working well.

Working well for what? Generative linguistics has certainly made progress in the past couple of decades, but it's not trying to solve engineering problems. If you think that generative linguistics and deep learning models are somehow competitors, you've probably misunderstood the former.


Also being able to count number of letters of a word is not required for language capability in the Chomskian sense at least.


I think this is greatly complicated by the fact that the human brain has been "pre-trained" (in the deep learning sense) by hundreds of millions of years of evolution.

A pre-trained LLM also can also learn new concepts from extremely few examples. Humans may still be much smarter but I think there's a lot of reason to believe that the mechanics are similar.


The poverty of the stimulus (POS) argument is that "evolutionary pre-training" in the form (recursive) grammar is fundamentally required and can not be inferred from the stimulus.

The argument is based on multiple questionable assumptions of Chomskian linguistics:

- Humans actually learn grammar in the Chomskian way - Syntax is separate from semantics, so only language (utterances) can be learned from uttrances, and not e.g. what is seen in the environment - At least in the Gold's formalization of the argument language is learned only from "positive examples", so e.g. the learner can't observe that some does not understand some utterance

One could argue for a (very) weak form of POS that there has to be some kind of "inductive bias" in the learning system, but this applies to all learning as shown by Kant. The inductive bias can be very generic.


>At least in the Gold's formalization

It seems to be a persistent myth (possibly revived more recently due to Norvig?) that Chomsky's POS argument has some interesting connection to Gold's theorem. The two things have only a very loose logical connection (Gold's theorem is in no sense a formalization of any claim of Chomsky's), and Chomsky himself never based any of his arguments for innateness on Gold's theorem. Here is a secondary source making the same point (search for 'Gold'): https://stevenpinker.com/files/pinker/files/jcl_macwhinney_c...

The assumption that syntax is 'separate from semantics' also does not figure in any of Chomsky's POS arguments. Chomsky argued that syntax was separate from semantics only in the fairly uncontroversial sense that there are properly syntactic primitives (e.g. 'noun', 'chain', 'c-command') that do not reduce entirely to semantic or phonological notions. But even if that were untrue, it would not undermine POS arguments, which for the most part can be run without any specific assumptions about the syntax/semantics boundary. Indeed, semantic and conceptual knowledge provides an equally fertile source of POS problems.


Yeah, I don't necessarily buy the whole Chomskian program. I'm willing to be persuaded that the reason kids learn to speak despite their individual poverty of stimulus is that there was sufficient empirically experience stimulus over evolutionary time. The Chomskian grammar stuff seems way too Platonic to be a description of human neuroanatomy. But be that as it may, it's clear the stimulus it takes to train an LLM is orders of magnitude greater than the stimulus necessary to train an individual child, so children must have a different process for language acquisition.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: