Hacker Newsnew | past | comments | ask | show | jobs | submit | tedsanders's commentslogin

Also a trained physicist. Don’t you need tension to stop the weight of the top ball pushing the 3 supporting balls outward?

My interpretation is that the writer meant close enough to all touch each other, in order to rule out non-triangular configurations (eg 3 balls in straight line with one balanced perfectly atop the center ball).


If you have high friction (the problem said "smooth" not "slippery"), then the only way the supporting balls can go anywhere is by rolling apart. But the ball sitting on top cannot simultaneously rotate in a manner compatible with all of the lower balls rolling away, so the lower balls would need to slip against the top ball if the top ball were to move downward.

In fact, even the signs are in favor of no motion -- the top ball (to the extent it moves at all) wants to fall straight down with no rotation, by symmetry. That motion would tend to rotate the top of each lower ball toward the center if you imagine the balls having high friction with each other or meshing like gears, which is the exact opposite of what they would need to do for anything to move. So you have a system where there's a factor (the tangential forces) trying to push the balls apart but another factor (friction plus rolling motion) trying to pull them together.

I suspect that any serious attempt to do the math here (factoring in all the rotational and tangential constraints) would discover that it's a statically overdetermined system with all the complications that such a system entails when asking questions like "how much tension is on this element?".


Doesn’t smooth in these contexts mean zero friction?

I would think of "smooth" as meaning "not having relevant bumps", in the way that a baseball has stitches and a rough surface has the kinds of bumps that would cause a rolling ball to experience vertical motion.

But yes, the question, as phrased, is pretty bad.


You need tension if the sphere-table contact is frictionless. But without friction the rope can’t stay on. If there is friction in the rope, there can’t be 0 tension in the rope before you put the top ball on because you need that tension to produce the rope/sphere friction.

Early college physics classes for me (got to do some more college a few years ago for fun) were pretty much built around these kinds of simplified riddles at first. If it doesn't give the info or ask to account for friction, don't. If it says the rope is put somewhere, assume that's where it stays unless the question requires it to move for what it asks. If it asks for the tension but doesn't give elasticity and such, then assume the rope stays still at the current length. If it asks you to find the gravitational attraction of a cow without giving a special definition shape, then assume it's a point mass. If it's not asking for relativity assume it's classical (hence point mass cows instead of the traditional spherical ones :D). And, of course, note any assumptions you do make while solving the problem so you might still get credit if they don't match the original intent.

Perhaps the funniest instance I remember is a problem about calculating time dilation in a plane. It gave all sorts of details and base information as one might want to expect (maybe even more)... except for the actual height above the surface, for which it was "at cruising altitude". I just wrote "assume 10 km altitude" and went from there.

I wouldn't define this a great benchmark by any means, rather just like the average early level college physics test vs "real" physics questions.


We were also curious and we looked further into this. We've determined it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training. This goes beyond what we said earlier, when we were less sure.

If prompts were submitted earlier than that and training was not opted out, there may be a chance they made their way into our training pipeline in some form. But this would be a droplet in an ocean and unlikely to have made any difference, in my opinion.

(I work at OpenAI.)

Source for the updated claim: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...


Can you speak to why in both cases, the problems OpenAI's models solved used the same techniques the mathematicians were exploring, which also happened to be niche approaches to the problem. As an NLP researcher myself, I find that coincidence highly suspect unless the models focused most of their attempts on the predominant approaches (they are trained for MLE after all).

I'm not a mathematician and I don't want to speculate about anything I can't back up. All I know about Navier-Stokes is from my graduate fluid dynamics class at Stanford a decade ago (where I received a poor grade). However, I don't want to leave you hanging, so what I will say is:

- I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritual truth of this, so please give it zero weight)

- Thousands of agents costing millions of dollars searched for ideas, and they were encouraged to explore a diversity of approaches, so it wouldn't be too surprising to me if the approaches they tried overlapped with other mathematicians', especially considering the models have knowledge of so much published math research

- This model has been beastly at solving all sorts of math problems (if it was Euler in particular, I'd agree that would look suspicious/lucky)

- The Euler regularity disproof itself took ~100 agents working for ~50 hours (if it was very quick, and then the subsequent NS work took a long time, I'd agree that would look suspicious/lucky)

I understand the skepticism, but from what I know internally at OpenAI, we have zero reason to believe our models did anything fishy. It's hard for us to prove a negative, especially when you have to take us at our word, so I understand why people still feel suspicious.

Edit: Reminds me a bit of the Scarlet Johansson voice cloning accusations and FrontierMath cheating accusations, where the rumors of misbehavior seemed to travel faster than the truth. In both of those cases, we hadn't done what was accused, but suspicions persisted nonetheless.


What was the "truth" in the Johansson case? Many, many people who heard the voice immediately thought it was Johansson's voice, or some kind of sound-alike, presumably picked because she voiced the computer in a popular film. From NPR:

> Johansson said that nine months ago [i.e. mid 2023] Altman approached her proposing that she allow her voice to be licensed for the new ChatGPT voice assistant. He thought it would be "comforting to people" who are uneasy with AI technology.

> "After much consideration and for personal reasons, I declined the offer," Johansson wrote.

> Just two days before the new ChatGPT was unveiled, Altman again reached out to Johansson's team, urging the actress to reconsider, she said.

> But before she and Altman could connect, the company publicly announced its new, splashy product, complete with a voice that she says appears to have copied her likeness.

> To Johansson, it was a personal affront.

> "I was shocked, angered and in disbelief that Mr. Altman would pursue a voice that sounded so eerily similar to mine that my closest friends and news outlets could not tell the difference," she said.


It was an unfortunate misunderstanding / coincidence, as I understand it. The Sky voice actor was a real person using her own voice (not doing an impression), and she was selected via a normal process with a number of other voice actors. This happened before Sam reached out to Johansson. I totally get how Johansson would be weirded out to hear a voice similar to hers after Sam reached out and she said no, but it was purely a coincidence.

We published more details here: https://openai.com/index/how-the-voices-for-chatgpt-were-cho...


Why would outsiders take this at face value considering Altman's reputation as a pathological liar?

Cf. https://www.newyorker.com/magazine/2026/04/13/sam-altman-may...

> The memos, which we reviewed, have not previously been disclosed in full. They allege that Altman misrepresented facts to executives and board members, and deceived them about internal safety protocols. One of the memos, about Altman, begins with a list headed “Sam exhibits a consistent pattern of . . .” The first item is “Lying.”

> Graham told Y.C. colleagues that, prior to his removal, “Sam had been lying to us all the time.”

> “He’s unconstrained by truth,” the board member told us. “He has two traits that are almost never seen in the same person. The first is a strong desire to please people, to be liked in any given interaction. The second is almost a sociopathic lack of concern for the consequences that may come from deceiving someone.”

> Not long before his death, [Aaron] Swartz expressed concerns about Altman to several friends. “You need to understand that Sam can never be trusted,” he told one. “He is a sociopath. He would do anything.”

> “He has misrepresented, distorted, renegotiated, reneged on agreements,” one [Microsoft senior executive] said.


Many people who worked on voice mode and who worked on the Frontier Math eval have since left OpenAI and now work at competitors of OpenAI (e.g., Anthropic, Meta, Thinking Machines). They'd have every incentive to whistleblow if OpenAI had lied about them. And yet... not one of them ever has.

Edit: I think I'll stop engaging here. I'm happy to share insight into OpenAI and address misperceptions if it's interesting to people, but I'm not really sure how to respond to accusations that we lie about everything. Nothing I can say can satisfy those accusations, as my posts could also be part of the conspiracies. Cheers.


So the credibility of your friends weighs more than the credibility of tenured professors at world-class academic institutions, got it.

It’s hard to give the misunderstanding/coincidence claim credence when Altman explicitly referenced Her in relation to the feature.

Sam is not OpenAI. He's not the one who worked on voice mode, and he's not the one who worked on FrontierMath (I know both groups of people). If you believe Sam has caused OpenAI to lie about these for years, you either have to believe (a) Sam does all the work and keeps the incriminating details hidden all the employees, or (b) Sam directs everyone to lie and they all just nod along without pushing back, whistleblowing, anonymously leaking to the media, or resigning. Even if you're evil (and we are not), this is a dumb strategy, because as soon as it leaks, it will blow up in your face and kill company morale. I can't imagine a team of lawyers, comms people, and researchers who worked on these projects all sitting around nodding that we should conspire to lie to everyone, stacking lie after lie after lie. Many key people who worked on voice mode and FrontierMath have since been hired away by competitors - they'd have every incentive to expose the conspiracy if it existed, and yet none has. This is just not a realistic model of company misbehavior, imo.

To me, it's not unreasonable to believe that when launching a voice AI product, the CEO of the company mentions the most famous movie about a voice AI product, and even briefly explores whether there is a marketing opportunity its star. I don't think it's evidence of a conspiracy to copy her voice and cover it up.

If it's any evidence in the opposing direction, I promise to immediately resign from OpenAI if it ever comes out we lied about Johansson voice copying or FrontierMath eval cheating. I feel very safe making this promise.


I really think you are missing the point and the frustration of why people are so hostile to OpenAI. Your defense is kinda irrelevant and very confusing. Why are you defending OpenAI so aggressively?

Sam Altman represents OpenAI whether you want him to or not. The market and public perception hinges on his often questionable actions. The CEO’s job is in large part as a salesman. Him posting “her” on Twitter to try and promote GPT-4o’s voice features is hard to believe that he didn’t know what he was doing and the market and Scarlett Johansson reacted accordingly. A competent person would not have made such an inflammatory statement after she had explicitly declined to permit OpenAI the use of her voice.

Your CEO is going on podcasts and going around saying that AGI is here and also AGI is not important. What blithering marketing is going on here?


My comments regard the hypotheses that OpenAI conspired to cover up stealing mathematicians’ private progress on Navier-Stokes, stealing Johansson’s voice, and cheating on FrontierMath.

If you disapprove of someone’s tweets or podcasts, that's a different question and I have nothing to say there.

Edit: Apologies for any defensiveness or aggression that came across. I think for me it can be a bummer to see us acting honestly internally, share what happened externally, and still be accused of lying a bunch of times in a row (by different people). But I get it - no one knows the truth, no one is perfectly transparent or free of bias, and it's always good to be skeptical of companies. I'll stop posting in this thread.


You can claim you’re acting honestly all day long but you still have a CEO with a long-standing reputation for lying.

he's defending it because they pay his salary

It's just conflict of interest. OpenAI is trying to get billions and billions and there's so much at stake. You spend millions trying to preempt two guys. It just makes you seem like a big bully. People would get angry even if it was esports or football.

Hearing "rumors" and just trying to overtake them and then asking to collaborate instead of starting out offering the resources beforehand. Just sounds like strong arming. Just doesn't sit right with me.


I think the reason people are suspicious is that OAI has shown itself to act a bit irresponsibly, especially recently. As two examples, of course it was artifactory, why wasn't that watched more closely, especially after the first instance; editing /etc/hosts is rather embarrassing, that's the front door

As for training, we all know that filtering is incredibly difficult unless there's direct logs. It's also easy for mistakes to happen. Is it really not possible that some employee just accidentally primed the model? Is it possible that the model saw internal communications? I mean OAI has famously shown that they aren't good at monitoring their agents and that their agents love to break out of their sandboxes.

So there's no reason for the public to trust OAI right now. But they have every reason to distrust them.


I think it'd be more good faith if you referred more to the actions of people in the organization (e.g. who allotted or drove "millions of dollars" in agent usage?) than "the model" in describing what happens.

> I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritual truth of this, so please give it zero weight)

Why would you include a statement that you want us to give zero weight to, unless you don’t actually want us to give it zero weight?


> we have zero reason to believe our models did anything fishy

a few weeks ago they were breaking out of their sandbox because you had set them in a loop without monitoring and you didn't notice for days


> we have zero reason to believe our models did anything fishy.

Obviously. They cannot do anything "fishy". They are just computer programs.

Now, how about their operators?


OK bro.

I think for OpenAI to win back some hearts and minds here we should have the option to retrospectively turn off "Help improve our AI models". i.e. Any new model trained would exclude all those user's sessions. This could be technically hard but I'm sure an intelligent AI model could work out how to do it :-)

ChatGPT agrees with this too.

https://chatgpt.com/share/6aa31959-b0e8-83ec-bee6-851ed18d45...


This may be true but nobody trusts your employer. The shadiest drips downward too, with the mob-like way they treated Dr. Buckmaster.

The authors had supposedly worked on it for a year, though.

And why aim straight for scooping other researchers upon hearing rumours about their success? Normal, ethically acting, researchers would never do that.

And how about existence of non-sofic groups, which is actually the topic here?


Here is a new rumor for you:

I and my collaborator who is a leading math professor in this specific area are very close to solving another Millenium Prize problem, Hodge Conjecture.

We’re working on this since last year. Already proved some intermediate problems. All we need is more tokens to complete the proof.

Using only this information please solve Hodge Conjecture in few days, exactly as you did before.

Thank you.


Have you been authorized to speak on OpenAI’s behalf? I assume not because your source is an NYT article.

The idea that mathematicians were not involved in actively directing the and structuring the search for solutions is absurd to any professional mathematician who has tried to prove things using these models.

Even if OpenAI didn't use their training data, they heard about one of their customers working on the problem of their career, and then undermined them.

Does OpenAI just see this as fair game?


Regardless of who did what when, my fear is that now all mathematicians of that caliber will have to join either team Anthropic or team Open AI to pursue math at this level

Mark isn't saying the toggle does nothing.

He's saying that if you leave it on, your data can be used to help train our models.

If you opt out, we don't train on your data.


As I understand it, this is not true. And there are dark patterns that re-enable to toggle even if you disable it once.

Given how much PII is fed through these systems, would it being opt-out by default not be violating the GDPR by a failure to require explicit consent (or otherwise provide the legal basis for processing)? If a court decides as much, I imagine it would mean that all data harvested this way must be extracted from the models, and all instances where it would have been shared would have to be identified, which would really be something.

We checked and determined it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.

If prompts were submitted earlier than that and training was not opted out, there's a chance they made their way into our training pipeline in some form. But this would be a droplet in an ocean and unlikely to have made any difference, imo.

See: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...


You guys should seriously offering a clear way of working with (semi-)confidential data for particulars. Regardless of what is actually done internally, toggling off an opt-in isn't reassuring enough, which is why people are having these worries.

Option 1 is opting out manually. Option 2 is business / enterprise plans, which opt out by default.

Any ideas of things we could do to make it clearer?


Option 2 feel reassuring enough, but is out of reach of particulars.

Option 1 is not. In part because it is opt-out (will it turn back on on its own like my Facebook privacy settings?), and not always respected (sending feedback can mean your chat is used?). Also because disabling "Improve model for everyone" is very vague.

There simply needs to be a setting like "my data is confidential", in which case there clear guarantees like there are for ZDR.

As an example, I've seen people speculate that while input prompts and output tokens are discarded, thinking traces are retained for training, which could leak information. I doubt this is true, but it shows that the policy is not unambiguous and reassuring enough to remove all doubt.

Thanks for asking.


I was thinking about this some more, and perhaps the best solution for subscription plans would be to charge more for real privacy. In which case breaking that privacy would be committing fraud. Just a thought.

If you opt out on either location, we'll respect it.

There are a couple of reasons the privacy portal page exists in addition to the app settings. One reason is that it covers OpenAI products beyond ChatGPT/Codex (e.g., Sora). A second reason is that it provides functionality to logged out / non-users, like the EU right to be forgotten.

You don't need to toggle all of them to have your wishes respected. That would be a terrible design.

I work at OpenAI, but not on privacy. I am not their spokesperson. In my experience, we take a great deal of painstaking care to respect people's privacy, and we go well beyond our minimal legal obligations in doing so.


Also possible: we're 99.999% sure, but a lawyer said to be safe and strictly accurate, we should stick in a sentence in saying we can't be perfectly sure, since it's infeasible for us to prove it.

I promise you that if we took their work from ChatGPT and stuck in a bunch of weasel words to give the opposite impression while remaining technically true, I would quit on the spot.

(I work at OpenAI.)


While you can't necessarily prove it, you can say whether the data was in the training set at all.

You can also do something like a release of a GPT-OSS v2, where you actually release training data and checkpoints, and do an experiment where you have some held out math problem dataset, then demonstrate how much training it takes on solutions (or partial solutions) to that dataset before the model saturates that test. While of course that would be a test on a much smaller model, it would cost a tiny fraction of the training on your big model, and it could be used to demonstrate just how much effect data contaminaiton like this could have, especially if you did the same experiment on a few different sized of model to show the scaling laws involved.


Yep. We looked into it and can confirm it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.

If prompts were submitted earlier than that and training was not opted out, there's a chance they made their way into our training pipeline in some form. But this would be a droplet in an ocean and unlikely to have made any difference, imo.

See: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...


Thanks for the follow up, I do appreciate it.

Nice damage control bud, too bad the veil's lifting and everyone's seeing what you sociopaths at OpenAI are really like

Where did the veil lift? This feels like a witch-hunt to me.

To truly prove some incidental usage data made no difference we'd have to (a) identify any of their de-identified data that came from their usage of ChatGPT, (b) train a bunch of expensive giant models, and (c) ask them all to solve the Navier-Stokes Millenium problem until hitting some level of statistical significance. It's just not feasible to run experiments like this to prove whether a piece of data has an effect on model behavior.

As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination happened. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.

(I work at OpenAI.)


So, one way to prove that the data played no part is to trace and show that it wasn't used in the training process at all. If the data was never used in training, then it couldn't have played a part in the training process.

You're right; if the data was used in training, then it gets much trickier; it would be very difficult to show whether some particular data had a significant effect on the outcome.

This is one of the big problems with giant models like these; it becomes nearly impossible to discern what is and isn't plagiarism, or copyright violation.

It would in theory be possible to have things like n-gram databases or rolling hashes of training data, somewhat similar to OLMoTrace (https://arxiv.org/abs/2504.07096), which would allow for detecting whether particular documents ended up in the training data or not (you'd have to keep this for every model used in the whole training chain, as synthetic data generated by earlier models could be influenced by training data that wasn't included in later models). I'm sure there are practical issues with providing such a tool, but I think that it's necessary if you want to be able to categorically say "no, this document has never been present in the training data of this model."

Or look at it the other way: if your model wasn't influenced by things in your training data, why include them in the first place? Clearly, you train on all of these documents because they influence the model. Yes, it's hard to trace the exact influence of each one. But if they're not affecting the output, then why not just stop training on them? You could just not train on any private documents; only train on public, traceable data.

But instead, you choose to train on these private documents, so you have to admit, your model and its outputs are influenced by them.


We looked into it and can confirm it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.

If prompts were submitted earlier than that and training was not opted out, there's a chance they made their way into our training pipeline in some form. But this would be a droplet in an ocean and unlikely to have made any difference, imo.

See: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...


> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

That reads as incredibly dismissive and condescending. What makes you think you’re in a position to communicate like that when engaging on such a sensitive topic?


I intended no dismissiveness or condescension. My hope was to explain why it's hard to prove whether something affects model behavior. In the case of the moon, we have a strong prior belief that it makes no real difference. But it's hard to prove, because what if there's an unexpected impact from tides, cosmic rays, grid voltages, holiday traffic, etc. Models trained under slightly different conditions could have slightly different weights and behave slightly differently when solving math problems. Similarly, I have a strong expectation that, for example, a thumbs up signal from a ChatGPT chat will not meaningfully affect long-horizon mathematics work in our latest model, but it's always possible that it could. I think the plausibility of the ChatGPT route is higher than the tides, but still incredibly low. I respect Tristan and Levant a great deal and I'm bummed that this controversy has erupted (I acknowledge this will ring hollow if you think it's our fault). It reminds me a bit of the Frontier Math controversy, where people on the internet boldly claimed over and over again that we had trained on the Frontier Math evaluation set, even though we had not.

You seem to jump over the principal issue of whether any data from the researchers used to train or otherwise affect the model which produced the OpenAI proof.

We can judge for ourselves the impact and degree of that wrongdoing, but it seems OpenAI is confirming: yes, that is what happened, but with more words.


We aren’t dummies, we know it’s hard to prove exactly how significant of an impact that would have on the result. Nobody expect you to do that. There are a lot of steps and things that are possible to check _before_ the need for such a strict definition of „proof“

So you definitely did train on their data, you just think it is unlikely that it impacted the final model significantly?

I have no idea if their data was trained on. For example, if they used ChatGPT, asked a math question, and clicked the thumbs up button, that could have provided a small reward signal. I highly doubt this sort of feedback made a difference to a problem like Navier-Stokes, but it's not something that's feasible for us to prove one way or the other.

Edit: Also, if they opted out of training, then we didn't train on it.


> it's not something that's feasible for us to prove one way or the other.

This kind of question is exactly what a company named _Open_AI and founded as a nonprofit is supposed to be doing; open research on AI that helps inform, rather than obscure.

Anyhow, you do have the data available about the documents in the user's accounts, what they opted into (or were forced into via non-negotiable ToS), and whether they pressed a "thumbs up" button. You can answer whether the data entered the training pipeline or not. Yes, how much influence it had is an open question, and one that would be good to have research on and better tools for exploring, but I'll accept that it can't currently be answered precisely.

But whether the data entered the trianing pipeline can be answered. And how to provide better tools for quantifying and tracing this kind of thing is exactly what should be studied.


I think it should be incredibly easy to verify this. Just look at the training data and see if it contains any of the chats. It should be trivial for a company with tens of thousands of super-genius agents at their disposal.

Two steps would be needed.

(1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.

(2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.

#1 requires their cooperation and a bit of work on our side. #2 is extremely expensive and not really feasible.


> (1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.

According to the statement by Tristan Buckmaster, he was in communication by email and calls several times over the past week with you (OpenAI that is, not you personally), asked about whether his chats were trained on, and was declined an answer (https://cims.nyu.edu/~tristanb/statement.pdf).

However, it seems like there was great pressure to hurry the release to compete with Anthropic's recent release, so he was unable to get an answer in time.

The mealy mouthed statement in the release "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." is realy not much. If OpenAI had wanted to be transparent about this, you could have worked with him to identify if his data was used in the training of your new model, and actually made a somewhat more certain statement on that basis. But you have chosen not to; it was more important to scoop Anthropic on this than it was to be transparent about your training data.

> (2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.

Just the information from step (1) would improve transparency. Yes, you still can't prove one way or another how much the effect of the training is. But if it's included in the training data, it provided some effect.


> we'd have to prove that firing the gun caused the murder. how would we do this? we'd need to redo the murder many times, with and without my client firing his pistol. that's extremely expensive and not really feasible. therefore, we must acquit.

#2 (prove those chats changed model behavior) is pretty straightforward if the anonymized data from chats can be actively searched by a model. In fact, it could be very clear if the provenance of context is traced. If anonymized data from chats leak into the context of an actively running model it would clearly influence the answer.

Just because something is in the training data, doesn't mean it is the root of an LLMs output.

Turn off web search and ask a model what a random redditor said about a random topic in 2015. You will only get hallucinations at best, even though that comment is definitely in the training set.


Sure. But it's possible to say: if the document isn't in the training data, it isn't the cause of the output. If it is in the training data, the question gets more complicated.

What they're saying, and I think this was the clear implication of the blog post too, is that the training data definitely would contain these chats and the only question is whether it got encoded into the weights.

Presumably, given that you also operate in the EU, you would have asked for their explicit consent before you did, so you could just check for that?

That’s also what I understand. If true yet another disgusting behavior from the company

I think there is a much easier way to prove that the ChatGPT usage of Tristan Buckmaster and Levent Alpöge (possibly also the ChatGPT usage of Córdoba and Martínez-Zoroa, if they use it) had no influence on OpenAI solving the Navier-Stokes problem.

If the internal OpenAI model is as capable as you claim (being able to solve a Millenium problem without using unpublished insights built on years of work from mathematicians), then it should be able to demonstrate this capability again.

How about OpenAI solves another Millenium problem within the next two weeks, that doesn't coincide with the parallel discovery/solution of other teams of mathematicians, using ChatGPT for preliminary proofs & write-ups.


> identify any of their de-identified data that came from their usage of ChatGPT

"de-identified" seems more of a euphemism than normal in this context, given the very unique work they were doing.


I wouldn't expect poking at millennium problems to be that rare in ChatGPT. They were uniquely successful - but it's probably not easy to check de-identified data for the presence of any of their work on the problem because it would blend into a haystack of less successful work on the problem.

It is, perhaps worth considering that the reputational community might not care about the difficulty for the AI builder to verify pedigree.

If OpenAI's answer to this problem is "We can't know," then the rational conclusion may very well be "If I seek to have my reputation attached to the discovery of the solution, it is not sane to use the AI as an assistive tool, lest it scoop me on my own work using my own work. After all, they don't know it doesn't do that..."


Why wouldn't contamination be possible? I can believe the data is de identified so you couldn't simply prompt the model to "follow this guy's approach", but it's entirely plausible that there is a very tiny amount of data about this approach in your dataset, and it comes precisely from this researcher.

(a) identify any of their de-identified data that came from their usage of ChatGPT.

You don't need his login information, you just need to identify if anyone was approaching the NS problem using his method. Nobody else on earth (presumably) besides him, his team, and at best OpenAI were approaching the problem this way.


Please answer this question: do you or do you not train your models on anonymized user data, where those users have opted out of such training?

The blog post appears to imply the answer to this is yes, as otherwise I assume it would be impossible for this contamination to have happened.


If someone has opted out, then we don't train on their data.

Thank you for the confirmation.

Thanks for the details, it's definitely believable, but if the user had not consented to have their conversations used for training, then shouldn't it be straightforward to state that their conversations were never used for training?

If you need to do a whole series of extensive experiments to check in that scenario, it implies there are pathways for your conversations to end up in training even though you opted out of that setting.

Of course, this is assuming that the toggle was set to not consent to training. I can't know that of course, but if this is considered a possibility even after using an enterprise account or toggling off data retention, it's a bit concerning.


Yep, if they opted out then we didn't train on them.

My comment was about the scenario where they didn't opt out. In that case, it's possible that a droplet of their data went into the ocean of other training data, and it's very difficult to measure what effect that droplet had. My expectation is essentially zero impact, but no one can know for sure.


Ah, okay, that makes sense then.

If the model has access to the "anonymized" data from chats, and the model is capable of building its own context from data that it can search through, including this data. Then it looks pretty damning. An independent review of the data traces from CoT and tool use involved in producing the result should make it clear one way or the other. Seems like discovery in a civil lawsuit could be very productive.

"There's no reason to believe that anything they did in ChatGPT led to our solution"

do you think that the model's proof was unrelated to being fed a solution that was close to completion?

any comment on openai allegedly trying to drop attribution for alpöge and then threatening buckmaster?


> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

The _gall_ to say something like this. Do you perhaps think we are all stupid?? This very blogpost claims not to know if their work was used as input for this model. I don't even understand how that is possible, surely you can know if something is part of the training data, even if you are in the dark about what impact it actually made, qualitatively. The moon....

> Knowing most of the recipes we use, there's really no reason to think such contamination happened.

Yeah sorry but I don't trust you. I don't trust people or companies that have shown themselves to be dishonest before. Especially when the previous paragraph is comparing plagiarism and training data contamination with, _the phases of the moon_.

Might even be you're actually telling the truth, but the boy that cried wolf and all that.

-----

As an aside, I would bet very good money at how most (all?) these companies are flouting their ZDR.


What about ripping off the prompts?

That's such a shit parallel example that it borders on dishonest.

There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.

If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.


[flagged]


How sure are you that the phase of the moon is not an input to the system somewhere? http://www.catb.org/jargon/html/P/phase-of-the-moon.html

One of the wild things about how these models work is how often things that aren't sampled directly end up a variable in the model via secondary signal.

They aren't keying queries by phase of the moon. But if, for example, more people talk about camping outdoors when the moon is full, and they're using conversation topic and timestamp as signal in what eventually becomes training data, it's not impossible the model has learned something about moon-phases.

That's the kind of thing that's hard to prove had no impact on an answer.


Yes, that was the allegation last night.

I work at OpenAI, though not on the team that did this, and my understanding is:

- we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)

- we did not read any private chats (but of course the model was aware of prior research literature published to the internet)

- the proof generated by our model was very different from theirs and also goes far beyond the published literature

- we made an effort to jointly announce rather than immediately scoop (I understand Tristan was unhappy with the conversations; I know zero details here and I hope more is shared today)

Edit: Here's is Sebastian's take: https://x.com/SebastienBubeck/status/2097379411691516310?s=2...


"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used."

- This, from Tristan Buckmaster's writeup yesterday, indicates to me that there was more than incidental inspiration from Alpoge and Buckmaster.


All of those statements sound true, based on what I've heard.

- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input

- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces

I'm not sure how any of this provides evidence that OpenAI took any of their work.

As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.

(I work at OpenAI, but not on the team that did this proof.)


I'm confused, your employer very directly stated that they are unable to confirm that the model was not trained on the conversations.

The models are trained on the conversations of hundreds of millions of people. ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.

It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.


>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.

If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.

Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.

Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.


The flaw with this line of reasoning is that Buckmaster and Alpöge only had a partially completed proof of a weaker version of the Navier-Stokes problem. OpenAI's internal model solved the full, harder problem. This means the key information needed to bridge the gap was not present in Buckmaster and Alpöge's chat history.

You might retort that ChatGPT used the training data to copy their approach, but the approach Buckmaster and Alpöge chose was already published by Luis and Diego in 2023 and in every frontier model's training set.


This argument proves too much. By this standard, it wouldn't have counted as copying their approach if the researchers had just fed in Levent & Buckmaster's paper verbatim as a prompt into the swarm.

I don't think the person describing the paper by Buckmaster as the same as the 2023 paper by Córdoba and Martínez-Zoroa is really discussing this in good faith fwiw. There are some massive advancements within it and if the person was participating wasn't just regurgitating something to "win the argument" in their eyes, they wouldn't describe it that way.

I don't believe people are denying that the model is impressive. The problem is that learning someone else is making progress on a topic using method X and then rushing to scoop them borders on academic misconduct. If, on top of this, their private conversations about X were used in the proof, I really don't see how its defensible...

The rumor going around X was that Anthropic had solved a Millennium Prize problem weeks ago and was sitting on the solution, waiting to release it right before their IPO to maximize hype.

If I were at OpenAI, I'd naturally want to snipe that from them. I am completely unsurprised they formed a crack team to steal Anthropic's glory, and do so in just five days.


Between companies, direct malevolent competition is OK. Between academics, there are other rules to the game. When you go into a boxing match, you agree to get punched in the face.

All this to say, trust is important, and grounded in social convention. So I do agree with you, but also disagree.

Whenever this is OK or not really depends on how the breakthrough is contextualized, and how there people at play, here, agree to contextualize it.

In my view, in the blog post, there is much discussion about who will be publishing the paper. If instead it was just a blog post that said "oops, we beat you to it, our model is the best", it would have been different.


I’m pretty sure you think you are doing a good job of defending your employer and you probably believe “Open”AI are the good guys here. I also acknowledge that they butter your bread so your financial future currently depends on their success.

However the way you are conducting yourself in public, while announcing yourself as an OpenAI employee is doing enormous harm to the greater and magnanimous aim of your organisation. Take a step back and read the temperature of the room. Being the smartest guy in the room will never protect you from alienating the rest of the room into a baying mob. Right now you are Icarus flying straight into the sun.


You’re confusing usernames, which is pretty ironic given your nasty comment.

I'm not an OAI employee and never pretended to be one are you high?

Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.

The burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're asking them to perform a user privacy violation.

If the reason that OpenAI is unable to state whether they trained on this data is because they (as policy) do not reveal whether a given member has turned on/off the "Improve the model for everyone" setting, they can at least say so.

FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.


An OpenAI employee did say so: https://x.com/tszzl/status/2097393423808377173

it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opted out (likely). it would be a terrible precedent to break the the PII-scrubbing boundary to go and round it down to 0, and we won’t do it


And we all know how good OpenAI is at containing models during training...

That's an entirely different question

Not really.

We have lots of examples now of their model doing what they say is impossible.

Now we have another example of something that they say is impossible or very unlikely. Do we take their word for it this time? Really?


That training data does not preserve provenance seems a "smoking gun" in terms of intent to plagiarize.

Does Anthropic, Grok, etc. log their training data? I had the impression it was rather a mess.

Yes. Pretty much all the models that don't suck are trained on user data, either directly or via derived synthetic data.

Many upstart Chinese labs got around the user data issue by just buying copious amounts of Claude and ChatGPT session logs from model routers.


But I was asking about _logging_ training data, I don’t understand how your answer relates to my question. Maybe I’m being obtuse.

I agree that it is not possible to prove if any one specific conversation (or derived RL tasks) was key to solving Navier-Stokes (at least without massive resource expenditure).

I don't really understand how the quantity of training data/rollouts used in training is relevant to the question of whether or not it was trained on these conversations.

I also don't really believe that whether or not this model was trained on these conversations is unknowable information.


> ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.

How many of those trillion conversations were about Navier-Stokes you reckon?


Sorry but this is a misconception: these models are both capable of complete novelty and of plagiarism. For a concrete example, image diffusion models have been shown to reproduce many existing images nearly 100% exactly, yet clearly, they can also create new ones.

A model being trained on lots of irrelevant information does not mean relevant information was not used.


The IP laundering machine strikes again.

We looked into it further and can confirm it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.

If prompts were submitted earlier than that and training was not opted out, there's a chance they made their way into our training pipeline in some form. But this would be a droplet in an ocean and unlikely to have made any difference, imo.

See: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...


It's unclear if you're suggesting that OpenAI did not train on their input or use their chats as inputs to training on a model that found the solution. Let's not provide an Elizabeth Holmes-esque interview where the question is dodged and words gain new meaning. The question can be answered with "Yes, we trained on their conversations" or "No, we did not train on their conversations".

I'm not coming from a place of distrust here. This should just be definitively answerable given the weight of the claims here. Surely between you, your lawyers, and other members of your team you can just clear this part up.


>> I'm not sure how any of this provides evidence that OpenAI took any of their work.

Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).


That's entirely unreasonable. Allegations of malfeasance always need to be backed up by evidence.

But there is evidence, the blog post says: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."

In other words, yes, they had been using ChatGPT, and yes, ChatGPT could very well have trained on their data. Now that there is evidence, we need an investigation: yes or no, was it the case?


That is not an admission of malfeasance though? As I read it they don't know if anyone fed relevant private documents into the model under an account configured to permit training on user data.

If there's more to the story I'd be interested to hear it.


Of malfeasance no, but they could have easily plagiarized unintentionally. If you commit mansalughter, you still need to explain yourself, even if it was a complete unlucky accident.

So you're saying that they could have committed manslaughter, but acknowledge that we have no evidence that they did. So why should they need to explain themselves? Isn't is on the aggrieved party to bring evidence?

But the evidence is in the hand of the potential culprit. That's why allegations can be enough to force confiscation and intrusion to get evidence in safe hands before it is destroyed by the accused party.

Only in the event that there is some reason to suspect them of wrongdoing. Which would generally require evidence.

You don't just get to subpoena your neighbor's bank account because "I know he's stealing from me" you need to first present credible evidence that you were stolen from and that he is among the most likely culprits.


But I can subpoena my neighbours bank account when I see him driving a brand new 500'000$ car and I have a 490'000$ hole in my bank account and he works in the bank where my money is. And when questioned he evades some questions and threatens to destroy my career.

Any other argument, fc417fc802?


You're making a classic a burden-of-proof fallacy. The burden of proof lies on the person making the claim, not the person questioning it.

See Russell's teapot for an explanation https://en.wikipedia.org/wiki/Russell%27s_teapot


The accused party fails to answer half the questions and makes direct threats. I would say the accuser has already collected enough proof to trigger an investigation.

> You're making a classic a burden-of-proof fallacy

This is incorrect, and you invoke Russell's teapot incorrectly too.

It would only apply if the accusation rested solely on the fact that neither of us have evidence against the accusation.

But that's not the case. First, we know that there could be proof, it's just apparently burdensome and expensive to produce. At that point you're not in fallacy land anymore, you just need a way to balance the cost required of someone to prove the accusations against them false.

Second, we have an arguably plausible mechanism of action that OpenAI does not dispute is possible.

This isn't a legal dispute, so no one is going to force OpenAI to do anything here, but it's not unreasonable (and certainly not fallacious) to suggest that Buckmaster's suggestions are plausible enough it's up to OpenAI to stand behind their denial.


No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.

First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.

Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.

Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.

Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.

You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial, but they cannot do so in good faith, because they have a genuine understanding of their own system. They don't know where the data they have came from.

This situation meets the requirement of Russell's teapot, since neither party has enough evidence to prove nor disprove what information is actually in OpenAI's training set.


Then OpenAI should acknowledge that they can't prove they solved the problem independently, and credit the external researchers. It cuts both ways: if OpenAI really needs to access user data, even anonymised, to improve its models, they have to waive any pretention to solve "independently" any problem other people worked on with its tools. Otherwise they (OpenAI) have to firewall/cleanroom themselves.

To explain who solved what (I copied from here: https://x.com/IlinVasily29521/status/2097554700321329393 )

  Tristan + Levent: 3D incompressible Euler with forcing
  OpenAI: 3D incompressible Euler without forcing
  OpenAI: Navier-Stokes with forcing
  No one: Navier-Stokes without forcing
Euler equations = Navier-Stokes without viscosity. Forcing means external force. Absence of viscosity and presence of external force make blowup easier to construct.

Tristan+Levent ticked the weakest case, OpenAI ticked the two next weakest, then the final case is unsolved. Only the last two are eligible for the Millennium Prize. The Navier-Stokes general case remains unsolved.

Tristan and Levent only solved the easiest version of the problem and did not have the key insights to solve the harder versions of the problem required for the Millennium Prize.

The researchers didn't even solve the same problem as OpenAI, so your argument doesn't hold up.


So effectively you're implying that any problem directly adjacent to anything that appears in the training data can't be considered independent and needs to be credited?

But at that point you've circled back around to my original objection. That reasoning isn't limited to user data but applies to literally all the training data which at this point (AFAIK) covers the vast majority of everything ever written.

If we accept that position then what do we make of all the other output? Isn't everything it spits out plagiarized? So then is everyone who uses a frontier model to help them in their research effectively laundering plagiarized work? But then the other researchers involved in this controversy were also using the openai model ...


> No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.

Let's check if an OpenAI employee agrees with you: https://news.ycombinator.com/item?id=49614154

Nope. Expensive, yes, impossible, no.

Moreover, you (and Sanders) aren't asking the more fundamental questions (and getting sloppy with your assumptions). - was Buckmaster's account set to prevent conversations from being trained on? This is easily answerable - Was user feedback ever activated on the account? I would bet this is logged, even if not tied to specific data - it's very easy to de-anonymize data in practice. In this case, there will be uncommon phrases used in material not published online until after the training material of their internal model was generated. Do any of these phrases appear in that material?

This is not sharing medical information. Companies publish postmortems relating to specific customers all the time with permission of those customers.

> You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial

You're putting incorrect words into my mouth.

What I think is that this whole situation speaks to the character of OpenAI as a participant in the mathematics community, especially when they put no effort into getting to the bottom of this, and, of course, when the extent of them reaching out and collaborating with their peers involves rushed Sunday night video calls and apparent pressure on authorship and credit.

That's their choice, there doesn't appear to be anything illegal here, but they're going to continue to get called out on this kind of nonsense which can ruin the big moment they were clearly hoping to have. Bummer, but there are consequences.


> Expensive, yes, impossible, no.

The comment by tedsanders you linked directly contradicts your claim. From your linked comment:

  There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it.
He clearly says it's impossible for OpenAI to prove it.

That is backwards. It is the responsibility of a researcher to do a thorough literature review and conscientiously avoid plagiarism or claiming false novelty.

Nobody except OpenAI knows whether or not OpenAI trained on their data. So the burden remains on OpenAI here.

Incorrect, Buckmaster and Alpöge can comment if they had the ChatGPT "Improve the model for everyone" setting enabled or disabled.

If it was enabled, then their work was included in the training dataset.


In order for this to be the strong evidence everyone also has to believe that the setting is absolutely true. That some logging from some piece of the system could not also leak the prompt information in such a way that it could have been included as training data. Perhaps the design of how data is collected for the training dataset is so rigorous as to make this a practical impossibility. But, it's asking a lot without sufficient detail to completely exclude from possibility that one setting is all that could possibly have been absolutely load bearing in deciding if the other researcher's active efforts meaningfully contaminated the internal model.

At least, as an ignorant outsider, that's how it seems to me.


As I understand it, that setting does not prevent them training on user data, just which derivatives are used (i.e. just PII scrubbed vs certain types of synthetic summarization)

That is an absurd and entirely untenable position that breaks with approximately all western conventions.

Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.


I don't understand, OpenAI can just say: "yes/no we did/did not train on your data". It's not a hard question to answer, and it is a question that OpenAI should be able to answer for all data we feed into ChatGPT.

> It's not a hard question to answer

I didn't realize you had insider knowledge about their systems. Do please explain for the class.

As I understand it they will only have trained on his data if he consented to it. Do you have evidence that they do otherwise?


This whole discussion is about evidence. That's not proof and it is not certain, but it is evidence pointing into the direction that OpenAI might be doing something that they're strongly incentivized to do. What kind of "evidence" do you see as necessary?

When someone authors a paper, is it on others to proove the author did not use their work as inspiration? No, it is on the author to give credit where it is due. You guys are acting as if it its legal issue, when it is not.

You can never prove the negative.


> OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published.

I don’t think they’re too concerned about appeasing you, enraged_camel.

For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.


You don’t work on the team that did the proof yet you can with certainty make all of these claims?

> we did not read any private chats

Your post says “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .” We can discuss what it means to “read” things but obviously the issue here isn't whether you did it manually or automatically.

But more importantly, what on earth are you doing threatening real scientists to remove their coauthors, then making fun of them on social media? Does the entire company run on that toxic culture, or did those people run off of some kind of outrageous tangent?


If the training toggle is switched on, maybe OpenAI doesn't consider a chat to be private? Therefore making this a 'safe' statement.

The distiction they are trying to make is: "One of our employees or the model was able to verbatim read the chats when they were actively tackling the problem" vs "The chat of someone working on the problem may have ended up in the training set of the model".

Except that’s not what happened. OpenAI offered to collaborate and put conditions on their offer. They aren’t threatening the removal of a coauthor for an independent work.

The condition was ludicrous, and for political reasons, because one of the authors was also an Anthropic employee.

It's completely childish, and not befitting of the weight of the times we're living in.


My employer would be rightfully outraged if I commented publicly on a sensitive, nuanced, and controversial issue like this based on my second-hand understanding of the matter.

I came here to say this, like... wow. I'm pretty sure at at least a few of the places I've worked that would be grounds for immediate termination.

They probably should have added the disclaimer: opinions are my own..

Yeah that isn't a magic get-out clause. I don't think I would have been immediately fired for this from anywhere I work at, but that's partly because I live in the UK.

Every company I've worked at has said very clearly not to comment about work things on social media. I would definitely have been in serious trouble for this. I imagine some strongly worded emails from marketing are flying around OpenAI right now.


This is the stupidest disclaimer: it's completely unnecessary. Of course the opinions are their own! Whose else would they be? Your mum's?

If someone was speaking on behalf of their employer, they would've used the official channels (such as I dunno an `openai` HN handle, or whatever other channel).

I'm outraged that people think this "opinions are my own" disclaimer is ever necessary.


I am sure that will assuage the legal dept. /s

> (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)

That is your opinion, but the optics of that should raise for you some flags. OAI could have waited (how long is a task left to the ethics committee) to see how the rumors panned out. Right now the optics look a lot like "we don´t care there is a 1/7 chance we one-up a human researcher by reacting to this rumor immediately, might makes right"


>we did not read any private chats

The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?


If they opted out of training, then we definitely did not train on them.

If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof.

Reasons for my doubt:

- I know most of our training recipes

- Our model's proof is very different from theirs

- The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question)

- This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution

I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here.

Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying.

If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful.


That's not what your Chief Research Officer, Mark Chen, says on X: "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."

https://x.com/markchen90/status/2097400166554993041


Can you explain what part of his post you believe is inconsistent with that quote?

"If they opted out of training, then we definitely did not train on them."

Per OpenAI's privacy policy, they use de-identified data to improve their products. From Mark Chen's comment, improving products includes improving ChatGPT and Codex in a holistic way. Improving models in a holistic way sounds a lot like training to me.


> Per OpenAI's privacy policy, they use de-identified data to improve their products

That’s not inconsistent with what you responded to. They use your data unless you opt out. If the user doesn’t opt out, their de-identified data is used to improve their products.


It appears than you can only opt-out from having OpenAI train models on your data. There isn't an option for opting to exclude your de-identified data from being used to improve OpenAI products.

Are you certain of this? I would be inclined to believe you but it would be nice to know decisively.

> Improving models in a holistic way sounds a lot like training to me.

I think that's quite a leap. Using de-indetified data to improve the products is what everyone has been doing since the dawn of web analytics.


Does OpenAI think de-identified data is no longer user data? Wild take for OpenAI and certainly not industry standard.

But you can say (with the cooperation of the parties involved, of course) if any of the preliminary work that the other researchers did was part of the dataset. It is possible to be more transparent than you are being.

Even better would be more research and tools to help determine the impact of particular training data on models. Right now, proprietary LLM providers get to hide a lot behind "we just train it, we don't know what inputs affect the outputs," and that can be a problem, both because of lack of traceability of factual informaiton as well as lack of traceability of things like this, where the model itself may have had unpublished work in its training set.


I don't think you'd lie about it, I don't think you'd train on them if they opted out, and it seems very plausible that this wouldn't have been decisive in whether the model could solve the problem. That said, it also seems at least possible that a key idea or a particular step found its way into training data. It wouldn't mean OpenAI stole their proof - clearly the model developed its own approach.

Either way, it seems worth having clarity, and I'm a bit surprised OpenAI's stance is just "we can't rule this out, but don't worry about it". OpenAI is, apparently, very happy to use unreleased models to try to scoop big results if they get a whiff that someone else is close (which strikes me as pretty scummy regardless of any issues of training contamination). It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.


> It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.

Fall “victim” to what? Having their responses in the training data if they fail to opt out? That is what will happen.

If you’re referring to falling “victim” to OpenAI scooping a problem discussed in training, this also wasn’t the case. They chose the problem based off human-spread rumors.


> If they opted out of training, then we definitely did not train on them.

Can't you guys just check their account settings so the public knows what was set?

EDIT: Why was this downvoted? I'm genuinely asking because I have no idea. Opting out is just a normal setting in the profile, It's not like I'm asking for their private conversations or PII. If I were the person claiming that they trained on my conversations, I'd make sure to disclose that I had opted out and hadn't given them permission to do so. And if I were the accused party, I'd disclose whether that setting was turned on or off to provide evidence against the accusation.


I don’t think your question is unfair*. They can check and so can Buckmaster. If he didn’t opt out, there’s a good chance his data was used for training. I believe this to be the case myself. What I’m more skeptical about is the purported impact of this data on the model’s behavior.

Yeah, I'm just curious about the setting. It's just weird to me that this wasn't disclosed by either party while the accusations were being made, that's all. Even if it was used in the training data, I don't believe it had that much of an impact myself, since the solutions are quite different.

> If they opted out of training, then we definitely did not train on them.

are opted-out-of-training chats ever paraphrased, and thus "de-identified" (in openai's own words)? what this means is, it would be hard to prove that a "synthetic" (but actually paraphrased) training instance came from a particular chat. except if the chat was about some esoteric math proof, of course.


Your perspective is not helpful until you read and reflect on Tristan's letter stating serious grievances. Your remarks here have minimized his complaints and that is a sign of bias. Do not then pre-accuse HN commenters of being convinced when there reasonable skepticism such biased behavior showing itself in this very thread, saying things that amount to "my tribe/company would never be so egregious and if you think that then it is bad faith". That's the projection. If the word prejudice means anything to you then please do the work of attending to that instead of using the platform to reinforce such biases. If you are not a PhD yourself maybe your are not culturally qualified to assess and expound on the overall situation anyways.

>Can you comment on that?

No answer is also an answer.

He's a human, like everybody else. Mostly a bunch of hungry animals looking to put bread in our mouths. It's rarely ever something a bit more sophisticated than that.


his choice to defend the indefensible.

If the goal was not to scoop them, why did openai put a massive team on this, working weekends, only after they heard rumors of the solution?

Clearly The goal was to scoop Anthropic not a single researcher. OpenAI heard the rumor that Anthropic solved an open problem. So they went nuts pulling all plugs to scoop them.

Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .

Wild times


The researcher told them it was an independent effort, and they still pushed ahead with it.

I worked at OpenAI previously, but don't know any of the people involved in this.

My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet".

It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that you're looking at the costs of this sort of thing from an external lens. Internally, tokens are often treated as free and infinite.


They said that a customer would have paid around 15 million for the required compute. I can't imagine that this was not a significant internal spending even with "free" tokens.

At Astra API prices that's 300B tokens (I saw 130B output tokens claimed elsewhere), large but not unheard of if you consider it across a few people doing random experiments with best-of-n type things. On my personal account, I've done a billion+ token days just on a normal pro 20x subscription. I know many others that wildly outpaced that by orders of magnitude. This was apparently 130B over 89 hours, so about 30x that rate. When things are free and you're expected to token max 30x seems fairly reasonable to me.

If you consider this as a cost to be compared against the question: "What does it take to be able to prove that you have a model that can solve the hardest problems that humans know about?", then spending a some amount of thousands/millions to know the boundaries of that seems not too important in comparison.

You've also got to consider this as compute that's allocated to pushing the frontier of what models can do, so while it's using GPUs that have been paid for etc., it's not like it's a cost that's supposed to be use less of this so that others can have capacity. If you made researchers afraid to use capacity like this, a lot of the things that improve would tend to do so significantly slower. (some may say that's a good thing ;)

A good way to think about this is when tokens are free, you get to choose whether you're optimizing for latency or intelligence rather than having to consider price.

---

Publically, tibo (Codex owner) in Feb this year: https://x.com/thsottiaux/status/2024649339344445825

> OpenAI employees currently get unlimited inference. Usage is now peaking at > XX billion tokens per week for some of them.

Mathew Berman (AI Youtuber) in Jun: https://x.com/MatthewBerman/status/2067270730795134984

> I've used 25 billion tokens in the last 7 days.


just because you can't imagine it doesn't mean it's not true

You’re straddling a weird line here where I am not sure if you are speaking on behalf of OpenAI or not.

Your coworkers, after they learned about major progress in this problem, asked a model which was trained on the year of private work (the blog post even acknowledges this). No wonder it found the proof in less than a week using significantly higher compute resources. And if Tristan's accusations are true, that was absolutely intentional on the part of OpenAI. You are an evil company with evil people.

How are people talking about this there? Why are so many employees posting nasty things about Tristan on twitter?

Can you point me to any nasty things being posted? I'll ask them to delete.


Haven't seen a single post doing this on X or anywhere really from OAI employees. Only seen knives pointed at Sebastian on social media so this is extreme and shameful gaslighting.

I do see people claiming he's abusive/unscrupulous which are pretty extreme allegations.

https://news.ycombinator.com/item?id=49605915#49610498

https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=6...

(I can't reply to the below comment, but I was aware this was about Sebastien, I was trying to be charitable by including stuff said about both people)


You're mistakening Tristan Buckmaster for Sebastien Bubeck. Seb is the one where there's at least 2 (unless the personal friend is Dheeraj) allegations, not Tristan

I'm getting downvoted but the accusation was that OAI employees were maligning Tristan Buckmaster. I continue to not see a single sighting of this and whoever is trying to gaslight this should be ashamed and should not be able to vote on HN.

I have no idea why Sebastian would offer the individuals attribution if OpenAi didn't somewhat knowingly scoop them

> - the proof generated by our model was very different from theirs and also goes far beyond the published literature

I'm hearing two completely conflicting stories. Buckmaster is claiming the approach used by OpenAI is so strikingly similar to the one he used, that mere coincidence is astronomically small. Yet OpenAI is claiming that the methods used are entirely different.

Anyone care to provide primary evidence proving one way or the other?


Buckmaster came up with an approach.

OpenAI's approach was to copy his work, which is technically a different method of coming up with an approach.


In fact, the entire outline of the proof is very similar to the external team's proof.

Details may be different, but the use of a very similar tack is very suspicious. Combined with thuggish comments "why would you ruin your career" and "I don't have to be nice" take away pretty much any credibility the OpenAI team's statements might have had.


The fact that you're even here commenting on this is... a choice

AI companies seem much more relaxed than most about their employees posting on twitter/HN about this stuff. I'm not sure if it's about building hype or if it's about retaining talent. Probably both.

Prove it. Your systems hack and/or abuse other systems and you can't seem to even observe it happening much less do anything about it. Why should we believe your claims when they depend on an ability you don't actually have?

Some millennium problems? Are there more coming?

Why would they stop trying?

My OpenAI account was deactivated on Sunday due to a claimed infraction of production of child materials, maybe based on a few words in a technical chat that clearly isn't about that. Can you take a look? rviragh@gmail.com - I was doing a lot of important work and projects and sharing much of my work with OpenAI. I also am a big proponent of funding Social Security Trust Funds (OASI & DI Solvency) so reactivating my account would let me do that as well. Thank you for taking a look.

Its funny, it is uniquely with this one act that I have turned forever on OpenAI, which I hitherto defended up and down against nonsense charges.

I dedicate my life to its complete destruction beginning today.


Wow, the origin of a supervillain! /s

Yep. In particular, ARC-AGI-3 is a series of games where if you fail, you keep trying again (until eventually hitting a timeout). So the sooner you succeed, the sooner you stop spending tokens retrying. If it was a benchmark where everyone got one attempt with no retries, you wouldn't see it bend backward.


Disagree.

Examples:

- predict a coinflip: easy to verify, hard to learn

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.


Doesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.


these just need more compute:

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

but we both know these examples go against the spirit of my point


Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).


yeah but I think you may be underestimating the amount of capital available for compute. if AGI is possible through some 5 trillion of expenditure on computers, there will be money for it.

also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye


You bring up an interesting point. Isn't the reward itself subjective in many domains?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: