But displaying regurgitations of very similar content may not be fair use. Fair use is a very delicate affair. One factor is whether the modified work poses as a market replacement for the original work.
The issue is, in part, a concern that ChatGPT responses are often just simple derivations of the original content in ways that wouldn’t be considered fair use.
Hmm, this is an interesting framing of the lawsuit. If it's about outputs and not just training, are the outputs really orthogonal to the training?
In traditional computer systems, no, outputs are always a function of inputs. LLMs throw a wrench into this reasoning because they apply opaque statistics to a combination of training data and the user prompt to produce outputs, so the input-output relationship is much less clear, but fundamentally it still holds.
So then this case should also be about training. The question then is: did OpenAI intend to have these models be able to regurgitate large amounts of content? Or is it yet another emergent property that nobody anticipated?
I would suspect the latter, because if you view these models as a lossy compression of the whole Internet (cf "Blurry JPEG of the Web" article) it is a surprising outcome that they are able to losslessly reproduce so much of the original content.
So this might come down to intent. Maybe the NYT would need to show that OpenAI intentionally designed for this property, e.g. by rewarding reproductions of entire segments of the original content in its training. In which case, it's looking in the wrong place for evidence.
>Hmm, this is an interesting framing of the lawsuit.
First, it's not a "framing" of the lawsuit. A lawsuit is a number of claims made by one party against the other. In the two California cases, there were no decisions made on claims relating to LLM outputs. In the NYT case, there are claims relating to LLM outputs.
Yes, it could also be about training. But the discovery pertains to the outputs, which is the issue in this case. So even if you apply the holding that training is fair, which I don't see likely to happen in the district courts of the second circuit, you still don't get the result that the person I responded to suggested, which was that this should all be moot because of two decisions in two different cases in California which are not binding precedent in the 2nd circuit, and which also would not dispose of all of NYT's claims.
>So then this case should also be about training. The question then is: did OpenAI intend to have these models be able to regurgitate large amounts of content? Or is it yet another emergent property that nobody anticipated?
Intent is not a required element of copyright infringement, so you'd be wrong there. Plaintiffs can use intent to evidence willful infringement, which they are entitled to do in statutory damages cases, and receive a damages multiplier, which this one is. So OpenAI can't avoid liability based on their intent or a lack thereof. They can only, at best, use 'intent' to establish that NYT's is not entitled to heightened damages.
>So this might come down to intent.
It's always amusing to see people apply completely made up rationales to legal cases based upon their own personal feelings about technologies while completely disregarding, lets say, 100 years of legal jurisprudence.
Oh I'm totally an armchair lawyer, so my ruminations were not grounded in laws or legal precedence :-) I do have some background on the patent side of things, where independent reinvention is also not a defence for infringement, but not so much in copyright, so this was educational.
However, has there been any case where the infringment was not only unintentional, but also unexpected?
That is, if you look at cases of uintentional infringement, these are typically cases where some the act of reproduction of content was intentional, but there was a lack of awareness or confusion about the copyright protections of that content. (This paper was useful for background: https://www.law.uci.edu/faculty/full-time/reese/reese_innoce...)
But I could not find a case where the act of copying itself was non-intentional.
In this case, looking at how LLM training works and what LLMs do, it is surprising that it could reproduce the training content verbatim. The fact that it reproduced those outputs is undeniable, but how does existing law and jurisprudence apply to an unprecedented case like this where the reproduction was through some magic black box that nobody can decipher?
These are interesting questions but they are not legal questions. Intent is not an element of infringement. It is only an element of willful infringement. Therefor it can never be used as a defense against infringement on its own.
>The fact that it reproduced those outputs is undeniable, but how does existing law and jurisprudence apply to an unprecedented case like this where the reproduction was through some magic black box that nobody can decipher?
People love to ponder... but ponder how the law should handle that... "Yes, your honor, our business has a magical black box that violates the law, we're just not sure how! Therefore we can't be liable" -- How does that even make sense? On what principle should that apply here and not elsewhere? Can your magic black box murder? Defame?
> On what principle should that apply here and not elsewhere? Can your magic black box murder? Defame?
Good questions, and I think relevant to the current point. We're already seeing cases like that pop up with the libel suits or the recent, tragic AI-assisted suicides.
It's very clear that these models were not designed to be "suicide-ideation machines", yet that turned out to be one of the things they do! In these cases the questions are definitely not going to be about whether the AI labs intended these outcomes, but whether they took sufficient precautions to anticipate and prevent such outcomes.
One possible defense for the AI labs could be "these machines have an unprecedented, possibly unlimited, range of capabilities, and we could not reasonably have anticipated this."
A smoking gun would be an email or report outlining just such a threat that they dismissed (which may well exist, given what I hear about these labs' "move fast, break people" approach to safety.) But without that it seems like a reasonable defense.
While that argument may not work for this or other cases, I think it will pop up as these models do more and more unexpected things, and the courts will have to grapple with it eventually.
Exactly. And the OpenAI corporates speak acting like they give a shit about our best interests. Give me a break, Sam Altman. How stupid do you think everyone is?
They have proven that they are the most untrustworthy company on the planet
And this isn't AI fear speaking. This is me raging at Sam Altman for spreading so much fear, uncertainty, and doubt just to get investments. The rest of us have to suffer for the last two years, worrying about losing our jobs, only to find out the AGI lie is complete bullsh*t.
To me, no company has the customers’ best interests in mind. This whole thing is akin to when Apple was refusing to unlock phones for the FBI. Of course, Apple profits by having people think that they take privacy seriously, and they demonstrate it by protecting users’ privacy. Same thing here; OpenAI needs chats to have some expectation of privacy, especially because a large use case of AI is personal advice on things. So they are fighting to make sure it's true.
Both OpenAI and NYT are bad. I don't know about NYT's privacy policy, because that's not really the industry they're in, but they did admit to fabricating a story that led to a now 2-year-long war, so.
Yes, but I think at least in this instance, OpenAI needs people to think that what they ask ChatGPT is private. They will have no business model if everyone thought that whatever private question they ask could fall into the hands of a media company and be used for anything. Also, at least when I signed up, you had to provide either a highly trusted email address or phone number to sign up, so your identity is definitely attached to whatever question you ask ChatGPT. They know how high the stakes are for them in this suit.
We need to be careful and mindful of our framing. Saying "X is bad" is a drastic oversimplification and not necessarily useful. Pointing at any one company and saying "bad" doesn't move the needle much in terms of figuring out how to steer us towards better outcomes. For that, we have to identify incentives and understand motivations.
It's weird. It went up and down and up and down. Controversial POV. But thanks for the support. Sam Altman's just too dishonest. It's been said time and time again by so many people, by Paul Graham, Ilya Sutskeve, everybody's telling everybody he's dishonest. When are we going to wake up and get this guy out of there?
> They should sell their stuff by mail if they hate open culture so much.
Does open culture mean free? Are you willing to work for free? It is perfectly OK to sell goods in exchange for money, which is what NYT is doing.
I dont know why you're so upset with it. You cant walk into Apple Store and except to walk away with a free iPhone. Then why are you expecting to "walk" into nytimes' website and walk away with free article?
The problem isn't that a news site is monetizing with a paywall. Totally fine, monetize how you want!
The problem is that prominent news orgs have lobbied governments all over the world to threaten google, apple, etc. for preferential treatment so these paywalled articles get prominent placement in various feeds and carousels and recommendation algos.
As a small publisher you'll never get this same preferential treatment if you throw up a paywall.
Creating the bizarre situation where big tech platforms feel they have to recommend paywalled articles from NYT/Bloomberg/etc, catfishing users right into a paywall when they click on headlines. This is essentially spam.
Open means open. Plenty of people make money in the open culture in way less obnoxious ways than NYT. What NYT does is crapping at the place where I am, but building a wall and charging for passage to a place that does stink little bit less. I don't mind them having such place, or even charging for access. What I mind is making mine actively worse. Do whatever you want and charge however much you want. But for the love of God don't advertise in my face using free space that I inhabit. My attention costs way more than your content. Don't be surprised that when you do I will disregard completely your wishful thinking about payment.
What I need is one checkbox in Google ecosystem (and/or my browser) that says "Never show links to paywalled content". Give me that and all my beef with NYT and similar garbage factories is gone in a blink of an eye.
Yeah, how terrible that you should be expected to spend /eleven minutes/ of the average U.S. tech worker's salary for a month of information. Perish the thought.
They should sell their stuff by mail
You're in luck! You can subscribe to the New York Times by mail, just like you want.
I wrote myself an extension to bypass all youtube adverts and used it for years. I'm perfectly capable of evading NYT garbage once the fury exceeds the lazyness. Still the issue remains. I'm not the only one bothered by paywalled links in search results, being linked from websites and suggested in feeds of mobile apps. Checkbox to filter them out was requested long time ago. Never implemented.
The issue isn't the Times, since you've admitted that you have a way to avoid them, and other people have suggested solutions.
The actual issue is that you enjoy being angry and expressing that anger in front of strangers on the internet, as if that somehow validates your anger, or makes you feel good, or gives you some other kind of reward for grinding your personal axe.
This is destructive behavior. I recommend introspection. Failing that, seek professional help.
Sure, that too. I just like fierce discussions about irrelevant, unchangeable things and they are easiest to find in the company of people with bland, mainstream opinions. Somehow they always try to defend them vehemently.
I have plenty of introspection. I know exactly what I am doing and why.
I'm glad the NYT is fighting them. They've infringed the rights of almost every news outlet but someone has to bring this case.