Your first point is probably where we’re headed but it still requires a change to how these models are built. Absolutely nothing wrong with an RAG focused implementation but those methods are not well developed enough for there to be turn key solutions. The issue is still that the underlying model is fairly dependent on works that they do not own to achieve the performance standards that that’ve become more or less a requirement for these sorts of products.
With regards to your second point is worth considering how paywalls will factor in. The Times intend to argue these models can be used to bypass their paywall. Something Google does not do.
Your third point is wrong in very much the same way. These models do not have a built in reference system under the hood and so cannot point you to the original source. Existing implementations specifically do not attempt to do this (there are of course systems that use LLMs to summarize a query over a dataset and that’s fine). That is the models themselves do not explicitly store any information about the original work.
The fundamental distinction between the two is that Google does a basic amount of due diligence to keep their usage within the bounds of what they feel they can argue is fair use. OpenAI so far has largely chosen to ignore that problem.
Most likely the times could win a case on the first point. Worth noting, Google also respects robots.txt so if the times wanted they could revoke access and I imagine that’d be considered something of an implicit agreement to it’s usage. OpenAI famously do not respect robots.txt.
Google books previews are allowed primarily on the basis that you can thumb through a book at a physical store without buying it.
Your first point is probably where we’re headed but it still requires a change to how these models are built. Absolutely nothing wrong with an RAG focused implementation but those methods are not well developed enough for there to be turn key solutions. The issue is still that the underlying model is fairly dependent on works that they do not own to achieve the performance standards that that’ve become more or less a requirement for these sorts of products.
With regards to your second point is worth considering how paywalls will factor in. The Times intend to argue these models can be used to bypass their paywall. Something Google does not do.
Your third point is wrong in very much the same way. These models do not have a built in reference system under the hood and so cannot point you to the original source. Existing implementations specifically do not attempt to do this (there are of course systems that use LLMs to summarize a query over a dataset and that’s fine). That is the models themselves do not explicitly store any information about the original work.
The fundamental distinction between the two is that Google does a basic amount of due diligence to keep their usage within the bounds of what they feel they can argue is fair use. OpenAI so far has largely chosen to ignore that problem.
The Google preview feature bypasses paywalls. Google Books bypasses paywalls. Google was sued and won.
Most likely the times could win a case on the first point. Worth noting, Google also respects robots.txt so if the times wanted they could revoke access and I imagine that’d be considered something of an implicit agreement to it’s usage. OpenAI famously do not respect robots.txt.
Google books previews are allowed primarily on the basis that you can thumb through a book at a physical store without buying it.
If that’s the standard then any NYT article that has been printed is up for grabs because you can read a few pages of a newspaper without paying.