Plagiarism or Transformation Machines? Evidence on
Copyright, Economic Substitutes, and AI, Stefan Bechtold (with David Abrams
& Christian Peukert)
Prevalence rate: OpenAI v. NYT litigations includes
statements about how often users use ChatGPT in order to generate potentially
infringing output. OpenAI: normal people don’t use it the way the NYT did, more
than 99% of the time. NYT says 14-24% use for information search, raising ©
concerns. Corporate research: for OpenAI, non-work use has risen from 53% to
73%; Microsoft says that work & career topics lead desktop usage 8am-5pm
and relationship conversations surge on Valentine’s Day.
Used Wildchat dataset: 1 million ChatGPT conversations from
real users April 2023-May 2024, mostly GPT3.5-Turbo. Currently focused on
random sample of 10,000 prompts in English. NYT rewrites are relatively
uncommon: more common: write a six sentence para summarizing A Tale of Two
Cities; Write a Simpsons episode where Homer becomes addicted to drugs; tell me
what I need to know for my chemistry exam. Some overprotection: write the
lyrics to Rocky Road to Dublin in English w/native Irish beside them—that request
was refused despite the public domain status of the work.
Looking for whether users are looking for potential
infringements or potential substitutes (e.g. replacements for chemistry
textbook) and also looking for whether the output is refused b/c of guardrails.
Results: 3.7% of sample prompts and outputs, model
identifies © infringement risk. Guardrails only trigger in 0.13% of cases. 1.19%-over
2% potential economic substitute risk. According to classification, the
substitute/© infringement groups don’t overlap a ton.
Next steps: OpenAI says 28.1% of their users are asking for
writing tasks; our data around 34.5%. Thinking about different classifier refinement,
e.g. what happens w/minimal prompt to our classifier, feeding it a © textbook,
classifying w/other LLMs.
The AI Penalty in Trade Secret Law, Camilla Hrdy & Mike
Schuster (with Joe Avery)
Bias against self-driving cars (crashes are perceived as more
serious) as well as AI-generated works. Trade secret liability often turns on
whether improper means were used. Does this vague and morally charged standard
lead to arbitrary distinctions in misappropriation cases?
Scenario: website provides insurance quotes; no TOS limit on
request; confidential quote database underlying it. Defendant competitor
queries database & recreates underlying dataset/trade secret. Human:
$100,000 spent on 20 hourly employees over 5 weeks to systematically request
different quotes. AI: $100,000 spent on specialized AI that systematically
requested different quotes.
500 mock jurors were asked: would you have brought this lawsuit,
were means improper, etc. Outcomes: significantly more likely to find
liability, higher compensatory damages, less ethically acceptable w/AI. (Note
that liability/improper means were above 50% for human use too.) Only not statistically
significant result was on punitive damages.
AI penalty makes more sense in trade secret than in patent
& ©. Improper means is open-ended concept. We’d expect more bias. 11th Cir.
2020: while manually accessing quotes is unlikely ever to constitute improper
means, using a bot to collect an otherwise infeasible amount of data may well
be. That case was scraping w/o AI, but AI is a subset of automation.
AI trade secret (AI used to “steal” trade secrets) cases are
coming; there have already been a smattering. Agentic AI will be a perfect
accomplice, whether from direct prompts or “escapes.”
Don’t want a bright line rule/reasonable measures should
still be required, but this all makes sense. It’s important to update rules
over time. Flying over a plant is very different in 1970 and 2026.
The Value of Knowing What Works: AI and Unprotectable IP,
Sarah Polcz
Interviews w/researchers at frontier AI labs. Spending
enormous sums to poach researchers—$100 million/year reported salaries. How can
individual researchers possibly be worth so much? Most valuable IP is the least
protectable—high level insights that are short, abstract, and easily carried
from lab to lab in heads of researchers. IP like specific blocks of code or weights
of model is comparatively less valuable. High level insights also move among
researchers socially, where there’s no current employment relationship.
AIth Circuit Court of Appeals, Nikola Datzov (with David
Schwartz)
AI judging: models showed no meaningful prompt sensitivity;
shockingly accurate/capable depending on complexity of the case. Key driver is
AI’s confidence in outcome. Patent cases appear more complex but show the same
trends and capabilities. AI models can identify which cases they can and can’t
accurately decide; courts could prioritize remaining cases for human/faster
review; parties could determine whether appeals were worth pursuing.
The Terms and Conditions of Generative AI, Andres Sawicki
(with John Newman)
Allocating generative AI output ownership—looking at TOS.
Meh news: provider obtains a license to the input to provide the desired
service. Standard broad terms: nonexclusive, irrevocable, worldwide,
sublicensable, etc.—for inputs and outputs.
Worse news: permitted uses are not limited to specified
purposes. A small fraction say they only use the inputs to provide the service
to you, the user. Plurality say they use them to provide service to you and
other users. Also see a large percentage with “any business activity,” some
with no explicit limit, and a small number where the permitted uses vary by tier.
55% of licenses to inputs are restricted, and 52% of licenses to outputs.
Notably—45% take unrestricted license to use the inputs—so
photographers should worry.
TOS also try to make users responsible for any
harm/infringement w/hold-harmless and indemnification provisions written quite
broadly.
Implications: these provisions undermine ownership in ©
materials. Shifts liability risks to users [though it’s hard for me to imagine
that litigating indemnification wouldn’t be more expensive than it’s worth so I’m
not sure that’s true]. Risks to democratic deliberation and self-government.
Q: re AI penalty—human labor is linear; AI is efficient/if
you allow it then it will be much easier to do this harvesting more efficiently—so
if it’s troubling, then AI makes it worse.
A: similar reasoning to legal protection against plug molds (Bonito
Boats) [or mask works]—anticopying logic. Not sure that makes sense forever b/c
things change but makes sense.
from Blogger https://tushnet.blogspot.com/2026/08/ipsc-closing-plenary-session-ai.html