Two different problems (do not mix them)
A. Rights in training corpora
- Was protected expression copied into a dataset?
- Under what license, exception, or contract?
- Who controls crawls, filters, and opt-outs?
B. Rights in outputs & hosted copies
- Does a published page copy someone else’s expression?
- Is there human authorship in what you published?
- Can an OSP remove a specific URL after a notice?
Parent overview: AI and copyright. Subject-matter basics: what is copyright?(ideas and facts vs protectable expression under frameworks such as 17 U.S.C. § 102).
How training datasets are usually built
At a high level, training stacks combine sources such as:
- Web crawls — pages collected under robots rules, ToS, and internal filters (or without them)
- Licensed corpora — books, news, code, or media sold or open-licensed for training
- Synthetic / model-generated text — used to expand or balance datasets
- Human-curated subsets — quality filters, safety filters, deduplication
Each path changes the risk profile. “It was online” is not the same as “we may reproduce distinctive expression elsewhere.” Public-domain myths also appear here—term expiry is country-specific; see copyright duration by country and public domain basics.
Publisher and site-owner risks
Even if you never train a model, AI-era copyright fights can still hit your business:
- Scraped mirrors of your articles on third-party domains, then reported or monetized by others
- False or overreaching notices against your original pages (including AI-assisted drafts you edited)
- Bulk automated complaints that ignore context—see copyright bots and automated DMCA
- Confusion over ownership when contractors use AI tools without clear IP assignment
Patterns of over-claiming: what is copyfraud?
Can you “DMCA a model” or an output?
The DMCA notice-and-takedown system used by hosts and search engines is built aroundmaterial stored or indexed at identifiable locations—typically URLs and files held by an online service provider—not an abstract set of model weights floating in research papers. Practically, site owners and operators focus on:
- Hosted pages or files that copy protected expression
- Search results removed after a copyright complaint
- Platform appeals and, where available, § 512(g) counter-notices
Whether training itself infringes is a separate litigation question. This page does not predict outcomes of any lawsuit; it maps what you can usually do when your index or hosting is on the line.
Compliance checklist for AI products (non-advice framing)
- Document dataset provenance and licenses for training and fine-tuning
- Honor robots/ToS policies you claim to respect; record opt-out handling
- Filter or review outputs for high-risk regurgitation of distinctive works
- Keep human review gates for customer-facing publications
- Define IP ownership in vendor and freelancer contracts when AI tools are used
- Maintain a notice-response playbook (legal + ops) before a multi-URL incident
US Copyright Office materials on AI and registration policy evolve—teams should track official guidance rather than blog rumors: copyright.gov/ai.
EU sketch (jurisdiction-specific): the EU Digital Single Market Directive (Directive (EU) 2019/790) includes text-and-data-mining exceptions (commonly discussed under Articles 3–4) with conditions such as lawful access and, for certain uses, rightholder opt-outs. This is not a US fair-use clone—do not paste a US training narrative into an EU product plan without mapping local implementation and counsel review.
If rankings drop after an AI-related copyright complaint
- Verify — Google Search Console legal removals / host tickets; run DMCA / index check.
- Preserve — notice, timestamps, full HTML, drafts, prompts, licenses.
- Classify — genuine similarity dispute vs wrong URL vs over-claim vs competitor abuse.
- Respond — form builder here, then counter-notice when you have good-faith grounds.
- Separate product risk — fix dataset/output process so the same complaint class does not recur.
Sources
FAQ
Is training on public web pages automatically fair use?
No automatic rule covers every crawl. US fair-use analysis is fact-specific (purpose, nature of the work, amount, market effect). Other countries use different exceptions, such as text-and-data-mining rules in the EU. Public availability of a page is not a free license to copy protected expression.
Does “AI-generated” mean no copyright and free to copy?
No. “AI-generated” is not a magic free-for-all label. Human-authored works remain protected under ordinary copyright rules. AI-assisted publications may still contain protectable human expression—and may also risk resembling someone else’s protected expression.
Can competitors claim my AI-assisted article?
They can send a notice; that does not mean the claim is correct. Preserve drafts, prompts, publication timestamps, and source materials. If Google removes the URL after a copyright complaint, evaluate a counter-notice when you have a good-faith right to keep the material online.
What evidence helps if Google removed my URL after a copyright complaint?
Save the notice and reference IDs, the full page as published, first-publication proof (CMS, Wayback, source files), and any licenses. Then verify index/legal-removal status and follow a structured counter-notice path when appropriate.
Can I DMCA the model weights themselves?
DMCA notice-and-takedown is aimed at material hosted by online service providers—typically specific URLs or files—not abstract model parameters in the abstract. Practical enforcement for site owners usually targets hosted copies and search listings, not “the model” as a single file you can easily identify on Google.
Related: AI and copyright ·What is copyright? ·Copyright bots ·Hyperlinking & framing ·Counter-notice guide ·Knowledge base