IP Law Daily, AI NEWS: Senate bill would help copyright owners obtain information about AI training data, (Aug 6, 2025)
Organizations Mentioned:SoundExchange
Aimed at adding “transparency” to emerging tech, the bipartisan bill would establish a subpoena process for copyright owners to determine whether developers have used their copyrighted works to train generative AI models.
Legislation introduced by a bipartisan group of Senators would create an administrative subpoena process to help copyright owners determine which of their copyrighted works have been used in the training of generative artificial intelligence models. Titled the “Transparency and Responsibility for Artificial Intelligence Networks (TRAIN) Act,” the measure proposes adding a new section to the Copyright Act to create this mechanism. The bill (S. 2455) is sponsored by Senator Peter Welch (D-Vt.) and co-sponsored by three of his Senate Judiciary Committee colleagues, Marsha Blackburn (R-Tenn.), Adam Schiff (D-Calif.), and Josh Hawley (R-Mo.). The proposal was originally introduced by Senator Welch in November 2024 (S. 5379, 118th Congress). After introduction, that bill was referred to the Judiciary Committee but was not considered. Senator Welch reintroduced the measure on July 24. Full text of the legislation is available here, and a section-by-section summary is available here.
Purpose of legislation. “Copyright owners—particularly small creators—are struggling to navigate novel legal issues posed by AI copying their work,” a news release by the office of Senator Blackburn asserts. “There are very few AI companies that share how their models were trained and nothing in current law requires them to disclose training materials to creators.” Senator Blackburn said that “The TRAIN Act would protect creators by allowing them to access the courts to find out if their work is being used to train generative AI models and seek compensation for that misuse.” In a press release, Senator Welch said, “This is simple: if your work is used to train AI, there should be a way for you, the copyright holder, to determine that it’s been used by a training model, and you should get compensated if it was.” While the proposed TRAIN Act would provide a new means for copyright owners to obtain information about potentially infringing uses of their copyrighted works, the bill does not provide for new enforcement methods or remedies.
Subpoena process. According to the sponsors, the TRAIN Act’s subpoena process is “modeled on the process used for matters of internet piracy.” This could refer to Section 512(h) of the Digital Millennium Copyright Act, 17 U.S.C. § 512(h), under which a copyright owner who has notified an online service provider that copyrighted works are being infringed through the service can request the clerk of any federal district court to issue a subpoena to the provider for purposes of identifying an alleged infringer. TRAIN Act subpoena requests would be granted only upon submission of a copyright owner’s sworn declaration of a good faith belief the owner’s work was used to train the AI model, and that the purpose of the request is to protect the owner’s rights. The subpoena’s reach would be limited to the requester’s own works. The measure is limited to generative AI models, defined as “an artificial intelligence model that emulates the structure and characteristics of input data in order to generate derived synthetic content, which may include images, videos, audio, text, and other digital content.” Subpoena requests made in “bad faith” could subject the requester to sanctions under Federal Rule of Civil Procedure 11.
Recipient’s obligations. An AI developer served with a subpoena under the TRAIN Act would be required to “expeditiously” disclose the requested information. The developer would be required to hand over training records “sufficient to identify with certainty” whether the requester’s copyrighted works were used to train the developer’s generative AI model. If a developer fails to comply with a subpoena, that failure would provide a rebuttable presumption that the developer made copies of the copyrighted work, the bill states.
Reaction. According to the bill’s sponsors, the measure has the support of many organizations serving the creative community, including the Recording Industry Association of America (RIAA), SAG-AFTRA National, the three major U.S. performing rights organizations, SoundExchange, the Authors Guild, Songwriters Guild of America (SGA), and the National Association of Voice Actors (NAVA).
Some organizations with ties to the tech sector expressed differing views. Re:Create, a coalition of groups advocating for “balanced copyright and a free and open internet”—including the Open Technology Institute, the Electronic Frontier Foundation, and the Consumer Technology Association—criticized the proposal in a statement and urged Congress to leave it to the courts to develop a framework for copyright and AI. “As two federal courts have recently ruled, fair use protects the right to train AI models,” Re:Create Executive Director Brandon Butler said. “The runaway TRAIN Act would nullify the crucial fair use rights of AI researchers, developers, and users, creating a tool for legal harassment that would chill free expression, create overly burdensome requirements for small AI developers, and stifle innovation.” Butler characterized the legislation as “a hunting license for trolls.”
California AI transparency legislation. Legislation with a similar purpose but a somewhat broader scope was introduced in California in February, but that measure has stalled after passing the Assembly in May. Assembly Bill 412, the proposed “AI Copyright Transparency Act,” would create a standardized process for copyright holders to verify if their work has been used in AI training datasets. Some of the bill’s provisions were softened during committee review after criticism by tech companies, but in July, the bill was delayed and converted to a two-year bill, meaning that it won’t go forward this year. The disclosures required by AB 412 would build upon the “California Training Data Transparency Act” (AB 2013, Chapter 817, Statutes of 2024), enacted in September 2024 and effective January 1, 2026, which requires developers who make their gen-AI systems or models available to Californians to post high-level, summary information about the sources of their training datasets.
News: AINews Copyright TechnologyInternet GCNNews