For decades, arXiv has been the cornerstone of open scientific communication. But today, open access for humans has become a free pass for commercial AI models to harvest our research for training data without consent or attribution.
Currently, authors posting to arXiv have no control over whether bots scrape their full-text PDFs and LaTeX source files.
We are calling on arXiv to implement a simple permission toggle during submission, allowing authors to decide what automated scrapers can access:
Full Access: Bots can read and scrape the full text and source files.
Metadata Only: Bots are restricted to the Title, Abstract, and Citation metadata.
Open science for human readers should not mean mandatory, unconsented data harvesting by commercial entities. Advancing scientific knowledge demands immense time, resources, and intellectual rigor — we deserve a say in whether our original work is fed into massive, opaque training sets.
We invite all researchers with an active arXiv account to add their name below in support of this petition.
Add your signature
Signatures are verified by email. You will receive a confirmation link at the address you provide — click it to make your name appear publicly.
A verification link has been sent to
Click the link in that email to confirm your signature. It expires in 72 hours. Check your spam folder if you don't see it within a minute or two.
Signatories
- Loading signatories…