AI companies destroy physical books – let's scan rare books before it's too late
annas-archive.gl478 points by Cider9986 20 hours ago
478 points by Cider9986 20 hours ago
I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.
https://en.wikipedia.org/wiki/Google_Books
https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included.
I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.
At one point Google Books was supposed to act as a clearinghouse for scans of out-of-print books. You could have purchased a scan of any book on the site for a reasonable price, and libraries could subscribe to a service where the full text of all books was available. This settlement then got shot down because some research libraries and authors argued that this was anti-competitive, and instead wanted Congress to pass a law to free up the rights to orphaned books so anyone could start a competing service. No progress on this was subsequently made because nobody in Congress cares enough about the rights to out-of-print books to get legislation passed. The whole reason why they're out of print when ebooks and print-on-demand exist is that they won't get enough sales to make it worth the time and money to figure out who the royalties should go to. It's not a flashy issue that would make a ton of people vote for you to get re-elected, and it won't create a ton of new jobs. The result is that now nobody outside of Google gets to see the full Google Books scans.
Yes, this was a great tragedy. I was very sad to see academics at the time arguing against Google providing what would have been one of the greatest storehouses of readily available knowledge in the world, in favor of an imaginary alternative that didn't exist and never would.
> Google providing what would have been one of the greatest storehouses of readily available knowledge in the world
I'd be worried about how much they'd be charging for access once they had the monopoly on so many rare books.
This is exactly the kind of pointless concern I'm talking about. How many legal digitial providers of those rare books are there now? Zero, I believe.
It's the logic of cutting off one's nose to spite one's face, which prefers that no one benefit rather than Google see any benefit.
Then at least there would be outrage to drive the passing of the needed legislation which otherwise hasn't come to pass anyway.
Copyright law desperately needs a production requirement or allowance.
The copyright owner must make new copies of the work available; the price must be no greater than the original price (not inflation adjusted). And if they fail to do so, anyone may produce copies and escrow the original price (less the cost of production) for collection by the copyright holder.
That means that orphan works are effectively in the public domain. Calculus professors can ask students to get the cheaper 2nd edition, not the latest 22nd edition. And a company like Google could make scanned works available in their entirety for a small amount of money for each work. And the copyright holder still gets their end, without having to arrange a printing or hold stock.
As a photographer, do I have to make every photograph that I've ever sold available to anyone to buy forever more? Can I refuse to sell a print to someone? I wasn't famous when I sold one for $20 back in the 90s... if I became famous, would I still need to sell that at $20 (inflation adjusted)?
What happens to limited editions of print runs? Can I not make a run of 200 prints anymore because the 201st will be something that someone could request?
Does a musician have to license any song they made to anyone who asks? Can they refuse to license a song to some organization they disagree with and not have it fall into the orphan works category?
---
Amending copyright to the way you describe requires a renegotiation of the TRIPS agreement ( https://en.wikipedia.org/wiki/TRIPS_Agreement ) with all the nations of the WTO (or withdrawing from the WTO).
Under my proposal, if someone wants to make prints of one of your old photographs, you are getting ~$20, which is ~$20 more than you are getting now. So it doesn't seem like a damage to you. If you get famous and now you could sell a print for $2,000... how is this helping society since the work has already been produced, so no new incentive is necessary?
I suppose the rule could make reference to a rival good, i.e. the 22nd and 23rd editions of the calculus textbook. It should not be reasonable for the publisher to make the 22nd edition only available for $1,000, and the 23rd edition for $250. But that definition would invite many lawsuits and chilling litigation in general. A clear definition based on the historic price is much simpler.
You can still number your limited runs, and your limited runs still have increased numismatic value over some other reproduction.
I don't envision licensing of performance rights in this system, just recordings or reproductions.
Is this a reduction of the rights granted by copyright? Yes, intentionally so. It is stripping copyright holders of the right not to copy, against the interests of society in granting that copyright in the first place.
> It's not a flashy issue
It's a very serious issue, very well known to the people with the connections and power to affect it.
> [ it wouldn't ] make a ton of people vote for you to get re-elected
The tons of people are moved by the media, people are oblivious to the tricks of that trade, for the same reasons, obviously. In other words, this issue isn't something that happens to slip below the radar, it's kept stealthy by well organized engineering and considerable expense.
> and it won't create a ton of new jobs
Nothing ever creates tons of new jobs, the "tons" are reserved for promises and other useless noise.
Quod licet Iovi, non licet bovi
Big companies will read up the books and make their AI recite them from memory, but Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)
> Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)
No, this is what they were doing before, but they explicitly started lending out "unlimited" copies, which is why they got sued.
That's why they got sued, but the suit is mainly over whether controlled digital lending is legal at all rather than their "emergency library". Archive.org lost the case on summary judgment, meaning that they could not come up with a single fair use argument for CDL that the judge found compelling enough to let the case go to trial. The full judgment is here https://storage.courtlistener.com/recap/gov.uscourts.nysd.53... but here's a couple excerpts:
> The crux of IA's first factor argument is that an organization has the right under fair use to make whatever copies of its print books are necessary to facilitate digital lending of that book, so long as only one patron at a time can borrow the book for each copy that has been bought and paid for. See Oral Arg. Tr. 31:10-15. But there is no such right, which risks eviscerating the rights of authors and publishers to profit from the creation and dissemination of derivatives of their protected works. See 17 U.S.C. §§ 106(1), (2). IA's wholesale copying and unauthorized lending of digital copies of the Publishers' print books does not transform the use of the books, and IA profits from exploiting the copyrighted material without paying the customary price. The first fair use factor strongly favors the Publishers.
> In this case, there is a "thriving ebook licensing market for libraries" in which the Publishers earn a fee whenever a library obtains one of their licensed ebooks from an aggregator like OverDrive. Pls.' 56.1 ¶¶ 577-578. This market generates at least tens of millions of dollars a year for the Publishers. Id. ¶¶ 170, 172. And IA supplants the Publishers' place in this market. IA offers users complete ebook editions of the Works in Suit without IA's having paid the Publishers a fee to license those ebooks, and it gives libraries an alternative to buying ebook licenses from the Publishers. Indeed, IA pitches the Open Libraries project to libraries in part as a way to help libraries avoid paying for licenses. See Pls.' 56.1 ¶ 383 (presentation IA gave to libraries asserting that pairing with IA means that "You Don't Have to Buy It Again!"); id. ¶ 382 (different presentation promising that the Open Libraries project "ensures that a library will not have to buy the same content over and over, simply because of a change in format"). IA thus "brings to the marketplace a competing substitute" for library ebook editions of the Works in Suit, "usurp[ing] a market that properly belongs to the copyright-holder."
> suit is mainly over whether controlled digital lending is legal at all
No, it was not, even supported by the quotes you pulled. Libraries right now, with publisher blessing, offer all manner of controlled digital lending. The suit was because IA did it buy undercutting the publishers copy rights to that legal market. Had IA simply done what every other library has done to provide controlled digital lending, there would be no suit.
"Controlled digital lending" is not a generic term for "lending digital items". It specifically refers to the practice of a library digitizing physical materials in its collection, then lending them digitally based on a 1:1 owned-to-loaned ratio. The idea is that the library should be able to treat digitized versions of a book the same way it treats the physical book, and the total number of physical and digital copies of the book that are lent out at once should never be more than the number of physical copies that the library has.
In contrast to this, the e-book lending practiced by most libraries with publisher blessing involves the library purchasing special library-specific e-book licenses from the publisher. These licenses contain various contractual restrictions, such as the library having to re-purchase the e-book after a certain amount of time or after a certain number of borrows.
There is so much misinformation/confusion about this... they go sued after lending "unlimited" copies, but they were sued (and lost) for lending exclusive copies (controlled digital lending):
> “At bottom, [the Internet Archive’s] fair use defense rests on the notion that lawfully acquiring a copyrighted print book entitles the recipient to make an unauthorized copy and distribute it in place of the print book, so long as it does not simultaneously lend the print book,” Judge John G. Koeltl of the U.S. District Court in Manhattan wrote. “But no case or legal principle supports that notion. Every authority points the other direction.” [0]
[0]: https://www.insidehighered.com/news/tech-innovation/teaching...
> Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included.
> I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.
Can you still search the restricted parts? If so there's still value to it: it helps you identify the book so you do an inter-library loan to get at the full content. Sure, it's not frictionless, but I wouldn't be all or nothing about it.
> The project was met with significant legal challenges from authors and publishers which was eventually overcome
I don't think they were overcome. As far as I remember Google couldn't make the books available so they abandoned the project. They possess the scans (if they didn't delete them) but they won't be made public.
IIRC the courts ruled that because it would be implausibe for a person to use google books previews to read an entire work (you'd have to make a whole bunch of separate accounts to do so), it could not plausibly impact the market for that work.
The consensus among copyright lawyers is Google lost Authors Guild v Google, but for some reason the media and Wikipedia do not clearly report it that way.
Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space.
It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.
> It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.
They never should have been trusted in the first place.
I think you're missing the point. Maybe I'm wrong, but I'm pretty sure they're trying to point out that this book scanning doesn't need to be destructive. I'm not sure if AI companies are using a scanning method that damages the book or not, but they destroy the books after scanning to avoid copyright issues (ie they aren't duplicating the books). This Google project seems to demonstrate that this isn't actually necessary, and that the AI companies are just doing it out of laziness.
If it isn't scanned destructively, is it a liability?
Can you do anything else with the book? What are its costs for storage in a way that retains the value of the book? If the assets of the warehouse are sold to another company (see also https://paizo.com/blog/paizo-restructuring-a-difficult-updat... ), what are your obligations for the format shifted copy that you retain?
These questions imply that there's a liability that exists when retaining the original that has little value to the company. And they (the books) aren't assets that can be resold.
It's easier (and cheaper), has no ongoing costs for physical storage, and answers those questions without creating legal entanglements for the future company.
Questions don't imply anything, you're just asking things. I don't have the answers to your questions, but the point is that Google was able to scan books without destroying them while still avoiding legal consequences (with some effort).
The physical books that were scanned by Google were returned to the libraries.
https://btaa.org/library/programs-and-services/book-search/f...
> Will scanning harm the books?
> No. Google developed innovative technology to scan the content without harming the books. Any book deemed too fragile will not be scanned by Google, but may be treated by expert library staff. Once scanned, all print volumes are returned to the library collections.
That was an inherently different goal (borrow the books from the library, scan them, and return them) than the Bartz v. Antrophic ruling.
https://cases.justia.com/federal/district-courts/california/...
> Storage and searchability are not creative properties of the copyrighted work itself but physical properties of the frame around the work or informational properties about the work. See Texaco, 802 F. Supp. at 14 (physical), aff’d, 60 F.3d at 919; Google, 804 F.3d at 225 (informational); Sony Corp. of Am. v. Universal City Studios, Inc. (“Sony Betamax”), 464 U.S. 417, 447 (1984) (rightful interests). In Texaco, the court reasoned that if a purchased scientific journal article had been copied “onto microfilm to conserve space, this might [have been] a persuasive transformative use.” 802 F. Supp. at 14 (Judge Pierre Leval), aff’d, 60 F.3d at 919 (reducing “bulk[ ]” “might suffice to tilt the first fair use factor in favor of Texaco if these purposes were dominant“). In Google Books, the court reasoned that a print-to-digital change to expose information about the work was transformative. Google, 804 F.3d at 225 (Judge Pierre Leval). And, in Sony Betamax, the Supreme Court held that making a recording of a television show in order to instead watch it at a later time was copying but did not usurp any rightful interest of the copyright owner. 464 U.S. at 447, 455. Important to the Supreme Court’s reasoning was the expectation that most such copiers would not distribute the permanent copies of the work. Finally, in A&M Records, Inc. v. Napster, Inc., our court of appeals recognized the reasoning just explained, and therefore rejected by contrast a digitization effort that was touted as space-shifting but in fact resulted in the multiplication of copies shared with outsiders through a file-sharing service. 239 F.3d 1004, 1019 (9th Cir. 2001), aff’g in this part 114 F. Supp. 2d 896, 912–13, 915–16 (N.D. Cal. 2000) (Judge Marilyn Hall Patel) (citing Sony Betamax and Texaco).
> Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others).
---
The AI training isn't borrowing from libraries and returning from libraries. Instead, it is buying a book, format shifting, and retaining that format shifted version from its own use. The company can't do anything else with the book once they've format shifted it. They can't donate it and they can't resell it. In that light, destructively scanning the book is the best option. There is no value in trying to non-destructively scan it because otherwise all it would do is sit in a warehouse and cost money to pay for storage of something they can't sell.
> The AI training isn't borrowing from libraries and returning from libraries
Perhaps that's an important distinction here, I was really just trying to explain what I thought the commenter was trying to say. Calm down with the copy pasta walls.
But I still think that AI companies could make a very similar argument to the one Google made in your copy-pasta. The laziness I refereed to earlier is the fact that they haven't even tried. They could donate the books afterwards which would go a long way towards helping that argument in court.
My copying and pasting is twofold.
First, in today's environment people will state AI hallucinations as fact or "google it yourself" as the reference. If someone doesn't know where to find that information they're left with the "someone on the net said XYZ".
Secondly, I'm not always certain that if I do provide a link to a large document that people will find the relevant section in there. And second and a halfly, if someone comes back to it in a year or two or five that the link will still be live. I've had situations in the past where I've provided a link and then the domain changes hands and the new owners of the site put up a retroactive robots.txt and make it inaccessible on the wayback machine.
---
The question for an AI company that can't return, sell, or donate the material after it has been scanned - are they then to pay to store it in a warehouse indefinitely?
After they've format shifted the content and retained the format shifted content for continued use in training, it becomes at least a very gray area to donate the original.
And even if the books were scanned like Google Books did, and they could donate them to libraries - the libraries don't want those books. These aren't rare books in that people would expect you to wear cotton gloves while handling them... they're books that have mostly disappeared from availability.
https://old.reddit.com/r/books/comments/1vugion/the_federal_... provides an example of what is being seen in a used book store:
> I work for a large used bookstore with an online component. We're getting slammed with orders for books like the proceedings of an obscure 1992 Dutch geology conference or $500 festschrifts about D-module applications we would have previously sold to some university library. We've never once had an order for anything anybody would actually want, and most of this shit has sat on our shelves for years, if not decades.
Neutron Radiography: Proceedings of the First World Conference San Diego, California, U.S.A. December 7–10, 1981 is technically a rare book. https://www.amazon.com/Neutron-Radiography-Proceedings-Confe...
If you had a copy of it, I would challenge you to find a library that would accept it as a donation.
In the event that you wanted to read a copy of it, there is a copy of it in the Library of Congress. https://search.catalog.loc.gov/instances/b220c9bd-63ad-5a8b-...
The answer to his question is "yes." If the book isn't destroyed, it becomes a liability to Anthropic, because then it's no longer format-shifting.
If you're going to destroy the book anyway you'll go with the easier method for a good scan(separate the pages from the spine)
They are also an AI company now. Why would they stop?
Also, why would they ever share their collection?
Book scans, secreted away, are worthless to the public.
They could use them as training data, without providing access the actual books.
At least we would all benefit from the books this way, so long as legal nonsense keeps the scans unavailable to the public.
I don’t think locking the content of rare books away in the hands of corporations who only give us access to tools trained on the books, and not the actual book, is a good path forward.
This doesn’t incentivize them to be good stewards of this data and making anything in the public domain available. It incentivizes less access to the source material, having to blindly trust their tools, and is effectively automating plagiarism.
The irony of countries blocking Anna's Archive (UK, Italy, Netherlands etc) but then it's the hackers who conserve and steward the books. Because governments can't stop AI companies from shredding history like they did with Google Books _because_ they actually tried to do it _by the book_.
Amazon also did this for their "Look Inside" feature. To do this they had to spin up massive digital infrastructure——then realized that they could sell that infrastructure and make more money than from the books they were scanning.
This is was what ended up spinning up AWS as a business.
I dislike the idea of destroying rare books--but how rare? A digital copy has a lot more benefit.
I've worked in academic libraries since 2010. I've always felt like it was a mistake for our digital library leadership to put so much trust in Google Books despite their promises to maintain the integrity of libraries (this mostly happened by the way).
At the time it was obvious and innovative but over time it was clear Google was establishing a technical precedent to corrode what libraries have the power to do. I'm hoping we can continue to do the good work but it's exhausting.
Internet Archive version: https://openlibrary.org/
Info on where to send books not yet in their collection: https://help.archive.org/help/how-do-i-make-a-physical-donat...
Mobile apps to determine if they need a book: https://help.archive.org/help/donate-books-app-for-ios-and-a...
Web app: https://archive.org/want/?mode=donation_book
For example, I donated a copy of Systems Bible (out of print, hard to find imho) and paid for it to jump the digitization queue (https://archive.org/details/systemsbiblebegi0000gall/). The original book will remain stored as a physical backup. It's not fully publicly available of course due to copyright (it will eventually be made public by the Internet Archive once its copyright expires ~2084 and it enters the public domain), which is where shadow libraries|archives like Anna's Archive and Z-Library fill the gap.
If you have rare books you would like digitized, archived, and distributed, I am very interested in providing assistance.
As an aside... https://www.google.com/books/edition/The_Systems_Bible/mrOsb... (and I haven't hit any "you can't read this" limits yet).
While the hard copy is a bit pricy for my shelf of curious books, it's also available on kindle. https://www.amazon.com/SYSTEMANTICS-SYSTEMS-BIBLE-John-Gall-...
If you read the judgement against (I think) openai, the judge said that it was OK to scan the books for LLMs, if they were destroyed afterwards. That is, only one copy of the data existed.
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
> a digital copy with the ability of doing millions of copies is stored somewhere
somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”
the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.
> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.
archivists keep everything, because we don't know right now what will be important 100 years from now.
> somewhere were we can't access it
By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.
Plus having the info part of a LLM makes it immediately available to literally billions.
I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.
>Plus having the info part of a LLM makes it immediately available to literally billions.
Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.
If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".
And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.
Depends upon what you want.
For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.
The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.
But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.
That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.
> But, that little bit of data is a bit more data than existed before,
No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.
> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.
That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.
I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.
> Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost
I understand what you mean but... "permanently lost" sounds dramatic. When I trow away old pictures, old drawings or pieces made by my son at school, they are lost as well. Not that I do that often, but it begs the question, should every 'ip' made by humans be preserved?
It sounds dramatic because it’s dramatic. You’re combining two things — whether something is, in fact, completely lost, and if it is something that should be kept. Something that should not be kept is still completely lost if it’s destroyed. It’s difficult to imagine they’d digitize it if it was worthless.
>all of these things are permanently lost
A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models).
That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home.
>They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine.
My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement.
>That’s a false dichotomy.
I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it.
> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. Or be useful and scan book for Anna archive and other shadow libraries.
Citizen's lobbying against megacorp is a mirage.
>But it's _closer_ to being widely available, not farther.
By what metric ? The copy is now guarded by a company instead of being on the second hand market.
> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed.
I don't if that is true: a lot of old books might still have copy or other rights associated to them, likely owned by author and/or publisher, directly or inherited, but often those who have the rights do not have digital or physical copies at hand anymore (some old books are, well, really old). Does Anthropic make sure to track down, contact and then share the digital copy they make with those who have rights on the work? If not, they are not making it in any way easier to re-print the books, while making their supply more scarce (they destroy existing embodiments).
> Is it "locked"? Yes, by copyright laws,
> Plus having the info part of a LLM makes it immediately available to literally billions.
Isn't this a contradiction? I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".
> permanently locking human knowledge inside private corporate servers
History tells us that very few "permanent" situations are truly permanent.
Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.
> History tells us that very few "permanent" situations are truly permanent.
If you destroy the only copy of a physical artifact, the situation is as permanent as it can get.
I mean books get destroyed all the time, really they are a major pain in the ass to keep together, especially as they age. Paper loves to crumble. Insects think they are tasty. Floods and fire love destroying them too.
So physical books are rather non-permanent themselves.
I wonder if Anthropic could rent or sell access to their collection to the internet archive? It would probably be a good PR move (which they probably need right now), but I'm not sure what type of legal shenanigans they would need to do in order to not violate copyright law.
> any important book has been duplicated by thousands, tens of thousands or even million of units.
Books in the former USSR display their print runs on the last page. "Important book" is a vague and arbitrary term, but rhere are works in whole fields (e.g. history, archaeology, linguistics, ethography) that any scholar would consider key references, and as few as 100 copies were printed.
The shadow libraries have made a lot available to the whole world. It would suck if private corporations scan and shred remaining copies of these before the shadow libraries can get a scan.
I think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands.
Do you think that you are somehow special and unique in needing these books? If the books cost in the 10s of thousands, then there's obviously value to other people. I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. All evidence I've seen is that they're scanning cheap books with no current value and no clear use to people today.
I have volunteered with a library, and probably threw hundreds of books over a couple week engagement from a university library into a shredder at the direction of a professional, academic librarian. Libraries are constantly culling books, the EXACT category books we're talking about here (old, never-read). This is happening at a much larger scale, so I would recommend railing against university librarians in addition to the AI juggernauts.
I have books, but they are just objects. They're nice objects, but just objects.
Fetishizing books isn't going to help.
In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful.
Exactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book.
Whenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems.
Somehow I don't think they're looking for the books that have been copied over and over.
Why do you think that? Do you have evidence, or is this just a hunch? Why would Amazon waste money on uber rare books when there are thousands and thousands of not-so-rare books that could serve the exact same purpose?
For one they've already used the entire library genesis. Anything not in there is going to be obscure in some capacity.
> any important book has been duplicated by thousands, tens of thousands or even million of units.
It is common for academic books to have publication runs in the low three digits.
You may argue these books are not important. But how do we know if we fail to preserve it?
Is there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning?
I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.
Beware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM.
Honestly, it’s easier to find a good movie than a good book, because books are way cheaper to publish. These days, the quality of pretty much all kinds of content has become a problem...
First of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans.
Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway.
Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why?
[0] https://www.bodleian.ox.ac.uk/services/research-partnerships...
Why wonder? The answer is abundantly clear if you follow the news. If a book is copyrighted under U.S. law, scanning and destroying counts as a format conversion which qualifies it as fair use, so there is no need to negotiate with copyright holders. See Judge William Alsup’s decision. If Anthropic did not destroy the books after scanning it would have not won the lawsuit, and scanning would be illegal. If a book is already out of copyright then of course they do not have to destroy it afterwards.
Why is it a big deal?
Do you think there's no difference between the books curated at Oxford University and the crap that Amazon is buying?
Why would amazon buy "crap"? Surely they want their model to succeed and they want to train it on valuable input, no? They have more than enough resources to determine whether or not the books are worth buying. They have been in book selling for a long time.
Read more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies.
I was also skeptical of this claim but managed to find someone who explains it:
https://downtownbrown.substack.com/p/five-fallacies-ai-and-d...
It's not that individual companies buy all copies of a given book, but that there's more than one book scanning company, and they aren't sharing the scans with each other. The result: books that were rare but nevertheless easy to find for purchase (thanks to the internet) are now vanishing off of the market, becoming de facto no longer accessible to the public.
Good article, but I still feel unsatisfied because even it cannot find an example of a book that’s actually been lost because of the destructive scanning frenzy (it only lists books that hypothetically could be lost because there’s not many physical copies available for sale online.).
If anyone has an example, I’d love to hear it.
Since we don't know what was actually purchased and what was actually destroyed, how do you expect us to furnish an example? This would require the destroyers to admit it, and beyond that it would require all of them to admit it since more than one of them might have been responsible for the extinction of one text. Seeing as they were already keeping this operation under wraps, I don't see that happening. "possibly extinct because no copies available online" is probably the best we can do.
The distributed nature of the problem and the utter lack of transparency are huge factors here too.
Here we have another entry in the long list of "things described on the Internet that never happened".
Most old books that are rare and unpreserved are so because their value is marginal, so nobody has bothered to collect and preserve them.
But where did you hear that they’re buying “all copies”? And to what end?
This is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable.
To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved.
Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
>Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.
OK... I'm going to assume good faith even though your wording makes it somewhat unlikely.
Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be.
Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera.
Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago.
I think there are two issues here:
1. If someone is acquiring books in bulk for bargain-bin prices and shredding them, they're books whose physical copies have essentially no market value and which, absent this buyer, were overwhelmingly headed for pulping or landfill anyway. Millions of books are destroyed every day.
Could one of these worthless looking books turn out to contain information historians care about in a 100 year? Sure. But that doesn't create an obligation for someone to pay to warehouse every extant copy forever. Physical Preservation has costs: space, cataloguing, handling, transportation etc. Archives and libraries have always had to make choices for this reason.
2. I'm not arguing that preservation has no value. The question is whether destroying a physical copy after digitizing it is a serious loss when talking about mass-produced printed material.
Was this the last surviving copy ? Is the information unavavilable in libraries, archives, other editions, scans, citations, contemporary works etc ? If not, nothing has been lost except one physical instance of a reproducible object.
Your film analogy worked a lot better because old film footage is often unique primary source material. A camera recording of a random brooklyn street in 1993 may literally be the only recording of those people, storefronts and circumstances. The nth printe dcopy of a technical manual is not analogous to that.
How would film be any different in any capacity whatsoever?
Believe it or not, what matters here is the message and access to the message, not the medium.
Where did you see they tend to buy all the copies? This comment is the first time I've heard of this.
I'd also be curious about the provenance of that statement. Why would they buy all copies? What would be the purpose of scanning multiple copies?
I could see buying multiple copies being useful to mitigate problems, like damage.
But buying all the copies is categorically different and I cannot imagine why they would attempt that, except perhaps to prevent their competition from getting a hold of the same text?
It’s just another lie of the type these threads tend to be filled with nowadays.
Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with.
I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation?
It really does seem like an organized operation, given that it's so prevalent and they all seem to be in lockstep with their specious claims.
I wouldn't be surprised to find out that this outrage is fueled by the same actors as the AI data center water outrage. Whoever they are
I think these type of lies usually come about as a result of a game of broken telephone and things get exaggerated
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.
The hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.
But that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book.
Dumpsters also don't typically come equipped with a robot scanner and network uplink built in.
Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.
That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
The Internet Archive tries to be that magical library, but they can only scan and physically archive what is sent to them.
> I really don't know what people objecting to this imagine typically happens to old, unwanted books.
In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.
Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.
So trashed, but with a bunch of extra steps then?
I would presume that the extra steps incrementally diminish the possibility of the book remaining unsold, and thus destroyed or sent elsewhere.