Rendered at 15:23:23 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
num42 24 hours ago [-]
Long live Anna’s Archive. I stand on the shoulders of the Internet, Wikipedia, Anna’s Archive, Z-Library, LibGen, YouTube, Hacker News, Reddit, and Sci-Hub.
I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference.
gruez 22 hours ago [-]
>Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it.
This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.
>and continue doing it
Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.
bix6 22 hours ago [-]
> From the start, Anthropic “ha[d] many places from which” it could have purchased
books, but it preferred to steal them to avoid “legal/practice/business slog,” as cofounder and
chief executive officer Dario Amodei put it (see Opp. Exh. 27).
They stopped because they already had everything in their model anyway.
austhrow743 12 hours ago [-]
Yes there are people who are both for and against copyright.
sam_lowry_ 22 hours ago [-]
[flagged]
yonran 22 hours ago [-]
> This is an unethical as a company may behave, short of killing people.
This is hysterical. No, format shifting old unwanted books is not unethical. The books still exist, in an internal digital library. If copyright law were to change to allow sharing orphan works some day, Anthropic could share them. But under current law, the books are preserved digitally and used for transformative uses that all Claude users benefit from.
empiricus 22 hours ago [-]
did not follow all the details, but my understanding is that some form of copyright law nudges in the direction of destroy after scan?
hintymad 19 hours ago [-]
> they can afford the penalties and continue doing it.
I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag.
Gud 18 hours ago [-]
That's what they're doing now, when they are established.
But when it was a proof of concept, they were using pirated data.
Just like Spotify did.
vaylian 22 hours ago [-]
> Just like they want it to be illegal to run local ML inference.
I see. The following two statements (headlines) stand out:
* We should not sell powerful chips or chipmaking equipment to China
* We should crack down on industrial-scale distillation operations
That's not quite a ban of local ML inference, but it basically says that he doesn't want companies in China to create their own state of the art models.
peri-cl 5 hours ago [-]
No; it's the other headline, mandatory US government certification, that would ban Americans from running their own Chinese large models.
bookofjoe 19 hours ago [-]
Over the past decade I've noticed on HN the following order of frequency in choice of words, most common to least:
1. Citation
2. Source
3. Reference
Long ago in a career based on original research, I/we ONLY used "reference."
vaylian 9 hours ago [-]
I could have written more words, but they would have not conferred more meaning.
tygon 19 hours ago [-]
While it is definitely over a decade at this point (over two in fact), some of this likely comes from the term [citation needed], that originated on Wikipedia, as a cynical backhanded response to unsourced claims. It has become a catch-all. Language and how it evolves is a pretty interesting subject.
bookofjoe 18 hours ago [-]
You're right. Has to be of Wikipedia use origin. Thanks!
vaylian 9 hours ago [-]
OP here. You came to the right conclusion. This is inspired by Wikipedia.
spwa4 23 hours ago [-]
In theory none of them actually got the right to train on illegally downloaded books. Anthropic was simply punished for doing it once.
One wonders if they're still doing it.
ungut 23 hours ago [-]
OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models.
There is just not enough non-copyrighted data out there.
yorwba 23 hours ago [-]
You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models
without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.
tygon 19 hours ago [-]
I wonder if anyone has run the numbers on what the actual cost, both in cash and logistical headache, contacting so many copyright holders would be. That seems like quite the feat to calculate.
yorwba 19 hours ago [-]
There's an established network of intermediaries that can supply a large variety of books for a few dollars apiece, so no need to contact copyright holders directly.
tygon 18 hours ago [-]
This is very true. As someone with quite the experience with materials published under Penguin, Scholastic, etc. you effectively have a "dictionary attack" on the matter, rather than true "brute force," but that still leaves quite a list to compile to send to each and is easier for larger titles than smaller ones. I wonder how that leads to a bias in what materials get used for training. You are not getting many local self-published books this way.
It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
Doesn't really matter. The incentive structure to steal clearly exists, so why would they even go through the trouble?
ungut 16 hours ago [-]
Pretty easy to assertain that they don't acquire them legally due to the plethora of evidence and court cases against them. No copyright holder would be sueing them if they knew they sold the works in the first place.
I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.
yorwba 10 hours ago [-]
I'm not aware of any successful accusations against OpenAI for illegally obtaining copyrighted material, in contrast to Anthropic, who settled for $3000 per work and then still had to buy legal copies to keep using them (likely for much less).
Instead, the ongoing lawsuits focus on the idea that AI training involves making additional copies, for which they would need a copyright license instead of just one legal copy.
ungut 8 hours ago [-]
You are kind of right, but you also did not look very hard. They deleted huge datasets in anticipation of lawsuits, at least that much is known.
Of course plaintiffs were unable to depose their internal lawyers (who apparently know why they were frantically deleted) due to 'attourney client priviledge' further refusing to provide any kind of transparency. But yeah, I guess they are they are better at covering their tracks and destroying evidence.
Also, as one more example, I find it hard to believe that their models could generate 'Studio Ghibli' style images without training on the movies. There is no licensing deal between them.
I think the real issues here are two-fold:
Firstly, Copyright is very ill equipped to handle these cases. Just because the model is tuned not to output the exact training data does not mean that compressing mostly-copyrighted datasets into a proprietary model is ethical, fair or /should/ be allowed, simply because they might destroy entire livelihoods. If you take those copyrighted works away you are left with, in OpenAIs own words, a cute little experiment.
Secondly, there is absolutely no transparency. Datasets are easily deleted and its impossible to tell what the models have been trained on, especially after fine tuning. Moreover, only the biggest most successfull works would be easily identifiable without the fine tuned model. Once again, sticking it to the little man.
spwa4 22 hours ago [-]
I'm pretty sure we would know if they did that. And we don't.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
spwa4 19 hours ago [-]
But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course.
You mean very likely the Anna's archive torrent dump because it's MUCH better quality than the general internet and beyond a certain amount of input data (which is a lot, but much less than the internet) the only thing that matters in training is the quality of the data, to the point that now many labs have thousands of people just making and improving essentially school exercises full time?
Hell, I know that for one "lab" (kindof AI lab) since 2020 or so has determined wikipedia quality is dropping fast. It was already dropping slowly before that, but now it's getting bad.
yorwba 6 hours ago [-]
No, I mean the WebText corpus whose construction from 45 million Reddit post with at least 3 karma is described in section 2.1 of the PDF I linked. They did remove all Wikipedia documents.
Anna's Archive didn't exist in 2019.
deadbunny 22 hours ago [-]
What tosh. It's copywrited material, they pay to access it like everyone else.
IncreasePosts 19 hours ago [-]
I thought the outcome of that was basically it's legal to train on books, but they acquired the books in the wrong way. If they went out and bought copies of them and trained it would have been fine
outside1234 23 hours ago [-]
Of course they are. They have just put on their Swiss Banker suit now and have all sorts of deflection techniques in place such that, of course, "the money has the stamps that says its clean" (when it it really blood money hidden behind a pretty wall).
toomuchtodo 24 hours ago [-]
No gain, all liability. Easier to cut them a check for access to training data and say nothing. Unless legal discovery was performed, the outside world would never know, and the payment records would roll off corporate records through a record retention schedule eventually. Could obfuscate it as a contractor consulting fee ("knowledge management subject matter expert") if you wanted to get tricky, depending on the risk appetite of whomever would receive the funds.
(not legal advice!)
p-e-w 24 hours ago [-]
There’s zero liability in a company stating publicly that they support Anna’s Archive. Zero. Free speech protections cover much more egregious statements than that.
pibaker 18 hours ago [-]
Your freedom of speech is your opponents' lawyers' wet dream. Your publicized support for a known piracy operation will not look very good in the court when you get sued by copyright holders.
toomuchtodo 24 hours ago [-]
I disagree. Anyone with even a hint of standing will sue, and keep suing. As someone who has to work with corporate counsel often, do not say anything you don't have to say. Only say what is absolutely necessary. Free speech protects you from your government. It does not shield you from civil suits, and the US is extremely litigious.
criddell 23 hours ago [-]
Maybe they are worried about claims of contributory infringement?
TZubiri 18 hours ago [-]
Those are all law-abiding organizations, which AA is not.
INB4: "Here is one time one of those organizations broke the law". Don't go there, absolute lowest level of conversation.
kmeisthax 20 hours ago [-]
I'm pretty sure[0] they're all using shadow libraries, and saying things in favor of them would increase their liability.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
mips_avatar 21 hours ago [-]
I wouldn’t include YouTube given how only Google is allowed to index it
xtracto 23 hours ago [-]
You missed Gigapedia (library.nu [1]) , which preceded most of the others.
Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
> Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
tentacleuno 21 hours ago [-]
Whilst it pays for the service (and, in that respect, may be a necessary evil), it's morally questionable (at minimum) to charge for things that, by law, aren't yours in the first place.
dessimus 21 hours ago [-]
Less morally questionable than claiming to users they are "buying" access to media that can be revoked at any point in the future with no recompense, of course referring to Sony and Amazon.
insane_dreamer 10 hours ago [-]
they're not charging for the content -- they're charging for the extra bandwidth if you want faster access
orbital-decay 21 hours ago [-]
BBSes were the first, of course. In particular, Libgen, Sci-Hub and others can be traced back through several generations of libraries to the SU.BOOKS FidoNet echo conference created in the early 90's.
wolvoleo 15 hours ago [-]
Yes and I remember a thriving community of IRC DCC file servers. I used them back in the day to read books on my palm pilot before wink existed.
dessimus 20 hours ago [-]
> Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.
Synaesthesia 1 days ago [-]
Incredible website. Really one of the dreams of the internet realised, the ability to access all knowledge at your fingertips.
simlevesque 24 hours ago [-]
I was blown away when I paid to get an API key for it. The process is intricate and state of the art.
linksbro 22 hours ago [-]
What's the process like? I wasn't aware there was a paid API key.
simlevesque 21 hours ago [-]
You buy an Amazon gift card as a gift and use the email Anna's Archive provide you as the recipient. The email is unique to you so they know when you paid them.
Yokolos 24 hours ago [-]
Probably the greatest achievement of modern society, a rebirth of the Library of Alexandria. Of course, only private for-business corporations are allowed to steal the world's knowledge, apparently, and only to be able to monetise it. There's something deeply wrong with our civilization that Anna's Archive is punished while companies like OpenAI, Google, Amazon, Anthropic, etc are just ... ignored when they do things like destructively digitise books or pirate things.
22 hours ago [-]
fragmede 7 hours ago [-]
By "ignored" do you mean everyone on the Internet yelling about it all over all the platforms?
throwatdem12311 24 hours ago [-]
Didn’t they have a default judgment against them because they didn’t show up to court? Do they even know who runs it?
They railroaded Aaron Swartz (which eventually lead to his suicide) for much, much less.
One side wants to freely share knowledge with all of humanity, the other wants to restrict it to make a buck. I know who I support.
puppycodes 24 hours ago [-]
Anna's Archive is a gift to humanity
weaksauce 21 hours ago [-]
Exactly. if AI gets a pass on stealing all of the world's information and gets a pass why shouldn't we get to enjoy the same benefits?
nekusar 24 hours ago [-]
Exactly. All libraries are worthy of being supported and grown. I make no distinction between a physical dead-tree library and a digital library.
The only reason we even have dead-tree libraries at all is because 100 years ago, that was what John Rockefeller and Andrew Carnegie put forth to whitewash their horrible capitalist behaviors across the USA. And because it was done by those generations' billionaires, public libraries because acceptable.
If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
MYEUHD 24 hours ago [-]
> If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
I wonder why nobody created a Netflix-like subscription for digital books
Ekaros 24 hours ago [-]
Amazon maybe? Kindle Unlimited. In some ways Audible with audiobooks. Which actually has lot of entrants though they tend to limit listening time on large part of their portfolio.
arrosenberg 24 hours ago [-]
Libby very much exists, you just need to sign up through your local library.
devilbunny 18 hours ago [-]
The problem with Libby is finding a local library that has a good selection. Mine does not, and my state does not have the feature of being able to get a library card from any library in the state. I would have to buy a membership to a non-local library, which can be surprisingly difficult to do if what you care about is a large selection of ebooks.
copper4eva 24 hours ago [-]
I believe you can get access to a lot of digital books with a library card online. Not something I have personally done, so can't really give details.
The reason I don't subscribe is it takes me a lot longer to finish a book/audiobook than a movie. I just get them from the library/Libby - even if it's a 4 month wait, there are plenty of available books to read while I wait.
There was already one that existed in the web 2.0 era. Founded in 2012, raised a $3M seed from Founders Fund and a $14M A. Completely failed though and was acquihired by Google to have the founders lead Google Play Books.
Amazon now has Kindle Unlimited but I think it's mostly romance slop and self-published books. Seems like publishing rightsholders were just too inflexible to let the business model take off.
TitaRusell 21 hours ago [-]
Yeah I like reading somewhat obscure history books- no public library carries those and they're extremely expensive if you can find them.
Usually funded by universities and foundations anyway because there's no real commercial market.
Mezzie 23 hours ago [-]
Because you have to pay a ton more for a lot of digital copies of books if you're doing any kind of lending and you're competing with existing libraries. So you both have higher costs (e.g. you have to pay more for a digital file that may have DRM, be required to self destruct after a certain number of uses, doesn't allow multiple lends at a time, etc.) and your prospective consumers have a free option so convincing them to pay is going to be difficult.
The economics aren't there.
Ajedi32 20 hours ago [-]
Public libraries lend DVDs and Blu-Rays yet Netflix still exists.
Mezzie 17 hours ago [-]
Netflix has a larger market base (far more people watch TV and movies than read books, particularly at a level high enough to justify a subscription) and a market base that is far less likely to already be aware of the free alternative (most serious readers of the sort that go through that amount of books are already aware of their local libraries, whereas serious movie/TV watchers are less likely to be so).
Additionally, modern Netflix is a streaming company, and even DVD Netflix was sending the materials directly to your house, which is a differentiator versus the library, whose materials need to be picked up. There's a convenience factor as an incentive to pay. Library streaming exists, but it's awful - very limited library and very limited watches - versus Netflix where once you sub you can watch as much as you want.
hydrogen7800 24 hours ago [-]
I consider libraries, public schools and maybe even public fire departments in the list of things that could never be proposed today if they didn't already exist.
te_chris 22 hours ago [-]
Public libraries pay for their books
dominick-cc 19 hours ago [-]
I don't understand why there is so much love for Anna's archive here. I feel like I'm getting whiplash because there is so much hate for AI companies training and profiting off of the worlds knowledge without licensing it. And yet, a site that directly facilitates that by taking payments from AI companies is lauded as this amazing and honorable thing. Can someone help me understand what im missing?
In general I don't have many qualms with modern piracy, it just seems very hypocritical and I'm confused.
pibaker 19 hours ago [-]
AI companies turn public data into closed commercial products. They are allowed to profit off your copyrighted data, but only they get to profit from their own models.
I'd imagine the backlash to AI companies from us white collar workers to be less severe if they have to publish their weights. In fact if you look closer you will see HN is actually pretty content with Chinese open models. It's the American AI corps with closed models that attract criticism.
keeda 16 hours ago [-]
> Can someone help me understand what im missing?
One directly threatens the livelihoods of the commenters here, the other doesn't ;-)
frgturpwd 4 hours ago [-]
I don't hate either particularly, but I care more about getting any book I want whenever, wherever, than I care about some abstract far away moralizing of something of consequence to me only ambiguously.
shaky-carrousel 24 hours ago [-]
And Google owes 20 decillion dollars to the Russian government. Who cares either way.
srean 22 hours ago [-]
I only wish they allowed browsing journals by year and volume. Libgen allowed that but war in Ukraine broken libgen. It lives but as a much shadier alter egos that does not support all that original libgen supported.
NooneAtAll3 24 hours ago [-]
"owes" seems like a wrong word here
profstasiak 20 hours ago [-]
Rembember to seed torrent kids
23 hours ago [-]
cloudie78 23 hours ago [-]
So when are Nvidia, Facebook, and everyone else going to chip in and pay up?
outside1234 23 hours ago [-]
Did you miss that they are huge corporations that the law doesn't apply to?
kingleopold 23 hours ago [-]
if this is true. what does this say about "rule of law"? is it all fiction?
gruez 22 hours ago [-]
You mean how anthropic lost a court case and now they're being made to pay?
22 hours ago [-]
kingleopold 21 hours ago [-]
they are paying from the revenue, they gained ton of valuation and revenue because of books? if you do that, you lose most of your life
gruez 20 hours ago [-]
Even in the infamous Aaron Swartz case he was offered 6 months in prison as a plea bargain. Even ignoring the difference in crime (torrenting vs CFAA), no one was going to "lose most of your life".
nemomarx 22 hours ago [-]
Like all things with society, it exists so long as we constantly create it. If you slack off many things will vanish faster than you think.
bethekidyouwant 21 hours ago [-]
I don’t get it. I am not a corp. i can download books from AA (who does the law apply to exactly?)
crimsdings 5 hours ago [-]
Have they released the Spotify scrape already?
Frieren 24 hours ago [-]
Anna's Archive is what TV told me in the 2000s that the future was. An online database with all published books one click away. Far from the dystopian reality than the corporate internet has become.
dcchambers 24 hours ago [-]
"If you own someone $340, that's your problem. If you owe someone $340,000,000, that's there problem."
Might as well be a trillion dollars as it's never getting paid.
qwertox 24 hours ago [-]
"when you're $100,000 in debt, it's your problem. But when you're $1 million in debt, it's the bank's" - Rosalie Goes Shopping (1989)
a4isms 23 hours ago [-]
An old variation:
Fred and Wilhelmina are in bed together. Fred tosses and turns, unable to sleep. "Fred, what's troubling you?"
"You know Bjarne from the bank? Well, a balloon payment is due tomorrow on my loan, and I don't have the cash flow to pay it."
Wilhelmina thinks for a bit, then reaches for her phone. "Betty? Yes, this is Wilhelmina. Sorry to call so late. Would you tell Bjarne that Fred can't make the loan payment tomorrow? Yes, that's all. Good night."
Fred states at Wilhelmina, aghast. "What did you do THAT for?" She smiles. "Now it's Bjarne's problem. Let him toss and turn, you can go to sleep."
ivberrOg 23 hours ago [-]
the flintstones?
a4isms 22 hours ago [-]
I chose the names from The Flintstones, although the joke is not from the show.
Ridiculously unimportant trivia: The Flintstones borrows heavily from The Honeymooners, was the first animated sitcom, and was also the first animated show aired in prime-time.
asdefghyk 19 hours ago [-]
I wonder what would happen if Annas archive announced
In response to the 340M fine, ....
"We have set up a LLM and are deleting all our books ? "
gitowiec 20 hours ago [-]
What are current domains? I tried the ones from article but got redirected to spam
wlonkly 15 hours ago [-]
Wikipedia usually has the most recent domains for sites like this. There's even a browser extension[1] that uses an `.idk` TLD to say "Look up the latest URL from this site's Wikipedia article".
I made a service specifically to translate Wikipedia entries to DNS with Anna's Archive as the main use case: https://whither.link/
barnabee 20 hours ago [-]
.gd works for me
sajithdilshan 24 hours ago [-]
Unfortunately it’s blocked in Germany
xiconfjs 23 hours ago [-]
only the local providers in Germany are forced to block it via DNS - 8.8.8.8 to the rescue
barnabee 20 hours ago [-]
Tor Browser is handy, if simply changing your DNS isn't enough
theideaofcoffee 21 hours ago [-]
Just make them owe $134 Trillion or whatever the evaluation was years ago suggested by the RIAA for their estimation of 'damages' for music piracy. It's about as meaningful.
coolThingsFirst 21 hours ago [-]
taking shots against the King?
better not miss!
expedition32 21 hours ago [-]
Thanks for reminding me I need to download some books for the holidays!
tokai 24 hours ago [-]
Just wished they didn't sell bulk access for LLM training.
nojvek 14 hours ago [-]
> U.S. courts can’t reach domains registered beyond their jurisdiction. That’s likely to increase calls for site-blocking legislation, a measure the industry has long favored and that remains high on the political agenda in the United States.
Eventually US will have its own great firewall like China.
Politics is about powergrab and what we’ve seen is more and more power grab.
It’s likely the billionaires are able to buy the govt goons to pass the laws that allow them to be the gatekeepers.
Anthropic, OpenAI and the model builders massively benefited from the archived information. Distill entirety of archived human knowledge.
nekusar 24 hours ago [-]
Might as well say "Annas Archive OWES ELEVENTY HUNDRED BILLIONTY-TRILLIONTY INFINITY DOLLARS!!!!!111" ala elementary school playground make-believe.
I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster. Or how Aaron Schwartz was executed by proxy by JSTOR and the feds, for what should have been free to access for all.
But hey, Anthropic, OpenAI, X, and others can pirate to their hearts content with for-profit piracy, but "we" (royal) are OK with that. We just cant have the poors have access to the sum of human knowledge.
mmooss 22 hours ago [-]
> I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster.
Why didn't Metallica's fans rebel? That would have stopped it quickly and set a precedent and example for the rest.
npongratz 21 hours ago [-]
Many of us did, and I certainly still do, but there obviously weren't enough of us rebelling to convince them of how reprehensible it is to sue your own fans and sic the guns of the state on them.
mmooss 17 hours ago [-]
> sue your own fans
What an incredible strategy - and it worked! And in something that's entertainment and fandom-driven; it's not like the fans need a medication, job, or their auto warranty. How pathetic. The day rock'n'roll died.
bethekidyouwant 21 hours ago [-]
[flagged]
brcmthrowaway 23 hours ago [-]
> elementary school playground make-believe
Oh yeah? Well, my dad works at Anna's Archive.
palmotea 24 hours ago [-]
[flagged]
tetris11 23 hours ago [-]
People can be predisposed towards something, and then have a tipping point brought by periods of great stress
Jason Arday recently committed suicide due to media pressure, and even when the media was told that he was not mentally well and should please back off, they didn't.
The man was accused of plagiarism, and there was no final verdict on if he did. The institutions that hired him clearly had no issues in vetting him. And yet he committed suicide.
There are absolutely horrible people out there right now who's entire careers are shams, and don't even come close to thinking about it. People are built different, and have different triggers
sejje 23 hours ago [-]
> The man was accused of plagiarism, and there was no final verdict on if he did.
Bullshit. It wasn't the first round of plagiarism, or redactions. His whole supposed life story was a fraud. He passed off other researchers' patient interviews as his own.
KennyBlanken 22 hours ago [-]
> accused of plagiarism
Weird way to spell "lied about nearly every facet of his life and anyone who questioned it was aggressively accused of being racist."
kenanfyi 24 hours ago [-]
Sure. That he was facing 35 years in prison for his “crimes“ had nothing to with his suicide.
KennyBlanken 22 hours ago [-]
Please educate yourself about how the US criminal justice system actually works, especially around plea deals. Also especially around sentencing for white collar crimes in terms of what white collar criminals are actually sentenced to, and what is typically taken into consideration by judges during sentencing phases (like, say, mental health issues.)
Please also educate yourself about how federal prosecutors treated Schwartz after he scraped PACER.
Please also educate yourself about how long prosecutors attempted to negotiate a plea deal with Schwartz, what their final offer was, what his family, friends, his girlfriend, and his team of top-shelf, famous lawyers advised him to do, and what he actually did.
nekusar 21 hours ago [-]
Please educate YOURSELF about criminal actions by a company is a fine (read: cost of doing business), whereas criminal actions by a person get years/decades long prison time and life-long ramifications after the actual punishment.
And PACER is federal court documents. Its a fucking scam and travesty that a system soo heavily built on precedent gatekeeps its own law at $.10/page.
Aaron is a motherfucking hero.
Catloafdev 24 hours ago [-]
If only you were able to take your perspective one step farther. Why did he commit suicide?
palmotea 21 hours ago [-]
> If only you were able to take your perspective one step farther. Why did he commit suicide?
If I get dumped and I commit suicide because of it, does that mean I was "executed by proxy" by my girlfriend/boyfriend?
"Executed by proxy" is ridiculous, overheated rhetoric.
sejje 23 hours ago [-]
Because something was wrong in his brain and he couldn't handle the stress of his situation.
Many people have faced similar or worse charges, and did not commit suicide, so we'll have to point at something else.
Fuck them for what they did to him, but we should be rallying to get him set free. We can't start pointing fingers at other people for suicides. We have agency.
NegativeLatency 24 hours ago [-]
Why did he commit suicide?
nekusar 24 hours ago [-]
"We are going to jail you for half or more of your life, restrict any medicines you might be on, treat you in deplorable conditions, constitute you as a slave of the state if we decide so". And there is no parole in federal prisons.
Feds did NOT have to choose to go the route they did. Their pursuance of a victimless "crime" lead to Aaron's only path was 'Exit'.
Sure, he committed suicide. Why? Because his life was already significantly threatened with what amounts to torture by prison.
sejje 23 hours ago [-]
JSTOR didn't want the criminal charges
rpdillon 22 hours ago [-]
But Carmen Ortiz did.
tokai 23 hours ago [-]
The feds are not JSTOR now are they? Your discourse is disgusting and only weakens the cause.
KennyBlanken 21 hours ago [-]
I started typing out something but they're foaming at the mouth and wildly ignorant about nearly everything ranging from the basics of how the US criminal trial system works, PACER fee policies, the fact that prosecutors spent a year negotiating with Schwartz over a plea deal with a final offer of 6 months...a deal he refused against the advice of friends, family, his partner,
Any time someone starts shrieking "half his life" you can assume they don't know a single factual thing about the case.
From the tone, ignorance, drama, and attitude I'm guessing they're about 22 years old.
plusfour 20 hours ago [-]
"owes"
surgical_fire 24 hours ago [-]
Reminder for me to make another donation.
Piracy is morally justified at this point.
voakbasda 24 hours ago [-]
I wonder by if the US considers this “supporting a terrorist organization”.
walrus01 24 hours ago [-]
Donate a significant amount of money, and you can get a free trip to the new gulag in El Salvador. Or maybe gitmo.
no_input 24 hours ago [-]
Don't give the government any ideas
voakbasda 24 hours ago [-]
Hardly a new idea: Trumplestiltskin suggested during this term that we should start detaining Americans who expressed their dissent against the government.
drstewart 23 hours ago [-]
I wonder if the EU or UK does.
Maybe if they use encryption, which would make them all pedophiles according to those governments. Hope you have your VPNs ready if visiting countries without a free internet, like North Korea, China, UK, or EU.
surgical_fire 23 hours ago [-]
Maybe.
Thankfully I don't live in the US, and likely won't set foot there ever again.
drstewart 23 hours ago [-]
[flagged]
surgical_fire 20 hours ago [-]
My labor allowed me to buy my house.
If anything, I was heavily underpaid.
drstewart 18 hours ago [-]
Nope, it was FAANG, as you admitted.
You're underpaid now at your euro job, that's a guarantee.
surgical_fire 17 hours ago [-]
I am where I want to be.
I also have no desire to return.
drstewart 12 hours ago [-]
No one wants you to
21 hours ago [-]
toomuchtodo 24 hours ago [-]
US probably considers anything impairing profits, stock prices, or enterprise value terrorist activity at this point.
Insimwytim 19 hours ago [-]
The term "piracy" is used by record companies to demonize sharing and cooperation by equating them to kidnaping, murder and theft.[1]
Yeah, truthfully. This type of article is generally a nice nudge to donate.
T3RMINATED 24 hours ago [-]
[dead]
cyanydeez 21 hours ago [-]
it "Owes" just like Anthropic and OpenAI "Own" their models trained on the worlds collective IP.
iLoveOncall 24 hours ago [-]
I really don't understand why HackerNews gets hard for Anna's Archive but you never see other illegal download websites praised here.
They're not more or less virtuous than any others, they're in it for the money and you're a fool if you think otherwise.
I'm not against piracy, it's great, but to claim they are moral saints is a joke.
Anyway, bring in the downvotes as I know will happen.
nemomarx 24 hours ago [-]
how's Anna make money? I know they offered to share the database for ai training (which I would criticize, that's very shady) but do they have donations or something
it doesn't seem like it would be high revenue at least
iLoveOncall 24 hours ago [-]
The same as most private Torrent tracker: via premium accounts masquerading as donations.
Anna's Archive is very limited in performance if you don't pay.
jbaiter 24 hours ago [-]
That's not how most established private torrent trackers operate? Account privileges are usually granted according to tenure, number of uploads and ratio requirements.
KennyBlanken 23 hours ago [-]
Not that the other guy is right, but most established private torrent trackers absolutely allow you to bypass requirements if you "donate."
HDThoreaun 23 hours ago [-]
> Anna's Archive is very limited in performance if you don't pay.
Maybe if youre trying to pull an AI lab and download their full corpus, although I suspect thats not too onerous either. I have never had any issues getting stuff from annas archive for free
24 hours ago [-]
KennyBlanken 24 hours ago [-]
Anyone paying for Annas Archive is a moron. It takes more time to find the right book/file than it does for the download, which is usually under ~5MB, to finish.
branon 23 hours ago [-]
I doubt most dontators think of it as "paying for Anna's Archive" - I pay to contribute to the institution more than for the minor convenience of the fast download links. I'd still donate even if I got nothing extra in return.
It also helps that AA has the best UI and search of any shadow library I've seen. Some really established bittorrent sites win out on curation but otherwise AA is top tier, not something I'm accustomed to seeing from a (free, no-signup!) clearnet/direct-download site
As far as GP's question I think you can be in it for the money (though I personally have my doubts that Anna is) and also be a force for moral good as well. Whether the operator's morals perfectly align with all of our ethics is an open question but the room at large generally tends to agree.
Also Torrentfreak's reporting has been extremely solid for an extremely long time so the link itself fits here too
wolvoleo 15 hours ago [-]
I've seen support here for the pirate party and other things like pirate bay.
barnabee 20 hours ago [-]
I praise all pirates and leakers
Maybe you can own (some) things, you can't own information
animuchan 23 hours ago [-]
Can't talk for all of HN, but I for one hereby praise most/all piracy websites. Anna's Archive is great, and Rutracker is also great.
It's not a competition, and being a moral saint isn't a good KPI; being a net positive for society is.
the_real_cher 24 hours ago [-]
As a hacker News viewer I'm for other illegal download sites as well.
If you just Google 'free media heck yeah' you'll see a great deal of them!
iLoveOncall 24 hours ago [-]
Yeah me too, but I don't claim that they're a moral good.
beej71 24 hours ago [-]
It depends on what you consider to be a moral good. Some people feel that the world's information is getting locked down in the name of profit to the detriment of humankind. AA is providing knowledge to people who would otherwise not be able to get it. From that viewpoint, it's immoral to try to shut AA down.
OTOH, if you think that AA is robbing creators and publishers of their hard-earned proceeds, and that loss is greater than the loss to humanity from the destruction of AA, then you wouldn't think AA was a moral good.
Of course, it's all gray in the middle.
iLoveOncall 21 hours ago [-]
I personally don't care about the morality, I'm just pointing out an absurd double standard.
You cannot claim that AI labs are bad for illegally downloading content to build their AI without paying creators, and that AA is good for illegally offering the same content.
I think both are immoral, but I definitely pirate all the content I can (except video games but only because it's unsafe, not to remunerate creators).
It is not grey at all, this is some bullshit that people tell themselves to feel better. Whether it's using the output of AI models or downloading a book on Anna's Archive, you are ultimately robbing the creator of profits.
Nobody on HackerNews would argue otherwise if their employer stole their code from their mind and didn't pay for it.
beej71 55 minutes ago [-]
> Whether it's using the output of AI models or downloading a book on Anna's Archive, you are ultimately robbing the creator of profits.
You're correct in the above, or at least I'm willing to accept the premise for the purposes of this argument.
However, it's not that simple when it comes to the morality of the actions since there are secondary effects.
You can't just make a blanket statement like "stealing is immoral", for example, when there are circumstances where stealing is absolutely the moral imperative (e.g. the food is held by barons who charge unaffordable prices and the populace is starving to death).
Or, to make an example closer to the topic, what about the stealing of scientific papers? There's a strong arguments to be made that millions of people have been helped through that theft. And yes, the gatekeepers of that knowledge have been harmed.
And it's like that with AA and AI (and their code) as well. People consider the secondary effects and whether or not the greatest good is served.
And, of course, "greater good" is gray.
ndriscoll 21 hours ago [-]
I agree that it's not grey at all, but because they're both so obviously good (particularly labs releasing open weight models). It's great that there are organizations making it ~free for anyone in the world to obtain information instantly. On the contrary, our policy of restricting something that naturally can be duplicated an unlimited number of times for free is obviously morally bad. People should be incentivized to create new ideas, not rent old ones for a century.
Of course on a related note, most work that's interesting to me was created by people who are dead anyway, so they don't mind.
Anyway, are you sure most people that support AA don't also support e.g. Deepseek and Qwen's efforts?
goolz 24 hours ago [-]
They provide the world with an invaluable service at the expense of their freedom. Seems incredibly noble and good to me. Especially considering the current climate.
iLoveOncall 21 hours ago [-]
They're a for profit organization and you're falling for their propaganda.
branon 21 hours ago [-]
Propaganda to what end? Are you referring to anything in particular?
Is your angle merely that shadow library hosts should be operating their services entirely for free (to the point of refusing payment) or is there something we're missing here?
iLoveOncall 20 hours ago [-]
No, just that they shouldn't claim that they have a vertous mission and that this is what your money is going towards.
barnabee 20 hours ago [-]
Freeing all information is virtuous, though…
zen928 17 hours ago [-]
Do you need to follow the law to have virtue? There's a correct answer to that question btw.
animuchan 23 hours ago [-]
We don't have to have the same exact moral systems of course, but this is curious: if we (collectively!) support the thing, why not claim it's good?
Auracle 23 hours ago [-]
I think the reason why it shouldn't be claimed as a moral good is pretty obvious - authors deserve to get paid for their work. I certainly can't claim I stand on a bedrock of morality as I've used it from time to time, but even though I don't have the money that a lot of HN commenters do I make sure to not use it for small-time authors. Even really successful authors I'll only use it about 50% of the time.
If you want information to be free? Great. But most of those authors wouldn't be putting the work in to making that information/literature in the first place if they know they aren't going to get paid.
rpdillon 21 hours ago [-]
> But most of those authors wouldn't be putting the work in to making that information/literature in the first place if they know they aren't going to get paid.
My understanding is that the number of authors that can make a full-time job of writing is a rounding error compared to the population of authors. I believe that people should be paid for their work, but I don't think the current copyright regime is actually very effective at paying authors for their work. So I value copyright enforcement proportionately less based on that observation.
Meanwhile, on the flip side of the coin, copyright rulings are causing companies like Atheropic to destroy books as they scan them, creating a rising sense of panic around knowledge scarcity. This panic is a direct result of the scarcity of the copyright intended to create in the first place. It's entirely artificial.
I've been saying this for 25 years, but copyright is essentially broken.
barnabee 20 hours ago [-]
> authors deserve to get paid for their work
I don't really agree
Information itself should be free, authors deserve the right to monetise other ways (merch, physical media, exhibitions/shows, etc.) but society should pay artists and creators to do their thing[0]
Not everything needs to be a business, and I think art is one of those things
[0] I don't know much about it but maybe Ireland's Basic Income for the Arts scheme is a model. I think some other European countries do similar-ish things, too
benj111 22 hours ago [-]
Personally I would say copyright should last somewhere less than 10 years. On the basis that business decisions about whether to fund a book/film/whatever today isn't based on earnings 10 years from now.
So I would say it is morally acceptable to download stuff older than that. Torrents offer brand new things, so I wouldn't say they are morally good. I don't think they're morally worse than those locking up old media though.
24 hours ago [-]
asdefghyk 19 hours ago [-]
when I read title .
"Annas archive owes $340 million absolutely first thing that came to mind is
" How much do all these Ai companies owe for their unauthorised use of books and other media?"
I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
Coming out would be a bold move for them.
This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.
>and continue doing it
Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.
https://cdn.arstechnica.net/wp-content/uploads/2025/06/Bartz...
Sounds pretty pro piracy to me.
This is hysterical. No, format shifting old unwanted books is not unethical. The books still exist, in an internal digital library. If copyright law were to change to allow sharing orphan works some day, Anthropic could share them. But under current law, the books are preserved digitally and used for transformative uses that all Claude users benefit from.
I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag.
But when it was a proof of concept, they were using pirated data.
Just like Spotify did.
Citation?
* We should not sell powerful chips or chipmaking equipment to China
* We should crack down on industrial-scale distillation operations
That's not quite a ban of local ML inference, but it basically says that he doesn't want companies in China to create their own state of the art models.
1. Citation
2. Source
3. Reference
Long ago in a career based on original research, I/we ONLY used "reference."
One wonders if they're still doing it.
It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.
Instead, the ongoing lawsuits focus on the idea that AI training involves making additional copies, for which they would need a copyright license instead of just one legal copy.
Also, as one more example, I find it hard to believe that their models could generate 'Studio Ghibli' style images without training on the movies. There is no licensing deal between them.
I think the real issues here are two-fold:
Firstly, Copyright is very ill equipped to handle these cases. Just because the model is tuned not to output the exact training data does not mean that compressing mostly-copyrighted datasets into a proprietary model is ethical, fair or /should/ be allowed, simply because they might destroy entire livelihoods. If you take those copyrighted works away you are left with, in OpenAIs own words, a cute little experiment.
Secondly, there is absolutely no transparency. Datasets are easily deleted and its impossible to tell what the models have been trained on, especially after fine tuning. Moreover, only the biggest most successfull works would be easily identifiable without the fine tuned model. Once again, sticking it to the little man.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
Hell, I know that for one "lab" (kindof AI lab) since 2020 or so has determined wikipedia quality is dropping fast. It was already dropping slowly before that, but now it's getting bad.
Anna's Archive didn't exist in 2019.
(not legal advice!)
INB4: "Here is one time one of those organizations broke the law". Don't go there, absolute lowest level of conversation.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
[1] https://en.wikipedia.org/wiki/Library.nu
This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.
They railroaded Aaron Swartz (which eventually lead to his suicide) for much, much less.
One side wants to freely share knowledge with all of humanity, the other wants to restrict it to make a buck. I know who I support.
The only reason we even have dead-tree libraries at all is because 100 years ago, that was what John Rockefeller and Andrew Carnegie put forth to whitewash their horrible capitalist behaviors across the USA. And because it was done by those generations' billionaires, public libraries because acceptable.
If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
I wonder why nobody created a Netflix-like subscription for digital books
The reason I don't subscribe is it takes me a lot longer to finish a book/audiobook than a movie. I just get them from the library/Libby - even if it's a 4 month wait, there are plenty of available books to read while I wait.
There was already one that existed in the web 2.0 era. Founded in 2012, raised a $3M seed from Founders Fund and a $14M A. Completely failed though and was acquihired by Google to have the founders lead Google Play Books.
Amazon now has Kindle Unlimited but I think it's mostly romance slop and self-published books. Seems like publishing rightsholders were just too inflexible to let the business model take off.
Usually funded by universities and foundations anyway because there's no real commercial market.
The economics aren't there.
Additionally, modern Netflix is a streaming company, and even DVD Netflix was sending the materials directly to your house, which is a differentiator versus the library, whose materials need to be picked up. There's a convenience factor as an incentive to pay. Library streaming exists, but it's awful - very limited library and very limited watches - versus Netflix where once you sub you can watch as much as you want.
In general I don't have many qualms with modern piracy, it just seems very hypocritical and I'm confused.
I'd imagine the backlash to AI companies from us white collar workers to be less severe if they have to publish their weights. In fact if you look closer you will see HN is actually pretty content with Chinese open models. It's the American AI corps with closed models that attract criticism.
One directly threatens the livelihoods of the commenters here, the other doesn't ;-)
Might as well be a trillion dollars as it's never getting paid.
Fred and Wilhelmina are in bed together. Fred tosses and turns, unable to sleep. "Fred, what's troubling you?"
"You know Bjarne from the bank? Well, a balloon payment is due tomorrow on my loan, and I don't have the cash flow to pay it."
Wilhelmina thinks for a bit, then reaches for her phone. "Betty? Yes, this is Wilhelmina. Sorry to call so late. Would you tell Bjarne that Fred can't make the loan payment tomorrow? Yes, that's all. Good night."
Fred states at Wilhelmina, aghast. "What did you do THAT for?" She smiles. "Now it's Bjarne's problem. Let him toss and turn, you can go to sleep."
Ridiculously unimportant trivia: The Flintstones borrows heavily from The Honeymooners, was the first animated sitcom, and was also the first animated show aired in prime-time.
In response to the 340M fine, .... "We have set up a LLM and are deleting all our books ? "
[1] https://github.com/aaronjanse/dns-over-wikipedia
better not miss!
Eventually US will have its own great firewall like China.
Politics is about powergrab and what we’ve seen is more and more power grab.
It’s likely the billionaires are able to buy the govt goons to pass the laws that allow them to be the gatekeepers.
Anthropic, OpenAI and the model builders massively benefited from the archived information. Distill entirety of archived human knowledge.
I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster. Or how Aaron Schwartz was executed by proxy by JSTOR and the feds, for what should have been free to access for all.
But hey, Anthropic, OpenAI, X, and others can pirate to their hearts content with for-profit piracy, but "we" (royal) are OK with that. We just cant have the poors have access to the sum of human knowledge.
Why didn't Metallica's fans rebel? That would have stopped it quickly and set a precedent and example for the rest.
What an incredible strategy - and it worked! And in something that's entertainment and fandom-driven; it's not like the fans need a medication, job, or their auto warranty. How pathetic. The day rock'n'roll died.
Oh yeah? Well, my dad works at Anna's Archive.
Jason Arday recently committed suicide due to media pressure, and even when the media was told that he was not mentally well and should please back off, they didn't.
The man was accused of plagiarism, and there was no final verdict on if he did. The institutions that hired him clearly had no issues in vetting him. And yet he committed suicide.
There are absolutely horrible people out there right now who's entire careers are shams, and don't even come close to thinking about it. People are built different, and have different triggers
Bullshit. It wasn't the first round of plagiarism, or redactions. His whole supposed life story was a fraud. He passed off other researchers' patient interviews as his own.
Weird way to spell "lied about nearly every facet of his life and anyone who questioned it was aggressively accused of being racist."
Please also educate yourself about how federal prosecutors treated Schwartz after he scraped PACER.
Please also educate yourself about how long prosecutors attempted to negotiate a plea deal with Schwartz, what their final offer was, what his family, friends, his girlfriend, and his team of top-shelf, famous lawyers advised him to do, and what he actually did.
And PACER is federal court documents. Its a fucking scam and travesty that a system soo heavily built on precedent gatekeeps its own law at $.10/page.
Aaron is a motherfucking hero.
If I get dumped and I commit suicide because of it, does that mean I was "executed by proxy" by my girlfriend/boyfriend?
"Executed by proxy" is ridiculous, overheated rhetoric.
Many people have faced similar or worse charges, and did not commit suicide, so we'll have to point at something else.
Fuck them for what they did to him, but we should be rallying to get him set free. We can't start pointing fingers at other people for suicides. We have agency.
Feds did NOT have to choose to go the route they did. Their pursuance of a victimless "crime" lead to Aaron's only path was 'Exit'.
Sure, he committed suicide. Why? Because his life was already significantly threatened with what amounts to torture by prison.
Any time someone starts shrieking "half his life" you can assume they don't know a single factual thing about the case.
From the tone, ignorance, drama, and attitude I'm guessing they're about 22 years old.
Piracy is morally justified at this point.
Maybe if they use encryption, which would make them all pedophiles according to those governments. Hope you have your VPNs ready if visiting countries without a free internet, like North Korea, China, UK, or EU.
Thankfully I don't live in the US, and likely won't set foot there ever again.
If anything, I was heavily underpaid.
You're underpaid now at your euro job, that's a guarantee.
I also have no desire to return.
They're not more or less virtuous than any others, they're in it for the money and you're a fool if you think otherwise.
I'm not against piracy, it's great, but to claim they are moral saints is a joke.
Anyway, bring in the downvotes as I know will happen.
it doesn't seem like it would be high revenue at least
Anna's Archive is very limited in performance if you don't pay.
Maybe if youre trying to pull an AI lab and download their full corpus, although I suspect thats not too onerous either. I have never had any issues getting stuff from annas archive for free
It also helps that AA has the best UI and search of any shadow library I've seen. Some really established bittorrent sites win out on curation but otherwise AA is top tier, not something I'm accustomed to seeing from a (free, no-signup!) clearnet/direct-download site
As far as GP's question I think you can be in it for the money (though I personally have my doubts that Anna is) and also be a force for moral good as well. Whether the operator's morals perfectly align with all of our ethics is an open question but the room at large generally tends to agree.
Also Torrentfreak's reporting has been extremely solid for an extremely long time so the link itself fits here too
Maybe you can own (some) things, you can't own information
It's not a competition, and being a moral saint isn't a good KPI; being a net positive for society is.
If you just Google 'free media heck yeah' you'll see a great deal of them!
OTOH, if you think that AA is robbing creators and publishers of their hard-earned proceeds, and that loss is greater than the loss to humanity from the destruction of AA, then you wouldn't think AA was a moral good.
Of course, it's all gray in the middle.
You cannot claim that AI labs are bad for illegally downloading content to build their AI without paying creators, and that AA is good for illegally offering the same content.
I think both are immoral, but I definitely pirate all the content I can (except video games but only because it's unsafe, not to remunerate creators).
It is not grey at all, this is some bullshit that people tell themselves to feel better. Whether it's using the output of AI models or downloading a book on Anna's Archive, you are ultimately robbing the creator of profits.
Nobody on HackerNews would argue otherwise if their employer stole their code from their mind and didn't pay for it.
You're correct in the above, or at least I'm willing to accept the premise for the purposes of this argument.
However, it's not that simple when it comes to the morality of the actions since there are secondary effects.
You can't just make a blanket statement like "stealing is immoral", for example, when there are circumstances where stealing is absolutely the moral imperative (e.g. the food is held by barons who charge unaffordable prices and the populace is starving to death).
Or, to make an example closer to the topic, what about the stealing of scientific papers? There's a strong arguments to be made that millions of people have been helped through that theft. And yes, the gatekeepers of that knowledge have been harmed.
And it's like that with AA and AI (and their code) as well. People consider the secondary effects and whether or not the greatest good is served.
And, of course, "greater good" is gray.
Of course on a related note, most work that's interesting to me was created by people who are dead anyway, so they don't mind.
Anyway, are you sure most people that support AA don't also support e.g. Deepseek and Qwen's efforts?
Is your angle merely that shadow library hosts should be operating their services entirely for free (to the point of refusing payment) or is there something we're missing here?
If you want information to be free? Great. But most of those authors wouldn't be putting the work in to making that information/literature in the first place if they know they aren't going to get paid.
My understanding is that the number of authors that can make a full-time job of writing is a rounding error compared to the population of authors. I believe that people should be paid for their work, but I don't think the current copyright regime is actually very effective at paying authors for their work. So I value copyright enforcement proportionately less based on that observation.
Meanwhile, on the flip side of the coin, copyright rulings are causing companies like Atheropic to destroy books as they scan them, creating a rising sense of panic around knowledge scarcity. This panic is a direct result of the scarcity of the copyright intended to create in the first place. It's entirely artificial.
I've been saying this for 25 years, but copyright is essentially broken.
I don't really agree
Information itself should be free, authors deserve the right to monetise other ways (merch, physical media, exhibitions/shows, etc.) but society should pay artists and creators to do their thing[0]
Not everything needs to be a business, and I think art is one of those things
[0] I don't know much about it but maybe Ireland's Basic Income for the Arts scheme is a model. I think some other European countries do similar-ish things, too
So I would say it is morally acceptable to download stuff older than that. Torrents offer brand new things, so I wouldn't say they are morally good. I don't think they're morally worse than those locking up old media though.