An Update on Wayback Machine Access

(blog.archive.org)

287 points | by ChrisArchitect 4 hours ago

26 comments

  • simonw 3 hours ago
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    • packetslave 3 hours ago
      This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
      • bsimpson 3 hours ago
        It's an open secret that you can often circumvent paywalls by searching Wayback.
        • gambiting 3 hours ago
          Every single paid article linked on HN has the way back machine link as the very first comment.
          • ValentineC 3 hours ago
            The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
            • eek2121 8 minutes ago
              Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.
            • petcat 2 hours ago
              [flagged]
              • sandcat_ 2 hours ago
                That isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.
                • petcat 2 hours ago
                  It's bad form to scrape the scrapers?
                  • sandcat_ 1 hour ago
                    Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)
                    • petcat 1 hour ago
                      You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason.

                      The end result is exactly the same.

              • organsnyder 2 hours ago
                They're different sites, with different goals, run by different people.
                • petcat 2 hours ago
                  That provide the same functional service....

                  Hence, distinction without a difference.

                  • celsoazevedo 2 hours ago
                    They are 2 different services, run by different people, one goes out of their way to bypass paywalls while the other doesn't, one is banned by Wikipedia and the other isn't, etc.

                    I think it's a distinction worth making.

                    Not to mention that the Wayback Machine itself isn't exactly a good tool to bypass paywalls as most paid sites don't let them archive paywalled content anyway.

                  • fluffybucktsnek 2 hours ago
                    Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine, it very much is a distinction with a difference.
                    • petcat 2 hours ago
                      Bot traffic or human traffic doesn't matter. The goal is to read websites without having your own access.

                      So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.

                  • rpdillon 1 hour ago
                    Yeah, you're mistaken. One archives web pages, the other maintains a list of paid-access accounts and fetches information from behind paywalls as a service.
    • autoexec 22 minutes ago
      I've personally been using the Wayback Machine more often because I increasingly find myself being blocked from websites who are trying to keep out scrapers even though I'm just a regular person with JS disabled (along with a bunch of other stuff)
    • pantsforbirds 3 hours ago
      We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice!

      Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.

      • subarctic 3 hours ago
        What if they charged money? Is it something you'd pay for?
        • bonestamp2 2 hours ago
          I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.
        • msephton 2 hours ago
          I'd pay for it, but only if they implemented the changes the community of users have been requesting for years.
          • carlosjobim 1 hour ago
            No matter what they did, you'd have a new excuse for why you won't pay.
            • msephton 53 minutes ago
              Ah, the old ad hominem attack. How refreshing.

              But anyway, no, I wouldn't keep finding reasons. I donate to them every year already. Somebody asked if I would be willing to pay and my answer was "yes, but".

              It would need to be improved because certain aspects of it suck right now, not only the error this post is about. They only need go as far as their forums and github repos to see the community feedback.

              • jakderrida 20 minutes ago
                Is it really an ad hom if he doesn't know the hom?

                Their reply is 100% based on the content of your post.

        • bee_rider 2 hours ago
          I wonder if there would be concern on their part about appearing to be a company that was basically offering paywall circumvention as a product.
          • cloakley 1 hour ago
            It wouldnt be a paywall, more like an option for companies to not pay scrappers. At least the payment deviates to the source.
    • bradly 3 hours ago
      Just yesterday from my one of my sessions with Sol:

      > Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

      • TeMPOraL 3 hours ago
        As it should.

        Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

        • bradly 2 hours ago
          Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.
          • aaron_m04 2 hours ago
            robots.txt?
            • bradly 2 hours ago
              Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.
              • bityard 2 minutes ago
                robots.txt applies (or should, in my opinion) to anything that automatically follows a link. Basically any software that is not a human-controlled web browser or single-shot curl command. Everything else: robot.
              • xena 1 hour ago
                AI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out.
                • ghaff 27 minutes ago
                  From the start, robots.txt has always been an indicator of a site's preference with no actual legal significance.
                • recursive 22 minutes ago
                  If a new directive was introduced that allows for an explicit setting in robots.txt, do you think the bros would follow it anyway? Something like `ALLOW AGENTS` or `DISALLOW AGENTS`
          • Analemma_ 1 hour ago
            I want agents to be able to act on my behalf, that’s the entire point. An agent should be able to do anything I can do sitting at my browser.
            • ryandrake 3 minutes ago
              Technically, even your browser is an agent. It says it in the HTTP: User-Agent. So is cURL. Every application the user runs is acting on the user's behalf.
            • compiler-guy 1 hour ago
              I suspect most people would be ok with this if they could only do it at the rate and frequency you yourself can do it. The problem is largely one of scale.
              • daveoc64 27 minutes ago
                Is scale what we're discussing though?

                e.g. a prompt of "fetch <article URL> and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.

              • cruffle_duffle 1 hour ago
                Then make agent friendly content. Take the text and make a markdown version.
                • fineIllregister 43 minutes ago
                  People doing this say it makes things worse because then the bots download both.
                  • compiler-guy 35 minutes ago
                    Not to mention that it solves none of the rate issues. If the scrapers are hitting your site 10,000 times a day, adding markdown isn’t going to change that at all.
            • cruffle_duffle 1 hour ago
              Dunno why the downvotes. I feel that is reasonable as well. Owners that block that stuff are doing so only to their detriment.
    • RobotToaster 2 hours ago
      Do they offer bulk torrent downloads as an alternative?
      • QuantumNomad_ 1 hour ago
        Once upon a time some people explored backing up the Internet Archive.

        However, that experiment ended. They mention there were some learnings and they then say:

        > The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use.

        https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK

        I would really like to know if any sort of thing like that is still ongoing and if it’s accessible to people in general. Would be nice to mirror some data from IA to my local drives, for example via BitTorrent or IPFS, to have it for offline exploration and personal archive.

        I know that individual items have torrents. And I’ve downloaded a few that way but always it ends up only using the “web seed” (i.e. the BitTorrent client is retrieving the files from IA via HTTP) because there are no one seeding some random single item I found. Plus, those torrents are unreliable sometimes because they include meta data files that were since updated but the torrent was not updated and so the web seed is giving the updated files that don’t match what the torrent says their hashes should be. So then you have to jump through some extra hoops to fix that and then resume the download, and all the while the HTTP connections to IA servers time out because their servers are overloaded. So when I say I wonder about possibilities of using BitTorrent I mean to retrieve whole collections of many items instead of individual ones, and with actual other peers instead of just having it put load on IA HTTP servers.

      • echelon 2 hours ago
        I would love to be able to download every page of a given domain as an archive, and I'd pay to do this.
        • msephton 2 hours ago
          They provide a free cli tool to do this.
        • petcat 2 hours ago
          isn't that what wget -m does? what is there to pay for?
        • carlosjobim 1 hour ago
          You'd pay the domain owner for it? How much?
    • luckylion 3 hours ago
      What sites would they be targeting? Generic "just give me anything"? Whenever I check regular sites on IA, the coverage is spotty -- they'll have the homepage and a few important pages, but it quickly fizzles out.

      Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.

    • jader201 2 hours ago
      > I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

      Appalling, yes. But also expected. I'm surprised they haven't been the target of scrapers for years. But sites putting their content behind login walls and other anti-bot mechanisms has certainly exacerbated this. But again, this isn't at all a surprising progression.

      > we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

      To be fair, another big motivation was likely users on sites like HN using archive.org (and similar sites) to get around their paywalls. In fact, I'd be surprised if this wasn't a big motivator.

      Again, it sucks, but it's not at all surprising to see it progress like this. I wouldn't be surprised to see similar blocks on other archive sites eventually.

    • toomuchtodo 3 hours ago
      It is. They will most likely eventually need to move to a walled model for Wayback due to scraper aggressiveness (like Reddit deprecating anonymous old.reddit.com), or behind Cloudflare for aggressive bot and scraping protection. Hard to defend against abuse of a public resource when its intent is public access with as little restriction as possible.

      https://en.wikipedia.org/wiki/Tragedy_of_the_commons

      (no affiliation)

      • ronsor 3 hours ago
        Reddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy.

        On the other hand, the Internet Archive is a non-profit offering a free public resource.

        • toomuchtodo 3 hours ago
          Examples provided as technical examples, strong feelings are out of scope for this thread.
          • itintheory 3 hours ago
            As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a temporary bandaid.

            The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons.

            [0] https://people.kernel.org/monsieuricon/creepy-crawlies

  • basilikum 2 hours ago
    Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access.

    The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.

    If you got some money to spare, consider donating to them. They need it.

    • ternaryoperator 1 hour ago
      I donate to them every year b/c I fully agree they’re doing a thankless critical job very well.
      • j79 34 minutes ago
        Thank you for the inspiration! I just made my first donation.
    • superxpro12 2 hours ago
      fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future.

      The future is bleak :\

      • mrguyorama 1 hour ago
        I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side.

        It's going to suddenly be extremely valuable that wikipedia didn't settle for having a small rainy day fund and instead ceaselessly grabbed every fucking donation they could for two decades so they can fight such a legal battle.

        • nephihaha 38 minutes ago
          Wikipedia is not a challenge to the system. If it were, it would be excluded from search results, much like blogs are now.
        • pvab3 38 minutes ago
          on what basis?
  • BeetleB 3 hours ago
    Wow, but I wonder if there's more to it.

    I've not been able to access web.archive.org from my work computer - I always get the 429 error.

    But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

    • flexagoon 3 hours ago
      I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP
      • jcrawfordor 1 hour ago
        There are definitely factors beyond IP being used. A week ago I found that all requests from Chrome-like browsers got a 429 across more than a half dozen networks and several machines, while Firefox reliably worked. I assume this was an overzealous policy on UAs.
        • BeetleB 1 hour ago
          In my work, it's failing on both Firefox and Chrome.
      • BeetleB 1 hour ago
        No - my phone is not connected to work's WiFi.

        Wonder who the bad actors in my company are...

        • flexagoon 1 hour ago
          Doesn't have to be someone at your work, it could be a block on a whole ISP network or at least an IP subnetwork that is shared between many clients

          You can try emailing the address mentioned in their post so they adjust their filters to match just the bot networks more precisely

        • iamacyborg 32 minutes ago
          It might just be your corporate VPN and whatever ASN it’s being routed through.
        • ButlerianJihad 54 minutes ago
          You should file a support ticket with your manager and the IT security or support desk. Show them the evidence of 429s that are blocking your assigned tasks during working hours. Also include the screenshots and files that you downloaded on your personal device in order to access your work-related materials. Be sure and thank them for adequately configuring the MDM on your personal mobile device so that you could do these work-related tasks. You should definitely also file an expense report to request reimbursement of your personal mobile bill, any data charges incurred, and the hours of networking or collaborating with external colleagues, while you were working on these work-related projects with your personal device.
    • novok 2 hours ago
      Your workplace is probably redirecting traffic through a datacenter IP range. Especially if they have their own datacenters like google, microsoft, oracle, amazon, etc.

      Try making a vpn via digital ocean for example and you'll see similar patterns.

    • dotmanish 3 hours ago
      Could be due to some scrapers from either your work ISP block, or the larger block which lends IPs to multiple workplaces.
  • emaro 2 hours ago
    It's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/

    I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.

    • zdragnar 2 hours ago
      I'm a little more skeptical that this is "AI is big so it is worse" issue. Yes, AI is big in scale, but this has been the case for almost every popular free service. They either start:

      - charging (news / journalist services)

      - gate-keeping (X forcing log-ins)

      - enshittifying (lots of ads and degraded service)

      The fact that the way back machine is incredibly useful but most people didn't know about it or use it very much doesn't change the fact that it has basically become very popular... only with LLM agents rather than humans. Ads alone aren't enough to support human traffic for many sites with human traffic.

  • CqtGLRGcukpy 3 hours ago
    > We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.
  • pelican0 2 hours ago
    Is it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies?

    Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.

  • lousken 3 hours ago
    AI companies should pay billions to wayback machine for access
    • KPGv2 3 hours ago
      I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.
      • autoexec 14 minutes ago
        They wouldn't be paying for the content, just the bandwidth. Like buying a linux OS on a CD ROM was about the cost of media not profiting off of the software.
      • roblh 3 hours ago
        Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?
        • Joel_Mckay 3 hours ago
          That is essentially what LLM vector search results are, but the misappropriated $9Tn worth of FOSS code "AI" scraped and compacted for isomorphic plagiarism tokens is harder to prove now with watermarking skewed outputs. =3
      • jMyles 3 hours ago
        It's time for copyright to end anyhow; that's what's gumming up the whole project in the first place.
        • autoexec 12 minutes ago
          I'd have a lot less of a problem with AI if everything that went into their training was public domain and made easily available to anyone for any use. It'd feel less like AI companies were just stealing the work of others and charging for it.
  • tgtweak 24 minutes ago
    Can't wayback machine just offer direct access to the archive for a premium and in doing so, pay for the service?
    • edelbitter 9 minutes ago
      Not while the new dukes of the internet wielding massive armies of hijacked smart TVs have a better time browsing the web than I have; as a mere peasant with just a few IP addresses. There would be no reason to sign up and pay up for bulk access, unless open access is shut down.
  • ilamont 1 hour ago
    Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website?

    My blogs are getting slammed and there are issues with cloudflare or captchas.

    • iamacyborg 30 minutes ago
      > Shouldn't the solution be to gate bulk access for automated services for a price?

      Fine in theory but determined scrapers will use residential proxies in bulk.

  • thimabi 2 hours ago
    I wonder why doesn’t the Internet Archive require logging-in prior to accessing the Wayback Machine. It would probably help them distinguish humans from bots, at a very little cost to humans.
    • extralongdivisi 2 hours ago
      Gatekeeping information is not the solution
      • hamandcheese 27 minutes ago
        Why not? If its the difference between the information being available at all, then I choose login any day of the week.
  • timpera 3 hours ago
    I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them.

    Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.

  • xacky 1 hour ago
    The anti virus industry needs to crack down on crawler and proxy malware, plus ISPs FINALLY need to replace CGNATs with iov6 to stop crawlers banning everyone behind a NAT.
  • int32_64 2 hours ago
    Are any AI companies using residential proxies to scrape?
    • xena 1 hour ago
      Yes. It's impossible to tell which because the split is residential proxies, dataset curators, and AI companies all being separate actors. However I fucking guarantee you it's out there and people are too cowardly to be honest about it so they don't get sued out of existence.
    • oasisbob 40 minutes ago
      Oh yeah, absolutely.
  • tech234a 3 hours ago
    I wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.
    • stickfigure 3 hours ago
      Plenty of threads on HN about this, Anubis does not work.
      • autoexec 10 minutes ago
        It always seems to keep me, a normal human, locked out of any site that uses it.
      • phendrenad2 9 minutes ago
        [delayed]
  • hubraumhugo 2 hours ago
    There is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement. So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race?

    Some approaches that I think are promising:

    - A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).

    - Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.

    - what else?

    [0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...

    [1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...

    • maxrev17 2 hours ago
      Yeah it’s kinda crazy to me that what was once a back alley python script is now accepted as ‘fine, free for all’. The new era of bros really are smth else.
  • brador 1 hour ago
    The only solution is to make visitors do compute. Compressing files for the archive to access other files would be perfect for this.

    Cross verify hashes to prevent cheating.

    Ez.

  • vlyan 3 hours ago
    unrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?
  • MattCruikshank 3 hours ago
    There was a feature on Amazon Web Services for a while, and I wish it was still there...

    Downloader pays.

    I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.

    I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.

  • ignoramous 3 hours ago
  • UltraSane 3 hours ago
    Why not put it in S3 with downloader pays?
    • charcircuit 2 hours ago
      S3 price gouges on bandwidth.
    • Kayvanian 3 hours ago
      As a public resource the hope is for Wayback to be free to access. I imagine putting up a paywall would be their last resort.
  • Onavo 3 hours ago
    Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon.

    It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.

    I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

    • oasisbob 33 minutes ago
      > It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs

      The problem with this perspective is that it ignores the victimization which is happening to all sorts of sites right now.

      On one hand, you have content owners/suppliers which are trying to place restrictions on how much free bulk use is allowed.

      When scrapers go to exotic lengths to evade the blocks, eg by using thousands of ephemeral IP addresses to collect an entire corpus, saying stuff like that makes it sound like it's all a wash.

      "Oh, what a silly situation... How did we ever end up like this? It's not good for anyone ..."

      No, there is a victim trying to defend themselves from rampant theft of resources, and a corporate asshole which doesn't care about the effects of their actions.

    • drdexebtjl 3 hours ago
      Sites would just block the Internet Archive crawler as well.
    • imglorp 3 hours ago
      Micropayments would solve so many Internet problems. It's not too late to adopt.

      Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage.

      The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.

      • novok 2 hours ago
        Micropayments are blocked by government money laundering regulations increasing the costs significantly to make them untenable.
      • Analemma_ 2 hours ago
        Micropayments would solve all the problems except for the problem that people absolutely loathe micropayments. Like, vein-popping furiously hate them.

        Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact that users despise micropayments and we'd rather blame credit card companies for the lack of adoption.

        • imglorp 1 hour ago
          It doesn't need to be crypto, or payment processor based.

          My ideal experience would be I load $20 into the browser somewhere like a wallet in one block (that could be a payment processor step). If I visit a participating page, it decrements my wallet $.01 or whatever.

          The downside is the possibility of abuse and tracking by governments, which would have to be handled at the source, not the symptom.

        • mindcandy 2 hours ago
          Micropayments would solve so many problems for the internet. And, cryptocurrencies would solve so many problems for micropayments. But, it's a non-starter because any proposal gets flooded with people popping veins about how crypto can't solve anything.
    • xp84 3 hours ago
      My guess? Because even with a paid endpoint, the type of unscrupulous yahoo that is DDOSing IA today would probably still abuse the free endpoints because they can. The revenue that might come from a paid endpoint could help to scale up, but with how slow IA usually seems, I suspect there is an upper limit to how much traffic they can serve without a LOT more revenue.

      This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.

    • KPGv2 3 hours ago
      > Why not just offer a paid endpoint for the crawlers?

      Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.

      • Ajedi32 2 hours ago
        What if you're not charging for the content, but as compensation for the network bandwidth / server resources consumed by serving that content? The idea isn't to profit from content (the IA is a nonprofit anyway), just to allow the IA to continue to serve its purpose as an archive of public data without being overwhelmed by bots.
        • Onavo 2 hours ago
          Exactly, it's a question for the lawyers to sort out.
    • croes 3 hours ago
      It’s one thing to archive other companies content, it’s another to sell the access to it
      • faefox 3 hours ago
        Yeah, who does the Internet Archive think it is, (insert literally any AI company here)?
        • bonoboTP 3 hours ago
          Which AI company is selling access to reliable verbatim copies of websites? I don't mean "it may regurgitate a paragraph", but as a reliable service where you can repeatably get website content snapshots to a reliability level that makes such a use case viable?

          Using the information for training purposes is not the same thing. Not legally the same and otherwise.

      • Onavo 3 hours ago
        That's for the lawyers to sort out, they have a lot of flexibility as a US nonprofit. The case law isn't that clear cut for this.
        • simonw 3 hours ago
          Internet Archive was almost destroyed by a copyright lawsuit from book publishers within the last few years. I expect they aren't excited to take on any additional risk of similar lawsuits right now.
          • quotemstr 1 hour ago
            They brought it on themselves by marketing a read-for-free product
        • celsoazevedo 3 hours ago
          They need access to sites to archive them. It's already hard to do it as it is, imagine if they start selling access to content. They'd be shooting themselves on the foot, independently of what the law says.
        • xp84 3 hours ago
          major [citation needed] on that. There are very limited exceptions to the massive power of copyright -- and they're mainly granted to libraries in the form of narrow waivers. And just the cost of fighting the most powerful copyright holders can bankrupt you -- especially if you're a relatively modestly-funded nonprofit.
  • josefritzishere 49 minutes ago
    [dead]
  • unkeen 3 hours ago
    [flagged]
    • tomhow 46 minutes ago
      We detached this subthread from https://news.ycombinator.com/item?id=49716735 and marked it off topic.
    • stronglikedan 2 hours ago
      Yes, that's one acceptable alternative, and another commonly accepted alternative is API's. Although, I'm not sure why you included the asterisk.
      • maxrev17 2 hours ago
        Unkeen on the apostrophe that’s why! Gotta keep HN proper and correct guize
  • xyst 3 hours ago
    [flagged]
    • plorkyeran 2 hours ago
      If you're in a room with a TV then literally yes, there's a good chance there's an abusive bot in the room.
    • gooeyblob 2 hours ago
      What reason do you have to doubt the claim?
    • alex1138 3 hours ago
      I mean there are people who have reported that with their own personal website Facebook's crawlers were essentially DDOSing them
  • swingandamiss 3 hours ago
    [flagged]
    • kg 3 hours ago
      Does xcancel scrape twitter? Isn't it more like a proxy for specific user requests to view tweets?
    • knowaveragejoe 3 hours ago
      Correct, and nothing wrong with that.
      • MadameMinty 3 hours ago
        "Kidnapping innocents bad but imprisoning criminals good?? Inconceivable!"
    • yifanl 3 hours ago
      It's almost as if moral values aren't assigned universally.
    • righthand 3 hours ago
      No one is upset that the AI companies are scraping the web, they’re upset how poorly implemented the scrapers, but the scraping itself is fine. Lots of people and businesses scrape the web.
      • akerl_ 3 hours ago
        There are people commenting parallel to you saying they are upset about AI companies scraping the web.
    • faefox 3 hours ago
      Yes, anything that potentially costs Elon Musk money is objectively a good thing. :)
    • dallen33 3 hours ago
      Yeah cuz X is fucking shitty, why would I want to give them any traffic?
      • xp84 3 hours ago
        Then... don't? If it sucks so much why do you need to read the tweets?

        Great take: "This private website is owned by a man I don't like, so I refuse to pay for it - or even give it the possibility to monetize my traffic with ads!"

        Still quite mainstream take: "... so I'll use an adblocker on it"

        Immature take: "This private website that I hate and boycott is also an important part of our culture, but the posts on it are too important and valuable to ignore, so I'll use a proxy to scrape it"

      • slig 3 hours ago
        You're giving them attention, thus validating their existence and their numbers.
      • qwerpy 3 hours ago
        “It’s ok to do bad things to people/things I don’t like”

        Feels good when you get to dish it out doesn’t it?

  • msephton 2 hours ago
    I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.
    • jolmg 2 hours ago
      > Asking users to email them with details of their OS, browser, IP address is just crazy.

      It's surely to serve as data to help tell humans apart from bots.

      > Changes made by IA shouldn't become my responsibility.

      They're a free service. It's ultimately not their responsibility to service you either.

      • msephton 46 minutes ago
        Imagine if Apple or Microsoft introduced a bug and said, ah yes we know about it we did that on purpose and we know it affects a huge number of people, if each of you could email us these details that'd be great. It's just such an insane request.

        IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.

        • jolmg 26 minutes ago
          No, it's more like you're requesting something from them and they're telling you they may need some technical, non-personally-identifiable info from you to fulfill your request.