Ask HN: Why are OpenAI, Claude, and Grok simultaneously down? Coincidence?

https://status.openai.com https://status.claude.com https://status.x.ai

113 points | by halcdev 1 hour ago

35 comments

  • kibae 44 minutes ago
    Cloudflare, Azure, AWS, and Google Cloud all have a similar uptick in reported errors around 7:30. I suspect an outage on Cloudflare or another load-bearing service cascaded through all the major cloud providers.

    https://downdetector.com/status/cloudflare/

    https://downdetector.com/status/windows-azure/

    https://downdetector.com/status/aws-amazon-web-services/

    https://downdetector.com/status/google-cloud/

    • oersted 22 minutes ago
      “load bearing” :)

      For once it’s appropriately used.

    • y-c-o-m-b 26 minutes ago
      Cloudfare did release an update for their "HTTP/3 issue affecting R2 custom domains" around that time

      https://www.cloudflarestatus.com/history?type=incident

    • martyfunkhouser 15 minutes ago
      Will the post-mortem reveal they all relied on a service running on a Macbook in a break room with a "Do not turn off" sign taped to it?
      • hejdidf 6 minutes ago
        you made an account just to write that tripe good lord on high have mercy on your children
        • esafak 5 minutes ago
          Evidently it annoyed you enough to create one too. Remember, guys, there's one more day until Friday!
      • hejdidf 6 minutes ago
        how do you handle having such a clever and original mind
    • KptMarchewa 35 minutes ago
      Not really. The impact isn't as big too - Codex for example did not stop working for me.

      https://updog.ai/

      • cromka 18 minutes ago
        > Not really. The impact isn't as big too - Codex for example did not stop working for me.

        Buddy, read the room. Just because it works for you doesn't mean it works for everyone. ESPECIALLY if the suggested issue here is, indeed, with Cloudflare.

      • taytus 7 minutes ago
        oh well, if it is working for you, then we are saved.
  • niobe 1 hour ago
    Well no one said it yet so I will, "international actors" is at least a possibility. And I don't mean any specific country because pretty much anyone is a potential these days, which makes it a perfect cover for different anyones. Demonstrating vulnerability in the US's AI boom can move the markets. That's a financial incentive and a strong geopolitical one.

    More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible"

    • gleenn 49 minutes ago
      Everyone is leasing datacenter space from some of Grok, Google, and Amazon aren't they? If it's hardware or DC level disruption I'm not too surprised it can affect multiple providers.
      • pixl97 0 minutes ago
        Also it's likely that more than one model use is common.

        Amazon starts going slow so some percentage switches to Google, some switch to Grok, now all of them are slow.

    • sixQuarks 18 minutes ago
      Except that the stock market is up today
  • juujian 16 minutes ago
    Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.
    • giancarlostoro 4 minutes ago
      I have a feeling this is part of it, especially when you consider how many services let you use any of many available AI providers.
  • Insanity 1 hour ago
    Think of it like one big distributed system. OpenAI is down, so people migrate to Claude, now this one gets overloaded and goes down, etc.

    So not a coincidence, one went down first and users migrated causing further DOS. At least that's my guess.

    • erdos_2 1 hour ago
      It'd be funny if this is true because that'd prolly mean nobody is touching Gemini even as a fallback.
      • nevir 17 minutes ago
        Or that Gemini is built to handle massive load spikes, and/or has a ton of excess capacity
      • sroussey 18 minutes ago
        I did, for stuff i do in cursor.

        i also finally installed opencode and switched its model to muse 1.3

        both are decent.

      • Insanity 1 hour ago
        Lol I didn't even think about Gemini missing from the list. Not sure what that says about Gemini or me :)
      • rtcoms 1 hour ago
        Just now I got this from gemini

        It looks like there's no response available for this search. Try asking something else.

        • exe34 1 hour ago
          I bet they had to implement that manually to make it look like they failed too!
      • giancarlostoro 3 minutes ago
        Someone noted Gemini was also having issues in another thread.
      • gleenn 45 minutes ago
        Google stopped putting so much money into SOTA models. All the hype has migrated. I was also frankly turned off when I got a popup from Gemein said I would either have to pay or have my conversations used for training. This may have always been true for other providers but when I declined, Gemini stopped remembering my conversations and that definitely made me move out.
        • ilaksh 43 minutes ago
          Gemini 3.8 which just came out sounds like it's very good and a great deal though.
      • benatkin 1 hour ago
        Not even the best agent that starts with a G
    • fny 1 hour ago
      I find it hard to believe that enough people would flock to from Claude and Chat to Grok to cause an outage. I feel like Gemini is the dominant release valve in this case especially for enterprise.
      • nevir 15 minutes ago
        Don't forget that there are a ton of tools out there that will automatically fall back in case of outage

        E.g. say you chose Sol as your default in Cursor, but Opus is your 2nd choice, it's going to give up on Sol after a few tries and switch to Opus

      • baq 51 minutes ago
        It’s cursor’s model so plausible, lots of folks use cursor still.
    • v3rm1n 17 minutes ago
      This is what Tibo posted on twitter in response
    • paxys 1 hour ago
      Especially considering memory/gpu/compute are scarce so these services are likely running with very little buffer.
    • toomuchtodo 1 hour ago
      https://en.wikipedia.org/wiki/Domino_effect

      Edit: Updated per valleyer's suggestion.

      • valleyer 1 hour ago
        "Domino effect" would probably be the more relevant named phenomenon there.
      • throwaway894345 1 hour ago
        This isn’t a thundering herd problem, it’s a cascading failure. (Thundering herd is about a bunch of workers waking up simultaneously)
  • docheinestages 1 hour ago
    My gut feeling tells me it has something to do with Cloudflare. Along with AWS, they're two of the main suspects in such incidents.
    • hosteur 55 minutes ago
      I thought OpenAI famously used Azure due to their partnership with Microsoft?
      • nullpoint420 51 minutes ago
        They use a lot of compute providers now, but they use Cloudflare for their networking
  • sebbul 35 minutes ago
    Traffic rerouting through NSA had a hiccup…
    • ibejoeb 2 minutes ago
      Room 641A is being cleaned, but we'll hold your bags for you.
  • Linello 1 hour ago
    What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?
  • codazoda 1 hour ago
    I kinda assume it's because one went down and a large amount of work shifted to another.

    I'm also aware that they have overlap in some areas on data centers.

  • faitswulff 1 hour ago
    Heard on the grapevine that the OpenAI blip was a cloudflare issue
  • Avicebron 1 hour ago
    I suspect Azure is having issues, Microsoft has had outages the paat two days, especially with email.
  • CSMastermind 1 hour ago
    I assume it cascaded from one provider to the other as people who lost claude access for instance moved to openai who moved to grok when it went down, etc.
  • maxbaines 1 hour ago
    They all rent compute from SpaceXAI
    • lavezzi 1 hour ago
      I don't believe OpenAI does
      • maxbaines 1 hour ago
        My mistake, in fact it was google not OpenAI, makes sense OpenAI doesn't.
    • halcdev 1 hour ago
      Surely it's a bit more distributed than that, right?
  • elar_verole 1 hour ago
    Pretty sure it's a US thing since it's available here in France. What exactly is down, idk
  • nmlt 46 minutes ago
    Somebody in another thread said gastown and wheelhouse automatically move to the next provider if one fails.
  • chasd00 1 hour ago
    claide.ai is working for me, so is chatgpt.com. grok still has a status message about issues, i can't try it without signing up.
  • dmillar 57 minutes ago
    Seeing 503s on Gemini via API as well
  • JackFr 58 minutes ago
    Obvious answer is it's the AI singularity. Been nice run for humanity. So long everyone.
  • dgorges 42 minutes ago
    It's always DNS
  • jedbrooke 1 hour ago
    according to https://downdetector.com/ Gemini is down too (and copilot, but that just uses ChatGPT right?)
  • convivialdingo 1 hour ago
    The Thundering Herd has thundered, apparently.
  • elorant 1 hour ago
    Some npm library that makes headers bold would be broken.
    • ibejoeb 1 hour ago
      Oh man. Some low effort supply chain attack that turns every GPU into a cryptominer. It's funny because it's plausible.
    • N_Lens 1 hour ago
      Ah yes ye olde bold-headers: ^3.13.31;
  • wejick 1 hour ago
    Probably same public cloud or CDN in front of them.
  • kocial 1 hour ago
    Maybe the stack behind it is down, like AWS or something
  • satvikpendem 1 hour ago
    They're all using Cloudflare.
  • dgellow 1 hour ago
    Too early to know, let’s wait and see
  • fidla 1 hour ago
    chatgpt is back
  • jauntywundrkind 51 minutes ago
    Fable 5.1 got released and generally I tend to think as soon as there's a new release there's this massive spike in people benchmarking & comparing, that services tend to go slow everywhere as everything gets super loaded. This should hypothetically be visible on OpenRouter too, so I guess someone could check and see if there's any merit to this idea.
  • aslkalska 1 hour ago
    they all rent compute from each other
  • fidla 1 hour ago
    ChatGPT is up
  • morkalork 1 hour ago
    Didn't SpaceX overbuilt infra and leases it out Anthropic? I f their dc goes down it probably takes a chunk out of Claude's capacity before even considering the flood of users switching over
  • Razengan 1 hour ago
    SkyNet is arming..
  • misano 1 hour ago
    The IRGC has cut the fiber-optic cables in the Strait of Hormuz. LOL
    • CamperBob2 21 minutes ago
      That's the Strait of Trump to you, peasant
  • ratelimitsteve 1 hour ago
    everything in this thread is raw speculation, obv, but if i had to put money on anything i'd say this is a left-pad incident. some piece of something or other that all of these services happen to depend on went down. Second most likely seems to be some random failure of one leading to an unexpected traffic spike in others, though it seems like we've been talking about automated scalability in web apps for so long that there should at least be a response to, if not a solution for, this sort of problem.
  • pwyq 1 hour ago
    [dead]
  • kregasaurusrex 51 minutes ago
    My guess is someone pulled the switch to go back to the Dark Ages. [0]

    [0] https://www.youtube.com/watch?v=YCzitO446ZY