AI Infrastructure

OpenAI details how Habitat storage scaled for more than one billion ChatGPT users

OpenAI described the evolution of Habitat, its online storage platform, from a Python client library into a massive service layer handling tens of millions of requests per second.

Published Updated
OpenAIChatGPTAI InfrastructureHabitat

OpenAI has published a detailed engineering account of how its online storage platform, Habitat, was reshaped to support the growth of ChatGPT, Codex and its API products. The September 11 article gives a rare look at the infrastructure pressure behind consumer AI services that now operate at internet scale. According to OpenAI, Habitat handles more than 70 million requests each second, supports products used by more than one billion people each week and manages more than 500 petabytes of data across almost 40 geographic regions.

The system began far more modestly. Habitat was originally built around DevDay 2023 as a Python client-side library connected to OpenAI's main product data stores. Its purpose was to spare product teams from repeatedly solving the same storage questions: where a piece of data should live, which request had permission to read it, how data residency rules should be applied and how traffic should be shaped. As ChatGPT usage expanded and product teams multiplied, that library model became hard to operate. Any protocol change required coordinating updates across many services, which slowed reliability work and made regional failover changes brittle.

OpenAI's answer was to pull Habitat into a central service. That move gave the company one place to enforce routing, authorization, rate limits, audit logging and access controls rather than scattering those decisions across clients. It also made storage behavior more predictable. Habitat deliberately avoids letting callers run arbitrary analytical queries against hot online stores. Instead, it exposes constrained object and edge operations that keep request cost clearer and reduce the chance that a single expensive query can overwhelm production systems.

The article is also candid about the limits of Python at this scale. OpenAI says the team chose a Python service as a temporary tradeoff because it let engineers stabilize the platform quickly, even though Python added latency and compute overhead. Tail latency became a constant theme. CPU-heavy work such as routing, compression, encryption, checksum handling and feature-flag parsing could delay asyncio scheduling even when downstream storage responded quickly. The team reduced those spikes by changing configuration refresh behavior, limiting per-process concurrency and spreading load across many workers.

Connection management produced another lesson. During traffic bursts, OpenAI found that LIFO connection reuse could feed more requests back to already overloaded server processes, creating a metastable failure pattern. A switch to FIFO reuse helped break that loop, while Envoy and HTTP/2 multiplexing reduced pressure on downstream services. These are not glamorous product features, but they are the machinery that determines whether AI tools feel instant or fragile to users.

The most consequential part of the post is the migration path. OpenAI says Habitat's Python service once handled more than 20 million requests per second at peak, but the company has now rewritten the service in Rust. In the second quarter of 2026, two engineers, working with Codex and GPT-5.5, rewrote the service; OpenAI says the Rust version now carries 95 percent of production requests and is six times more CPU efficient and fifteen times more memory efficient than the Python system. The story shows how AI scale is forcing companies to rebuild basic application infrastructure, and also how the same AI tools are beginning to assist the engineers doing that rebuilding.