The reporting

1 articles

Topic: latency

Closeup photograph of the rear of a server rack at the NERSC data center, showing blue LED screens and network cables on dense compute hardware.

NEWS AI Infrastructure

OpenAI's GPT-6 caching stores more of what an AI already read, cutting repeated input costs by up to 90% and trimming wait times

OpenAI's September 22, 2026 announcement details a rebuilt caching system for GPT-6 that stores more of the context an agent already processed, discounts reused input tokens by up to 90 percent within a 30-minute window, and gives developers new dashboards for diagnosing cache misses. Here is what that means for the cost and speed of long AI sessions.