The headline number is a vendor estimate. The playbook is the story.

LegalOn Technologies, a Tokyo-based legal technology company, published a cost-management account on OpenAI's startup site on October 8, 2026 describing how it reduced its estimated daily costs of running OpenAI's Codex coding assistant by approximately 65 percent without, it says, slowing development. The method is not a technical trick. It is a governance system: three models tiered by task complexity, Fast mode restricted by default, and budget caps set at the department, group, and individual levels.

Read this piece as analysis, not news verification. Every quantitative figure below is vendor-reported and self-described. The word "estimated" appears in the source itself, and it matters: these are LegalOn's own estimates of daily costs, not audited invoices reviewed by anyone outside the two companies. What is genuinely documented is the structure of the controls, and that structure is the part other organizations can actually learn from.

What the primary source actually documents

The company began, in its telling, the way most adoption stories begin: developers had unlimited access to its main model, GPT-5.5 in Fast mode, and used it across design, implementation, and everyday work. Then the arithmetic intervened. Continuing to use high-performance models without limits, the account states, would inevitably exceed the annual budget.

LegalOn's answer was to centralize the decision framework and decentralize the decisions. Its AI-powered Development CoE tested and monitored models and shared findings with managers, who passed them to teams, so that each engineer could independently choose the most suitable model for each task. The tier assignments described in the primary source are specific: GPT-6 Luna, the lightest model, handles code implementation with clear requirements and everyday automations; GPT-6.1 Sol supports standard design, data analysis, and document preparation; and GPT-6 Astra, the most capable model, handles advanced judgment including architecture design and orchestration of other agents.

This is the concrete answer to a question that tiered pricing alone cannot answer. OpenAI's tiered lineup, including Sol at a reported one-fifth of Astra's price, only broadens practical access if organizations can operationalize model choice. A price list does not tell a team which task deserves the expensive model. LegalOn's guidelines, backed by administrator settings with monthly usage limits for departments and individuals, attempt exactly that mapping.

Fast mode off by default, and budgets that distinguish mature businesses from new ones

Alongside model selection, the company began restricting Fast mode by default, allowing individual requests only when needed. The source acknowledges this raised concerns about development speed and says teams used approaches such as running tasks in parallel to maintain performance. It also introduced budget caps at the department, group, and individual levels, with an efficiency target the source describes as up to approximately 20 percent for the established business, while new businesses in the launch phase received generous budgets to encourage active use of AI.

That asymmetry is worth pausing on. Cost governance that treats a young product the same way it treats a mature cash generator would quietly tax exactly the experimentation the company says it wants. LegalOn's stated aim, per the account, was to direct resources toward businesses it wanted to grow. In my reading, this is the most transferable idea in the piece: a budget cap is a policy instrument, and like any instrument its value depends on calibration. A flat cap is a blunt constraint; a staged cap is a portfolio decision.

Measuring ROI per feature release: a metric being built, not built

The company's Senior Engineering Manager, Yuta Tokitake, is quoted in the OpenAI-published profile questioning whether speed itself proves anything.

"It is almost a given that AI can accelerate system development and updates," says Tokitake. "What we really want to understand is whether that faster development actually translates into value for customers."

Yuta Tokitake, Senior Engineering Manager at LegalOn Technologies, quoted in the OpenAI-published company profile "LegalOn halves Codex costs while maintaining development speed," published October 8, 2026, on openai.com. This quotation reaches readers through an OpenAI-published company profile, not an independent transcript or recording, so wording is exactly as published by OpenAI rather than independently verified.

In the same source, Tokitake explained that while AI use lets the company track the cost of individual tasks, the total cost of a complete piece of work, a feature release, often remains a black box. He said the company therefore designed its measurement around each feature release as a single unit, aiming to link the customer value a release delivers to the actual AI costs invested in it so the return on investment becomes visible and accurate. (Paraphrase with attribution; the corresponding passage in the OpenAI profile is a verbatim quotation that contains em dashes, which this article's style prohibits outside unaltered source wording. The verbatim fragment "often remains a black box" is quoted exactly as published on the OpenAI page.)

The metric is described as in development: the company says it is currently building the pipeline to calculate it, and that once complete it will make the true ROI in terms of customer value visible. That means the most interesting claim in the piece, that AI spending can be tied to customer value per feature release, is a plan, not a result. The distinction matters. Observed and documented in the source: the tiering structure, the administrator controls, the budget caps, and the existence of a metric under construction. Vendor claim: the 65 percent and 20 percent figures. Not yet demonstrated: any link between AI cost and delivered customer value.

The governance balance, in the company's own words

The piece closes with a second Tokitake quotation that frames the governance philosophy.

"Excessive restrictions through rules and budgets can undermine an organization's momentum. What we need is a flexible operating model that effectively balances risk control with the speed teams need."

Yuta Tokitake, Senior Engineering Manager, LegalOn Technologies, closing quotation in the OpenAI-published profile, October 8, 2026, via openai.com. Wording exactly as published by OpenAI.

My opinion, clearly marked as such: this is the correct instinct and the hardest part to copy. Every organization adopting AI coding tools faces the same failure modes in both directions. Unlimited access to the most capable model produces cost blowouts and habituates teams to spending that cannot survive a budget review. Blanket restrictions, conversely, push engineers back toward unassisted workflows and erase the productivity gains that justified the adoption in the first place. The tiering-and-caps approach LegalOn describes is a genuine third path, but it requires something most companies lack: a central function, here the AID CoE, with the credibility and time to test models and publish selection criteria engineers actually trust. Without that, tier guidance decays into either paperwork or ritual.

What this source cannot tell us, and why it matters for access

It is also worth stating what this source cannot tell us. The 65 percent figure is an estimate of daily costs calculated by the vendor's customer, published by the vendor, with no methodology, baseline period, or audit disclosed. The 20 percent efficiency target for mature businesses is a target, not a measured outcome. Development speed being "maintained" is the company's own assessment. None of this means the numbers are wrong; it means we do not know how right they are. An organization considering copying this playbook should treat the structural elements as the verified content and the savings figures as a hypothesis to test against its own invoices.

One provenance note for transparency: my assignment brief referenced a stored capture of this page under one hash, while my retrieval on October 9, 2026 returned a different hash, consistent with a page update or re-capture between storage and retrieval. The text I verified is the one cited here, retrieved October 9, 2026 from the primary URL. If the page changes materially after this date, the quotations and figures should be re-checked against the archived version.

The larger access question sits behind all of this. Tiered pricing is often framed as a democratizing move: cheaper tokens mean more organizations can afford capable AI. But price is only one layer of access. The other layer is organizational capability: who decides which task deserves the expensive model, whether engineers have the information to make that call, and whether budgets are structured to protect experimentation. LegalOn's account suggests that for a small-to-mid company, that second layer is now the real bottleneck. Capability arrived. The pricing structure arrived. What most organizations have not yet built is the governance to use either well.