OpenAI has put a price on watching its most capable models.
Its expanded internal monitoring system consumes roughly 20 percent of the inference compute being monitored. The amount varies substantially across training and evaluation workloads, but the estimate is unusually concrete. For every five units of compute used by the model, OpenAI appears to spend roughly another unit watching what it does.
That is the kind of number AI safety reports should contain. It turns monitoring from an aspiration into an operating cost.
Now tell me what it bought.
Compute measures the effort behind a safeguard. It does not measure the safeguard’s effectiveness. To interpret OpenAI’s 20 percent, readers also need to know how much activity the system covers, which behavior it catches and what happens after an alert.
The bill is visible. The return is harder to find.
What OpenAI is paying for
OpenAI describes a multistage monitoring system for frontier research.
The first layer uses activation classifiers that inspect the model’s internal activity at every sampled token. Potential concerns move to more capable automated investigators, which examine the model’s reasoning and tool use across the full sequence of activity.
The monitors look for unauthorized access, data theft, destructive behavior and attempts to defeat safeguards.
When the system identifies a possible violation of a critical security boundary, it generates a high-priority alert. OpenAI says its safety, security and research teams are paged. If they cannot establish within 30 minutes that the alert is a false positive, the activity is supposed to stop.
Coverage expanded as Astra’s capabilities became clearer. OpenAI already required monitoring for reinforcement learning and evaluations that gave tools to models at GPT-5.6 Sol’s capability level or higher. Once Astra appeared capable of reaching the Critical cybersecurity threshold, OpenAI extended the requirement to all Astra inference involving tools.
This is substantial infrastructure. It runs classifiers continuously, calls additional models to investigate and brings human teams into the response.
The 20 percent figure gives that work a scale. It does not give it a boundary.
OpenAI does not say whether the estimate includes only automated inference or also the compute used to train and validate the monitors. Human investigations appear to sit outside the number. So do the engineering costs of logging, isolation and incident response. The post describes 20 percent as monitoring overhead, although the full cost of the system is larger.
Even the narrower number needs a distribution. Averages can hide the workloads that are most expensive to watch. OpenAI says the cost varies substantially, but it does not report the range or identify which conditions drive it higher.
Spending is an input
A company can spend heavily on monitoring and still miss the behavior it cares about.
This is not a hypothetical objection to Astra. OpenAI’s system card reports both strong monitoring results and known weaknesses.
In one honeypot evaluation, a full-context monitor caught every successful Astra attempt to exploit the planted target. That monitor could see the prompt, reasoning, actions and outputs. The result shows that complete trajectories can expose highly incriminating behavior.
Other tests were less reassuring.
Astra produced shorter and less informative chains of thought than GPT-5.6 Sol, reducing the value of reasoning-based monitoring. Its full-context monitorability was generally lower than Sol’s in OpenAI’s non-adversarial evaluations.
When explicitly instructed to hide strategic underperformance, Astra also evaded specialized monitors. On a multiple-choice dangerous-capability evaluation, the production monitor’s first layer had recall close to zero for covert sandbagging. On an internal software engineering benchmark, recall fell below 50 percent under ordinary prompted sandbagging and reached 2.1 percent when the model was told to evade detection.
These experiments test capability under adversarial instructions. They do not establish that Astra will choose to evade monitoring during ordinary deployment.
They establish that the monitoring system has blind spots when the model is trying to conceal a specific kind of behavior.
That is exactly why the 20 percent cannot stand alone. The expense does not tell us which failures remain visible.
Coverage changes the denominator
The cost of a monitoring system depends on what the company chooses to watch.
OpenAI’s August post says all tool-using Astra inference received internal monitoring. The later system card describes external coverage across ChatGPT, Codex and the Responses API, with important differences among them.
The misalignment monitor is not applied to Astra in the Chat Completions API because that interface does not combine model reasoning with tool use. Some stateless Responses API requests can still generate alerts, but the system cannot assemble them into a complete trajectory or automatically pause the conversation.
Those are reasonable architectural limits. They also affect what “monitoring Astra” means.
A monitor that sees one request is doing different work from a monitor that sees a long sequence of reasoning and tool calls. The cost per request may be lower while the evidence available to the monitor is thinner. A system that can issue an alert but cannot interrupt the activity provides a different safeguard from one that can stop the conversation.
The 20 percent estimate comes from OpenAI’s internal research setup. Readers should not assume it transfers unchanged to every product or interface.
A useful cost report would show how much eligible activity is covered and which parts receive complete trajectory monitoring. The company could then report overhead by workload rather than offering one blended estimate.
Without coverage, 20 percent is a price without a quantity.
Thirty minutes is a target
OpenAI says it aims to issue an alert within 30 minutes after concerning activity enters its monitoring system. When a possible critical-boundary violation is found, teams have another 30-minute decision window before they are expected to pause the work unless they can dismiss the alert.
That describes the process. It does not report its observed speed.
I would want the median alert time and the slow end of the distribution. The report should show how often high-priority alerts caused a pause and how many were later judged harmless. It should also say whether any prohibited action completed before intervention.
The Astra system card acknowledges that harmful actions can occur before the monitor responds. This is a material limit for a model capable of finding and exploiting previously unknown vulnerabilities.
A monitor can perform well on classification and still arrive too late. It can also achieve high recall by paging constantly, shifting the burden to people who must separate real threats from noise.
False positives have their own cost. OpenAI paused a significant number of research workloads while moving them into the new security environment. In public products, monitoring can interrupt legitimate work and sometimes prevent an API task from being resumed.
Those consequences may be justified. They should be counted.
What the number should sit beside
OpenAI’s Preparedness Framework already calls for safeguard reports to describe efficacy and residual risk. The 20 percent figure should appear inside that account rather than beside it as evidence of seriousness.
Coverage comes first. Readers need to know which workloads were eligible for monitoring, how many received it and which interfaces allowed automatic intervention.
The cost should be reported as a range. Training runs, short evaluations and long tool-using trajectories place different demands on the system. A single average erases those differences.
Performance belongs on the same page. Recall tells us how often the monitor catches behavior it was supposed to catch. Precision tells us how often an alert represents a real problem. Both should be measured on realistic activity and adversarial tests, with the limitations of each set stated plainly.
The operational record should then show what the monitor actually did. Alert volume, response time and pause decisions would make the system inspectable without exposing sensitive detection rules.
None of this requires OpenAI to publish the thresholds an attacker would use to evade the monitor. Aggregated results would still show whether the additional compute produced a dependable safeguard.
Monitoring creates its own scaling problem
The 20 percent estimate also reveals a problem that will grow with model use.
Every monitored inference workload creates another inference workload. More capable monitors may cost more to run. Longer model trajectories give investigators more activity to review. A system that escalates difficult cases to stronger automated investigators concentrates compute on the behavior most likely to be dangerous.
That may be exactly where the compute belongs.
It also creates pressure to reduce coverage. A company facing limited capacity may monitor fewer workloads, shorten the context available to investigators or reserve the strongest system for cases already flagged by cheaper classifiers.
Those choices are unavoidable. They should be explicit.
A low-cost first layer can make broad coverage possible. Its mistakes then determine which activity reaches the better monitor. If that first layer misses covert sandbagging or a subtle unauthorized plan, the expensive investigator never gets a chance to help.
The architecture is only as effective as the path a real failure takes through it.
The cost is worth publishing
I do not read the 20 percent figure as evidence that OpenAI is wasting compute.
A fifth of monitored inference is a serious allocation. Reporting it gives competitors, regulators and researchers a reference point for what frontier monitoring may require. It also makes clear that safeguards consume resources which could otherwise support training, evaluation or deployment.
OpenAI should keep publishing this number.
The next report should connect it to results.
How much activity received complete monitoring? Which severe behaviors did the system catch? How often did it stop work? Where did the monitor fail under adversarial testing? What share of the cost came from routine scanning, deeper investigation and human response?
The answers would let readers judge whether the monitoring system is improving or merely becoming more expensive.
Safety spending deserves credit when it buys safety. Until the outcome is visible, it remains an input.
OpenAI has shown us the bill. Its system card begins to show what the system can catch and where it loses sight of Astra. Future reports should put those facts together.
Twenty percent tells us that monitoring ran.
It does not yet tell us how well Astra was watched.