The unpublished internal app is not blocked because someone proved it dangerous. It is blocked because nobody can say it is safe.

Why AI Agent Sandboxing Matters Now

Every enterprise that has let employees build their own agents has met the same queue. A business analyst assembles an agent in an afternoon. It touches a spreadsheet, a ticketing system, maybe a customer extract. Then it waits for review, and the review has no natural end date, because the reviewer is being asked to certify the absence of risk, which no one can do from a diagram.

Block-until-ready sounds conservative. In practice it is an open-ended position with a carrying cost: the work does not stop, it moves somewhere you cannot see. People route around the queue with personal accounts and unsanctioned tools, which is the BYOAI problem in its natural habitat. The review queue does not remove the exposure. It relocates it off your books.

Signal Figure Source
Time for a licensed user to build and publish an agent-made app internally Roughly two weeks Ospiri field notes, see the Copilot velocity gap post
Typical change-management review for the same app Two to three months Ospiri field notes
Reviewer’s actual question “Can you prove it is safe?” Unanswerable from a design document

The gap between a two-week build and a three-month review is a structural mismatch, not a staffing problem. Hiring more reviewers shortens the queue and leaves the question unanswerable.

Two Defaults, Priced Like Positions

Let’s step back and treat the default state as a position you hold on every unreviewed agent. There are two to choose from.

Dimension Block until ready Safe state until approved
Default state of a new agent Cannot run Runs, contained
What the reviewer must prove That it is safe That it should be promoted
Cost of a wrong approval Full exposure, immediately Bounded by the sandbox until promotion
Shadow workaround incentive High: the queue is the enemy Low: the sanctioned path is already working
Reversibility Low: once promoted, hard to claw back High: state changes are staged and revocable
Time to first value Months Days

The first model charges the business a certain, recurring cost (delay and shadow usage) to avoid an uncertain loss. The second accepts a small, capped exposure up front and buys an option: the right to promote, or to discard, with evidence in hand. A trader would recognise this immediately. One is paying full premium to hold no position. The other is holding a position with a stop-loss.

Anatomy of a Review Queue That Never Drains

The failure is rarely a single bad decision. It is a pattern that repeats:

  1. The agent is submitted with a description, not a behavior. The reviewer reads intent, but risk lives in what the agent actually reaches.
  2. The reviewer cannot bound the blast radius, so every unknown becomes a blocking question: which tables, which folders, which hosts, which skills files.
  3. Questions go back to the builder, who answers from memory, and the answers cannot be verified without running the agent.
  4. Running the agent requires approval, which is the thing under review. The loop closes.
  5. The submitter gives up on the sanctioned path and ships the same logic through a personal account, where no one is reviewing anything.

Step four is the tell. A process that forbids the one experiment that would answer its own question is not conservative. It is stuck.

The Formula: Exposure While You Wait

Here is the bet, stated as a number you can put in front of a risk committee.

Unreviewed Exposure = (Agents Waiting × Average Wait in Weeks) × (Shadow Migration Rate + Residual Sandbox Risk)

Factor What it measures How a sandbox moves it
Agents Waiting Depth of the queue Falls, because approval becomes promotion, not permission to run
Average Wait in Weeks Months today, days with a contained default Falls, since the reviewer observes behavior instead of inferring it
Shadow Migration Rate Share of waiting work that leaves the sanctioned path Falls sharply, as the safe path is already usable
Residual Sandbox Risk Exposure inside the contained default Small and capped by construction, if the containment is real

The last factor is where honesty matters. A sandbox that is only a policy statement has a Residual Sandbox Risk equal to the original risk. Containment has to be enforced outside the agent’s own process, or the number you just computed is fiction.

What Safe State Actually Requires

“Safe state until approved” is an architectural claim, and it only holds if specific control points exist. This is also where we are careful about scope: a sandbox is a statement about what an agent can reach and change, not a verdict on whether its logic is good.

Control point What it does What it is not
Copy-on-write file access Agent works on a staged copy; originals stay untouched until promotion A read-only block that breaks legitimate work
Egress allowlisting Agent reaches only declared hosts; everything else is observed and held A prompt filter that reads content
Credential scoping Agent sees only the secrets its declared purpose needs Identity issuance alone
Skills and MCP inventory Every capability the agent loads is enumerated at run time A self-reported manifest
Promotion record An auditor-readable trail of what the agent did in the sandbox and who approved the move A ticket comment

The distinction that matters is block-on-deny versus copy-on-write. Block-on-deny stops the agent and asks a human what to do. Copy-on-write lets the agent finish the task against a copy, so the reviewer evaluates a completed, inspectable outcome. That is the difference between a reviewer reading a plan and a reviewer reading a result. We have gone deeper on that tradeoff in the copy-on-write post linked above, and the enforcement layer behind it is what an agent firewall is for.

So, what’s the moral? Review is cheap when it evaluates evidence and expensive when it evaluates promises. A sandbox is a machine for producing evidence.

What CISOs Should Do This Quarter

Step Action Output Effort
1. Measure the queue Count agents and apps waiting on review and their age in weeks A baseline Agents Waiting and Average Wait figure 1 week
2. Find the leak Sample where waiting work is going: personal accounts, unsanctioned tools An estimate of Shadow Migration Rate 2 weeks
3. Define the safe state Write the containment spec: file, network, credential and skills boundaries A one-page spec reviewers can sign 2 weeks
4. Pilot one business unit Run its queue through a contained default, with promotion as the gate Before-and-after wait times and a promotion record First quarter

The Bottom Line

Blocking an agent until someone proves it safe is a standing position that pays premium for no protection, while a contained default with evidence-based promotion caps the downside and shortens the wait from months to days. The reviewer’s job should change from certifying safety to deciding on promotion, and that only works if the sandbox is enforced below the agent, not declared by it. Governance that cannot say yes quickly gets routed around, and what it loses is visibility, not risk.

If your team is sizing this for the next budget cycle, request a working session. We will walk through your current review queue, map where waiting work is leaking to unsanctioned tools, and scope a contained-default pilot for one business unit. Plan on 90 minutes for the session and the first quarter for the pilot.