Product

Runtime Enforcement Is Table Stakes. Can You Prove What Your Agent Was Allowed to Do?

Blocking a bad action is now the baseline for AI agent security. The harder question is whether you can show, after the fact, what the agent was authorized to do and what the control decided before it acted.

Otis

The first generation of AI security programs asked a set of inventory questions. Which models, agents and copilots exist? Who built them? What data can they reach? Can we red-team them before deployment? Are they vulnerable to prompt injection, data leakage or tool abuse?

Those questions still matter. Discovery, testing and model evaluation remain necessary parts of a serious program. But they are no longer the whole control story.

As agents move from answering questions to taking actions, the decisive question becomes: what happens immediately before an agent causes a consequential side effect?

Runtime enforcement is now the baseline

The market has answered that question with runtime enforcement, and it has answered it in nearly the same words.

Zenity says it “secures long-horizon AI agents at the decision, the exact point where context and intent come together,” and its runtime engine “decides to let it through, block it, or shut the agent down.” Noma’s runtime layer “alerts, blocks, masks data, or routes to a human,” with access policies enforced “before actions are carried out.” Check Point, which completed its acquisition of Lakera in October 2025, describes “real-time runtime enforcement” alongside continuous red teaming. HiddenLayer’s AI Runtime Security offers “inline enforcement” that will “detect, block, or redact unsafe actions.” Palo Alto Networks, which absorbed Protect AI into Prisma AIRS, describes its AI Gateway as “a single enforcement point.”

Gartner’s September 2026 Emerging Market Quadrant for AI Application Security named Zenity and Noma Market Shapers among startup vendors. And in a press release dated 26 August 2026, Gartner predicts that “over half of successful cyberattacks on AI agents are expected to exploit access control weaknesses and prompt injections by 2029.”

So runtime enforcement is no longer a differentiator on its own. Every serious vendor claims it, most with overlapping verbs: block, mask, alert, route to a human. The next question is whether the enforcement decision is explicit, portable and verifiable.

Detection is not authorization

Much of the market frames the runtime problem as comparing an agent’s intent with its behavior. That is a detection problem, and a useful one. Detection asks whether an event looks suspicious.

Authorization asks a different question: is this specific agent permitted to perform this specific operation, against this specific target, with these arguments, under this identity, right now?

Those questions are related, but they are not the same control. Detection scores likelihood. Authorization returns a decision.

Consider an engineering agent with access to a repository and a deployment system. The agent may be allowed to read source code, run tests and open a change for review. The same agent may not be allowed to merge to a protected branch, read production secrets, change deployment configuration or release directly to production.

A monitoring product may flag unusual behavior. An observer may log the call. But a production-grade authorization layer must make the decision before the action executes, and that decision must be understandable to the operator and enforceable by the runtime.

Gartner made the same distinction in a 26 May 2026 press release on agent governance: “Failures are most likely to occur when organizations fail to distinguish between an agent’s ability to act and the scope of access it is granted.” Access lets an agent in. Authority determines what it can do.

The control point that matters

The useful architecture is straightforward.

  1. An agent proposes a tool call, a retrieval, a delegation to another agent, or some other consequential action.
  2. The authorization layer receives the proposed action and its available context.
  3. Policy evaluates the agent, its identity, the target, the requested operation and its arguments, and the relevant conditions.
  4. The runtime returns an explicit result: Permit, Mask, Block, or Escalate to a named human.
  5. Only a permitted, masked or approved action proceeds.
  6. The decision is preserved as evidence.

The important detail is not merely that policy exists. It is that policy sits on the execution path. That is the line between guidance and control.

VisIQ’s position: authorization that can be examined later

VisIQ Labs is built around AI agent security and pre-execution enforcement. Its operating model is intentionally concrete: Permit. Mask. Block. Escalate. Prove.

In the product those outcomes are permit; mask for tool arguments and redact for retrieved documents; deny; approval_required, which pauses the call for a human; and signed decision receipts. On the website the same action reads permitted, redacted, refused, or held for a human.

The VisIQ harness wraps the agent framework’s tool dispatch, so every call is evaluated before the tool function body runs, in-process, against a locally cached rule bundle. A deny is returned to the model as the tool’s output, so the agent can adapt rather than crash. Sensitive actions can be denied outright or paused for a named person, who is “deciding one concrete action in the shape it will execute.” Where the retriever is directly accessible to the harness, retrieval governance evaluates each returned document and allows, denies or redacts it before the agent reads it. On the retrieval path an escalation is recorded for review without pausing, because a retrieval cannot wait on a human.

Every decision, including each human grant or rejection, is recorded into the same signed decision chain, “cryptographically tied to the execution event it governed.”

This creates a broader control story than a dashboard or a threat score:

  • What was the agent trying to do?
  • What was it allowed to see?
  • Which policy version evaluated the request?
  • Was the action permitted, masked, blocked or escalated?
  • Did a human approve it, and what exactly did they approve?
  • What evidence remains after the event, and who can verify it?

None of this replaces network controls, identity systems, sandboxing or application security. A runtime authorization layer works alongside those controls and governs the action paths that cross it.

Three layers of proof

Action proof

The runtime should preserve the decision attached to the attempted action, not a later summary of what the system believes happened. For example:

  • get_account: permitted
  • place_hold: human approval required
  • block_card: denied

That is more useful than a statement that the agent was “monitored.” In VisIQ’s model, each enforcement receipt “names the posture, policy version, and enforcement basis behind the decision it records.”

Context proof

An action cannot be evaluated well if the model has already absorbed context it should never have seen. As VisIQ’s retrieval-governance paper puts it, “once prohibited material reaches the model’s context window, the application has violated the intended data boundary.”

Retrieval governance is a second boundary: for retrievers directly accessible to the harness, each document is allowed, denied or redacted by classification and trust tier before it becomes part of the agent’s working context. It complements source-system permissions rather than replacing them. The security question is therefore both: what may the agent do, and what may the agent know before deciding what to do?

Note what this is not. It is not a classifier guessing whether a prompt is malicious. Detecting injection is probabilistic; a classifier will sometimes be wrong in both directions. Authorization at the retrieval and the tool call is deterministic: injected intent cannot become unauthorized access or an unauthorized action without an affirmative decision.

Decision proof

A dashboard is useful for operations. It is not the same thing as independently reviewable evidence.

VisIQ’s receipts are Ed25519 signatures over the canonical decision payload, batched into Merkle trees, with each batch root countersigned by an RFC 3161 timestamp authority VisIQ does not control. A receipt verifies offline with standard cryptographic libraries and its own fields; no call to VisIQ is required. Signing happens after the decision, off the hot path, so the agent never waits on cryptography.

The boundary is published with the design. The guarantee is tamper-evidence against outside parties and against accident. A verifier who needs protection against the operator itself holds a checkpoint of their own. And a receipt attests to decisions made on the instrumented path, not to actions taken outside it.

That precision is what makes the evidence useful during incident response, internal control reviews, regulated oversight, customer assurance, and disputes about whether a control actually operated.

Portability matters more as agents multiply

Organizations rarely standardize on one agent framework. They may run an internal orchestration layer, a coding assistant, a customer-support agent, a data-analysis agent and agents embedded in vendor platforms. If every environment requires a different policy language, approval model and evidence format, security fragments again.

VisIQ’s harness is one call where the agent is constructed, and the same rules, approvals and receipts apply across the supported frameworks and coding-agent hosts listed in the documentation, “carried unchanged from testing into production.” The implementation surface differs by framework, but the security questions stay constant:

  1. Identify the agent.
  2. Describe the proposed action.
  3. Evaluate policy before execution.
  4. Return an explicit decision.
  5. Preserve the evidence.

The same contract extends to agents that delegate to other agents. A child agent’s authority is the parent’s authority intersected with an explicit, revocable grant, and revoking a grant cascades to everything beneath it.

Portability does not mean claiming universal support. It means treating authorization as a reusable contract rather than a feature trapped inside one vendor’s runtime.

What buyers should ask every AI security vendor

Before buying another AI security dashboard, ask:

  1. Can the product make a decision before a consequential tool call executes?
  2. Can it distinguish read, write, delete, export, administrative and delegation actions?
  3. Can policy use agent identity, target, arguments, trust tier and the delegation grant the agent is acting under?
  4. Can it mask arguments before the tool receives them, and redact documents before the model receives them?
  5. Can it require human approval for a defined class of actions, and what does the approver actually see?
  6. What happens when no rule matches, when masking fails, when an approval goes unanswered, or when the control path is unavailable?
  7. Does the product govern the actual runtime path, or record activity afterward?
  8. Can the resulting decision be independently verified, and by whom?
  9. Which agent frameworks and tool surfaces are genuinely supported today?
  10. Which unmanaged paths remain outside the control boundary?

The last question is as important as the first. Any runtime control governs only the paths that cross it. An action taken with a raw credential, a direct network call, a compromised workload or an unmanaged third-party system is outside that boundary, whoever the vendor is. Buyers should demand that boundary map explicitly, and vendors should publish it.

The practical starting point

Most organizations do not need to govern every agent action on day one. They need to find the actions that matter most: money movement, customer-impacting writes, production changes, sensitive data export, privileged access, external communications, and agent-to-agent delegation.

VisIQ Lab Services runs that assessment. It maps the agent’s authority surface, identifies the highest-consequence transitions, determines which paths are instrumented, and defines where policy must execute before the side effect. The output is not a maturity score. It is a control map:

  • which actions are permitted
  • which actions require approval
  • which actions must be blocked
  • which context must be redacted
  • which decisions produce verifiable evidence
  • which paths still require network, identity, sandbox or application controls

Runtime security needs a stronger finish line

Discovery tells you what exists. Red teaming shows how an agent can fail. Model evaluation measures behavior under test. Runtime enforcement constrains what happens when the agent acts.

Verifiable authorization answers the question that follows: can you prove what the agent was allowed to do, and what the control decided before it did it?

That is the standard AI security programs should be moving toward.

Map your agent’s highest-consequence actions with VisIQ Lab Services, and find out which decisions are actually governed before execution.

Sources (accessed 25 September 2026). Zenity: zenity.io homepage and Runtime Boundaries page. Noma Security: noma.security homepage and platform page. Check Point: press release “Check Point Acquires Lakera to Deliver End-to-End AI Security for Enterprises,” 16 September 2025; Form 6-K, 2 December 2025. HiddenLayer: AI Runtime Security page. Palo Alto Networks: “Announcing the General Availability of Prisma AIRS AI Gateway,” 16 July 2026. Gartner Press Release, “Gartner Forecasts the Market for Securing AI Will Reach $4.8 Billion in 2027,” 26 August 2026. Gartner Press Release, “Gartner Says Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure,” 26 May 2026. Vendor placement announcements: Zenity and Noma press releases, 23 September 2026. Gartner does not endorse any company, vendor, product or service depicted in its publications. VisIQ Labs makes no claim to inclusion in, or endorsement by, Gartner research. VisIQ product statements are drawn from visiqlabs.com and docs.visiqlabs.com.

More posts

Newest first