AI security began as two activities. Test the model before you ship it. Filter the prompts and responses once it is live.
Both still matter. Neither answers the question a CISO now gets asked after an agent has done something in production: who authorized that, under what policy, and can we prove it?
That question is why the market is reorganizing. Gartner’s September 2026 Emerging Market Quadrant for AI Application Security describes technologies that combine security testing, exposure management and runtime defense to detect, alert on or block threats against enterprise-built AI applications and agents, and identifies three core capabilities: automated discovery and inventory, automated security testing, and runtime defense. Vendors, including us, now describe some version of a discover, test, protect loop.
Drawing the loop is easy. This post is about what it takes for the loop to be real.
Six questions, in order
An organization running its own agents has to answer six connected questions.
- What AI systems, agents, tools, data connections and delegated authorities exist?
- How do those systems behave under expected conditions?
- How do they fail under adversarial pressure?
- Does a changed configuration hold up under attack before it reaches production?
- What is allowed to execute right now, when an agent proposes a consequential action?
- What happened, and can someone outside the vendor verify it?
Each question has a stage of the lifecycle attached to it. Discovery tells the organization what exists. Evaluation shows how systems behave. Red teaming exposes how they fail. Sandbox testing proves a configuration before it ships. Runtime authorization determines what may execute. Evidence proves what was tested, permitted, denied, masked or escalated.
The order matters, and so does the direction of flow. Each stage should produce something the next stage consumes.
Discovery is the beginning, not the deliverable
You cannot secure what you cannot see, so the lifecycle starts with an inventory: the agent frameworks in use, the MCP servers and the tools they expose, local model runtimes, coding and CLI agents on developer machines, provider API keys sitting in environment files, and the vector stores and databases the code can reach.
A good inventory is honest about its own gaps. A scan that could not complete should say so rather than report a clean result. Secrets should be fingerprinted, never collected. The sensor should read, not modify.
But an inventory, however honest, does not tell you whether an agent will behave safely when it is handed an unexpected instruction, a poisoned document, or more authority than its task requires. It tells you where to look. The next stages tell you what you will find.
Evaluation and red teaming show behavior, not intent
An agent is a combination: this model, this system prompt, these tool schemas, this business function. Testing the model alone tests a component. The risk lives in the combination you shipped.
Evaluation asks how that combination behaves under the conditions you expect. Red teaming asks how it behaves under the conditions an adversary will create: instructions smuggled into retrieved content, persuasion campaigns aimed at the agent’s goal, tool calls with over-broad arguments, delegation beyond the agent’s authority, attempts to talk the agent out of its own constraints.
Two design choices separate useful red teaming from theater. First, attack a faithful clone in a sealed sandbox with mocked tools, so nothing the attack achieves touches a real system, and so a new skill can be detonated against canary tokens before it ever runs live. Second, judge the outcome from what the agent actually did, not from what it says about itself. An agent that reports “I refused” while its tool call went through has failed the test.
Sandbox testing turns findings into decisions
A red-team finding is only useful if it changes something. The change might be a stricter policy, a narrower tool surface, a different system prompt, a different model, or a new skill the agent has picked up. Before any of those reaches production, the changed configuration should face the same adversarial cases in the same sealed sandbox, and a new rule should be simulated against the agent’s real traffic before it takes effect.
This is the step that makes AI security an engineering discipline rather than a periodic review. The question shifts from “did we test it” to “did this version hold, against which cases, and what would this rule have done to last week’s traffic,” and the answer is observed rather than asserted.
Production changes the question
A configuration that survives the sandbox still needs a control in production, because production is where the agent meets the systems that matter: customer records, repositories, deployment pipelines, payment rails, privileged APIs, other agents, and the outside world.
The production question is not “is this agent secure.” It is narrower and harder:
What should happen when this agent attempts this action against this target, right now?
Detection after the fact does not answer it. An observer that sees the export after it left the building has produced an incident report. The answer requires an authorization point on the path between the agent’s proposal and the side effect, evaluated before the side effect occurs.
At that point the decision can take into account which agent is acting, what class of operation it is attempting (a read, a write, a delete, an administrative change, a funds transfer, a delegation), the actual target and arguments, the credentials and identity in play, the agent’s trust tier and business function, and the delegation grant it is acting under if another agent handed it the task. The outcome is one of four: permit the call as proposed, mask sensitive arguments and let it proceed, deny it and tell the agent why, or pause it for a named human to approve the concrete action in the shape it will execute.
Three properties make such a control worth trusting.
It has to be on the dispatch path. A callback that observes a tool call cannot stop it; a control that wraps the dispatch method evaluates before the function body runs.
It has to fail in a known way. Buyers should be able to read, in the vendor’s documentation, what happens when policy cannot be evaluated, when an approval expires, or when an action matches no rule. Where it is supported and configured, the control should fail closed. Where it does not, the vendor should say so.
It has to start in monitor mode. The first weeks of enforcement should record what would have been decided, simulate new rules against real traffic, and refuse to enable a rule that would block more than a small fraction of legitimate work. Enforcement is then turned on per agent and per operation, beginning with the actions that cannot be undone.
Evidence closes the loop
Every stage should leave evidence, and the evidence should not depend on trusting the agent or the vendor.
A runtime decision should produce a signed receipt that binds the request, the outcome, the policy version and the approval, if there was one, to the execution event. Receipts should be batched into a tamper-evident structure, timestamped by an independent authority, and verifiable offline with standard cryptographic libraries. A vendor who publishes the verification procedure, and states plainly what the receipts do not prove, has given the buyer something an audit log cannot: a way to check.
Evidence also has to flow backward. A red-team verdict should become a proposed rule that a reviewer accepts or rejects, in the same policy language the runtime enforces. A runtime anomaly should become a draft rule with a dry run over the agent’s own recent decisions. That is the difference between a dashboard that displays the lifecycle and a control plane that runs it.
How VisIQ connects the stages
VisIQ Labs builds the lifecycle around the runtime decision, because that is the point at which AI behavior becomes a real system side effect.
- Discover. A read-only sensor inventories agent frameworks, MCP servers and their tool surfaces, local model runtimes, coding agents, provider keys and reachable data stores, reports its own coverage gaps, and marks each agent as ungoverned, harnessed or governed.
- Evaluate. Every new agent is provisioned in monitor mode, so every tool call and the decision it would have received is recorded before anything is enforced. New and edited rules are simulated against recent real traffic, with the outcome distribution shown before a rule takes effect.
- Red-team. VisIQ clones the agent into a sealed sandbox with mocked tools and drives prompt injection, social engineering and tool abuse against the clone, with each technique carrying its MITRE ATLAS id. New skills are detonated against canary tokens before they run live. Verdicts are recomputed from the agent’s actual behavior. A finding becomes a proposed rule that nothing applies until a reviewer accepts it.
- Test before rollout. The clone is built from the combination you actually ship, this model, these tool schemas, this system prompt, so a changed configuration is attacked as shipped. A skill can be held in a fail-closed state until it clears the test cell.
- Authorize. The VisIQ harness wraps the framework’s tool dispatch, so each action is evaluated in-process before the tool function runs, against a locally cached policy bundle. Outcomes are permit, mask, deny or approval required; retrievals are allowed, redacted, denied or escalated for review. Agent-to-agent delegation is governed as its own authority decision: a child agent’s power is the parent’s authority intersected with an explicit, revocable grant, and revocation cascades. The failure behavior is published.
- Prove. Every material decision emits a signed receipt, batched into a Merkle tree, sealed with a hardware-held key, countersigned by an independent timestamp authority, and verifiable offline. The docs state what the receipts do not prove.
The same rule shape governs actions, retrievals and delegations, and each receipt names the policy version behind the decision. That shared vocabulary is what lets a finding in one stage become a control in the next without a translation layer in between.
The boundary
Runtime authority applies where the action passes through the governed path. It does not replace identity and access management, network segmentation, secrets management, admission control, endpoint detection, or the application security work you already do, and a mature deployment documents both its governed paths and the paths that bypass the agent framework. A vendor that will not have that conversation is selling a dashboard.
The question to ask
The buying question is no longer “can this platform detect AI risk?”
It is: “Can it discover our agents, evaluate and red-team them, prove a changed configuration before rollout, decide what executes in production, and prove what happened, in a way our auditors can check without calling the vendor?”
That is the AI application security lifecycle. VisIQ was built to run it end to end, with the runtime decision as the control point.
See it on your own agent. Install the harness in monitor mode and watch the decisions your agent would have received, no card and no sales call: docs.visiqlabs.com. Prefer a guided review? Score your current deployment with the 30-question scorecard, then request a technical review and bring one agent and one consequential action.
Product statements describe VisIQ as documented at docs.visiqlabs.com on the publication date. Gartner does not endorse any vendor, product or service depicted in its research publications. VisIQ Labs has not been evaluated in Gartner’s Emerging Market Quadrant for AI Application Security, and nothing here should be read to imply otherwise.