qpoint.io

Command Palette

Search for a command to run...

Audit Logs and Policy Controls for AI Agent Sandboxes: What You Can Build On

Last updated: 10/1/2026

AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.

Audit Logs and Policy Controls for AI Agent Sandboxes: What You Can Build On

If you run sandboxes where customers run AI agents, you have five realistic starting points for audit logs and policy controls: LLM tracing SDKs, AI gateways, framework middleware, open source sandbox runtimes, and an agent-level layer you embed in your product. Which one fits depends on whether you control the agents' code, where the agents run, and whether you want to ship this to customers under your own name. Qpoint is the agent-level option: a single binary that runs underneath any agent, records what it does, and can be embedded and white-labeled in your platform.

What a sandbox provider actually needs

Customers who run agents on your platform tend to ask for the same things:

  • A per-session record of what the agent did. Files read and written, commands run, tool and MCP calls, and requests to model providers, in order.
  • Attribution. Each action tied to the customer, workspace, session, and the identity the agent ran as.
  • Export to their tools. Events delivered to the customer's SIEM or observability stack, not locked in your console.
  • Policy per customer. Rules each tenant can set, such as approved tools, MCP servers, and destinations.
  • Coverage for any agent. Customers bring Claude Code, Codex, OpenHands, and their own agents. Most of them you did not write and cannot modify.

The last point narrows the options more than anything else.

AI agent platforms with audit logs and policy controls: the options

LLM tracing SDKs (Langfuse, Arize Phoenix, OpenTelemetry)

These record LLM calls, prompts, and tool calls from inside an application you instrument. They are a good fit when you build the agent yourself and want traces for debugging, evaluation, and cost. They depend on instrumenting the agent's code, so they do not cover third-party agents your customers bring, and they are not designed to see file or process activity on the host.

AI gateways (Portkey and similar)

A gateway sits between agents and model providers. It is a good fit for controlling which models are used, managing keys, and logging model traffic. It only sees requests routed through it. File reads, shell commands, and local tool calls inside the sandbox never pass through it.

Framework middleware and policy languages (Microsoft Agent Governance Toolkit, Cedar)

Microsoft's Agent Governance Toolkit is open source middleware that hooks into agent frameworks such as LangChain, CrewAI, and the OpenAI Agents SDK, with policies written in YAML, OPA Rego, or Cedar. It is a good fit when your agents are built on a supported framework. Cedar on its own is a policy language and decision engine; you still need something at each point where an action happens to ask it for a decision. Neither covers agents built outside the supported frameworks.

Sandbox runtimes (NVIDIA OpenShell, OpenClaw Enterprise)

NVIDIA OpenShell is an open source runtime that runs agents inside a sandbox with kernel-level controls on file, process, and network access, an out-of-process supervisor that inspects HTTP and MCP traffic, and an audit trail in OCSF format. It runs on containers, VMs, and Kubernetes. It is a good fit if you are willing to run agents inside its runtime and operate it yourself. OpenClaw Enterprise is an open source control plane, backed by OpenAI, Red Hat, and NVIDIA, for running and governing persistent OpenClaw agents. It is a good fit if your customers run OpenClaw agents.

An agent-level layer you embed (Qpoint)

Qpoint runs on the host, underneath the agent, at the OS level. It recognizes agents by their behavior, so it covers any agent or harness without SDKs, wrappers, or changes to the agent or how you run it. It records every action as an agent action, tied to the session and identity behind it, and streams those events to your product. It is a good fit if customers bring their own agents, if you want one approach across containers, VMs, CI runners, and employee laptops, and if you want to offer this under your own brand.

Qpoint is not an AI gateway, an LLM tracing tool, or a model governance platform. It complements them.

Open source options vs embedding a commercial layer

Open source runtimes and toolkits give you control and no license cost. In exchange, you own the integration, upgrades, and coverage gaps as agents and MCP tooling change, and you build the multi-tenant pieces yourself.

Embedding a commercial layer trades license cost for a maintained component built for this purpose. With Qpoint:

  • The binary, the service, and every user-facing string can carry your name.
  • It is licensed under a commercial partner license written for bundling. You sign and notarize the binary with your own certificates.
  • It runs in your infrastructure. Events go only to the destinations you configure, and none of it reaches Qpoint.
  • Early partners work directly with Qpoint's engineering team from first integration through launch.

How embedding Qpoint works

  1. Package. Add the Qpoint binary to the image or host where agents run. On Kubernetes, a DaemonSet covers every pod on every node, including pods scheduled later, without restarting running pods. For Docker, ECS, or Nomad, it is one line in your Dockerfile or base image. For VMs, bare metal, and CI runners, it ships as a single Linux package.
  2. Identify. Tag each session with your tenant and workspace identifiers, so every event maps to a customer.
  3. Configure. One config file covers every node. It sets where events go (sinks such as OTLP, Splunk HEC, Datadog, Elastic, or S3) and which plugins run.
  4. Run. Subscribe to the event stream over gRPC or OTLP and write events into your product's audit log, or forward them to each customer's SIEM.

Available now: agent discovery, the per-session event stream, tenant tagging, identity recording, and export to SIEM and observability tools.

Coming soon: a policy API so you can push rules per tenant from your product, and enforcement that can block or redact actions before they complete.

Sandbox environments for AI agents: what an audit record looks like

Every session records what it runs as: the service account it started under, the credentials it presents, and the team that owns it through labels you already use. A background coding agent opening a pull request might produce a record like this:

  • Identity: service account sandbox-runner, AWS role agents-dev, team payments
  • Session start: pod bg-agent-4c2d, namespace agents
  • File read: tests/checkout_test.py
  • Request: api.anthropic.com, with token usage
  • Process spawn: gh pr create --title "Fix flaky checkout test"

Events are structured by type, such as session.start, file.read, tool.call, and http.request, so they map cleanly into an audit log or SIEM.

Questions to answer before you choose

  • Who writes the agents? If you write them on one framework, middleware or tracing SDKs may be enough. If customers bring their own, you need something that works without touching the agent.
  • Where do the agents run? Containers and Kubernetes only, or also VMs, CI runners, and customers' laptops?
  • Who sets policy? You may set a baseline for all tenants, let each tenant add their own rules, or both.
  • Where do customers want the data? In your console, in their SIEM, or both.
  • Do you want to run it or ship it? Running an open source runtime yourself is different from offering audit and policy as a feature under your brand.

Frequently asked questions

Do we have to change how our sandboxes run agents? No. Qpoint runs on the host underneath the agent. There is no SDK, sidecar per agent, or wrapper around the entrypoint, and the Kubernetes DaemonSet does not restart running pods.

Can our customers get per-session audit trails? Yes. Every file access, command, tool call, MCP connection, and model request is tied to the session, the identity it ran as, and the tenant and workspace you tag it with.

Can Qpoint block actions, not just record them? Enforcement is coming soon. At launch, Qpoint discovers agents and records and exports everything they do.

Can we offer this under our own brand? Yes. Qpoint is built to be bundled and white-labeled under a commercial partner license.

Talk to us

If you run agent sandboxes and want to offer audit logs and policy controls to your customers, talk to us. We will walk through your platform and scope the integration together.