Monitoring AI Agents in Kubernetes and CI Runners: What to Use
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
Monitoring AI Agents in Kubernetes and CI Runners: What to Use
There are two kinds of monitoring for AI agents, and most teams running agents on their own infrastructure need both. LLM observability tools such as Langfuse, LangSmith, Arize Phoenix, and Braintrust show whether an agent is working: its traces, prompts, latency, cost, and quality. Agent activity monitoring shows what the agent actually did on your nodes and runners: the files it touched, the commands it ran, the tools and MCP servers it called, the requests it sent, and the identity it used. Infrastructure tools like Datadog and Grafana and runtime security tools like Falco and Tetragon cover part of the second question but do not know which processes are agents. Qpoint is built for that second question and sends its events into the tools you already use.
Two questions, two kinds of tools
Is the agent working? This is about quality and performance: what the model was asked, what it answered, how long it took, what it cost, and whether the output was good. LLM observability tools answer this by tracing the agent's own code.
What did the agent do on our infrastructure? This is about operations and security: which files it read, which commands and processes it ran, which tools and MCP servers it called, what it sent to model providers, and which service account or credentials it used. This needs visibility on the node, underneath the agent.
Monitoring AI agents in Kubernetes: the options
LLM observability (Langfuse, LangSmith, Arize Phoenix, Braintrust)
These record traces from inside an agent you instrument with an SDK or OpenTelemetry: LLM calls, prompts, tool calls the framework knows about, evaluations, and cost. They are a good fit when you build the agents and want to debug and improve them. They only see what the instrumented code reports, so they do not cover third-party agents like Claude Code or Codex running in your pods, or what a spawned script does on the node.
Infrastructure monitoring (Datadog, Grafana, Kubernetes audit logs)
These show pod health, resource use, logs, and API server activity. Datadog also offers LLM observability for instrumented applications. They are a good fit for running the platform. They see containers and processes, not agent sessions, so they cannot tell you which agent read a secret or why a request went out.
Runtime security (Falco, Cilium Tetragon)
These watch system calls, file access, process execution, and network connections on each node, and can alert on or block suspicious behavior. They are a good fit for general workload security. They do not know which process is an AI agent, which session or prompt an action belongs to, or what a tool call or MCP connection means, so agent activity looks like any other process activity.
Agent activity monitoring (Qpoint)
Qpoint runs on each node, underneath the agents, and recognizes agents by their behavior. It records every action as part of an agent session and ties it to the identity the session ran as. It is a good fit when agents you did not write run on your infrastructure, when security needs to answer what an agent did, and when you want that data in Datadog, Grafana, or your SIEM rather than another console. See Qpoint for hosted agents.
What to watch for each kind of agent workload
Agents on your infrastructure usually fall into three groups, and each raises a different question:
- Background agents. Engineers hand a task to Claude Code, Codex, or OpenHands running in sandbox pods or VMs and get a pull request back. The question is what the agent touched while nobody was watching.
- Pipeline agents. Agents run as a step in GitHub Actions, GitLab CI, or Buildkite, reviewing code, fixing tests, or writing release notes. The question is whether a pull request comment could steer the agent into the runner's secrets.
- Agent services. Long-running agents answer customers or internal users. The question is what the agent can do on a customer's behalf.
Monitoring AI agents in CI runners
CI runners are where agents meet credentials: the runner's token, the repository's deploy keys, and any cloud role the workflow assumes. They are also short-lived, so anything that depends on registering each workload misses most of them.
For GitHub Actions, StepSecurity Harden-Runner monitors what each job sends over the network and can restrict it, which is useful for any workflow, with or without AI. It does not know which steps are agent activity or what an agent's tool calls were.
Qpoint covers CI runners from the host. A pipeline agent's session looks like this:
| Time | Event | Detail |
|---|---|---|
| 00:00 | Identity | runner token ci-7f3a · OIDC role ci-deploy |
| 00:01 | Session start | runner ci-7f3a · workflow ai-review.yml |
| 00:04 | File read | diff · 14 files |
| 00:18 | File read · blocked | /var/run/secrets/…/serviceaccount/token |
| 00:22 | Request | api.openai.com · 8.1k tokens |
| 00:40 | Tool call | post_review_comment · 6 comments |
The record shows what the agent read, the identity it ran as, what it sent to the model, and the attempt to read the service account token, which was blocked.
How Qpoint deploys
Pick the method that fits each workload. All of them report to the same control plane, which runs in your environment.
- Kubernetes: a DaemonSet, from a Helm chart or manifest. It covers every pod on every node, including pods scheduled later, without restarting running pods.
- Containers: one line in your Dockerfile or base image, for Docker, ECS, and Nomad.
- VMs, bare metal, and CI runners: a single Linux package that covers every process on the host.
One config file covers every node. It sets where events go and which agents to cover. Events can go to OTLP, Splunk HEC, Datadog, Elastic, or S3, so agent activity shows up in the dashboards and alerts your platform and security teams already use. Coverage comes from the node, so pods that reschedule and runners that live for minutes are covered without registering anything. Token usage and spend roll up by service, job, provider, and model, and follow your team labels. There is no SDK, sidecar per agent, or wrapper around the entrypoint, and your data stays on your infrastructure. The hosted agents page has more on each deployment method, and Qpoint Monitor covers what gets recorded.
Frequently asked questions
We already use Langfuse or LangSmith. Do we need this? They answer different questions. Keep your LLM observability for traces and evals of agents you build. Qpoint covers what any agent, including ones you did not build, does on the node.
Does it work with Datadog and Grafana? Yes. Qpoint sends events over OTLP or directly to Datadog, so agent activity appears next to your existing metrics, logs, and traces.
Does it work with GitHub-hosted runners? Qpoint runs on the host, so it is designed for runners you operate: self-hosted runners, and runners on your own VMs or Kubernetes. If you use GitHub-hosted runners, talk to us about your setup.
Do we have to change our agents or pipelines? No. Qpoint runs underneath the agent. Workloads run exactly as they did before.
What about pods and runners that only exist for a few minutes? Coverage comes from the node, so short-lived pods and runners are recorded without registering them.
See it on your clusters
Book a demo and we will walk through your clusters and pipelines and show you an agent session recorded end to end.