Watch InnerWarden stop a hijacked AI agent. On camera, not slides.
A web page or a tool result can trick an AI agent into running a dangerous command. Most guardrails live inside the agent, so a hijacked agent can switch them off. InnerWarden enforces one level deeper, in the Linux kernel, where the agent can't reach. Every clip below is a real recording from our lab, trimmed for length. Three closing captions the recording script typed into the terminal are cut, because they claimed more than the clips show; nothing a program printed was removed.
- 01Can it block a program the agent was never allowed to run?
- 02Can it stop malware from phoning home?
- 03Can it guard just the AI agent and leave the rest of the server alone?
- 04In a shared fleet, can it name and cage the one bad agent?
Four real recordings from our lab, trimmed for length. Green means allowed, red means blocked.
The threat
A web page, a file, or a tool result hides an instruction that tricks your AI agent into running a dangerous command.
The catch
InnerWarden enforces one level below the agent, in the Linux kernel. Inside the agent's scope, a program that was not approved does not start, whatever the agent was told.
The safe part
It installs disarmed, refuses a host-wide enforce, and refuses to arm while a program your machine needs to boot is missing from its allowlist. Turning it off is always instant.
- 01
Can it block a program the agent was never allowed to run?
What you'll see: An approved command runs normally. Then an unknown program tries to start, and is stopped the instant it launches.
UbuntuLinux kernelExecution GateResult: The unknown program was blocked by the kernel before it ran a single instruction. Approved programs kept working, and turning the guard off is instant.
For engineers: how it works
Enforcement is an eBPF LSM hook on bprm_check_security. Any binary whose path-hash is not in the pre-authorized EXEC_ALLOWLIST map is denied at exec with -EPERM, below the agent process: the same bytes at an unknown path are denied, because the kernel compares the path, not the content. Approve a binary and the in-kernel map updates in-process so a re-exec passes; `exec-gate disarm` clears the policy key with no licence and no restart. The clip was recorded on 0.15.24 with a development licence key, which is why its log warns that the signature check was skipped. That build armed the gate host-wide; the current release refuses a host-wide enforce, and the gate is scoped to an agent first, as in clip 03.
Ubuntu · Linux kernel · Execution Gatereal recordingstarting recording…loading…raw .cast - 02
Can it stop malware from phoning home?
What you'll see: In observe, a known-bad domain still resolves and is only recorded. After a rehearse the operator arms enforce, and the same domain is answered NXDOMAIN while a clean domain keeps resolving. Disarmed, it resolves again.
UbuntuDNS Guard (local resolver)Result: In enforce, a lookup of a denylisted domain is answered NXDOMAIN, so a callback to it fails at the lookup step, and example.com, the clip's clean domain, kept resolving. The clip's rehearse also lists example.com as a denylisted would-block, and the recording does not show why, so read the clip as proof that denylisted lookups are refused, not that the rehearse list is exact. This is a lab recording. No host we run has DNS Guard in enforce today: on our public challenge box it is set to observe, and only the challenge agent is pointed at it.
For engineers: how it works
DNS Guard is a forwarding resolver that refuses to resolve malicious domains: in enforce, any domain on the threat-feed denylist is answered NXDOMAIN, and everything else is forwarded upstream. It runs in observe by default. The operator arms enforce; we recommend a rehearse first, and arming does not require one. In the clip, the arm step only asks whether a rehearse was run. Enforce needs a paid licence; disarm needs none.
Ubuntu · DNS Guard (local resolver)real recordingstarting recording…loading…raw .cast - 03
Can it guard just the AI agent and leave the rest of the server alone?
What you'll see: The gate is scoped to a real OpenClaw agent's cgroup and armed in enforce, with no rehearse. An unknown program runs on the host outside it; the same program, run inside the agent's cgroup, is blocked, while an approved one still runs there.
UbuntuExecution Gate scoped to the agentResult: Inside the agent's cgroup only pre-approved programs ran, enforced below the agent. On the host outside it, the same unknown program ran as normal.
For engineers: how it works
The gate scopes to a single cgroup, the agent's process tree, via the EXEC_GATE_SCOPE map. Only exec inside that cgroup is checked; in the clip the block is shown by a shell that joins the agent's cgroup and runs the program, and the same program runs on the host outside it. The clip arms enforce with the lower-level `config-sign exec-gate arm --apply`, which does not run a rehearse. After the refusal the shell prints '/tmp/rogue-tool: Success': the exec did fail, which is why the shell printed an error at all, but that line names the wrong reason; clip 01 shows the same refusal with the right one, 'Operation not permitted'.
Ubuntu · Execution Gate scoped to the agentreal recordingstarting recording…loading…raw .cast - 04
In a shared fleet, can it name and cage the one bad agent?
What you'll see: Two teams share one cloud server; one agent pod goes rogue and runs its attacks. InnerWarden names the exact team and pod behind each incident. The Execution Gate is then scoped to that pod alone, watched in observe, and switched to enforce.
AzureKubernetesper-tenant attribution + agent-scoped gateResult: The rogue pod's first attacks ran, and were detected and attributed to it. Once the gate was enforcing on that pod, a new miner binary was denied in the kernel, while the pod's approved programs, the other team and the host kept running.
For engineers: how it works
Pod identity is read straight from the kernel cgroup, so the agent cannot spoof it, and each incident is attributed to its tenant and pod. The gate runs on a real cloud kernel (7.0, k3s). In this recording its scope and mode are written straight into the gate's maps with bpftool, to show the kernel side; that skips the product's own arm step and its refusals. The pod's baseline allowlists /usr/bin/bash, so a script handed to bash would still run: the gate decides which programs start, not what an allowed one does.
Azure · Kubernetes · per-tenant attribution + agent-scoped gatereal recordingstarting recording…loading…raw .cast
“Won't a kernel guard risk breaking my machine?”
That is the first thing we engineered against. Arming refuses a host-wide enforce, refuses while a boot-essential program is missing from the allowlist, and refuses where the kernel cannot run the gate. On its very first real run it caught several boot-essential programs and refused to arm. We fixed the coverage, and the same check now passes clean. A rehearse also checks the allowlist against the programs running now, and innerwarden-ctl exec-gate enforce switches an agent's gate to enforce only after an observe window in which the agent ran nothing unlisted. The lower-level arm command does not wait for that window: clip 03 arms enforce with it and no rehearse, and clip 04 writes the gate's maps by hand to show the kernel side, which skips the arm step altogether.
It installs disarmed; nothing changes until you arm it.
No host-wide enforce, and no arm while a boot essential is missing; the enforce command waits for a clean observe window.
Disarm needs no licence and never locks you out.
Try the guard in your own terminal
The free tier screens your agent's commands and tool calls out of the box. The kernel Execution Gate you just watched is part of the paid host stack (Pro and Enterprise): it installs disarmed, and you arm it in observe first.
npm install -g innerwardenPrebuilt binary with signed npm provenance from a trusted registry. No postinstall script, nothing runs at install time · Linux, macOS, Windows. On Linux the npm global prefix is root-owned, so this one needs sudo; the signed installer at innerwarden.com/free does not. Prefer not to install? npx innerwarden runs it once.
Running agents in production? Talk to security about arming the gate on your hosts.
Other install methods
curl -fsSL https://innerwarden.com/free | shSigned release, no Node required · macOS, Linux.
irm https://innerwarden.com/free.exe -OutFile innerwarden.exeWindows PowerShell.
cargo install --git https://github.com/InnerWarden/inner-warden innerwardenFrom source, Rust.