Skip to content
See it work

Watch InnerWarden stop a hijacked AI agent. On camera, not slides.

A web page or a tool result can trick an AI agent into running a dangerous command. Most guardrails live inside the agent, so a hijacked agent can switch them off. InnerWarden enforces one level deeper, in the Linux kernel, where the agent can't reach. Every clip below is a real recording from our lab, trimmed for length. Three closing captions the recording script typed into the terminal are cut, because they claimed more than the clips show; nothing a program printed was removed.

The threat

A web page, a file, or a tool result hides an instruction that tricks your AI agent into running a dangerous command.

The catch

InnerWarden enforces one level below the agent, in the Linux kernel. Inside the agent's scope, a program that was not approved does not start, whatever the agent was told.

The safe part

It installs disarmed, refuses a host-wide enforce, and refuses to arm while a program your machine needs to boot is missing from its allowlist. Turning it off is always instant.

How to read the clips:you don't need to follow every line;green = allowed,red = blocked. Watch for the red line where the attack is denied.
  1. 01

    Can it block a program the agent was never allowed to run?

    What you'll see: An approved command runs normally. Then an unknown program tries to start, and is stopped the instant it launches.

    UbuntuLinux kernelExecution Gate

    Result: The unknown program was blocked by the kernel before it ran a single instruction. Approved programs kept working, and turning the guard off is instant.

    For engineers: how it works

    Enforcement is an eBPF LSM hook on bprm_check_security. Any binary whose path-hash is not in the pre-authorized EXEC_ALLOWLIST map is denied at exec with -EPERM, below the agent process: the same bytes at an unknown path are denied, because the kernel compares the path, not the content. Approve a binary and the in-kernel map updates in-process so a re-exec passes; `exec-gate disarm` clears the policy key with no licence and no restart. The clip was recorded on 0.15.24 with a development licence key, which is why its log warns that the signature check was skipped. That build armed the gate host-wide; the current release refuses a host-wide enforce, and the gate is scoped to an agent first, as in clip 03.

    Ubuntu · Linux kernel · Execution Gatereal recording
    starting recording…
    loading…raw .cast
  2. 02

    Can it stop malware from phoning home?

    What you'll see: In observe, a known-bad domain still resolves and is only recorded. After a rehearse the operator arms enforce, and the same domain is answered NXDOMAIN while a clean domain keeps resolving. Disarmed, it resolves again.

    UbuntuDNS Guard (local resolver)

    Result: In enforce, a lookup of a denylisted domain is answered NXDOMAIN, so a callback to it fails at the lookup step, and example.com, the clip's clean domain, kept resolving. The clip's rehearse also lists example.com as a denylisted would-block, and the recording does not show why, so read the clip as proof that denylisted lookups are refused, not that the rehearse list is exact. This is a lab recording. No host we run has DNS Guard in enforce today: on our public challenge box it is set to observe, and only the challenge agent is pointed at it.

    For engineers: how it works

    DNS Guard is a forwarding resolver that refuses to resolve malicious domains: in enforce, any domain on the threat-feed denylist is answered NXDOMAIN, and everything else is forwarded upstream. It runs in observe by default. The operator arms enforce; we recommend a rehearse first, and arming does not require one. In the clip, the arm step only asks whether a rehearse was run. Enforce needs a paid licence; disarm needs none.

    Ubuntu · DNS Guard (local resolver)real recording
    starting recording…
    loading…raw .cast
  3. 03

    Can it guard just the AI agent and leave the rest of the server alone?

    What you'll see: The gate is scoped to a real OpenClaw agent's cgroup and armed in enforce, with no rehearse. An unknown program runs on the host outside it; the same program, run inside the agent's cgroup, is blocked, while an approved one still runs there.

    UbuntuExecution Gate scoped to the agent

    Result: Inside the agent's cgroup only pre-approved programs ran, enforced below the agent. On the host outside it, the same unknown program ran as normal.

    For engineers: how it works

    The gate scopes to a single cgroup, the agent's process tree, via the EXEC_GATE_SCOPE map. Only exec inside that cgroup is checked; in the clip the block is shown by a shell that joins the agent's cgroup and runs the program, and the same program runs on the host outside it. The clip arms enforce with the lower-level `config-sign exec-gate arm --apply`, which does not run a rehearse. After the refusal the shell prints '/tmp/rogue-tool: Success': the exec did fail, which is why the shell printed an error at all, but that line names the wrong reason; clip 01 shows the same refusal with the right one, 'Operation not permitted'.

    Ubuntu · Execution Gate scoped to the agentreal recording
    starting recording…
    loading…raw .cast
  4. 04

    In a shared fleet, can it name and cage the one bad agent?

    What you'll see: Two teams share one cloud server; one agent pod goes rogue and runs its attacks. InnerWarden names the exact team and pod behind each incident. The Execution Gate is then scoped to that pod alone, watched in observe, and switched to enforce.

    AzureKubernetesper-tenant attribution + agent-scoped gate

    Result: The rogue pod's first attacks ran, and were detected and attributed to it. Once the gate was enforcing on that pod, a new miner binary was denied in the kernel, while the pod's approved programs, the other team and the host kept running.

    For engineers: how it works

    Pod identity is read straight from the kernel cgroup, so the agent cannot spoof it, and each incident is attributed to its tenant and pod. The gate runs on a real cloud kernel (7.0, k3s). In this recording its scope and mode are written straight into the gate's maps with bpftool, to show the kernel side; that skips the product's own arm step and its refusals. The pod's baseline allowlists /usr/bin/bash, so a script handed to bash would still run: the gate decides which programs start, not what an allowed one does.

    Azure · Kubernetes · per-tenant attribution + agent-scoped gatereal recording
    starting recording…
    loading…raw .cast
The honest question

“Won't a kernel guard risk breaking my machine?”

That is the first thing we engineered against. Arming refuses a host-wide enforce, refuses while a boot-essential program is missing from the allowlist, and refuses where the kernel cannot run the gate. On its very first real run it caught several boot-essential programs and refused to arm. We fixed the coverage, and the same check now passes clean. A rehearse also checks the allowlist against the programs running now, and innerwarden-ctl exec-gate enforce switches an agent's gate to enforce only after an observe window in which the agent ran nothing unlisted. The lower-level arm command does not wait for that window: clip 03 arms enforce with it and no rehearse, and clip 04 writes the gate's maps by hand to show the kernel side, which skips the arm step altogether.

Off by default

It installs disarmed; nothing changes until you arm it.

Refuses an unsafe arm

No host-wide enforce, and no arm while a boot essential is missing; the enforce command waits for a clean observe window.

Instant off-switch

Disarm needs no licence and never locks you out.

Try the guard in your own terminal

The free tier screens your agent's commands and tool calls out of the box. The kernel Execution Gate you just watched is part of the paid host stack (Pro and Enterprise): it installs disarmed, and you arm it in observe first.

npm install -g innerwarden

Prebuilt binary with signed npm provenance from a trusted registry. No postinstall script, nothing runs at install time · Linux, macOS, Windows. On Linux the npm global prefix is root-owned, so this one needs sudo; the signed installer at innerwarden.com/free does not. Prefer not to install? npx innerwarden runs it once.

Protect an agent

Running agents in production? Talk to security about arming the gate on your hosts.

Other install methods

curl -fsSL https://innerwarden.com/free | sh

Signed release, no Node required · macOS, Linux.

irm https://innerwarden.com/free.exe -OutFile innerwarden.exe

Windows PowerShell.

cargo install --git https://github.com/InnerWarden/inner-warden innerwarden

From source, Rust.

All install methods and how to verify →