I run my own server — nginx, Docker, a database, a handful of systemd units. Most of my day already happens in Claude Code, so the obvious next step was to point the agent at that server and let it answer questions like “what filled the disk?” or “why is nginx throwing 502s?” without me pasting logs into a chat window.

Every option I found offered some version of the same deal:

Here is a tool that takes a command string and runs it over SSH. Commands that look dangerous are blocked.

That is a root shell with a regex in front of it, and the regex is the weak part — because the agent is not the only author of that string. An agent tailing logs reads text an attacker wrote. An agent summarizing a web page or a commit message reads text an attacker wrote. A prompt-injected agent is a perfectly well-behaved process carrying out someone else’s intentions, and a shell is all it needs.

So I built opsgate: an MCP server that gives agents real capability on real servers, without ever giving them a string to fill in.

▸ nginx_test(host: "web1")
  nginx: configuration file /etc/nginx/nginx.conf test is successful

▸ nginx_reload(host: "web1")

  ┌─ opsgate: approval required ────────────────────────────────┐
  │ allow nginx_reload on host "web1"?                          │
  │                                                             │
  │ The nginx configuration passed validation.                  │
  │ This will reload nginx.                                     │
  │                                                             │
  │ Command: 'nginx' '-s' 'reload'                              │
  │                                       [ Approve ]  [ Deny ] │
  └─────────────────────────────────────────────────────────────┘

Deny, and nothing happens — the agent is told the operator refused. Approve, and it runs and lands in the audit log. Either way, you saw the exact command first.

Denylists are guesswork Link to heading

The standard mitigation is a denylist: refuse anything that matches rm -rf, mkfs, dd, and friends. It fails the moment someone spells the same thing differently:

rm -rf /                              →  blocked
r''m -rf /                            →  not blocked
echo cm0gLXJmIC8K | base64 -d | sh    →  not blocked

You can keep adding patterns, but you are playing defense against an adversary with infinite phrasings and unlimited retries. And “dangerous-looking” was never the real problem. systemctl restart nginx contains nothing a filter would flag, and it is catastrophic at 3pm on Black Friday. A safety model has to reason about what an action does and where — not about what its text looks like.

Safe by construction, not by regex Link to heading

opsgate does not filter strings. The agent gets 25 typed tools — service_status(name), docker_logs(container), journal_tail(unit, since, priority), file_read(path) — and opsgate builds every command itself, as an argv array:

// What opsgate does — the name is one argv element, whatever it contains
argv := []string{"systemctl", "restart", in.Name}

// What it never does
argv := []string{"sh", "-c", "systemctl restart " + in.Name}

On the local machine, argv goes straight to exec.Command — no shell process exists at all. Over SSH a remote shell is unavoidable, so the executor single-quotes every element, which keeps each one exactly one shell word. Send nginx; rm -rf / as a service name and systemctl receives it as a single literal argument and fails as an unknown unit. Nothing else runs. The test suite proves this by round-tripping injection payloads through a real shell, rather than trusting the quoting function on inspection.

A second, independent layer sits behind it: service, container, and unit names are restricted to a conservative character set, and file_grep runs grep -F, so a search pattern is always a literal string and never a regex.

The design baseline is blunt: opsgate assumes every tool argument is attacker-controlled. Nothing in the safety model depends on the model behaving well.

Deny by default, per host Link to heading

Every tool is classified as observe, mutate, or shell when it is registered, and every call is checked against the mode of the host it targets:

mode: operate            # reads run freely; every write asks a human first

hosts:
  prod:
    addr: 10.0.0.1
    mode: observe        # prod is read-only, whatever the default says
  staging:
    addr: 10.0.0.2       # inherits operate

tools:
  service_restart:
    allow_targets: ["nginx", "myapp*"]   # postgresql is not restartable, period
  service_stop:
    enabled: false                       # the agent never even sees this tool

observe is the default, and it is not “mutation, but discouraged” — there is no code path from a read-only mode to a state-changing command. allow_targets narrows a tool to specific objects in every mode and at every approval level. Setting it also makes the target mandatory, because leaving it out is usually broader than any allowed value: journal_tail with no unit reads the entire journal, so an unconstrained call is refused instead of silently permitted.

That last rule is typical of the whole project. When a call is ambiguous, opsgate fails closed.

A human sees the exact command Link to heading

In operate mode, every mutating call stops and asks — showing the tool, the host, a plain-language summary, and the literal argv that will run. Approval is per call; there is no “remember this choice”.

This is the control that actually addresses prompt injection. A compromised agent can request a restart. It cannot perform one without a person reading the request first. And if your MCP client cannot show an approval prompt at all, the call is refused with an explanation — it never silently proceeds.

The prompt uses MCP elicitation, in both the older and the newer protocol style, and the tests drive a real MCP client to prove that only an explicit accept executes.

Receipts: a hash-chained audit log Link to heading

Every call — including every refusal — appends one JSON line that embeds the hash of the previous record:

{"seq":42,"host":"web1","tool":"file_read","args":{"path":"/root/.ssh/id_rsa"},"decision":"deny","error":"outside the allowed paths",…}
{"seq":43,"host":"web1","tool":"service_restart","phase":"intent","decision":"needs_approval","approved":true,…}
{"seq":44,"host":"web1","tool":"service_restart","phase":"outcome","decision":"needs_approval","approved":true,"exit_code":0,…}

Three details matter more than they look:

  • Refusals are recorded. A blocked attempt to read a private key is exactly the line you want to find later.
  • Intent is written before execution. A state-changing call writes phase: intent before the command runs and phase: outcome after it. If the intent record cannot be written, opsgate refuses to act — so an action that happened can never be missing from the log.
  • Secrets stay out. Values under keys like password and token are redacted, and only a SHA-256 of command output is stored, never the output itself.

Edit, delete, or reorder any line, and verification fails at exactly that point:

$ opsgate audit verify
opsgate: audit chain INVALID after 41 records: line 42: record tampered (hash mismatch)

Break it yourself Link to heading

I would rather you try to break it than take my word for it. Run opsgate in observe mode with files.allow_paths: [/var/log] and throw these at it:

Try thisWhat happens
service_status(name: "nginx; touch /tmp/pwned")Refused — invalid characters in a target name
file_read(path: "/var/log/../../etc/shadow")Refused — the path is cleaned first, so it lands outside the allowlist
file_grep(pattern: "x$(touch /tmp/pwned)")Searched as a literal string; nothing executes
service_restart(name: "nginx")Refused — observe mode cannot mutate
shell_exec(command: "id")Rejected — the tool does not exist unless you opt in

Afterwards, /tmp/pwned does not exist, and opsgate audit verify reports an intact chain holding a record of each refused call. If something does get through, that is a security bug — please report it privately through the repository’s security advisories.

What opsgate deliberately is not Link to heading

  • Not a sandbox. Approved commands run for real, with the privileges of the SSH user you configured. If that user is root, an approved call is a root call — give opsgate its own least-privilege user.
  • Not a reason to trust an untrusted agent with production. A prompt-injected agent can still call allowed tools with plausible arguments. The policy is what bounds the damage, which is why tight allow_targets matter.
  • Not complete. The tool set is finite. shell_exec exists for the genuine long tail: disabled by default, approval-gated when enabled, and arbitrary code execution by design once approved.

Two more limits are worth stating plainly. The path check is lexical, so a symlink inside an allowed directory that points outside it will be followed. And the audit chain is tamper-evident, not tamper-proof: someone with write access to the file can rewrite the entire chain, so ship it off the box if that is your threat. All of this is written down in the threat model, next to the controls that address the rest.

How I run it Link to heading

I run opsgate from Claude Code against my own server, in observe mode. Most of what you want from an agent on a server is diagnosis — which process holds a port, what crashed overnight, what is eating the disk — and observe mode answers all of it while being structurally unable to change anything. When a change is genuinely needed, operate mode turns it into one approved command instead of a session in a root shell.

CI runs govulncheck on every push. When two denial-of-service advisories landed in golang.org/x/crypto/ssh — the library opsgate uses to reach your hosts — the next build went red, and v0.1.1 shipped the fix the same day.

Try it Link to heading

opsgate runs with no config file at all — local machine, read-only:

go install github.com/polymatx/opsgate/cmd/opsgate@latest
opsgate serve
opsgate dev | mode=observe | default_host=local | audit=~/.opsgate/audit.jsonl

Add it to Claude Code, then ask what is using disk space on your machine:

claude mcp add opsgate -- opsgate serve

When you are ready for real hosts, opsgate init writes a config. SSH host keys are verified against known_hosts; unknown keys are refused rather than trusted on first use. Cloud or sandboxed clients that cannot launch a local process can reach opsgate over HTTP with a bearer token — read the warning in the README before you expose that endpoint to any network.

The code is MIT licensed at github.com/polymatx/opsgate. New tools are around twenty lines each, and the one rule that matters is in CONTRIBUTING: commands are argv, never concatenated strings. Kubernetes, PostgreSQL status, and certificate expiry would all make good first contributions.