← All guides

ZeroClaw Security: Workspace Scoping, Command Allowlists and Sandboxing

ZeroClaw.net · Updated 2026-10-08

ZeroClaw Security: Workspace Scoping, Command Allowlists and Sandboxing

Version notice — 8 October 2026: Verified against ZeroClaw v0.8.5. Security boundaries, autonomy levels, command policies, and sandboxing are configured under [risk_profiles.<alias>] and per-agent workspace definitions, replacing legacy top-level [security] keys. Absolute safety claims have been replaced with tested runtime controls and actual enforcement limits.

An AI agent is a program that decides what to do next based on text it was given, and then does it. If that program can read files and execute commands, anyone who can influence that text can influence what runs on your host machine. This is not a hypothetical failure mode — it is the normal operating condition of every autonomous agent runtime.

ZeroClaw's approach is layered containment: several independent controls, each of which limits the blast radius when others fail. This guide covers what each control does in ZeroClaw v0.8.5, how to configure it in config.toml, what each boundary genuinely prevents, and where its real limits lie.

The threat model

Before configuring anything, be clear about what you are defending against. There are three distinct risks:

  1. The agent makes an inference error. The model misreads an instruction and attempts to delete or overwrite the wrong directory. No malice is involved, just a faulty inference executed with actual system permissions.
  2. Unauthorized access to an agent endpoint. Your gateway or bot is discovered by an unauthorized party. Anyone who reaches it can issue tasks unless strong authentication blocks them.
  3. Prompt injection. The agent reads a webpage, an email, or a document containing adversarial text crafted to look like instructions. Frontier models cannot reliably distinguish data they were asked to process from instructions they were asked to follow.

The goal of ZeroClaw's security model is not an agent that can never be tricked. It is an agent whose worst possible day is strictly bounded and survivable.

Control 1: Workspace scoping and forbidden paths

Workspace scoping restricts filesystem tool operations to a specific folder hierarchy.

[agents.assistant.workspace]
path = "/var/zeroclaw/workspace"
unrestricted_filesystem = false

[risk_profiles.hardened]
workspace_only = true

When workspace_only = true is enabled on the active risk profile, ZeroClaw's internal file tools (file_read, file_write, file_list) resolve relative paths against the workspace and verify that canonicalized paths do not escape the designated directory tree.

Built-in forbidden paths

In ZeroClaw v0.8.5, the runtime enforces an automatic denylist of sensitive operating system paths, regardless of whether workspace scoping is enabled:

[risk_profiles.hardened]
forbidden_paths = [
  "/etc",
  "/root",
  "~/.ssh",
  "~/.gnupg",
  "~/.aws",
  "~/.config"
]

On Windows, the default denylist blocks C:\Windows, C:\Program Files, ~/.ssh, ~/.aws, and similar system paths.

Realistic limits of filesystem scoping

Filesystem path checking is not an impenetrable OS container:

  • Child process bypass: Workspace path validation applies to ZeroClaw's built-in file tools. If the agent is allowed to execute shell commands (git, cargo, bash), those processes run under the host operating system user and can access any file that user can access — unless restricted by OS-level sandboxing (Control 4).
  • Network exfiltration: Workspace scoping bounds what files can be read; it does not restrict where read data can be sent. An agent with network access (http_request or curl) can exfiltrate workspace contents.
  • Symlinks: Symlinks created inside the workspace pointing outside can create traversal vectors. ZeroClaw canonicalizes paths to detect symlink escapes, but nested archive extraction should still be supervised.

Control 2: Autonomy levels and approval gates

Autonomy is not a binary switch. In v0.8.5, autonomy is configured per risk profile via [risk_profiles.<alias>].level. Each agent references a profile via agents.<alias>.risk_profile = "<alias>".

The runtime accepts three distinct levels:

[risk_profiles.reader]
level = "readonly"      # "readonly", "supervised", or "full"

Note: The string is readonly without an underscore. Using read_only fails strict schema validation at startup.

1. readonly

The agent can inspect its environment but cannot produce persistent side effects. Allowed tools are strictly read-only: file_read, file_list, memory_search, http (GET only), web_search, and time. File writes, shell executions, and HTTP POSTs are blocked.

2. supervised (recommended default)

Actions are classified by risk score:

Risk levelOperationsRuntime behavior
Lowfile_read, memory_search, web_search, time, http GETExecutes automatically
Mediumfile_write within workspace, allowed shell commands, http POSTPauses and requests operator approval
HighDisallowed commands, writes outside workspace, high-risk flagsBlocked by default

When a medium-risk action occurs, ZeroClaw sends an approval request over the active channel (such as an inline keyboard in Telegram, an interactive prompt in the CLI, or a session/request_permission JSON-RPC call in the Web UI/ACP).

  • Timeout behavior: Unanswered approval requests expire after the channel's approval_timeout_secs (default 120 seconds). Timeouts fail closed and are treated as explicit denials.
  • Cross-channel approval routing: You can route approvals to a dedicated operator channel instead of the initiating chat:
    [risk_profiles.frontline.approval_route]
    approver_channel = "matrix.ops"
    

3. full

The agent executes low, medium, and high-risk tool operations without prompting. workspace_only is relaxed, though forbidden_paths and kernel sandboxing remain active unless explicitly turned off (such as in YOLO mode). Use full only in disposable CI environments or trusted local development.

Control 3: Command allowlists and risk scoring

Shell command execution is deny-by-default. If an agent attempts to invoke a command not present in allowed_commands, the invocation is immediately rejected:

[risk_profiles.hardened]
allowed_commands = ["git", "ls", "cat"]
require_approval_for_medium_risk = true
block_high_risk_commands = true
shell_timeout_secs = 60

The limits of command allowlists

Allowlisting an executable allowlists everything that executable can execute:

  • git can trigger arbitrary script execution through git hooks and repository configuration (git config core.hooksPath).
  • cargo and npm execute build scripts (build.rs) and lifecycle hooks (postinstall) defined in external project files.
  • Interpreters (python, node, bash) provide complete shell escape capabilities.
  • Setting allowed_commands = ["*"] completely removes command denylisting.

Start with an empty allowlist or restrict it to non-interpreting read tools (ls, cat, date). If build tools are required, back them up with kernel sandboxing and run ZeroClaw under a restricted operating system user.

Control 4: OS-level sandboxing (Landlock and Firejail)

To contain child processes executed by shell tools, ZeroClaw v0.8.5 integrates operating system kernel sandboxing:

[risk_profiles.hardened]
sandbox_enabled = true
sandbox_backend = "landlock"   # "landlock" (Linux) or "firejail"
  • Landlock (Linux 5.13+): Uses the Linux kernel's Landlock LSM (Linux Security Module). It enforces unprivileged filesystem access restrictions directly on the child process, preventing spawned tools from reading files outside permitted directories even if the process runs as your user.
  • Firejail: Wraps command execution inside Linux namespaces and seccomp filters, isolating filesystem access, network interfaces, and process tables.

Control 5: Action tracking and cost ceilings

Autonomous loops can run away if an agent gets trapped in a retry cycle. ZeroClaw limits operational velocity and financial spend:

[risk_profiles.hardened]
max_actions_per_hour = 20
max_cost_per_day_cents = 500   # $5.00 daily limit
  • Action rate limiting: Tracks successful operations within a sliding one-hour window. If the agent exceeds 20 operations in an hour, subsequent actions are throttled.
  • Cost ceiling: Tracks token spend against model provider pricing caches. When cumulative daily spend reaches 500 cents ($5.00), further inference calls are halted.

Control 6: Gateway and channel pairing

To prevent unauthorized parties from sending instructions to your agent:

  • Pairing required by default: The gateway listener ([gateway] require_pairing = true) requires a one-time 8-character pairing code to authorize any browser or API client.
  • Public bind protection: The gateway binds to 127.0.0.1:42617. Setting host = "0.0.0.0" is rejected at startup unless allow_public_bind = true is explicitly configured.
  • Remote admin lock: Administrative endpoints such as reload and shutdown reject non-loopback connections unless allow_remote_admin = true and valid pairing credentials are provided.

Control 7: Encrypted secret storage

API keys are encrypted at rest using local machine keyfiles:

zeroclaw quickstart

Credentials stored in configuration files are encrypted before writing to disk, protecting against accidental exposure in git commits or public backup archives.

Realistic limit: The decryption key lives on the host filesystem so the unattended service can restart. Anyone with read access to the user's filesystem can decrypt the secrets. Defense-in-depth requires file permission restrictions (chmod 700 ~/.zeroclaw) and provider-side budget caps.

Control 8: WebAssembly plugin sandboxing

ZeroClaw plugins compile to WebAssembly components (wasm32-wasip2) executed in Wasmtime:

  • Deny-by-default capabilities: Plugins possess zero ambient authority. They cannot open sockets or access filesystem directories unless explicitly declared in the plugin's manifest.toml and approved by the operator.
  • Integrity and authenticity: Plugin installation checks transport integrity via SHA-256 and verifies author authenticity using Ed25519 digital signatures.

A tested hardened security policy

Here is a tested, schema-3-compliant configuration implementing these controls:

schema_version = 3

[providers.models.ollama.local]
model = "qwen2.5:0.5b"
uri = "http://127.0.0.1:11434"
temperature = 0.0
num_ctx = 4096
num_predict = 128
native_tools = false

[agents.assistant]
model_provider = "ollama.local"
risk_profile = "hardened"

[agents.assistant.workspace]
path = "/var/zeroclaw/workspace"
unrestricted_filesystem = false

[risk_profiles.hardened]
level = "supervised"
workspace_only = true
allowed_commands = ["git", "ls", "cat"]
forbidden_paths = [
  "/etc",
  "/root",
  "~/.ssh",
  "~/.gnupg",
  "~/.aws",
  "~/.config"
]
require_approval_for_medium_risk = true
block_high_risk_commands = true
max_actions_per_hour = 20
max_cost_per_day_cents = 500
shell_timeout_secs = 60
allowed_tools = ["file_read", "file_list", "time", "memory_search"]
excluded_tools = ["shell"]
sandbox_enabled = true

You can download this tested file directly: zeroclaw-v0.8.5-hardened-security.toml.

Verify the policy against your installed ZeroClaw binary before starting the agent:

zeroclaw --config-dir /path/to/config config migrate --json

Defense-in-depth outside ZeroClaw

The runtime controls above must be reinforced by host-level defenses:

  • Run as a dedicated, unprivileged user: Run ZeroClaw under a service user (zeroclaw) that owns only the agent workspace directory and has no sudo privileges. This OS user boundary is the backstop the model cannot negotiate around.
  • Restrict outbound network traffic: If the agent only requires access to a local Ollama instance or a specific model API, block all other outbound network traffic using host firewall rules (nftables or pfctl).
  • Store backups outside the workspace: Ensure system backups and configuration archives reside on storage inaccessible to the agent user.
  • Review audit logs: Regularly inspect the runtime logs (zeroclaw-events.jsonl or gateway.log) for unexpected tool calls or approval denials.

What none of this fixes

Prompt injection is an architectural property of Large Language Models. When an agent processes untrusted external content (such as internet search results, customer emails, or parsed documents), that content can guide model decisions.

The fundamental rule of agent security: never grant an autonomous agent a capability whose worst-case execution you cannot afford. If an agent has the ability to send emails or delete database records, assume that untrusted input could eventually trigger that behavior. Build your containment boundaries with that assumption in mind.

Related reading