This page states what drukbox protects, what it deliberately does not, and the tradeoffs behind each default. For how sandboxes are reached read Networking; for the configuration knobs named here read Deploy.
Drukbox has one trust tier. A valid service token is full control — it can create, list, get, and delete any host and manage every HTTP proxy. There is no per-token scoping, per-tenant isolation, or ownership check between token holders. Treat every token as an operator-level credential.
The asymmetry that is part of the model: callers are trusted, the sandboxes they provision are not. Drukbox hands back SSH coordinates and stops — it never runs a sandbox's code and owns no runtime inside the VM. The hardening below is about keeping an untrusted workload on a sandbox from reaching back into drukbox's credentials or its cloud account, not about isolating one token holder from another.
This is the right model for a single team standing up sandboxes behind their own API. It is not a multi-tenant boundary: do not hand drukbox tokens to mutually distrusting users.
Every endpoint except GET /healthz requires
Authorization: Bearer <service-token>. Tokens come from
SERVICE_TOKENS (comma-separated) and are compared in constant time,
so a wrong token leaks no timing signal. /healthz is unauthenticated
by design and returns only {"status": "ok"} — no version, config, or
dependency detail. /doctor is authenticated and is the only endpoint
that reports dependency state.
Rotate a token by adding the new value to SERVICE_TOKENS, moving
callers over, then dropping the old one. Multiple tokens are accepted
at once precisely so rotation needs no downtime.
The API holds provider credentials and mints cloud resources, so the
process is a high-value target. It binds 0.0.0.0 by default for
container friendliness. When only a co-located or host-networked
caller reaches it, set UVICORN_HOST=127.0.0.1 to keep the control
plane off other interfaces. When it must be remote, front it with TLS
termination and treat the token as the only thing standing between the
internet and your cloud account — drukbox does no TLS itself and adds
no rate limiting (see Resource exhaustion).
How a caller reaches a sandbox, and the tradeoffs of each path, are covered in Networking. The security-relevant summary:
- Per-VM keys. On AWS (Tailscale off) and Hetzner, drukbox mints a
fresh ed25519 keypair per VM, returns the private half once in
the create response, and never persists it — a later
GET /hosts/{id}returnsprivate_key: null. The key is the auth boundary; password auth is never enabled. - AWS ingress fail-open. The managed
drukbox-managedsecurity group opens SSH to the detected egress/32, or to whateverAWS_SSH_CIDRSspecifies. If egress detection fails and no CIDRs are set, ingress falls back to0.0.0.0/0with a warning log. The per-VM key remains the boundary in that case, but setAWS_SSH_CIDRSexplicitly in any environment where world-open port 22 is unacceptable. - Hetzner has no firewall. A fresh server exposes port 22 to the internet; the per-VM key is the only boundary. There is no ingress configuration to manage.
- First-keyscan MITM window. With Tailscale off, the
known_hostsmaterial is scanned over the public network and carries the usual trust-on-first-use window. Enable Tailscale to run the scan over the authenticated overlay.
Provider tokens (EXE_API_TOKEN, HETZNER_API_TOKEN, Tailscale OAuth)
and AWS credentials are read from the environment / the AWS SDK default
chain and never written to the database or returned by the API. Caller
env is write-only: it is delivered to the VM but never echoed in any
response, and reserved keys (TAILSCALE_AUTHKEY) are rejected at the
schema.
Two pieces of material reach the VM through its provider's user-data / setup-script mechanism, and that channel is the relevant exposure:
- Tailscale auth key. Minted per host, ephemeral, tag-scoped, and short-lived; it is not persisted in drukbox's database. It is delivered to the VM via user-data, so a process on the box can read it — acceptable given its single-use, ephemeral nature.
- AWS IMDS. Sandboxes run untrusted code, so launched EC2
instances require IMDSv2 (
HttpTokens: required) with a put-response hop limit of 1. This stops an in-VM SSRF or stray process from reading instance metadata over the legacy unauthenticated IMDSv1 path — which would otherwise expose the user-data auth key and, ifAWS_INSTANCE_PROFILEis set, live IAM role credentials. The hop limit keeps a containerized workload one network hop from the endpoint; if you run a sandbox payload in a container that genuinely needs IMDS, raise it deliberately. Residual: a local root on the box can still read its own user-data, so scopeAWS_INSTANCE_PROFILEtightly (or leave it unset) and keep nothing in callerenvthat the sandbox workload should not see.
Provisioning failures are stored on the host as a concrete summary
(exception type and message), not a raw Python traceback. That summary
is what HostOut.last_error and the POST /hosts 502 detail return to
callers; the full traceback stays in the server log only. Keep log
sinks access-controlled — they hold the detail the API withholds.
Drukbox adds no quota or rate limit of its own: a valid token can
provision paid VMs without bound, so a leaked token is a cost-DoS as
well as a control-plane compromise. The controls that exist are
operational — expires_at plus the janitor reap idle hosts,
PROVISIONING_GRACE_SECONDS bounds strands, and
POOL_MAX_CREATES_PER_TICK caps pool over-provision. Put per-caller
quotas and rate limiting in the layer that issues and fronts tokens.
- A service token can delete any host. There is no second factor for destructive calls — the token is the boundary.
- Drukbox never opens an SSH session, runs sandbox code, or creates Linux users. Everything past the returned SSH coordinates is the caller's responsibility.
private_keyappearing once in the create response is intentional; callers must capture it then, because it is never recoverable later.
See SECURITY.md for private reporting.