CL-0022: tmpfs exec — “Permission denied” running from tmpfs¶
Severity: LOW
Derivation (see severity model):
- Baseline: A — the attacker already has code execution in this container, as the workload uid
- Precondition: Removes a mitigation —
execandsuidon a tmpfs grant no privilege. They remove two defaults Docker applies to every tmpfs - Impact: Single container
- Qualifier/modifier: none
- Derived: Removes a mitigation × Single container = LOW
- Shipped: LOW
- Scoping assumptions: scored with
read_only: truepresent (the secure default of CL-0007) and a non-root workload. Withoutread_only, the rootfs is itself writable and executable and the tmpfs flag protects nothing; as root,/devis already a writable tmpfs withoutnoexec - Evidence:
_cl0022and_t22_exec_tmpfs— a tmpfs isnoexecby default and:execremoves it, re-proven against a live container on every CI run._cl0022_dev_inertproves the converse for thedevoption this rule deliberately does not flag (see below). Thesuidleg is a captured observation, not a check: with:exec,suid, a setuid-rootbashstaged by root runs aseuid=0fornobody; with:execalone it does not (Docker 29.1.3,debian:bookworm-slim, 2026-08-23)
References:
- OWASP Docker Security Rule #8 — cited for the read_only + tmpfs pattern this rule protects. It recommends the pattern; it says nothing about mount options. The option semantics are grounded by the live checks above, under ADR-016's runtime arm
- Docker docs — tmpfs mounts — the option vocabulary
What it detects¶
A tmpfs: entry that explicitly passes exec or suid — each of which removes a security default Docker applies to every tmpfs:
tmpfs:
- /tmp:exec # flagged: re-enables execution (default is noexec)
- /tmp:suid # flagged: re-enables setuid (default nosuid)
- /scratch:exec,suid # flagged: names both
- /tmp # clean: gets noexec,nosuid,nodev by default
- /run:size=64m # clean: size doesn't remove the defaults
- /tmp:noexec # clean: that's the secure value
- /run:dev # clean: removes nodev, which grants nothing (below)
Matching is on whole comma-separated tokens, so the secure noexec is never mistaken for exec. The long volumes: [{type: tmpfs}] form is out of scope — it keeps the secure defaults and can't express these tokens.
Why it matters¶
Docker mounts every tmpfs with noexec,nosuid,nodev by default — verified across the short, list, and long mount forms, and the defaults survive even when other options like size= are set. The only way to lose them is to pass the opposite option explicitly:
| Option | Removes | Effect | Flagged |
|---|---|---|---|
exec |
noexec |
binaries can be executed from the mount | yes |
suid |
nosuid |
setuid/setgid bits are honored | yes |
dev |
nodev |
device nodes on the mount can be opened — by the mount; the device cgroup still refuses them | no |
A tmpfs is always writable and in-memory. Under read_only: true (CL-0007) it is often the only writable path, and noexec is what stops the commodity playbook — download to /tmp, chmod +x, run. Re-enabling execution hands that path back. suid additionally lets a setuid-root binary staged on the mount by a root process run as root for any other uid — a persistence vector on an image whose entrypoint drops privileges, and one no-new-privileges (CL-0003) closes. Because the secure posture is the default, an explicit exec/suid is a deliberate weakening that should be justified.
This is why the rule flags the presence of these options rather than the absence of noexec/nosuid: a plain tmpfs: [/tmp] is already hardened by Docker, so flagging it would be noise.
What noexec does not stop. The mitigation is shallow, which is why the rule is LOW and not higher. Measured on Docker 29.1.3 at defaults: memfd_create + execve runs a dropped binary without touching any mount, for any uid, under the default seccomp profile; interpreters read scripts from a noexec mount as usual; and a root workload always has /dev — a writable tmpfs mounted rw,nosuid with no noexec, even under read_only: true. The ld.so loader trick, by contrast, is dead on both libcs (musl refuses; glibc cannot mmap PROT_EXEC from a noexec mount). noexec narrows a non-root attacker's options; it does not remove them.
Why dev is not flagged. Removing nodev changes the error, not the outcome. At default capabilities a root process can mknod (MKNOD is a Docker default), and opening the node is then refused by the device cgroup — Operation not permitted — on a tmpfs:dev, on the rootfs (mounted rw,relatime, no nodev), and in /dev (no nodev either). Where the cgroup is off (privileged: true, CL-0002), /dev already permits the node, so the tmpfs adds nothing there too; a non-root workload cannot mknod anywhere. Two further legs were measured and agree: a non-root workload cannot mknod at all on a tmpfs:dev (EPERM, even for an allow-listed node, and even with cap_add: MKNOD — a non-root process's CapEff is empty), and for the devices Docker does allow-list the option changes nothing that matters — /dev/net/tun created on a tmpfs:dev opens and then refuses TUNSETIFF with EPERM exactly as the same node created in /dev does, with NET_ADMIN (CL-0011) the gate in both cases. A finding on dev would describe a configuration that changes nothing — the failure mode that removed CL-0023. _cl0022_dev_inert re-proves the cgroup refusal on both legs on every CI run, and ADR-028 records the decision.
Fix¶
Remove the offending option to fall back to Docker's secure default:
# Before
tmpfs:
- /tmp:exec
# After
tmpfs:
- /tmp
compose-lint fix makes this edit. It removes the exec and suid tokens from the flagged entry and leaves every other option in place (/run:exec,size=64m becomes /run:size=64m); when nothing survives, the colon goes too. Both spellings the rule checks are handled — a list item and the scalar tmpfs: /tmp:exec form. The fix is behavior-changing and says so in the dry run: a workload that really executes a binary, or relies on a setuid bit, from that mount fails once the default is back. That is the case where the option was set deliberately, and the answer there is a suppression with a reason, not the fix.
It refuses, and leaves the entry for you, when the service is built from a YAML anchor or merge key, when the list is written in flow style (tmpfs: [/tmp:exec]), when the entry carries ${VAR} interpolation, or when the line does not show the whole value (a wrapped or block scalar). See the fix contract.
Reading the failure¶
If you arrived here from an error, it is probably this one — a binary or script refusing to run from a tmpfs. The busybox wording is re-proven against a live container on every CI run (scripts/validate_rule_premises.py); the exact prefix varies by shell.
| Symptom in logs (verbatim) | What it means | Action — in this order |
|---|---|---|
sh: line 0: /scratch/busybox: Permission denied executing a file that is executable (x bit set) from a tmpfs |
the mount's default noexec blocked the exec — not a file-permission problem |
1. relocate the executable into the image (bake it in at build time) or onto a named volume; 2. only if the workload genuinely must write-then-execute on that path, add :exec — which is exactly what this rule flags, so pair it with a suppression carrying a documented reason |
The trap is that the obvious fix (:exec) is the finding: a writable, executable, in-memory mount is the classic staging ground for a dropped payload, and doubly so under read_only: true (CL-0007), where tmpfs is often the only writable path. Treat :exec as a justified exception, not a fix. A related red herring: chmod +x appears to succeed on the tmpfs file and changes nothing — noexec is a property of the mount, not the file.
ATT&CK coverage¶
No adversary technique maps to this rule. That is a finding about the rule, not an omission — see below.