Skip to content

CL-0016: Dangerous host devices exposed via devices: or device_cgroup_rules:

Severity: CRITICAL

Derivation (see severity model):

  • Baseline: A — the attacker already has code execution in this container, as the workload uid
  • Precondition: Direct — a mapped block device is read and written with ordinary file operations, at default capabilities. No technique, no second defect. A device_cgroup_rules: entry is Direct under one stated condition: the container can obtain the node, through Docker's default MKNOD capability or a /dev bind mount. Without either, the rule permits a device the container cannot open, and it is not flagged
  • Impact: Host — raw access to the disk backing the host filesystem reads and rewrites host state regardless of file permissions
  • Qualifier/modifier: none
  • Derived: Direct × Host = CRITICAL
  • Shipped: CRITICAL
  • Evidence: _cl0016 — devices: exposes the host device, verified live. The grant is capability-independent: a raw read of the host's NVMe device through --device succeeded at default capabilities, where the same read through a /dev bind mount was refused by the device cgroup. _cl0016_cgroup_rule — a b <major>:* r rule plus mknod reads the host disk at default capabilities; _cl0016_cgroup_rule_gates — dropping MKNOD or granting only m reads nothing, while a /dev bind mount reads the disk at --cap-drop ALL

References: - CIS Docker Benchmark 5.18 — Ensure that host devices are not directly exposed to containers

What it detects

Services that expose dangerous host devices via devices:. The host path is normalized before matching, so equivalent spellings of the same device node are all detected — /dev/sda, //dev/sda, /dev/./sda and /dev/../dev/sda name one device to the kernel and Compose passes each through verbatim.

Both devices: syntaxes are read: the short host:container[:permissions] string and the long mapping, whose source: is the host path — the form docker compose config itself renders. The /dev directory is flagged too: Docker walks a directory source and maps every device node beneath it, so /dev grants every disk the host has. A symlink named directly (/dev/disk/by-id/..., /dev/mapper/<name>) is resolved to its disk and flagged by the rows below.

Pattern Risk
/dev The whole device directory — every host disk at once
/dev/sd* SCSI/SATA block devices — raw disk read/write
/dev/nvme* NVMe block devices — raw disk read/write
/dev/vd* virtio block devices — the host root disk on KVM and Proxmox guests
/dev/xvd* Xen block devices — the host root disk on EC2
/dev/mmcblk* SD/eMMC block devices — the host root disk on a Raspberry Pi
/dev/md* Linux software RAID arrays
/dev/disk/* Block device symlinks — same risk as raw devices
/dev/loop* Loop devices — mount arbitrary disk images
/dev/dm-*, /dev/mapper/* Device mapper — encrypted/LVM volumes, raw block access
/dev/zfs ZFS pool control — create/destroy/import any pool the host can see
/dev/rbd* Ceph RBD — remote block device access
/dev/zd*, /dev/zvol/* ZFS zvols — a VM's whole virtual disk on Proxmox and TrueNAS
/dev/nbd* Network block devices — raw read/write of whatever is attached
/dev/mtdblock* MTD flash — the root filesystem on embedded boards
/dev/hd* Legacy IDE block devices — raw disk read/write
/dev/kmsg Kernel log buffer — read or inject kernel messages

Safe devices like /dev/net/tun are not flagged.

device_cgroup_rules:

A device cgroup rule opens the same gate a devices: entry does, without mapping a node. The container supplies the node itself: Docker's default MKNOD capability lets it run mknod /dev/sda b 8 0, and a bind mount of /dev (or of a node beneath it) hands over the host's node directly. Either way the rule turns into a raw read or write of the disk, which is the same outcome and the same severity as a mapped disk.

Rule Flagged Why
a *:* rwm Yes Every device, every disk included
b 8:* rwm, b 259:0 r, b *:* w Yes A block device read or written. Any major counts: NVMe, device-mapper, zvol and nbd disks sit on dynamically allocated majors, so a table of disk majors would miss the host disk itself
b 8:* m No m permits creating the node, not reading or writing it
c ... No Character devices are not claimed by this key; the devices: rows above cover /dev/kmsg

A rule is flagged only when the container can obtain the node: it keeps MKNOD (not dropped, or restored by cap_add), or it bind-mounts /dev or a path beneath it. A service that drops MKNOD and mounts nothing from /dev holds a rule with no node to use. All of these were measured on a live daemon at Docker's defaults; the premise checks named above re-run the decisive ones on every CI run.

The finding's evidence is the rule's type and numbers (b 8:*), without the access letters: r alone is already a full read, so narrowing rwm to r is the same finding. A service with both devices: [/dev/sda] and b 8:* rwm gets two findings, because each grant has to be removed on its own.

Devices this rule deliberately does not claim

A device that is live only alongside a capability another rule already flags belongs to that rule (ADR-020): it could otherwise fire only beside a strictly higher finding, or alone on a configuration that grants nothing. All verified at default capabilities on Docker 29.1.3:

Device Why not here
/dev/mem, /dev/port Operation not permitted without CAP_SYS_RAWIO, which CL-0024 flags CRITICAL. /dev/mem is further bounded to the sub-1MB region even with the capability, because CONFIG_STRICT_DEVMEM refuses the rest — offsets at 1 MiB and 4 GiB both fail
/dev/fuse mount(2) needs CAP_SYS_ADMIN (CL-0024, CRITICAL), which is the gate that holds on every host. On an AppArmor host it also needs an unconfined profile (CL-0009) — measured, SYS_ADMIN alone mounts where no AppArmor policy is loaded, so that second gate is posture-specific (ADR-020) and the drop does not rest on it
/dev/kmem, /dev/raw Docker refuses to create the container at all, so a finding could never describe a running service. CONFIG_DEVKMEM is also off on modern kernels
/dev/disk, /dev/mapper, /dev/md as whole directories Docker's directory walk maps real nodes and skips symlinks, and these hold symlinks. Measured on two hosts: --device /dev/disk is refused (not a device node), and --device /dev/mapper maps only control, whose ioctls need CAP_SYS_ADMIN (CL-0024). /dev/md was not measurable (no software RAID) and is left unclaimed rather than assumed. A symlink inside them, named directly, is resolved and stays flagged. Premise check: _cl0016_symlink_dirs_grant_no_disk

/dev/kmsg stays despite needing CAP_SYSLOG on a host with kernel.dmesg_restrict=1: that is a host sysctl whose upstream default is 0, where the read needs no capability. (Granting cap_add: SYSLOG is CL-0030's finding; this row is the device node at the upstream default, where no capability is involved.)

/dev/loop* stays too, and for a stronger reason than it was first given. The loop family was assumed to need mount(2), and so CAP_SYS_ADMIN — which would have grouped it with /dev/fuse above. That is wrong for the step that matters. Captured observation (Docker 29.4.3, no premise check — see below): a container given only --device /dev/loop-control, at --cap-drop ALL, issued LOOP_CTL_GET_FREE and allocated a loop device on the host. Attaching a backing file is a further step, but the allocation itself needs no capability, which puts the loop family on the same footing as the raw block devices above. The pattern matches /dev/loop-control as well as /dev/loop0, which is what this observation exercised.

It is a capture rather than a check on purpose: automating it would create and remove a host loop device on every CI run, and the premise suite is non-destructive by design. (The removal path works — LOOP_CTL_REMOVE restored the host — it is simply not something to do on a schedule.)

Why it matters

Exposing raw memory or block devices bypasses all container isolation. Block device access enables reading and overwriting the host filesystem directly, regardless of mount permissions; loop and device-mapper access lets a container construct arbitrary block devices and mount them. ZFS and Ceph control devices grant pool/cluster-level operations that affect everything that runtime can see.

These are equivalent to giving the container full root access to the host's storage — and unlike most escape paths, this one is capability-independent. Verified on Docker 29.1.3: a container with Docker's default capabilities, holding no cap_add at all, read the first sector of the host's NVMe device through a devices: mapping. Nothing in the default capability set, the default seccomp profile, or docker-default AppArmor stands in the way; the device cgroup is the gate, and devices: is what opens it.

Fix

Remove the dangerous device. If you need specific hardware access, prefer narrower mechanisms:

  • device_cgroup_rules: — express exactly which devices the container may access, kept to the specific character device the workload needs. A rule that opens a block major (b 8:* rwm), or every device (a), is flagged exactly like the disk it grants — see above
  • A single specific device entry (e.g. /dev/ttyUSB0) instead of a wildcard or whole subtree
  • Kernel-level offload (e.g. user-space NVMe via SPDK) where a container shouldn't need block-device access at all

Where the container only needs the data on a disk rather than the disk itself, mount the filesystem instead:

# Before
services:
  app:
    image: myapp:1.0
    devices:
      - /dev/sda:/dev/sda

# After — use volumes for disk data access
services:
  app:
    image: myapp:1.0
    volumes:
      - /data:/app/data:ro

If raw device access is genuinely required (e.g., for disk management tooling), the workload should not run in a container.

ATT&CK coverage

Remediating this finding contributes to mitigating the following MITRE ATT&CK techniques (pinned to ATT&CK v18). compose-lint is a static analyser, so this is mitigation coverage — it detects nothing at runtime.

Technique Tactic
T1611 Escape to Host Privilege Escalation
T1552.001 Unsecured Credentials: Credentials In Files Credential Access

See also

  • CL-0013 — bind-mounting /dev whole (parallel risk at the path-mount level)
  • CL-0002 — privileged: true (grants every device the host has)