Runc

2 posts

datadogOriginal article

Using the Dirty Pipe vulnerability to break out from containers | Datadog (opens in new tab)

The Dirty Pipe vulnerability (CVE-2022-0847) is a critical Linux kernel flaw that allows unprivileged processes to write data to any file they can read, effectively bypassing standard write permissions. This primitive is particularly dangerous in containerized environments like Kubernetes, where it can be leveraged to overwrite the host’s container runtime binary. By exploiting how the kernel manages page caches, an attacker can achieve a full container breakout and gain administrative privileges on the underlying host. ## Container Runtimes and the OCI Specification * Kubernetes utilizes the Container Runtime Interface (CRI) to manage containers via high-level runtimes like containerd or CRI-O. * These high-level runtimes rely on low-level Open Container Interface (OCI) runtimes, most commonly runC, to handle the heavy lifting of namespaces and control groups. * Isolation is achieved by runC setting up a restricted environment before executing the user-supplied entrypoint via the `execve` system call. ## Evolution of runC Vulnerabilities * A historical vulnerability, CVE-2019-5736, previously allowed escapes by overwriting the host’s runC binary through the `/proc/self/exe` file descriptor. * To mitigate this, runC was updated to either clone the binary before execution or mount the host's runC binary as read-only inside the container. * While the read-only mount improved performance through kernel cache page sharing, it created a target for the Dirty Pipe vulnerability, which specifically targets the kernel page cache. ## The Dirty Pipe Exploitation Primitive * Dirty Pipe allows an attacker to overwrite any file they can read, including read-only files, by manipulating the kernel's internal pipe-buffer structures. * The exploit targets the page cache, meaning the overwrite is non-persistent and resides only in memory; the original file on disk remains unchanged. * In a container escape scenario, the attacker waits for a runC process to start (triggered by actions like `kubectl exec`) and targets the file descriptor at `/proc/<runC-pid>/exe`. ## Proof-of-Concept Escape Walkthrough * The attack begins with a standard, unprivileged pod running a malicious script that monitors the system for new runC processes. * Once a `kubectl exec` command is issued by an administrator, the script identifies the runC PID and applies the Dirty Pipe exploit to the associated executable. * The exploit overwrites the runC binary in the kernel page cache with a malicious ELF binary. * Because the host kernel is executing this hijacked binary with root privileges to manage the container, the attacker’s malicious code (e.g., a reverse shell or administrative command) runs with full host-level authority. To protect against this attack vector, it is essential to patch the Linux kernel to a version that includes the fix for CVE-2022-0847 and ensure that container nodes are running updated distributions.

datadog3 min readCurated summary

Escaping containers using the Dirty Pipe vulnerability | Datadog Security Labs

The post demonstrates how the Linux Dirty Pipe vulnerability can enable an unprivileged process to escape a container and gain administrative privileges on the host. The exploit abuses runC’s execution model and its host binary, which is exposed read-only inside the container but can still be modified through the kernel page cache. A proof of concept shows how a compromised Kubernetes pod can overwrite runC with a malicious executable when an administrator runs `kubectl exec`. ## Container Runtimes and runC - Kubernetes commonly uses containerd or CRI-O through the Container Runtime Interface (CRI). - These runtimes rely on lower-level OCI runtimes, most notably runC, to create isolated Linux processes. - runC configures namespaces, cgroups, and the container environment before executing the supplied entrypoint with `execve`. - During execution, `/proc/self/exe` inside the container can refer to an open descriptor for the runC binary on the host. ## Earlier runC Escape Vulnerability - CVE-2019-5736 exploited this `/proc/self/exe` behavior: - A malicious container entrypoint could write to the host’s runC binary. - Overwriting runC enabled subsequent container operations to execute attacker-controlled code with host-level privileges. - runC initially mitigated the issue by cloning its binary before execution. - It later changed the design to mount the runC binary read-only inside the container, improving performance through kernel page-cache sharing. - That optimization created conditions in which Dirty Pipe could bypass the apparent read-only protection. ## Dirty Pipe as a Container Escape Primitive - Dirty Pipe allows an unprivileged process to overwrite files it can read, even without write permission. - The modification occurs in the kernel page cache rather than persistent storage: - The original file remains intact on disk. - Dropping caches or rebooting can restore the original contents. - Despite being temporary, the overwrite is sufficient to execute malicious code when the modified binary is run. - In this case, the attacker targets the host’s runC binary through `/proc/<runC-pid>/exe`. ## Kubernetes Proof of Concept - The demonstration starts an ordinary, unprivileged pod using an attacker-controlled container image. - Its entrypoint script: - Replaces `/bin/sh` with a launcher referencing `/proc/self/exe`. - Waits for a runC process to appear. - Invokes the Dirty Pipe exploit against that process’s executable. - An administrator running `kubectl exec` causes runC to execute inside the container, triggering the overwrite. - The modified runC is replaced with a malicious ELF binary that runs commands such as `id` and `hostname`, recording their output in `/tmp/hacked`. - The exploit is adapted from the original Dirty Pipe proof of concept and the earlier runC escape technique. The attack illustrates that kernel vulnerabilities can undermine container isolation even when the container is unprivileged and the target binary is mounted read-only. Systems should promptly patch vulnerable Linux kernels and container runtimes, while treating compromised containers as potential paths to host compromise.

Read original(opens in new tab)