Skip to main content
Redhat Developers  Logo
  • AI

    Get started with AI

    • Red Hat AI
      Accelerate the development and deployment of enterprise AI solutions.
    • AI learning hub
      Explore learning materials and tools, organized by task.
    • AI interactive demos
      Click through scenarios with Red Hat AI, including training LLMs and more.
    • AI/ML learning paths
      Expand your OpenShift AI knowledge using these learning resources.
    • AI quickstarts
      Focused AI use cases designed for fast deployment on Red Hat AI platforms.
    • No-cost AI training
      Foundational Red Hat AI training.

    Featured resources

    • OpenShift AI learning
    • Open source AI for developers
    • AI product application development
    • Open source-powered AI/ML for hybrid cloud
    • AI and Node.js cheat sheet

    Red Hat AI Factory with NVIDIA

    • Red Hat AI Factory with NVIDIA is a co-engineered, enterprise-grade AI solution for building, deploying, and managing AI at scale across hybrid cloud environments.
    • Explore the solution
  • Learn

    Self-guided

    • Documentation
      Find answers, get step-by-step guidance, and learn how to use Red Hat products.
    • Learning paths
      Explore curated walkthroughs for common development tasks.
    • Guided learning
      Receive custom learning paths powered by our AI assistant.
    • See all learning

    Hands-on

    • Developer Sandbox
      Spin up Red Hat's products and technologies without setup or configuration.
    • Interactive labs
      Learn by doing in these hands-on, browser-based experiences.
    • Interactive demos
      Click through product features in these guided tours.

    Browse by topic

    • AI/ML
    • Automation
    • Java
    • Kubernetes
    • Linux
    • See all topics

    Training & certifications

    • Courses and exams
    • Certifications
    • Skills assessments
    • Red Hat Academy
    • Learning subscription
    • Explore training
  • Build

    Get started

    • Red Hat build of Podman Desktop
      A downloadable, local development hub to experiment with our products and builds.
    • Developer Sandbox
      Spin up Red Hat's products and technologies without setup or configuration.

    Download products

    • Access product downloads to start building and testing right away.
    • Red Hat Enterprise Linux
    • Red Hat AI
    • Red Hat OpenShift
    • Red Hat Ansible Automation Platform
    • See all products

    Featured

    • Red Hat build of OpenJDK
    • Red Hat JBoss Enterprise Application Platform
    • Red Hat OpenShift Dev Spaces
    • Red Hat Developer Toolset

    References

    • E-books
    • Documentation
    • Cheat sheets
    • Architecture center
  • Community

    Get involved

    • Events
    • Live AI events
    • Red Hat Summit
    • Red Hat Accelerators
    • Community discussions

    Follow along

    • Articles & blogs
    • Developer newsletter
    • Videos
    • Github

    Get help

    • Customer service
    • Customer support
    • Regional contacts
    • Find a partner

    Join the Red Hat Developer program

    • Download Red Hat products and project builds, access support documentation, learning content, and more.
    • Explore the benefits

Why your non-root container dropped its capabilities (and how to fix it)

A debugging story that will change how you think about container capabilities

October 5, 2026
Nandan Hegde
Related topics:
Containers
Related products:
Red Hat OpenShift Container Platform

    You are a platform engineer running Red Hat OpenShift. A development team runs a monitoring sidecar as a non-root user that needs to perform ICMP ping health checks. They need CAP_NET_RAW, the capability required for raw socket access. Straightforward enough, and the security context constraints (SCC) is configured to allow the capability:

    apiVersion: security.openshift.io/v1
    kind: SecurityContextConstraints
    metadata:
      name: net-raw-scc
    allowedCapabilities:
      - NET_RAW
    requiredDropCapabilities:
      - ALL
    runAsUser:
      type: MustRunAsRange
      uidRangeMin: 1000
      uidRangeMax: 65534
    fsGroup:
      type: MustRunAs
    seLinuxContext:
      type: MustRunAs

    And the PodSpec requests it:

    apiVersion: v1
    kind: Pod
    metadata:
      name: ping-test
    spec:
      containers:
      - name: monitor
        image: registry.example.com/monitor:latest
        securityContext:
          runAsUser: 1000
          capabilities:
            add:
            - NET_RAW
            drop:
            - ALL

    The pod is admitted, the container starts successfully, but inside the container there's this error:

    $ ping -c 1 10.0.0.1
    ping: Operation not permitted

    The developer checks the process capabilities:

    $ cat /proc/1/status | grep Cap
    CapInh: 0000000000000000
    CapPrm: 0000000000000000
    CapEff: 0000000000000000
    CapBnd: 0000000000002000
    CapAmb: 0000000000000000

    CapEff is empty. The capability is not effective, and ping fails with Operation not permitted.

    Everything looks correct. SCC allows it. PodSpec requests it. CRI-O starts the container successfully. But the privilege is not there at runtime.

    If you've never traced the full lifecycle from SCC admission through CRI-O's capability computation, this seems like a bug.

    It is not.

    To understand why, we need to follow the entire security pipeline from the OpenShift API server's admission decision down to the capability evaluation in CRI-O.

    What SCC actually does (and what it does not)

    Security context constraints are OpenShift's mechanism for controlling what security configuration is allowed to enter the cluster. SCC predates Kubernetes Pod Security Admission (PSA) and provides significantly more granular control. While PSA operates on a coarse 3-tier model (privileged, baseline, restricted), SCCs allow precise field-level policies, including which capabilities are allowed, UIDs, volume types, whether privilege escalation is permitted, and so on.

    SCC selection and admission

    When a pod is created, the OpenShift API server's admission controller evaluates it against all SCCs available to the requesting user's service account. For a deeper treatment of SCC management, read Managing SCCs in OpenShift.

    The critical distinction

    Here is what engineers frequently misunderstand: SCC does not enforce runtime privileges. SCC is an admission-time gate. It determines whether a particular security configuration is allowed to be submitted to the cluster. Once the pod passes admission, the SCC's job is done. It has no runtime component. It does not communicate with CRI-O. It does not write iptables rules, adjust kernel parameters, or configure capabilities on processes.

    The actual runtime enforcement happens in an entirely separate stack: SCC admits the pod. The pod security configuration is evaluated and capability sets are derived by CRI-O. The Linux kernel, finally, ensures the capability is enforced during run time.

    How SCC translates into runtime configuration

    CRI-O has no concept of SCCs. It receives a CreateContainerRequest and it translates that into a runtime specification. Let's trace the key SCC-derived fields through this pipeline.

    For SCC fields like allowHostNetwork, allowHostPID, and allowHostIPC, the SCC-to-runtime translation is straightforward. CRI-O receives namespace mode flags and directly configures whether the container joins the host's namespace or gets its own. The mapping is direct and predictable.

    Capabilities are where the translation becomes non-obvious.

    Privilege and capability configuration

    This is where the pipeline becomes consequential for our debugging scenario. SCC fields like allowPrivilegedContainer, allowedCapabilities, defaultAddCapabilities, and requiredDropCapabilities restrict what can appear in the PodSpec.

    CRI-O receives the security configs from CreateContainerRequest and processes them in SpecSetPrivileges():

    func (c *container) SpecSetPrivileges(ctx context.Context, securityContext *types.LinuxContainerSecurityContext, cfg *config.Config) error {
        specgen := c.Spec()
        if c.Privileged() {
            specgen.SetupPrivileged(true)
        } else {
            caps := securityContext.GetCapabilities()
            if err := c.SpecSetupCapabilities(caps, cfg.DefaultCapabilities, cfg.AddInheritableCapabilities); err != nil {
                return err
            }
        }
        // ...
    }

    This is the fork point. Privileged and non-privileged containers receive fundamentally different treatments. To understand what happens in each branch and why our non-root container loses its capability, we need to understand how Linux capabilities actually work at the kernel level.

    How CRI-O handles capabilities

    Linux capabilities decompose the traditional superuser model into distinct privilege units. Capabilities exist in 5 distinct sets, and the kernel evaluates each set differently:

    • Bounding (CapBnd): Upper limit on what capabilities a process can ever acquire.
    • Permitted (CapPrm): The pool from which effective capabilities can be drawn. A capability must be permitted before it can be effective.
    • Effective (CapEff): What the kernel actually checks during privileged operations. If CAP_NET_RAW is absent from this set, creating a raw socket for ping returns Permission Denied regardless of what the other sets contain.
    • Inheritable (CapInh): Capabilities that can carry across execve() only when the executed binary also has matching file inheritable capabilities set. execve() is a system call that replaces the current process with a new program (for example, when a shell runs /usr/bin/ping, it calls execve() to load and execute the ping binary).
    • Ambient (CapAmb): Introduced in Linux 4.3 specifically to solve the non-root capability problem. Ambient capabilities automatically populate both the Permitted and Effective sets after execve(), without requiring file capabilities on the binary.

    For the full specification of these sets, see capabilities(7).

    With this context, let's examine what CRI-O actually writes into each set and why the kernel's treatment of those sets differs between root and non-root processes.

    Privileged containers

    When a container is privileged, CRI-O calls SetupPrivileged(true) on the spec generator. This populates all 5 capability sets with every known capability (41 at the time of writing):

    func (g *Generator) SetupPrivileged(privileged bool) {
        if privileged {
            // ...
            g.Config.Process.Capabilities.Bounding = append(..., finalCapList...)
            g.Config.Process.Capabilities.Effective = append(..., finalCapList...)
            g.Config.Process.Capabilities.Inheritable = append(..., finalCapList...)
            g.Config.Process.Capabilities.Permitted = append(..., finalCapList...)
            g.Config.Process.Capabilities.Ambient = append(..., finalCapList...)
        }
    }

    It is worth noting that dropCapabilities in the PodSpec's security context has no effect (at the time of this writing) when a container runs in privileged mode. The entire add/drop logic is bypassed. All capabilities are unconditionally granted across all 5 sets, regardless of what the PodSpec requests.

    Non-privileged containers

    For non-privileged containers, CRI-O first builds the full list of capabilities to apply. Before processing the PodSpec's add/drop lists, CRI-O merges in a set of 9 default capabilities unless the PodSpec specifies drop: ALL:

    if !addAll && !dropAll {
        caps.AddCapabilities = append(caps.AddCapabilities, defaultCaps...)
    }

    These defaults are defined in CRI-O's configuration and include: CHOWN, DAC_OVERRIDE, FSETID, FOWNER, SETGID, SETUID, SETPCAP, NET_BIND_SERVICE, and KILL. This means that even if a PodSpec adds no capabilities at all, the container still receives these 9 by default. They can be overridden with the default_capabilities setting in crio.conf.

    When PodSpec.SecurityContext.Capabilities specifies drop: ALL, CRI-O clears the entire list first including these defaults and then applies individual adds on top. This is why the drop: ALL along with add: [specific cap] pattern is common in production: It gives you a clean, minimal capability set.

    After building the final list, CRI-O populates three capability sets Bounding, Effective, and Permitted for each capability:

    if err := specgen.AddProcessCapabilityBounding(capPrefixed); err != nil {
        return err
    }
    if err := specgen.AddProcessCapabilityEffective(capPrefixed); err != nil {
        return err
    }
    if err := specgen.AddProcessCapabilityPermitted(capPrefixed); err != nil {
        return err
    }

    The capability section created by CRI-O is identical regardless of whether the container runs as root or non-root. But the runtime outcome is drastically different.

    Root user (UID 0): Capabilities work

    There is no UID transition. The process starts as UID 0 and stays as UID 0. The kernel never clears any capability sets. What CRI-O writes into the OCI spec is exactly what the process ends up with: CapEff contains the requested capabilities.

    Non-root user (UID != 0): Capabilities vanish

    The OCI runtime starts the process as root, applies the capability sets, then eventually switches to the target UID (for example, 1000). During this UID transition from 0 to non-zero, the kernel clears the Permitted and Effective sets:

    • Capbnd: Preserved
    • Capprm: Cleared by kernel
    • Capeff: Cleared by kernel
    • Capinh: Empty (never set by CRI-O)
    • Capamb: Empty (never set by CRI-O)

    The only mechanism that survives a UID transition and populates the Effective set for a non-root process is ambient capabilities. CRI-O intentionally does not set them:

    // Make sure to remove all ambient capabilities. Kubernetes is not yet ambient capabilities aware
    // and pods expect that switching to a non-root user results in the capabilities being
    // dropped. This should be revisited in the future.

    This is a deliberate security decision: Kubernetes assumes runAsNonRoot drops capabilities. Setting ambient capabilities would break that contract.

    There is an existing enhancement proposal to add ambient capability support to Kubernetes, which would allow selectively granting effective capabilities to non-root container processes.

    What can you actually do about it?

    Knowing that the kernel clears CapEff during the UID transition, and that CRI-O intentionally omits ambient capabilities, what options does an engineer have to make CAP_NET_RAW work for UID 1000?

    File capabilities (setcap)

    This is the most targeted approach. Apply the capability directly to the binary during image build:

    setcap 'cap_net_raw=+ep' /usr/bin/ping

    This sets file-level Permitted and Effective bits. When the kernel executes this binary, it sees the file capability and promotes it into the process's Permitted and Effective sets, even for non-root users as long as the capability remains in the Bounding set (which it does, as we saw in the CapBnd: 0000000000002000 output at the very beginning of this debugging story).

    Tradeoff: Requires modifying the container image. File capabilities are stored as extended attributes, so your build pipeline and storage driver must preserve them. Also, because any capability can be embedded in the image filesystem, this shifts the trust boundary to image provenance. You need to trust that the image hasn't been tampered with.

    Running as root

    Running as UID 0 avoids the UID transition entirely. No transition, no clearing.

    Tradeoff: The most common SCC in production (restricted-v2) prohibits this. Running as root while explicitly dropping all unnecessary capabilities can reduce the exposure.

    Privileged containers

    Privileged mode populates all 5 sets including Ambient, and disables SELinux, AppArmor, and seccomp.

    Tradeoff: This is using a sledgehammer for a thumbtack. As we saw in the CRI-O code, privileged mode bypasses the entire capability add/drop logic. You cannot selectively restrict a privileged container. It gets everything unconditionally.

    Conclusion

    Let's return to our debugging scenario. Your SCC was correct. The PodSpec was correct. CRI-O did exactly what it was supposed to, and wrote CAP_NET_RAW into the Bounding, Effective, and Permitted sets of the OCI spec. But none of that mattered, because:

    • SCC admitted the pod and its job ended there.
    • CRI-O translated the CRI request into an OCI spec with the capability in 3 sets, but not in Ambient.
    • The OCI runtime applied the spec, then switched to UID 1000.
    • The kernel cleared Permitted and Effective during the UID transition.
    • The process started with CapEff: 0000000000000000.

    Once you understand where SCC ends and where Linux capability semantics begin, container privilege debugging becomes far less mysterious. The kernel is not ignoring SCC, it's faithfully applying capability transition rules exactly as designed.

    The next time you see CapEff: 0000000000000000 in a non-root container, you now know exactly where to look.

    Related Posts

    • Configure a pod security context with Cryostat Operator

    • Understanding OpenShift Security Context Constraints

    • Harden local container base images in Podman Desktop

    • Building rootless containers for JavaScript front ends

    • Rootless containers with Podman: The basics

    Recent Posts

    • Why your non-root container dropped its capabilities (and how to fix it)

    • Add NeMo Guardrails to a LangGraph agent on OpenShift AI

    • Benchmarking AI decision models against traditional guardrails

    • Kube AuthKit: Unified Kubernetes and OpenShift auth in Python

    • Smarter GPU sharing: How Red Hat build of Kueue works with dynamic resource allocation

    What’s up next?

    Learning Path red_hat_build_of_podman_desktop feature image

    Build and run a bootable container image with image mode for RHEL and Podman Desktop

    Learn how to locally build and run a bootable container (bootc) image in...
    Red Hat Developers logo LinkedIn YouTube Twitter Facebook

    Platforms

    • Red Hat AI
    • Red Hat Enterprise Linux
    • Red Hat OpenShift
    • Red Hat Ansible Automation Platform
    • See all products

    Build

    • Developer Sandbox
    • Developer tools
    • Interactive tutorials
    • API catalog

    Quicklinks

    • Learning resources
    • E-books
    • Cheat sheets
    • Blog
    • Events
    • Newsletter

    Communicate

    • About us
    • Contact sales
    • Find a partner
    • Report a website issue
    • Site status dashboard
    • Report a security problem

    RED HAT DEVELOPER

    Build here. Go anywhere.

    We serve the builders. The problem solvers who create careers with code.

    Join us if you’re a developer, software engineer, web designer, front-end designer, UX designer, computer scientist, architect, tester, product manager, project manager or team lead.

    Sign me up

    Red Hat legal and privacy links

    • About Red Hat
    • Jobs
    • Events
    • Locations
    • Contact Red Hat
    • Red Hat Blog
    • Inclusion at Red Hat
    • Cool Stuff Store
    • Red Hat Summit
    © 2026 Red Hat

    Red Hat legal and privacy links

    • Privacy statement
    • Terms of use
    • All policies and guidelines
    • Digital accessibility
    Ask AI