Skip to content

Memory Exhausting in a container causes performance issue for other containers on the same host #431

Description

@horner

> User: I am interested in why high memory in a single container with only 4gb or ram would cause this load spike in proxmox

Memory and CPU pressure

GitHub Copilot: Short answer: because Linux "load" isn't CPU usage — it counts both runnable tasks and tasks stuck in uninterruptible sleep (D-state), and a container thrashing against its memory cgroup limit generates lots of both.

What happens when an LXC container hits its 4GB cap:

  1. Cgroup memory pressure → direct reclaim. LXC memory limits are enforced by the host kernel via cgroups. When the container approaches 4GB, every allocation inside it triggers direct reclaim — the allocating process itself runs kernel reclaim code on the host CPU. That's real host CPU/kernel work, not contained "inside" the CT.

  2. Swap thrashing / page cache refaults. The kernel starts swapping the container's pages out and dropping its page cache. Processes then immediately fault those pages back in. Each fault means synchronous disk I/O, and the process waits in D-state. Load average counts every one of those blocked tasks, so a container with 20 thrashing threads can add ~20 to the node's load while CPU utilization looks modest. (This matches your graph: tall load spikes in a plateau/spike pattern typical of thrash cycles, not sustained CPU work.)

  3. Kernel threads run at host level. kswapd, writeback threads, and (if you use ZFS) ARC reclaim/arc_prune all do the shoveling on behalf of the container — that work shows up as node load, not container CPU.

  4. OOM-kill cycles. If the workload keeps growing, the cgroup OOM killer kills it, the service restarts, re-allocates, and the cycle repeats — producing repeated spikes like the separated humps in your chart.

How to confirm on the node:

# Per-cgroup memory pressure (PSI) — the smoking gun
cat /sys/fs/cgroup/lxc/<CTID>/memory.pressure
cat /proc/pressure/memory /proc/pressure/io

# Swap-in/out activity during a spike
vmstat 1        # watch si/so columns

# D-state processes during a spike
ps -eo state,pid,comm | awk '$1=="D"'

# OOM events for that container
journalctl -k | grep -i oom

Fixes: raise the CT's memory limit, reduce or disable the CT's swap (swap: 0 in the CT config stops thrash-via-swap and fails fast to OOM instead), or fix the app's memory footprint. If the node uses ZFS, also check that ARC isn't competing for the same RAM (arc_summary, cap via zfs_arc_max).

If you share the Proxmox browser tab or give me SSH access to the node from this terminal, I can help dig into the actual PSI/OOM evidence.

> User: lets get into opensource-phxdc-pve1 and look at what is going on.

GitHub Copilot: Ran terminal command: ssh -o ConnectTimeout=10 -o BatchMode=yes root@opensource-phxdc-pve1 "hostname && uptime && free -h"

Ran terminal command: ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new root@opensource-phxdc-pve1 "hostname && uptime && free -h"

The SSH session is waiting for the root password. Please type it directly into the terminal (I can't handle secrets) — then I'll continue investigating.

> User: I just sudo su -

GitHub Copilot: Run in Terminal

> User: now we can jump betwen hosts

GitHub Copilot: Ran terminal command: whoami && ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new root@opensource-phxdc-pve1 "hostname && uptime && free -h && echo --- && cat /proc/pressure/memory"

Ran terminal command: ps -eo state,pid,ppid,wchan:30,comm | awk '$1=="D"' | head -40; echo ===; ps -eo state | grep -c D

Ran terminal command: for d in /sys/fs/cgroup/lxc/*/; do ctid=$(basename $d); p=$(grep '^full' $d/memory.pressure 2>/dev/null | awk '{print $2}' | cut -d= -f2); [ -n "$p" ] && echo "$ctid full_avg10=$p"; done | sort -t= -k2 -rn | head; echo ===; top -bn1 -o %CPU | head -25

Found it: container 392 is at 83% memory-stall (full_avg10=83.63) — it's the one thrashing. Let's see what's inside it.

Ran terminal command: pct config 392 | head -20; echo ===; cat /sys/fs/cgroup/lxc/392/memory.current /sys/fs/cgroup/lxc/392/memory.max /sys/fs/cgroup/lxc/392/memory.swap.current 2>/dev/null; echo ===; grep -E 'oom|refault' /sys/fs/cgroup/lxc/392/memory.events /sys/fs/cgroup/lxc/392/memory.stat 2>/dev/null | grep -Ev ' 0$' | head

Ran terminal command: export PATH=$PATH:/usr/sbin; pct config 392; echo ===; grep -cE '^oom' /sys/fs/cgroup/lxc/392/memory.events; cat /sys/fs/cgroup/lxc/392/memory.events

Ran terminal command: cat /etc/pve/lxc/392.conf 2>&1 | head; echo ===; systemctl is-active pve-cluster; echo ===; for p in $(cat /sys/fs/cgroup/lxc/392/ns/cgroup.procs 2>/dev/null || cat /sys/fs/cgroup/lxc/392/cgroup.procs); do echo "$p $(cat /proc/$p/comm 2>/dev/null) rss=$(awk '/VmRSS/{print $2,$3}' /proc/$p/status 2>/dev/null)"; done | sort -k3 -rn -t= | head -20

Ran terminal command: pvecm status 2>&1 | tail -15; echo ===; ps --ppid $(head -1 /proc/$(pgrep -f "lxc-start.392" | head -1)/task//children 2>/dev/null) 2>/dev/null; find /sys/fs/cgroup/lxc/392 -name cgroup.procs -exec cat {} ; 2>/dev/null | sort -u | wc -l

Ran terminal command: sudo -n true 2>&1 && echo PASSWORDLESS_OK || echo NEEDS_PASSWORD

Checked terminal output

Ran terminal command: whoami; hostname; sudo -n whoami 2>&1

Passwordless sudo works. Continuing:

Ran terminal command: sudo cat /etc/pve/lxc/392.conf; echo ===; sudo grep -E '^(anon|file|slab_reclaimable|kernel|inactive_file|active_file) ' /sys/fs/cgroup/lxc/392/memory.stat

Ran terminal command: sudo find /sys/fs/cgroup/lxc/392 -name cgroup.procs -exec cat {} ; 2>/dev/null | sort -un | while read p; do rss=$(awk '/VmRSS/{print $2}' /proc/$p/status 2>/dev/null); [ -n "$rss" ] && echo "$rss $p $(tr '\0' ' ' < /proc/$p/cmdline 2>/dev/null | cut -c1-100)"; done | sort -rn | head -15

Culprit identified. CT 392 is a code-server dev workspace and it's simply oversubscribed:

RSS Process
1.78 GB code-server VS Code extension host
678 MB second code-server bootstrap process
418 MB vite --port 3000 dev server (support-crm)
357 MB kerebron-server (node)
187 MB code-server entry
93 + 74 MB two more duplicate vite instances

That's ~3.6 GB of unswappable anon memory in a 4 GB container. The kernel can't reclaim any of it (swap: 0), so it strangles the remaining ~80 MB of file cache — every process constantly refaults its own code pages from disk (1.8B refaults, cgroup hit memory.max 3,889 times). Those stalls are what's pumping the node load to ~53.

Recommendations, in order of impact:

  1. Kill the duplicate vite dev servers — pids 1527703/1527685 and 1378423/1378404 look like abandoned npm run dev sessions alongside the active one (1353407).
  2. Bump the CT to 6–8 GB — code-server + extension host + a Vite dev server realistically needs ~6 GB. It's an interactive dev box on a host with 430 GB available; 4 GB is just too tight.
  3. The 1.78 GB extension host suggests a heavy extension (Copilot/language servers) or a leak — a code-server "Reload Window" would reset it.

Also flagging: that container config embeds LDAP bind and Wazuh registration passwords in plaintext lxc.environment.runtime lines (readable by anyone with node shell access, and now in this terminal's scrollback). Worth rotating/moving to a secrets mechanism.

Want me to kill the stale vite processes and/or raise the memory limit (pct set 392 -memory 8192)?

> User: not yet. I am glad we have a root cause. I want to discuss ideas for detection and mitigation for future events. When a memory exhaustion occurs what should we do on the host computer to resuce the situation and prevent other containers from paying a price for the thrashing?

GitHub Copilot: Good question — this is a classic "cgroups contain memory, but not the side effects of memory pressure" problem. Here's how I'd think about it in three layers:

Why other containers pay a price

The thrash in CT 392 escapes containment via three channels:

  • Disk I/O: 1.8B refaults means the CT hammers Ceph/pool0 with reads — shared with everyone.
  • Kernel CPU time: direct reclaim runs in-kernel on the host's cores.
  • Load-driven scheduling noise and pressure on shared daemons (Ceph OSDs on the same node).

So mitigation = detect early, and convert "slow thrash" into either "reclaim gently" or "fail fast (OOM kill)".

1. Detection (know within seconds, not when the graph spikes)

PSI is the right signal — it's exactly what caught this today:

  • Alert on per-cgroup full avg10 > 10–20 sustained for 30–60s: /sys/fs/cgroup/lxc/<CTID>/memory.pressure
  • Alert on host-level memory full avg10 > ~5 (host of this size should be near zero).
  • Watch memory.events max counter deltas — CT 392 hit its limit 3,889 times silently, no OOM ever fired.

Practical options:

  • Prometheus: node_exporter with --collector.pressure + cadvisor-style cgroup metrics, or a tiny textfile-collector script that dumps per-CT PSI. Alertmanager rule on full avg10.
  • You already run Wazuh — a custom command check that scans /sys/fs/cgroup/lxc/*/memory.pressure each minute is a 10-line script.
  • Key insight: workingset_refault rate + PSI full is the definitive "thrashing, not just full" signature. memory.current ≈ memory.max alone is normal and not alertable.

2. Automatic mitigation (make it fail fast instead of thrash)

The core problem: with swap: 0 and mostly-anon memory, the kernel can't reclaim, but it also never OOMs because it can technically keep evicting the tiny file cache. That's the thrash trap. Fixes:

  1. memory.oom.group=1 on the CT cgroup — when OOM finally fires, kill the whole cgroup atomically instead of picking off one process and re-thrashing.

  2. Set memory.high below memory.max (e.g. high=3.5G, max=4G). This throttles + reclaims the offender early and gradually, before it reaches the death-spiral zone. Proxmox doesn't expose this, but a hookscript or systemd drop-in can set it at CT start:

    lxc.hook.start-host: /usr/local/sbin/set-memory-high.sh
    
  3. Run a userspace PSI-driven killer — this is the industry answer (Facebook's oomd, systemd's systemd-oomd, or earlyoom). Configure it per-cgroup: "if lxc/392 has memory full avg10 > 40 for 30s, kill the cgroup." It acts in seconds, whereas the kernel OOM killer can stall for many minutes while thrash destroys the node. On PVE, systemd-oomd with ManagedOOMMemoryPressure=kill on the LXC slice is the least-custom path.

  4. Cap the collateral I/O: set an io.max / blkio limit on CT cgroups (or at least the known dev-box class) so a refault storm can't saturate the Ceph pool. This directly protects the other containers.

3. Manual rescue playbook (when it's happening now)

In order of preference — least to most destructive:

# 1. Confirm the offender (10 seconds)
grep -H . /sys/fs/cgroup/lxc/*/memory.pressure | grep full | sort -t= -k2 -rn | head

# 2. Freeze it — instantly stops the thrash, buys time to think, fully reversible
echo 1 > /sys/fs/cgroup/lxc/392/cgroup.freeze
# ... investigate, then: echo 0 > .../cgroup.freeze

# 3. Give it headroom live (no restart needed) — often the real fix
pct set 392 -memory 8192   # applies to running CT

# 4. Force targeted reclaim (kernel ≥5.19) instead of waiting for kswapd
echo 512M > /sys/fs/cgroup/lxc/392/memory.reclaim

# 5. Kill just the hog inside the CT (from host)
pct exec 392 -- kill -9 <pid>

# 6. Last resort
pct stop 392

cgroup.freeze is the underused gem — it's the SIGSTOP of cgroups. The instant you freeze the CT, all its reclaim/refault activity ceases, node load drains, other containers recover, and you've lost nothing: you can inspect its processes while frozen and thaw after raising the limit.

Suggested minimal package for your fleet

  1. Wazuh/Prometheus alert on per-CT PSI full avg10 > 20 for 60s.
  2. Hookscript on dev-class CTs: set memory.high = 90% of max and memory.oom.group=1.
  3. A rescue-ct helper script on each node implementing steps 1–4 above.
  4. Policy: interactive code-server CTs get ≥6 GB.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions