> User: I am interested in why high memory in a single container with only 4gb or ram would cause this load spike in proxmox
GitHub Copilot: Short answer: because Linux "load" isn't CPU usage — it counts both runnable tasks and tasks stuck in uninterruptible sleep (D-state), and a container thrashing against its memory cgroup limit generates lots of both.
What happens when an LXC container hits its 4GB cap:
-
Cgroup memory pressure → direct reclaim. LXC memory limits are enforced by the host kernel via cgroups. When the container approaches 4GB, every allocation inside it triggers direct reclaim — the allocating process itself runs kernel reclaim code on the host CPU. That's real host CPU/kernel work, not contained "inside" the CT.
-
Swap thrashing / page cache refaults. The kernel starts swapping the container's pages out and dropping its page cache. Processes then immediately fault those pages back in. Each fault means synchronous disk I/O, and the process waits in D-state. Load average counts every one of those blocked tasks, so a container with 20 thrashing threads can add ~20 to the node's load while CPU utilization looks modest. (This matches your graph: tall load spikes in a plateau/spike pattern typical of thrash cycles, not sustained CPU work.)
-
Kernel threads run at host level. kswapd, writeback threads, and (if you use ZFS) ARC reclaim/arc_prune all do the shoveling on behalf of the container — that work shows up as node load, not container CPU.
-
OOM-kill cycles. If the workload keeps growing, the cgroup OOM killer kills it, the service restarts, re-allocates, and the cycle repeats — producing repeated spikes like the separated humps in your chart.
How to confirm on the node:
# Per-cgroup memory pressure (PSI) — the smoking gun
cat /sys/fs/cgroup/lxc/<CTID>/memory.pressure
cat /proc/pressure/memory /proc/pressure/io
# Swap-in/out activity during a spike
vmstat 1 # watch si/so columns
# D-state processes during a spike
ps -eo state,pid,comm | awk '$1=="D"'
# OOM events for that container
journalctl -k | grep -i oom
Fixes: raise the CT's memory limit, reduce or disable the CT's swap (swap: 0 in the CT config stops thrash-via-swap and fails fast to OOM instead), or fix the app's memory footprint. If the node uses ZFS, also check that ARC isn't competing for the same RAM (arc_summary, cap via zfs_arc_max).
If you share the Proxmox browser tab or give me SSH access to the node from this terminal, I can help dig into the actual PSI/OOM evidence.
> User: lets get into opensource-phxdc-pve1 and look at what is going on.
GitHub Copilot: Ran terminal command: ssh -o ConnectTimeout=10 -o BatchMode=yes root@opensource-phxdc-pve1 "hostname && uptime && free -h"
Ran terminal command: ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new root@opensource-phxdc-pve1 "hostname && uptime && free -h"
The SSH session is waiting for the root password. Please type it directly into the terminal (I can't handle secrets) — then I'll continue investigating.
> User: I just sudo su -
GitHub Copilot: Run in Terminal
> User: now we can jump betwen hosts
GitHub Copilot: Ran terminal command: whoami && ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new root@opensource-phxdc-pve1 "hostname && uptime && free -h && echo --- && cat /proc/pressure/memory"
Ran terminal command: ps -eo state,pid,ppid,wchan:30,comm | awk '$1=="D"' | head -40; echo ===; ps -eo state | grep -c D
Ran terminal command: for d in /sys/fs/cgroup/lxc/*/; do ctid=$(basename $d); p=$(grep '^full' $d/memory.pressure 2>/dev/null | awk '{print $2}' | cut -d= -f2); [ -n "$p" ] && echo "$ctid full_avg10=$p"; done | sort -t= -k2 -rn | head; echo ===; top -bn1 -o %CPU | head -25
Found it: container 392 is at 83% memory-stall (full_avg10=83.63) — it's the one thrashing. Let's see what's inside it.
Ran terminal command: pct config 392 | head -20; echo ===; cat /sys/fs/cgroup/lxc/392/memory.current /sys/fs/cgroup/lxc/392/memory.max /sys/fs/cgroup/lxc/392/memory.swap.current 2>/dev/null; echo ===; grep -E 'oom|refault' /sys/fs/cgroup/lxc/392/memory.events /sys/fs/cgroup/lxc/392/memory.stat 2>/dev/null | grep -Ev ' 0$' | head
Ran terminal command: export PATH=$PATH:/usr/sbin; pct config 392; echo ===; grep -cE '^oom' /sys/fs/cgroup/lxc/392/memory.events; cat /sys/fs/cgroup/lxc/392/memory.events
Ran terminal command: cat /etc/pve/lxc/392.conf 2>&1 | head; echo ===; systemctl is-active pve-cluster; echo ===; for p in $(cat /sys/fs/cgroup/lxc/392/ns/cgroup.procs 2>/dev/null || cat /sys/fs/cgroup/lxc/392/cgroup.procs); do echo "$p $(cat /proc/$p/comm 2>/dev/null) rss=$(awk '/VmRSS/{print $2,$3}' /proc/$p/status 2>/dev/null)"; done | sort -k3 -rn -t= | head -20
Ran terminal command: pvecm status 2>&1 | tail -15; echo ===; ps --ppid $(head -1 /proc/$(pgrep -f "lxc-start.392" | head -1)/task//children 2>/dev/null) 2>/dev/null; find /sys/fs/cgroup/lxc/392 -name cgroup.procs -exec cat {} ; 2>/dev/null | sort -u | wc -l
Ran terminal command: sudo -n true 2>&1 && echo PASSWORDLESS_OK || echo NEEDS_PASSWORD
Checked terminal output
Ran terminal command: whoami; hostname; sudo -n whoami 2>&1
Passwordless sudo works. Continuing:
Ran terminal command: sudo cat /etc/pve/lxc/392.conf; echo ===; sudo grep -E '^(anon|file|slab_reclaimable|kernel|inactive_file|active_file) ' /sys/fs/cgroup/lxc/392/memory.stat
Ran terminal command: sudo find /sys/fs/cgroup/lxc/392 -name cgroup.procs -exec cat {} ; 2>/dev/null | sort -un | while read p; do rss=$(awk '/VmRSS/{print $2}' /proc/$p/status 2>/dev/null); [ -n "$rss" ] && echo "$rss $p $(tr '\0' ' ' < /proc/$p/cmdline 2>/dev/null | cut -c1-100)"; done | sort -rn | head -15
Culprit identified. CT 392 is a code-server dev workspace and it's simply oversubscribed:
| RSS |
Process |
| 1.78 GB |
code-server VS Code extension host |
| 678 MB |
second code-server bootstrap process |
| 418 MB |
vite --port 3000 dev server (support-crm) |
| 357 MB |
kerebron-server (node) |
| 187 MB |
code-server entry |
| 93 + 74 MB |
two more duplicate vite instances |
That's ~3.6 GB of unswappable anon memory in a 4 GB container. The kernel can't reclaim any of it (swap: 0), so it strangles the remaining ~80 MB of file cache — every process constantly refaults its own code pages from disk (1.8B refaults, cgroup hit memory.max 3,889 times). Those stalls are what's pumping the node load to ~53.
Recommendations, in order of impact:
- Kill the duplicate vite dev servers — pids 1527703/1527685 and 1378423/1378404 look like abandoned
npm run dev sessions alongside the active one (1353407).
- Bump the CT to 6–8 GB — code-server + extension host + a Vite dev server realistically needs ~6 GB. It's an interactive dev box on a host with 430 GB available; 4 GB is just too tight.
- The 1.78 GB extension host suggests a heavy extension (Copilot/language servers) or a leak — a code-server "Reload Window" would reset it.
Also flagging: that container config embeds LDAP bind and Wazuh registration passwords in plaintext lxc.environment.runtime lines (readable by anyone with node shell access, and now in this terminal's scrollback). Worth rotating/moving to a secrets mechanism.
Want me to kill the stale vite processes and/or raise the memory limit (pct set 392 -memory 8192)?
> User: not yet. I am glad we have a root cause. I want to discuss ideas for detection and mitigation for future events. When a memory exhaustion occurs what should we do on the host computer to resuce the situation and prevent other containers from paying a price for the thrashing?
GitHub Copilot: Good question — this is a classic "cgroups contain memory, but not the side effects of memory pressure" problem. Here's how I'd think about it in three layers:
Why other containers pay a price
The thrash in CT 392 escapes containment via three channels:
- Disk I/O: 1.8B refaults means the CT hammers Ceph/pool0 with reads — shared with everyone.
- Kernel CPU time: direct reclaim runs in-kernel on the host's cores.
- Load-driven scheduling noise and pressure on shared daemons (Ceph OSDs on the same node).
So mitigation = detect early, and convert "slow thrash" into either "reclaim gently" or "fail fast (OOM kill)".
1. Detection (know within seconds, not when the graph spikes)
PSI is the right signal — it's exactly what caught this today:
- Alert on per-cgroup
full avg10 > 10–20 sustained for 30–60s: /sys/fs/cgroup/lxc/<CTID>/memory.pressure
- Alert on host-level memory full avg10 > ~5 (host of this size should be near zero).
- Watch
memory.events max counter deltas — CT 392 hit its limit 3,889 times silently, no OOM ever fired.
Practical options:
- Prometheus:
node_exporter with --collector.pressure + cadvisor-style cgroup metrics, or a tiny textfile-collector script that dumps per-CT PSI. Alertmanager rule on full avg10.
- You already run Wazuh — a custom command check that scans
/sys/fs/cgroup/lxc/*/memory.pressure each minute is a 10-line script.
- Key insight:
workingset_refault rate + PSI full is the definitive "thrashing, not just full" signature. memory.current ≈ memory.max alone is normal and not alertable.
2. Automatic mitigation (make it fail fast instead of thrash)
The core problem: with swap: 0 and mostly-anon memory, the kernel can't reclaim, but it also never OOMs because it can technically keep evicting the tiny file cache. That's the thrash trap. Fixes:
-
memory.oom.group=1 on the CT cgroup — when OOM finally fires, kill the whole cgroup atomically instead of picking off one process and re-thrashing.
-
Set memory.high below memory.max (e.g. high=3.5G, max=4G). This throttles + reclaims the offender early and gradually, before it reaches the death-spiral zone. Proxmox doesn't expose this, but a hookscript or systemd drop-in can set it at CT start:
lxc.hook.start-host: /usr/local/sbin/set-memory-high.sh
-
Run a userspace PSI-driven killer — this is the industry answer (Facebook's oomd, systemd's systemd-oomd, or earlyoom). Configure it per-cgroup: "if lxc/392 has memory full avg10 > 40 for 30s, kill the cgroup." It acts in seconds, whereas the kernel OOM killer can stall for many minutes while thrash destroys the node. On PVE, systemd-oomd with ManagedOOMMemoryPressure=kill on the LXC slice is the least-custom path.
-
Cap the collateral I/O: set an io.max / blkio limit on CT cgroups (or at least the known dev-box class) so a refault storm can't saturate the Ceph pool. This directly protects the other containers.
3. Manual rescue playbook (when it's happening now)
In order of preference — least to most destructive:
# 1. Confirm the offender (10 seconds)
grep -H . /sys/fs/cgroup/lxc/*/memory.pressure | grep full | sort -t= -k2 -rn | head
# 2. Freeze it — instantly stops the thrash, buys time to think, fully reversible
echo 1 > /sys/fs/cgroup/lxc/392/cgroup.freeze
# ... investigate, then: echo 0 > .../cgroup.freeze
# 3. Give it headroom live (no restart needed) — often the real fix
pct set 392 -memory 8192 # applies to running CT
# 4. Force targeted reclaim (kernel ≥5.19) instead of waiting for kswapd
echo 512M > /sys/fs/cgroup/lxc/392/memory.reclaim
# 5. Kill just the hog inside the CT (from host)
pct exec 392 -- kill -9 <pid>
# 6. Last resort
pct stop 392
cgroup.freeze is the underused gem — it's the SIGSTOP of cgroups. The instant you freeze the CT, all its reclaim/refault activity ceases, node load drains, other containers recover, and you've lost nothing: you can inspect its processes while frozen and thaw after raising the limit.
Suggested minimal package for your fleet
- Wazuh/Prometheus alert on per-CT PSI
full avg10 > 20 for 60s.
- Hookscript on dev-class CTs: set
memory.high = 90% of max and memory.oom.group=1.
- A
rescue-ct helper script on each node implementing steps 1–4 above.
- Policy: interactive code-server CTs get ≥6 GB.
> User: I am interested in why high memory in a single container with only 4gb or ram would cause this load spike in proxmox
GitHub Copilot: Short answer: because Linux "load" isn't CPU usage — it counts both runnable tasks and tasks stuck in uninterruptible sleep (D-state), and a container thrashing against its memory cgroup limit generates lots of both.
What happens when an LXC container hits its 4GB cap:
Cgroup memory pressure → direct reclaim. LXC memory limits are enforced by the host kernel via cgroups. When the container approaches 4GB, every allocation inside it triggers direct reclaim — the allocating process itself runs kernel reclaim code on the host CPU. That's real host CPU/kernel work, not contained "inside" the CT.
Swap thrashing / page cache refaults. The kernel starts swapping the container's pages out and dropping its page cache. Processes then immediately fault those pages back in. Each fault means synchronous disk I/O, and the process waits in D-state. Load average counts every one of those blocked tasks, so a container with 20 thrashing threads can add ~20 to the node's load while CPU utilization looks modest. (This matches your graph: tall load spikes in a plateau/spike pattern typical of thrash cycles, not sustained CPU work.)
Kernel threads run at host level.
kswapd, writeback threads, and (if you use ZFS) ARC reclaim/arc_pruneall do the shoveling on behalf of the container — that work shows up as node load, not container CPU.OOM-kill cycles. If the workload keeps growing, the cgroup OOM killer kills it, the service restarts, re-allocates, and the cycle repeats — producing repeated spikes like the separated humps in your chart.
How to confirm on the node:
Fixes: raise the CT's memory limit, reduce or disable the CT's swap (
swap: 0in the CT config stops thrash-via-swap and fails fast to OOM instead), or fix the app's memory footprint. If the node uses ZFS, also check that ARC isn't competing for the same RAM (arc_summary, cap viazfs_arc_max).If you share the Proxmox browser tab or give me SSH access to the node from this terminal, I can help dig into the actual PSI/OOM evidence.
> User: lets get into opensource-phxdc-pve1 and look at what is going on.
GitHub Copilot: Ran terminal command: ssh -o ConnectTimeout=10 -o BatchMode=yes root@opensource-phxdc-pve1 "hostname && uptime && free -h"
Ran terminal command: ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new root@opensource-phxdc-pve1 "hostname && uptime && free -h"
The SSH session is waiting for the root password. Please type it directly into the terminal (I can't handle secrets) — then I'll continue investigating.
> User: I just sudo su -
GitHub Copilot: Run in Terminal
> User: now we can jump betwen hosts
GitHub Copilot: Ran terminal command: whoami && ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new root@opensource-phxdc-pve1 "hostname && uptime && free -h && echo --- && cat /proc/pressure/memory"
Ran terminal command: ps -eo state,pid,ppid,wchan:30,comm | awk '$1=="D"' | head -40; echo ===; ps -eo state | grep -c D
Ran terminal command: for d in /sys/fs/cgroup/lxc/*/; do ctid=$(basename$d); p=$ (grep '^full' $d/memory.pressure 2>/dev/null | awk '{print $2}' | cut -d= -f2); [ -n "$p" ] && echo "$ctid full_avg10=$p"; done | sort -t= -k2 -rn | head; echo ===; top -bn1 -o %CPU | head -25
Found it: container 392 is at 83% memory-stall (full_avg10=83.63) — it's the one thrashing. Let's see what's inside it.
Ran terminal command: pct config 392 | head -20; echo ===; cat /sys/fs/cgroup/lxc/392/memory.current /sys/fs/cgroup/lxc/392/memory.max /sys/fs/cgroup/lxc/392/memory.swap.current 2>/dev/null; echo ===; grep -E 'oom|refault' /sys/fs/cgroup/lxc/392/memory.events /sys/fs/cgroup/lxc/392/memory.stat 2>/dev/null | grep -Ev ' 0$' | head
Ran terminal command: export PATH=$PATH:/usr/sbin; pct config 392; echo ===; grep -cE '^oom' /sys/fs/cgroup/lxc/392/memory.events; cat /sys/fs/cgroup/lxc/392/memory.events
Ran terminal command: cat /etc/pve/lxc/392.conf 2>&1 | head; echo ===; systemctl is-active pve-cluster; echo ===; for p in $(cat /sys/fs/cgroup/lxc/392/ns/cgroup.procs 2>/dev/null || cat /sys/fs/cgroup/lxc/392/cgroup.procs); do echo "$p $(cat /proc/$p/comm 2>/dev/null) rss=$(awk '/VmRSS/{print $2,$3}' /proc/$p/status 2>/dev/null)"; done | sort -k3 -rn -t= | head -20
Ran terminal command: pvecm status 2>&1 | tail -15; echo ===; ps --ppid$(head -1 /proc/$ (pgrep -f "lxc-start.392" | head -1)/task//children 2>/dev/null) 2>/dev/null; find /sys/fs/cgroup/lxc/392 -name cgroup.procs -exec cat {} ; 2>/dev/null | sort -u | wc -l
Ran terminal command: sudo -n true 2>&1 && echo PASSWORDLESS_OK || echo NEEDS_PASSWORD
Checked terminal output
Ran terminal command: whoami; hostname; sudo -n whoami 2>&1
Passwordless sudo works. Continuing:
Ran terminal command: sudo cat /etc/pve/lxc/392.conf; echo ===; sudo grep -E '^(anon|file|slab_reclaimable|kernel|inactive_file|active_file) ' /sys/fs/cgroup/lxc/392/memory.stat
Ran terminal command: sudo find /sys/fs/cgroup/lxc/392 -name cgroup.procs -exec cat {} ; 2>/dev/null | sort -un | while read p; do rss=$(awk '/VmRSS/{print $2}' /proc/$p/status 2>/dev/null); [ -n "$rss" ] && echo "$rss$p $ (tr '\0' ' ' < /proc/$p/cmdline 2>/dev/null | cut -c1-100)"; done | sort -rn | head -15
Culprit identified. CT 392 is a code-server dev workspace and it's simply oversubscribed:
vite --port 3000dev server (support-crm)That's ~3.6 GB of unswappable anon memory in a 4 GB container. The kernel can't reclaim any of it (
swap: 0), so it strangles the remaining ~80 MB of file cache — every process constantly refaults its own code pages from disk (1.8B refaults, cgroup hitmemory.max3,889 times). Those stalls are what's pumping the node load to ~53.Recommendations, in order of impact:
npm run devsessions alongside the active one (1353407).Also flagging: that container config embeds LDAP bind and Wazuh registration passwords in plaintext
lxc.environment.runtimelines (readable by anyone with node shell access, and now in this terminal's scrollback). Worth rotating/moving to a secrets mechanism.Want me to kill the stale vite processes and/or raise the memory limit (
pct set 392 -memory 8192)?> User: not yet. I am glad we have a root cause. I want to discuss ideas for detection and mitigation for future events. When a memory exhaustion occurs what should we do on the host computer to resuce the situation and prevent other containers from paying a price for the thrashing?
GitHub Copilot: Good question — this is a classic "cgroups contain memory, but not the side effects of memory pressure" problem. Here's how I'd think about it in three layers:
Why other containers pay a price
The thrash in CT 392 escapes containment via three channels:
So mitigation = detect early, and convert "slow thrash" into either "reclaim gently" or "fail fast (OOM kill)".
1. Detection (know within seconds, not when the graph spikes)
PSI is the right signal — it's exactly what caught this today:
full avg10 > 10–20sustained for 30–60s:/sys/fs/cgroup/lxc/<CTID>/memory.pressurememory.eventsmaxcounter deltas — CT 392 hit its limit 3,889 times silently, no OOM ever fired.Practical options:
node_exporterwith--collector.pressure+cadvisor-style cgroup metrics, or a tiny textfile-collector script that dumps per-CT PSI. Alertmanager rule onfull avg10./sys/fs/cgroup/lxc/*/memory.pressureeach minute is a 10-line script.workingset_refaultrate + PSI full is the definitive "thrashing, not just full" signature.memory.current ≈ memory.maxalone is normal and not alertable.2. Automatic mitigation (make it fail fast instead of thrash)
The core problem: with
swap: 0and mostly-anon memory, the kernel can't reclaim, but it also never OOMs because it can technically keep evicting the tiny file cache. That's the thrash trap. Fixes:memory.oom.group=1on the CT cgroup — when OOM finally fires, kill the whole cgroup atomically instead of picking off one process and re-thrashing.Set
memory.highbelowmemory.max(e.g. high=3.5G, max=4G). This throttles + reclaims the offender early and gradually, before it reaches the death-spiral zone. Proxmox doesn't expose this, but a hookscript or systemd drop-in can set it at CT start:Run a userspace PSI-driven killer — this is the industry answer (Facebook's
oomd, systemd'ssystemd-oomd, orearlyoom). Configure it per-cgroup: "iflxc/392has memoryfull avg10 > 40for 30s, kill the cgroup." It acts in seconds, whereas the kernel OOM killer can stall for many minutes while thrash destroys the node. On PVE,systemd-oomdwithManagedOOMMemoryPressure=killon the LXC slice is the least-custom path.Cap the collateral I/O: set an
io.max/ blkio limit on CT cgroups (or at least the known dev-box class) so a refault storm can't saturate the Ceph pool. This directly protects the other containers.3. Manual rescue playbook (when it's happening now)
In order of preference — least to most destructive:
cgroup.freezeis the underused gem — it's the SIGSTOP of cgroups. The instant you freeze the CT, all its reclaim/refault activity ceases, node load drains, other containers recover, and you've lost nothing: you can inspect its processes while frozen and thaw after raising the limit.Suggested minimal package for your fleet
full avg10 > 20for 60s.memory.high= 90% of max andmemory.oom.group=1.rescue-cthelper script on each node implementing steps 1–4 above.