Support fork in LiteBox userland - #1483
Merged
Merged
Conversation
Single-threaded Linux x86-64 guests can now call fork, or clone without CLONE_VM, when process duplication is allowed. The parent shim uploads its memory to the pending child's broker-held image with a new WriteChildMemory operation and starts the child with a fork startup carrying its registers, FS base, program break, memory regions, descriptors, signal state, credentials, working directory, and umask. The userland broker passes the sealed image memfd to the child runner and launches runners without address randomization, so the child runner restores each region at its parent's address before enabling seccomp and resumes the child from fork. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Release the broker's fork image once the child runner is set up, copy execute-only regions by making them readable only while snapshotting, and make the fork tests check the child's restored cwd and execute-only code. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Fork is always available in the shim; whether a process may be duplicated is decided by the broker, which denies it with EPERM unless process duplication is allowed. Add PageManagementProvider::PLACEMENT_ADDR_MAX, an upper bound for addresses the memory manager chooses itself, while fixed-address mappings may still reach TASK_ADDR_MAX. On x86-64 Linux userland it is 0x7000_0000_0000, hints are placed exactly with MAP_FIXED_NOREPLACE, and native remaps use MREMAP_FIXED into a destination claimed first, so the host kernel never moves guest memory into the host's mmap area. This lets a forked child restore every region at the parent's address without racing the fresh runner's own host mappings. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Move the loop that maps the child's regions and fills them from the parent's image into process.rs as Task::restore_fork_image, beside write_fork_image whose layout it reads, and have set_forked_init_state make the child return zero from fork. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Map the fork image into the child instead of reading it, read guest files in 512 KiB chunks, keep the tail of grown file mappings as fresh anonymous memory in remap_pages, and poll runner startup every 1 ms. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
The forking thread interrupts the other guest threads and waits until each parks at a safe point (before re-entering the guest or before blocking in a wait) while the parent is snapshotted; the child gets only the calling thread. fork now restarts with ERESTARTNOINTR if interrupted while pausing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
- Fall back cleanly when a native mremap move fails instead of panicking. - Remember whether an inaccessible area ever held data (VM_HAS_CONTENTS) so fork copies PROT_NONE pages with contents while leaving bare reservations uncopied. - Count pending child images against a broker-wide budget (4 GiB default). - Write child images straight from the shared transfer buffer and merge adjacent shared-buffer slots into one range. - Wait for runner exit and control connections with a pidfd and the listener on Linux instead of polling every millisecond. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Copying remaps pass only access permissions to the platform, waits recompute their timeout after the interrupt check, growing remaps treat areas differing only in whether they hold data as one, the fork image scan folds pages and trims zero pages, and child memory writes use up to 16 shared-buffer slots. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Growing a mapping now only checks that areas differing in whether they may hold data count as one, without merging them, so failed or in-place growth leaves them apart. Moving such a range uses the same check, so an overflowing size fails instead of panicking. The threaded fork test spins and forks by system call so its forkers enter fork at once. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Weidong Cui (wdcui)
force-pushed
the
wdcui/ulitebox/fork
branch
from
October 3, 2026 02:04
0500add to
ea3a623
Compare
…panic Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c83ae9b-67f4-4c14-9f8b-e4b3fba456fe
|
🤖 SemverChecks 🤖 Click for details |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds
forksupport to LiteBox userland for Linux x86-64 guests when process duplication is allowed:fork(orclonewithoutCLONE_VM) pauses the parent's other threads, streams its memory into a broker-held image through a new write-child-memory operation, and starts a fresh runner that maps the image at the parent's addresses and resumes only the calling thread. Parents with shared mappings or runner-local descriptors getEAGAIN. Supporting changes keep LiteBox-placed guest memory below a newPLACEMENT_ADDR_MAX, cap child images with a broker budget, raise shared-buffer transfers to 1 MiB, and have the userland broker launch runners without ASLR and wait on them with a pidfd.