Page MenuHomeFreeBSD

Jail init process and virtual reboot
Needs ReviewPublic

Authored by jamie on Wed, Sep 16, 11:17 PM.
Tags
None
Referenced Files
Unknown Object (File)
Fri, Oct 9, 10:57 PM
Unknown Object (File)
Tue, Oct 6, 12:03 PM
Unknown Object (File)
Mon, Oct 5, 8:14 PM
Unknown Object (File)
Sun, Oct 4, 7:33 PM
Unknown Object (File)
Sun, Oct 4, 5:13 PM
Unknown Object (File)
Fri, Oct 2, 12:23 AM
Unknown Object (File)
Thu, Oct 1, 8:11 PM
Unknown Object (File)
Thu, Oct 1, 7:15 PM

Details

Reviewers
olce
Group Reviewers
Restricted Owners Package(Owns No Changed Paths)
Jails
Summary

Prisons can currently run /sbin/init, which acts a reaper for processes started under it. This extends and formalizes the virtual init concept.

The "init" jail parameter specifies a jail with virtual init. When the parameter is set, an init process is forked off, similar to how the real init (process 1) is created on bootup. The prison field pr_initproc points to this process, and is used in some places that used to check global initproc. prison0.pr_initproc points to the same process as global initproc.

A lot of system management depends on init being pid 1, and a lot of the changes here are so looking up pid 1 in a jail returns its init process instead of the global one, and also any reporting about that process (from within the jail) shows its pid as 1.

When reboot(2) is called from withinga jail with init, its runs a virtual halt: all processes in that jail are killed, an OSD PR_METHOD_REBOOT call is made (which can reset things to a newish state), and then either a new init process is forked off (virtual reboot) or the prison is marked as no longer having an init process. A virtually halted jail may go away once its processes are dead, though it may not, depending on such things as child jails, persistent status, or non-init-related processes.

At this time, there is no concept of a virtual single-user mode, or a virtual console, both of which are necessary for a proper init-based system. The virtual init is also not a universal parent process in the prison, as any process attached to prison keep their own parent processes and reapers.

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped

Event Timeline

Owners added a reviewer: Restricted Owners Package.Wed, Sep 16, 11:17 PM

This is interesting. I asked Claude Code to help a bit with the review...

The patch makes prison_priv_check() grant PRIV_REBOOT to any jail that has an init. That privilege is also the only check in sys_kexec_load(), apart from securelevel_gt(cred, 0). Root in such a jail, at the default securelevel of -1, can load a kernel image for the host, replace one that is already loaded, or unload it. The host boots that image on its next RB_KEXEC reboot.

ki_jid regression (fill_kinfo_proc_only, around kern_proc.c:1163). kp->ki_jid = pr->pr_id now takes the id of the viewer's jail instead of the target's (cred->cr_prison). From the host, every jailed process would report jid 0 (ps -o jid, jls-style tools).

tdfind() is given the translated pid in thr_kill, thr_wake and thr_set_name (kern_thr.c). tdfind() compares against the real p_pid, so for the jailed init these calls should/would fail with ESRCH. libc's raise() is thr_self + thr_kill, so raise() and abort() in the jailed init should break, I think.

prison_boot() never refreshes the child thread's credential after do_jail_attach(prit, …). create_init() does this explicitly (crcowfree and td_realucred = crcowget(p_ucred)). As written, start_init() runs kern_execve() with the host caller's unjailed td_ucred, so exec-time MAC and permission checks see the host credential.

The host's boothowto leaks into jails. start_init() still adds -s when boothowto & RB_SINGLE, with no pr == NULL guard. After a host boot -s, every init jail boots single-user.

The host's kenv reaches the jailed init. init.c honours init_exec, init_script, init_chroot, init_shell, init_rc and init_path from kenv(2). KENV_GET is open to jails because unprivileged_kenv_read defaults to 1, so the host's boot-time settings would apply inside every jail. The patch already computes jailed in init.c, so skipping kenv there is cheap.

Some replaced checks look at the caller's jail when the question is about the target's.

  • PT_TRACE_ME: p->p_pptr == initproc became == caller's pr_initproc. In a jail without init, that is NULL. A jailed orphan whose parent is the host's init can then PT_TRACE_ME, which main refuses today.
  • The same applies to PT_DETACH resetting p_sigparent.
  • The same applies to p_candebug()'s "can't trace init at securelevel > 0" check.
  • The predicate should probably be p == initproc || p == <target's prison>->pr_initproc.

reboot -p and Linux REBOOT_POWEROFF restart the jail instead of stopping it. prison_boot() only halts when RB_HALT is set, so RB_POWEROFF without RB_HALT becomes a reboot.

prison_boot() errors are ignored in kern_jail_set(). If fork1() fails, jail -c init … still reports success, and a non-persistent jail silently disappears.

Failure path in prison_boot(): when do_jail_attach() fails, the child gets SIGKILL but is still scheduled into start_init(). It then runs an exec of the init path on the host root with host credentials before the signal lands.

The kill loop is incomplete. It walks pr_proclist only and skips PRS_NEW. A process that is mid-fork survives a virtual reboot, and child jails are left alone. prison_proc_iterate() already exists for this job. The loop also doesn't wait for anything to exit before the PR_METHOD_REBOOT cleanup and the new init start.

jail.h leaves a dangling lock key. The patch deletes the (A) allproc_lock entry, but pr_proclist on main still uses (A).

init.shutdown_timeout = 0 can't be set. Zero is treated as "unset" and replaced from an ancestor whenever init is set, and the value is not validated (negative values pass).

int jailed in init.c is read uninitialised if sysctlbyname() fails.

Exec args leak on the jail error path in start_init() (exit1 without exec_free_args).

goto panic_or_exit jumps backwards into an if block.

A "jail mutex" entry was added to the witness order list, did I overlook a path in the diff that takes pr_mtx under a proc lock?

_pfind() reads pr_initproc->p_pid and then drops the lock, so if init exits and its pid is reused, "pid 1" can resolve to an unrelated process.

The pid-1 illusion is spread across about 100 call sites, and it is incomplete:

  • Init is now both pid 1 and its real pid inside the jail.
  • ps shows 1, but /proc/1 (procfs and linprocfs) doesn't resolve, because procfs lookups still use pfind().
  • getsid, getpgid results, F_GETOWN, tcgetpgrp, kevent data, ktrace (your XXX) and ki_reaper (your XXX) still show the real pid.
  • Every pfind() or pget() call added later will quietly break it again.
  • Almost every *_cred call passes curthread's credential. Doing the translation inside pfind() and pget() from curthread would remove most of the 30-file churn and cover future callers automatically.

In base, the userland that depends on init being pid 1 is 7 call sites in 3 programs: init.c (getpid() != 1, kill(1, …)), reboot.c and shutdown.c. A per-jail "init pid" readable from userland would fix those with no kernel aliasing. The counter-argument is third-party software that checks getppid() == 1. I'm not sure if this would fit your goals.

Init's starting state differs from the real init's:

  • The jailed init inherits the operator's audit session, login class, rlimits, nice value and MAC label from the jail(8) process.
  • The real init gets mac_cred_create_init() and audit_cred_proc1().

This is interesting. I asked Claude Code to help a bit with the review...

Ah, my first encounter with Claude. So how much of this was Claude Code, and how much was you?

The patch makes prison_priv_check() grant PRIV_REBOOT to any jail that has an init. That privilege is also the only check in sys_kexec_load(), apart from securelevel_gt(cred, 0). Root in such a jail, at the default securelevel of -1, can load a kernel image for the host, replace one that is already loaded, or unload it. The host boots that image on its next RB_KEXEC reboot.

So sys_kexec_load needs to check !jailed as well as PRIV_REBOOT. That's the only other place besides sys_reboot that uses that permission (except a strange bit in linux_reboot that does no more than return a result code).

ki_jid regression (fill_kinfo_proc_only, around kern_proc.c:1163). kp->ki_jid = pr->pr_id now takes the id of the viewer's jail instead of the target's (cred->cr_prison). From the host, every jailed process would report jid 0 (ps -o jid, jls-style tools).

Yep.

tdfind() is given the translated pid in thr_kill, thr_wake and thr_set_name (kern_thr.c). tdfind() compares against the real p_pid, so for the jailed init these calls should/would fail with ESRCH. libc's raise() is thr_self + thr_kill, so raise() and abort() in the jailed init should break, I think.

OK that and sys_thr_kill2 only get the pid to turn it pack into a process pointer later, not to report back to the user. So yeah, it makes sense to keep the real pid.

prison_boot() never refreshes the child thread's credential after do_jail_attach(prit, …). create_init() does this explicitly (crcowfree and td_realucred = crcowget(p_ucred)). As written, start_init() runs kern_execve() with the host caller's unjailed td_ucred, so exec-time MAC and permission checks see the host credential.

I'll test that, but it makes sense. All other uses of do_jail_attach return to the user, and don't need to worry about the thread cred until it's reset (by whatever keeps thread creds in sync).

The host's boothowto leaks into jails. start_init() still adds -s when boothowto & RB_SINGLE, with no pr == NULL guard. After a host boot -s, every init jail boots single-user.

Sounds like I need an init.boothowto jail parameter. Notably, I haven't yet done anything with how init gets a console equivalent, which might use something there.

The host's kenv reaches the jailed init. init.c honours init_exec, init_script, init_chroot, init_shell, init_rc and init_path from kenv(2). KENV_GET is open to jails because unprivileged_kenv_read defaults to 1, so the host's boot-time settings would apply inside every jail. The patch already computes jailed in init.c, so skipping kenv there is cheap.

init.c is hardly touched, so this isn't surprising. I already have a jailed init.path; perhaps I need the others as well. Even if I end up with jail parameters for these, it's al open question what the defaults should be, because like boothowto, what you want to do for the real boot likely doesn't change what you want to do for jails; but maybe it does?

Some replaced checks look at the caller's jail when the question is about the target's.

  • PT_TRACE_ME: p->p_pptr == initproc became == caller's pr_initproc. In a jail without init, that is NULL. A jailed orphan whose parent is the host's init can then PT_TRACE_ME, which main refuses today.
  • The same applies to PT_DETACH resetting p_sigparent.
  • The same applies to p_candebug()'s "can't trace init at securelevel > 0" check.
  • The predicate should probably be p == initproc || p == <target's prison>->pr_initproc.

I want to be able to debug an inferior jail's init process, which to my view is just a regular process. p_cansee should stop be from debugging a superior jail's init, so the only thing left to check is the caller's (debugging process's) init process. The PT_DETACH case makes sense though, since that really doesn't have anything to do with the caller at that point.

PT_TRACE_ME is not so much about the target's inittproc (especially since the target is defined as the caller), but about the parent's initproc, i.e. is the parent its own initproc.

reboot -p and Linux REBOOT_POWEROFF restart the jail instead of stopping it. prison_boot() only halts when RB_HALT is set, so RB_POWEROFF without RB_HALT becomes a reboot.

Yeah, I made an assumption based on the code I looked at, that any of the halt-related options always had at least DB_HALT set. I didn't look at Linux code, which apparently does things differently. So I could definitely add the RB_POWEROFF bit to the test. I should also look at RB_SINGLE, in conjunction with the virtualized boothowto mentioned above. I might also end up disallowing a bunch of bits instead of ignoring them.

prison_boot() errors are ignored in kern_jail_set(). If fork1() fails, jail -c init … still reports success, and a non-persistent jail silently disappears.

Yes, I should return the error from prison_boot. The jail will still disappear of course, but at least with a proper error. Of course, other errors can easily occur in the virtual init process itself, starting with the exec. Since that's all in a detached process, the jail creator won't see it. Jail descriptors are likely useful in such a case.

Failure path in prison_boot(): when do_jail_attach() fails, the child gets SIGKILL but is still scheduled into start_init(). It then runs an exec of the init path on the host root with host credentials before the signal lands.

Ah, there's the answer to my "XXX Test if I can do this". Perhaps start_init needs some code to check for a pending signal, and not cleanly die if it exists.

The kill loop is incomplete. It walks pr_proclist only and skips PRS_NEW. A process that is mid-fork survives a virtual reboot, and child jails are left alone. prison_proc_iterate() already exists for this job. The loop also doesn't wait for anything to exit before the PR_METHOD_REBOOT cleanup and the new init start.

The child jails are an open question, and have been discussed a little, Ideally (in my mind), it would also kill child jails that were created by the rebooting jail, but not those that were created by a parent jail. There's no current record of a jail's creator.

prison_proc_iterate notably also skips PRS_NEW processes, and they have been a problem already e.g. in jail_remove. There needs to be a way to kill a new process, so it stops running as soon as it's able to cleanly do so.

jail.h leaves a dangling lock key. The patch deletes the (A) allproc_lock entry, but pr_proclist on main still uses (A).

Uh, oops.

init.shutdown_timeout = 0 can't be set. Zero is treated as "unset" and replaced from an ancestor whenever init is set, and the value is not validated (negative values pass).

Yeah, that test should be replaced by "!gotinitto". I chose not to validate the number, because it's also not validated in the original sysctl.

int jailed in init.c is read uninitialised if sysctlbyname() fails.

I suppose that could theoretically fail. I'm not sure how, but I could check it and set jailed to ... something.

Exec args leak on the jail error path in start_init() (exit1 without exec_free_args).

Real init never has to worry about that, so virtual init didn't either. But of course it's not free to ignore that.

goto panic_or_exit jumps backwards into an if block.

Is that a problem aside from bad form? I could certainly move it elsewhere.

A "jail mutex" entry was added to the witness order list, did I overlook a path in the diff that takes pr_mtx under a proc lock?

I don't think so. I recall finding something in testing that suggested this was worth including, but either not part of the patch or perhaps something that didn't make the final cut. So I think that ordering is proper but don't recall why, and if I add it, I should do so independent of all this.

But in looking for that, I noticed that I set pr_initproc protected by the prison lock in one place, and then by the process lock in another.

_pfind() reads pr_initproc->p_pid and then drops the lock, so if init exits and its pid is reused, "pid 1" can resolve to an unrelated process.

Generally not a problem, because init is not expected to exit. I would consider that equivalent to any non-1 pid having a race between the time the system call enters and the lookup is done.

The pid-1 illusion is spread across about 100 call sites, and it is incomplete:

  • Init is now both pid 1 and its real pid inside the jail.

True. Even though you don't see its real pid, if you happen to know it you can use it. This seems harmless.

  • ps shows 1, but /proc/1 (procfs and linprocfs) doesn't resolve, because procfs lookups still use pfind().

I should fix procfs. Probably linprocfs as well, since one of the goals is for even a linux emulation userspace to work.

  • getsid, getpgid results, F_GETOWN, tcgetpgrp, kevent data, ktrace (your XXX) and ki_reaper (your XXX) still show the real pid.

I haven't changed things that report the pgid. Since it's defined as a process id, it would make sense to do so. That would obviate where I currently skip init's check that it could get sid 1.

kevent and friends are definitely a big XXX. There's this disconnect between where the pid is recorded and where I have a cred I can work with. And I'm not sure there's even a way for it to appear as one pid to some readers and as another to others.

  • Every pfind() or pget() call added later will quietly break it again.
  • Almost every *_cred call passes curthread's credential. Doing the translation inside pfind() and pget() from curthread would remove most of the 30-file churn and cover future callers automatically.

Yes, I could do things the other way around: add a call for "get only the actual pid" and have the default do the translation. The cost is using curthread instead of an already-fetch thread value, but that's probably not a big deal. It surprised me how many places want to to this conversion. When I have first done this back in the FreeBSD 6 days, there's weren't quite so many.

In base, the userland that depends on init being pid 1 is 7 call sites in 3 programs: init.c (getpid() != 1, kill(1, …)), reboot.c and shutdown.c. A per-jail "init pid" readable from userland would fix those with no kernel aliasing. The counter-argument is third-party software that checks getppid() == 1. I'm not sure if this would fit your goals.

the main reason isn't the userland programs, which I could find some way to change (perhaps an initpid sysctl). It's the mindshare that doing stuff to pid 1 will affect init. Old sysadmins like me might be tempted to "kill -1 1" or something like that (I used to do that back when I ran a system with actual TTYs). And also back to the case of userland we don't have control over, notably linux emulation.

Init's starting state differs from the real init's:

  • The jailed init inherits the operator's audit session, login class, rlimits, nice value and MAC label from the jail(8) process.
  • The real init gets mac_cred_create_init() and audit_cred_proc1().

I originally wanted to fork the jailed init through the same kind of immaculate conception that real init gets, but I soon found out that fork1 (via something at the machine-dependent level) had to operate on curproc. I still may explore a way of doing this, perhaps leaving some task for a kernel thread.

I think this is a nice feature !

Some thinking about this:

  1. Actually prison0 is almost identical to other unprivileged prisons. After we have virtual init, then is it possible to fast reboot prison0 ( let's forget kill 1 right now ) ?
  2. A vnet prison has some network related state in kernel. Is it possible to extend this virtual reboot feature to clean up the vnet network stack so we can have a neat clean network state ? For example flush all routing table, deleting all IP addresses on interfaces, and / or detaching all cloned interfaces and recreate on virtual reboot .

This is interesting. I asked Claude Code to help a bit with the review...

Ah, my first encounter with Claude. So how much of this was Claude Code, and how much was you?

Hard to say. I have a "knowledgebase" of 1.4MB which Claude is using. This KB is having instructions for various cases and situations, and knowledge areas and instructions how to behave and what is important to me. This KB is the result of about 3 months of various projects with Claude and training it to "behave" and "do what I mean". The "goto backwards into an conditional" is a case which is based on my preferences (it works, you can do that, there are cases where it may make sense, but ... "goto is harmful, specially the backwards case", so my pragmatic stand is to at least flag it and make up my mind myself).

I gave Claude (paid version, with the current best model IMO) your patch and asked it to review it. This triggered surely into a lot of the KB entries. I have then read all the feedback and picked the one which made sense to me. I did not look at every code line, also not at every code line Claude pointed out. For some parts it surely surpasses my low-level knowledge, but the high level aspects made sense at least to mention. What you got was the output filtered by me. And some rewording where it made sense to me.

Based upon the KB, I would say maybe 80:20 or 70:30 from Claude:Me. From the actual work spend on it, it is more 95:05.

I have also an extensive testing KB part. When you consider this patchset test-ready, feel free to ping me with a high level description of what you want to have tested. I would throw that at my automated test-rig (I found the recent hwpmc issue which resulted in a security advisory on this test-rig).

I have an old (non-stack) stack open (I believe from 2018) of which you would have been a reviewer:

D15556 Initial vps (virtual process space) framework for jails.
D15570 Virtualization of basic variables and locks for jail+vps.
D15865 Provide process space virtualisation functionality for jails.
D15906 Implement "global" process (and zombie) lists loops iterating over all vps instances and all processes in each.

Just in case it helps taking this one level further; a rebase will certainly not be easy.

Some thinking about this:

  1. Actually prison0 is almost identical to other unprivileged prisons. After we have virtual init, then is it possible to fast reboot prison0 ( let's forget kill 1 right now ) ?

It could be added with some new flag to specify it. But I don't know about the utility: usually when you're rebooting the whole system, you want to reset hardware state, or change a kernel, or something else more than this virtual reboot provides.

  1. A vnet prison has some network related state in kernel. Is it possible to extend this virtual reboot feature to clean up the vnet network stack so we can have a neat clean network state ? For example flush all routing table, deleting all IP addresses on interfaces, and / or detaching all cloned interfaces and recreate on virtual reboot .

I'd definitely like to add such things, Anything that keeps per-jail state would be better off reset. None of that work has been done though.

In D59748#1380508, @bz wrote:

I have an old (non-stack) stack open (I believe from 2018) of which you would have been a reviewer:
D15556 Initial vps (virtual process space) framework for jails.
D15570 Virtualization of basic variables and locks for jail+vps.
D15865 Provide process space virtualisation functionality for jails.
D15906 Implement "global" process (and zombie) lists loops iterating over all vps instances and all processes in each.

Just in case it helps taking this one level further; a rebase will certainly not be easy.

Yes, I should be taking a new look at those, just for questions on the pid mapping interface. I don't have nearly the same level of work of course, but it would make sense to do things in an API way that makes sense for both.