[Devel] [PATCH vz10 1/2] fs/kernfs, ve: hide entries from a VE without invalidating the dentry
Pavel Tikhomirov
ptikhomirov at virtuozzo.com
Tue Aug 25 17:59:03 MSK 2026
On 8/25/26 15:29, Mirian Shilakadze wrote:
> kernfs_dop_revalidate() ends with a per VE visibility check and answers
> it with the same "return 0" that the staleness checks above it use.
> Those checks are properties of the kernfs node and hold for every
> observer: the node was deactivated, moved, renamed, or retagged.
> Visibility is a property of the calling task's VE, so one host dentry
> answers "valid" to a ve0 task and "stale" to a task inside a Container.
>
> The VFS reads 0 as a global fact and calls d_invalidate(), which walks
> the subtree and hands every mountpoint it finds to __detach_mounts().
> The mountpoint hash is not scoped to a mount namespace, and m_list holds
> every mount attached at that dentry in any of them, so a Container's
> lookup unmounts the host's mounts.
>
> One lookup of /sys/fs/bpf from a task that only did setns() into a
> Container's ve namespace, staying in the host mount namespace, both
> hides the entry from the caller and destroys the host's bpffs. A
> Container start reaches the same path on its own: libvzctl stats every
> mount point in the namespace to collect the mount flags of a bindmount
> source, and does it after CLONE_NEWVE and before pivot_root, so the host
> loses bpffs and tracefs on the way. libvzctl needs bpffs for the cgroup
> v2 device controller, so no Container on the node can be managed
> afterwards, and the damage outlives the failed start.
>
> Return -ENOENT instead. The caller that cannot see the entry is told
> the name is missing, which is what the check is for, and the dentry
> stays valid for everyone else. No caller of ->d_revalidate() reaches
> d_invalidate() with a negative return: lookup_dcache(), lookup_fast(),
> __lookup_slow() and lookup_open() in fs/namei.c all gate it on exactly
> 0, ovl_revalidate_real() gates it the same way, and
> ecryptfs_d_revalidate() hands the value back without invalidating
> anything itself. kernfs_iop_lookup() already answers this same
> condition with a plain "not found".
>
> Feature: kernfs: per-CT entries visibility and permissions configuration
> https://virtuozzo.atlassian.net/browse/VSTOR-142552
> Fixes: 3dd8c2499df6 ("ve/kernfs: hide forbidden entries in container")
> Signed-off-by: Mirian Shilakadze <mirian.shilakadze at virtuozzo.com>
> ---
> fs/kernfs/dir.c | 13 +++++++++++--
> 1 file changed, 11 insertions(+), 2 deletions(-)
>
> diff --git a/fs/kernfs/dir.c b/fs/kernfs/dir.c
> index be680eb98ed4..7947c49ed1a6 100644
> --- a/fs/kernfs/dir.c
> +++ b/fs/kernfs/dir.c
> @@ -1199,8 +1199,17 @@ static int kernfs_dop_revalidate(struct dentry *dentry, unsigned int flags)
> kernfs_info(dentry->d_sb)->ns != kn->ns)
> goto out_bad;
>
> - if (!kernfs_d_visible(kn, kernfs_info(dentry->d_sb)))
> - goto out_bad;
> + if (!kernfs_d_visible(kn, kernfs_info(dentry->d_sb))) {
> + /*
> + * The node is fine, it is only outside this VE's view.
> + * Returning 0 would tell the VFS that the dentry is stale, and
> + * it answers that with d_invalidate(), which detaches every
> + * mount on that dentry in every mount namespace. Report the
> + * name as missing to this caller instead.
> + */
> + up_read(&root->kernfs_rwsem);
> + return -ENOENT;
> + }
Not fully sure, but probably worth preserving behavior in CT as much as possible:
diff --git a/fs/kernfs/dir.c b/fs/kernfs/dir.c
index 7947c49ed1a6f..daa4bc264e6bb 100644
--- a/fs/kernfs/dir.c
+++ b/fs/kernfs/dir.c
@@ -1201,12 +1201,11 @@ static int kernfs_dop_revalidate(struct dentry *dentry, unsigned int flags)
if (!kernfs_d_visible(kn, kernfs_info(dentry->d_sb))) {
/*
- * The node is fine, it is only outside this VE's view.
- * Returning 0 would tell the VFS that the dentry is stale, and
- * it answers that with d_invalidate(), which detaches every
- * mount on that dentry in every mount namespace. Report the
- * name as missing to this caller instead.
+ * Invalidate the node in ve owned mount namespace,
+ * or report no entry in other mount namespace.
*/
+ if (current->nsproxy->mnt_ns->ve_owner == get_exec_env())
+ return 0;
up_read(&root->kernfs_rwsem);
return -ENOENT;
}
This way we still umount and invalidate dentry in CT filesystem, making it look a bit more clean.
>
> up_read(&root->kernfs_rwsem);
> return 1;
--
Best regards, Pavel Tikhomirov
Senior Software Developer, Virtuozzo.
More information about the Devel
mailing list