[Devel] [PATCH vz9 3/3] eventpoll: fix ep_remove struct eventpoll / struct file UAF

Vladimir Riabchun vladimir.riabchun at virtuozzo.com
Wed Jul 8 13:12:53 MSK 2026



On 7/7/26 16:06, Eva Kurchatova wrote:
> [ Upstream commit a6dc643c69311677c574a0f17a3f4d66a5f3744b ]
> 
> ep_remove() (via ep_remove_file()) cleared file->f_ep under
> file->f_lock but then kept using @file inside the critical section
> (is_file_epoll(), hlist_del_rcu() through the head, spin_unlock).
> A concurrent __fput() taking the eventpoll_release() fastpath in
> that window observed the transient NULL, skipped
> eventpoll_release_file() and ran to f_op->release / file_free().
> 
> For the epoll-watches-epoll case, f_op->release is
> ep_eventpoll_release() -> ep_clear_and_put() -> ep_free(), which
> kfree()s the watched struct eventpoll. Its embedded ->refs
> hlist_head is exactly where epi->fllink.pprev points, so the
> subsequent hlist_del_rcu()'s "*pprev = next" scribbles into freed
> kmalloc-192 memory.
> 
> In addition, struct file is SLAB_TYPESAFE_BY_RCU, so the slot
> backing @file could be recycled by alloc_empty_file() --
> reinitializing f_lock and f_ep -- while ep_remove() is still
> nominally inside that lock. The upshot is an attacker-controllable
> kmem_cache_free() against the wrong slab cache.
> 
> Pin @file via epi_fget() at the top of ep_remove() and gate the
> critical section on the pin succeeding. With the pin held @file
> cannot reach refcount zero, which holds __fput() off and
> transitively keeps the watched struct eventpoll alive across the
> hlist_del_rcu() and the f_lock use, closing both UAFs.
> 
> If the pin fails @file has already reached refcount zero and its
> __fput() is in flight. Because we bailed before clearing f_ep,
> that path takes the eventpoll_release() slow path into
> eventpoll_release_file() and blocks on ep->mtx until the waiter
> side's ep_clear_and_put() drops it. The bailed epi's share of
> ep->refcount stays intact, so the trailing ep_refcount_dec_and_test()
> in ep_clear_and_put() cannot free the eventpoll out from under
> eventpoll_release_file(); the orphaned epi is then cleaned up
> there.
> 
> A successful pin also proves we are not racing
> eventpoll_release_file() on this epi, so drop the now-redundant
> re-check of epi->dying under f_lock. The cheap lockless
> READ_ONCE(epi->dying) fast-path bailout stays.
> 
> Fixes: 58c9b016e128 ("epoll: use refcount to reduce ep_mutex contention")
> Reported-by: Jaeyoung Chung <jjy600901 at snu.ac.kr>
> Link: https://patch.msgid.link/20260423-work-epoll-uaf-v1-6-2470f9eec0f5@kernel.org
> Signed-off-by: Christian Brauner (Amutable) <brauner at kernel.org>
> Signed-off-by: Quentin Schulz <quentin.schulz at cherry.de>
> Signed-off-by: Sasha Levin <sashal at kernel.org>
> (cherry picked from commit 9324de74a3a59b9fde9b62ee45ebaa71458ba2e5)
> 
> Signed-off-by: Eva Kurchatova <eva.kurchatova at virtuozzo.com>
> 
> https://virtuozzo.atlassian.net/browse/VSTOR-137490
> Feature: fix epoll cve
> ---
>   fs/eventpoll.c | 36 +++++++++++++++++++++++++++++-------
>   1 file changed, 29 insertions(+), 7 deletions(-)
> 
> diff --git a/fs/eventpoll.c b/fs/eventpoll.c
> index 7ac22aadaf8b..c937325c614b 100644
> --- a/fs/eventpoll.c
> +++ b/fs/eventpoll.c
> @@ -715,6 +715,8 @@ static void ep_free(struct eventpoll *ep)
>   	kfree_rcu(ep, rcu);
>   }
>   
> +static struct file *epi_fget(const struct epitem *epi);
> +
>   /*
>    * Removes a "struct epitem" from the eventpoll RB tree and deallocates
>    * all the associated resources. Must be called with "mtx" held.
> @@ -725,7 +727,7 @@ static void ep_free(struct eventpoll *ep)
>    */
>   static bool __ep_remove(struct eventpoll *ep, struct epitem *epi, bool force)
>   {
> -	struct file *file = epi->ffd.file;
> +	struct file *file = NULL;
>   	struct epitems_head *to_free;
>   	struct hlist_head *head;
>   
> @@ -736,13 +738,29 @@ static bool __ep_remove(struct eventpoll *ep, struct epitem *epi, bool force)
>   	 */
>   	ep_unregister_pollwait(ep, epi);
>   
> -	/* Remove the current item from the list of epoll hooks */
> -	spin_lock(&file->f_lock);
> -	if (epi->dying && !force) {
> -		spin_unlock(&file->f_lock);
> -		return false;
> +	if (!force) {
> +		/* cheap sync with eventpoll_release_file() */
> +		if (unlikely(READ_ONCE(epi->dying)))
> +			return false;
> +
> +		/*
> +		 * If we manage to grab a reference it means we're not in
> +		 * eventpoll_release_file() and aren't going to be.
> +		 */
> +		file = epi_fget(epi);
> +		if (!file)
> +			return false;
> +	} else {
> +		/*
> +		 * For force (eventpoll_release_file()), the caller already owns
> +		 * the file and set epi->dying; the file is safe to access
> +		 * directly but its refcount is zero so epi_fget() would fail.
> +		 */
> +		file = epi->ffd.file;
>   	}

I don't like this. It seems to be a correct backport of original patch,
but it is very hard to verify it.

I'd prefer taking stable backport patchsets, check this thread:
https://lore.kernel.org/all/20260628161933.532572-1-guanwentao@uniontech.com/

This will make the review and following fixes of the code much easier.

Two other patches are fine.

>   
> +	/* Remove the current item from the list of epoll hooks */
> +	spin_lock(&file->f_lock);
>   	to_free = NULL;
>   	head = file->f_ep;
>   	if (head->first == &epi->fllink && !epi->fllink.next) {
> @@ -776,6 +794,10 @@ static bool __ep_remove(struct eventpoll *ep, struct epitem *epi, bool force)
>   	call_rcu(&epi->rcu, epi_rcu_free);
>   
>   	percpu_counter_dec(&ep->user->epoll_watches);
> +	if (!force) {
> +		/* Put down the pinned file reference */
> +		fput(file);
> +	}
>   	return true;
>   }
>   
> @@ -993,7 +1015,7 @@ void eventpoll_release_file(struct file *file)
>   	spin_lock(&file->f_lock);
>   	if (file->f_ep && file->f_ep->first) {
>   		epi = hlist_entry(file->f_ep->first, struct epitem, fllink);
> -		epi->dying = true;
> +		WRITE_ONCE(epi->dying, true);
>   		spin_unlock(&file->f_lock);
>   
>   		/*

-- 
Best regards, Riabchun Vladimir
Linux Kernel Developer, Virtuozzo



More information about the Devel mailing list