A per-thread CPU timer holds a reference to the PID of the thread it is attached to and, while it is armed, its node is queued in that thread's posix_cputimers. The task is looked up by that PID. When a non-leader thread exec()s, de_thread() changes which task owns that PID. pid_task(timer->it.cpu.pid, PIDTYPE_PID) then returns NULL, but the node is still queued on tsk, which is alive. timer_lock_sighand() takes a failed lookup to mean that the node is already dequeued, so it has nothing to undo. begin_new_exec() calls posix_cpu_timers_exit(me) right after exec_task_namespaces() and that removes the leftover node, so the state normally stays invisible. But bprm->point_of_no_return is set before de_thread(), so if unshare_files(), set_mm_exe_file(), exec_mmap() or exec_task_namespaces() fails, the task dies before it gets there. exit_itimers() then frees the k_itimer while its node is still queued, and reaping tsk later erases that freed node from the rbtree. In short: the non-leader thread B the parent timer_create(CLOCK_THREAD_CPUTIME_ID) timer_settime() arm_timer() // the node is queued on B execve() de_thread(B) exchange_tids(B, leader) // B's PID now belongs to the leader release_task(leader) __exit_signal(leader) posix_cpu_timers_exit(leader) // cleans leader's queue, not B's __unhash_process(leader) // that PID has no task anymore exec_mmap() mmap_read_lock_killable(old_mm) kill(B, SIGKILL) // -EINTR get_signal() do_exit() exit_itimers() posix_timer_delete() posix_cpu_timer_del() posix_timer_unhash_and_free() // freed while still queued wait4() release_task(B) posix_cpu_timers_exit(B) cleanup_timerqueue() timerqueue_del() // use-after-free KASAN log: BUG: KASAN: slab-use-after-free in timerqueue_del+0x59/0x80 Write of size 8 at addr ffff88801090e7e0 by task poc/99 ... Call Trace: timerqueue_del+0x59/0x80 posix_cpu_timers_exit+0xa8/0xe0 release_task+0x1c6/0xa40 wait_consider_task+0x922/0x15f0 __do_wait+0x3bc/0x3d0 do_wait+0xba/0x1d0 kernel_wait4+0xfe/0x1d0 __do_sys_wait4+0x10b/0x120 do_syscall_64+0xf9/0x540 entry_SYSCALL_64_after_hwframe+0x77/0x7f ... Allocated by task 114: kmem_cache_alloc_noprof+0x143/0x3d0 do_timer_create+0x156/0x890 __x64_sys_timer_create+0x113/0x130 ... Freed by task 104: __kasan_slab_free+0x43/0x70 __rcu_free_sheaf_prepare+0x5c/0x260 rcu_free_sheaf+0x1a/0xc0 rcu_core+0x3d1/0xdf0 handle_softirqs+0x133/0x400 ... The buggy address belongs to the object at ffff88801090e700 which belongs to the cache posix_timers_cache of size 384 The buggy address is located 224 bytes inside of freed 384-byte region [ffff88801090e700, ffff88801090e880) Dequeue the per-thread CPU timers of tsk before the PID changes hands, so that a failed lookup again implies a dequeued node. Process-wide timers are looked up with PIDTYPE_TGID and transfer_pid() moves that link to tsk, so they are left alone. Fixes: 55e8c8eb2c7b ("posix-cpu-timers: Store a reference to a pid not a task") Cc: stable@vger.kernel.org Signed-off-by: Hyunwoo Kim --- Changes in v2: - Add the trigger sequence and the KASAN log to the commit message. - v1: https://lore.kernel.org/all/anfgrsPlUdwBhdrp@v4bel/ --- fs/exec.c | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/fs/exec.c b/fs/exec.c index 745f6eb5279e6..9b2d140bd9f2a 100644 --- a/fs/exec.c +++ b/fs/exec.c @@ -1003,6 +1003,18 @@ static int de_thread(struct task_struct *tsk) * the former thread group leader: */ +#ifdef CONFIG_POSIX_TIMERS + /* + * exchange_tids() hands this thread's PID to the old leader, + * which is reaped right after. The PID lookup in + * timer_lock_sighand() then fails while the per thread CPU + * timers are still queued here, so dequeue them first. + */ + spin_lock(lock); + posix_cpu_timers_exit(tsk); + spin_unlock(lock); +#endif + /* Become a process group leader with the old leader's pid. * The old leader becomes a thread of the this thread group. */ -- 2.43.0