From: "Metzger, Markus T" <markus.t.metzger@intel.com>
To: Simon Marchi <simark@simark.ca>
Cc: "gdb-patches@sourceware.org" <gdb-patches@sourceware.org>
Subject: RE: [PATCH] gdb: fix an issue with thread list corruption
Date: Tue, 5 May 2026 05:04:12 +0000 [thread overview]
Message-ID: <DM8PR11MB57494F32508897DA39ABD143DE3E2@DM8PR11MB5749.namprd11.prod.outlook.com> (raw)
In-Reply-To: <f8b93a2c-f249-47d1-8d2d-57d76893feeb@simark.ca>
Hello Simon,
>-----Original Message-----
>From: Simon Marchi <simark@simark.ca>
>Sent: Monday, May 4, 2026 5:30 PM
>To: Metzger, Markus T <markus.t.metzger@intel.com>; gdb-
>patches@sourceware.org
>Subject: Re: [PATCH] gdb: fix an issue with thread list corruption
>
>On 5/4/26 3:16 AM, Markus Metzger wrote:
>> When resuming a target in non-stop mode with 'c -a', the continue
>command
>> uses for_each_thread() to proceed each stopped thread individually. This
>> uses an all_threads_safe() iteration.
>>
>> If one of the stopped threads does an inline step-over, since the target
>> is non-stop, we stop_all_threads(), which involves update_thread_list(),
>> which, in turn, may delete_thread().
>>
>> If this deleted the thread pointed to by the m_next safe iterator member,
>> the above all_threads_safe() iteration will be corrupted.
>>
>> The thread we're proceeding is stopped and there is no reason to delete
>> it. Consequently, there is no reason for all_threads_safe(), which isn't
>> that safe in this scenario.
>>
>> Iterate using all_threads() and inline proceed_thread_callback().
>
>Just to try to make sure I understand the circumstances that lead the
>the failure correcty (and make sure this doesn't just cover up other
>problems):
>
> - thread 1 is stopped on a breakpoint
> - thread 2 is executing
> - displaced stepping is disabled
> - the target doesn't report thread events to the core
> - you "continue -a"
> - thread 1 starts an inline step-over, which calls stop_all_threads,
> which calls update_thread_list
> - meanwhile, thread 2 exits
> - update_thread_list causes thread 2 to be deleted
> - the safe iterator's m_next field now points to a deleted thread info
>
>I thought: don't we now ask the target to report thread exit events when
>doing steps now (to handle step over exit)? But we only ask for to
>report the exit event for the stepping thread, so the target wouldn't
>report the exit of thread 2.
I saw this with our new GPU target. Even if we asked for exit events,
our target would only be able to provide them when single-stepping
(i.e. for thread 1) and not for running threads (i.e. for thread 2, assuming
GDB even asked for it).
I believe that AMD GPU targets suffer from the same limitation.
>If my understand above is corerct, then I agree with your reasoning.
>The currently iterated on thread is stopped and should not disappear,
>except maybe on a misbehaving target (in which case we'd fix the
>target). Maybe it could happen with a remote target where communication
>breaks during this update_thread_list, but that is notoriously difficult
>to handle correctly.
Our GPU target is a non-stop remote target. I wouldn't rule out that it still
misbehaves somewhat, but the scenario described above is the normal behavior.
We will be able to step over breakpoints inside the target in a future version, so
stopping all threads for an inline step-over will not be necessary anymore. But
I'd argue that this will simply hide the bug.
>
>> @@ -762,7 +739,30 @@ continue_1 (int all_threads)
>> scoped_disable_commit_resumed disable_commit_resumed
>> ("continue all threads in non-stop");
>>
>> - for_each_thread (proceed_thread_callback);
>> + /* Do not use all_threads_safe in case threads get removed while
>> + resuming THREAD. */
>
>I understand this comment because I just read your commit message, but I
>don't think I would understand it in isolation. Because
>all_threads_safe is usually used to handle cases where threads (the
>current one) get removed while iterating. I suggest:
>
> /* Do not use all_threads_safe, because it's possible for the next thread
> to get removed while resuming THREAD. We know that thread is
>stopped
> and should not disappear under our feet. */
>
>Then if I wanted to understand more about the circumstances where this
>check was added, I would do a git blame and find the commit message.
I added your comment to v2.
Thanks,
Markus.
Intel Deutschland GmbH
Registered Address: Dornacher Strasse 1, 85622 Feldkirchen, Germany
Tel: +49 89 991 430, www.intel.de
Managing Directors: Harry Demas, Jeffrey Schneiderman, Yin Chong Sorrell
Chairperson of the Supervisory Board: Nicole Lau
Registered Seat: Munich
Commercial Register: Amtsgericht Muenchen HRB 186928
next prev parent reply other threads:[~2026-05-05 5:05 UTC|newest]
Thread overview: 36+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-05-04 7:16 [PATCH] gdb, btrace: fix wrong #endif comments Markus Metzger
2026-05-04 7:16 ` [PATCH] gdb: fix an issue with thread list corruption Markus Metzger
2026-05-04 15:30 ` Simon Marchi
2026-05-05 5:04 ` Metzger, Markus T [this message]
2026-05-05 15:51 ` Simon Marchi
2026-05-06 5:43 ` Metzger, Markus T
2026-05-06 18:36 ` Simon Marchi
2026-05-07 7:36 ` Metzger, Markus T
2026-05-07 15:11 ` Simon Marchi
2026-05-11 7:25 ` Metzger, Markus T
2026-05-11 10:05 ` Metzger, Markus T
2026-05-12 6:36 ` Metzger, Markus T
2026-05-13 14:33 ` Tom Tromey
2026-05-13 17:05 ` Tom Tromey
2026-05-14 22:07 ` Thiago Jung Bauermann
2026-05-04 7:16 ` [PATCH] gdb, remote: fix notify debug nullptr dereference Markus Metzger
2026-05-04 15:40 ` Simon Marchi
2026-05-04 7:16 ` [PATCH] gdb, testsuite: increase timeout in gdb.threads/attach-non-stop.exp Markus Metzger
2026-05-04 15:49 ` Simon Marchi
2026-05-05 6:15 ` Metzger, Markus T
2026-05-05 15:54 ` Simon Marchi
2026-05-06 6:33 ` Metzger, Markus T
2026-05-04 7:16 ` [PATCH] gdb: update store_integer's comment Markus Metzger
2026-05-04 15:52 ` Simon Marchi
2026-05-04 7:16 ` [PATCH] gdb: use correct target in notify_thread_exited() Markus Metzger
2026-05-04 16:06 ` Simon Marchi
2026-05-05 7:56 ` Metzger, Markus T
2026-05-05 16:10 ` Simon Marchi
2026-05-06 6:56 ` Metzger, Markus T
2026-05-04 7:16 ` [PATCH v2] gdb, btrace: support libipt v2.2 events Markus Metzger
2026-05-04 12:08 ` Eli Zaretskii
2026-05-04 14:00 ` Metzger, Markus T
2026-05-04 15:04 ` Eli Zaretskii
2026-05-05 4:32 ` Metzger, Markus T
2026-05-05 5:12 ` Eli Zaretskii
2026-05-04 14:44 ` [PATCH] gdb, btrace: fix wrong #endif comments Simon Marchi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DM8PR11MB57494F32508897DA39ABD143DE3E2@DM8PR11MB5749.namprd11.prod.outlook.com \
--to=markus.t.metzger@intel.com \
--cc=gdb-patches@sourceware.org \
--cc=simark@simark.ca \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox