From: Simon Marchi <simon.marchi@efficios.com>
To: Matthieu Longo <matthieu.longo@arm.com>
Cc: "gdb-patches@sourceware.org" <gdb-patches@sourceware.org>
Subject: Re: [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files
Date: Mon, 21 Sep 2026 10:57:36 -0400 [thread overview]
Message-ID: <de68525a-151d-4298-88c0-400a0815fe83@efficios.com> (raw)
In-Reply-To: <016bc26f-7889-4267-9a7e-00d152bca960@arm.com>
On 9/21/26 10:26 AM, Matthieu Longo wrote:
> On 18/09/2026 13:21, Simon Marchi wrote:
>> On Linux, /proc/<pid> is keyed by the thread-group leader PID. When the
>> leader has exited, some /proc/<pid>/... entries become unavailable even
>> though another thread is still alive. This can happen, for instance,
>> when the main thread calls pthread_exit() and another thread continues
>> the execution (existing test: gcore-stale-thread, renamed below).
>>
>> This causes GDB to fail to read procfs entries such as cmdline, cwd,
>> exe, maps, and smaps when it builds the paths from the inferior pid
>> after the thread-group leader has exited.
>>
>> A /proc/<lwp> entry exists for every live LWP of the process, and for
>> the entries listed above its contents are the same as those of
>> /proc/<pid>. Fix this by building the procfs paths from the LWP id of a
>> thread GDB believes is alive, rather than from the inferior pid.
>>
>> Add get_ptid_for_slash_proc, which returns the PTID to read from, and
>> use it in places that read /proc:
>>
>> - linux_info_proc
>> - linux_process_address_in_memtag_page
>> - linux_find_memory_regions_full
>> - linux_fill_prpsinfo
>> - linux_address_in_shadow_stack_mem_range
>>
>> linux_vsyscall_range_raw also reads /proc, but I did not change it, as
>> it uses a different pattern, "/proc/<pid>/task/<pid>/maps". It could in
>> theory suffer from the same problem.
>>
>> The thread is chosen with any_non_exited_thread_of_inferior, which gives
>> preference to the selected thread. I think this is actually a good
>> feature. Some entries in /proc (like stat and status) show some
>> thread-specific content, so this lets the user pick which thread "info
>> proc" reports on. Most entries are per-process, the choice of thread
>> makes no difference for them.
>>
>> linux_fill_prpsinfo is an exception to the above. It populates the
>> NT_PRPSINFO note of a core file, which the kernel (for kernel-generated
>> core files) fills in from the thread group leader, even when that leader
>> is a zombie:
>>
>> https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1901
>>
>> To keep GDB-generated cores looking like kernel-generated ones, it keeps
>> reading stat and status from the leader's entry, so the values taken
>> from those match what the kernel would write (stat and status are still
>> readable when the thread is zombie). Only the cmdline is read from a
>> live LWP, since trying to read that one from a zombie leader doesn't
>> work. This mirrors the kernel's fill_psinfo function, which takes the
>> program name and arguments from the mm of the thread doing the dump:
>>
>> https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1530-L1539
>>
>> This also fixes a bug, in that core dumps produced while the leader is
>> zombie would be missing the NT_PRPSINFO note. linux_fill_prpsinfo would
>> read an empty string from cmdline and return false. A user-visible
>> consequence of that is that loading back the core would not show the
>> expected "Core was generated by..." line.
>>
>> This is still a best-effort: GDB's view of the thread list may be stale,
>> so the thread selected to do the /proc accesses may actually have
>> exited but GDB does not know it.
>>
>> Since the entry "info proc" reads is no longer necessarily that of the
>> thread group leader, replace its "process <pid>" header with:
>>
>> Reading /proc for process <pid>
>>
>> or, when the LWP differs from the PID:
>>
>> Reading /proc for process <pid> (LWP <lwp>)
>>
>> Since "info proc" now stores the pid given on the command line in a
>> ptid_t, whose pid field is an int, reject a value that would not survive
>> the conversion, rather than printing one value and reading /proc for
>> another.
>>
>> Update gdb.base/info-proc.exp for the new header.
>>
>> Rename gdb.threads/gcore-stale-thread.{exp,c} to
>> gdb.threads/thread-leader-exited.{exp,c} and modify it in a few ways:
>>
>> - give the program a second worker thread
>>
>> - check that we're able to read /proc on Linux even when the leader has
>> exited (and is selected)
>>
>> - check that "info proc stat" reports each worker thread's own LWP id
>> in its "Process:" field, confirming that "info proc" prioritizes the
>> selected thread
>>
>> - load the generated core back to verify that we see the "Core was
>> generated by ..." line, and therefore the NT_PRPSINFO note was
>> generated
>>
>> - enable non-stop using the usual save_vars / append to GDBFLAGS
>> pattern, which fixes a pre-existing bug of not being able to run the
>> test on the native-extended-gdbserver board.
>>
>> Bug: https://sourceware.org/bugzilla/show_bug.cgi?id=31207
>> Change-Id: I1e6ccfae90551b8e96a5644b51c7652bbda593fa
>> Co-Authored-By: Matthieu Longo <matthieu.longo@arm.com>
> It looks good to me.
>
> I ran the 2 patches on the top of master, on both AArch64 and x86_64 on our internal CI and I found
> 2 failing tests on AArch64: step-over-process-exit.exp and missing-thread.exp.
> I re-ran them manually and both of them are passing.
> This has nothing to do with your patch, but are those tests known for being flickering ?
>
> This is the logs for step-over-process-exit.exp:
> ```
> maint show target-non-stop
> Whether the target is always in non-stop mode is auto (currently on).
> (gdb) next
> [Thread 0xfffff7fc0020 (LWP 291431) (id 1) exited]
> [New LWP 291431 (id 3)]
> [Thread 0xfffff7e0f120 (LWP 292049) (id 2) exited]
> warning: error removing breakpoint 0 at 0xaaaaaaaa087c
> Command aborted, thread exited.
> Cannot remove breakpoints because program is no longer writable.
> Further execution is probably impossible.
> (gdb) [Inferior 1 (process 291431) exited normally]
> FAIL: gdb.threads/step-over-process-exit.exp: which=other: next (timeout)
This one looks odd, you get a notification that LWP 291431 exits, then a
notification that there is a "new" LWP 291431. It sounds like maybe we
processed an exit event for that thread, followed by another event, that
made GDB think there was a new thread.
> ```
> VS when it passed:
> ```
> maint show target-non-stop
> Whether the target is always in non-stop mode is auto (currently on).
> (gdb) next
> [LWP 1540 (id 2) exited]
> [Inferior 1 (process 1539) exited normally]
> (gdb) PASS: gdb.threads/step-over-process-exit.exp: which=other: next
> ```
> It looks like there is an issue while removing breakpoints. At first glance, I would say that there
> is a race condition in the test when thread 1 exist before 2.
> And this, somehow creates a time-out ?
>
> Regarding logs for missing-thread.exp:
> ```
> continue
> Continuing.
> [New Thread 217413.218469 (id 2)]
> [Thread 217413.218469 (id 2) exited]
> warning: command aborted, Thread 217413.218469 unexpectedly exited after signal stop event
> Remote communication error. Target disconnected: error while reading: Connection reset by peer.
> (gdb) FAIL: gdb.replay/missing-thread.exp: non_stop=on: missing 1 thread log: replay_with_log: continue
> ```
> vs when it passed:
> ```
> continue
> Continuing.
> [New Thread 1692.1693 (id 2)]
>
> Thread 2 "missing-thread" received signal SIGTRAP, Trace/breakpoint trap.
> 0x0000fffff7eb6ce0 in clock_nanosleep () from /lib/aarch64-linux-gnu/libc.so.6
> (gdb) PASS: gdb.replay/missing-thread.exp: non_stop=on: with unmodified log: replay_with_log: continue
> ```
The message in the failure in this one sounds like the remote side
(gdbreplay) suddenly exiting. Not sure if it's just exiting cleanly
because it's done, or if it's crashing.
Some threads tests are known to be racy, ideally it would take someone
to invest time in each of them to understand and fix the race
conditions.
Could you please open some bugs to document those failures in Bugzilla
(if there aren't already)?
Simon
prev parent reply other threads:[~2026-09-21 15:06 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 20:16 [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Simon Marchi
2026-09-17 20:16 ` [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files Simon Marchi
2026-09-21 10:22 ` [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Matthieu Longo
2026-09-21 13:46 ` Simon Marchi
[not found] ` <20260917201755.589524-2-simon.marchi@efficios.com_5b6b04a9_release>
2026-09-21 14:26 ` [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files Matthieu Longo
2026-09-21 14:57 ` Simon Marchi [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=de68525a-151d-4298-88c0-400a0815fe83@efficios.com \
--to=simon.marchi@efficios.com \
--cc=gdb-patches@sourceware.org \
--cc=matthieu.longo@arm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox