Mirror of the gdb-patches mailing list
 help / color / mirror / Atom feed
From: Matthieu Longo <matthieu.longo@arm.com>
To: Simon Marchi <simon.marchi@efficios.com>
Cc: "gdb-patches@sourceware.org" <gdb-patches@sourceware.org>
Subject: Re: [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files
Date: Mon, 21 Sep 2026 15:26:01 +0100	[thread overview]
Message-ID: <016bc26f-7889-4267-9a7e-00d152bca960@arm.com> (raw)
In-Reply-To: <20260917201755.589524-2-simon.marchi@efficios.com_5b6b04a9_release>

On 18/09/2026 13:21, Simon Marchi wrote:
> On Linux, /proc/<pid> is keyed by the thread-group leader PID.  When the
> leader has exited, some /proc/<pid>/... entries become unavailable even
> though another thread is still alive.  This can happen, for instance,
> when the main thread calls pthread_exit() and another thread continues
> the execution (existing test: gcore-stale-thread, renamed below).
> 
> This causes GDB to fail to read procfs entries such as cmdline, cwd,
> exe, maps, and smaps when it builds the paths from the inferior pid
> after the thread-group leader has exited.
> 
> A /proc/<lwp> entry exists for every live LWP of the process, and for
> the entries listed above its contents are the same as those of
> /proc/<pid>.  Fix this by building the procfs paths from the LWP id of a
> thread GDB believes is alive, rather than from the inferior pid.
> 
> Add get_ptid_for_slash_proc, which returns the PTID to read from, and
> use it in places that read /proc:
> 
>  - linux_info_proc
>  - linux_process_address_in_memtag_page
>  - linux_find_memory_regions_full
>  - linux_fill_prpsinfo
>  - linux_address_in_shadow_stack_mem_range
> 
> linux_vsyscall_range_raw also reads /proc, but I did not change it, as
> it uses a different pattern, "/proc/<pid>/task/<pid>/maps".  It could in
> theory suffer from the same problem.
> 
> The thread is chosen with any_non_exited_thread_of_inferior, which gives
> preference to the selected thread.  I think this is actually a good
> feature.  Some entries in /proc (like stat and status) show some
> thread-specific content, so this lets the user pick which thread "info
> proc" reports on.  Most entries are per-process, the choice of thread
> makes no difference for them.
> 
> linux_fill_prpsinfo is an exception to the above.  It populates the
> NT_PRPSINFO note of a core file, which the kernel (for kernel-generated
> core files) fills in from the thread group leader, even when that leader
> is a zombie:
> 
>   https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1901
> 
> To keep GDB-generated cores looking like kernel-generated ones, it keeps
> reading stat and status from the leader's entry, so the values taken
> from those match what the kernel would write (stat and status are still
> readable when the thread is zombie).  Only the cmdline is read from a
> live LWP, since trying to read that one from a zombie leader doesn't
> work.  This mirrors the kernel's fill_psinfo function, which takes the
> program name and arguments from the mm of the thread doing the dump:
> 
>   https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1530-L1539
> 
> This also fixes a bug, in that core dumps produced while the leader is
> zombie would be missing the NT_PRPSINFO note.  linux_fill_prpsinfo would
> read an empty string from cmdline and return false.  A user-visible
> consequence of that is that loading back the core would not show the
> expected "Core was generated by..." line.
> 
> This is still a best-effort: GDB's view of the thread list may be stale,
> so the thread selected to do the /proc accesses may actually have
> exited but GDB does not know it.
> 
> Since the entry "info proc" reads is no longer necessarily that of the
> thread group leader, replace its "process <pid>" header with:
> 
>   Reading /proc for process <pid>
> 
> or, when the LWP differs from the PID:
> 
>   Reading /proc for process <pid> (LWP <lwp>)
> 
> Since "info proc" now stores the pid given on the command line in a
> ptid_t, whose pid field is an int, reject a value that would not survive
> the conversion, rather than printing one value and reading /proc for
> another.
> 
> Update gdb.base/info-proc.exp for the new header.
> 
> Rename gdb.threads/gcore-stale-thread.{exp,c} to
> gdb.threads/thread-leader-exited.{exp,c} and modify it in a few ways:
> 
>  - give the program a second worker thread
> 
>  - check that we're able to read /proc on Linux even when the leader has
>    exited (and is selected)
> 
>  - check that "info proc stat" reports each worker thread's own LWP id
>    in its "Process:" field, confirming that "info proc" prioritizes the
>    selected thread
> 
>  - load the generated core back to verify that we see the "Core was
>    generated by ..." line, and therefore the NT_PRPSINFO note was
>    generated
> 
>  - enable non-stop using the usual save_vars / append to GDBFLAGS
>    pattern, which fixes a pre-existing bug of not being able to run the
>    test on the native-extended-gdbserver board.
> 
> Bug: https://sourceware.org/bugzilla/show_bug.cgi?id=31207
> Change-Id: I1e6ccfae90551b8e96a5644b51c7652bbda593fa
> Co-Authored-By: Matthieu Longo <matthieu.longo@arm.com>
It looks good to me.

I ran the 2 patches on the top of master, on both AArch64 and x86_64 on our internal CI and I found
2 failing tests on AArch64: step-over-process-exit.exp and missing-thread.exp.
I re-ran them manually and both of them are passing.
This has nothing to do with your patch, but are those tests known for being flickering ?

This is the logs for step-over-process-exit.exp:
```
maint show target-non-stop
Whether the target is always in non-stop mode is auto (currently on).
(gdb) next
[Thread 0xfffff7fc0020 (LWP 291431) (id 1) exited]
[New LWP 291431 (id 3)]
[Thread 0xfffff7e0f120 (LWP 292049) (id 2) exited]
warning: error removing breakpoint 0 at 0xaaaaaaaa087c
Command aborted, thread exited.
Cannot remove breakpoints because program is no longer writable.
Further execution is probably impossible.
(gdb) [Inferior 1 (process 291431) exited normally]
FAIL: gdb.threads/step-over-process-exit.exp: which=other: next (timeout)
```
VS when it passed:
```
maint show target-non-stop
Whether the target is always in non-stop mode is auto (currently on).
(gdb) next
[LWP 1540 (id 2) exited]
[Inferior 1 (process 1539) exited normally]
(gdb) PASS: gdb.threads/step-over-process-exit.exp: which=other: next
```
It looks like there is an issue while removing breakpoints. At first glance, I would say that there
is a race condition in the test when thread 1 exist before 2.
And this, somehow creates a time-out ?

Regarding logs for missing-thread.exp:
```
continue
Continuing.
[New Thread 217413.218469 (id 2)]
[Thread 217413.218469 (id 2) exited]
warning: command aborted, Thread 217413.218469 unexpectedly exited after signal stop event
Remote communication error.  Target disconnected: error while reading: Connection reset by peer.
(gdb) FAIL: gdb.replay/missing-thread.exp: non_stop=on: missing 1 thread log: replay_with_log: continue
```
vs when it passed:
```
continue
Continuing.
[New Thread 1692.1693 (id 2)]

Thread 2 "missing-thread" received signal SIGTRAP, Trace/breakpoint trap.
0x0000fffff7eb6ce0 in clock_nanosleep () from /lib/aarch64-linux-gnu/libc.so.6
(gdb) PASS: gdb.replay/missing-thread.exp: non_stop=on: with unmodified log: replay_with_log: continue
```

  parent reply	other threads:[~2026-09-21 14:27 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-17 20:16 [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Simon Marchi
2026-09-17 20:16 ` [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files Simon Marchi
2026-09-21 10:22 ` [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Matthieu Longo
2026-09-21 13:46   ` Simon Marchi
     [not found] ` <20260917201755.589524-2-simon.marchi@efficios.com_5b6b04a9_release>
2026-09-21 14:26   ` Matthieu Longo [this message]
2026-09-21 14:57     ` [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files Simon Marchi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=016bc26f-7889-4267-9a7e-00d152bca960@arm.com \
    --to=matthieu.longo@arm.com \
    --cc=gdb-patches@sourceware.org \
    --cc=simon.marchi@efficios.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox