* [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior
@ 2026-09-17 20:16 Simon Marchi
2026-09-17 20:16 ` [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files Simon Marchi
` (2 more replies)
0 siblings, 3 replies; 6+ messages in thread
From: Simon Marchi @ 2026-09-17 20:16 UTC (permalink / raw)
To: gdb-patches; +Cc: Simon Marchi
any_non_exited_thread_of_inferior prefers to return the selected thread:
/* Prefer the current thread, if there's one. */
if (inf == current_inferior () && inferior_ptid != null_ptid)
return inferior_thread ();
... but the selected thread can be an exited thread. This can happen in
non-stop when the selected thread exits, in which case it is kept around
but marked as exited (see thread_info::deletable).
Check the current thread's state before returning it, and fall through
to the existing loop otherwise.
Note that we also have any_live_thread_of_inferior, which is similar but
differs in subtle ways. We should probably check if we can merge them
at some point.
Change-Id: I620ee200037c6a3fd3d59b96f16750d6b962c724
---
gdb/thread.c | 8 ++++++--
1 file changed, 6 insertions(+), 2 deletions(-)
diff --git a/gdb/thread.c b/gdb/thread.c
index 0f86f6efe380..e14254619e34 100644
--- a/gdb/thread.c
+++ b/gdb/thread.c
@@ -670,9 +670,13 @@ any_non_exited_thread_of_inferior (inferior *inf)
{
gdb_assert (inf->pid != 0);
- /* Prefer the current thread, if there's one. */
+ /* Prefer the current thread, if there's one and it hasn't exited. */
if (inf == current_inferior () && inferior_ptid != null_ptid)
- return inferior_thread ();
+ {
+ if (thread_info *curr_thr = inferior_thread ();
+ curr_thr->state () != THREAD_EXITED)
+ return curr_thr;
+ }
for (thread_info &tp : inf->non_exited_threads ())
return &tp;
base-commit: 0595410d007e5f628e14717cd268e557e166675f
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files
2026-09-17 20:16 [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Simon Marchi
@ 2026-09-17 20:16 ` Simon Marchi
2026-09-21 10:22 ` [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Matthieu Longo
[not found] ` <20260917201755.589524-2-simon.marchi@efficios.com_5b6b04a9_release>
2 siblings, 0 replies; 6+ messages in thread
From: Simon Marchi @ 2026-09-17 20:16 UTC (permalink / raw)
To: gdb-patches; +Cc: Simon Marchi, Matthieu Longo
On Linux, /proc/<pid> is keyed by the thread-group leader PID. When the
leader has exited, some /proc/<pid>/... entries become unavailable even
though another thread is still alive. This can happen, for instance,
when the main thread calls pthread_exit() and another thread continues
the execution (existing test: gcore-stale-thread, renamed below).
This causes GDB to fail to read procfs entries such as cmdline, cwd,
exe, maps, and smaps when it builds the paths from the inferior pid
after the thread-group leader has exited.
A /proc/<lwp> entry exists for every live LWP of the process, and for
the entries listed above its contents are the same as those of
/proc/<pid>. Fix this by building the procfs paths from the LWP id of a
thread GDB believes is alive, rather than from the inferior pid.
Add get_ptid_for_slash_proc, which returns the PTID to read from, and
use it in places that read /proc:
- linux_info_proc
- linux_process_address_in_memtag_page
- linux_find_memory_regions_full
- linux_fill_prpsinfo
- linux_address_in_shadow_stack_mem_range
linux_vsyscall_range_raw also reads /proc, but I did not change it, as
it uses a different pattern, "/proc/<pid>/task/<pid>/maps". It could in
theory suffer from the same problem.
The thread is chosen with any_non_exited_thread_of_inferior, which gives
preference to the selected thread. I think this is actually a good
feature. Some entries in /proc (like stat and status) show some
thread-specific content, so this lets the user pick which thread "info
proc" reports on. Most entries are per-process, the choice of thread
makes no difference for them.
linux_fill_prpsinfo is an exception to the above. It populates the
NT_PRPSINFO note of a core file, which the kernel (for kernel-generated
core files) fills in from the thread group leader, even when that leader
is a zombie:
https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1901
To keep GDB-generated cores looking like kernel-generated ones, it keeps
reading stat and status from the leader's entry, so the values taken
from those match what the kernel would write (stat and status are still
readable when the thread is zombie). Only the cmdline is read from a
live LWP, since trying to read that one from a zombie leader doesn't
work. This mirrors the kernel's fill_psinfo function, which takes the
program name and arguments from the mm of the thread doing the dump:
https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1530-L1539
This also fixes a bug, in that core dumps produced while the leader is
zombie would be missing the NT_PRPSINFO note. linux_fill_prpsinfo would
read an empty string from cmdline and return false. A user-visible
consequence of that is that loading back the core would not show the
expected "Core was generated by..." line.
This is still a best-effort: GDB's view of the thread list may be stale,
so the thread selected to do the /proc accesses may actually have
exited but GDB does not know it.
Since the entry "info proc" reads is no longer necessarily that of the
thread group leader, replace its "process <pid>" header with:
Reading /proc for process <pid>
or, when the LWP differs from the PID:
Reading /proc for process <pid> (LWP <lwp>)
Since "info proc" now stores the pid given on the command line in a
ptid_t, whose pid field is an int, reject a value that would not survive
the conversion, rather than printing one value and reading /proc for
another.
Update gdb.base/info-proc.exp for the new header.
Rename gdb.threads/gcore-stale-thread.{exp,c} to
gdb.threads/thread-leader-exited.{exp,c} and modify it in a few ways:
- give the program a second worker thread
- check that we're able to read /proc on Linux even when the leader has
exited (and is selected)
- check that "info proc stat" reports each worker thread's own LWP id
in its "Process:" field, confirming that "info proc" prioritizes the
selected thread
- load the generated core back to verify that we see the "Core was
generated by ..." line, and therefore the NT_PRPSINFO note was
generated
- enable non-stop using the usual save_vars / append to GDBFLAGS
pattern, which fixes a pre-existing bug of not being able to run the
test on the native-extended-gdbserver board.
Bug: https://sourceware.org/bugzilla/show_bug.cgi?id=31207
Change-Id: I1e6ccfae90551b8e96a5644b51c7652bbda593fa
Co-Authored-By: Matthieu Longo <matthieu.longo@arm.com>
---
gdb/linux-tdep.c | 104 ++++++++++-----
gdb/testsuite/gdb.base/info-proc.exp | 3 +-
.../gdb.threads/gcore-stale-thread.exp | 57 --------
...-stale-thread.c => thread-leader-exited.c} | 30 ++++-
.../gdb.threads/thread-leader-exited.exp | 122 ++++++++++++++++++
5 files changed, 222 insertions(+), 94 deletions(-)
delete mode 100644 gdb/testsuite/gdb.threads/gcore-stale-thread.exp
rename gdb/testsuite/gdb.threads/{gcore-stale-thread.c => thread-leader-exited.c} (64%)
create mode 100644 gdb/testsuite/gdb.threads/thread-leader-exited.exp
diff --git a/gdb/linux-tdep.c b/gdb/linux-tdep.c
index 05ee71fc96f4..2086dd2c1fd7 100644
--- a/gdb/linux-tdep.c
+++ b/gdb/linux-tdep.c
@@ -456,6 +456,30 @@ linux_has_shared_address_space (struct gdbarch *gdbarch)
return linux_is_uclinux ();
}
+/* Return the ptid of a thread of the current inferior whose lwp id can be
+ used to build "/proc/<lwp>" paths.
+
+ Through any_non_exited_thread_of_inferior, this function prefers the
+ currently selected thread, if it has not exited. For most information in
+ /proc, it does not matter which thread of a process we pick. But a few
+ entries (like stat counters) are thread-specific. This allows the user
+ to focus a particular thread and get specific information about that
+ thread.
+
+ Throw an error if the inferior has no thread that GDB believes is
+ alive. */
+
+static ptid_t
+get_ptid_for_slash_proc ()
+{
+ if (thread_info *thread
+ = any_non_exited_thread_of_inferior (current_inferior ());
+ thread != nullptr)
+ return thread->ptid;
+
+ error (_("Could not find a non-exited thread to read /proc from."));
+}
+
/* This is how we want PTIDs from core files to be printed. */
static std::string
@@ -841,9 +865,7 @@ static void
linux_info_proc (struct gdbarch *gdbarch, const char *args,
enum info_proc_what what)
{
- /* A long is used for pid instead of an int to avoid a loss of precision
- compiler warning from the output of strtoul. */
- long pid;
+ ptid_t ptid;
int cmdline_f = (what == IP_MINIMAL || what == IP_CMDLINE || what == IP_ALL);
int cwd_f = (what == IP_MINIMAL || what == IP_CWD || what == IP_ALL);
int environ_f = (what == IP_ENVIRON || what == IP_ALL);
@@ -858,7 +880,16 @@ linux_info_proc (struct gdbarch *gdbarch, const char *args,
{
char *tem;
- pid = strtoul (args, &tem, 10);
+ /* A long is used for pid instead of an int to avoid a loss of precision
+ compiler warning from the output of strtoul. */
+ long pid = strtoul (args, &tem, 10);
+
+ /* ptid_t holds the pid as an int, so reject anything that would not
+ fit an int. */
+ if (pid > INT_MAX)
+ error (_("Invalid process id: %s"), args);
+
+ ptid = ptid_t (pid, pid);
args = tem;
}
else
@@ -868,17 +899,22 @@ linux_info_proc (struct gdbarch *gdbarch, const char *args,
if (current_inferior ()->fake_pid_p)
error (_("Can't determine the current process's PID: you must name one."));
- pid = current_inferior ()->pid;
+ ptid = get_ptid_for_slash_proc ();
}
args = skip_spaces (args);
if (args && args[0])
error (_("Too many parameters: %s"), args);
- gdb_printf (_("process %ld\n"), pid);
+ if (ptid.pid () == ptid.lwp ())
+ gdb_printf (_("Reading /proc for process %d\n"), ptid.pid ());
+ else
+ gdb_printf (_("Reading /proc for process %d (LWP %ld)\n"),
+ ptid.pid (), ptid.lwp ());
+
if (cmdline_f)
{
- xsnprintf (filename, sizeof filename, "/proc/%ld/cmdline", pid);
+ xsnprintf (filename, sizeof filename, "/proc/%ld/cmdline", ptid.lwp ());
gdb_byte *buffer;
LONGEST len = target_fileio_read_alloc (nullptr, filename, &buffer);
@@ -900,7 +936,7 @@ linux_info_proc (struct gdbarch *gdbarch, const char *args,
}
if (cwd_f)
{
- xsnprintf (filename, sizeof filename, "/proc/%ld/cwd", pid);
+ xsnprintf (filename, sizeof filename, "/proc/%ld/cwd", ptid.lwp ());
std::optional<std::string> contents
= target_fileio_readlink (NULL, filename, &target_errno);
if (contents.has_value ())
@@ -910,7 +946,7 @@ linux_info_proc (struct gdbarch *gdbarch, const char *args,
}
if (environ_f)
{
- xsnprintf (filename, sizeof filename, "/proc/%ld/environ", pid);
+ xsnprintf (filename, sizeof filename, "/proc/%ld/environ", ptid.lwp ());
gdb_byte *buffer;
LONGEST len = target_fileio_read_alloc (nullptr, filename, &buffer);
@@ -934,7 +970,7 @@ linux_info_proc (struct gdbarch *gdbarch, const char *args,
}
if (exe_f)
{
- xsnprintf (filename, sizeof filename, "/proc/%ld/exe", pid);
+ xsnprintf (filename, sizeof filename, "/proc/%ld/exe", ptid.lwp ());
std::optional<std::string> contents
= target_fileio_readlink (NULL, filename, &target_errno);
if (contents.has_value ())
@@ -944,7 +980,7 @@ linux_info_proc (struct gdbarch *gdbarch, const char *args,
}
if (mappings_f)
{
- xsnprintf (filename, sizeof filename, "/proc/%ld/maps", pid);
+ xsnprintf (filename, sizeof filename, "/proc/%ld/maps", ptid.lwp ());
gdb::unique_xmalloc_ptr<char> map
= target_fileio_read_stralloc (NULL, filename);
if (map != NULL)
@@ -989,7 +1025,7 @@ linux_info_proc (struct gdbarch *gdbarch, const char *args,
}
if (status_f)
{
- xsnprintf (filename, sizeof filename, "/proc/%ld/status", pid);
+ xsnprintf (filename, sizeof filename, "/proc/%ld/status", ptid.lwp ());
gdb::unique_xmalloc_ptr<char> status
= target_fileio_read_stralloc (NULL, filename);
if (status)
@@ -999,7 +1035,7 @@ linux_info_proc (struct gdbarch *gdbarch, const char *args,
}
if (stat_f)
{
- xsnprintf (filename, sizeof filename, "/proc/%ld/stat", pid);
+ xsnprintf (filename, sizeof filename, "/proc/%ld/stat", ptid.lwp ());
gdb::unique_xmalloc_ptr<char> statstr
= target_fileio_read_stralloc (NULL, filename);
if (statstr)
@@ -1670,9 +1706,8 @@ linux_process_address_in_memtag_page (CORE_ADDR address)
if (current_inferior ()->fake_pid_p)
return false;
- pid_t pid = current_inferior ()->pid;
-
- std::string smaps_file = string_printf ("/proc/%d/smaps", pid);
+ ptid_t ptid = get_ptid_for_slash_proc ();
+ std::string smaps_file = string_printf ("/proc/%ld/smaps", ptid.lwp ());
gdb::unique_xmalloc_ptr<char> data
= target_fileio_read_stralloc (NULL, smaps_file.c_str ());
@@ -1730,7 +1765,6 @@ linux_find_memory_regions_full (struct gdbarch *gdbarch,
linux_dump_mapping_p_ftype *should_dump_mapping_p,
linux_find_memory_region_ftype func)
{
- pid_t pid;
/* Default dump behavior of coredump_filter (0x33), according to
Documentation/filesystems/proc.txt from the Linux kernel
tree. */
@@ -1743,12 +1777,12 @@ linux_find_memory_regions_full (struct gdbarch *gdbarch,
if (current_inferior ()->fake_pid_p)
return false;
- pid = current_inferior ()->pid;
+ ptid_t ptid = get_ptid_for_slash_proc ();
if (use_coredump_filter)
{
std::string core_dump_filter_name
- = string_printf ("/proc/%d/coredump_filter", pid);
+ = string_printf ("/proc/%ld/coredump_filter", ptid.lwp ());
gdb::unique_xmalloc_ptr<char> coredumpfilterdata
= target_fileio_read_stralloc (NULL, core_dump_filter_name.c_str ());
@@ -1762,7 +1796,7 @@ linux_find_memory_regions_full (struct gdbarch *gdbarch,
}
}
- std::string maps_filename = string_printf ("/proc/%d/smaps", pid);
+ std::string maps_filename = string_printf ("/proc/%ld/smaps", ptid.lwp ());
gdb::unique_xmalloc_ptr<char> data
= target_fileio_read_stralloc (NULL, maps_filename.c_str ());
@@ -1770,7 +1804,7 @@ linux_find_memory_regions_full (struct gdbarch *gdbarch,
if (data == NULL)
{
/* Older Linux kernels did not support /proc/PID/smaps. */
- maps_filename = string_printf ("/proc/%d/maps", pid);
+ maps_filename = string_printf ("/proc/%ld/maps", ptid.lwp ());
data = target_fileio_read_stralloc (NULL, maps_filename.c_str ());
if (data == nullptr)
@@ -2273,8 +2307,6 @@ linux_fill_prpsinfo (struct elf_internal_linux_prpsinfo *p)
const char *prog_state;
/* The state of the process. */
char pr_sname;
- /* The PID of the program which generated the corefile. */
- pid_t pid;
/* Process flags. */
unsigned int pr_flag;
/* Process nice value. */
@@ -2284,9 +2316,20 @@ linux_fill_prpsinfo (struct elf_internal_linux_prpsinfo *p)
gdb_assert (p != nullptr);
+ /* The kernel fills NT_PRPSINFO from the thread group leader, even it is a
+ zombie. Do the same here, read the process state from the leader's proc
+ entries.
+
+ However, we can't read cmdline from the a zombie leader thread (it would
+ read empty). Read that one from a thread that is still alive (this is
+ what the kernel does too). It's the same value for the whole process, so
+ the choice of thread does not matter otherwise. */
+ const int leader_id = current_inferior ()->pid;
+ const ptid_t live_ptid = get_ptid_for_slash_proc ();
+
/* Obtaining PID and filename. */
- pid = inferior_ptid.pid ();
- xsnprintf (filename, sizeof (filename), "/proc/%d/cmdline", (int) pid);
+ xsnprintf (filename, sizeof (filename), "/proc/%ld/cmdline",
+ live_ptid.lwp ());
/* The full name of the program which generated the corefile. */
gdb_byte *buf = nullptr;
LONGEST buf_len = target_fileio_read_alloc (nullptr, filename, &buf);
@@ -2309,7 +2352,7 @@ linux_fill_prpsinfo (struct elf_internal_linux_prpsinfo *p)
memset (p, 0, sizeof (*p));
/* Defining the PID. */
- p->pr_pid = pid;
+ p->pr_pid = leader_id;
/* Copying the program name. Only the basename matters. */
basename = lbasename (fname.get ());
@@ -2326,7 +2369,7 @@ linux_fill_prpsinfo (struct elf_internal_linux_prpsinfo *p)
strncpy (p->pr_psargs, psargs.c_str (), sizeof (p->pr_psargs) - 1);
p->pr_psargs[sizeof (p->pr_psargs) - 1] = '\0';
- xsnprintf (filename, sizeof (filename), "/proc/%d/stat", (int) pid);
+ xsnprintf (filename, sizeof (filename), "/proc/%d/stat", leader_id);
/* The contents of `/proc/PID/stat'. */
gdb::unique_xmalloc_ptr<char> proc_stat_contents
= target_fileio_read_stralloc (NULL, filename);
@@ -2404,7 +2447,7 @@ linux_fill_prpsinfo (struct elf_internal_linux_prpsinfo *p)
/* Finally, obtaining the UID and GID. For that, we read and parse the
contents of the `/proc/PID/status' file. */
- xsnprintf (filename, sizeof (filename), "/proc/%d/status", (int) pid);
+ xsnprintf (filename, sizeof (filename), "/proc/%d/status", leader_id);
/* The contents of `/proc/PID/status'. */
gdb::unique_xmalloc_ptr<char> proc_status_contents
= target_fileio_read_stralloc (NULL, filename);
@@ -3209,9 +3252,8 @@ linux_address_in_shadow_stack_mem_range
if (!target_has_execution () || current_inferior ()->fake_pid_p)
return false;
- const int pid = current_inferior ()->pid;
-
- std::string smaps_file = string_printf ("/proc/%d/smaps", pid);
+ ptid_t ptid = get_ptid_for_slash_proc ();
+ std::string smaps_file = string_printf ("/proc/%ld/smaps", ptid.lwp ());
gdb::unique_xmalloc_ptr<char> data
= target_fileio_read_stralloc (nullptr, smaps_file.c_str ());
diff --git a/gdb/testsuite/gdb.base/info-proc.exp b/gdb/testsuite/gdb.base/info-proc.exp
index 89eed2ac815a..5310341264eb 100644
--- a/gdb/testsuite/gdb.base/info-proc.exp
+++ b/gdb/testsuite/gdb.base/info-proc.exp
@@ -63,7 +63,8 @@ if {![runto_main]} {
return
}
-gdb_test "info proc" "process ${decimal}.*" "info proc with process"
+gdb_test "info proc" "Reading /proc for process ${decimal}.*" \
+ "info proc with process"
gdb_test "info proc mapping" \
".*Mapped address spaces:.*${hex}${ws}${hex}${ws}${hex}${ws}${hex}.*"
diff --git a/gdb/testsuite/gdb.threads/gcore-stale-thread.exp b/gdb/testsuite/gdb.threads/gcore-stale-thread.exp
deleted file mode 100644
index d0465c87a4ad..000000000000
--- a/gdb/testsuite/gdb.threads/gcore-stale-thread.exp
+++ /dev/null
@@ -1,57 +0,0 @@
-# Copyright 2014-2026 Free Software Foundation, Inc.
-
-# This program is free software; you can redistribute it and/or modify
-# it under the terms of the GNU General Public License as published by
-# the Free Software Foundation; either version 3 of the License, or
-# (at your option) any later version.
-#
-# This program is distributed in the hope that it will be useful,
-# but WITHOUT ANY WARRANTY; without even the implied warranty of
-# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
-# GNU General Public License for more details.
-#
-# You should have received a copy of the GNU General Public License
-# along with this program. If not, see <http://www.gnu.org/licenses/>.
-
-standard_testfile
-set corefile [standard_output_file ${testfile}.core]
-
-if {[gdb_compile_pthreads "${srcdir}/${subdir}/${srcfile}" "${binfile}" executable debug] != ""} {
- return
-}
-
-clean_restart ${testfile}
-
-gdb_test_no_output "set non-stop on"
-
-if {![runto_main]} {
- return
-}
-
-gdb_test_multiple "info threads" "threads are supported" {
- -re ".* main .*\r\n$gdb_prompt $" {
- # OK, threads are supported.
- }
- -re "\r\n$gdb_prompt $" {
- unsupported "gdb does not support threads on this target"
- return
- }
-}
-
-gdb_breakpoint ${srcfile}:[gdb_get_line_number "break-here"]
-# gdb_continue_to_breakpoint does not work as it uses "$gdb_prompt $" regex
-# which does not work due to the output: (gdb) [Thread ... exited]
-set name "continue to breakpoint: break-here"
-gdb_test_multiple "continue" $name {
- -re "Breakpoint .* (at|in) .* break-here .*\r\n$gdb_prompt " {
- pass $name
- }
-}
-
-gdb_gcore_cmd "$corefile" "save a corefile"
-
-# Do not run "info threads" before "gcore" as it could workaround the bug
-# by discarding non-current exited threads.
-gdb_test "info threads" \
- {The current thread <Thread ID 1> has terminated\. See `help thread'\.} \
- "exited thread is current due to non-stop"
diff --git a/gdb/testsuite/gdb.threads/gcore-stale-thread.c b/gdb/testsuite/gdb.threads/thread-leader-exited.c
similarity index 64%
rename from gdb/testsuite/gdb.threads/gcore-stale-thread.c
rename to gdb/testsuite/gdb.threads/thread-leader-exited.c
index 9dbeeca5e0e1..998ff4d41d60 100644
--- a/gdb/testsuite/gdb.threads/gcore-stale-thread.c
+++ b/gdb/testsuite/gdb.threads/thread-leader-exited.c
@@ -19,29 +19,49 @@
#include <assert.h>
static pthread_t main_thread;
+static pthread_barrier_t barrier;
+
+/* First worker. */
static void *
start (void *arg)
{
- int i;
-
- i = pthread_join (main_thread, NULL);
+ int i = pthread_join (main_thread, NULL);
assert (i == 0);
+ i = pthread_barrier_wait (&barrier);
+ assert (i == 0 || i == PTHREAD_BARRIER_SERIAL_THREAD);
+
return arg; /* break-here */
}
+/* Second worker. */
+
+static void *
+start_second (void *arg)
+{
+ int i = pthread_barrier_wait (&barrier);
+ assert (i == 0 || i == PTHREAD_BARRIER_SERIAL_THREAD);
+
+ return arg; /* break-here-second */
+}
+
int
main (void)
{
- pthread_t thread;
- int i;
+ pthread_t thread, thread_second;
main_thread = pthread_self ();
+ int i = pthread_barrier_init (&barrier, NULL, 2);
+ assert (i == 0);
+
i = pthread_create (&thread, NULL, start, NULL);
assert (i == 0);
+ i = pthread_create (&thread_second, NULL, start_second, NULL);
+ assert (i == 0);
+
pthread_exit (NULL);
assert (0);
return 0;
diff --git a/gdb/testsuite/gdb.threads/thread-leader-exited.exp b/gdb/testsuite/gdb.threads/thread-leader-exited.exp
new file mode 100644
index 000000000000..1c493c69242e
--- /dev/null
+++ b/gdb/testsuite/gdb.threads/thread-leader-exited.exp
@@ -0,0 +1,122 @@
+# Copyright 2014-2026 Free Software Foundation, Inc.
+
+# This program is free software; you can redistribute it and/or modify
+# it under the terms of the GNU General Public License as published by
+# the Free Software Foundation; either version 3 of the License, or
+# (at your option) any later version.
+#
+# This program is distributed in the hope that it will be useful,
+# but WITHOUT ANY WARRANTY; without even the implied warranty of
+# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
+# GNU General Public License for more details.
+#
+# You should have received a copy of the GNU General Public License
+# along with this program. If not, see <http://www.gnu.org/licenses/>.
+
+# Check that some things work (like "info proc" and generating core files) when
+# the thread group leader has exited.
+
+standard_testfile
+set corefile [standard_output_file ${testfile}.core]
+
+if {[gdb_compile_pthreads "${srcdir}/${subdir}/${srcfile}" "${binfile}" executable debug] != ""} {
+ return
+}
+
+save_vars { GDBFLAGS } {
+ append GDBFLAGS " -ex \"set non-stop on\""
+ clean_restart ${testfile}
+}
+
+if {![runto_main]} {
+ return
+}
+
+gdb_test_multiple "info threads" "threads are supported" {
+ -re ".* main .*\r\n$gdb_prompt $" {
+ # OK, threads are supported.
+ }
+ -re "\r\n$gdb_prompt $" {
+ unsupported "gdb does not support threads on this target"
+ return
+ }
+}
+
+gdb_breakpoint ${srcfile}:[gdb_get_line_number "break-here"]
+gdb_breakpoint ${srcfile}:[gdb_get_line_number "break-here-second"]
+
+gdb_test -no-prompt-anchor "continue &" "Continuing\\."
+
+# Wait for both workers to stop and the leader to exit.
+set num_events 0
+gdb_test_multiple "" "wait for the workers to stop" -lbl {
+ -re "break-here \\*/|break-here-second \\*/|\\(id 1\\) exited\\]" {
+ incr num_events
+ if { $num_events < 3 } {
+ exp_continue
+ }
+ }
+}
+
+gdb_assert { $num_events == 3 } "both workers stopped, leader exited"
+
+# Check that "info proc" reads /proc through a thread that is still alive. The
+# format is Linux-specific.
+if { [istarget "*-*-linux*"] } {
+ set binfile_re [string_to_regexp $binfile]
+ gdb_test "info proc" \
+ [multi_line \
+ "Reading /proc for process $decimal \\(LWP $decimal\\)" \
+ "cmdline = '$binfile_re'" \
+ "cwd = '.+'" \
+ "exe = '$binfile_re'"]
+}
+
+gdb_gcore_cmd "$corefile" "save a corefile"
+
+# Do not run "info threads" before "gcore" as it could workaround the bug
+# by discarding non-current exited threads.
+gdb_test "info threads" \
+ {The current thread <Thread ID 1> has terminated\. See `help thread'\.} \
+ "exited thread is current due to non-stop"
+
+# Verify that "info proc stat" reads the /proc entry for the selected thread.
+proc check_info_proc_stat { thread_id } {
+ # Capture the thread LWP id while switching thread. The message is
+ # different depending on native vs remote.
+ set lwp ""
+ gdb_test_multiple "thread $thread_id" "switch to thread $thread_id" {
+ -re -wrap "Switching to thread $thread_id \\(Thread \[^\r\n\]*\\(LWP ($::decimal)\\)\\).*" {
+ set lwp $expect_out(1,string)
+ pass $gdb_test_name
+ }
+ -re -wrap "Switching to thread $thread_id \\(Thread $::decimal\\.($::decimal)\\).*" {
+ set lwp $expect_out(1,string)
+ pass $gdb_test_name
+ }
+ }
+
+ if { $lwp eq "" } {
+ error "unable to capture LWP id for thread $thread_id"
+ return
+ }
+
+ gdb_test "info proc stat" \
+ [multi_line \
+ "Reading /proc for process $::decimal \\(LWP $lwp\\)" \
+ "Process: $lwp" \
+ ".*"] \
+ "info proc stat for thread $thread_id"
+}
+
+if { [istarget "*-*-linux*"] } {
+ check_info_proc_stat 2
+ check_info_proc_stat 3
+}
+
+# Load the core back to verify that GDB managed to fill the NT_PRPSINFO note.
+# A user-visible consequence of not having this note is that GDB doesn't print
+# the "Core was generated by..." line. gdb_core_cmd needs to see this line in
+# order to emit a pass, and it emits a fail in all other cases.
+clean_restart ${testfile}
+gdb_core_cmd $corefile "reload the corefile"
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior
2026-09-17 20:16 [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Simon Marchi
2026-09-17 20:16 ` [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files Simon Marchi
@ 2026-09-21 10:22 ` Matthieu Longo
2026-09-21 13:46 ` Simon Marchi
[not found] ` <20260917201755.589524-2-simon.marchi@efficios.com_5b6b04a9_release>
2 siblings, 1 reply; 6+ messages in thread
From: Matthieu Longo @ 2026-09-21 10:22 UTC (permalink / raw)
To: Simon Marchi; +Cc: gdb-patches
On 17/09/2026 21:16, Simon Marchi wrote:
> any_non_exited_thread_of_inferior prefers to return the selected thread:
>
> /* Prefer the current thread, if there's one. */
> if (inf == current_inferior () && inferior_ptid != null_ptid)
> return inferior_thread ();
>
> ... but the selected thread can be an exited thread. This can happen in
> non-stop when the selected thread exits, in which case it is kept around
> but marked as exited (see thread_info::deletable).
>
> Check the current thread's state before returning it, and fall through
> to the existing loop otherwise.
>
> Note that we also have any_live_thread_of_inferior, which is similar but
> differs in subtle ways. We should probably check if we can merge them
> at some point.
>
> Change-Id: I620ee200037c6a3fd3d59b96f16750d6b962c724
For my curiosity, what is this Change-Id for ?
I saw several ones in the git history, but I have no idea why it is there.
Matthieu
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior
2026-09-21 10:22 ` [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Matthieu Longo
@ 2026-09-21 13:46 ` Simon Marchi
0 siblings, 0 replies; 6+ messages in thread
From: Simon Marchi @ 2026-09-21 13:46 UTC (permalink / raw)
To: Matthieu Longo; +Cc: gdb-patches
On 9/21/26 6:22 AM, Matthieu Longo wrote:
> On 17/09/2026 21:16, Simon Marchi wrote:
>> any_non_exited_thread_of_inferior prefers to return the selected thread:
>>
>> /* Prefer the current thread, if there's one. */
>> if (inf == current_inferior () && inferior_ptid != null_ptid)
>> return inferior_thread ();
>>
>> ... but the selected thread can be an exited thread. This can happen in
>> non-stop when the selected thread exits, in which case it is kept around
>> but marked as exited (see thread_info::deletable).
>>
>> Check the current thread's state before returning it, and fall through
>> to the existing loop otherwise.
>>
>> Note that we also have any_live_thread_of_inferior, which is similar but
>> differs in subtle ways. We should probably check if we can merge them
>> at some point.
>>
>> Change-Id: I620ee200037c6a3fd3d59b96f16750d6b962c724
>
> For my curiosity, what is this Change-Id for ?
> I saw several ones in the git history, but I have no idea why it is there.
>
> Matthieu
It's a Gerrit change id, added by the Gerrit hook. That's because I use
Gerrit to track patches when I develop them. Having the Change-Id in
the upstream commits makes it so that Gerrit auto-closes the changes
when I pull origin/master into Gerrit's master branch, so it helps me
with tracking what I have pending.
Simon
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files
[not found] ` <20260917201755.589524-2-simon.marchi@efficios.com_5b6b04a9_release>
@ 2026-09-21 14:26 ` Matthieu Longo
2026-09-21 14:57 ` Simon Marchi
0 siblings, 1 reply; 6+ messages in thread
From: Matthieu Longo @ 2026-09-21 14:26 UTC (permalink / raw)
To: Simon Marchi; +Cc: gdb-patches
On 18/09/2026 13:21, Simon Marchi wrote:
> On Linux, /proc/<pid> is keyed by the thread-group leader PID. When the
> leader has exited, some /proc/<pid>/... entries become unavailable even
> though another thread is still alive. This can happen, for instance,
> when the main thread calls pthread_exit() and another thread continues
> the execution (existing test: gcore-stale-thread, renamed below).
>
> This causes GDB to fail to read procfs entries such as cmdline, cwd,
> exe, maps, and smaps when it builds the paths from the inferior pid
> after the thread-group leader has exited.
>
> A /proc/<lwp> entry exists for every live LWP of the process, and for
> the entries listed above its contents are the same as those of
> /proc/<pid>. Fix this by building the procfs paths from the LWP id of a
> thread GDB believes is alive, rather than from the inferior pid.
>
> Add get_ptid_for_slash_proc, which returns the PTID to read from, and
> use it in places that read /proc:
>
> - linux_info_proc
> - linux_process_address_in_memtag_page
> - linux_find_memory_regions_full
> - linux_fill_prpsinfo
> - linux_address_in_shadow_stack_mem_range
>
> linux_vsyscall_range_raw also reads /proc, but I did not change it, as
> it uses a different pattern, "/proc/<pid>/task/<pid>/maps". It could in
> theory suffer from the same problem.
>
> The thread is chosen with any_non_exited_thread_of_inferior, which gives
> preference to the selected thread. I think this is actually a good
> feature. Some entries in /proc (like stat and status) show some
> thread-specific content, so this lets the user pick which thread "info
> proc" reports on. Most entries are per-process, the choice of thread
> makes no difference for them.
>
> linux_fill_prpsinfo is an exception to the above. It populates the
> NT_PRPSINFO note of a core file, which the kernel (for kernel-generated
> core files) fills in from the thread group leader, even when that leader
> is a zombie:
>
> https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1901
>
> To keep GDB-generated cores looking like kernel-generated ones, it keeps
> reading stat and status from the leader's entry, so the values taken
> from those match what the kernel would write (stat and status are still
> readable when the thread is zombie). Only the cmdline is read from a
> live LWP, since trying to read that one from a zombie leader doesn't
> work. This mirrors the kernel's fill_psinfo function, which takes the
> program name and arguments from the mm of the thread doing the dump:
>
> https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1530-L1539
>
> This also fixes a bug, in that core dumps produced while the leader is
> zombie would be missing the NT_PRPSINFO note. linux_fill_prpsinfo would
> read an empty string from cmdline and return false. A user-visible
> consequence of that is that loading back the core would not show the
> expected "Core was generated by..." line.
>
> This is still a best-effort: GDB's view of the thread list may be stale,
> so the thread selected to do the /proc accesses may actually have
> exited but GDB does not know it.
>
> Since the entry "info proc" reads is no longer necessarily that of the
> thread group leader, replace its "process <pid>" header with:
>
> Reading /proc for process <pid>
>
> or, when the LWP differs from the PID:
>
> Reading /proc for process <pid> (LWP <lwp>)
>
> Since "info proc" now stores the pid given on the command line in a
> ptid_t, whose pid field is an int, reject a value that would not survive
> the conversion, rather than printing one value and reading /proc for
> another.
>
> Update gdb.base/info-proc.exp for the new header.
>
> Rename gdb.threads/gcore-stale-thread.{exp,c} to
> gdb.threads/thread-leader-exited.{exp,c} and modify it in a few ways:
>
> - give the program a second worker thread
>
> - check that we're able to read /proc on Linux even when the leader has
> exited (and is selected)
>
> - check that "info proc stat" reports each worker thread's own LWP id
> in its "Process:" field, confirming that "info proc" prioritizes the
> selected thread
>
> - load the generated core back to verify that we see the "Core was
> generated by ..." line, and therefore the NT_PRPSINFO note was
> generated
>
> - enable non-stop using the usual save_vars / append to GDBFLAGS
> pattern, which fixes a pre-existing bug of not being able to run the
> test on the native-extended-gdbserver board.
>
> Bug: https://sourceware.org/bugzilla/show_bug.cgi?id=31207
> Change-Id: I1e6ccfae90551b8e96a5644b51c7652bbda593fa
> Co-Authored-By: Matthieu Longo <matthieu.longo@arm.com>
It looks good to me.
I ran the 2 patches on the top of master, on both AArch64 and x86_64 on our internal CI and I found
2 failing tests on AArch64: step-over-process-exit.exp and missing-thread.exp.
I re-ran them manually and both of them are passing.
This has nothing to do with your patch, but are those tests known for being flickering ?
This is the logs for step-over-process-exit.exp:
```
maint show target-non-stop
Whether the target is always in non-stop mode is auto (currently on).
(gdb) next
[Thread 0xfffff7fc0020 (LWP 291431) (id 1) exited]
[New LWP 291431 (id 3)]
[Thread 0xfffff7e0f120 (LWP 292049) (id 2) exited]
warning: error removing breakpoint 0 at 0xaaaaaaaa087c
Command aborted, thread exited.
Cannot remove breakpoints because program is no longer writable.
Further execution is probably impossible.
(gdb) [Inferior 1 (process 291431) exited normally]
FAIL: gdb.threads/step-over-process-exit.exp: which=other: next (timeout)
```
VS when it passed:
```
maint show target-non-stop
Whether the target is always in non-stop mode is auto (currently on).
(gdb) next
[LWP 1540 (id 2) exited]
[Inferior 1 (process 1539) exited normally]
(gdb) PASS: gdb.threads/step-over-process-exit.exp: which=other: next
```
It looks like there is an issue while removing breakpoints. At first glance, I would say that there
is a race condition in the test when thread 1 exist before 2.
And this, somehow creates a time-out ?
Regarding logs for missing-thread.exp:
```
continue
Continuing.
[New Thread 217413.218469 (id 2)]
[Thread 217413.218469 (id 2) exited]
warning: command aborted, Thread 217413.218469 unexpectedly exited after signal stop event
Remote communication error. Target disconnected: error while reading: Connection reset by peer.
(gdb) FAIL: gdb.replay/missing-thread.exp: non_stop=on: missing 1 thread log: replay_with_log: continue
```
vs when it passed:
```
continue
Continuing.
[New Thread 1692.1693 (id 2)]
Thread 2 "missing-thread" received signal SIGTRAP, Trace/breakpoint trap.
0x0000fffff7eb6ce0 in clock_nanosleep () from /lib/aarch64-linux-gnu/libc.so.6
(gdb) PASS: gdb.replay/missing-thread.exp: non_stop=on: with unmodified log: replay_with_log: continue
```
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files
2026-09-21 14:26 ` [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files Matthieu Longo
@ 2026-09-21 14:57 ` Simon Marchi
0 siblings, 0 replies; 6+ messages in thread
From: Simon Marchi @ 2026-09-21 14:57 UTC (permalink / raw)
To: Matthieu Longo; +Cc: gdb-patches
On 9/21/26 10:26 AM, Matthieu Longo wrote:
> On 18/09/2026 13:21, Simon Marchi wrote:
>> On Linux, /proc/<pid> is keyed by the thread-group leader PID. When the
>> leader has exited, some /proc/<pid>/... entries become unavailable even
>> though another thread is still alive. This can happen, for instance,
>> when the main thread calls pthread_exit() and another thread continues
>> the execution (existing test: gcore-stale-thread, renamed below).
>>
>> This causes GDB to fail to read procfs entries such as cmdline, cwd,
>> exe, maps, and smaps when it builds the paths from the inferior pid
>> after the thread-group leader has exited.
>>
>> A /proc/<lwp> entry exists for every live LWP of the process, and for
>> the entries listed above its contents are the same as those of
>> /proc/<pid>. Fix this by building the procfs paths from the LWP id of a
>> thread GDB believes is alive, rather than from the inferior pid.
>>
>> Add get_ptid_for_slash_proc, which returns the PTID to read from, and
>> use it in places that read /proc:
>>
>> - linux_info_proc
>> - linux_process_address_in_memtag_page
>> - linux_find_memory_regions_full
>> - linux_fill_prpsinfo
>> - linux_address_in_shadow_stack_mem_range
>>
>> linux_vsyscall_range_raw also reads /proc, but I did not change it, as
>> it uses a different pattern, "/proc/<pid>/task/<pid>/maps". It could in
>> theory suffer from the same problem.
>>
>> The thread is chosen with any_non_exited_thread_of_inferior, which gives
>> preference to the selected thread. I think this is actually a good
>> feature. Some entries in /proc (like stat and status) show some
>> thread-specific content, so this lets the user pick which thread "info
>> proc" reports on. Most entries are per-process, the choice of thread
>> makes no difference for them.
>>
>> linux_fill_prpsinfo is an exception to the above. It populates the
>> NT_PRPSINFO note of a core file, which the kernel (for kernel-generated
>> core files) fills in from the thread group leader, even when that leader
>> is a zombie:
>>
>> https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1901
>>
>> To keep GDB-generated cores looking like kernel-generated ones, it keeps
>> reading stat and status from the leader's entry, so the values taken
>> from those match what the kernel would write (stat and status are still
>> readable when the thread is zombie). Only the cmdline is read from a
>> live LWP, since trying to read that one from a zombie leader doesn't
>> work. This mirrors the kernel's fill_psinfo function, which takes the
>> program name and arguments from the mm of the thread doing the dump:
>>
>> https://elixir.bootlin.com/linux/v7.2.5/source/fs/binfmt_elf.c#L1530-L1539
>>
>> This also fixes a bug, in that core dumps produced while the leader is
>> zombie would be missing the NT_PRPSINFO note. linux_fill_prpsinfo would
>> read an empty string from cmdline and return false. A user-visible
>> consequence of that is that loading back the core would not show the
>> expected "Core was generated by..." line.
>>
>> This is still a best-effort: GDB's view of the thread list may be stale,
>> so the thread selected to do the /proc accesses may actually have
>> exited but GDB does not know it.
>>
>> Since the entry "info proc" reads is no longer necessarily that of the
>> thread group leader, replace its "process <pid>" header with:
>>
>> Reading /proc for process <pid>
>>
>> or, when the LWP differs from the PID:
>>
>> Reading /proc for process <pid> (LWP <lwp>)
>>
>> Since "info proc" now stores the pid given on the command line in a
>> ptid_t, whose pid field is an int, reject a value that would not survive
>> the conversion, rather than printing one value and reading /proc for
>> another.
>>
>> Update gdb.base/info-proc.exp for the new header.
>>
>> Rename gdb.threads/gcore-stale-thread.{exp,c} to
>> gdb.threads/thread-leader-exited.{exp,c} and modify it in a few ways:
>>
>> - give the program a second worker thread
>>
>> - check that we're able to read /proc on Linux even when the leader has
>> exited (and is selected)
>>
>> - check that "info proc stat" reports each worker thread's own LWP id
>> in its "Process:" field, confirming that "info proc" prioritizes the
>> selected thread
>>
>> - load the generated core back to verify that we see the "Core was
>> generated by ..." line, and therefore the NT_PRPSINFO note was
>> generated
>>
>> - enable non-stop using the usual save_vars / append to GDBFLAGS
>> pattern, which fixes a pre-existing bug of not being able to run the
>> test on the native-extended-gdbserver board.
>>
>> Bug: https://sourceware.org/bugzilla/show_bug.cgi?id=31207
>> Change-Id: I1e6ccfae90551b8e96a5644b51c7652bbda593fa
>> Co-Authored-By: Matthieu Longo <matthieu.longo@arm.com>
> It looks good to me.
>
> I ran the 2 patches on the top of master, on both AArch64 and x86_64 on our internal CI and I found
> 2 failing tests on AArch64: step-over-process-exit.exp and missing-thread.exp.
> I re-ran them manually and both of them are passing.
> This has nothing to do with your patch, but are those tests known for being flickering ?
>
> This is the logs for step-over-process-exit.exp:
> ```
> maint show target-non-stop
> Whether the target is always in non-stop mode is auto (currently on).
> (gdb) next
> [Thread 0xfffff7fc0020 (LWP 291431) (id 1) exited]
> [New LWP 291431 (id 3)]
> [Thread 0xfffff7e0f120 (LWP 292049) (id 2) exited]
> warning: error removing breakpoint 0 at 0xaaaaaaaa087c
> Command aborted, thread exited.
> Cannot remove breakpoints because program is no longer writable.
> Further execution is probably impossible.
> (gdb) [Inferior 1 (process 291431) exited normally]
> FAIL: gdb.threads/step-over-process-exit.exp: which=other: next (timeout)
This one looks odd, you get a notification that LWP 291431 exits, then a
notification that there is a "new" LWP 291431. It sounds like maybe we
processed an exit event for that thread, followed by another event, that
made GDB think there was a new thread.
> ```
> VS when it passed:
> ```
> maint show target-non-stop
> Whether the target is always in non-stop mode is auto (currently on).
> (gdb) next
> [LWP 1540 (id 2) exited]
> [Inferior 1 (process 1539) exited normally]
> (gdb) PASS: gdb.threads/step-over-process-exit.exp: which=other: next
> ```
> It looks like there is an issue while removing breakpoints. At first glance, I would say that there
> is a race condition in the test when thread 1 exist before 2.
> And this, somehow creates a time-out ?
>
> Regarding logs for missing-thread.exp:
> ```
> continue
> Continuing.
> [New Thread 217413.218469 (id 2)]
> [Thread 217413.218469 (id 2) exited]
> warning: command aborted, Thread 217413.218469 unexpectedly exited after signal stop event
> Remote communication error. Target disconnected: error while reading: Connection reset by peer.
> (gdb) FAIL: gdb.replay/missing-thread.exp: non_stop=on: missing 1 thread log: replay_with_log: continue
> ```
> vs when it passed:
> ```
> continue
> Continuing.
> [New Thread 1692.1693 (id 2)]
>
> Thread 2 "missing-thread" received signal SIGTRAP, Trace/breakpoint trap.
> 0x0000fffff7eb6ce0 in clock_nanosleep () from /lib/aarch64-linux-gnu/libc.so.6
> (gdb) PASS: gdb.replay/missing-thread.exp: non_stop=on: with unmodified log: replay_with_log: continue
> ```
The message in the failure in this one sounds like the remote side
(gdbreplay) suddenly exiting. Not sure if it's just exiting cleanly
because it's done, or if it's crashing.
Some threads tests are known to be racy, ideally it would take someone
to invest time in each of them to understand and fix the race
conditions.
Could you please open some bugs to document those failures in Bugzilla
(if there aren't already)?
Simon
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-09-21 15:06 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-17 20:16 [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Simon Marchi
2026-09-17 20:16 ` [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files Simon Marchi
2026-09-21 10:22 ` [PATCH 1/2] gdb: don't return an exited thread from any_non_exited_thread_of_inferior Matthieu Longo
2026-09-21 13:46 ` Simon Marchi
[not found] ` <20260917201755.589524-2-simon.marchi@efficios.com_5b6b04a9_release>
2026-09-21 14:26 ` [PATCH 2/2] gdb: use the selected thread's LWP ID when reading Linux procfs files Matthieu Longo
2026-09-21 14:57 ` Simon Marchi
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox