Page MenuHomeFreeBSD

D59365.diff
No OneTemporary

D59365.diff

diff --git a/lib/libpmc/pmc.3 b/lib/libpmc/pmc.3
--- a/lib/libpmc/pmc.3
+++ b/lib/libpmc/pmc.3
@@ -384,6 +384,17 @@
.It Fn pmc_set
Set the reload value for a sampling PMC.
.El
+.It "PMC Grouping"
+.Bl -tag -width 6n -compact
+.It Fn pmc_allocate_group
+Allocate a PMC that will join a group.
+.It Fn pmc_group_create , Fn pmc_group_add , Fn pmc_group_commit
+Create a group, add members to it, and freeze its membership.
+.It Fn pmc_group_read
+Read every member of a committed group in one consistent snapshot.
+.It Fn pmc_read_pair
+Read one PMC together with the times that scale it.
+.El
.It "Queries"
.Bl -tag -width 6n -compact
.It Fn pmc_capabilities
@@ -414,6 +425,96 @@
instruction to directly read the contents of the PMC.
.El
.El
+.Ss PMC Groups
+Events measured separately are measured at different times, so their
+ratios are not trustworthy.
+A group is a set of PMCs the hardware programs together: every member
+counts over exactly the same interval, or none of them does.
+.Pp
+Members are allocated with
+.Fn pmc_allocate_group ,
+which defers the choice of hardware counter, then collected with
+.Fn pmc_group_create
+and
+.Fn pmc_group_add
+and frozen by
+.Fn pmc_group_commit .
+One member is designated the leader, and it stands for the group
+thereafter:
+.Fn pmc_attach ,
+.Fn pmc_start ,
+.Fn pmc_stop
+and
+.Fn pmc_release
+act on the whole group through it, and the same call on another member
+fails with
+.Er ENOTTY .
+.Pp
+When more events are requested than the hardware can hold at once, a
+group may be multiplexed: it holds counters for part of the time and is
+evicted for the rest.
+.Fn pmc_group_read
+therefore returns, alongside the member values, the time the group was
+enabled and the time it was actually running.
+A caller scales a raw count by
+.Va enabled Ns / Ns Va running
+to estimate what the count would have been with the group resident
+throughout.
+Such a value is an estimate, and should be presented as one.
+Sample counts must never be scaled: a multiplexed sampling member simply
+does not observe what happens while it is evicted.
+.Ss Probing For Group Support
+There is no operation that asks whether grouping is available.
+A program probes by attempting the allocation it intended to perform
+anyway, and treating
+.Er EOPNOTSUPP
+from
+.Fn pmc_allocate_group
+as
+.Dq this kernel cannot group these events ,
+falling back to ungrouped allocation.
+A throwaway probe is a poor idea, since each one consumes a deferred
+handle and a group slot.
+.Pp
+An
+.Er EINVAL
+from
+.Fn pmc_group_commit
+is never a probe.
+It reports a usage error with several possible causes \(em no leader,
+two leaders, mixed virtual and system modes, system members bound to
+different CPUs, or a leader-only flag on another member \(em and should
+be shown to the user rather than treated as absent support.
+.Ss Deferred Handle Budget
+A grouped PMC keeps a stable identifier from allocation until release,
+even as multiplexing moves it between hardware counters.
+These identifiers come from a pool of 32 per owning process per CPU.
+.Pp
+Process-scope groups are not bound to a CPU, so all of an owner's
+process-scope members draw from a single pool of 32, however many groups
+they belong to.
+One maximum-size group therefore exhausts the entire process-scope
+budget, and a further allocation fails with
+.Er EMFILE
+until that group is released.
+System-scope groups draw from a separate pool for each CPU they are
+bound to.
+.Ss Reading A Group
+.Fn pmc_group_read
+fills a caller-supplied array.
+Pass a count of zero to ask how many members the group has; the call
+succeeds and reports the count without touching the array.
+Passing a nonzero count smaller than the membership fails with
+.Er E2BIG ,
+and again reports the required count, so a caller can size its array in
+one retry.
+.Pp
+.Fn pmc_read_pair
+reads a single PMC with the times that scale it.
+It allocates internally, so it may fail for reasons unrelated to the
+PMC, and it applies only to a group leader:
+on any other member it fails with
+.Er EOPNOTSUPP .
.Ss Signal Handling Requirements
Applications using PMCs are required to handle the following signals:
.Bl -tag -width ".Dv SIGBUS"
diff --git a/lib/libpmc/pmclog.3 b/lib/libpmc/pmclog.3
--- a/lib/libpmc/pmclog.3
+++ b/lib/libpmc/pmclog.3
@@ -21,7 +21,7 @@
.\" OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
.\" SUCH DAMAGE.
.\"
-.Dd March 26, 2006
+.Dd August 22, 2026
.Dt PMCLOG 3
.Os
.Sh NAME
@@ -93,6 +93,7 @@
struct pmclog_ev_pmcallocatedyn pl_ad;
struct pmclog_ev_pmcattach pl_t;
struct pmclog_ev_pmcdetach pl_d;
+ struct pmclog_ev_pmcgroupinheritmiss pl_gim;
struct pmclog_ev_proccsw pl_c;
struct pmclog_ev_procexec pl_x;
struct pmclog_ev_procexit pl_e;
@@ -201,6 +202,9 @@
A record describing a PMC attach operation.
.It Dv PMCLOG_TYPE_PMCDETACH
A record describing a PMC detach operation.
+.It Dv PMCLOG_TYPE_PMCGROUPINHERITMISS
+This record shows a failed group inheritance.
+The record gives the group-leader PMC and the child PID that missed it.
.It Dv PMCLOG_TYPE_PROCCSW
A record describing a PMC reading at the time of a process context switch.
.It Dv PMCLOG_TYPE_PROCEXEC
@@ -222,6 +226,14 @@
A record containing user data.
.El
.Pp
+PMC-specific identifier fields hold the stable handle from PMC allocation.
+For a grouped PMC, multiplexing can move the PMC between hardware rows.
+The handle stays the same in this case.
+Do not read the handle as the current hardware row.
+A record not tied to one PMC may use a record-defined sentinel value
+such as
+.Dv PMC_ID_INVALID .
+.Pp
Function
.Fn pmclog_feed
is used with parsers configured to parse memory based event streams.
diff --git a/usr.sbin/pmcstat/pmcstat.8 b/usr.sbin/pmcstat/pmcstat.8
--- a/usr.sbin/pmcstat/pmcstat.8
+++ b/usr.sbin/pmcstat/pmcstat.8
@@ -248,6 +248,15 @@
Allocate a system mode sampling PMC measuring hardware events
specified in
.Ar event-spec .
+A system mode PMC is bound to a single CPU: the lowest-numbered one in
+the set given by
+.Fl c ,
+or in the process's root CPU set if
+.Fl c
+was not given.
+To profile several CPUs, run one
+.Nm
+per CPU.
.It Fl T
Use a
.Xr top 1 Ns -like
@@ -309,6 +318,83 @@
saved with the
.Fl O
option.
+.It Fl b
+Enable PMU event grouping.
+When this flag is set, an argument to
+.Fl p ,
+.Fl P ,
+.Fl s
+or
+.Fl S
+may be a brace-list of the form
+.Sq Brq Ar ev1 , Ar ev2 , Ar ev3 :
+all listed events are programmed atomically as a single PMU group, so
+the counter values for siblings are read at consistent program points.
+Every brace-mode group leader is allocated with
+.Dv PMC_F_GROUP_MUX
+set, which has no effect on groups that fit in hardware but lets the
+kernel transparently time-multiplex any group that overflows the
+.Em remaining
+hardware counter budget.
+For example, on a 6-counter Zen3 the invocation
+.Bd -literal -offset indent
+pmcstat -b -p '{instructions,unhalted-cycles}' \\
+ -p '{ls_alloc_mab_count,ls_not_halted_cyc,ls_dispatch.all}' \\
+ -p '{ls_smi_rx,ls_int_taken,instructions}' sleep 60
+.Ed
+.Pp
+attaches three groups (2+3+3 = 8 events) to one target.
+The first two groups consume 5 of 6 counters, leaving 1.
+The third group is then committed in multiplex mode so its 3 events
+rotate through the remaining slot rather than failing with
+.Er ENOSPC .
+.Pp
+While a group is being multiplexed its counting-mode columns show
+.Em scaled estimates
+(interval count scaled by the group's enabled/running time ratio),
+and the group's hardware residency for the interval is appended in a
+trailing
+.Ql res%
+column that stays blank while the group is fully resident:
+.Bd -literal -offset indent
+# p/instructions p/unhalted-cycles p/branches res%
+ 24000210 51003001 92010 50.2%
+ <not counted> <not counted> <not counted> 0.0%
+ 23890110 50700323 91500
+.Ed
+.Pp
+An interval in which the group never reached hardware prints
+.Ql <not counted>
+for every member; when residency falls below 1/10
+.Pq Dv PMC_SCALE_MAX
+the raw interval count is shown instead of an extrapolated estimate.
+With
+.Fl C ,
+cumulative values are running sums of the per-interval estimates.
+.Pp
+A scaled value is an estimate of what the count would have been had the
+group stayed on hardware, obtained by assuming the events the group did
+not see resemble the ones it did.
+Over a short interval that assumption is weak: a group that was resident
+for one rotation window in three has sampled the interval three times,
+and the resulting estimate is correspondingly rough.
+Choose a display interval spanning at least about twenty rotation
+periods \(em with the default
+.Va kern.hwpmc.mux_period_ms
+of 50, roughly one second \(em before reading much into a scaled number.
+.Pp
+For system-wide sampling, multiplexing costs coverage rather than
+accuracy.
+A CPU whose group is evicted produces no samples at all for that window,
+and nothing reconstructs them: the profile is of the time the group was
+resident, not of the whole run.
+Sample counts are therefore never scaled.
+At the end of a run,
+.Nm
+sends this warning if a multiplexed group has a sampling period that is
+longer than the event budget of one rotation window.
+.Nm
+does not send this warning for a fully resident group.
.It Fl c Ar cpu-spec
Set the cpus for subsequent system mode PMCs specified on the
command line to
@@ -326,11 +412,32 @@
the target process.
The default is to measure events for the target process alone.
(it has to be passed in the command line prior to
-.Fl p ,
-.Fl s ,
-.Fl P ,
+.Fl p
or
-.Fl S ) .
+.Fl P ) .
+.Pp
+A system mode PMC is bound to a CPU and has no process target to follow
+through a fork, so combining
+.Fl d
+with
+.Fl s
+or
+.Fl S
+is rejected.
+.Pp
+With
+.Fl b ,
+a brace-list group follows a fork as a whole: every member attaches to
+the child, or none does.
+If the kernel cannot attach the group to a new child \(em a fork cannot
+be failed, so under memory pressure the child is simply not measured \(em
+its work is missing from the totals, which otherwise look complete.
+.Nm
+reports that at the end of the run:
+.Bd -literal -offset indent
+pmcstat: WARNING: 2 descendant attaches failed; their contributions
+are not in these totals.
+.Ed
.It Fl e
Specify that the gprof profile files will use a wide history counter.
These files are produced in a format compatible with

File Metadata

Mime Type
text/plain
Expires
Fri, Oct 2, 6:48 AM (2 h, 35 m)
Storage Engine
blob
Storage Format
Raw Data
Storage Handle
40066447
Default Alt Text
D59365.diff (10 KB)

Event Timeline