Page Menu
Home
FreeBSD
Search
Configure Global Search
Log In
Files
F174192818
D59365.diff
No One
Temporary
Actions
View File
Edit File
Delete File
View Transforms
Subscribe
Mute Notifications
Flag For Later
Award Token
Size
10 KB
Referenced Files
None
Subscribers
None
D59365.diff
View Options
diff --git a/lib/libpmc/pmc.3 b/lib/libpmc/pmc.3
--- a/lib/libpmc/pmc.3
+++ b/lib/libpmc/pmc.3
@@ -384,6 +384,17 @@
.It Fn pmc_set
Set the reload value for a sampling PMC.
.El
+.It "PMC Grouping"
+.Bl -tag -width 6n -compact
+.It Fn pmc_allocate_group
+Allocate a PMC that will join a group.
+.It Fn pmc_group_create , Fn pmc_group_add , Fn pmc_group_commit
+Create a group, add members to it, and freeze its membership.
+.It Fn pmc_group_read
+Read every member of a committed group in one consistent snapshot.
+.It Fn pmc_read_pair
+Read one PMC together with the times that scale it.
+.El
.It "Queries"
.Bl -tag -width 6n -compact
.It Fn pmc_capabilities
@@ -414,6 +425,96 @@
instruction to directly read the contents of the PMC.
.El
.El
+.Ss PMC Groups
+Events measured separately are measured at different times, so their
+ratios are not trustworthy.
+A group is a set of PMCs the hardware programs together: every member
+counts over exactly the same interval, or none of them does.
+.Pp
+Members are allocated with
+.Fn pmc_allocate_group ,
+which defers the choice of hardware counter, then collected with
+.Fn pmc_group_create
+and
+.Fn pmc_group_add
+and frozen by
+.Fn pmc_group_commit .
+One member is designated the leader, and it stands for the group
+thereafter:
+.Fn pmc_attach ,
+.Fn pmc_start ,
+.Fn pmc_stop
+and
+.Fn pmc_release
+act on the whole group through it, and the same call on another member
+fails with
+.Er ENOTTY .
+.Pp
+When more events are requested than the hardware can hold at once, a
+group may be multiplexed: it holds counters for part of the time and is
+evicted for the rest.
+.Fn pmc_group_read
+therefore returns, alongside the member values, the time the group was
+enabled and the time it was actually running.
+A caller scales a raw count by
+.Va enabled Ns / Ns Va running
+to estimate what the count would have been with the group resident
+throughout.
+Such a value is an estimate, and should be presented as one.
+Sample counts must never be scaled: a multiplexed sampling member simply
+does not observe what happens while it is evicted.
+.Ss Probing For Group Support
+There is no operation that asks whether grouping is available.
+A program probes by attempting the allocation it intended to perform
+anyway, and treating
+.Er EOPNOTSUPP
+from
+.Fn pmc_allocate_group
+as
+.Dq this kernel cannot group these events ,
+falling back to ungrouped allocation.
+A throwaway probe is a poor idea, since each one consumes a deferred
+handle and a group slot.
+.Pp
+An
+.Er EINVAL
+from
+.Fn pmc_group_commit
+is never a probe.
+It reports a usage error with several possible causes \(em no leader,
+two leaders, mixed virtual and system modes, system members bound to
+different CPUs, or a leader-only flag on another member \(em and should
+be shown to the user rather than treated as absent support.
+.Ss Deferred Handle Budget
+A grouped PMC keeps a stable identifier from allocation until release,
+even as multiplexing moves it between hardware counters.
+These identifiers come from a pool of 32 per owning process per CPU.
+.Pp
+Process-scope groups are not bound to a CPU, so all of an owner's
+process-scope members draw from a single pool of 32, however many groups
+they belong to.
+One maximum-size group therefore exhausts the entire process-scope
+budget, and a further allocation fails with
+.Er EMFILE
+until that group is released.
+System-scope groups draw from a separate pool for each CPU they are
+bound to.
+.Ss Reading A Group
+.Fn pmc_group_read
+fills a caller-supplied array.
+Pass a count of zero to ask how many members the group has; the call
+succeeds and reports the count without touching the array.
+Passing a nonzero count smaller than the membership fails with
+.Er E2BIG ,
+and again reports the required count, so a caller can size its array in
+one retry.
+.Pp
+.Fn pmc_read_pair
+reads a single PMC with the times that scale it.
+It allocates internally, so it may fail for reasons unrelated to the
+PMC, and it applies only to a group leader:
+on any other member it fails with
+.Er EOPNOTSUPP .
.Ss Signal Handling Requirements
Applications using PMCs are required to handle the following signals:
.Bl -tag -width ".Dv SIGBUS"
diff --git a/lib/libpmc/pmclog.3 b/lib/libpmc/pmclog.3
--- a/lib/libpmc/pmclog.3
+++ b/lib/libpmc/pmclog.3
@@ -21,7 +21,7 @@
.\" OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
.\" SUCH DAMAGE.
.\"
-.Dd March 26, 2006
+.Dd August 22, 2026
.Dt PMCLOG 3
.Os
.Sh NAME
@@ -93,6 +93,7 @@
struct pmclog_ev_pmcallocatedyn pl_ad;
struct pmclog_ev_pmcattach pl_t;
struct pmclog_ev_pmcdetach pl_d;
+ struct pmclog_ev_pmcgroupinheritmiss pl_gim;
struct pmclog_ev_proccsw pl_c;
struct pmclog_ev_procexec pl_x;
struct pmclog_ev_procexit pl_e;
@@ -201,6 +202,9 @@
A record describing a PMC attach operation.
.It Dv PMCLOG_TYPE_PMCDETACH
A record describing a PMC detach operation.
+.It Dv PMCLOG_TYPE_PMCGROUPINHERITMISS
+This record shows a failed group inheritance.
+The record gives the group-leader PMC and the child PID that missed it.
.It Dv PMCLOG_TYPE_PROCCSW
A record describing a PMC reading at the time of a process context switch.
.It Dv PMCLOG_TYPE_PROCEXEC
@@ -222,6 +226,14 @@
A record containing user data.
.El
.Pp
+PMC-specific identifier fields hold the stable handle from PMC allocation.
+For a grouped PMC, multiplexing can move the PMC between hardware rows.
+The handle stays the same in this case.
+Do not read the handle as the current hardware row.
+A record not tied to one PMC may use a record-defined sentinel value
+such as
+.Dv PMC_ID_INVALID .
+.Pp
Function
.Fn pmclog_feed
is used with parsers configured to parse memory based event streams.
diff --git a/usr.sbin/pmcstat/pmcstat.8 b/usr.sbin/pmcstat/pmcstat.8
--- a/usr.sbin/pmcstat/pmcstat.8
+++ b/usr.sbin/pmcstat/pmcstat.8
@@ -248,6 +248,15 @@
Allocate a system mode sampling PMC measuring hardware events
specified in
.Ar event-spec .
+A system mode PMC is bound to a single CPU: the lowest-numbered one in
+the set given by
+.Fl c ,
+or in the process's root CPU set if
+.Fl c
+was not given.
+To profile several CPUs, run one
+.Nm
+per CPU.
.It Fl T
Use a
.Xr top 1 Ns -like
@@ -309,6 +318,83 @@
saved with the
.Fl O
option.
+.It Fl b
+Enable PMU event grouping.
+When this flag is set, an argument to
+.Fl p ,
+.Fl P ,
+.Fl s
+or
+.Fl S
+may be a brace-list of the form
+.Sq Brq Ar ev1 , Ar ev2 , Ar ev3 :
+all listed events are programmed atomically as a single PMU group, so
+the counter values for siblings are read at consistent program points.
+Every brace-mode group leader is allocated with
+.Dv PMC_F_GROUP_MUX
+set, which has no effect on groups that fit in hardware but lets the
+kernel transparently time-multiplex any group that overflows the
+.Em remaining
+hardware counter budget.
+For example, on a 6-counter Zen3 the invocation
+.Bd -literal -offset indent
+pmcstat -b -p '{instructions,unhalted-cycles}' \\
+ -p '{ls_alloc_mab_count,ls_not_halted_cyc,ls_dispatch.all}' \\
+ -p '{ls_smi_rx,ls_int_taken,instructions}' sleep 60
+.Ed
+.Pp
+attaches three groups (2+3+3 = 8 events) to one target.
+The first two groups consume 5 of 6 counters, leaving 1.
+The third group is then committed in multiplex mode so its 3 events
+rotate through the remaining slot rather than failing with
+.Er ENOSPC .
+.Pp
+While a group is being multiplexed its counting-mode columns show
+.Em scaled estimates
+(interval count scaled by the group's enabled/running time ratio),
+and the group's hardware residency for the interval is appended in a
+trailing
+.Ql res%
+column that stays blank while the group is fully resident:
+.Bd -literal -offset indent
+# p/instructions p/unhalted-cycles p/branches res%
+ 24000210 51003001 92010 50.2%
+ <not counted> <not counted> <not counted> 0.0%
+ 23890110 50700323 91500
+.Ed
+.Pp
+An interval in which the group never reached hardware prints
+.Ql <not counted>
+for every member; when residency falls below 1/10
+.Pq Dv PMC_SCALE_MAX
+the raw interval count is shown instead of an extrapolated estimate.
+With
+.Fl C ,
+cumulative values are running sums of the per-interval estimates.
+.Pp
+A scaled value is an estimate of what the count would have been had the
+group stayed on hardware, obtained by assuming the events the group did
+not see resemble the ones it did.
+Over a short interval that assumption is weak: a group that was resident
+for one rotation window in three has sampled the interval three times,
+and the resulting estimate is correspondingly rough.
+Choose a display interval spanning at least about twenty rotation
+periods \(em with the default
+.Va kern.hwpmc.mux_period_ms
+of 50, roughly one second \(em before reading much into a scaled number.
+.Pp
+For system-wide sampling, multiplexing costs coverage rather than
+accuracy.
+A CPU whose group is evicted produces no samples at all for that window,
+and nothing reconstructs them: the profile is of the time the group was
+resident, not of the whole run.
+Sample counts are therefore never scaled.
+At the end of a run,
+.Nm
+sends this warning if a multiplexed group has a sampling period that is
+longer than the event budget of one rotation window.
+.Nm
+does not send this warning for a fully resident group.
.It Fl c Ar cpu-spec
Set the cpus for subsequent system mode PMCs specified on the
command line to
@@ -326,11 +412,32 @@
the target process.
The default is to measure events for the target process alone.
(it has to be passed in the command line prior to
-.Fl p ,
-.Fl s ,
-.Fl P ,
+.Fl p
or
-.Fl S ) .
+.Fl P ) .
+.Pp
+A system mode PMC is bound to a CPU and has no process target to follow
+through a fork, so combining
+.Fl d
+with
+.Fl s
+or
+.Fl S
+is rejected.
+.Pp
+With
+.Fl b ,
+a brace-list group follows a fork as a whole: every member attaches to
+the child, or none does.
+If the kernel cannot attach the group to a new child \(em a fork cannot
+be failed, so under memory pressure the child is simply not measured \(em
+its work is missing from the totals, which otherwise look complete.
+.Nm
+reports that at the end of the run:
+.Bd -literal -offset indent
+pmcstat: WARNING: 2 descendant attaches failed; their contributions
+are not in these totals.
+.Ed
.It Fl e
Specify that the gprof profile files will use a wide history counter.
These files are produced in a format compatible with
File Metadata
Details
Attached
Mime Type
text/plain
Expires
Fri, Oct 2, 6:48 AM (2 h, 35 m)
Storage Engine
blob
Storage Format
Raw Data
Storage Handle
40066447
Default Alt Text
D59365.diff (10 KB)
Attached To
Mode
D59365: pmc, pmcstat: document PMU event grouping
Attached
Detach File
Event Timeline
Log In to Comment