The mbm_event mode assigns hardware MBM counters to RMID/event pairs. It is intended for deployments that need to manage counter assignment on platforms where the number of monitoring groups exceeds the available hardware counters. Hardware counters are scarce on these platforms, so mbm_event mode suits workflows that monitor a subset of groups at a time and rotate assignments as needed. resctrl enables the mbm_event mode by default on hardware that supports it. This breaks the pqos tool [1], which assumes the historical default mode. It creates 16 or more monitoring groups and uses two counters per group. On a platform with 32 counters per domain, that exhausts the mbm_event counters pool, causing additional groups to read "Unassigned". pqos interprets a non-numeric event read as zero bandwidth and reports 0 MB/s for those groups: $sudo pqos -m all:0-4 CORE IPC MISSES LLC[KB] MBL[MB/s] MBR[MB/s] 0 0.69 16k 192.0 0.0 0.0 1 0.40 3k 0.0 0.0 0.0 2 0.44 1k 64.0 0.0 0.0 3 0.39 1k 64.0 0.0 0.0 4 0.43 1k 0.0 0.0 0.0 Leave mbm_assign_mode in default mode during initialization. Default mode uses one counter per monitoring group. On existing AMD platforms, the default active counters pool is 64 and may be larger on newer hardware. The active counters pool is the number of counters that the hardware can actively track in default mode. Deployments within the active counter pool receive accurate bandwidth measurements in the default mode. Users can create more monitoring groups (4096) than there are active counters in the pool. Beyond that pool, readings may be misleading or report as "Unavailable", with no user-visible indication. Default mode is still limited once the number of monitoring groups exceeds that active counters pool. After hardware re-allocates a counter, a read may return "Unavailable". pqos still treats "Unavailable" as zero. Successive reads can be a count, "Unavailable", then another count. pqos treats those jumps as wraparound and can report an inconsistent rate: $sudo pqos -m all:0-4 CORE IPC MISSES LLC[KB] MBL[MB/s] MBR[MB/s] 0 0.76 60k 32.0 0.0 0.0 1 0.45 1k 32.0 0.0 0.0 2 1.57 107k 64.0 0.0 17592186044184.9 3 1.62 276k 4928.0 3.2 0.5 Users that need stable measurements beyond the active countes pool should use mbm_event mode and rotate assignments as needed. Enable it with: $echo mbm_event > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode Fixes: 0f1576e43adc ("x86/resctrl: Configure mbm_event mode if supported") Closes: https://github.com/intel/intel-cmt-cat/issues/311 Signed-off-by: Babu Moger Cc: stable@vger.kernel.org Link: https://github.com/intel/intel-cmt-cat # [1] --- v4: Introduced mbm_event counter pool and default active counter pool. Removed duplicate texts in resctrl.rst. Re-arranged the tag order. v3: Added pqos output to show the problem. More changelog to detail the issue. Updated the resctrl.rst to mention changes. v2: Added documentation describing the known issue with the default mode. Will add cc to stable once we have all the things in order. Let me know if I missed anything. v1: https://lore.kernel.org/lkml/8cb66e18e32e4087a9712c1e68ee6da614efe244.1784322818.git.babu.moger@amd.com/ --- Documentation/filesystems/resctrl.rst | 43 +++++++++++++++++---------- arch/x86/kernel/cpu/resctrl/monitor.c | 1 - 2 files changed, 28 insertions(+), 16 deletions(-) diff --git a/Documentation/filesystems/resctrl.rst b/Documentation/filesystems/resctrl.rst index c8507580474a..a896bcae3319 100644 --- a/Documentation/filesystems/resctrl.rst +++ b/Documentation/filesystems/resctrl.rst @@ -354,8 +354,8 @@ with the following files: :: # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode - [mbm_event] - default + [default] + mbm_event "mbm_event": @@ -376,20 +376,35 @@ with the following files: "mbm_L3_assignments" after switching to "mbm_event" mode for counter assignment states of all monitoring groups. - The mode is beneficial for AMD platforms that support more CTRL_MON - and MON groups than available hardware counters. By default, this - feature is enabled on AMD platforms with the ABMC (Assignable Bandwidth - Monitoring Counters) capability, ensuring counters remain assigned even - when the corresponding RMID is not actively used by any processor. + It is intended for deployments that need to manage counter assignment on + platforms where the number of monitoring groups exceeds the available + hardware counters. Hardware counters are scarce on these platforms, so + mbm_event mode suits workflows that monitor a subset of groups at a time and + rotate assignments as needed. The mbm_event mode ensures counters remain + assigned even when the corresponding RMID is not actively monitored. "default": In default mode, resctrl assumes there is a hardware counter for each - event within every CTRL_MON and MON group. On AMD platforms, it is - recommended to use the mbm_event mode, if supported, to prevent reset of MBM - events between reads resulting from hardware re-allocating counters. This can - result in misleading values or display "Unavailable" if no counter is assigned - to the event. + event within every CTRL_MON and MON group. This mode is enabled by default. + + On AMD platforms that support more CTRL_MON and MON groups than hardware + counters, hardware dynamically shares a smaller pool of counters among + RMIDs. One counter from this pool is used per monitoring group and counts + every MBM event of that group's RMID. The size of the pool is not + enumerated to software (unlike "num_mbm_cntrs"), and "num_rmids" may be + much larger. For example, 64 counters in this pool monitor 64 groups. + Newer hardware may provide a larger pool. + + While the number of monitoring groups does not exceed that pool, a counter + stays attached to each RMID and readings remain accurate. Creating more + groups than the pool can cause hardware to re-allocate those counters + between successive reads of an event. Bandwidth values may then be + misleading, or a read may return "Unavailable" if no counter is allocated + to the RMID. There is no user-visible indication when this begins. + Users who need stable readings beyond that pool should switch to mbm_event + mode, if supported, and assign counters to the groups of interest + (rotating assignments as needed). * To enable "mbm_event" counter assignment mode: :: @@ -1787,11 +1802,9 @@ View the llc occupancy snapshot:: Examples on working with mbm_assign_mode ======================================== -a. Check if MBM counter assignment mode is supported. +a. Check if MBM counter assignment mode is supported and enabled. :: - # mount -t resctrl resctrl /sys/fs/resctrl/ - # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode [mbm_event] default diff --git a/arch/x86/kernel/cpu/resctrl/monitor.c b/arch/x86/kernel/cpu/resctrl/monitor.c index 5d0d3b18f9b8..a6a9090c4808 100644 --- a/arch/x86/kernel/cpu/resctrl/monitor.c +++ b/arch/x86/kernel/cpu/resctrl/monitor.c @@ -472,7 +472,6 @@ int __init rdt_get_l3_mon_config(struct rdt_resource *r) cpuid_count(0x80000020, 5, &eax, &ebx, &ecx, &edx); /* cntr_id is 12 bits and can only encode 4096 counters. */ r->mon.num_mbm_cntrs = min((ebx & GENMASK(15, 0)) + 1, BIT(12)); - hw_res->mbm_cntr_assign_enabled = true; } r->mon_capable = true; -- 2.43.0