From: Ionut Nechita The managed IRQ affinity spreading introduced by the isolcpus= managed_irq_strict series restricts the affinity of managed interrupt vectors to housekeeping CPUs. However, the reserved vectors that do not participate in managed affinity spreading - the affd->pre_vectors (e.g. the NVMe admin completion queue) and the trailing affd->post_vectors - are still initialised with irq_default_affinity, which by default spans all CPUs including the isolated ones. These reserved vectors are not marked managed, so their affinity is only an initial hint. Still, on a freshly probed device the admin queue interrupt can be delivered to an isolated CPU until userspace (if ever) re-steers it. On PREEMPT_RT this is especially undesirable: a stray threaded interrupt waking on an isolated core directly perturbs the latency-critical workload, and there is no clean boot-time ordering guarantee for a userspace fixup. Confine these reserved vectors to the housekeeping mask when HK_TYPE_MANAGED_IRQ_STRICT is enabled, so that even the unmanaged admin/reserved interrupts are kept off isolated CPUs from the moment of allocation. When strict isolation is not configured the behaviour is unchanged (irq_default_affinity is used as before). This closes the remaining admin-queue gap left by the managed_irq_strict series for NVMe and other multiqueue drivers that reserve a pre_vector. Signed-off-by: Ionut Nechita --- Hi Aaron, This is the follow-up I referenced in section 3 of the v16 test report (the one applied as "0010" when gathering those results). It is functionally identical to the diff you sketched in your reply; the only differences are the local variable name (reserved_mask vs def_mask) and an explanatory comment. Posting it here as you suggested so you can fold it into patch 8 ("genirq/affinity: Restrict managed IRQ affinity to housekeeping CPUs") for v17, or take it as a standalone patch - whichever you prefer. Feel free to adjust the variable name to match your style if you squash it. The pre/post-vector affinities shown in section 3 of the report (nvme0q0 eff = 79, nvme1q0 eff = 39) were produced with this applied. Thanks, Ionut kernel/irq/affinity.c | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/kernel/irq/affinity.c b/kernel/irq/affinity.c index 8a7e68843c161..29eb22725cfd5 100644 --- a/kernel/irq/affinity.c +++ b/kernel/irq/affinity.c @@ -29,6 +29,7 @@ irq_create_affinity_masks(unsigned int nvecs, struct irq_affinity *affd) unsigned int affvecs, curvec, usedvecs, i, j; struct irq_affinity_desc *masks = NULL; const struct cpumask *hk_mask = housekeeping_cpumask(HK_TYPE_MANAGED_IRQ_STRICT); + const struct cpumask *reserved_mask; bool hk_enabled = housekeeping_enabled(HK_TYPE_MANAGED_IRQ_STRICT); /* @@ -63,9 +64,18 @@ irq_create_affinity_masks(unsigned int nvecs, struct irq_affinity *affd) if (!masks) return NULL; + /* + * Vectors that do not need managed affinity (pre_vectors such as the + * NVMe admin queue, and trailing post_vectors) default to + * irq_default_affinity, which may include isolated CPUs. Under strict + * isolation, confine them to housekeeping CPUs so that even these + * unmanaged interrupts never land on an isolated core. + */ + reserved_mask = hk_enabled ? hk_mask : irq_default_affinity; + /* Fill out vectors at the beginning that don't need affinity */ for (curvec = 0; curvec < affd->pre_vectors; curvec++) - cpumask_copy(&masks[curvec].mask, irq_default_affinity); + cpumask_copy(&masks[curvec].mask, reserved_mask); /* * Spread on present CPUs starting from affd->pre_vectors. If we @@ -104,7 +114,7 @@ irq_create_affinity_masks(unsigned int nvecs, struct irq_affinity *affd) else curvec = affd->pre_vectors + usedvecs; for (; curvec < nvecs; curvec++) - cpumask_copy(&masks[curvec].mask, irq_default_affinity); + cpumask_copy(&masks[curvec].mask, reserved_mask); /* Mark the managed interrupts */ for (i = affd->pre_vectors; i < nvecs - affd->post_vectors; i++) -- 2.47.3