[PATCH 2/2] sched/fair: Honor asymmetric SMT priority in idle selection
Andrea Righi
arighi at nvidia.com
Tue Sep 8 13:49:41 PDT 2026
Hi Prateek,
On Wed, Sep 09, 2026 at 01:10:48AM +0530, K Prateek Nayak wrote:
> Hello Andrea,
>
> On 9/8/2026 1:53 PM, Andrea Righi wrote:
> > @@ -9747,8 +9796,10 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags)
> > }
> >
> > /* Slow path */
> > - if (unlikely(sd))
> > - return sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag);
> > + if (unlikely(sd)) {
> > + new_cpu = sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag);
> > + return select_idle_smt_cpu(p, new_cpu);
>
> nit. I personally feel this can be better integrated into the
> sched_balance_find_dst_cpu(). Something like the following:
>
> (Only build tested)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index b8bd308c2d5b..1012dfb33f08 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -12353,6 +12353,17 @@ static inline void update_sg_wakeup_stats(struct sched_domain *sd,
>
> }
>
> + /*
> + * If we are on a SD_SHARE_CPUCAPACITY | SD_ASYM_PACKING
> + * domain, use the group_asym_packing classification to
> + * decide placement based on rankings of idle siblings.
> + */
> + if (unlikely(sched_smt_asym_active() &&
> + (sd->flags & SD_SHARE_CPUCAPACITY) &&
> + (sd->flags & SD_ASYM_PACKING) &&
> + sgs->idle_cpus))
> + sgs->group_asym_packing = 1;
Integrating the preference in the slow-path selection sounds appealing, but I
don't think group_asym_packing can be used as a destination classificaiton here.
The intended policy is to prefer PE0 over PE1 when both siblings of the selected
SMT core are idle. And if PE0 is busy, PE1 should remain a valid destination. It
shouldn't make a busy PE0 preferable to an idle PE1.
IIUC group_type is ordered for busiest-group selection, group_asym_packing
describes a source group whole load should be moved to a "more preferred" CPU.
Marking an idle SMT group as group_asym_packing could make it rank worse than a
fully busy group.
Example: a fork on SMT2 can have the busy local PE0 classified as
group_has_spare or group_fully_busy, while the idle PE1 is forced to
group_asym_packing, sched_balance_find_dst_group() can then consider the busy
local group the better destination and stack the new task on PE0. That may
preserve one-thread mode for a short task, but it can also reduce throughput for
sustained work.
> +
> sgs->group_capacity = group->sgc->capacity;
>
> sgs->group_weight = group->group_weight;
> @@ -12393,9 +12404,15 @@ static bool update_pick_idlest(struct sched_group *idlest,
> return false;
> break;
>
> + case group_asym_packing:
> + /*
> + * Only possible for sched_smt_asym_active().
> + * Select the idle SMT that is more preferred.
> + */
> + return sched_asym_prefer(idlest->asym_prefer_cpu,
> + group->asym_prefer_cpu);
update_pick_idlest() returns true when @group should replace @idlest, so I think
the operands would need to be reversed.
But even with that, the local-versus-idlest comparison still returns NULL when
both sides are group_asym_packing. Therefore, if the slow path initially lands
on an idle PE1 while PE0 is also idle, it would not switch to PE0.
> case group_llc_balance:
> case group_imbalanced:
> - case group_asym_packing:
> case group_smt_balance:
> /* Those types are not used in the slow wakeup path */
> return false;
> ---
>
> It leads to slightly more branches but they are super predictable when
> iterating at a particular sched_domain level so the overhead should be
> negligible.
>
> I don't have any strong feelings either ways. Thoughts?
A deeper integration could preserve the normal group_has_spare classification
and use SMT priority only as a tie-breaker between otherwise equivalent
available siblings. It'd also need to handle the local-versus-idlest comparison
and preserve choose_idle_cpu() semantics I think.
For now, applying select_idle_smt_cpu() after the existing slow-path selection
seems simpler. It lets the existing load and capacity logic choose the core
first, then applies the preference only among available siblings within that
core.
Thanks,
-Andrea
More information about the linux-arm-kernel
mailing list