[PATCH v4 0/2] sched: Enable preferred SMT siblings on NVIDIA Olympus
Dietmar Eggemann
dietmar.eggemann at arm.com
Wed Sep 9 00:20:35 PDT 2026
On 08.09.26 10:23, Andrea Righi wrote:
[...]
> The series was tested on a two-node Vera system using an 88-thread
> single-precision GEMM on the 88 physical cores of NUMA node 0.
Can we use 'OpenBLAS benchmark/sgemm.goto' as an open alternative for
your NVIDIA internal single-precision GEMM benchmark?
IIUC, you used it for the 'Prefer fully idle cores for NOHZ balancing'
work: https://lore.kernel.org/r/anIq6pU5KXTTFCDN@gpd4
If yes, I assume you would run something like:
export OMP_NUM_THREADS=88
numactl -C XXX --membind=0 ./benchmark/sgemm.goto 16384 16384 16384
Essentially you want to show that those 88 compute intensive tasks each
runs on his own core alone and so you get a higher TFLOPS value.
[...]
More information about the linux-arm-kernel
mailing list