[PATCH v2] random: vDSO: Avoid call to memset() when zeroing reserved in __cvdso_getrandom_data()

Jason A. Donenfeld Jason at zx2c4.com
Thu Oct 1 02:25:35 PDT 2026


On Thu, Oct 01, 2026 at 06:41:20AM +0200, Christophe Leroy (CS GROUP) wrote:
> Hi,
> 
> Le 30/09/2026 à 17:16, Jason A. Donenfeld a écrit :
> > On Wed, Sep 30, 2026 at 5:13 PM Nathan Chancellor <nathan at kernel.org> wrote:
> >>
> >> On Wed, Sep 30, 2026 at 04:44:29PM +0200, Jason A. Donenfeld wrote:
> >>> Nathan, would this be okay with you?
> >>> https://eur01.safelinks.protection.outlook.com/?url=https%3A%2F%2Fgit.zx2c4.com%2Flinux-rng%2Fcommit%2F%3Fid%3Dd216701724b7d8209ff42150658ad5c712bdb503&data=05%7C02%7Cchristophe.leroy2%40cs-soprasteria.com%7C8f69935bffb344bf0ddc08df1f05e358%7C8b87af7d86474dc78df45f69a2011bb5%7C0%7C0%7C639263782324080837%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=BA7Fyr0%2FynggvBnR%2FpiCvNCQHtNBagqvSb0i11AvoLU%3D&reserved=0
> >>
> >> Can you stick
> >>
> >>    Cc: stable at vger.kernel.org # v6.12+
> >>    Closes: https://eur01.safelinks.protection.outlook.com/?url=https%3A%2F%2Fgithub.com%2FClangBuiltLinux%2Flinux%2Fissues%2F2183&data=05%7C02%7Cchristophe.leroy2%40cs-soprasteria.com%7C8f69935bffb344bf0ddc08df1f05e358%7C8b87af7d86474dc78df45f69a2011bb5%7C0%7C0%7C639263782324110679%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=6C8%2FAEQoh%2B8pfgfYQkL7qf5Ue%2F5sAGtfqRptzsSG0CM%3D&reserved=0
> >>
> >> on that? Otherwise, looks good to me, that's basically what I had for my
> >> v3 locally.
> > 
> > Sure, done. Also removed the now-unused array_size.h include.
> 
> I'm still very sceptic with this patch. You are degrading the behaviour 
> with GCC for a problem with CLANG. Why ?
> 
> Before the patch, with both GCC 13 and GCC 16 on powerpc32 I get a 
> pretty standard optimised loop that clears words 4 by 4 (with auto 
> increment of pointer) which is the most optimal on powerpc:
> 
>   3f0:	39 00 00 0c 	li      r8,12
>   3f4:	35 08 ff fc 	addic.  r8,r8,-4
>   3f8:	91 49 00 04 	stw     r10,4(r9)
>   3fc:	91 49 00 08 	stw     r10,8(r9)
>   400:	91 49 00 0c 	stw     r10,12(r9)
>   404:	95 49 00 10 	stwu    r10,16(r9)
>   408:	40 82 ff ec 	bne     3f4 <__c_kernel_getrandom+0x3f4>
> 
> With the patch,
> 
> With GCC 13 I get a very suboptimal loop copying bytes one by one
> 
>   3d8:	39 40 00 34 	li      r10,52
> ...
>   3e4:	39 20 00 00 	li      r9,0
>   3e8:	7d 49 03 a6 	mtctr   r10
>   3ec:	9d 3e 00 01 	stbu    r9,1(r30)
>   3f0:	42 00 ff fc 	bdnz    3ec <__c_kernel_getrandom+0x3ec>
> 
> With GCC 16 I get something a bit better but not as good as before, it 
> is a loop clearing words only one by one and incrementing pointer with 
> an additional insn instead of using auto-increment instruction stwu.
> 
>   3e0:	39 40 00 0d 	li      r10,13
> ...
>   3f0:	7d 49 03 a6 	mtctr   r10
>   3f4:	91 3f 00 00 	stw     r9,0(r31)
>   3f8:	3b ff 00 04 	addi    r31,r31,4
>   3fc:	42 00 ff f8 	bdnz    3f4 <__c_kernel_getrandom+0x3f4>
> 
> Please restrict the patch to clang builds.

Darn. Yea. The naive memset kills optimizations.

Okay, new strategy:

- on clang, pass `-mllvm -max-store-memset=4294967295`
- on gcc, pass `-finline-stringops=memset`

And keep the same code. (Or, better, see if those options generate good
code with b7bad082e113640fc81200ff869e5c2d7a9c29a2 reverted; I would
prefer that simpler initializer.)

Nathan, do these work?

Jason



More information about the linux-riscv mailing list