[PATCH v6 6/8] dma: swiotlb: Centralize memory-encryption pool sizing

Aneesh Kumar K.V aneesh.kumar at kernel.org
Thu Oct 8 07:32:12 PDT 2026


Catalin Marinas <catalin.marinas at arm.com> writes:

> On Thu, Oct 08, 2026 at 11:03:27AM +0530, Aneesh Kumar K.V wrote:
>> Aneesh Kumar K.V <aneesh.kumar at kernel.org> writes:
>> > Will Deacon <will at kernel.org> writes:
>> >> On Wed, Oct 07, 2026 at 11:04:05AM +0100, Catalin Marinas wrote:
>> >>> On Tue, Oct 06, 2026 at 10:49:17PM +0100, Will Deacon wrote:
>> >>> > On Thu, Sep 24, 2026 at 11:37:54AM +0530, Aneesh Kumar K.V (Arm) wrote:
>> >>> > > @@ -496,7 +516,8 @@ swiotlb_select_pool_policy(unsigned int flags)
>> >>> > >  	if (swiotlb_force_disable)
>> >>> > >  		return SWIOTLB_POOL_NONE;
>> >>> > >  
>> >>> > > -	if (cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT))
>> >>> > > +	if (cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT) &&
>> >>> > > +	    !restricted_dma_pool_present)
>> >>> > >  		return SWIOTLB_POOL_CC_GUEST;
>> >>> > 
>> >>> > I think this check on the restricted DMA pool is too general -- the pool
>> >>> > could be tied to a specific DMA-capable peripheral and so treating its
>> >>> > presence as a global property isn't right.
>> >>> 
>> >>> I agree it's a hack but that was the simplest way to avoid the pVMs
>> >>> getting a bounce buffer after this patch. More than happy to leave it
>> >>> out and reduce the buffer on cmdline or we come up with some better
>> >>> heuristics.
>> >>
>> >> Hrm, that does mean that reverting just this part will regress pVMs
>> >> because they'll suddenly be allocating a tonne more memory for an
>> >> entirely unused swiotlb buffer. So I think I'd prefer to drop the entire
>> >> series until this has been worked out properly.
>> >>
>> >>> Another option could be the arch code passing another flag that it
>> >>> doesn't want an encrypted pool (e.g. when running in a pKVM guest) but I
>> >>> don't particularly this either. The arch code doesn't know whether
>> >>> there's an alternative pool.
>> >>
>> >> At that point, the default size may as well be driven by the
>> >> drivers/virt/coco driver.
>> >>
>> >>> That said, such heuristics should have been a separate patch to make it
>> >>> easier to review/drop.
>> >>
>> >> I think the only right way to get a semi-accurate heuristic is to take
>> >> into account the set of dma-capable devices that will use the swiotlb
>> >> pool, but that's fiddly and should probably be tackled as a separate
>> >> series. Maybe a simpler hack in that direction would be to take the
>> >> SWIOTLB_POOL_CC_GUEST if _any_ device is going to use swiotlb? You'll
>> >> run into the usual problem of not being able to tell if a device is
>> >> DMA-capable or not, but you could probably look for a global restricted
>> >> DMA pool and, if that doesn't exist, check for per-device restricted pools
>> >> on dma-coherent devices (since restricted DMA isn't supported by ACPI) as
>> >> a reasonable approximation.
>> >
>> > So, something like this?
>> >
>> > 	if (cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT) &&
>> > 	    swiotlb_cc_guest_needs_default_pool())
>> > 		return SWIOTLB_POOL_CC_GUEST;
>> >
>> 
>> Detecting a DMA-capable device is not straightforward, and if we get it
>> wrong, we will enable SWIOTLB_POOL_CC_GUEST unnecessarily. Would the
>> code below be a reasonable approximation of what you suggested?
>> 
>> Another option would be to make swiotlb_cc_guest_needs_default_pool() a
>> weak function that architectures can override. arm64 pKVM could then use
>> a different scheme (for this patch series default to false). Would that
>> be preferable?
>
> Even the rmem check for each device is still a hack that may bite us in
> the future (private devices for example would not need swiotlb). I'm
> thinking more and more of leaving the sizing an arch-specific decision,
> don't bother generalising it at all.
>
> On pKVM vs CCA guests, there's really nothing specific here to pKVM
> guests. The only difference is that confidential guests that so far have
> run without a swiotlb buffer will regress if their memory is tight. For
> confidential guests without dedicated rmem (either CCA or pKVM), I think
> our options are either command line swiotlb sizing or dynamic swiotlb.
>
> Could you respin your series while leaving out the generic sizing? IOW,
> no x86 code generalisation. We can discuss the best strategy on sizing
> later (I haven't checked how much of this series still makes sense
> without the generic sizing).
>

It would mostly consist of the first three cleanup patches, followed by
three patches that replace the addressing_limit argument with flags.

#define SWIOTLB_VERBOSE	(1 << 0) /* verbose initialization */
/* Initialize a pool for devices with limited DMA addressing. */
#define SWIOTLB_INIT_ADDRESSING_LIMIT	(1 << 1)
/* Initialize a pool that requires architecture remapping. */
#define SWIOTLB_INIT_REMAP		(1 << 2)
/* Initialize a pool for DMA to memory-encrypted host or guest memory. */
#define SWIOTLB_INIT_MEM_ENCRYPT		(1 << 3)
/* Initialize a pool for unaligned kmalloc bouncing. */
#define SWIOTLB_INIT_KMALLOC		(1 << 4)
/* Do not initialize a pool unless SWIOTLB is explicitly required. */
#define SWIOTLB_INIT_DEFAULT_OFF		(1 << 5)

This results in the large change below. I'm not sure we want to do this
for no real benefit other than making swiotlb_should_init() slightly
easier to follow.

static bool __init swiotlb_should_init(unsigned int flags)
{
	if (swiotlb_force_disable)
		return false;

	if (swiotlb_force_bounce)
		return true;

	if (flags & (SWIOTLB_INIT_REMAP | SWIOTLB_INIT_MEM_ENCRYPT)))
		return true;

	/* Explicit requirements override an architecture's default opt-out. */
	if (flags & SWIOTLB_INIT_DEFAULT_OFF)
		return false;

	return flags & (SWIOTLB_INIT_ADDRESSING_LIMIT | SWIOTLB_INIT_KMALLOC);
}

Marek,

Patch 3 is a fix, so you may want to take it even if we drop the rest of
the series. Perhaps the first three patches could be taken together?

modified   arch/arm/mm/init.c
@@ -223,7 +223,11 @@ static inline void poison_init_mem(void *s, size_t count)
 void __init arch_mm_preinit(void)
 {
 #ifdef CONFIG_ARM_LPAE
-	swiotlb_init(max_pfn > arm_dma_pfn_limit, SWIOTLB_VERBOSE);
+	unsigned int flags = SWIOTLB_VERBOSE;
+
+	if (max_pfn > arm_dma_pfn_limit)
+		flags |= SWIOTLB_INIT_ADDRESSING_LIMIT;
+	swiotlb_init(flags);
 #endif
 
 #ifdef CONFIG_SA1111
modified   arch/arm64/mm/init.c
@@ -351,7 +351,10 @@ void __init arch_mm_preinit(void)
 		swiotlb_adjust_size(min(swiotlb_default_pool_size(), size));
 	}
 
-	swiotlb_init(true, flags);
+	if (max_pfn > PFN_DOWN(arm64_dma_phys_limit))
+		flags |= SWIOTLB_INIT_ADDRESSING_LIMIT;
+
+	swiotlb_init(flags);
 
 	/*
 	 * Check boundaries twice: Some fundamental inconsistencies can be
modified   arch/loongarch/kernel/setup.c
@@ -404,7 +404,7 @@ static void __init arch_mem_init(char **cmdline_p)
 
 	memblock_set_bottom_up(true);
 
-	swiotlb_init(true, SWIOTLB_VERBOSE);
+	swiotlb_init(SWIOTLB_VERBOSE | SWIOTLB_INIT_ADDRESSING_LIMIT);
 
 	dma_contiguous_reserve(PFN_PHYS(max_low_pfn));
 
modified   arch/mips/cavium-octeon/dma-octeon.c
@@ -235,5 +235,5 @@ void __init plat_swiotlb_setup(void)
 #endif
 
 	swiotlb_adjust_size(swiotlbsize);
-	swiotlb_init(true, SWIOTLB_VERBOSE);
+	swiotlb_init(SWIOTLB_VERBOSE | SWIOTLB_INIT_ADDRESSING_LIMIT);
 }
modified   arch/mips/loongson64/dma.c
@@ -25,5 +25,5 @@ phys_addr_t dma_to_phys(struct device *dev, dma_addr_t daddr)
 
 void __init plat_swiotlb_setup(void)
 {
-	swiotlb_init(true, SWIOTLB_VERBOSE);
+	swiotlb_init(SWIOTLB_VERBOSE | SWIOTLB_INIT_ADDRESSING_LIMIT);
 }
modified   arch/mips/sibyte/common/dma.c
@@ -10,5 +10,5 @@
 
 void __init plat_swiotlb_setup(void)
 {
-	swiotlb_init(true, SWIOTLB_VERBOSE);
+	swiotlb_init(SWIOTLB_VERBOSE | SWIOTLB_INIT_ADDRESSING_LIMIT);
 }
modified   arch/powerpc/kernel/dma-swiotlb.c
@@ -14,8 +14,10 @@ unsigned int ppc_swiotlb_flags;
 
 void __init swiotlb_detect_4g(void)
 {
-	if ((memblock_end_of_DRAM() - 1) > 0xffffffff)
+	if ((memblock_end_of_DRAM() - 1) > 0xffffffff) {
 		ppc_swiotlb_enable = 1;
+		ppc_swiotlb_flags |= SWIOTLB_INIT_ADDRESSING_LIMIT;
+	}
 }
 
 static int __init check_swiotlb_enabled(void)
modified   arch/powerpc/mm/mem.c
@@ -287,6 +287,19 @@ void __init arch_mm_preinit(void)
 	BUILD_BUG_ON(MMU_PAGE_COUNT > 16);
 
 #ifdef CONFIG_SWIOTLB
+	if (is_secure_guest()) {
+
+		/* Don't release the SWIOTLB buffer. */
+		ppc_swiotlb_enable = 1;
+
+		/*
+		 * Since the guest memory is inaccessible to the host,
+		 * devices always need to use the SWIOTLB buffer for DMA
+		 * even if dma_capable() says otherwise.
+		 */
+		ppc_swiotlb_flags |= SWIOTLB_ANY;
+	}
+
 	/*
 	 * Some platforms (e.g. 85xx) limit DMA-able memory way below
 	 * 4G. We force memblock to bottom-up mode to ensure that the
@@ -295,7 +308,7 @@ void __init arch_mm_preinit(void)
 	 * back to to-down.
 	 */
 	memblock_set_bottom_up(true);
-	swiotlb_init(ppc_swiotlb_enable, ppc_swiotlb_flags);
+	swiotlb_init(ppc_swiotlb_flags);
 #endif
 
 	kasan_late_init();
modified   arch/powerpc/platforms/pseries/svm.c
@@ -21,16 +21,6 @@ static int __init init_svm(void)
 	if (!is_secure_guest())
 		return 0;
 
-	/* Don't release the SWIOTLB buffer. */
-	ppc_swiotlb_enable = 1;
-
-	/*
-	 * Since the guest memory is inaccessible to the host, devices always
-	 * need to use the SWIOTLB buffer for DMA even if dma_capable() says
-	 * otherwise.
-	 */
-	ppc_swiotlb_flags |= SWIOTLB_ANY;
-
 	/* Share the SWIOTLB buffer with the host. */
 	swiotlb_update_mem_attributes();
 
modified   arch/powerpc/sysdev/fsl_pci.c
@@ -444,6 +444,7 @@ static void setup_pci_atmu(struct pci_controller *hose)
 	if (hose->dma_window_size < mem) {
 #ifdef CONFIG_SWIOTLB
 		ppc_swiotlb_enable = 1;
+		ppc_swiotlb_flags |= SWIOTLB_INIT_ADDRESSING_LIMIT;
 #else
 		pr_err("%pOF: ERROR: Memory size exceeds PCI ATMU ability to "
 			"map - enable CONFIG_SWIOTLB to avoid dma errors.\n",
modified   arch/riscv/mm/init.c
@@ -165,14 +165,17 @@ static void print_vm_layout(void) { }
 
 void __init arch_mm_preinit(void)
 {
-	bool swiotlb = max_pfn > PFN_DOWN(dma32_phys_limit) &&
-		       memblock_start_of_DRAM() < dma32_phys_limit;
 	unsigned int swiotlb_flags = SWIOTLB_VERBOSE;
 #ifdef CONFIG_FLATMEM
 	BUG_ON(!mem_map);
 #endif /* CONFIG_FLATMEM */
 
-	if (IS_ENABLED(CONFIG_DMA_BOUNCE_UNALIGNED_KMALLOC) && !swiotlb &&
+	if (max_pfn > PFN_DOWN(dma32_phys_limit) &&
+	    memblock_start_of_DRAM() < dma32_phys_limit)
+		swiotlb_flags |= SWIOTLB_INIT_ADDRESSING_LIMIT;
+
+	if (IS_ENABLED(CONFIG_DMA_BOUNCE_UNALIGNED_KMALLOC) &&
+	    !(swiotlb_flags & SWIOTLB_INIT_ADDRESSING_LIMIT) &&
 	    dma_cache_alignment != 1) {
 		/*
 		 * No 32-bit DMA bouncing needed (either all DRAM is within
@@ -186,11 +189,10 @@ void __init arch_mm_preinit(void)
 		unsigned long size =
 			DIV_ROUND_UP(memblock_phys_mem_size(), 1024);
 		swiotlb_adjust_size(min(swiotlb_default_pool_size(), size));
-		swiotlb = true;
-		swiotlb_flags |= SWIOTLB_ANY;
+		swiotlb_flags |= SWIOTLB_INIT_KMALLOC | SWIOTLB_ANY;
 	}
 
-	swiotlb_init(swiotlb, swiotlb_flags);
+	swiotlb_init(swiotlb_flags);
 
 	print_vm_layout();
 }
modified   arch/s390/mm/init.c
@@ -166,7 +166,7 @@ static void __init pv_init(void)
 	virtio_set_mem_acc_cb(virtio_require_restricted_mem_acc);
 
 	/* make sure bounce buffers are shared */
-	swiotlb_init(true, SWIOTLB_VERBOSE | SWIOTLB_ANY);
+	swiotlb_init(SWIOTLB_VERBOSE | SWIOTLB_ANY);
 	swiotlb_update_mem_attributes();
 }
 
modified   arch/x86/kernel/pci-dma.c
@@ -44,8 +44,10 @@ static unsigned int x86_swiotlb_flags;
 static void __init pci_swiotlb_detect(void)
 {
 	/* don't initialize swiotlb if iommu=off (no_iommu=1) */
-	if (!no_iommu && max_possible_pfn > MAX_DMA32_PFN)
+	if (!no_iommu && max_possible_pfn > MAX_DMA32_PFN) {
 		x86_swiotlb_enable = true;
+		x86_swiotlb_flags |= SWIOTLB_INIT_ADDRESSING_LIMIT;
+	}
 
 	/*
 	 * Set swiotlb to 1 so that bounce buffers are allocated and used for
@@ -81,8 +83,10 @@ static void __init pci_xen_swiotlb_init(void)
 	if (!xen_swiotlb_enabled())
 		return;
 	x86_swiotlb_enable = true;
-	x86_swiotlb_flags |= SWIOTLB_ANY;
-	swiotlb_init_remap(true, x86_swiotlb_flags, xen_swiotlb_fixup);
+	/* Xen can use a SWIOTLB pool anywhere in directly mapped memory. */
+	x86_swiotlb_flags &= ~SWIOTLB_INIT_ADDRESSING_LIMIT;
+	x86_swiotlb_flags |= SWIOTLB_INIT_REMAP | SWIOTLB_ANY;
+	swiotlb_init_remap(x86_swiotlb_flags, xen_swiotlb_fixup);
 	dma_ops = &xen_swiotlb_dma_ops;
 	if (IS_ENABLED(CONFIG_PCI))
 		pci_request_acs();
@@ -103,7 +107,7 @@ void __init pci_iommu_alloc(void)
 	gart_iommu_hole_init();
 	amd_iommu_detect();
 	detect_intel_iommu();
-	swiotlb_init(x86_swiotlb_enable, x86_swiotlb_flags);
+	swiotlb_init(x86_swiotlb_flags);
 }
 
 static __init int iommu_setup(char *p)
@@ -149,8 +153,10 @@ static __init int iommu_setup(char *p)
 			return 1;
 		}
 #ifdef CONFIG_SWIOTLB
-		if (!strncmp(p, "soft", 4))
+		if (!strncmp(p, "soft", 4)) {
 			x86_swiotlb_enable = true;
+			x86_swiotlb_flags |= SWIOTLB_INIT_ADDRESSING_LIMIT;
+		}
 #endif
 		if (!strncmp(p, "pt", 2))
 			iommu_set_default_passthrough(true);
modified   include/linux/swiotlb.h
@@ -16,6 +16,14 @@ struct scatterlist;
 
 #define SWIOTLB_VERBOSE	(1 << 0) /* verbose initialization */
 #define SWIOTLB_ANY	(1 << 1) /* allow any memory for the buffer */
+/* Initialize a pool for devices with limited DMA addressing. */
+#define SWIOTLB_INIT_ADDRESSING_LIMIT	(1 << 2)
+/* Initialize a pool that requires architecture remapping. */
+#define SWIOTLB_INIT_REMAP		(1 << 3)
+/* Initialize a pool for DMA to memory-encrypted host or guest memory. */
+#define SWIOTLB_INIT_MEM_ENCRYPT		(1 << 4)
+/* Initialize a pool for unaligned kmalloc bouncing. */
+#define SWIOTLB_INIT_KMALLOC		(1 << 5)
 
 /*
  * Maximum allowable number of contiguous slabs to map,
@@ -39,8 +47,8 @@ struct scatterlist;
 #endif
 
 unsigned long swiotlb_default_pool_size(void);
-void __init swiotlb_init_remap(bool addressing_limit, unsigned int flags,
-	int (*remap)(void *tlb, unsigned long nslabs));
+void __init swiotlb_init_remap(unsigned int flags,
+			       int (*remap)(void *tlb, unsigned long nslabs));
 int swiotlb_init_late(size_t size, gfp_t gfp_mask,
 	int (*remap)(void *tlb, unsigned long nslabs));
 extern void __init swiotlb_update_mem_attributes(void);
@@ -183,7 +191,7 @@ static inline bool is_swiotlb_force_bounce(struct device *dev)
 	return mem && mem->force_bounce;
 }
 
-void swiotlb_init(bool addressing_limited, unsigned int flags);
+void swiotlb_init(unsigned int flags);
 void __init swiotlb_exit(void);
 void swiotlb_dev_init(struct device *dev);
 size_t swiotlb_max_mapping_size(struct device *dev);
@@ -193,7 +201,7 @@ void __init swiotlb_adjust_size(unsigned long size);
 phys_addr_t default_swiotlb_base(void);
 phys_addr_t default_swiotlb_limit(void);
 #else
-static inline void swiotlb_init(bool addressing_limited, unsigned int flags)
+static inline void swiotlb_init(unsigned int flags)
 {
 }
 
modified   kernel/dma/swiotlb.c
@@ -354,24 +354,15 @@ static void swiotlb_mark_pool_used(struct io_tlb_pool *pool)
 void __init swiotlb_update_mem_attributes(void)
 {
 	struct io_tlb_pool *mem = &io_tlb_default_mem.defpool;
-	unsigned long bytes;
-
-	/*
-	 * if platform support memory encryption, swiotlb buffers are
-	 * shared by default.
-	 */
-	if (cc_platform_has(CC_ATTR_MEM_ENCRYPT))
-		io_tlb_default_mem.cc_shared = true;
-	else
-		io_tlb_default_mem.cc_shared = false;
 
 	if (!mem->nslabs || mem->late_alloc)
 		return;
-	bytes = PAGE_ALIGN(mem->nslabs << IO_TLB_SHIFT);
 
 	if (io_tlb_default_mem.cc_shared) {
 		int ret;
+		unsigned long bytes;
 
+		bytes = PAGE_ALIGN(mem->nslabs << IO_TLB_SHIFT);
 		ret = set_memory_decrypted((unsigned long)mem->vaddr,
 					   bytes >> PAGE_SHIFT);
 		if (ret) {
@@ -462,12 +453,32 @@ static void __init *swiotlb_memblock_alloc(unsigned long nslabs,
 	return tlb;
 }
 
+static bool __init swiotlb_kmalloc_needs_bounce(void)
+{
+	return IS_ENABLED(CONFIG_DMA_BOUNCE_UNALIGNED_KMALLOC) &&
+	       (dma_get_cache_alignment() > 1);
+}
+
+static bool __init swiotlb_should_init(unsigned int flags)
+{
+	if (swiotlb_force_disable)
+		return false;
+
+	if (swiotlb_force_bounce)
+		return true;
+
+	return (flags & (SWIOTLB_INIT_ADDRESSING_LIMIT |
+			 SWIOTLB_INIT_REMAP |
+			 SWIOTLB_INIT_MEM_ENCRYPT |
+			 SWIOTLB_INIT_KMALLOC));
+}
+
 /*
  * Statically reserve bounce buffer space and initialize bounce buffer data
  * structures for the software IO TLB used to implement the DMA API.
  */
-void __init swiotlb_init_remap(bool addressing_limit, unsigned int flags,
-		int (*remap)(void *tlb, unsigned long nslabs))
+void __init swiotlb_init_remap(unsigned int flags,
+			       int (*remap)(void *tlb, unsigned long nslabs))
 {
 	struct io_tlb_pool *mem = &io_tlb_default_mem.defpool;
 	unsigned long nslabs;
@@ -475,9 +486,12 @@ void __init swiotlb_init_remap(bool addressing_limit, unsigned int flags,
 	size_t alloc_size;
 	void *tlb;
 
-	if (!addressing_limit && !swiotlb_force_bounce)
-		return;
-	if (swiotlb_force_disable)
+	if (cc_platform_has(CC_ATTR_MEM_ENCRYPT))
+		flags |= SWIOTLB_INIT_MEM_ENCRYPT;
+	if (swiotlb_kmalloc_needs_bounce())
+		flags |= SWIOTLB_INIT_KMALLOC;
+
+	if (!swiotlb_should_init(flags))
 		return;
 
 	io_tlb_default_mem.force_bounce = swiotlb_force_bounce;
@@ -491,6 +505,10 @@ void __init swiotlb_init_remap(bool addressing_limit, unsigned int flags,
 		io_tlb_default_mem.phys_limit = ARCH_LOW_ADDRESS_LIMIT;
 #endif
 
+	/* if we have host or guest memory encryption */
+	if (cc_platform_has(CC_ATTR_MEM_ENCRYPT))
+		io_tlb_default_mem.cc_shared = true;
+
 	if (!default_nareas)
 		swiotlb_adjust_nareas(num_possible_cpus());
 
@@ -531,9 +549,9 @@ void __init swiotlb_init_remap(bool addressing_limit, unsigned int flags,
 		swiotlb_print_info();
 }
 
-void __init swiotlb_init(bool addressing_limit, unsigned int flags)
+void __init swiotlb_init(unsigned int flags)
 {
-	swiotlb_init_remap(addressing_limit, flags, NULL);
+	swiotlb_init_remap(flags, NULL);
 }
 
 /*


-aneesh



More information about the linux-riscv mailing list