[RFC PATCH v4 03/16] iommu/arm-smmu-v3: Add initial pSMMU realm viommu plumbing

Aneesh Kumar K.V aneesh.kumar at kernel.org
Wed Sep 2 09:39:15 PDT 2026


Aneesh Kumar K.V <aneesh.kumar at kernel.org> writes:

> Jason Gunthorpe <jgg at ziepe.ca> writes:
>
>> On Wed, Sep 02, 2026 at 02:30:00PM +0530, Aneesh Kumar K.V wrote:
>>> Jason Gunthorpe <jgg at ziepe.ca> writes:
>>> 
>>> > On Tue, Sep 01, 2026 at 03:36:37PM +0530, Aneesh Kumar K.V wrote:
>>> >
>>> >> @@ -463,14 +460,13 @@
>>> >>  	vsmmu->vmid = s2_parent->s2_cfg.vmid;
>>> >>  
>>> >>  	if (viommu->type == IOMMU_VIOMMU_TYPE_ARM_SMMUV3) {
>>> >> +		if (arm_smmu_is_realm_viommu(viommu))
>>> >> +			return arm_realm_smmu_v3_init(viommu, user_data);
>>> >> +
>>> >
>>> > I think the realm vsmmu is going to require a different info struct
>>> > than the normal psmmu case, isn't it?
>>> >
>>> > If so it needs its own enum value.
>>> >
>>> > It would be nice to see a draft patch showing how the real vsmmu works
>>> > on top of the RMM spec for it. If we are using a viommu object then
>>> > non-vsmmu case should be identical just with an option in the info
>>> > struct to not create the vsmmu object.
>>> 
>>> Based on feedback on other emails in this thread, I have now implemented
>>> this without using a vdevice or viommu. This should make the CCA and
>>> non-CCA cases similar.
>>
>> That wasn't the feedback. The feedback was to use the viommu and not
>> make a bunch of new stuff..
>>
>
> That rework was done before I saw your discussion with Nicolin.
>
> To reiterate, for this configuration:
>
> - The viommu will use a stage-1 bypass configuration.
> - A new IOMMU_VIOMMU_TYPE_ARM_REALM_SMMUV3 type will create the viommu.
>   The psmmu will be activated at this point to avoid creating psmmu
>   objects early. We will reference-count it to ensure that the same
>   psmmu is shared across realm guests.
> - Creating a vdevice will invoke SMC_RMI_PSMMU_ST_L2_CREATE.
>
> I am unclear about the vdev_create suggestion. Creating a vdevice
> requires an RD, which is created later in the flow above. How do you
> suggest linking vdevice_alloc to vdev_create?
>


I was able to prototype the following flow:

 - The viommu uses a stage-1 bypass configuration.
 - A new IOMMU_VIOMMU_TYPE_ARM_REALM_SMMUV3 type creates the viommu.
   The pSMMU is activated at this point to avoid creating pSMMU objects
   early. It is reference-counted so that the same pSMMU can be shared
   across realm guests.
 - The realm is created during viommu allocation, ensuring that the
   vSMMU can be created here.
 - Creating a vdevice invokes SMC_RMI_PSMMU_ST_L2_CREATE.
 - This is followed by tsm_bind() if the device has an established TSM
   link.

static int arm_realm_smmu_v3_vdevice_init(struct iommufd_vdevice *vdev)
{
	struct device *dev = iommufd_vdevice_to_device(vdev);
	struct kvm *kvm = vdev->viommu->kvm_file->private_data;
	struct arm_smmu_device *smmu;
	struct arm_smmu_stream *stream;
	struct arm_smmu_master *master;
	unsigned long rmi_ret = 0;
	unsigned long l2_sid;
	int ret;

	if (!tsm_is_configured(dev))
		return 0;

	master = dev_iommu_priv_get(dev);
	/* FIXME which stream to pick */
	/* At this moment, iommufd only supports PCI device that has one SID */
	stream = &master->streams[0];
	smmu = master->smmu;

	l2_sid = ALIGN_DOWN(stream->id, STRTAB_NUM_L2_STES);

	{
		guard(mutex)(&smmu->realm.mutex);

		if (!arm_realm_smmu_active(smmu))
			return -EINVAL;

		ret = rmi_psmmu_st_l2_create(smmu->base_phys, l2_sid,
					     &rmi_ret);
		if (ret || rmi_ret) {
			if (!ret)
				return -EIO;
			if (RMI_RETURN_STATUS(rmi_ret) != RMI_ERROR_PSMMU_ST ||
			    RMI_RETURN_INDEX(rmi_ret) != 2) {
				dev_warn(dev, "failed to create realm stream mapping\n");
				return -EIO;
			}
			/* The L2 stream table already exists. */
		}
	}

	vdev->destroy = arm_realm_smmu_v3_vdevice_destroy;
	return tsm_bind(dev, kvm, vdev->virt_id);
}

 - tsm_bind() calls cca_tsm_bind(), which in turn calls vdev_create().
 - After boot, the guest locks the device. This generates an RHI request
   that reaches cca_tsm_guest_req() with RHI_DA_TDI_CONFIG_LOCKED.
 - cca_tsm_guest_req() now handles TSM_REQ_SET_TDI_STATE requests for
   the unlocked, locked, and running states.

@@ -513,10 +514,16 @@ static ssize_t cca_tsm_guest_req(struct pci_tdi *tdi,
 		if (copy_from_user((void *)&req_obj, req.user, req_len))
 			return -EFAULT;
 
-		if (req_obj.tdi_state != RHI_DA_TDI_CONFIG_RUN)
+		switch (req_obj.tdi_state) {
+		case RHI_DA_TDI_CONFIG_UNLOCKED:
+			return cca_vdev_device_unlock(pdev);
+		case RHI_DA_TDI_CONFIG_LOCKED:
+			return cca_vdev_device_lock(pdev);
+		case RHI_DA_TDI_CONFIG_RUN:
+			return cca_vdev_device_start(pdev);
+		default:
 			return -EINVAL;
-
-		return cca_vdev_device_start(pdev);
+		}
 	}
 
-aneesh



More information about the linux-arm-kernel mailing list