When Azure Managed Identities Fail at Scale: Terraform and 100+ VMs

Reading Time: 6 minutes

The Issue

When deploying Azure infrastructure at scale, you eventually encounter issues that are much easier to hit when working with larger resource counts. I recently ran into exactly such a problem while provisioning Azure Virtual Desktop session hosts with Terraform.

The same Terraform module had been working reliably for our AVD host pools. However, when provisioning larger host pools with more than 100 virtual machines, we started seeing the following error more frequently:

Error: creating Windows Virtual Machine

Subscription: "<SUBSCRIPTION-ID>"
Resource Group Name: "rg-weu-avd-sessionhosts-prod-0001"
Virtual Machine Name: "vm-weu-avd-sessionhost-prod-0093"

polling after CreateOrUpdate: polling failed:
the Azure API returned the following error:

Status: "ManagedServiceIdentityNotFound"
Code: ""

Message:
"Managed service identity referenced with URL
'https://control-westeurope.identity.azure.net/subscriptions/<SUBSCRIPTION-ID>/
resourcegroups/rg-weu-avd-sessionhosts-prod-0001/providers/Microsoft.Compute/
virtualMachines/vm-weu-avd-sessionhost-prod-0093/credentials/v2/identities'
does not exist.

Additional details on error message:
No managed service identities are associated with resource
'/subscriptions/<SUBSCRIPTION-ID>/resourcegroups/
rg-weu-avd-sessionhosts-prod-0001/providers/Microsoft.Compute/
virtualMachines/vm-weu-avd-sessionhost-prod-0093'.

For more information, please follow
https://aka.ms/managed-identity/knownissues"

Terraform reported the failure against the azurerm_windows_virtual_machine resource:

with module.session_host.azurerm_windows_virtual_machine.this[
  "vm-weu-avd-sessionhost-prod-0093"
]

on .terraform/modules/session_host/main.tf line 267,
in resource "azurerm_windows_virtual_machine" "this":

At first glance, this looked like a Terraform or configuration issue. It wasn’t…

The interesting part: scale increases the chance of hitting it

What made the issue particularly interesting was that it wasn’t deterministic. One run could fail while creating:

vm-weu-avd-sessionhost-prod-0006

while another run might fail much later with:

vm-weu-avd-sessionhost-prod-0093

Most VMs were provisioned successfully using exactly the same Terraform module and managed identity configuration. We also saw the issue with smaller host pools, so this isn’t something that only happens once you cross 100 VMs. The difference is simply probability.

With fewer session hosts, there are fewer VM provisioning and managed identity assignment operations, so the chance of encountering the issue during a single Terraform run is lower.

With 100+ session hosts, there are significantly more opportunities for one of those operations to fail.

Conceptually:

Smaller Host Pool
-----------------
VM 001  ✓
VM 002  ✓
VM 003  ✓
...
VM 020  ✓

Fewer operations
→ Lower chance of encountering the issue


Large Host Pool
---------------
VM 001  ✓
VM 002  ✓
...
VM 057  ✗ ManagedServiceIdentityNotFound
...
VM 100  ✓
VM 101  ✓

More operations
→ Higher chance of encountering the issue

This was also an important indication that the Terraform configuration itself was unlikely to be the root cause. If the user-assigned managed identity or its configuration had been incorrect, I would have expected the failure to occur consistently. Instead, most VMs were successfully provisioned while individual VMs occasionally failed.

What is actually happening?

The important part of the Azure error message is:

ManagedServiceIdentityNotFound

No managed service identities are associated with resource '<VM RESOURCE ID>'.

Our virtual machines use user-assigned managed identities. This distinction is important because the managed identity already exists before the virtual machines are provisioned. Azure does not need to create a new Microsoft Entra service principal for every VM. Instead, the existing user-assigned managed identity is associated with each VM during provisioning.

When provisioning a large AVD host pool, many VM creation and managed identity association operations happen within a relatively short period of time. The identity itself exists and works. Most VMs are provisioned successfully using the same user-assigned managed identity, while individual VMs can fail with: ManagedServiceIdentityNotFound

That makes a missing or incorrectly configured user-assigned identity unlikely to be the root cause. If the identity ID or Terraform configuration were incorrect, we would expect the problem to affect the VMs consistently. Instead, the failure is intermittent.

The hint was actually in the error message

After spending time analyzing the provisioning process, I noticed something that, in hindsight, I probably should have followed earlier. Azure literally tells us where to look:

For more information, please follow
https://aka.ms/managed-identity/knownissues

Following that link takes you to Microsoft’s documentation for known issues with managed identities.

There is a dedicated section called: Error during managed identity assignment operations

One of Microsoft’s documented error messages is essentially exactly what we were seeing:

No managed service identities are associated with resource 'azure-resource-id'

So this wasn’t an undocumented Terraform oddity. It is a known Azure Managed Identity issue.

Microsoft documents that managed identity assignment operations can fail in rare cases. For a user-assigned managed identity, Microsoft’s documented workaround is to remove the user-assigned managed identity from the resource and add it again.

That sounds relatively straightforward when working with a single resource manually. In an Infrastructure as Code deployment containing dozens or even hundreds of virtual machines, things become more complicated.

Why does scale matter?

There doesn’t appear to be a magic threshold where VM number 101 suddenly breaks managed identities, and our experience confirms that. We have also encountered the issue with smaller host pools.

The important difference is that larger host pools perform significantly more VM provisioning and managed identity association operations during a single Terraform run. We’re not creating 100+ managed identities.

We’re associating an existing user-assigned managed identity with 100+ virtual machines.

Existing UAMI
     |
     +----> VM 001
     +----> VM 002
     +----> VM 003
     +----> ...
     +----> VM 100+

Most of these operations complete successfully. Occasionally, however, an individual VM provisioning operation fails because Azure reports that no managed identity is associated with that VM. Another VM being provisioned during the same Terraform run, using the same module and the same user-assigned identity, can complete successfully. This points away from a static configuration problem and toward an intermittent Azure-side managed identity assignment issue.

The important point here is:

100 VMs is not a documented Azure Managed Identity limit, nor is it the point at which the issue starts occurring.

We have seen the issue with smaller host pools as well. A host pool containing 100+ VMs simply gives the issue more opportunities to occur during a single provisioning run.

What can we do about it?

Unfortunately, there isn’t a particularly elegant workaround.

1. Retry the failed provisioning

Because the issue is intermittent, a subsequent Terraform run may succeed. However, in practice, “just retry” can be painful. Terraform isn’t necessarily dealing with an operation that cleanly failed before anything was created. Depending on where Azure failed during VM provisioning, parts of the VM and its dependent resources may already exist. A failed session host can involve resources such as:

Session Host VM
   |
   +-- Network Interface
   +-- OS Disk
   +-- Monitoring resources
   +-- Other dependent resources

Depending on the resulting Terraform and Azure state, recovering from the failed deployment can require destroying already provisioned resources for the affected VM before it can be provisioned cleanly again.

That costs time….

And with a host pool containing more than 100 session hosts, having the deployment fail near the end because one VM hits a transient managed identity issue is particularly frustrating. So while retrying is technically a workaround, it isn’t necessarily a cheap one in a large Infrastructure as Code deployment.

2. Reduce Terraform concurrency

If you’re provisioning a large number of VMs simultaneously, reducing Terraform parallelism is worth testing. For example:

terraform apply -parallelism=10

Instead of aggressively increasing parallelism to reduce provisioning time. There is obviously a trade-off here. Lower parallelism means fewer simultaneous operations against Azure, but it also means a longer overall provisioning time. For large AVD host pools, that difference can become significant.

More parallelism isn’t always better, but less parallelism isn’t free either…

3. Provision session hosts in batches

Another potential approach is to deliberately split very large AVD host pools into controlled provisioning batches. For example:

Batch 1: VM 001 - 025
Batch 2: VM 026 - 050
Batch 3: VM 051 - 075
Batch 4: VM 076 - 100
Batch 5: VM 101 - 125

This reduces the number of VM creation and managed identity association operations happening at the same time. Again, the trade-off is provisioning time and additional orchestration complexity. For an automated Terraform deployment, introducing batching solely to work around an intermittent Azure platform issue isn’t necessarily something you want to do unless testing shows that it meaningfully reduces the failure rate.

Don’t immediately blame Terraform

This incident was another good reminder that an Infrastructure as Code deployment can fail even when the Infrastructure as Code is perfectly valid. When Terraform returns something like: ManagedServiceIdentityNotFound

The first instinct is often to investigate:

Terraform module
      ↓
AzureRM provider
      ↓
resource configuration
      ↓
dependencies

And those are absolutely worth checking. But sometimes the actual path looks more like this:

Terraform
    ↓
AzureRM Provider
    ↓
Azure Compute API
    ↓
Managed Identity infrastructure
    ↓
intermittent assignment problem

The fact that dozens of other VMs are successfully provisioned using the exact same Terraform configuration and user-assigned managed identity is an important troubleshooting signal.

Lessons learned

There are a few things I took away from this issue.

First, always read the complete Azure error message. Azure literally included the managed identity known-issues documentation in the error. That was a pretty strong hint.

Second, don’t confuse correlation with a hard platform limit. We initially noticed the issue much more prominently with our 100+ VM host pools. But the issue can also occur with smaller host pools. Larger host pools simply increase the number of operations and therefore the probability of encountering an intermittent failure.

Third, test infrastructure at realistic scale. A Terraform module successfully provisioning 10 or 20 session hosts doesn’t necessarily tell you how often you’ll encounter transient platform issues when provisioning 100 or more.

Fourth, retry isn’t always cheap with Infrastructure as Code. A transient Azure error might last only seconds, but recovering from it can take significantly longer when Terraform has already provisioned parts of the affected VM and dependent resources need to be cleaned up before another attempt.

And finally, don’t immediately assume Terraform is the problem.

If 99 VMs successfully use the same module, the same identity, and the same configuration while one VM fails with a documented Azure Managed Identity issue, that’s a strong indication to look beyond your Terraform code. Sometimes Infrastructure as Code at scale doesn’t expose a bug in your code. It exposes the rare platform issue that you simply weren’t hitting often enough before.

I really hope Microsoft addresses this issue in the future. Managed Identity assignment is a fundamental Azure capability and shouldn’t cause this kind of headache when provisioning resources at scale. I opened a Microsoft support case to investigate the issue further, but as of October 6, 2026, there is unfortunately no ETA for a fix.

References

Microsoft Learn – Known issues with managed identities for Azure resources
https://learn.microsoft.com/en-us/entra/identity/managed-identities-azure-resources/known-issues#error-during-managed-identity-assignment-operations

Leave a Reply

Your email address will not be published. Required fields are marked *