Deployment troubleshooting
Failures seen against real deployments, what they mean, and how to get past them.
"is a trademarked or reserved word"#
The resource name 'xbox5fafitglmazscripts' or a part of the name is a
trademarked or reserved word. (Code:DeploymentScriptOperationFailed)
Fix: redeploy, setting Deployment retry suffix to any short value — 2, b,
retry. On the CLI, pass deploymentScriptNameSuffix.
Why this happens#
The deployment runs four short-lived setup tasks (Azure calls them deployment scripts).
Each one provisions its own temporary working storage account, and Azure generates that
account's name, not us — a random-looking fragment with azscripts appended. Very
occasionally the generated fragment contains a word Azure refuses in a storage account
name. It rejects the name, and the whole deployment fails.
Nothing you configured caused it, and the resource it names is one you never asked for.
Why retrying unchanged does not help#
The generated name is deterministic. Measured on 2026-08-01 by deploying two probe scripts, deleting the resources outright, and redeploying:
| Script | Run 1 | Run 2, after deletion |
|---|---|---|
probe-alpha |
jjyswwuliwxpoazscripts |
jjyswwuliwxpoazscripts |
probe-beta-differently-named |
st2h6icyqt4yyazscripts |
st2h6icyqt4yyazscripts |
Same inputs, same name, every time. So a plain retry reproduces the failure exactly.
The inputs are the script name and the resource group. Measured the same day by deploying one identically-named script into two resource groups:
| Resource group | Generated account |
|---|---|
eidprobe-a |
ajz6lxwzt5qo2azscripts |
eidprobe-b |
qc57udtiebybyazscripts |
Because the resource group is an input, this is a per-deployment lottery rather than something every customer hits — and because the script name is also an input, changing it is a reliable way out. That is all the suffix does: it appends your value to the five script names, which produces five different generated storage names.
Which tasks are affected#
Five, each an independent chance per deployment:
| Task | What it does |
|---|---|
stage-functionapp-package |
Uploads the web application package (runs before the Function App is deployed) |
read-dashboard-settings |
Reads the settings you chose in the dashboard (schedule, retention, alerting, verification), so an upgrade keeps them |
preserve-wizard-auth |
Re-applies your sign-in settings during an upgrade |
revoke-legacy-role-assignments |
Removes over-broad permissions from older versions |
prune-stale-schedules |
Removes backup schedules that no longer apply |
Three of those matter only on upgrade, so this can hit an existing deployment as readily as a new one.
read-dashboard-settings has one failure of its own worth knowing about. If your deployment
uses private networking, the data storage account has public access disabled, and this task
cannot read your settings — so it stops the upgrade rather than guess, because guessing
would put the deployment's original settings back over the ones you chose.
Re-enable public network access on the data storage account for the duration of the upgrade
and retry; see private-endpoints.md.
If it happens again with a suffix set#
Change the suffix to a different value. Each distinct value produces a different set of generated names.
Why we do not just supply our own storage account#
Deployment scripts can be pointed at a storage account you provide, which would remove this failure mode entirely. We do not, and the reason is a deliberate trade:
- Azure can only mount that account into the script container using an access key, so
the account would need
allowSharedKeyAccess = true. This product is otherwise entirely keyless — Entra identity everywhere, no account keys. - It would need public network access in the default topology, because the container mounts the file share over the public endpoint.
- It would persist between deployments. The per-run file share is cleaned up automatically; the account is not, and cannot delete itself mid-deployment.
- It would conflict with restricting network access, so hardening the deployment would break its own upgrades.
That is a permanent, publicly reachable, shared-key storage account working against a security control — traded against a failure that has a one-field workaround. The account would hold no backups and sit empty between runs, so its own exposure is small, but the trade is still poor.
If Azure adds managed-identity mounting (Azure/bicep#17093), this becomes worth revisiting.
The Key Vault name is still held by a deleted vault#
The vault name '<prefix>-kv-<hash>' is already in use. (Code:VaultAlreadyExists)
The wizard now catches this before you deploy: the Basics page lists soft-deleted vaults in your subscription and warns if one already uses your name prefix. If you get here from the ARM error instead, the wizard message is the same advice.
Why the name is held#
Deleting a Key Vault does not release its name. The vault goes soft-deleted for 90 days, and this deployment turns on purge protection, which means it cannot be purged early — by you, by us, or by support. Confirmed on 2026-08-02 against a torn-down rehearsal stack:
az keyvault purge --name eidrh1736-kv-324jiterpsr
ERROR: (MethodNotAllowed) Operation 'DeletedVaultPurge' is not allowed.
That is the control working as intended: purge protection is what stops an attacker
destroying your backup certificates by deleting the vault. The operational cost is that
the name is reserved until scheduledPurgeDate.
Key Vault names are globally unique, so the name is held in every region, not just the one the vault was deleted from.
What to do#
az keyvault list-deleted -o table shows what is held, in which region, and until when.
Add --query "[].{name:name,loc:properties.location,from:properties.vaultId}" to see which
resource group each one came from — that is the field that decides the remedy below.
Changing the Resource name prefix always works. It is the answer unless you specifically want the old certificates back.
Recovery is narrower, and getting it wrong turns a working deployment into a failed one:
| Situation | Remedy |
|---|---|
| You are redeploying the same stack — same resource group name, same subscription, same region — and want its certificates back | Tick Recover a soft-deleted Key Vault of the same name on the Backup configuration page (recoverKeyVault on the CLI). The vault comes back intact, so onboarded tenants keep working. |
| The deleted vault merely shares your prefix — it came from a different resource group | Leave recovery off. Your deployment generates a different vault name, so there is nothing to collide with and nothing to recover; ticking the box would fail a deployment that was going to succeed. |
| The name matches but the vault is in a different region | Change the prefix. Switching your deployment to that region does not on its own fix anything — the name is held globally, so createMode: default hits the same conflict there. Getting the certificates back takes deploying into that region and ticking recovery, together. |
The distinction is the hash. Recovery needs the vault name to match exactly, and the
name is <prefix>-kv-<hash> where the hash comes from the resource group id — so
"same prefix" is not the same thing as "same vault". The wizard warns on the prefix
because that is all it can see (below); deciding between these rows is yours, and
properties.vaultId is how you decide it.
The different-region case is the one worth knowing about, because it is reached by an ordinary reaction: a deployment fails on regional capacity, so the operator retries somewhere else, and the vault name follows them there while the recovery option does not.
What the wizard check cannot see#
Two limits, both deliberate, because neither has a fix available inside a createUiDefinition:
- It matches on the prefix, not the whole name. The vault name is
<prefix>-kv-<hash>and the hash isuniqueString(resourceGroup().id), which the wizard has no way to compute — there is nouniqueStringfunction, and the resource group does not exist yet. So the check can tell you a collision is possible, not that it is certain. - It reads one page of results.
Microsoft.KeyVault/deletedVaultspaginates with anextLinkand accepts no$topor$filter, andMicrosoft.Solutions.ArmApiControlissues exactly one request with no way to follow a continuation. If your subscription holds enough deleted vaults to page, the wizard says so and asks you to runaz keyvault list-deleted -o tableyourself, rather than reporting a clean check it did not actually perform.
In both cases the fallback is the behaviour that existed before: ARM raises the name conflict. The check makes the common case visible early; it is not a guarantee.
Why a plain redeploy usually works anyway#
The vault name is <prefix>-kv-<hash>, where the hash derives from the resource
group id. A new deployment into a new resource group produces a different hash
and therefore a different vault name. The collision needs the same prefix and the same
resource group name in the same subscription — which is precisely what a retry
after a failed deployment tends to be.
Observed 2026-08-04: the same-resource-id retry recovers the vault by itself#
Measured against the template's own compiled keyvault.bicep (api-version 2024-11-01,
westus2), with a vault soft-deleted from resource group eidretry5-rg and that group
recreated under the same name:
- A
createMode: defaultPUT to the same resource id — same subscription, same resource group name, same vault name — succeeded and implicitly recovered the soft-deleted vault: the deployment completed, the soft-deleted entry disappeared, and the vault came back with its originalsystemData.createdAt. NorecoverKeyVaultneeded. - The same PUT from a different resource group failed with the verbatim
VaultAlreadyExistserror above. So the implicit recovery does not extend past a matching resource id, and the name hold is real from any other resource group. A cross-region retry that keeps the same subscription, resource-group name and vault name shares the resource id (region is not part of an ARM id) and was not tested — it may recover, conflict, or fail on a location mismatch; the decision table above remains the guidance for that case. createMode: recover(recoverKeyVault) against the soft-deleted vault also succeeded, recovering it intact — the first live observation of that path.
One observation each, one region, one API version — treat the implicit recovery as how that API version behaved on that day, not a contract. The operative advice above is unchanged (recovery ticked when you want the certificates back is now proven to work); what softens is the failure mode: the exact-match retry this section warns about came back on its own rather than failing.
A restarted Function App keeps running the OLD package#
Symptom: you replace the package blob in the app container, restart the Function App,
and the old code keeps serving. No error anywhere.
WEBSITE_RUN_FROM_PACKAGE points at a blob URL, and the platform does not re-download
on restart alone — the URL is treated as the cache key. Overwriting the blob leaves the URL
identical, so the running app keeps what it already has.
The marketplace template cache-busts by design (issue #49). Each package build stamps
its own identity into mainTemplate, and both halves of the deployment derive the blob
name from it: the staging script uploads functionapp-<stamp>.zip, and the Function App's
WEBSITE_RUN_FROM_PACKAGE names that same blob. A new build is a new name, so it is a new
URL, so ARM writes a changed app setting and the platform downloads. Upgrading a
marketplace deployment therefore delivers the new API and SPA without any manual step.
Two consequences worth knowing:
- Staging runs before the site is deployed. The URL must never be published ahead of the blob it names — a host that recycles onto a missing package fails to boot, and that is not something a later restart reliably repairs.
- Old package blobs are kept. They are small next to the storage they sit in, and each
is a rollback point: point
WEBSITE_RUN_FROM_PACKAGEback at an earlierfunctionapp-<stamp>.zipand restart. Nothing prunes them automatically, because pruning from a deployment path means deleting the running app's own package if an upgrade only partly succeeds. Delete them by hand if they ever matter.
The manual remedy still applies to a hand-built stack, a deployment from before #49, or any time you overwrite a package blob yourself: change the URL, then restart. A query string is enough:
az functionapp config appsettings set -g <rg> -n <app> --settings "WEBSITE_RUN_FROM_PACKAGE=https://<account>.blob.core.windows.net/app/<blob>.zip?v=$(date +%H%M%S)"
az functionapp restart -g <rg> -n <app>
This matters most when verifying a code change against a live stack. Hit on
2026-08-01 while proving the EIDB_AUTH_DENIED logging: the first restart silently kept
the old package, the marker never appeared in FunctionAppLogs, and the natural reading
was "the fix does not work". It did — the fix was never running.
If a live check of new code produces no evidence, confirm the package actually reloaded
before concluding anything about the code. FunctionAppLogs is a good tell: user output
lands in a Function.<name>.User category, so if that category does not exist for the
function you changed, your code is not running.
Every request 503s or times out, while the app reports Running#
Symptom: a freshly deployed (marketplace/Linux) Function App answers no request —
/ and /api/* alike hang for ~100–230 s and then return 503, or the client gives up
first — while az functionapp show reports state: Running, availabilityState: Normal, and the deployment itself succeeded with no ARM failure.
This is not the setup-pending lockdown. Deny-by-default is enforced by application
code (Assert-ApiRole → 401/403), so it requires the app to be serving: an auth denial
comes back in milliseconds with a status the app chose. An indefinite hang followed by
a 503 means no function execution ever completed — the 503 is the platform giving up,
not the app answering.
Root cause, found live on eiddet803-rg (issue #20): the payload's host.json shipped
with managedDependency.enabled: true. Managed dependencies are a Windows-only feature
of the PowerShell worker, but the Linux worker still honours the flag — it starts a
PSGallery download that never completes there, and holds every function execution
behind it. The package build vendors the Az modules precisely because managed
dependencies do not exist on Linux, so the download was pure downside. Fixed by
Disable-PackagedManagedDependency in modules/WebPackaging.psm1: a vendored package
now ships the flag disabled. (The source web/api/host.json keeps it enabled — an
unvendored payload carries no Az modules, so the managed-dependency download is its
only source of them. Only the vendored zip is patched.)
The tells, all in FunctionAppLogs (measured over 9 hours on the affected stack):
Executing Functions.<name>rows with zero matchingExecutedrows — invocations start and never finish. On the affected stack: 113Executing, 0Executed, ever.- The worker says so outright, level Warning:
The first managed dependency download is in progress, function execution will continue when its done.— still calling itself "first" 9 hours after deployment, because each ~35-minute host recycle starts it over. - The
KeepWarmtimer fires withUnscheduledInvocationReason: IsPastDue, OriginalSchedule: <deployment time>— it has never once completed since deployment. - No
ErrororCriticalrows at all. The host is healthy; the executions are merely parked forever.
To confirm on a live stack:
FunctionAppLogs
| where Message startswith 'Executed '
| count // 0 = nothing has ever finished
Remediating an already-deployed stack: patch host.json inside the staged package blob
(set managedDependency.enabled to false), upload it under a new URL, point
WEBSITE_RUN_FROM_PACKAGE at it and restart — the URL is the cache key, so overwriting
the blob in place changes nothing (see the previous section).
Onboarding a tenant returns 403 tenant_limit_reached#
403 { "code": "tenant_limit_reached",
"error": "Your Starter plan protects up to 3 tenants. Offboard a tenant,
or upgrade your plan in the Azure portal, to onboard another." }
What it means: this is a marketplace solution-template deployment, and the number of onboarded External ID tenants has reached the number the current plan includes (Starter 3, Standard 5, Enterprise unlimited). It is not an error in the deployment and nothing is broken.
What is unaffected — all of it. Every tenant already onboarded keeps backing up on schedule, keeps its recovery points, and stays restorable, comparable and exportable. That holds when a plan downgrade leaves the deployment over the limit, too: "5 of 3" backs up all five. Only the onboarding commit is refused.
Re-onboarding a tenant that is already in the list is never refused. The commit is an upsert keyed on the normalized tenant id, so the count does not change and the check returns clean — including over the limit. Repairing a tenant that already appears on the dashboard (deleted app registration, re-issued certificate) is therefore always allowed, at any tenant count.
The exemption is membership, not intent. It applies only when the tenant id is already in the deployment's tenant list. A wizard that never reached its final commit never added the tenant, so at the cap the retry is refused exactly like a brand-new one — even though the earlier attempt may already have created app registrations and certificates in the target tenant. The rule of thumb: if the tenant is not on your dashboard, the exemption does not apply. Free a slot or change plan first, then re-run the wizard.
Two ways a run can fail to commit, and they leave different debris:
- Refused at the cap. The wizard reverses its own work on the spot — it deletes the app registrations it created and revokes the Graph roles it granted, using the delegated token it still holds — then lists anything it could not remove. Read that list: an app that already existed is adopted, never deleted, and a certificate attached to a surviving app is a live credential, so the screen tells you to revoke it promptly. Nothing is stranded silently.
- Interrupted (browser closed, session lost) before the commit. No rollback runs, so whatever the flow had created stays. Re-running onboarding once you have capacity adopts those objects rather than duplicating them.
Resolve it by either:
- changing the plan on your EIDGuard SaaS subscription in the Azure portal (Marketplace → your subscription → change plan). The deployment picks the new limit up on its next check-in with the licensing service — within six hours, usually sooner — with nothing redeployed. Redeploying the solution template does not change the limit, and deploying under a different resource name prefix builds a second, separately billed estate while leaving this one blocked; or
- offboarding a tenant you no longer need to protect, which frees a slot immediately.
Confirming it server-side: the API writes a stable marker on every denial —
EIDB_QUOTA_DENIED onboarding blocked at plan limit (N) — in the
Function.TenantCommit.User category of FunctionAppLogs. Query for the marker
string rather than the prose. No alert is wired to it.
A garbled or negative EIDB_TENANT_LIMIT cannot produce this response — the
parse fails open to unlimited by design, so a mangled setting never bricks
onboarding. See plans-and-limits.md.
Export refused with 409 WorkerOffline#
409 { "code": "export_start_failed",
"error": "WorkerOffline: this deployment runs jobs on a private hybrid
worker that is currently powered off, so an export started now
would queue and never run. …" }
Only private-networking deployments see this. With private networking, every job runs on the Hybrid Runbook Worker VM inside the VNet — the Automation cloud sandbox cannot reach private-only Key Vault and storage. That VM is powered on only around each scheduled backup run (auto start/stop, which is most of what keeps the private option affordable; see private-endpoints.md), so outside those windows there is nothing to run the job.
Why it fails fast instead of queueing. Automation would happily accept the
start and hand back a job id; the job would then sit queued until the worker
came up — indefinitely, if auto start/stop was disabled or the VM was stopped
deliberately. For an evacuate-before-you-delete flow that is the worst possible
outcome: a job id that looks like a successful export and is not. The API
therefore probes the worker's lastSeenDateTime heartbeat first and refuses on a
positive determination only — an undetermined probe never blocks.
Fix: start the export during a backup window (from 15 minutes before a scheduled run until 2 hours after), or power the worker VM on first and retry. The probe needs both hybrid-worker read actions (see permissions.md); if the probe cannot read them it returns undetermined and the export proceeds unguarded, which looks like this section not applying.
Raising deletion protection fails right after an upgrade#
The Raise-ExternalID-Immutability job reports not raised with an
authorization error (AuthorizationFailed on reading the policy, its update
history, the lifecycle policy or a job record, or on
immutabilityPolicies/extend/action) shortly after an upgrade.
Fix: wait a few minutes and raise again. The upgrade that adds the
EIDGuard Immutability Extender role and the two Reader grants assigns them to
the Automation account's identity, and a new role assignment can take a few
minutes to apply. Nothing was changed — the job checks everything it can before
it extends anything — and a raise that changed nothing does not count towards
the one-raise-a-day limit, which is counted from the policy's own record of
extensions.
If the job instead says the policy has already been extended as many times as Azure allows, that is permanent: Azure lets a locked policy be extended only a limited number of times, and the window stays where it is.
Custom role definitions "left behind" after a teardown#
The deployment creates five CustomRole definitions whose only assignable scope is the
resource group: EIDGuard Retention Manager, EIDGuard Automation Operator,
EIDGuard Upgrade Cleanup and EIDGuard Immutability Extender
(marketplace/bicep/modules/roles.bicep), and
EIDGuard Immutability Locker (marketplace/bicep/modules/stagingroles.bicep), each
suffixed with uniqueString(resourceGroup().id). The cleanup below matches on the
EIDGuard prefix, so it finds all five.
Deleting the resource group removes them, and did in every teardown measured: on 2026-08-02 (a minimal probe, succeeded and deliberately failed, and two full-stack deliberately failed deployments) and again on 2026-09-13, when two solution-template stacks whose groups had held only the data storage account for two days were deleted — their six remaining definitions were gone in four of five samples taken over the following eighty seconds, and the fifth sample returned all of them, which is the read-plane trap described next.
Before concluding they were stranded, read twice. The RBAC read plane is eventually consistent, and it lies in the direction that looks alarming:
- 6 of 150 GETs of six provably deleted definitions returned HTTP 200 with the complete object.
- 1 of 15
Get-AzRoleDefinition -Customcalls returned all six at once — which is exactly what "three definitions are still there" looks like. - In a full-stack teardown the list showed all three at T+90s while a GET-by-id of those same three ids returned 404 in the same second.
Two more traps worth knowing:
DELETEon a role definition that is already gone returns 204, so a successful delete is not proof there was anything to delete.Remove-AzRoleDefinition -IdthrowsNotFoundand is the honest probe.- The cascade's deletion of role definitions is not written to the activity log (assignments are). Absence there means nothing.
If a definition really is present in every sample, clear it with the Azure CLI (Owner or User Access Administrator on the subscription). Nothing here needs anything from this repository. Match on the suffix in the role name, which is the deployment's own: a second EIDGuard deployment in the same subscription has its own five, and they must stay.
# 1. Find the five definitions whose assignable scope is the deleted group.
az role definition list --custom-role-only true \
--query "[?starts_with(roleName,'EIDGuard ')].{roleName:roleName, id:name, scopes:assignableScopes}" -o table
# 2. A definition cannot be deleted while an assignment references it. Any
# assignment listed here is an orphan of the deleted group; remove it.
az role assignment list --all --role "<roleName>" --query "[].id" -o tsv \
| xargs -r -n1 az role assignment delete --ids
# 3. Delete the definition. A delete of one that is already gone returns
# success too, so re-run step 1 a few times, a minute apart, to confirm.
az role definition delete --name "<roleName>"
While the group still exists, the definitions cannot be deleted by hand — each has
a live assignment (the deployment identities on the storage account, the container and
the group) and Remove-AzRoleDefinition answers BadRequest until those go. They go
with the group. So the order is: finish deleting the group (see the next section for the
one thing that stops it), then check the definitions, then clean up only what is still
listed in every sample.
Deleting the resource group fails with AccountProtectedFromDeletion#
eidg1sa7yveiftimhs6y is protected from deletion. The container(s) backups have object
level immutability enabled.
The group delete removed everything except the data storage account. Azure refuses to
delete an account while its immutability-enabled backups container holds any blob or
blob version — including ones whose immutable window has expired. This is the
protection for recovery points, and it does not distinguish an uninstall from an
attack.
Fix: empty the container, then delete the group again. The full order — export, offboard, wait out the window, purge, delete — is in plans-and-limits.md, with the commands. The two things that trip people up, both measured on 2026-09-13:
- The purge needs Storage Blob Data Owner on the account. Storage Blob Data
Contributor deletes the current blobs, then is refused on every expired previous
version with
OperationNotAllowedOnAutomaticSnapshot("The specified operation is not allowed on version"). - Deleting a current blob leaves its previous version in place, so one pass over the listing is never enough: list with versions, delete, list again, until empty. A version still inside its window is refused with a 409, and that is correct.
Once the container is empty it deletes normally despite its locked policy, and the account and group follow.
