Monitoring and alerting
What the deployment sends where, which alerts ship, and the two detections you have to add yourself because a deployment cannot deploy them.
Where logs go#
The template routes diagnostics from four resources into a Log Analytics workspace:
| Resource | What is captured | Notes |
|---|---|---|
| Key Vault | AuditEvent |
Every secret and certificate read, including the one the backup runbook makes each run |
| Storage (blob) | Write and delete | Read operations are off by default — see cost below |
| Automation | Job logs and job streams | The same signal the failed-job alert uses |
| Function app | Application and authentication logs | Authentication logs are what the forged-principal alert reads |
By default the template creates a workspace in the resource group named
<prefix>-auto-<hash>-logs, where <hash> is the 13-character uniqueString of the
resource group id that the Automation account name carries (issue #12) —
eidguard-auto-7uh3nvw6e4v7g-logs, for example. You can supply your own instead with Existing Log Analytics
workspace in the wizard, and there is a good reason to: a workspace inside the deployment's
resource group is operated by the same identities that operate the service. Putting the
audit trail somewhere else puts it outside the blast radius of the thing it audits.
All four settings pin logAnalyticsDestinationType: Dedicated, so rows land in
resource-specific tables (AZKVAuditLogs, StorageBlobLogs,
AppServiceAuthenticationLogs, FunctionAppLogs) rather than the legacy shared
AzureDiagnostics table. The shipped alert queries union across both shapes, so a
deployment upgraded from an earlier version whose rows already went to AzureDiagnostics
keeps working.
Cost#
Storage read logging is the largest single cost driver and is off by default. On a backup product every snapshot restore, every compare and every UI listing is a blob read, so on a large estate this table dwarfs everything else. Turn it on with Log storage read operations if you need read-plane forensics and have budgeted for the ingestion.
Retention on a template-created workspace defaults to 30 days. A customer-supplied workspace keeps whatever retention you have already set.
Alerts that ship#
Ten rules, all delivered to the action group that emails the address given at deployment.
Rule names below are the suffix; the deployed name is prefixed with your Automation
account, which is <prefix>-auto-<hash> — so eidguard-auto-7uh3nvw6e4v7g-failed-jobs,
not eidguard-failed-jobs. Severity 1 is the highest.
| Rule | Severity | Fires when |
|---|---|---|
failed-jobs |
1 | An Automation job fails. (Until 2026-09 a completed backup that found an expiring certificate also failed the job to ride this alert; that warning now has its own rule below.) |
drift |
2 | A compare run publishes a non-zero DriftDifferences metric |
certificate-secret-read |
1 | A certificate private key is read from Key Vault by any principal other than the backup automation identity |
mass-blob-deletion |
1 | More than 50 blobs deleted from backup storage in an hour by anything other than the retention lifecycle sweep |
forged-principal-attempts |
2 | More than 25 failed authentication events against the API in an hour — successful sign-ins are not counted |
licence-contact-stale |
2 | This deployment has not reached the licensing service recently. Scheduled backups stop after 60 days without contact; restore, compare and export are never affected |
cert-expiring |
3 | A backup certificate is inside its warning window. Backups still run and upload normally — rotate from the dashboard before it lapses |
cert-expired |
1 | A certificate has expired. If it is the restore certificate, restore and rotation are unavailable until tenant onboarding is re-run — backups can still upload, so the dashboard may look healthy while recovery is not |
worker-window-miss |
2 | A backup or verification run Azure will fire falls outside the private Hybrid Worker VM's start/stop window, so it would run against a deallocated worker — which does not fail a job, it never starts one. Checked on every deployment with a Hybrid Worker, UTC schedules included (each backup run reports a miss still there a minute later; the dashboard checks every few minutes, confirms a miss under the schedule lock and re-aims the window); if it repeats, save the backup schedule. Deployed everywhere, but never fires on a deployment without a Hybrid Worker (no network.hybridWorkerGroup): nothing emits it there |
auth-mode-fallback |
3 | The API authorised on the WEBSITE_AUTH_ENABLED fallback because EIDB_AUTH_MODE is unset |
Switching alerts off under the dashboard's Configuration → Notifications domain disables exactly three
of these — failed-jobs, drift and cert-expiring, the operational ones. The
other six are security controls, and cert-expired and worker-window-miss are
recovery-at-risk signals (an expired restore certificate; backups that silently
never start); they cannot be switched off from the dashboard. A version upgrade keeps the
toggle where you left it (it reads your choice back before it declares the
rules), and keeps the alert address you set on the Alerting page.
Every query is scoped by _ResourceId to this deployment's own resources. That matters
if you supplied a shared workspace: without it, an unrelated Key Vault read elsewhere in
your estate would raise an alert naming this product.
The same toggle also switches every per-tenant backup-failure rule (below) on and off with those three.
Per-tenant recipients#
A tenant can have its own alert recipient. When it does, the dashboard creates two more resources in the deployment's resource group, named by a stable hash of the tenant id (tenant ids are domains and do not fit Azure's naming rules):
- an action group
<automation account>-alerts-<hash>emailing only that recipient; - a metric alert
<automation account>-backup-failed-<hash>, severity 1, taggedeidgTenantId = <tenant id>so you can see in the portal which tenant it belongs to.
The rule watches the custom metric ExternalIDBackup/BackupFailures on the backup
storage account, filtered to that tenant's TenantId dimension (lower-cased on both sides, so the
casing of the id in your configuration does not matter). The scheduled backup
publishes it for each tenant it backs up: the number of resource types that failed
for that tenant, 0 on a clean run, and 1 if the tenant failed outright. The rule
fires when the maximum over its window is above zero.
This is in addition to failed-jobs, never instead of it. The deployment-wide
rule still covers every tenant and still emails the deployment address; a tenant
recipient gets a second, tenant-scoped email on top. Removing a tenant's recipient
deletes both resources (rule first, then action group).
The window is one day (P1D), evaluated hourly. A tenant that failed stays due
but is only retried at the next scheduled run, which can be up to 24 hours later, and
a tenant that is not due publishes nothing. With the one-hour window the other metric
alerts use, this rule would auto-resolve an hour after a failure and email
"resolved" while the tenant was still unprotected. One day is the longest window
Azure Monitor allows. The consequence runs the other way too: a failure keeps the
rule fired for up to a day even if the next run succeeds sooner.
Anything that stops the job before the tenant loop is not per-tenant. No
BackupFailures value is published for anyone in these cases, so no per-tenant
rule fires, and a tenant recipient hears nothing:
- The job cannot authenticate to Azure or cannot resolve its configuration. The job
fails, so the deployment-wide
failed-jobsalert covers it — one reason that rule cannot be replaced by per-tenant recipients. - The licence gate stops scheduled backups after 60 days without contact with the
licensing service. The job ends successfully by design
(
automation/runbook-main.ps1,exit 0afterEIDB_LICENCE_BACKUP_STOPPED), sofailed-jobsdoes not fire either;licence-contact-staleis the alert for it. - Another backup job is already running, so this one stands down. Also a successful
end (
returnafter "Skipping this run"), and not a failure: the running job backs the tenants up and publishes their metric itself.
The dashboard alert toggle governs these rules. A rule is created enabled or
disabled to match the toggle at the time, and switching alerting off or on
afterwards flips every <automation account>-backup-failed-* rule along with
failed-jobs, drift and cert-expiring. The security rules and cert-expired
are unaffected, as above.
How quickly a certificate-read alert reaches you#
Measured end to end on 2026-08-03, twice, from an unauthorised read of a certificate's backing secret through to the notification actually arriving at the action group:
| Stage | Sample 1 | Sample 2 |
|---|---|---|
| Read → audit row queryable in Log Analytics | 132 s | 198 s |
| Read → alert raised | 240 s | 264 s |
| Read → notification delivered | 241 s | 264 s |
So about four minutes in both runs. Treat that as a typical observed latency, not an upper bound. Two things stack on top of it and only one of them is bounded:
- the rule evaluates every 15 minutes, so a row can wait that long to be looked at — bounded
- ingestion has no published SLA. It was 45–314 s across every vault measured here, but nothing guarantees that range, and a new workspace's stream can be down entirely (below)
A rough planning figure is ~20 minutes (observed ingestion plus one full evaluation period). It is an estimate built on measured ingestion, not a worst case Azure commits to — if ingestion runs long, the end-to-end delay runs long with it. Action-group dispatch itself is 2–3 seconds and is never the constraint.
If you are checking whether an alert fired, note that the Azure Monitor alerts API lagged the webhook receiver by 33 and 34 seconds in these runs, so an empty alerts blade in the first half-minute does not mean nothing fired.
That ordering was measured for the webhook leg only — a Logic App recorded the arrival.
The default deployment is email-only, and mailbox delivery was never confirmed here; a
separate action-group test reported the email receiver dispatched at Succeeded within 2
seconds, which is dispatch, not receipt. So do not assume email beats the alerts blade until
someone times the mail path.
The audit stream can start late, and drops what it misses#
certificate-secret-read depends on Key Vault AuditEvent rows reaching Log Analytics. In
some deployments that stream does not start immediately, and events occurring before it starts
are discarded, not delayed — they never appear, at any later point.
Measured 2026-08-03, timed from the diagnostic setting being created:
| Workspace | Region | A read this late was still lost | Stream confirmed live by |
|---|---|---|---|
| A | eastus2 |
+21 min 47 s | +37 min 23 s |
| B | eastus2 |
+3 min 4 s | +35 min 51 s |
| C | eastus |
(nothing lost) | first event, +8 s |
It is not universal, and it is not simply "new workspaces". Workspace C was created the same day, just as new, and captured the read issued 8 seconds after its diagnostic setting existed. What A and B have in common is their region.
That is two workspaces in one region on one afternoon. Earlier centralus runs also saw
zero rows, but they read once and then only queried, so they never showed a stream starting
late and cannot be counted as the same behaviour — only as unexplained zero-row observations.
So there is no established regional rule here, and none of this is advice to pick a
region. Assume it can happen to you and check.
For A and B the exact start moment is unknown: nothing touched either vault between its last lost read and its first captured event, so the start could be anywhere in that gap. Only the bounds are established, and they differ between the two: A was still dropping a read at +21m47s, while B's latest proven loss was only +3m04s, because nothing was read between 3 and 36 minutes. Both were working by ~37. Treat ~22 minutes as a single observation, not a replicated floor.
So a deployment may have an opening window, plausibly 20–40 minutes, in which this alert cannot fire — while still showing as a live rule on the portal blade. You cannot tell which case you are in without checking, which is the point of the query below.
What a certificate read in that window loses is the audit record and the caller identity:
there is no AZKVAuditLogs row, so nothing says who read the key, from where, or whether it
succeeded. The vault's platform metrics do still count the call — ServiceApiHit on the
vault recorded the reads that never reached the table, which is how this was confirmed. So if
you are investigating a suspected read inside that window, that metric is the one piece of
evidence left; it gives you a count and a timestamp, and nothing about the actor.
The delay is in the platform rather than anything this deployment configures: the identical diagnostic setting shape delivered immediately in the region that was unaffected. No way to shorten or predict it has been found.
Verify the stream is live before relying on the control. The check has to keep making new reads, not wait on one:
- a read made before the stream starts is discarded permanently, so polling for that operation can never succeed no matter how long you wait — it would stall, then look like a failure
- a single early check proves nothing either, because rows took 45–314 s to become queryable
- an unscoped count is worthless in a shared workspace: another vault's rows satisfy it
So: read the secret again every few minutes, and poll for any row newer than the last one you know was dropped. Allow at least 45 minutes before concluding — the affected streams here were not confirmed live until ~36–37 minutes.
# Each pass makes a NEW read, so the check cannot be defeated by the first read
# having been discarded. It STOPS as soon as a row appears.
#
# START is fixed once, before the loop. Do NOT use a rolling ago(10m): ingestion
# has no SLA, so a read can become queryable only after it is already older than
# a short window, and the loop would then run forever against a live stream.
START=$(date -u -d '-2 minutes' +%Y-%m-%dT%H:%M:%SZ)
check() {
az monitor log-analytics query -w <workspace-guid> --analytics-query "
AZKVAuditLogs
| where tolower(_ResourceId) == tolower('<key-vault-resource-id>')
| where TimeGenerated > datetime($START)
| project TimeGenerated, OperationName, ResultSignature
| order by TimeGenerated desc" -o tsv
}
# Phase 1 - keep generating probes right up to the 45-minute deadline.
for i in $(seq 1 15); do
az keyvault secret show --vault-name <vault> --name <cert> --query value -o tsv >/dev/null
sleep 180
ROWS=$(check); [ -n "$ROWS" ] && { echo "stream is live:"; echo "$ROWS"; exit 0; }
echo "probe $i: nothing yet"
done
# Phase 2 - one probe ON the deadline, then ingestion grace.
#
# The boundary probe is not optional. Phase 1's last read lands at ~minute 42,
# so a stream starting between then and minute 45 would have NO probe after it
# and could never be detected however long you polled. This read is the one that
# covers that gap; the grace window then just lets it ingest, since observed lag
# reached 314 s and a single 180 s wait would miss it.
az keyvault secret show --vault-name <vault> --name <cert> --query value -o tsv >/dev/null
for i in $(seq 1 5); do
sleep 120
ROWS=$(check); [ -n "$ROWS" ] && { echo "stream is live (late):"; echo "$ROWS"; exit 0; }
echo "grace $i: nothing yet"
done
echo "no rows after 45 min of probes plus 10 min grace - treat as unresolved, not as proof"
The break matters. Every pass reads the certificate's backing secret, which is exactly
what certificate-secret-read alerts on — a loop that keeps going after the answer is in
raises a Sev1 alert every evaluation until somebody stops it. Expect the reads you do make
during this check to alert, and say so in advance if your action group pages anyone.
The first row that appears tells you the stream is live from at least that point. It tells you nothing about whether earlier reads were captured — if they happened before the stream started they are gone, and no amount of waiting will surface them.
Supplying your own established workspace is not a reliable way around this. A workspace already carrying Key Vault logs did capture five newly added vaults from their first minute — but a brand-new workspace elsewhere did the same, so prior use is not what makes the difference. Run the check either way.
Why the retention sweep does not trip mass-blob-deletion#
The lifecycle rule that expires snapshots issues the same DeleteBlob operation an
attacker would, so a daily sweep on a large estate would exceed the 50-blob threshold and
fire the alert on the product doing exactly what you configured it to do. The rule
therefore excludes the sweep — on the one field a real sweep was observed to carry, not
on a guess.
Observed 2026-08-04, by inducing a 60-blob sweep and reading its rows back from
StorageBlobLogs in the same workspace the rule queries:
| Deletion path | AuthenticationType |
RequesterObjectId |
Counted by the rule |
|---|---|---|---|
| Lifecycle sweep (60/60 rows) | TrustedAccessSas |
(empty) | no — excluded |
| AAD data plane (the product path) | OAuth |
populated | yes |
| User-delegation SAS | DelegationSas |
populated | yes |
| Unauthenticated attempt | Anonymous |
(empty) | yes |
AuthenticationType is asserted by the storage service from how the request actually
authenticated. The sweep's value, TrustedAccessSas, is a system-key SAS the platform
issues to itself (sk=system-1 in the logged URI) — it is not something a caller picks.
The exclusion also coalesces authenticationType_s, the plausible legacy
AzureDiagnostics flattening of the same field, for deployments whose rows still take
that route — plausible rather than observed, because no storage row has ever been seen
in AzureDiagnostics here; a wrong spelling leaves rows counted, so the rule fails
noisy, never blind.
The sweep's user-agent header (ObjectLifeCycleScanner) also identifies it, but headers
are client-controlled, so the exclusion deliberately does not use it. Note the empty
RequesterObjectId distinguishes nothing: anonymous attempts carry an empty requester
too, which is why the exclusion keys on the authentication type instead. The deployment
disables shared-key access on the account, so a key-signed SAS never authenticates at
all, and the only SAS a caller can mint logs as DelegationSas — still counted.
Without the exclusion, the shipped query over the observed sweep hour counted
Deletions = 60, past the threshold — the alert would have fired on routine retention.
With it, the same hour counts 0 and a simulated 60-deletion OAuth attack in the same
window still counts 60 and fires. The threshold stays at 50 rather than being raised:
sweep volume scales with your estate, an attack does not have to.
One path is deliberately left outside this alert: something with management-plane
write access on the storage account could rewrite the lifecycle policy itself and have
the platform delete blobs under TrustedAccessSas. That is exactly what the
retention policy changed detection below watches (MANAGEMENTPOLICIES/WRITE in the
Activity Log) — one more reason to add it — and deleted recovery points remain
recoverable through versioning and soft delete.
If this alert fires anyway, the deletions were not the retention sweep. Do not disable the rule.
If the certificate-read alert fires constantly#
It is designed to. The rule extracts the calling principal's object id and alerts when it is anything other than the backup automation identity — and when extraction yields nothing, the empty value does not match, so it alerts. A constantly firing rule therefore means the identity field name no longer resolves against your table schema, not that you are under attack.
Fix the field name; do not mute the rule. It is the compensating control for restore-credential isolation, and a muted alert is not a control.
The two detections you have to add#
Role-assignment changes and storage retention-policy changes are both worth alerting on,
and neither ships. Both live in AzureActivity, and populating that table requires a
diagnostic setting at subscription scope. This template deploys into a single
resource group and its grants are scoped to that group, so it cannot create a
setting at subscription scope.
Earlier versions shipped these two rules anyway. They queried a table that received no data, so they looked like controls on the portal blade and reported nothing, forever. They were deleted rather than left in place — a documented gap is honest, a dead rule is not.
1. Route Activity Log to the workspace#
Once per subscription:
az monitor diagnostic-settings subscription create \
--name eidb-activity \
--location <region> \
--workspace <workspace-resource-id> \
--logs '[{"category":"Administrative","enabled":true},{"category":"Security","enabled":true}]'
2. Add the rules#
Both are scheduled query rules against the same workspace. Substitute the id of the resource group you deployed EIDGuard into, and the id of its backup storage account — both are on that resource group's Overview blade in the portal (Properties → Resource ID, and the storage account's Properties → Resource ID).
Role assignment changed — someone granted or revoked access to the resource group. On a correctly operating deployment this happens at deployment time and never again.
AzureActivity
| where OperationNameValue in~ (
"MICROSOFT.AUTHORIZATION/ROLEASSIGNMENTS/WRITE",
"MICROSOFT.AUTHORIZATION/ROLEASSIGNMENTS/DELETE")
| where ActivityStatusValue =~ "Success"
| where tolower(_ResourceId) startswith tolower("<managed-resource-group-id>")
Retention policy changed — the lifecycle rule that expires snapshots was rewritten outside the product. The API enforces a floor at the soft-delete window; a change made directly against the storage account does not go through it.
AzureActivity
| where OperationNameValue in~ (
"MICROSOFT.STORAGE/STORAGEACCOUNTS/MANAGEMENTPOLICIES/WRITE",
"MICROSOFT.STORAGE/STORAGEACCOUNTS/MANAGEMENTPOLICIES/DELETE")
| where ActivityStatusValue =~ "Success"
| where tolower(_ResourceId) startswith tolower("<storage-account-id>")
Set both to Count > 0, evaluated every 15 minutes over a 1 hour window, pointed at the
same action group the deployment created (<prefix>-auto-<hash>-alerts — the Automation
account name plus -alerts; read it off the resource group rather than composing it).
Both queries are syntactically validated against a live workspace, but the volume and
exact _ResourceId shape of Activity Log rows depends on your subscription. Confirm each
returns rows for a change you make deliberately before relying on it.
