Skip to content
PagesMonitoring and alerting

EIDGuard docs

On this page

Monitoring and alerting

What the deployment sends where, which alerts ship, and the two detections you have to add yourself because a deployment cannot deploy them.

Where logs go#

The template routes diagnostics from four resources into a Log Analytics workspace:

Resource What is captured Notes
Key Vault AuditEvent Every secret and certificate read, including the one the backup runbook makes each run
Storage (blob) Write and delete Read operations are off by default — see cost below
Automation Job logs and job streams The same signal the failed-job alert uses
Function app Application and authentication logs Authentication logs are what the forged-principal alert reads

By default the template creates a workspace in the resource group named <prefix>-auto-<hash>-logs, where <hash> is the 13-character uniqueString of the resource group id that the Automation account name carries (issue #12) — eidguard-auto-7uh3nvw6e4v7g-logs, for example. You can supply your own instead with Existing Log Analytics workspace in the wizard, and there is a good reason to: a workspace inside the deployment's resource group is operated by the same identities that operate the service. Putting the audit trail somewhere else puts it outside the blast radius of the thing it audits.

All four settings pin logAnalyticsDestinationType: Dedicated, so rows land in resource-specific tables (AZKVAuditLogs, StorageBlobLogs, AppServiceAuthenticationLogs, FunctionAppLogs) rather than the legacy shared AzureDiagnostics table. The shipped alert queries union across both shapes, so a deployment upgraded from an earlier version whose rows already went to AzureDiagnostics keeps working.

Cost#

Storage read logging is the largest single cost driver and is off by default. On a backup product every snapshot restore, every compare and every UI listing is a blob read, so on a large estate this table dwarfs everything else. Turn it on with Log storage read operations if you need read-plane forensics and have budgeted for the ingestion.

Retention on a template-created workspace defaults to 30 days. A customer-supplied workspace keeps whatever retention you have already set.

Alerts that ship#

Ten rules, all delivered to the action group that emails the address given at deployment.

Rule names below are the suffix; the deployed name is prefixed with your Automation account, which is <prefix>-auto-<hash> — so eidguard-auto-7uh3nvw6e4v7g-failed-jobs, not eidguard-failed-jobs. Severity 1 is the highest.

Rule Severity Fires when
failed-jobs 1 An Automation job fails. (Until 2026-09 a completed backup that found an expiring certificate also failed the job to ride this alert; that warning now has its own rule below.)
drift 2 A compare run publishes a non-zero DriftDifferences metric
certificate-secret-read 1 A certificate private key is read from Key Vault by any principal other than the backup automation identity
mass-blob-deletion 1 More than 50 blobs deleted from backup storage in an hour by anything other than the retention lifecycle sweep
forged-principal-attempts 2 More than 25 failed authentication events against the API in an hour — successful sign-ins are not counted
licence-contact-stale 2 This deployment has not reached the licensing service recently. Scheduled backups stop after 60 days without contact; restore, compare and export are never affected
cert-expiring 3 A backup certificate is inside its warning window. Backups still run and upload normally — rotate from the dashboard before it lapses
cert-expired 1 A certificate has expired. If it is the restore certificate, restore and rotation are unavailable until tenant onboarding is re-run — backups can still upload, so the dashboard may look healthy while recovery is not
worker-window-miss 2 A backup or verification run Azure will fire falls outside the private Hybrid Worker VM's start/stop window, so it would run against a deallocated worker — which does not fail a job, it never starts one. Checked on every deployment with a Hybrid Worker, UTC schedules included (each backup run reports a miss still there a minute later; the dashboard checks every few minutes, confirms a miss under the schedule lock and re-aims the window); if it repeats, save the backup schedule. Deployed everywhere, but never fires on a deployment without a Hybrid Worker (no network.hybridWorkerGroup): nothing emits it there
auth-mode-fallback 3 The API authorised on the WEBSITE_AUTH_ENABLED fallback because EIDB_AUTH_MODE is unset

Switching alerts off under the dashboard's Configuration → Notifications domain disables exactly three of these — failed-jobs, drift and cert-expiring, the operational ones. The other six are security controls, and cert-expired and worker-window-miss are recovery-at-risk signals (an expired restore certificate; backups that silently never start); they cannot be switched off from the dashboard. A version upgrade keeps the toggle where you left it (it reads your choice back before it declares the rules), and keeps the alert address you set on the Alerting page.

Every query is scoped by _ResourceId to this deployment's own resources. That matters if you supplied a shared workspace: without it, an unrelated Key Vault read elsewhere in your estate would raise an alert naming this product.

The same toggle also switches every per-tenant backup-failure rule (below) on and off with those three.

Per-tenant recipients#

A tenant can have its own alert recipient. When it does, the dashboard creates two more resources in the deployment's resource group, named by a stable hash of the tenant id (tenant ids are domains and do not fit Azure's naming rules):

  • an action group <automation account>-alerts-<hash> emailing only that recipient;
  • a metric alert <automation account>-backup-failed-<hash>, severity 1, tagged eidgTenantId = <tenant id> so you can see in the portal which tenant it belongs to.

The rule watches the custom metric ExternalIDBackup/BackupFailures on the backup storage account, filtered to that tenant's TenantId dimension (lower-cased on both sides, so the casing of the id in your configuration does not matter). The scheduled backup publishes it for each tenant it backs up: the number of resource types that failed for that tenant, 0 on a clean run, and 1 if the tenant failed outright. The rule fires when the maximum over its window is above zero.

This is in addition to failed-jobs, never instead of it. The deployment-wide rule still covers every tenant and still emails the deployment address; a tenant recipient gets a second, tenant-scoped email on top. Removing a tenant's recipient deletes both resources (rule first, then action group).

The window is one day (P1D), evaluated hourly. A tenant that failed stays due but is only retried at the next scheduled run, which can be up to 24 hours later, and a tenant that is not due publishes nothing. With the one-hour window the other metric alerts use, this rule would auto-resolve an hour after a failure and email "resolved" while the tenant was still unprotected. One day is the longest window Azure Monitor allows. The consequence runs the other way too: a failure keeps the rule fired for up to a day even if the next run succeeds sooner.

Anything that stops the job before the tenant loop is not per-tenant. No BackupFailures value is published for anyone in these cases, so no per-tenant rule fires, and a tenant recipient hears nothing:

  • The job cannot authenticate to Azure or cannot resolve its configuration. The job fails, so the deployment-wide failed-jobs alert covers it — one reason that rule cannot be replaced by per-tenant recipients.
  • The licence gate stops scheduled backups after 60 days without contact with the licensing service. The job ends successfully by design (automation/runbook-main.ps1, exit 0 after EIDB_LICENCE_BACKUP_STOPPED), so failed-jobs does not fire either; licence-contact-stale is the alert for it.
  • Another backup job is already running, so this one stands down. Also a successful end (return after "Skipping this run"), and not a failure: the running job backs the tenants up and publishes their metric itself.

The dashboard alert toggle governs these rules. A rule is created enabled or disabled to match the toggle at the time, and switching alerting off or on afterwards flips every <automation account>-backup-failed-* rule along with failed-jobs, drift and cert-expiring. The security rules and cert-expired are unaffected, as above.

How quickly a certificate-read alert reaches you#

Measured end to end on 2026-08-03, twice, from an unauthorised read of a certificate's backing secret through to the notification actually arriving at the action group:

Stage Sample 1 Sample 2
Read → audit row queryable in Log Analytics 132 s 198 s
Read → alert raised 240 s 264 s
Read → notification delivered 241 s 264 s

So about four minutes in both runs. Treat that as a typical observed latency, not an upper bound. Two things stack on top of it and only one of them is bounded:

  • the rule evaluates every 15 minutes, so a row can wait that long to be looked at — bounded
  • ingestion has no published SLA. It was 45–314 s across every vault measured here, but nothing guarantees that range, and a new workspace's stream can be down entirely (below)

A rough planning figure is ~20 minutes (observed ingestion plus one full evaluation period). It is an estimate built on measured ingestion, not a worst case Azure commits to — if ingestion runs long, the end-to-end delay runs long with it. Action-group dispatch itself is 2–3 seconds and is never the constraint.

If you are checking whether an alert fired, note that the Azure Monitor alerts API lagged the webhook receiver by 33 and 34 seconds in these runs, so an empty alerts blade in the first half-minute does not mean nothing fired.

That ordering was measured for the webhook leg only — a Logic App recorded the arrival. The default deployment is email-only, and mailbox delivery was never confirmed here; a separate action-group test reported the email receiver dispatched at Succeeded within 2 seconds, which is dispatch, not receipt. So do not assume email beats the alerts blade until someone times the mail path.

The audit stream can start late, and drops what it misses#

certificate-secret-read depends on Key Vault AuditEvent rows reaching Log Analytics. In some deployments that stream does not start immediately, and events occurring before it starts are discarded, not delayed — they never appear, at any later point.

Measured 2026-08-03, timed from the diagnostic setting being created:

Workspace Region A read this late was still lost Stream confirmed live by
A eastus2 +21 min 47 s +37 min 23 s
B eastus2 +3 min 4 s +35 min 51 s
C eastus (nothing lost) first event, +8 s

It is not universal, and it is not simply "new workspaces". Workspace C was created the same day, just as new, and captured the read issued 8 seconds after its diagnostic setting existed. What A and B have in common is their region.

That is two workspaces in one region on one afternoon. Earlier centralus runs also saw zero rows, but they read once and then only queried, so they never showed a stream starting late and cannot be counted as the same behaviour — only as unexplained zero-row observations. So there is no established regional rule here, and none of this is advice to pick a region. Assume it can happen to you and check.

For A and B the exact start moment is unknown: nothing touched either vault between its last lost read and its first captured event, so the start could be anywhere in that gap. Only the bounds are established, and they differ between the two: A was still dropping a read at +21m47s, while B's latest proven loss was only +3m04s, because nothing was read between 3 and 36 minutes. Both were working by ~37. Treat ~22 minutes as a single observation, not a replicated floor.

So a deployment may have an opening window, plausibly 20–40 minutes, in which this alert cannot fire — while still showing as a live rule on the portal blade. You cannot tell which case you are in without checking, which is the point of the query below.

What a certificate read in that window loses is the audit record and the caller identity: there is no AZKVAuditLogs row, so nothing says who read the key, from where, or whether it succeeded. The vault's platform metrics do still count the call — ServiceApiHit on the vault recorded the reads that never reached the table, which is how this was confirmed. So if you are investigating a suspected read inside that window, that metric is the one piece of evidence left; it gives you a count and a timestamp, and nothing about the actor.

The delay is in the platform rather than anything this deployment configures: the identical diagnostic setting shape delivered immediately in the region that was unaffected. No way to shorten or predict it has been found.

Verify the stream is live before relying on the control. The check has to keep making new reads, not wait on one:

  • a read made before the stream starts is discarded permanently, so polling for that operation can never succeed no matter how long you wait — it would stall, then look like a failure
  • a single early check proves nothing either, because rows took 45–314 s to become queryable
  • an unscoped count is worthless in a shared workspace: another vault's rows satisfy it

So: read the secret again every few minutes, and poll for any row newer than the last one you know was dropped. Allow at least 45 minutes before concluding — the affected streams here were not confirmed live until ~36–37 minutes.

# Each pass makes a NEW read, so the check cannot be defeated by the first read
# having been discarded. It STOPS as soon as a row appears.
#
# START is fixed once, before the loop. Do NOT use a rolling ago(10m): ingestion
# has no SLA, so a read can become queryable only after it is already older than
# a short window, and the loop would then run forever against a live stream.
START=$(date -u -d '-2 minutes' +%Y-%m-%dT%H:%M:%SZ)

check() {
  az monitor log-analytics query -w <workspace-guid> --analytics-query "
    AZKVAuditLogs
    | where tolower(_ResourceId) == tolower('<key-vault-resource-id>')
    | where TimeGenerated > datetime($START)
    | project TimeGenerated, OperationName, ResultSignature
    | order by TimeGenerated desc" -o tsv
}

# Phase 1 - keep generating probes right up to the 45-minute deadline.
for i in $(seq 1 15); do
  az keyvault secret show --vault-name <vault> --name <cert> --query value -o tsv >/dev/null
  sleep 180
  ROWS=$(check); [ -n "$ROWS" ] && { echo "stream is live:"; echo "$ROWS"; exit 0; }
  echo "probe $i: nothing yet"
done

# Phase 2 - one probe ON the deadline, then ingestion grace.
#
# The boundary probe is not optional. Phase 1's last read lands at ~minute 42,
# so a stream starting between then and minute 45 would have NO probe after it
# and could never be detected however long you polled. This read is the one that
# covers that gap; the grace window then just lets it ingest, since observed lag
# reached 314 s and a single 180 s wait would miss it.
az keyvault secret show --vault-name <vault> --name <cert> --query value -o tsv >/dev/null
for i in $(seq 1 5); do
  sleep 120
  ROWS=$(check); [ -n "$ROWS" ] && { echo "stream is live (late):"; echo "$ROWS"; exit 0; }
  echo "grace $i: nothing yet"
done
echo "no rows after 45 min of probes plus 10 min grace - treat as unresolved, not as proof"

The break matters. Every pass reads the certificate's backing secret, which is exactly what certificate-secret-read alerts on — a loop that keeps going after the answer is in raises a Sev1 alert every evaluation until somebody stops it. Expect the reads you do make during this check to alert, and say so in advance if your action group pages anyone.

The first row that appears tells you the stream is live from at least that point. It tells you nothing about whether earlier reads were captured — if they happened before the stream started they are gone, and no amount of waiting will surface them.

Supplying your own established workspace is not a reliable way around this. A workspace already carrying Key Vault logs did capture five newly added vaults from their first minute — but a brand-new workspace elsewhere did the same, so prior use is not what makes the difference. Run the check either way.

Why the retention sweep does not trip mass-blob-deletion#

The lifecycle rule that expires snapshots issues the same DeleteBlob operation an attacker would, so a daily sweep on a large estate would exceed the 50-blob threshold and fire the alert on the product doing exactly what you configured it to do. The rule therefore excludes the sweep — on the one field a real sweep was observed to carry, not on a guess.

Observed 2026-08-04, by inducing a 60-blob sweep and reading its rows back from StorageBlobLogs in the same workspace the rule queries:

Deletion path AuthenticationType RequesterObjectId Counted by the rule
Lifecycle sweep (60/60 rows) TrustedAccessSas (empty) no — excluded
AAD data plane (the product path) OAuth populated yes
User-delegation SAS DelegationSas populated yes
Unauthenticated attempt Anonymous (empty) yes

AuthenticationType is asserted by the storage service from how the request actually authenticated. The sweep's value, TrustedAccessSas, is a system-key SAS the platform issues to itself (sk=system-1 in the logged URI) — it is not something a caller picks. The exclusion also coalesces authenticationType_s, the plausible legacy AzureDiagnostics flattening of the same field, for deployments whose rows still take that route — plausible rather than observed, because no storage row has ever been seen in AzureDiagnostics here; a wrong spelling leaves rows counted, so the rule fails noisy, never blind. The sweep's user-agent header (ObjectLifeCycleScanner) also identifies it, but headers are client-controlled, so the exclusion deliberately does not use it. Note the empty RequesterObjectId distinguishes nothing: anonymous attempts carry an empty requester too, which is why the exclusion keys on the authentication type instead. The deployment disables shared-key access on the account, so a key-signed SAS never authenticates at all, and the only SAS a caller can mint logs as DelegationSas — still counted.

Without the exclusion, the shipped query over the observed sweep hour counted Deletions = 60, past the threshold — the alert would have fired on routine retention. With it, the same hour counts 0 and a simulated 60-deletion OAuth attack in the same window still counts 60 and fires. The threshold stays at 50 rather than being raised: sweep volume scales with your estate, an attack does not have to.

One path is deliberately left outside this alert: something with management-plane write access on the storage account could rewrite the lifecycle policy itself and have the platform delete blobs under TrustedAccessSas. That is exactly what the retention policy changed detection below watches (MANAGEMENTPOLICIES/WRITE in the Activity Log) — one more reason to add it — and deleted recovery points remain recoverable through versioning and soft delete.

If this alert fires anyway, the deletions were not the retention sweep. Do not disable the rule.

If the certificate-read alert fires constantly#

It is designed to. The rule extracts the calling principal's object id and alerts when it is anything other than the backup automation identity — and when extraction yields nothing, the empty value does not match, so it alerts. A constantly firing rule therefore means the identity field name no longer resolves against your table schema, not that you are under attack.

Fix the field name; do not mute the rule. It is the compensating control for restore-credential isolation, and a muted alert is not a control.

The two detections you have to add#

Role-assignment changes and storage retention-policy changes are both worth alerting on, and neither ships. Both live in AzureActivity, and populating that table requires a diagnostic setting at subscription scope. This template deploys into a single resource group and its grants are scoped to that group, so it cannot create a setting at subscription scope.

Earlier versions shipped these two rules anyway. They queried a table that received no data, so they looked like controls on the portal blade and reported nothing, forever. They were deleted rather than left in place — a documented gap is honest, a dead rule is not.

1. Route Activity Log to the workspace#

Once per subscription:

az monitor diagnostic-settings subscription create \
  --name eidb-activity \
  --location <region> \
  --workspace <workspace-resource-id> \
  --logs '[{"category":"Administrative","enabled":true},{"category":"Security","enabled":true}]'

2. Add the rules#

Both are scheduled query rules against the same workspace. Substitute the id of the resource group you deployed EIDGuard into, and the id of its backup storage account — both are on that resource group's Overview blade in the portal (Properties → Resource ID, and the storage account's Properties → Resource ID).

Role assignment changed — someone granted or revoked access to the resource group. On a correctly operating deployment this happens at deployment time and never again.

AzureActivity
| where OperationNameValue in~ (
    "MICROSOFT.AUTHORIZATION/ROLEASSIGNMENTS/WRITE",
    "MICROSOFT.AUTHORIZATION/ROLEASSIGNMENTS/DELETE")
| where ActivityStatusValue =~ "Success"
| where tolower(_ResourceId) startswith tolower("<managed-resource-group-id>")

Retention policy changed — the lifecycle rule that expires snapshots was rewritten outside the product. The API enforces a floor at the soft-delete window; a change made directly against the storage account does not go through it.

AzureActivity
| where OperationNameValue in~ (
    "MICROSOFT.STORAGE/STORAGEACCOUNTS/MANAGEMENTPOLICIES/WRITE",
    "MICROSOFT.STORAGE/STORAGEACCOUNTS/MANAGEMENTPOLICIES/DELETE")
| where ActivityStatusValue =~ "Success"
| where tolower(_ResourceId) startswith tolower("<storage-account-id>")

Set both to Count > 0, evaluated every 15 minutes over a 1 hour window, pointed at the same action group the deployment created (<prefix>-auto-<hash>-alerts — the Automation account name plus -alerts; read it off the resource group rather than composing it).

Both queries are syntactically validated against a live workspace, but the volume and exact _ResourceId shape of Activity Log rows depends on your subscription. Confirm each returns rows for a change you make deliberately before relying on it.