Skip to content
PagesPrivate endpoints & Hybrid Worker deployment

EIDGuard docs

On this page

Private endpoints & Hybrid Worker deployment

Optionally wire the backup solution's Azure resources — Key Vault, Storage account, and the Automation account — behind Azure Private Link, so their data planes are reachable only from inside a virtual network.

Why a Hybrid Worker VM is required#

Azure Automation runs scheduled runbooks in a shared cloud sandbox that has no route into your VNet. If Key Vault and Storage are private-only, a cloud sandbox job cannot fetch the backup certificate or upload blobs. The supported way to run a job with private-network access is a Hybrid Runbook Worker — a VM inside the VNet that executes the runbook.

So enabling private endpoints also deploys a small Linux Hybrid Worker VM and links the backup schedule to run -RunOn that worker.

What stays public, and why#

disablePublicAccess (default true) disables public network access on Key Vault and Storage only — the data plane that actually holds your tenant backups and certificate.

The Automation account keeps public access enabled. This is deliberate: the cheap auto start/stop of the worker VM is driven by tiny cloud runbooks (Start-HybridWorkerVM / Stop-HybridWorkerVM) that only call ARM. If the Automation account were private too, those cloud jobs couldn't run — and they're what powers the worker VM on before the backup. The Automation account still gets a private endpoint (DSCAndHybridWorker + Webhook) for the worker's own communication.

Resources created#

Resource Purpose
VNet + snet-pe, snet-hw Private-endpoint subnet and hybrid-worker subnet
NAT gateway (+ Standard public IP) on worker/app subnets that lack one Outbound internet for the worker and route-all dashboard. Existing customer NAT gateways are reused; creation can be disabled for subnets whose 0.0.0.0/0 route already uses an NVA, Azure Firewall, or other egress service
Private DNS zones privatelink.vaultcore.azure.net, privatelink.blob.core.windows.net, privatelink.azure-automation.net (linked to the VNet)
Private endpoints Key Vault (vault), Storage (blob), Automation (DSCAndHybridWorker, Webhook)
Linux VM Burstable Hybrid Runbook Worker (system-assigned identity, no public IP, SSH-key auth, encryption at host when the subscription feature is registered — see below)
Worker runtime PowerShell 7 + Az.Accounts/Az.KeyVault/Az.Storage + Microsoft.Graph.Authentication installed on the VM (AllUsers) — a Linux hybrid worker, unlike the Azure cloud sandbox, ships none of these. Installed automatically via a Run Command
Role assignments Automation account identity → Key Vault Secrets User + Storage Blob Data Contributor (on an extension-based worker the runbook's Connect-AzAccount -Identity binds to the Automation account MI, not the VM's), and Virtual Machine Contributor on the VM for start/stop. The VM identity gets the same KV/Storage roles as a fallback.
Start/stop runbooks + schedules Deallocate the VM outside the backup windows — see The worker's start/stop window

Cost#

The Hybrid Worker is a burstable B-series Linux VM (default Standard_B2s) — no Windows license. It is deallocated except for a short window around each scheduled backup, so at a daily cadence you pay for a couple of hours of compute per day plus a small OS disk. A shorter RPO means more windows, and at a 1h or 2h cadence the VM stays on (see below). Standard_B1ms (2 GB) is cheaper but tight for the Az + Microsoft.Graph modules the runbook loads; Standard_B2s (4 GB) is the safer default.

The worker's start/stop window#

Scheduled jobs on a private deployment run on the worker, and a job routed to a worker group with no online worker never runs. So the VM is started before every scheduled backup and deallocated after it:

  • Start 15 minutes before each backup run, stop 2 hours after, both repeating at the backup schedule's cadence, linked to the cloud runbooks Start-HybridWorkerVM / Stop-HybridWorkerVM. The cadence is the schedule's real one — the GCD of the default RPO and every tenant's RPO — not the default RPO alone. Each schedule's name records its cadence and first time of day, for example Start-HybridWorker-6h-0145 and Stop-HybridWorker-6h-0400.

  • Always on at a 1h or 2h cadence, where the windows would overlap: there is no stop schedule, and a daily start brings the VM back if someone stops it.

  • Scheduled verification runs on the worker too. At the backup hour or the hour after, the backup window already covers it. Two hours after, every backup window stays up an extra hour. An hour before, a daily Start-HybridWorker-Verify-24h-<HHmm> schedule starts the VM early. Anywhere else it gets its own daily Start-HybridWorker-Verify-… / Stop-HybridWorker-Verify-… pair. A stop never fires inside another job's window.

  • With a schedule time zone (Configuration → Protection), the backup and verification schedules run on that zone's clock, but these power schedules stay in UTC with the same names. Each window is widened to cover every UTC instant its job can fall on through the year — the local time at the zone's standard and its daylight-saving offset — so in a zone with daylight saving the VM is up about an hour longer per window (for 2:00 AM New York: 05:45 to 09:00 UTC), and none of it depends on how Azure treats the nights the clocks change. Changing the zone re-aims the window like any other schedule save.

    Three things guard against the zone data the window is built from disagreeing with Azure's (Get-HybridWorkerWindow, modules/ExternalIDBackup.psm1):

    • Azure's own instants count. The next run and start time Azure reports for the backup and verification schedules are read back, and the offset each implies is added to the zone's, so the window covers the moment Azure will actually fire (Get-ScheduleObservedRun). The dashboard (Sync-HybridWorkerWindow) and the networking deploy read them the same way and build the same window.
    • Zones whose clock changes are in dispute are not trusted. The dashboard's API is the one place this is decided (Get-ScheduleZoneDataProblem, modules/ExternalIDBackup.psm1). It compares the zone with the reference zone of the Windows time zone it maps to, over the coming 405 days; on the Linux host that is the IANA zone database on both sides, so it catches only zones whose Windows reference zone has different offsets — a place that changed its rules apart from its neighbours (parts of Greenland and Antarctica, a Canadian province) or a Windows zone shared by places that no longer agree. Which zones that is depends on the host's zone data, so it is computed, not listed. Saving such a zone on a private deployment is refused with the reason (Assert-ScheduleZoneForWorker, web/api/Modules/WebApi/WebApi.psm1); one already set when private networking is enabled keeps the VM always on from the next schedule save, with a warning saying why. A zone the check cannot read is treated as unknown — neither refused nor always on — and left to Azure's own instants and the re-check below, which is also what covers Azure's data simply being older or newer than the host's (Morocco's Ramadan rules, for example).
    • Every run is re-checked, on every deployment with a Hybrid Worker, UTC schedules included. Each backup run compares the next run of each backup and verification schedule with the worker's start/stop schedules (Find-HybridWorkerWindowMiss); a miss still there a minute later is a warning in the job output and the always-on worker-window-miss alert (monitoring.md). Every few minutes the dashboard makes the same check without taking any lock; only a miss takes the schedule lock and is read again under it — skipping that cycle if a schedule save holds the lock, so a save in progress never reads as a miss and a check never blocks a save. A confirmed miss shows under Needs attention (on every refresh until the next check) and the window is re-aimed (Invoke-HybridWorkerWindowCheck, at most hourly).

These power jobs appear on the dashboard's Jobs page as Worker VM power on / Worker VM power off, so a failed power-on (and the backups that then never ran) is visible there. They are listed, not offered as operations: the job-start allowlist (Get-WebRunbookNames) excludes them, and the only dashboard path that starts one is the schedule-save re-aim described below.

The networking deploy creates these schedules, and every schedule save re-aims them: changing the RPO (globally or for one tenant), the backup hour, or verification. The rule lives in Get-HybridWorkerWindow (modules/ExternalIDBackup.psm1); the dashboard applies it through Sync-HybridWorkerWindow (web/api/Modules/WebApi/WebApi.psm1), and the networking runbook carries a copy that a test holds identical to it.

A changed window gets new schedule names. They are created and linked first, and only then are the other Start-HybridWorker* / Stop-HybridWorker* schedules removed, including the plain Start-HybridWorker / Stop-HybridWorker pair older deployments have. If a new schedule cannot be created, nothing is removed and the previous window keeps running. An unchanged window creates nothing.

The dashboard's identity cannot start, stop or modify the VM. It can read it: Monitoring Contributor, granted at resource-group scope for the alert action groups (functionMonitoring, lines 442-449 of marketplace/bicep/modules/roles.bicep), carries */read and monitoring resources (alerts, diagnostic settings), but no Microsoft.Compute write or action. Its only Automation grant is the custom EIDGuard Automation Operator role (functionAutomationActions, lines 384-407, assigned at lines 431-439 and scoped to the Automation account). That role allows schedule write/delete, job-schedule links and starting jobs, with no Microsoft.Compute actions either. Virtual Machine Contributor appears only in automationNetworkingRoles (lines 185-192), which is assigned to the Automation identity (automationPrincipalId, lines 194-203). So when the VM has to come up immediately (switching to always-on, or a backup or verification due before the new start schedule first fires), the dashboard starts a Start-HybridWorkerVM job, which runs as the Automation identity. The networking runbook, which already runs as that identity, starts a deallocated VM directly.

New stop schedules never fire within 2 hours of being created, so re-aiming the window cannot deallocate the VM under a job that is already running. If the window cannot be updated, the save still stands, and the dashboard shows a warning asking you to save again.

Template upgrades leave these schedules alone: the upgrade prune removes only stale ExternalID-* schedules.

Deploying it#

Dashboard → Networking. Everything below is deployed by the Deploy-ExternalID-Networking runbook (automation/runbook-networking.ps1), started from that page. Deployment is estate-wide, so only one such job may run at a time.

Two modes:

  • New network (recommended). The VNet is created inside the resource group, where the deployment identity already has rights. Nothing to grant, nothing to prepare — but the network is isolated, so reaching it from your own machines needs peering or a VPN.
  • Your existing network. Pick one of your VNets and grant the deployment's automation identity Network Contributor on it. The page offers this as a one-click grant, made browser-direct with your own Azure token.

You can also optionally put the web app itself behind a private endpoint, in which case the dashboard is only reachable from inside the network too.

The first deploy is slow (~15 min): the VM boot, the Hybrid Worker extension install, and the worker-runtime bootstrap (PowerShell 7 + Az/Graph modules) each take a few minutes.

Encryption at host#

Encryption at host encrypts the worker VM's OS/temp disks and its disk caches on the Azure host itself, before anything reaches storage. It is free, but it needs a one-time subscription-level feature registration (Microsoft.Compute/EncryptionAtHost), and it is applied when the VM is created — so the Networking page checks it before you deploy, and offers to do it:

  • Registered — nothing to do. The worker is created with it.

  • Not registered — the page offers Register encryption at host, made browser-direct with your own Azure token. The deployment itself cannot do it: registering a provider feature is subscription-scoped, while every role the deployment holds is scoped to the resource group (marketplace/bicep/modules/roles.bicep). Registration takes ~15 minutes to reach Registered, and the VM must be created after that — so wait, re-check, then deploy.

  • Registration refused — your account may not register subscription features. A subscription Owner or Contributor runs this once:

    Register-AzProviderFeature -ProviderNamespace Microsoft.Compute -FeatureName EncryptionAtHost
    # ~15 minutes to reach 'Registered'; check with Get-AzProviderFeature
    
  • State could not be determined (no permission to read subscription features, or the probe failed) — the page says so and claims nothing.

None of these block the deployment: a browser-side check that can be wrong in your disfavour must not stop you using private networking, which is valuable with or without this layer. Registering is the primary action; going ahead regardless is a separate, explicit Continue without encryption at host choice.

If the feature is not registered when the runbook runs, the VM is created without encryption at host and the deployment continues. Disks are still encrypted at rest by platform-managed SSE (always on for managed disks); the host-level layer is what is absent. The job output then says which of the two things happened — registration accepted (~15 minutes to propagate), or refused with the command for a subscription Owner or Contributor to run.

Azure can also refuse encryption at host at creation time even with the feature registered — the flag may not have propagated yet, and legacy VM sizes do not support it. The runbook retries the worker once without encryption at host and logs a NOTE: naming the rejection, rather than failing a deployment that has already created the subnets, any required NAT gateway, DNS zones and private endpoints. Any other VM creation failure is not retried and stops the job.

Turning it on for a worker that already exists#

A worker created before the feature was registered keeps standard SSE, and redeploying private networking does not change that — the deploy is idempotent and reuses the existing worker. Microsoft's documented route for an existing VM is that it "must be deallocated and reallocated in order to be encrypted".

The Networking page offers exactly that, once private networking is on and the feature reads Registered: Enable encryption at host. It runs as an Automation job that stops the worker, sets the property, and starts it again.

  • The worker restarts. For a few minutes it cannot run jobs. Scheduled backups routed to it during that window fail and run at the next slot.
  • It refuses before touching anything if a backup, restore or comparison is already running on the worker — or if it cannot tell whether one is. That is checked twice, the second time immediately before the stop, because a scheduled backup does not pass through the dashboard and can start while the first check is running. Every refusal says the worker was not stopped.
  • The job runs in the Azure cloud sandbox, not on the worker, for the obvious reason: a job running on the VM it stops cannot start it again. This is the same route the automatic start/stop schedules use.
  • The prerequisite is checked with your credentials on the page before the job starts. The job re-checks, but the deployment's own identity is scoped to the resource group and usually cannot read subscription features — so when that read is denied the job says so and continues rather than refusing every time. If the feature really is missing, Azure rejects the update and the worker is restarted unchanged.
  • It is a no-op on a worker that already has encryption at host.
  • If the update fails, the worker is still restarted and the job says so. In the rare case the restart itself fails, the job reports CRITICAL and names the VM to start from the Azure portal — scheduled jobs fail until it is running.

Registering the feature before the first deployment is still the better path: it avoids the restart entirely.

Ordering, and what goes private last#

The deploy uploads/creates everything first and disables public access on Key Vault + Storage last, so it cannot lock itself out mid-run.

After lockdown, anything that talks to Key Vault or Storage directly from outside the VNet — an administrator's workstation reading a recovery point with Storage Explorer, for instance — needs network line-of-sight to the VNet (VPN, peering, or a jumpbox). The scheduled hybrid worker always has it.

Outbound connectivity (the worker DOES need internet)#

The Hybrid Worker has no public IP and accepts no inbound traffic, but it requires outbound internet — provided by an existing subnet NAT gateway, the EIDGuard NAT gateway, or another egress path such as an NVA/Azure Firewall. Private endpoints only cover Key Vault, Storage, and the Automation account; several dependencies have no private-link path and must be reached over the internet:

Destination Why
login.microsoftonline.com Entra token acquisition (managed identity + app-only)
graph.microsoft.com reading the External ID tenant — Microsoft Graph has no Private Link
packages.microsoft.com, PowerShell Gallery one-time worker runtime bootstrap (PowerShell 7 + modules)

So the worker cannot run in a fully air-gapped subnet. You can, however, restrict egress to just the FQDNs above with an Azure Firewall / NSG in front of the NAT gateway if your policy requires it. Key Vault and Storage traffic stays on the private endpoints (10.x) regardless.

The table above is the short version, and it is not the whole list — it omits ARM, the Automation service endpoints, the Functions host's queue/table storage, and the destinations that only appear in particular configurations. Before you write firewall rules, use outbound-connectivity.md: the full inventory, split by which component makes the call, with the symptom you will see when each one is blocked.

Ordering & the workstation caveat#

The deploy sequence uploads the certificate to Key Vault before locking it down, then disables public access on Key Vault + Storage last.

After lockdown, any machine that reaches Key Vault or Storage directly must have network line-of-sight to the VNet (VPN, peering, or a jumpbox). The scheduled Hybrid Worker always does. If you deploy from a machine outside the VNet, do it before lockdown or re-enable public access temporarily — and see "Upgrading a private-networked deployment" below for version upgrades.

Upgrading a private-networked deployment#

A version upgrade redeploys the template, and the template's deployment scripts run in Azure container instances that are not covered by the storage account's "trusted Azure services" bypass. After lockdown the data storage account has public access disabled, so those scripts cannot reach it:

  • read-dashboard-settings reads the settings you chose in the dashboard (schedule, retention, alerting, verification) so the upgrade keeps them. When it cannot read them it fails the upgrade on purpose — falling back to the template's original values would silently put the deployment's settings back over yours.
  • stage-functionapp-package writes the new application package to the same account, and preserve-wizard-auth reads the same settings blob.

So, before a version upgrade of a private-networked deployment:

  1. Re-enable public network access on the data storage account (the one holding the backups and webconfig containers), for example az storage account update -n <data-account> -g <rg> --public-network-access Enabled.
  2. Run the upgrade.
  3. Disable public access again: --public-network-access Disabled. The private endpoints and the Hybrid Worker's routing are untouched by the upgrade, so nothing else needs re-doing.

The window is the length of the deployment. Snapshots stay protected throughout by the locked immutability policy, which does not depend on network posture.

Verifying#

From the Hybrid Worker VM (or any host in/peered to the VNet):

nslookup <your-vault>.vault.azure.net      # resolves to a 10.x private address
nslookup <your-storage>.blob.core.windows.net

Then start the backup runbook and confirm the job runs on the worker group and succeeds end-to-end:

Then start a backup from the dashboard (Operations → Backup) and confirm the job runs on the worker group and succeeds end to end.

A direct run from outside the VNet should now fail to reach Key Vault/Storage — confirming the isolation.

Outbound egress#

By default the NAT gateway (eidb-natgw, in the resource group) is associated with each of these subnets only when that subnet has no NAT gateway already attached:

Subnet Why it needs egress
Hybrid worker No public IP, so no default outbound. Needs Microsoft Graph and extension installs
App / site integration The site runs vnetRouteAllEnabled = true, so all of its outbound is forced through the VNet. Needs the tenant JWKS endpoint for bearer-token validation, ARM, and the Automation control plane

NAT associations already present on customer subnets are preserved and reused; the worker and app subnets may therefore use different customer gateways. Where EIDGuard attaches its own gateway, those subnets share one stable public egress address. Existing-VNet mode also offers an opt-out for NAT-less subnets whose 0.0.0.0/0 route already reaches an NVA, Azure Firewall, or another egress service. Opting out without such a path can break Internet connectivity.

Automation stays deliberately public (control plane only; backup data moves through private endpoints). Key Vault and storage are the services whose public access is disabled by the lockdown step.

Before this was fixed the app subnet had no explicit egress and the web tier relied on platform-provided SNAT (finding P-01 in the product's internal security hardening register).

Teardown note, existing-VNet mode only: an existing customer NAT gateway is only reused and is not deleted or detached. Where EIDGuard attached eidb-natgw, the gateway lives in the resource group while the subnet lives in the customer's VNet, and the deployment delete cascade cannot detach across that boundary. Clear that stale association after deletion. Dedicated-VNet mode (the recommended default, everything in the deployment's resource group) is unaffected.