Skip to content
PagesBacking up multiple External ID tenants

EIDGuard docs

On this page

Backing up multiple External ID tenants

One installation backs up multiple External ID tenants — up to the number your plan includes (Starter 3, Standard 5, Enterprise unlimited; see plans-and-limits.md). Each tenant gets its own read-only app registration + certificate (app-only auth is per-tenant), but they share the storage account, Key Vault, and — for the scheduled path — a single Automation account / Hybrid Worker.

You onboard each tenant from the dashboard's onboarding wizard; everything else on this page happens in the dashboard too.

Recovery points are namespaced by tenant (<tenantId>/<timestamp>/… in the backups container), so tenants never collide.

At the plan limit, onboarding another tenant is refused with a 403 until you change plan or offboard one. Nothing else is affected: tenants already onboarded keep backing up even if a plan change leaves the deployment over the limit, and re-onboarding an existing tenant is never blocked. See plans-and-limits.md.

What the deployment keeps per tenant#

For each onboarded tenant the deployment records the tenant id, the client ids of its two app registrations (backup, read-only; restore, read-write) and the names of their certificates in the deployment's Key Vault, plus any protection override an admin has set for it — its own RPO, retention window or alert recipient (see Per-tenant protection). A tenant without an override inherits the deployment-wide value for each. You never edit this record by hand: onboarding writes it, offboarding removes it, and re-onboarding a tenant refreshes its app registrations and certificates while keeping its protection override.

What stays deployment-wide: the storage account and Key Vault, the one Automation schedule and the backup hour that anchors it, private networking and its Hybrid Worker, and the backups container's immutable window.

When a tenant is offboarded its own alert rule is removed, and the retention policy and schedule are recomputed without it. Its snapshots stay in storage and age out under the deployment's catch-all rule — the longest retention window still configured — not under the tenant's former override. While any tenant or the default keeps snapshots forever there is no catch-all, so an offboarded tenant's snapshots are kept forever too. If any of that cleanup could not be completed the offboarding still succeeds and says what was left, and the Tenants page offers that Admin Retry cleanup for the tenant (Dismiss hides it). Retry cleanup re-runs only the deployment-side cleanup — the alert rule, the retention policy and the schedule — by calling the offboard API again in cleanup-only mode (web/api/TenantOffboard/run.ps1, safe to repeat). It never signs in to the tenant, never touches its app registrations and never changes certificates. It is not offered for a tenant that is configured again, and the server refuses it (changing nothing) if the tenant was onboarded again in the meantime. A full offboard leaves enabled any certificate another configured tenant uses.

Per-tenant protection#

An Admin sets a tenant's own protection on Tenants → the tenant's row → Settings (web/app/src/components/TenantProtectionPanel.tsx, saved by web/api/TenantProtection/run.ps1). Each field is optional; leave it blank to inherit the deployment value from Configuration.

Setting Values How it is applied
RPO 1, 2, 4, 6, 8, 12 or 24 hours The tenant is backed up by the scheduled run once it is due (below)
Retention 0 (keep forever), or from the retention floor up to 3650 days A prefix rule for the tenant's snapshots in the storage lifecycle policy
Alert recipient One email address A tenant-scoped backup-failure alert, in addition to the deployment address — see monitoring.md

RPO. There is still one backup schedule. It fires at the greatest common divisor of the default RPO and every tenant's effective RPO (Set-BackupSchedule in web/api/Modules/WebApi/WebApi.psm1, via Get-BackupScheduleInterval in modules/ExternalIDBackup.psm1) — not the shortest, because a 4h schedule could only serve a 6h tenant every 8h; tenants at 4h and 6h give a 2h schedule. Each scheduled run backs up only the tenants that are due (automation/runbook-main.ps1, its per-tenant RPO section, which judges them against that same cadence — the default RPO and every tenant's). The selection is Select-DueBackupTenant, which asks Test-TenantBackupDue (both in modules/ExternalIDBackup.psm1) about each tenant: a tenant is due when waiting for the next scheduled run would take its newest recovery point past its RPO (with 15 minutes of slack for job start-time jitter). So a recovery point taken on schedule is followed exactly one RPO later, and one taken at any other time — an on-demand backup, onboarding — no later than 15 minutes past its RPO. A tenant with no recovery point yet, or whose last one cannot be read, is always due. A shorter RPO on one tenant therefore makes the schedule run more often, but the other tenants are still backed up at their own RPO.

Retention. All windows compile into the one lifecycle policy on the backups container (ConvertTo-LifecycleRuleSet in modules/ExternalIDBackup.psm1): a catch-all rule at the longest finite window, plus prefix rules (backups/<tenantId>/) for every tenant with a shorter one. If any tenant or the default keeps snapshots forever (0), there is no catch-all, and only tenants with a finite window get a rule. Azure caps a lifecycle policy at 100 rules; tenants that share a window share rules (ten prefixes per rule), so the limit is reached only with many distinct windows. The compiler refuses before anything is written, so the live policy stays as it was. What happens to the value you saved depends on where you saved it:

  • A tenant save that would exceed the limit is stored but not applied: the override is kept on the tenant, the policy is not written, and the dashboard reports retention as not applied.
  • A default retention save (Configuration) that would exceed it is refused, and nothing is stored.

The retention floor is max(soft-delete window, the immutable window Azure currently enforces on the container) (Get-RetentionFloor in web/api/Modules/WebApi/WebApi.psm1; the immutable window is the one marketplace/bicep/modules/wormlock.bicep locks). A new tenant window below it is refused before it is stored (the retentionDays check in web/api/TenantProtection/run.ps1), and so is any finite one while the immutable window cannot be read.

A window you saved earlier can end up below the floor later — an upgrade that lengthens soft delete, for instance. It is not refused and it does not block anything:

  • Below the soft-delete window, both the upgrade and every dashboard retention write (a tenant save, onboarding, offboarding, a default save) apply it at the soft-delete window instead, until you raise or clear it. Each of those dashboard writes shows a warning naming the tenant and the window applied. Both use the same raise: ConvertTo-FlooredTenantRetention (modules/ExternalIDBackup.psm1), called by Set-RetentionPolicy (web/api/Modules/WebApi/WebApi.psm1) and inlined in the readback script of marketplace/bicep/modules/settingsreadback.bicep. The Tenants page still shows the value you saved: the raise never rewrites it. The one thing that does is the immutable window rising during a retention save, when every window it overtook is recorded at the corrected value (Set-RetentionPolicy, web/api/Modules/WebApi/WebApi.psm1) — only ever lengthened, never shortened.
  • Below the immutable window only, it is applied as saved, with the same kind of warning. It loses nothing: Azure refuses every delete inside the immutable window, so the rule cannot act until the window has passed.

Upgrades keep all three per-tenant settings: see plans-and-limits.md.

Onboarding tenants#

Onboard each tenant from the dashboard: Tenants → Onboard a tenant, once per External ID tenant. See setup-wizard.md for the full flow and the admin rights each tenant needs.

How many tenants you may onboard is set by your plan — Starter 3, Standard 5, Enterprise unlimited. See plans-and-limits.md; note that reaching the limit only ever blocks onboarding a new tenant, never backups of the ones you already have.

Single shared Key Vault#

All tenants' certificates live in one Key Vault, the one created inside the deployment's resource group. To keep them from colliding, each tenant's certificate is named per-tenant — ExternalID-Backup-ReadOnly-<tenant-slug> and ExternalID-Restore-ReadWrite-<tenant-slug>, where the slug is the first segment of the tenant's domain. Onboarding handles this automatically; there is nothing to choose.

Running backups#

A backup run is one job that loops over the tenants and writes one recovery point per tenant under the same timestamp. A run started from the dashboard's Run backup button backs up every onboarded tenant; a scheduled run backs up only the tenants that are due under their RPO (see Per-tenant protection). A failure in one tenant is logged and the others still run; the job reports which tenants succeeded and which did not. Onboarding finishes with a validation backup of the new tenant, so you do not have to wait for the next run to see it work.

Restoring#

Restore always targets one tenant at a time. In the dashboard, open Recovery points, pick the tenant, then one of its recovery points, and choose what to restore — a resource type, or named objects within it. The restore runs as a preview first (nothing changes) and applies only when you confirm. If the recovery point was taken from a different tenant than the one you are restoring into, the dashboard warns before anything runs. See web-ui.md for the restore wizard in detail.

Scheduled multi-tenant backups#

The scheduled backup runs from one Automation account and one schedule (and one hybrid worker, in private mode): each run is one job that loops over the tenants due at that run. The tenant list comes from the webconfig projection, which onboarding updates — so a newly onboarded tenant joins the next scheduled run with nothing further to do. Setting, changing or clearing a tenant's RPO, changing the default RPO or backup hour, and offboarding a tenant re-aim the schedule to the new cadence; a newly onboarded tenant inherits the default RPO, which the cadence already serves.