Backup resilience: separate production, protection, and approval

Conceptual control boundaries; verify workload support and vault settings before adopting this pattern.

Production workload · application ownerProtected VM, database, or supported service sends scheduled backups to an Azure Backup vault.
Vaulted recovery points · backup platform teamSoft delete delays permanent removal; vault immutability blocks destructive operations when enabled and locked.
Independent authorization · security teamResource Guard and multi-user authorization gate protected operations; monitoring signals unusual changes.
Restore environment · recovery teamValidate a selected recovery point and application behavior before reconnecting restored systems.

Recovery path: detect an incident → authorize response → select a viable recovery point → restore into a controlled environment → validate data and service.

Executive Context

A backup job reporting success does not prove that a business service can survive ransomware. An attacker who gains administrative access might stop protection, shorten retention, delete data, or simply wait until compromised data replaces useful recovery points. A resilient design must preserve an earlier recovery point, make destructive changes difficult, tell responders when protection changes, and demonstrate that restoration actually works. Azure Backup offers several complementary controls, but none independently establishes a recovery-time or recovery-point objective.

This article focuses on vaulted Azure Backup recovery points, not every operational backup or snapshot. Microsoft describes vaulted backup data as isolated from the production environment in Microsoft-managed storage. That boundary is valuable, but it does not replace access governance or a response plan. The pattern below is an architectural proposal: adapt its retention, vault type, supported workloads, and approval process to your own recovery requirements. See Microsoft's security overview for the underlying controls.

Constraints and Assumptions

Start with the recovery contract

Inventory applications, dependencies, data stores, and their owners before creating vault policies. For each service, document the most data it can afford to lose, the time allowed to restore it, the minimum clean history required to outlast detection delays, and who declares data safe to use. A long retention period cannot compensate for infrequent backups; frequent backups cannot compensate for retained copies that an attacker can erase. Choose schedules and retention against those distinct requirements, and account for the possibility that the most recent recovery point contains encrypted or otherwise corrupted data.

Group backup items by recovery needs rather than imposing one policy across all workloads. Separate the question “does this vault hold enough history?” from “could an operator destroy that history?” The first is a policy and capacity decision; the second calls for protection controls. Record which recovery method each workload supports and whether the test environment has the networking, keys, identities, and application dependencies needed for a meaningful restore. Treat those prerequisites as part of the recovery design, not a detail to discover during an incident.

Target Architecture

Use soft delete as a recovery window, not a permanent lock

Soft delete delays permanent deletion so a responder can recover deleted backup data during the configured window. Microsoft's secure-by-default guidance documents a default 14-day soft-delete retention period and a configurable range up to 180 days. It also notes that the first 14 days have no additional charge while longer periods can incur charges. Choose a window that covers likely detection and escalation time; do not assume the default is long enough for an organization that discovers attacks weeks later.

Check the actual region, vault type, and workload instead of treating “on by default” as a universal compatibility claim. The same guidance distinguishes general availability from preview for Backup vaults by region and describes exceptions for operational backups: operational blob and Azure Files backups do not have this soft-delete coverage, while VM instant-restore snapshots can be directly accessed and deleted before their soft-delete period expires. That is why the diagram identifies the protected boundary as vaulted recovery points. Document precisely what is and is not inside it.

Soft delete is a second chance after deletion, not authorization to make destructive changes. A privileged user may still attempt to alter protection, and an incident that goes undetected beyond the window can exhaust that chance. Assign responders the task of checking for soft-deleted items after suspicious activity and practicing recovery of one in a non-production drill. Couple this control with immutable retention and independent approval rather than presenting it as “ransomware-proof.”

Key Architecture Decisions

Lock immutability only after testing its consequences

Azure Backup's immutable vault guidance explains that immutability blocks operations that could cause loss of recovery points, including stopping protection while deleting data and reducing retention under the documented rules. The setting has three states: disabled; enabled but reversible; and enabled and locked. Only the locked state makes the setting irreversible and enables WORM-backed storage where available. “Enabled” and “locked” are therefore not interchangeable evidence in an audit.

Decide which retention behavior is appropriate before locking. Microsoft documents both immutability based on backup-policy retention and a specific immutable duration independent of overall policy retention. With a specific duration, recovery points can remain under ordinary policy retention after immutability expires; with policy-based immutability, policy retention governs the protection interval. Do not equate a longer policy with a longer immutable period without checking the selected mode. Test policy changes, decommissioning, and exception handling on representative items first. Because locking cannot be reversed, require a documented review of retention obligations, storage cost, and operational consequences before approval.

Scope matters here too: immutability applies to all data in the vault, but Microsoft's support matrix differs by vault and workload, and the guidance explicitly excludes operational backups of blobs, files, and disks. Build a coverage register showing each protected item, vault type, configured state, and immutable duration. If an application depends on an excluded operational copy, design an appropriate separate protection route rather than assuming the vault setting extends to it.

DevSecOps Control Model

Separate routine backup work from destructive authority

Microsoft's Azure Backup security overview describes Azure role-based access control and multi-user authorization (MUA). MUA uses a separate Azure Resource Guard so selected critical operations require applicable authorization. Design the Resource Guard ownership boundary so the same person or automation identity cannot both operate the backup vault and unilaterally approve its most consequential changes. Review access at the resource scope and confirm the exact operations protected by your vault's configuration.

Give backup operators only the permissions necessary to run and troubleshoot backups; reserve changes to retention, security settings, and destructive workflows for reviewed roles. Protect emergency access without letting it become a standing shortcut around approval. An access review should include service principals and deployment identities, not just human users. Rotation and offboarding matter because an unused but privileged identity remains a possible path to disable protections. MUA complements immutability: approval slows or stops sensitive changes, while locked immutability constrains destructive outcomes even after credentials are compromised.

Operational Evidence and SLOs

Make changes and missing backups visible

Monitor both the assets that are protected and those that should be protected but are not. Microsoft's overview points to Azure Backup's built-in monitoring, alerts, and Backup Reports for backup and restore activity. Give someone ownership of failed or missing jobs, changed policies, stopped protection, deleted items, and suspicious vault operations. Route actionable signals into the incident process with the application owner, backup operator, and security responder identified in advance. An alert with no accountable receiver is not a recovery control.

Use a daily coverage check against the workload inventory and periodically reconcile the list of vault-protected items. Record the last successful backup and the last successful tested restore separately. During an incident, preserve the timeline of jobs and configuration changes so responders can identify when recovery points might have become untrustworthy. Monitoring detects drift; it does not itself preserve data or confirm application correctness. Those jobs belong to retention safeguards and restore exercises.

Implementation Blueprint

Exercise the complete restore, not just the button

A restore drill should begin with a specific incident scenario: a deleted backup item, compromised application data, or loss of a production environment. Select a recovery point from before the simulated compromise, restore to a controlled non-production location when supported, then test application-level consistency and access. Do not reconnect restored systems to production until they have passed security and data checks. Document the steps to recover associated secrets, identity dependencies, networking, and downstream services without exposing the restored copy to the original threat.

Time the exercise from declaration through service validation. Compare observed recovery time and the age of the usable recovery point with the agreed objectives; record blockers and repeat until they are fixed. Also rehearse escalation for any protected operation that needs Resource Guard authorization, and confirm the approver is reachable when the primary team is unavailable. A successful file restore is useful evidence, but it cannot by itself certify a multi-tier service. Keep a small evidence packet: selected recovery point, job outcome, integrity checks, elapsed time, and approved exceptions.

Adoption Roadmap

Start by identifying vault-backed versus operational copies and measuring current protection coverage. Next, review soft-delete retention against detection time and cost; then test immutability in a pilot vault with realistic retention changes. Establish the independent Resource Guard approval path, review role assignments, and only then lock immutability after the irreversible decision has been explicitly accepted. Add alerts and an owned coverage review, followed by scheduled restore drills for representative applications. Revisit policy whenever a workload, threat model, or compliance requirement changes.

Failure Modes and Trade-offs

The principal trade-off is intentional friction: long-lived, locked recovery points reduce destructive flexibility and can increase storage costs. Independent approval can slow emergency changes if its owners and break-glass process are poorly designed. Conversely, omitting those controls to speed routine operations leaves the backup estate dependent on the security of a single administrative identity. A quiet backup failure can erase the value of every configured safeguard; a restore that succeeds technically but fails application checks can be just as disruptive.

Conclusion

No single Azure Backup setting promises immunity from ransomware. The defensible outcome is a set of independently owned boundaries—recoverable deleted data, appropriately locked vaulted recovery points, reviewed critical operations, visible coverage, and exercised recovery—plus evidence that an authorized team can restore a clean service within its agreed objectives.

Sources

  1. Microsoft Learn: Overview of security features in Azure Backup
  2. Microsoft Learn: Secure by default with soft delete for Azure Backup
  3. Microsoft Learn: Immutable vault for Azure Backup