From Azure signals to service ownership

A baseline is a collection and response path, not simply a shared dashboard.

  1. CollectPlatform metrics and activity logs are available; select resource logs with diagnostic settings.
  2. InvestigateRoute selected logs to a Log Analytics workspace; retain relevant metrics for rapid health checks.
  3. RespondConnect symptom alerts and supporting evidence to an owning service team and a tested action.

Executive Context

An Azure subscription can contain thousands of metrics and logs yet still leave an on-call engineer unable to answer a basic question: which customer-facing service is impaired? A platform observability baseline should make that answer faster. It defines which signals are collected, where they live, who may inspect them, and which changes in those signals deserve a human response. It does not promise that every incident will be detected automatically.

Azure Monitor brings together metrics, logs, traces, and events for investigation. Its data platform distinguishes Log Analytics workspaces, used for log and trace analysis with Kusto Query Language, from Azure Monitor workspaces, used for Prometheus and OpenTelemetry metrics with PromQL. An organization should not treat those similar names as interchangeable storage choices. This article concentrates on the Log Analytics and platform-signal path; teams adopting Prometheus metrics should design that workspace path separately.

Constraints and Assumptions

Consider a platform team supporting several subscriptions and application teams. Production services have different owners, while platform operators must see cross-service symptoms. Some logs contain sensitive operational information. Regions, data-access boundaries, retention requirements, and a finite telemetry budget influence the design. The baseline therefore needs a default that is easy to operate and an explicit process for exceptions, rather than a separate monitoring stack for every deployment.

Assume the application teams can identify critical user journeys and nominate a responder, even if they have not yet formalized service-level objectives. Start with a small set of failure and latency symptoms that a responder can act on. A diagnostic setting that copies every available category does not supply application context by itself. Equally, a dashboard that looks healthy while collection is broken is not an acceptable operational control.

Target Architecture

The default design is one appropriately located Log Analytics workspace for a coherent operational boundary, with diagnostic settings on resources whose logs are needed for investigation. Microsoft recommends starting with a single workspace and adding only the workspaces required by concrete business criteria. Data volume alone is not a reason to split: the workspace guidance says volume does not impose a workspace performance limit. Cross-resource queries and common operating procedures are simpler when related data remains together.

A workspace is not a universal dumping ground. Regional storage obligations, different data owners or billing parties, distinct access requirements, and incompatible retention needs may justify another workspace. For example, a separate security boundary may be appropriate when operational data and Microsoft Sentinel data have different owners or cost implications. Document the exception and its cross-workspace investigation cost before creating it. Set retention deliberately at workspace or table level where the requirement applies, and review who has access to the resulting data.

The diagram above describes three connected planes. Azure supplies platform metrics and activity-log data without first creating resource diagnostic settings; selected resource logs need settings to be collected. Log Analytics provides a queryable investigation path for routed logs. Service teams then attach alert rules, ownership and incident procedures to the signals they can actually use. Keep the metric and log roles distinct: a fast metric symptom can prompt an investigation, while resource logs may supply the context needed to explain it.

DevSecOps Control Model

Make the baseline part of service onboarding. A service specification should name its critical operation, owning team, expected telemetry sources, relevant resource-log categories, workspace destination, alert contact, and evidence of a tested response. Infrastructure review should check both the existence of required settings and whether they point to the intended destination. Settings must be defined for each resource that produces the selected data; merely creating a workspace does not turn on resource-log collection.

Separate authorization to change collection from authorization to read telemetry. Monitoring data can expose identifiers and operational details, so workspace placement and data-access decisions belong in the same review as retention. Platform engineers can maintain a reusable collection pattern, while application owners decide what constitutes a harmful user-visible symptom. A security team may require its own data boundary; that is a documented design choice, not a reason to duplicate all ingestion without review. Record exceptions with an owner and expiration or review date.

Key Architecture Decisions

Choose collection by question

Begin with questions an engineer will ask during a real failure: did requests fail, when did a resource change, and which component was affected? Diagnostic settings can route resource logs, platform metrics, and activity-log entries to supported destinations, including a Log Analytics workspace, a storage account, or an event hub. Resource logs are not collected by default; platform metrics and the activity log are automatically collected, but diagnostic settings are needed to send them to another destination. Select categories that answer investigation questions, then check actual availability and ingestion after rollout.

Separate live queries from archival needs

Choose the destination for the intended use. A Log Analytics workspace supports queries, workbooks, and log-based alerting. Storage can serve longer-lived archival or audit requirements, while an event hub can support downstream processing. These destinations have different operational characteristics; copying the same stream everywhere multiplies data movement and governance work. The Microsoft guidance also notes destination constraints: a destination must exist before configuration, and regional resources sending to storage or event hubs require destinations in the same region. Check those constraints before declaring a shared routing template portable.

Keep the exception budget small

Multiple workspaces increase query and configuration complexity. On the other hand, a single workspace should not override residency or ownership requirements. Write down the reason for each split, the teams allowed to query it, its retention policy, and the way responders will correlate incidents across boundaries. Revisit the decision when a new region, team, or security requirement changes the operating model.

Implementation Blueprint

  1. Inventory the services. Map subscriptions and resources to customer-facing services and responders. Mark the resources that participate in a critical journey, not just those easiest to monitor.
  2. Establish workspace boundaries. Start with the smallest practical number; review region, tenant, data ownership, access, billing and retention criteria before accepting an additional workspace.
  3. Build a collection matrix. For each resource type, identify the platform metrics to watch, the activity-log events worth investigating, and the resource-log categories required to diagnose a failure. State the destination and why each stream is needed.
  4. Apply settings and verify receipt. Create the destinations first, then configure per-resource diagnostic settings for selected streams. Generate a safe test event or inspect a known event, and verify that it arrives in the expected place. Allow for source-specific ingestion delay before interpreting an empty query as a failure.
  5. Bind alerts to actions. For each critical service, choose a symptom threshold or condition with a clear owner, an escalation route and a runbook. Exercise the alert path, not just the rule definition.

Treat collection templates and service ownership as versioned configuration. A platform change that creates a new resource type should trigger a review of its available diagnostic categories. A service release that changes traffic patterns should prompt a review of its alert thresholds. This is a repeatable control loop, not a one-time portal setup.

Operational Evidence and SLOs

Measure the baseline itself. Useful evidence includes the proportion of in-scope resources with the intended settings, whether expected events reach their destination, and the age of the latest representative log. These checks detect gaps in collection; they are different from a service-health objective. For the service, define a user-centered indicator, such as successful requests or time spent responding, and agree on an objective appropriate to that service. A platform-wide CPU alert is rarely a substitute for a failing user journey.

Review alert usefulness after incidents. Can the responder identify the service, impact, relevant time window, and next action from the notification? Which notifications were noise, and which incidents went undetected? Track acknowledgement and investigation time as operational feedback, not as proof that a workload met an availability objective. Workbooks or dashboards can help responders correlate signals, but they should follow clear questions and ownership rather than replace them.

Failure Modes and Trade-offs

A common failure is to confuse a reachable workspace with complete coverage. Missing per-resource settings leave resource logs absent even though metrics appear in the portal. Another is to route all categories without checking volume, usefulness or access; that raises cost and obscures important events. A third is to create many workspaces for organizational convenience and then discover that incident responders need a cross-workspace view. The antidote is a collection matrix, test events, explicit boundaries, and periodic review of ingestion against actual investigation needs.

Lifecycle matters as well. Microsoft warns that diagnostic settings should be removed when a resource is deleted, renamed, or moved across groups or subscriptions; otherwise, a recreated resource can inherit a stale setting. Incorporate that cleanup in decommissioning and migration procedures. Do not assume every resource supports identical log categories or that data appears instantaneously. Test the behavior of each onboarded resource type and state what delayed or absent telemetry means for incident response.

Adoption Roadmap

First, select one representative production service and document the journey, owner, workspace decision, required logs, and two or three actionable symptoms. Next, implement its resource settings and confirm that a responder can travel from alert to evidence. Record gaps found in a tabletop incident. Then convert the working pattern into a standard for other teams, allowing justified exceptions for region, access and retention. Finally, review coverage, alert noise, and ingestion choices on a regular cadence with both platform and service owners.

The rollout gate should not be “a dashboard exists.” It should be “we can recognize a real service problem, reach an owner, and find the evidence to investigate it.” That test makes observability a platform capability rather than a collection of disconnected charts.

Conclusion

An Azure observability baseline is a small set of accountable choices: which questions matter, which signals answer them, where those signals are stored, and who responds when they change. Begin with a simple workspace design, collect resource logs intentionally, and verify the complete path from symptom to investigation. Expand only when a documented requirement or an incident demonstrates the need. The result is not maximum telemetry; it is evidence that an owner can use when the service needs attention.

Sources

  1. Microsoft Learn: Azure Monitor overview
  2. Microsoft Learn: Design a Log Analytics workspace architecture
  3. Microsoft Learn: Diagnostic settings in Azure Monitor