Requirements - Non-Functional Profile - NF01 - Observability

Solutions should incorporate workload observability and understand service health.


Requirement description

This requirement is about ensuring that the solution can be monitored effectively and that teams understand how the service is performing.

The solution should provide sufficient observability to understand service health, identify issues, investigate incidents and support operational decision-making. Observability should cover the workload being delivered, rather than relying solely on infrastructure metrics.

In simple terms:
Teams should be able to see what the service is doing, understand whether it is healthy, and quickly identify and investigate problems.

Scoring rubric table – NF01 Workload Observability & Service Health

Score What it looks like Typical evidence Key gaps / risks
0 No evidence that workload observability exists. Service health is not measured or understood. No monitoring tools, dashboards, health metrics, telemetry or alerting. Significant risk that service failures, performance degradation or outages will not be detected promptly.
1 Limited evidence of monitoring. Monitoring is largely focused on infrastructure or ad-hoc checks. Basic infrastructure monitoring, manual operational checks, supplier assurances without supporting evidence. High-risk gaps. Limited visibility of user experience, application behaviour or operational performance.
2 Some elements of observability are implemented. Monitoring exists for parts of the solution but coverage is incomplete. Application logs, basic dashboards, limited alerting, selected service health indicators or partial telemetry. Significant notable gaps in monitoring coverage, alerting quality, ownership or operational processes.
3 Much of the requirement is met. Key workload and service health indicators are monitored and operational teams can investigate most incidents. Application monitoring, dashboards, health checks, structured alerting, logging platforms, operational runbooks and incident investigation evidence. Notable gaps remain in end-to-end visibility, dependency monitoring, business-level metrics or operational maturity. Mitigating actions are required.
4 Most of the requirement is met with good levels of evidence. Service health is actively monitored through defined indicators with clear operational ownership. Maintained dashboards, service-level indicators (SLIs), alert thresholds, operational reviews, monitoring standards, trend reporting and ownership documentation. Minor gaps only. Remaining risks are documented, understood and actively managed.
5 Comprehensive evidence that observability is mature, proactive and embedded within operational management. Monitoring capability exceeds normal expectations. End-to-end observability strategy, workload telemetry, proactive alerting, business and technical health indicators, post-incident learning, trend analysis and continuous improvement activities. Minimal gaps. Monitoring capability is routinely reviewed and drives service and architectural improvements.

What assessors should look for

  1. Service health monitoring
    Evidence that service availability, reliability and operational health can be measured and understood.
  2. Workload observability
    Evidence that application behaviour, transactions and workloads can be observed and analysed.
  3. Alerting and incident support
    Monitoring should help identify, diagnose and resolve incidents efficiently.
  4. Operational ownership
    Clear ownership of monitoring, dashboards, alerts and supporting processes.
  5. Dependency coverage
    Critical integrations, external services and supporting platforms are monitored where appropriate.
  6. Use of monitoring data
    Evidence that telemetry and monitoring outputs are used to support operational decisions and service improvements.

What separates a 3 from a 4 or 5

A score of 3 generally means the service is observable and operational teams can investigate issues, but there are notable gaps in coverage, ownership, dependency monitoring or monitoring maturity.

A score of 4 requires evidence that monitoring is actively maintained, governed and regularly reviewed. Service health indicators are defined, ownership is clear and observability is embedded in operational processes.

A score of 5 requires evidence that observability is used proactively rather than reactively. Monitoring informs operational improvement, incident reduction, capacity planning and architectural decision-making. Coverage is comprehensive across technical, operational and business perspectives.

Suggested examples of evidence (not SAF-mandated artefacts):

  • Observability strategy.
  • Application Performance Monitoring (APM) dashboards.
  • Service health dashboards.
  • Logging and telemetry standards.
  • Alert definitions and escalation procedures.
  • Service Level Indicators (SLIs).
  • Operational runbooks.
  • Incident reports and post-incident reviews.
  • Monitoring architecture diagrams.
  • Trend, capacity and performance reports.

Updated: 07 August 2026 (SAF Version 1.1)