Site Reliability | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Mon, 09 Feb 2026 12:34:55 +0000 en hourly 1 Cloud-native observability made seamless with OpenPipeline and AI-driven observability https://www.dynatrace.com/news/blog/simplify-cloud-native-environments-ai-driven-observability/ https://www.dynatrace.com/news/blog/simplify-cloud-native-environments-ai-driven-observability/#respond Thu, 03 Oct 2024 14:08:48 +0000 https://www.dynatrace.com/news/?p=65815 cloud-native observability made seamless with OpenPipeline

Dynatrace helps cloud-native teams monitor, manage, and troubleshoot complex systems built on Kubernetes® and cloud technologies. With new bespoke apps for Ops and SRE teams, Dynatrace provides familiar experiences with capabilities beyond those offered by popular open-source solutions. Immediate observability insights into workloads and platforms without configuration are complemented with automatic security assessments and AI analytics for predictive operations.

The post Cloud-native observability made seamless with OpenPipeline and AI-driven observability appeared first on Dynatrace news.

]]>
cloud-native observability made seamless with OpenPipeline

The complexity of modern cloud-native environments is ever-increasing. The latest State of Observability 2024 report shows that 86% of interviewed technology leaders see an explosion of data beyond humans’ ability to manage it. On average, organizations have twelve different platforms and services embedded in their multicloud environments.

To avoid drowning in data, it’s critical to ensure that organizations can seamlessly collect data from any source in a single place and in context. Visualizing data in context while supporting and automating decisions with causal, predictive, and generative AI—all while providing a seamless experience—is where the future of cloud-native observability lies.

With new and updated experiences for Ops and SRE teams, proactively managing your cloud environments has never been easier.

Cloud-native observability for monitoring your cloud

OpenPipeline™ is the Dynatrace platform data-handling solution designed to seamlessly ingest and process data from any source, regardless of scale or format. With OpenPipeline, you can effortlessly collect data from Dynatrace OneAgent®, open-source collectors such as OpenTelemetry, or other third-party tools. OpenPipeline then filters and preprocesses the data.

OpenPipeline also incorporates data contextualization technology, enriching data with metadata and linking it to other relevant data sources. By contextualizing data, OpenPipeline enhances the Dynatrace platform’s ability to offer AI-driven insights, analytics, and automation across observability, security, software lifecycle, and business domains.

Furthermore, OpenPipeline is designed to collect and process data securely and in compliance with industry standards. It features high-performance filtering, masking, routing, and encryption capabilities that are easy to configure and operate.

In the latest enhancements of Dynatrace Log Management and Analytics, Dynatrace extends coverage for

  • Native Syslog support: Use Dynatrace ActiveGate to automatically add context and optimize network traffic to your Syslog messages.
  • Seamless integration with AWS Data Firehose: address high-impact issues quickly through real-time, high-frequency log analytics. Dynatrace support for AWS Data Firehose includes AWS Lambda logs, Amazon Virtual Private Cloud (VPC) flow logs, Amazon S3 logs, and Amazon CloudWatch.
  • Kubernetes log monitoring with Fluent Bit

In an effort to further democratize data, Dynatrace provides a curated and supported OpenTelemetry collector. The Dynatrace Otel Collector includes collector components verified by Dynatrace for seamless operation. This removes the burden of manually validating each component and use case. The list of use cases is actively extended and includes batching and ingesting data from Syslog, Fluentd®, Jaeger™, Prometheus®, StatsD, and more.

Dynatrace OTel Collector
Figure 1. Dynatrace OTel Collector

Understand and secure your cloud

It’s critical to have easy and intuitive access to contextually relevant answers when working within complex cloud-native environments and the gold mine of information they provide. Dynatrace is essential for unlocking that gold mine of data, allowing you to enhance application performance, deliver better experiences, and optimize operational efficiency.

Kubernetes

The Dynatrace Kubernetes experience for Site Reliability Engineers (SREs) and Platform Engineers focuses on providing insights into the health and performance of multicloud Kubernetes environments in the tailor-made Kubernetes app. Powered by Davis® AI, the Kubernetes app offers proactive monitoring and analysis, allowing automated monitoring and optimization of health and performance and providing simple and easy troubleshooting. In addition, ready-made dashboards are available for a quick and easy overview, allowing you to see the Kubernetes data you want alongside the cloud-native observability data you need, all in one place.

Kubernetes Cluster Dashboard
Figure 2. Kubernetes Cluster Dashboard

Clouds

Quickly onboard and manage cloud monitoring in one place across different cloud providers and observe multiple cloud environments, including their instances, resources, cost analysis, health, and optimization. Ingest data remotely through cloud integrations covering Amazon CloudWatch, Azure Monitor, Azure Liftr, and Google Cloud™ Kubernetes with GKE™ AutoPilot cluster.

Vulnerabilities

Prioritize and sort vulnerabilities based on Davis Security Score, detection time, or the number of affected entities. Quickly identify if public exploits, internet exposure, or reachable data assets are exposed. Get powerful insights into the true impact of vulnerabilities in your environment. The embedded Davis Security Advisor can help you enhance remediation for third-party vulnerabilities with AI-assisted and precise recommendations on remediation actions, allowing you to address several critical vulnerabilities simultaneously. Davis AI uses security intelligence and runtime context to determine risk and remediation based on criteria like internet exposure and access to sensitive data.

Prioritization of vulnerabilities and recommendations by Davis Security Advisor
Figure 3. Prioritization of vulnerabilities and recommendations by Davis Security Advisor

Service-Level Objectives

Measuring a system’s reliability can be complex and overwhelming. Indicators that provide insights into the environment’s health and performance vary across industries, products, and applications. Applying Service Level Objectives (SLO) to track these indicators is a common best practice within site reliability engineering. Still, an SLO’s quality lies in the significance of the underlying service-level indicator.

Dynatrace guides you in quickly setting up the most valuable SLOs—considering the typically used metrics but providing the freedom to use any data stored in Dynatrace Grail™ data lakehouse, such as logs and events.

It doesn’t matter if you need the typically used failure rate or response-time metrics to ensure your system’s availability and performance or if you need to rely on abnormal log drops to gain insights into raising problems—SLOs leveraged with Grail provide all the information you need.

Automate your cloud

Answer-driven Dynatrace Automation is further extended by providing direct interaction with your cloud and cloud-native ecosystem:

  • Kubernetes: Read manifests, free up resources (for example, delete failed terminations), and restart deployments.
  • AWS: Automate your AWS infrastructure with actions across EC2, S3, Lambda, and more.
  • GitHub®: Integrate with your GitHub repositories. Automate issues and pull requests (for example, to change configuration files for sizing).
  • GitLab™: Integrate with your GitLab projects. Automate issues and merge requests (for example, to change configuration files for sizing).

Example: Predictive auto-scaling for Kubernetes workloads

Kubernetes provides flexibility but also introduces complexity. For instance, manual scaling is time-consuming, reactive, and prone to errors. Using Dynatrace Automation and Davis AI helps you predict bottlenecks and automatically open pull requests to scale applications. This proactive, automated approach minimizes downtime, ensures your applications perform at their best, and helps optimize resource utilization and cost.

The Auto-Scaling tutorial provides a step-by-step guide to automatically scaling Kubernetes workloads up or down, horizontally or vertically.

Workflow in Dynatrace video thumbnail

Try out cloud-native observability yourself

Proactively manage your environments to increase performance and reduce cost. It’s never been easier to truly own your cloud!

The capabilities highlighted in this blog post will be available in Dynatrace SaaS environments in the coming weeks.

The post Cloud-native observability made seamless with OpenPipeline and AI-driven observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/simplify-cloud-native-environments-ai-driven-observability/feed/ 0
Implementing AWS Well-Architected pillars with automated workflows https://www.dynatrace.com/news/blog/implementing-aws-well-architected-pillars/ https://www.dynatrace.com/news/blog/implementing-aws-well-architected-pillars/#respond Wed, 13 Sep 2023 16:52:01 +0000 https://www.dynatrace.com/news/?p=59517 Dynatrace | AWS

If you use AWS cloud services to build and run your applications, you may be familiar with the AWS Well-Architected Framework. This is a set of best practices and guidelines that help you design and operate reliable, secure, efficient, cost-effective, and sustainable systems in the cloud. The framework comprises six pillars: Operational Excellence, Security, Reliability, […]

The post Implementing AWS Well-Architected pillars with automated workflows appeared first on Dynatrace news.

]]>
Dynatrace | AWS

If you use AWS cloud services to build and run your applications, you may be familiar with the AWS Well-Architected Framework. This is a set of best practices and guidelines that help you design and operate reliable, secure, efficient, cost-effective, and sustainable systems in the cloud. The framework comprises six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, and Sustainability.

Six AWS well architected pillars

But how can you ensure that your applications meet these pillars and deliver the best outcomes for your business? And how can you verify this performance consistently across a multicloud environment that also uses Microsoft Azure and Google Cloud Platform frameworks? Because Google offers its own Google Cloud Architecture Framework and Microsoft its Azure Well-Architected Framework, organizations that use a combination of these platforms triple the challenge of integrating their performance frameworks into a cohesive strategy.

This is where unified observability and Dynatrace Automations can help by leveraging causal AI and analytics to drive intelligent automation across your multicloud ecosystem. The Dynatrace platform approach to managing your cloud initiatives provides insights and answers to not just see what could go wrong but what could go right. For example, optimizing resource utilization for greater scale and lower cost and driving insights to increase adoption of cloud-native serverless services.

In this blog post, we’ll demonstrate how Dynatrace automation and the Dynatrace Site Reliability Guardian app can help you implement your applications according to all six AWS Well-Architected pillars by integrating them into your software development lifecycle (SDLC).

Dynatrace AutomationEngine workflows automate release validation using AWS Well-Architected pillars

With Dynatrace, you can create workflows that automate various tasks based on events, schedules or Davis problem triggers. Workflows are powered by a core platform technology of Dynatrace called the AutomationEngine. Using an interactive no/low code editor, you can create workflows or configure them as code. These workflows also utilize Davis®, the Dynatrace causal AI engine, and all your observability and security data across all platforms, in context, at scale, and in real-time.

One of the powerful workflows to leverage is continuous release validation. This process enables you to continuously evaluate software against predefined quality criteria and service level objectives (SLOs) in pre-production environments. You can also automate progressive delivery techniques such as canary releases, blue/green deployments, feature flags, and trigger rollbacks when necessary.

This workflow uses the Dynatrace Site Reliability Guardian application. The Site Reliability Guardian helps automate release validation based on SLOs and important signals that define the expected behavior of your applications in terms of availability, performance errors, throughput, latency, etc. The Site Reliability Guardian also helps keep your production environment safe and secure through automated change impact analysis.

But this workflow can also help you implement your applications according to each of the AWS Well-Architected pillars. Here’s an overview of how the Site Reliability Guardian can help you implement the six pillars of AWS Well-Architected.

AWS well-architected six pillars workflow
A Dynatrace Workflow that uses Dynatrace Site Reliability Guardian to implement the six AWS well-architected pillars

AWS Well-Architected pillar #1: Performance efficiency

The performance efficiency pillar focuses on using computing resources efficiently to meet system requirements, maintaining efficiency as demand changes, and evolving technologies.

A study by Amazon found that increasing page load time by just 100 milliseconds costs 1% in sales. Storing frequently accessed data in faster storage, usually in-memory caching, improves data retrieval speed and overall system performance. Beyond efficiency, validating performance thresholds is also crucial for revenues.

Once configured, the continuous release validation workflow powered by the Site Reliability Guardian can automatically do the following:

  • Validate if service response time, process CPU/memory usage, and so on, are satisfying SLOs
  • Stop promoting the release into production if the error rate in the logs is too high
  • Notify the SRE team using communication channels
  • Create a Jira ticket or an issue on your preferred Git repository if the release is violating the set thresholds for the performance SLOs
Pillar #1 performance efficiency of the AWS well architected pillars
The continuous release validation workflow powered by Dynatrace Site Reliability Guardian automatically verifies performance efficiency validation success and threshold violation cases

SLO examples for performance efficiency

The following examples show how to define an SLO for performance efficiency in the Site Reliability Guardian using Dynatrace Query Language (DQL).

Validate if response time is increasing under high load utilizing OpenTelemetry spans

fetch spans 
| filter endpoint.name == "/api/getProducts" 
| filter k8s.namespace.name == "catalog" 
| filter k8s.container.name == "product-service" 
| filter http.status_code == 200 
| summarize avg(duration) // in milliseconds
fetch spans result for AWS well architected pillar #1
* Please note that the Traces on Grail feature is currently in private preview, and the DQL syntax is subject to change.

Check if process CPU usage is in a valid range

timeseries val = avg(dt.process.cpu.usage) 
,filter in(dt.entity.process_group_instance, "PROCESS_GROUP_INSTANCE-ID") 
| fields avg = arrayAvg(val) // in percentage

CPU result

AWS Well-Architected pillar #2: Security

The security pillar focuses on protecting information system assets while delivering business value through risk assessment and mitigation strategies.

The continuous release validation workflow powered by Site Reliability Guardian can automatically do the following:

  • Check for vulnerabilities across all layers of your application stack in real-time, getting help from Dynatrace Davis Security Score as a validation metric
  • Block releases if they do not meet the security criteria
  • Notify the security team of the vulnerabilities in your application and create an issue/ticket to track the progress
Davis Security Score for AWS well architected pillar #2, security
Davis Security Score against third-party vulnerabilities

SLO examples for security

The following examples show how to define an SLO for security in the Site Reliability Guardian using DQL.

Runtime Vulnerability Analysis for a Process Group Instance – Davis Security Assessment Score

fetch events 
| filter event.kind == "SECURITY_EVENT" 
| filter event.type == "VULNERABILITY_STATE_REPORT_EVENT" 
| filter event.level == "ENTITY" 
| filter in("PROCESSGROUP_INSTANCE_ID",affected_entity.affected_processes.ids) 
| sort timestamp, direction:"descending" 
| summarize  
{  
status=takeFirst(vulnerability.resolution.status), 
score=takeFirst(vulnerability.davis_assessment.score), 
affected_processes=takeFirst(affected_entity.affected_processes.ids) 
}, 
by: {vulnerability.id, affected_entity.id} 
| filter status == "OPEN"  
| summarize maxScore=takeMax(score)

security score result

AWS Well-Architected pillar #3: Cost optimization

The cost optimization pillar focuses on avoiding unnecessary costs and understanding managing tradeoffs between cost capacity performance.

The continuous release validation workflow powered by the Site Reliability Guardian can automatically do the following:

  • Detect underutilized and/or overprovisioned resources in Kubernetes deployments considering the container limits and requests
  • Determine the non-Kubernetes-based applications that underutilize CPU, memory, and disk
  • Simultaneously validate if performance objectives are still in the acceptable range when you reduce the CPU, memory, and disk allocations
A graphic that shows the cost optimization without affecting the application performance for AWS well-architected pillar #3, cost performance
A graph that shows the cost optimization without affecting the application performance

SLO examples for cost optimization

The following examples show how to define an SLO for cost optimization in the Site Reliability Guardian using DQL.

Reduce CPU size and cost by checking CPU usage

To reduce CPU size and cost, check if CPU usage is below the SLO threshold. If so, test against the response time objective under the same Site Reliability Guardian. If both objectives pass, you have achieved your cost reduction on CPU size.

Screenshot of cost performance objective SLO in support of the AWS well-architected pillars

Here are the DQL queries from the image you can copy:

timeseries cpu = avg(dt.containers.cpu.usage_percent), 
filter: in(dt.containers.name, "CONTAINER-NAME") 
| fields avg = arrayAvg(cpu) // in percentage
fetch logs 
| filter k8s.container.name == "CONTAINER-NAME" 
| filter k8s.namespace.name == "CONTAINER-NAMESPACE" 
| filter matchesPhrase(content, "/api/uri/path") 
| parse content, "DATA '/api/uri/path' DATA 'rt:' SPACE? FLOAT:responsetime "  
| filter isNotNull(responsetime) 
| summarize avg(responsetime) // in milliseconds

Reduce disk size and re-validate

If the SLO specified below is not met, you can try reducing the size of the disk and then validating the same objective under the performance efficiency validation pillar. If the objective under the performance efficiency pillar is achieved, it indicates successful cost reduction for the disk size.

timeseries disk_used = avg(dt.host.disk.used.percent), 
filter: in(dt.entity.host,"HOST_ID") 
| fields avg = arrayAvg(disk_used) // in percentage

Disk used results

AWS Well-Architected pillar #4: Reliability

The reliability pillar focuses on ensuring a system can recover from infrastructure or service disruptions and dynamically acquire computing resources to meet demand and mitigate disruptions such as misconfigurations or transient network issues.

The continuous release validation workflow powered by the Site Reliability Guardian can automatically do the following:

  • Monitor the health of your applications across hybrid multicloud environments using Synthetic Monitoring and evaluate the results depending on your SLOs
  • Proactively identify potential availability failures before they impact users on production
  • Simulate failures in your AWS workloads using Fault Injection Simulator (FIS) and test how your applications handle scenarios such as instance termination, CPU stress, or network latency. SRG validates the status of the resiliency SLOs for the experiment period.
World map showing reliability metrics in support of AWS well architected pillar #4, reliability
Application availability validation across the world using Dynatrace Synthetic monitoring

SLO examples for reliability

The following examples show how to define an SLO for reliability in the Site Reliability Guardian using DQL.

Success Rate – Availability Validation with Synthetic Monitoring

fetch logs 
| filter log.source == "logs/requests" 
| parse content,"JSON:request" 
| fieldsAdd httpRequest = request[httpRequest] 
| fieldsAdd httpStatus = httpRequest[status] 
| fieldsAdd success = toLong(httpStatus < 400) 
| summarize successRate = sum(success)/count() * 100 // in percentage

logs request result

Number of Out of memory (OOM) kills of a container in the pod to be less than 5

timeseries oom_kills = avg(dt.kubernetes.container.oom_kills), 
filter: in(k8s.cluster.name,"CLUSTER-NAME") and in(k8s.namespace.name,"NAMESPACE-NAME") and in(k8s.workload.kind,"statefulset") and in (k8s.workload.name,"cassandra-workload-1") 
| fields sum = arraySum(oom_kills) // num of oom_kills

Reliability OOM kills result

AWS Well-Architected pillar #5: Operational excellence

The operational excellence pillar focuses on running and monitoring systems to deliver business value and continually improve supporting processes and procedures.

With continuous release validation workflow powered by the Site Reliability Guardian, you can:

  • Automatically verify service or application changes against key business metrics such as customer satisfaction score, user experience score, and Apdex rating
    Apdex rating in support of AWS well-architected pillar #5, operational excellence
  • Enhance collaboration with targeted notifications of relevant teams using the Ownership feature
  • Create an issue on your preferred Git repository to track and resolve the invalidated SLOs
  • Trigger remediation workflows based on events such as service degradations, performance bottlenecks, security vulnerabilities
  • Validate your CI/CD performance over time, considering the execution times, pipeline performance, failure rate, etc., which shows your operational efficiency in your software delivery pipeline.Screenshot of pipeline metrics in support of AWS well-architected pillar #5, operational excellence

SLO examples for operational excellence

The following examples show how to define an SLO for operational excellence in the Site Reliability Guardian.

Apdex rating validation of a web application

  1. Navigate to “Service-level objectives” and click on “Add new SLO” button
  2. Select “User experience” as a template. It will auto-generate the metric expression as the following:
    (100)*(builtin:apps.web.actionCount.category:filter(eq("Apdex category",SATISFIED)):splitBy())/(builtin:apps.web.actionCount.category:splitBy())
  3. Replace your application name in the entityName attribute:
    type("APPLICATION"),entityName("APPLICATION-NAME")
  4. Add a success criteria depending on your needs
    Success criteria for operational excellence
  5. Reference this SLO in your Site Reliability Guardian objective

AWS Well-Architected pillar #6: Sustainability

The sustainability pillar focuses on minimizing environmental impact and maximizing the social benefits of cloud computing.

The continuous release validation workflow powered by the Site Reliability Guardian can automatically do the following:

  • Measure and evaluate carbon footprint emissions associated with cloud usage
  • Leverage observability metrics to identify underutilized resources for reducing energy consumption and waste emissions
Sustainability dashboard supporting AWS well-architected pillar #6, Sustantainability
The Dynatrace Carbon Impact Dashboard evaluates the carbon impact of the resources in the cloud

SLO examples for sustainability

The following examples show how to define an SLO for sustainability in the Site Reliability Guardian using DQL.

Carbon emission total of the host running the application for the last 2 hours

fetch bizevents, from: -2h 
| filter event.type == "carbon.report" 
| filter dt.entity.host == "HOST-ID" 
| summarize toDouble(sum(emissions)), alias:total // total CO2e in grams

carbon emissions results supporting the AWS well-architected pillar #6, sustainability

Under-utilized memory resource validation

timeseries memory=avg(dt.containers.memory.usage_percent), by:dt.entity.host 
| filter dt.entity.host == "HOST-ID" 
| fields avg = arrayAvg(memory) // in percentage

memory resource validation for AWS well-architected pillar #6, sustainability

Implementing AWS Well-Architected Framework with Dynatrace: A practical guide

To help you get started with implementing the AWS Well-Architected Framework using Dynatrace, we’ve provided a sample workflow and SRGs in our official repository. This resource offers a step-by-step guide to quickly set up your validation tools and integrate them into your software development lifecycle.

Quick Implementation: Follow the instructions in our Dynatrace Configuration as Code Samples repository to deploy the sample workflow and SRGs. This will enable you to immediately begin validating your applications against the AWS Well-Architected pillars.

For more about how Site Reliability Guardian helps organizations automate change impact analysis, performance, and service level objectives, join us for the on-demand Observability Clinic, Site Reliability Guardian with DevSecOps activist Andreas Grabner.

The post Implementing AWS Well-Architected pillars with automated workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/implementing-aws-well-architected-pillars/feed/ 0
Site reliability done right: 5 SRE best practices that deliver on business objectives https://www.dynatrace.com/news/blog/site-reliability-done-right/ https://www.dynatrace.com/news/blog/site-reliability-done-right/#respond Wed, 31 May 2023 15:56:55 +0000 https://www.dynatrace.com/news/?p=57959 Site Reliability Guardian, CrowdStrike outage

Site reliability engineering has emerged as a critical discipline for organizations seeking the benefits of digital transformation.

The post Site reliability done right: 5 SRE best practices that deliver on business objectives appeared first on Dynatrace news.

]]>
Site Reliability Guardian, CrowdStrike outage

Keeping pace with modern digital transformation requires ensuring that applications are responsive, resilient, and always available amid increased complexity. As a result, site reliability has emerged as a critical success metric for many organizations.

Site reliability engineering (SRE) has recently become a critical discipline in recent years as the world has shifted in favor of web-based interactions. Mobile retail e-commerce spending in the U. S. surpassed $387 billion in 2022, more than double the figure of three years earlier. The volume of travel spending booked online is expected to reach nearly $1.5 trillion by 2027, up from $800 billion in 2021. With so many of their transactions occurring online, customers are becoming more demanding, expecting websites and applications to always perform perfectly. One recent report found that 32% would leave their favorite brand after just one bad experience. Website load times have been found to have a direct correlation with conversion rates.

This shift is leading more organizations to hire site reliability engineers to guarantee the reliability and resiliency of their services. But the transition to SRE maturity is not always easy.

How site reliability engineering affects organizations’ bottom line

SRE applies the disciplines of software engineering to infrastructure management, both on-premises and in the cloud. The practice uses continuous monitoring and high levels of automation in close collaboration with agile development teams to ensure applications are highly available and perform without friction.

According to the emerging trends from the global shift towards web-based interactions, IT infrastructure performance has a dramatic impact on the organizations’ bottom-line business goals. Uptime Institute’s 2022 Outage Analysis report found that over 60% of system outages resulted in at least $100,000 in total losses, up from 39% in 2019. More than one in seven outages cost more than $1 million.

Maintaining reliable uptime and consistent service quality has become more complex as organizations expand their computing footprints across multiple data centers and in the cloud. Microservices-based architectures and software containers enable organizations to deploy and modify applications with unprecedented speed. However, cloud complexity has made software delivery challenging. There are now many more applications, tools, and infrastructure variables that impact an application’s performance and availability.

Understanding the interactions between these factors heavily influences decisions about whether and when to promote a new release into production. That’s why good communication between SREs and DevOps teams is important. By automating and accelerating the service-level objective (SLO) validation process and quickly reacting to regressions in service-level indicators (SLIs), SREs can speed up software delivery and innovation.

Understanding the goal of “five-nines” availability

The guiding principle of SRE has long been “five-nines” availability, meaning systems are operative 99.999% of the time. As organizations distribute workloads among a greater number of cloud environments, that goal has become harder to attain because more variables are involved in the computing equation. The growing amount of data processed at the network edge, where failures are more difficult to prevent, magnifies complexity.

Visibility and automation are two of the most important SRE tools. The Dynatrace 2022 Global CIO Report found that 71% of top IT executives say the explosion of data produced by cloud-native technology stacks is beyond human ability to manage, and more than three-quarters say their IT environment changes once every minute or less. The takeaway is clear: IT environments are now too complex to manage without automation and AI. Without these capabilities, achieving five-nines availability will become close to impossible.

Aligning site reliability goals with business objectives

Because of this, SRE best practices align objectives with business outcomes. The following three metrics are commonly used to measure success:

  • Service-level agreements (SLAs). These metrics are the product of an agreement between the service provider and customer that certain measurable levels of service will be delivered.
  • Service-level objectives (SLOs). These metrics are the factors and service levels that must be achieved for each activity, function, and process to deliver on the SLA. These can include business metrics, such as conversion rates, as well as technical measures like underlying CPU availability. They’re typically expressed as percentages, such as 99.5% availability.
  • Service-level indicators (SLIs). At the lowest level, SLIs provide a view of service availability, latency, performance, and capacity across systems.

5 SRE best practices

Let’s break down SRE best practices into the following five major steps:

1. Start looking for signals

Begin by monitoring the “four golden signals” that were originally outlined in Google’s SRE handbook:

  • Latency: the time it takes to serve a request
  • Traffic: the total number of requests across the network
  • Errors: the number of requests that fail
  • Saturation: the load on the network and servers

2. Identify KPIs

Next, create a list of the key performance indicators (KPIs) that are important to the business. These may include technical metrics such as response times to search query referrals, page load times, and error message frequency. They may also include business metrics influenced by performance, such as shopping cart abandonment and page views per customer.

3. Establish SLOs

Drawing on the KPIs, you can now create a list of SLOs. Remember that less is more. SLOs should directly relate to the SLA or business objective. Defining too many SLOs creates more work without a clear business impact.

Make SLOs realistic. If they’re intentionally set low to avoid SLA violations, they won’t provide an accurate picture of how systems are impacting user experience. If set too high, they can drive higher costs and effort for little incremental gain.

4. Identify key stakeholders

Once you’ve ensured SLAs and SLOs are realistic, begin recruiting stakeholders. For example, a panel of customers may occasionally provide feedback on service quality and performance. Business stakeholders can understand how SLOs relate to business results using historical data trends. Everyone needs to agree on SLO targets; otherwise, the organization will risk failing to deliver on its SLAs.

5. Automate workflows wherever possible

Automation is critical to achieving the business agility that digital transformation demands. A powerful observability solution collects relevant SLIs and evaluates SLOs automatically. It can also automatically generate alerts before SLO violation and even automatically repair many problems. Automation also enables tools to move into developers’ hands so they can make decisions about deploying code without needing to involve operations teams.

Automating SRE best practices with Site Reliability Guardian

As part of the early 2023 rollout of its new AppEngine low-code toolset for creating custom, compliant, and intelligent data-driven applications, Dynatrace introduced Site Reliability Guardian (SRG) for automated change impact analysis

AppEngine is a unique solution that consolidates observability, security, and business data with full context and dependency mapping to simplify the creation of intelligent applications and integrations. It enables teams, for the first time, to leverage causal AI for customized applications that address specific use case requirements.

SRG enriches the Dynatrace platform’s value for SREs by automating change impact analysis. It detects regressions and deviations from previously observed behavior across metrics such as latency, traffic, error rates, saturation, security coverage, vulnerability risk levels, and memory consumption.

SRG also enriches the Dynatrace platform’s value to DevOps teams who can automate release validation in pre-production environments to ensure only high-quality, highly secure software moves to production. Should the SRG detect any violations of key objectives or metrics, the CI/CD pipeline tool can halt the delivery of the artifact.

For each critical service or application, SRG can automatically monitor for golden signals, validate SLOs, and probe for security vulnerabilities before and after deployment or configuration changes. It can also send targeted notifications to the service and application owners as needed. The result is safer, more secure releases for DevOps teams and less overhead for SREs.

Keeping the focus on site reliability

SRE is a critical discipline for organizations to master as they navigate the transition to digital business powered by data-driven decisions. Tight integration with business objectives and automation now makes it possible for organizations to proactively monitor their digital presence to ensure the highest levels of availability, responsiveness, and customer experience.

To learn more about the challenges SREs are facing and how organizations are tackling them, download the free State of SRE Report.

The post Site reliability done right: 5 SRE best practices that deliver on business objectives appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/site-reliability-done-right/feed/ 0
What is SRE (site reliability engineering)? And what do site reliability engineers do? https://www.dynatrace.com/news/blog/what-is-site-reliability-engineering/ https://www.dynatrace.com/news/blog/what-is-site-reliability-engineering/#respond Thu, 26 Jan 2023 20:28:44 +0000 https://www.dynatrace.com/news/?p=42456 Site Reliability Engineering highlights reliability, scalability, and efficiency.

As more organizations adopt cloud-based computing and the demand for digital services increases, site reliability engineering (SRE) practices have become essential. These practices help organizations meet service level agreements (SLAs) for availability, performance, user experience, and business KPIs. But what exactly is SRE, and what do site reliability engineers do? What is site reliability engineering? […]

The post What is SRE (site reliability engineering)? And what do site reliability engineers do? appeared first on Dynatrace news.

]]>
Site Reliability Engineering highlights reliability, scalability, and efficiency.

As more organizations adopt cloud-based computing and the demand for digital services increases, site reliability engineering (SRE) practices have become essential. These practices help organizations meet service level agreements (SLAs) for availability, performance, user experience, and business KPIs.

But what exactly is SRE, and what do site reliability engineers do?

What is site reliability engineering?

Site reliability engineering (SRE) is the practice of applying software engineering principles to operations and infrastructure processes to help organizations create highly reliable and scalable software systems. As a discipline, SRE focuses on improving software system reliability across key categories, including availability, performance, latency, efficiency, capacity, and incident response. Those who perform the tasks involved are known as site reliability engineers.

The term “site reliability engineering” was coined in 2003 by Google VP of Engineering Ben Sloss, who famously noted on his LinkedIn profile, “If Google ever stops working, it’s my fault.” According to Google, “SRE is what you get when you treat operations as a software problem.”

Although every organization and software system is unique, it’s important to understand the fundamentals of SRE and the skills and mindset of its engineers as you think about optimizing the reliability and overall quality of your software.

The benefits of site reliability engineering

Site reliability engineering highlights reliability, scalability, and efficiency. The benefits of SRE include:

  • Increased reliability and uptime: SRE focuses on preventing and mitigating incidents to ensure that systems and applications are always available and performant.
  • Improved scalability: By optimizing resource usage and minimizing waste, SRE can help organizations scale their infrastructure and applications more efficiently.
  • Improved user experience: SRE can ensure that applications and services are always available and responsive, which can directly impact customer satisfaction, brand reputation, and revenue.
  • Continuous improvement: SRE highlights the use of data and metrics to identify areas for improvement and drive ongoing optimization and innovation.
  • Increased security: SRE can help to guarantee that systems and applications are secure and compliant with industry standards and regulations.
  • Predictable performance: By monitoring and analyzing usage patterns, SRE can help to predict and prevent performance issues before they occur, ensuring that systems and applications perform predictably and consistently.
  • Cost savings: SRE can reduce costs by automating routine tasks and optimizing resource usage, reducing the need for manual intervention and saving time and money.
  • Collaboration between development and operations teams: SRE emphasizes cross-functional teams and shared ownership of reliability and performance, promoting a culture of collaboration and accountability.

Diagram showing the 8 benefits of site reliability engineering

Five things to know about site reliability engineering

Implementing SRE may seem like a large undertaking, but the operational shifts it requires offer lasting benefits that promote greater efficiency and provide a positive return on investment.

1. SRE focuses on automation

A major goal of SRE is to reduce duplication or redundancy of effort as much as possible. SRE teams focus on automating manual tasks, such as provisioning access and infrastructure, setting up accounts, and building self-service tools. As a result, development teams can focus on delivering features, and operations teams can focus on managing infrastructure.

Automating processes is even more critical as organizations speed up delivery of new features into production. On one hand, speed comes from DevOps teams who leverage automation to increase continuous integration and continuous delivery (CI/CD). On the other hand, the move to microservice architectures and the adoption of cloud-native technology, containers, Kubernetes, and serverless architectures offer even more ways to delivery smaller changes faster. These methods increase efficiency and speed, but also demand consistent, repeatable processes that reduce risk and provide feedback loops for measuring operations, so teams can identify areas for improvement.

2. SRE bridges the gap between Dev and Ops

Everything the organization does in the value stream process should answer the question “how do we ensure this runs in production reliably?” SREs drive resiliency-based engineering. They can become mentors and ensure that resiliency is a top priority for both developers and operations.

Applying the DevOps mindset and skills to software reliability helps reduce silos between development and operations teams by sharing responsibility for detecting reliability and performance issues early in the development life cycle. Collaboration between developers, operations, and product owners enables site reliability engineers to define and meet uptime and availability targets.

3. SRE drives a “shift-left” mindset

SRE is a constantly evolving discipline, presenting opportunities to build methods, policies, and processes into the delivery pipeline that allow applications to “auto-remediate” or users to solve their own problems. A shift-left mindset means SREs can embed reliability principles from Dev to Ops, baking reliability and resiliency into each process, app, and code change to improve the quality of software that goes to production.

Here are some ways SRE helps to drive a “shift-left” mindset:

  • Develop quality gates based on production-level service level objectives (SLOs) to detect issues earlier in the development cycle.
  • Automate build testing and validation using service-level indicators (SLIs) and SLOs
  • Influence architectural decisions during initial design stages to ensure resiliency and scale at the outset of software development.

The goal is to take early, proactive steps to ensure quality and reliability are built-in from the beginning. SRE can influence processes more broadly and expand to coordinating testing across the enterprise in support of CI/CD practices.

To learn more about how Dynatrace enables SRE with “shift-left SLIs,” join us for the on-demand performance clinic Automated SRE-driven performance engineering with Dynatrace.

4. SRE builds services and tools to help operations and support

Traditionally, a major goal of operations teams is to improve uptime. This single-dimensional approach looks for the coveted “five nines” of uptime, or 99.999%, which translates to just over five minutes of downtime per year.

But the higher frequency of change in distributed cloud-native environments requires a multi-dimensional approach.

The goal of SRE is to enable higher change rates while maintaining resiliency and that coveted 99.999% uptime. In multicloud environments, resiliency is measured across key metrics such as performance, user experience, responsiveness, conversion rates, and so on. SRE teams need to build and implement services that improve operations and facilitate the release process across all these areas. This can be anything from adjusting monitoring and alerting to making code changes in production. Site reliability engineers often build custom tooling from scratch to meet specific needs in the software delivery or incident management workflow.

Adopting an SRE approach also requires standardizing the technologies and tools teams use. Standardization makes it easier to manage operations and reduces the burden of managing incompatible technologies, which gives teams more time to collaborate and innovate.

5. SRE requires a cultural change

Because SRE is a practice, it requires changing how teams across multiple disciplines communicate, solve problems, and implement solutions. To adopt a successful SRE culture, organizations must adopt new approaches to managing risk. It also means they must adapt governance processes, invest in hiring, and educate a collaborative workforce that’s versed in engineering and operations and learns and adapts quickly.

Organizations can then integrate these skilled engineers at key points in the DevOps lifecycle. In development and testing teams, SRE specialists develop automation that helps developers test early and often without impeding agile delivery schedules. At a system level, SRE specialists develop tooling that coordinates releases and launches, evaluates system architecture readiness, and meets system-wide SLOs. At a governance level, SRE specialists help to define and oversee enterprise architecture, establish best practices, and select tools and resources that support company-wide site reliability.

What does a site reliability engineer do?

To get an expert’s view on what site reliability engineers do, I asked our DevOps Activist, Andi Grabner.

“Site reliability engineers use good practices around software engineering to provide resilient infrastructure and resilient services to their organizations and the people that actually deliver new applications,” he explains. He also notes that SREs often come from traditional operations roles, such as systems engineers who keep systems up and running. “Site reliability engineers ensure systems stay reliable, resilient, and available,” he adds.

State of SRE Report

We asked 450 SREs across various industries to share their unfiltered perspective into how site reliability engineering (SRE) is evolving as a discipline. The report uncovers the challenges SREs must overcome, and what the future of SRE looks like.

Typical expectations for SREs

Typically, SREs are tasked with ensuring that the speed of delivery doesn’t result in security, service, or solution interruption. But as Grabner notes, “Expectations are a little different for every company. There’s no golden rule. Many are responsible for monitoring and observability and maintaining systems — and providing automation to spin up required environments.”

Grabner highlights the role of SREs in providing the frameworks and platforms for service and application deployment. “When things go wrong, SREs often take on the role of first-line defenders if there’s an alert,” he says. “In a great organization, they don’t do it alone — they constantly work with and within individual application teams to deal with apps that are under fire.”

Perhaps the most important role of SREs is architecting resiliency. “You can’t buy resiliency as-a-service,” Grabner observes. “You have to architect for it by building systems that are resilient by design.” This architectural approach recently helped Dynatrace to withstand an AWS outage in Germany. Automatic delivery, resiliency, and auto-remediation helped to ensure critical systems weren’t affected.

Roles and responsibilities of an SRE

The role of an SRE has become invaluable as technology continues to evolve and businesses become more dependent on digital infrastructure. Here are some of the standard responsibilities of an SRE:

  1. Monitoring and Alerting
    One of the primary duties of an SRE is to monitor a company’s digital infrastructure. This involves setting up monitoring tools and systems to detect concerns before they become significant problems. SREs set up alert systems that notify the appropriate people when issues are detected.
  2. Incident Response
    The SRE responds quickly and effectively when issues are detected by identifying the root cause, developing and implementing a plan, and communicating with relevant stakeholders.
  3. Automation and Tooling
    SREs develop and maintain the tools and systems used to manage a company’s digital infrastructure. This includes developing automation scripts to streamline processes and reduce the risk of human error. SREs also identify areas where tooling can be improved and create new tools to meet the changing needs of the business.
  4. Capacity Planning
    SREs ensure that a company’s digital infrastructure can meet the needs of the business. This involves analyzing usage patterns to predict and guarantee the required capacity to meet future demand.
  5. Collaboration
    SREs work closely with other teams to ensure the company’s digital infrastructure is reliable, scalable, and secure.
Roles and Responsibilities of an SRE
Roles and Responsibilities of an SRE

What makes a great SRE?

Great SREs are risk-takers, tinkerers, and innovators. They figure out what it takes to scale a system from 100 users to 100,000 users to 1,000,000 users while maintaining uptime and resiliency. They’re systems thinkers who consider how decisions made in development affect production environments, and how the needs of production systems can influence design.

This requires constant testing, accepting failure, and adapting, automating repeatable processes. Successful SREs bring a resiliency and adaptation mindset to every situation.

Grabner highlights the need for SREs to learn from their mistakes. “Some companies run ‘chaos days’ to deal with worst-case scenarios to understand what could happen and how to deal with it,” he says.

Automation is another marker of SRE success. “People excel in this role when they try to automate all the tasks that can and must be automated,” says Grabner. “This frees them up to deliver true innovation.” He notes that while anyone can innovate under the right circumstances, teams are often held back by the “toil” of manual and repetitive tasks. “Your goal is to automate yourself out of your current role and into your next role.”

Finally, Grabner made it clear that SREs can’t operate in isolation. “You need to let people educate themselves with new technologies and practices,” he says. “Show the world. Don’t keep secrets — be open and share your own learnings, along with learning from others. There are a lot of great conferences out there — it’s worth getting inspired by what others do and inspiring others with what you do.”

DevOps vs. SRE

Where DevOps teams focus on streamlining change, SREs help ensure these changes don’t increase overall failure rates. In effect, they’re two sides of the same coin: DevOps automates speed, while SRE automates reliability. “It’s a balance between speed and safety,” Grabner says.

He sees DevOps processes as moving left to right along the development life cycle, using automation to speed up new capabilities that are typically measured by deployment frequency and lead time for changes. In comparison, SRE moves right to left using production-level requirements in development, with a focus on limiting failure rates and reducing the time required to restore service. “SRE is about making sure that even though there is a lot of change, these changes don’t break things.”

Grabner sees SRE and DevOps overlapping when it comes to SLOs. “SLOs are all about supporting business goals,” he said. “Companies may need systems at 99% reliability. They may want to increase their user base or improve the end-user experience.” Satisfying these goals is the role of DevOps. “But underneath these goals are technical goals that are specific to your objectives,” he says. “They contribute to business success with the right features at the right time and help you cope with change.” Delivering on these goals is the job of SRE staff. As a result, “SLOs are a great way to bring DevOps and SREs together.”

In another blog post you can learn more about the differences and similarities of SRE vs DevOps.

Solving for site reliability

Site reliability isn’t and never will be a “solved problem”. New services and applications combined with evolving enterprise demands mean there’s always work for SRE teams, and there’s always room for improvement.

As Grabner notes, when it comes to improving SRE impact, “The biggest thing is to be open and share your own learnings, along with learning from others. There are a lot of great conferences out there — it’s worth getting inspired by what others do and inspiring others with what you do.” He also highlights the need to learn from your mistakes. “Some companies run ‘chaos days’ to deal with worst-case scenarios to understand what could happen and how to deal with it.” Finally, Grabner made it clear that SREs can’t operate in isolation. “You need to let people educate themselves with new technologies and practices. Show the world. Don’t keep secrets — don’t see it as a silo.”

Looking for a solution to up-level your SRE practices? Dynatrace can help. Providing automatic and intelligent observability for even the most complex distributed cloud environments, the Dynatrace Software Intelligence Platform empowers SRE and DevOps teams to identify problems before they occur. Driven by continuous automation with AI at its core, Dynatrace delivers precise root-cause answers to site reliability issues at every step of the software development lifecycle. From early development in pre-production environments through delivery and operations in production environments, Dynatrace helps SRE teams improve reliability, availability, and latency, and mitigate the business impact of service outages and slowdowns.

Watch webinar

You can also join us for the on-demand webinar The State of SRE in 2022 on to learn more about SRE trends in 2022, how to become a better SRE and hear from three panelists about the state of SRE in their organizations.

The post What is SRE (site reliability engineering)? And what do site reliability engineers do? appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/what-is-site-reliability-engineering/feed/ 0
Site reliability engineering: Six SRE trends to unleash DevOps innovation https://www.dynatrace.com/news/blog/six-site-reliability-engineering-trends/ https://www.dynatrace.com/news/blog/six-site-reliability-engineering-trends/#respond Thu, 02 Jun 2022 08:31:17 +0000 https://www.dynatrace.com/news/?p=51742 site reliability engineering and sre for DevOps maturity

Site reliability engineering (SRE) continues to gain popularity as organizations embrace hybrid cloud strategies and IT automation at scale. By applying software engineering principles to operations and infrastructure practices, SRE enables organizations to streamline and automate IT processes. In the Dynatrace State of SRE Report: 2022 Edition, 88% of surveyed site reliability engineers (SREs) said […]

The post Site reliability engineering: Six SRE trends to unleash DevOps innovation appeared first on Dynatrace news.

]]>
site reliability engineering and sre for DevOps maturity

Site reliability engineering (SRE) continues to gain popularity as organizations embrace hybrid cloud strategies and IT automation at scale. By applying software engineering principles to operations and infrastructure practices, SRE enables organizations to streamline and automate IT processes.

In the Dynatrace State of SRE Report: 2022 Edition, 88% of surveyed site reliability engineers (SREs) said there is now greater recognition of the strategic importance of their role in business success than there was three years ago. SRE is becoming an essential discipline in organizations that use DevOps (the combination of development and operations) and agile methodologies.

SRE adoption is growing, yet gaps remain. In order to unleash the innovation organizations need to evolve their SRE approaches. The report uncovers six site reliability engineering trends that will help organizations get the most from DevOps practices.

1. MTTR reduction remains top of the list for SREs

SREs improve the reliability of production systems, and reducing mean-time-to-repair (MTTR) is their top priority. But research shows that 60% of SREs find they spend most of their time building and maintaining automation code. While increasing automation is a key goal, organizations will lose their derived efficiency if enablement is arduous and time-consuming.

Much of the problem stems from how site reliability engineering teams build automation for DevOps workflows. Often, teams handle this on a case-by-case basis because their tooling doesn’t come with automation built in and doesn’t offer everything-as-code capabilities. As a result, they must build a layer of automation on top of their tooling.

Over time, this creates a complex web of code that becomes more difficult to scale across the DevOps pipeline. If SRE teams don’t identify a more efficient approach to DevOps from day one, they will undoubtedly find more of their time drained in the future. That time sink results in developers chasing down bugs and other code issues rather than focusing on strategic, revenue-generating work.

This underscores the need for SREs to work with DevOps teams, developers, and architects to ensure that software not only meets a business need but also is resilient and automatable by default. Enabling teams to easily integrate new automation capabilities with existing tools and workflows reduces manual effort and improves engineering practices.

2. A shift to SRE-driven engineering takes hold

More than half (51%) of SREs say they dedicate significant time to influencing architectural design decisions to improve reliability. This suggests that all departments have made progress toward SRE-driven engineering across to improve reliability, resiliency, and security. But there’s still a long way to go.

The most mature organizations embrace site reliability engineering practices that include developers who have fought the battles. They understand what it takes to build systems that can scale from 10 users to 1,000, or from 1 million to 10 million users. Integrating these developers with the design process provides insight that enables architects to incorporate reliability from day one.

3. Security is a core pillar of site reliability engineering

SREs are also making progress in extending organization-wide DevSecOps approaches to ensure organizations can restore systems quickly after discovering a vulnerability. More than two-thirds (68%) of SREs say they expect their role in security to become even more central in the future. This trend will increase as organizations continue using third-party libraries for cloud-native application development.

As the Log4Shell vulnerability demonstrated when it emerged in December 2021, third-party code libraries face significant security risks. Site reliability engineering teams are critical to identifying those flaws and eliminating them for fail-safe IT protection and to minimize cloud and third-party risks.

4. SREs need the freedom to experiment

While more than half (52%) of SREs dedicate significant of time to designing experiments and tests to reduce the risk of production failure, only 1 in 10 highlights this as their top priority.

Experimentation is critical to SRE. Teams still need to make progress to ensure they have more time available for these tasks. For SREs to deliver more strategic business value, engineers must streamline tasks that involve intensive manual effort.

5. SREs need the license to prioritize strategic work

Although experimentation falls relatively low on their priority list, 51% of SREs say they’re encouraged to experiment. In addition, only a quarter (26%) of organizations see incremental project failure as OK. The weak emphasis on experimentation highlights the many distractions SREs face, limiting the time they have to focus on it.

Organizations must consider new strategies that enable site reliability engineering teams to focus on more strategic tasks. That way, teams have more time to play. Team leaders also need to foster a culture that accepts failure and understands that the principle “fail fast, fail often” provides the greatest competitive edge.

Unshackle SRE teams from traditional organizational structures that view IT as a cost center and a burden. Organizations can’t benefit from lessons gained from IT mistakes without openness to failure.

6. Organizations recognize and reward site reliability engineering teams

SREs must be free to challenge accepted norms and set new benchmarks for innovation-led design and engineering practices. Many organizations are making strides in this direction and have methods for rewarding SRE teams’ successes.

Nearly a third (31%) use hackathons to devise new ways to improve reliability, offering prizes to winning SRE teams. These approaches are key to encouraging a culture of experimentation that promotes the strategic value of site reliability engineering for the business.

All six of these SRE trends can help your organization advance its SRE practices. Embrace experimentation and unleash exciting innovation.

The post Site reliability engineering: Six SRE trends to unleash DevOps innovation appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/six-site-reliability-engineering-trends/feed/ 0
Shift-Left SRE: Building Self-Healing into your Cloud Delivery Pipeline https://www.dynatrace.com/news/blog/shift-left-sre-building-self-healing-into-your-cloud-delivery-pipeline/ https://www.dynatrace.com/news/blog/shift-left-sre-building-self-healing-into-your-cloud-delivery-pipeline/#respond Tue, 04 Dec 2018 13:21:49 +0000 https://www.dynatrace.com/news/?p=29539 Dynatrace employee on stage

Site Reliability Engineering is an exciting discipline in our industry. There are many aspects to building reliable and resilient systems which I won’t be able to cover in a single blog post. What I learned from many conversations with enterprise organizations is that the term “self-healing” is getting everybody excited. That’s also why I put […]

The post Shift-Left SRE: Building Self-Healing into your Cloud Delivery Pipeline appeared first on Dynatrace news.

]]>
Dynatrace employee on stage

Site Reliability Engineering is an exciting discipline in our industry. There are many aspects to building reliable and resilient systems which I won’t be able to cover in a single blog post. What I learned from many conversations with enterprise organizations is that the term “self-healing” is getting everybody excited. That’s also why I put it into the title, to get you “excited”! However – I have at least two major problems with the way self-healing is sold and thought of:

#1: Self-healing as a term is misleading!

I don’t think that there are any real self-healing systems. I think we have gotten better to automate remediation actions – and by leveraging modern monitoring tools we can execute specific remediation actions in a much smarter and efficient way. Hence, I believe we should talk about Smart- or Auto-Remediation and not necessarily Self-Healing as it is misleading!

#2: Self-healing is not just an operations activity!

In most of my discussions around self-healing, I talk with Ops teams that try to fix issues that materialized in production. While it is very important to keep our systems running even under “chaotic” conditions in production, I strongly believe we should invest just as much time and effort in preventing these issues ever making it into production. That’s where my “Shift-Left SRE” term comes from as I believe we must bake resiliency into our delivery pipelines.

I presented my thoughts on this topic at AWS re:Invent 2018. It was called Shift-Left SRE: Self-Healing with Lambda!

While my talk was targeted for an AWS crowd the concepts, I explained are applicable for any platform, cloud or technology stack.

I structured my presentation in 4 parts which is the same way I’d like to organize this blog:

  • Remediation Use Cases
  • PREVENT in CI/CD vs REPAIR in Prod
  • “Auto Remediation as Code”
  • The “Unbreakable Delivery Pipeline”

Let’s dig into it!

Remediation Use Cases

I have to be honest. I never worked in Operations, neither did I work as an SRE. I always worked in Performance Engineering, where I ended up breaking a lot of environments for different reasons, e.g: hitting the system with too much load, not cleaning up log directories before a new run or improper configuration of my test or test environment. When talking to my friends, customers and partners that work in Operations, they tell me similar root causes they see in a production environment.

Let me therefore start with a quick overview of common problem scenarios that impact production environments and the remediation actions that can be taken to bring the system back to normal:

Some of the remediation use cases that I have seen in recent months with our users
Some of the remediation use cases that I have seen in recent months with our users

This is a very short list of remediation actions – but – a list that addresses very common and basic issues in todays environments: Process Restarts, Resource (e.g: Disk) Cleanup, Revert Bad Configuration Changes, Scale Up or Down, Switch Blue vs Green, …

While these use cases have been known for years by many teams I talk with, only a handful have invested in some sort of automation. While many have built scripts – these typically get executed manually when a problem is detected. That means – there is a lot of room for improvement!

PREVENT in CI/CD vs REPAIR in Prod

While we must invest in automating these remediation scripts in production we should learn from every problem we detect in production, figure out what the root cause is, how to replicate, measure or detect that root cause in a lower level environment and automate that into the CI/CD Delivery Pipeline as a Quality Gate:

Shift-Left SRE: Add CI/CD Quality Gate checks for every production remediation action. This allows us to prevent vs repair!
Shift-Left SRE: Add CI/CD Quality Gate checks for every production remediation action. This allows us to prevent vs repair!

This approach of adding these checks in a pre-production environment requires the teams (SRE Teams, Ops, …) that currently work on auto-remediation scripts for production to also think about how these situations can be detected in a lower level environment as part of the automated delivery pipeline. Here are the tasks that from my perspective define Shifting-Left SRE:

  • Take the lessons learned from production
  • Reproduce them in Pre-Production through e.g: Chaos Monkey Tests
  • Automate the checks into CI/CD Quality Gates

In my presentation at re:Invent, I listed examples on how to detect certain use cases, what the metrics are to watch out for and how to capture these metrics:

List of Use Cases, Metrics and how to Query these metrics in your CI/CD Pipeline
List of Use Cases, Metrics and how to Query these metrics in your CI/CD Pipeline

Once we have the metrics and the tests to replicate these scenarios in a lower level environment we can add it to the pipeline as quality gate checks:

Simple add these tests and metric checks into your pipeline. This will prevent issues from ever entering production!
Simple add these tests and metric checks into your pipeline. This will prevent issues from ever entering production!

If you are interested in integrating Dynatrace captured metrics into your Jenkins, Azure, Concourse, Bamboo … pipelines check out our my recent blog posts and YouTube tutorials: Jenkins Performance Signature, Azure DevOps for Dynatrace, AWS DevOps Tutorial!

Auto-Remediation As Code

When using Dynatrace for full-stack monitoring of platforms we do not just receive alerts based on individual metrics that violated a static threshold or baseline, e.g: CPU is high or Disk is Full. The Smartscape Model, the PurePaths and the high-fidelity monitoring data we ingest and forward to our AI allows Dynatrace to detect problems and correlate all relevant events and data points to a single problem instance. All correlated information on that problem is made available to the auto-remediation as code scripts that SRE engineers wrote to handle specific problem root causes. The auto-remediation scripts do not have to do any correlation work but simply focus on fixing the root cause of a single problem vs chasing many loose ends of unrelated events you get from other monitoring or alerting tools.

The following animation visualizes the full problem/incident response workflow Once Dynatrace detects a problem we set the correct incident response actions in motion to inform the impacted teams as well as call auto-remediation as code scripts that address the root causes rather than just symptoms:

Self-Healing with Dynatrace AIOps and Auto-Remediation as Code enables the Path to Autonomous Ops
Self-Healing with Dynatrace AIOps and Auto-Remediation as Code enables the Path to Autonomous Ops

In my demo at re:Invent, I showed the self-healing part of my AWS DevOps Tutorial where I deploy a faulty version of a Service that results in a high failure rate impacting end users. During the deployment I pass the link to an AWS Lambda function that the developer (in this case me) wrote to be called in case of an issue in this service (=Auto-Remediation As Code). Once Dynatrace AI detects the issue, it calls the Lambda function that is associated with the latest deployment of the faulty service!

I encourage Operations or SRE Teams to collaborate with the engineering teams so that they can start writing their own remediation code. This code should be battle tested in a pre-prod environment when those problems get simulated that have been seen in production before. This is a great way to build resilience into the code package that comes from engineering and as it has been battle tested in CI/CD the chances of a real problem in production are very small 😊

The “Unbreakable Delivery Pipeline”

Now – if we put all the discussed concepts together in an automated delivery pipeline we can speak of what I coined “The Unbreakable Delivery Pipeline”:

  1. We stop potential negative impacting changes early
  2. We auto-validate “Auto-Remediation as Code” via a Quality Gate
  3. We use “Auto-Remediation as Code” to “self-heal” production

We “Shift-Left” every problem we haven’t yet auto-remediated

The Unbreakable Delivery Pipeline with Shift-Left SRE Quality Gates and Self-Healing in Production
The Unbreakable Delivery Pipeline with Shift-Left SRE Quality Gates and Self-Healing in Production

What do you “Auto-Remediate as Code” now?

I hope you got a couple of new ideas on how to Shift-Left SRE (Site Reliability Engineering). Working closely with engineers, providing them the frameworks to develop and extend the delivery pipeline to automatically validate and deploy your auto-remediation scripts allows you to build more and better resilient systems.

Leave me a message, drop me an email or message me on LinkedIn or Twitter in case you want to have a share your thoughts on this topic.

The post Shift-Left SRE: Building Self-Healing into your Cloud Delivery Pipeline appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/shift-left-sre-building-self-healing-into-your-cloud-delivery-pipeline/feed/ 0