Site Reliability Guardian | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Wed, 04 Mar 2026 13:46:41 +0000 en hourly 1 Self-service observability: Empower engineers to get the most value out of your observability data https://www.dynatrace.com/news/blog/self-service-observability-empower-engineers-to-get-the-most-value-out-of-your-observability-data/ https://www.dynatrace.com/news/blog/self-service-observability-empower-engineers-to-get-the-most-value-out-of-your-observability-data/#respond Wed, 01 Oct 2025 16:47:24 +0000 https://www.dynatrace.com/news/?p=71195 Observability data

What good is your observability data if the engineers who need it have to spend a lot of time looking for it? Why is it important to deliver data right into the engineering workflows? Events, logs, traces, and metrics are only valuable if the right people can use them effectively.

The post Self-service observability: Empower engineers to get the most value out of your observability data appeared first on Dynatrace news.

]]>
Observability data

The Dynatrace unified platform experience delivers actionable insights to engineers without the hassle. With customizable launchpads, seamless Backstage integration, the RedHat DevHub plugin, and automated deployment validation with Site Reliability Guardian, engineers get the most important information front and center, allowing them to prioritize work and make informed decisions.

In this blog post, you’ll learn how platform teams can bring observability to developers, right where they need it.

Focus on what matters with launchpads

Dynatrace launchpads let you create customizable home pages that offer an easily digestible view of your environment, tailored to the needs of engineering teams. Launchpads can address various use cases, from onboarding and learning, serving as an entry point for occasional users, or providing an opinionated view for daily operations across different users, teams, and departments.

You can pin quick access to relevant Dynatrace® Apps, link to external tools or specific entries in Dashboards, Notebooks, Workflows, Problems, and more.

For example, the launchpad below was built for a developer team. It covers their daily routines, relevant content, and documents. Using Launchpads this way ensures that even occasional users know where to find key apps, how to resolve common issues, and how to debug problems—using Dynatrace and external resources.

Read our Launchpads blog and explore different Launchpads on our Dynatrace Playground tenant!

This home page is the entry point for a team of developers who work with Dynatrace on a daily basis. It includes links to daily routines, relevant content, and documents.
Figure 1. This home page is the entry point for a team of developers who work with Dynatrace on a daily basis. It includes links to daily routines, relevant content, and documents.

Bring observability to developers with developer portal integrations

Developer portals like Backstage have become increasingly popular in recent years, based on their ability to centralize access to tools, services, and documentation. By offering a consistent interface, they help developers navigate complex ecosystems more effectively and reduce time spent on context switching.

The seamless integration between Dynatrace and Backstage allows developers to pull observability and security data from Dynatrace and display it in software components you manage through the Backstage Software Catalog. This plugin allows you to provide smart links to Dynatrace apps, context-rich overview tables, and Site Reliability Guardian results and logs, all directly in Backstage!

Read our documentation on Backstage integration, and to get started, check out the Backstage Dynatrace plugins monitoring & observability page on the Dynatrace Hub.

This is a team’s Backstage developer portal, directly accessed from Dynatrace through deep links. Data from Dynatrace is displayed directly in the portal.
Figure 2. This is a team’s Backstage developer portal, directly accessed from Dynatrace through deep links. Data from Dynatrace is displayed directly in the portal.

Teams using the RedHat Developer Hub can also benefit from the Dynatrace integration. This plugin seamlessly integrates observability and security data from Dynatrace into the portal, enabling teams to monitor and operate their software components more effectively.

You can display real-time insights alongside managed components and add links to Dynatrace Apps for deeper analysis and root cause investigation.

This page comes from the Red Hat Developer Hub, showing data from Dynatrace displayed directly in the portal.
Figure 3. This page comes from the Red Hat Developer Hub, showing data from Dynatrace displayed directly in the portal.

Validate deployments automatically with Site Reliability Guardian

Dynatrace’s Site Reliability Guardian (SRG) embeds observability directly into the engineering workflow, allowing for fast, automated validation of service health during every change.

Automated change impact analysis

SRG automatically analyses the impact of deployments on performance, availability, and capacity. Powered by Davis® AI, it detects regressions and anomalies early—before they hit production.

SLO validation

Engineers can define service-level objectives for their software components to ensure they meet expected quality standards. See how your software components rate against company benchmarks and get insights on where and how to improve to deliver the best quality. SRG validates these objectives regularly or on demand, ensuring services stay performant and resilient.

Unified overview

SRG validation results are stored in Dynatrace Grail® data lakehouse and visualized in a single-page summary, showing the pass/fail status of the last validations and a detailed evaluation of individual objectives. This allows you to easily spot trends and degressions over time. Visualization can be enriched with custom metadata, such as release versions, build IDs, and environment identifiers.

Automated release validation, CI/CD integrated

Trigger SRG validations directly from your CI/CD pipeline, to get instant feedback and observability insights into the current health and performance state of your component.

For more information, check out our documentation and then go to the Hub and set up your first Site Reliability Guardian! If you’re looking for inspiration, check out our Site Reliability Guardian config-as-code samples.

Conclusion

Observability as a self-service is a reality with Dynatrace. Our unified platform experience allows platform teams to connect observability with developer portals, workflows, and automated deployment validation, giving engineers direct access to the data they need—without bottlenecks.

By empowering engineers to curate and automate their data according to their needs, teams can work more effectively and bring reliability checks earlier into the development lifecycle.

Next steps

Ready to try out these self-service observability enhancements yourself? Explore Dynatrace integrations and start customizing your observability and developer experience!

The post Self-service observability: Empower engineers to get the most value out of your observability data appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/self-service-observability-empower-engineers-to-get-the-most-value-out-of-your-observability-data/feed/ 0
Enhance efficiency and compliance with automated AWS tag change triggers: A step-by-step guide https://www.dynatrace.com/news/blog/srg-aws-tag-changes/ https://www.dynatrace.com/news/blog/srg-aws-tag-changes/#respond Wed, 02 Apr 2025 15:52:01 +0000 https://www.dynatrace.com/news/?p=68498 Site Reliability Guardian

Streamlining site reliability at scale can be daunting, particularly with large-scale AWS environments and architecture that rely on hundreds—or even thousands—of Amazon EC2 instances. However, you can simplify the process by automating guardians in the Site Reliability Guardian (SRG) to trigger whenever there are AWS tag changes, helping teams improve compliance and effectively manage system […]

The post Enhance efficiency and compliance with automated AWS tag change triggers: A step-by-step guide appeared first on Dynatrace news.

]]>
Site Reliability Guardian

Streamlining site reliability at scale can be daunting, particularly with large-scale AWS environments and architecture that rely on hundreds—or even thousands—of Amazon EC2 instances. However, you can simplify the process by automating guardians in the Site Reliability Guardian (SRG) to trigger whenever there are AWS tag changes, helping teams improve compliance and effectively manage system performance.

This step-by-step guide will show you how to configure your architecture to trigger guardians whenever EC2 tags are updated. Note that EC2 is an example; this guide can be made to work generically for tag changes on any AWS resource. By the end of this guide, you’ll be ready to automate guardians at scale and optimize Amazon EC2 management with ease.

Why automate guardians for AWS tag changes?

Before diving into the technical setup, here’s why automating guardians whenever EC2 tags change is beneficial for your organization:

  • Greater efficiency: Automatically triggering guardians removes the need for manual intervention, saving time for DevOps or site reliability engineering (SRE) teams and allowing for more efficient resource management at scale.
  • Better compliance: Automating guardians ensures critical policies and checks are consistently applied after changes across your architecture, improving security and compliance efforts.
  • Cost optimization: Immediate responses to tag changes lead to informed decisions about scaling, shutting down unused instances, or fine-tuning resource efficiency.
  • Proactive site reliability: Automated guardians can monitor the four golden signals, enabling proactive reliability measures.

Now, let’s get started with the setup!

Step 1: Create an API token

Step 1: Create an API token

First, create an API token to integrate AWS services with Dynatrace for guardian automation.

  1. Log into your Dynatrace tenant

Log in to your Dynatrace tenant and note the first part of the URL (for instance, “abc12345”), which is your tenant ID.

  1. Access token settings

Press Ctrl + K or CMD + K and search for “Access Tokens” within Dynatrace.

  1. Generate a new token

Create a new access token and assign it “bizevents.ingest” permissions.

  1. Save the token

Copy and securely store the token, which looks like “dt0c01.*****.*****”. You’ll use this later during configuration.

Step 2: Create the EventBridge connection

Create the EventBridge connection

Configure invocation

Create the EventBridge connection

Amazon EventBridge acts as the bridge between AWS and Dynatrace. Here’s how to set it up:

  1. Navigate to Amazon EventBridge

Log in to your AWS Management Console and go to EventBridge > Connections.

  1. Recreate the cURL command

You can use this cURL command as a reference to establish your connection:

curl -X POST \
'https://abc12345.live.dynatrace.com/api/v2/bizevents/ingest' \
-H 'Authorization: Api-Token dt0c01.*****.*****' \
-H 'Content-Type: application/cloudevent+json' \
-d '{…}'
  1. Set the Authorization Method

Create a new EventBridge connection with the Authorization Method set to “API Key” and use the API token from Step 1 as the value (i.e., “Api-Token dt0c01.*****.****”).

Reminder: The API token is a sensitive value and should be stored in an encrypted format using a tool like AWS Secrets Manager. When following this guide, AWS Secrets Manager is already used.

Step 3: Define and configure EventBridge rules

Event pattern

EventBridge rules define the exact conditions for triggering guardians:

  1. Specify the input template:

Create an input template to modify your event data:

{
  "specversion": "1.0",
  "id": "<id>",
  "source": "aws.<source>",
  "type": "ec2.tag.change",
  "time": "<time>",
  "aws.region": "<region>",
  "aws.eventbridge.rule.arn": "<aws.events.rule-arn>",
  "aws.resources": <resources>,
  "data": <detail>
}
  1. Set the event pattern

Create a rule in EventBridge with the following event pattern:

{
  "source": ["aws.tag"],
  "detail-type": ["Tag Change on Resource"],
  "detail": {
    "service": ["ec2"],
    "resource-type": ["instance"]
  }
}
  1. Apply targets and permissions

Apply targets and permissions

Apply targets and permissions

Assign targets and permissions to ensure successful data ingestion into Dynatrace. Use an IAM role to permit EventBridge to call your API destination.

Input transformer

The input transformer should be set as follows:

{"detail":"$.detail","id":"$.id","region":"$.region","resources":"$.resources","source":"$.source","time":"$.time"}

Step 4: Test tag changes on Amazon EC2 instances

To validate your configuration, perform the following:

  1. Change a tag

Modify a tag by going to your Amazon EC2 instances in the AWS Management Console. For instance, update the “Environment” tag with a new value.

  1. Verify event logging

Check the EventBridge console to ensure your tag change triggered the appropriate event.

  1. Confirm data in Dynatrace

Within Dynatrace, press CMD/Ctrl + K and search for “Notebooks.” Create a new notebook and run the following query:

fetch bizevents | filter event.type == "ec2.tag.change"

If the query returns results, your configuration is working correctly.

Test tag changes on Amazon EC2 instances

Step 5: Set up the guardian

  1. Create a new guardian

Set up the guardian

In Dynatrace, search for “Site Reliability Guardian” (`CMD/Ctrl + K`) and create a new guardian. For best practices, use the “Four Golden Signals” template.

  1. Automate the workflow

Set up the guardian

Either on the overview page showing all guardians or on the analysis page of a selected guardian, click the Automate button. This will generate a workflow that triggers the guardian based on incoming bizevents (Business events). Configure the event type as `bizevent` and set the filter query to:

event.type == "ec2.tag.change"

Set up the guardian

  1. Add a pause

Set up the guardian

To allow your systems to stabilize before triggering the guardian, add a “wait before” step. For example, set a delay of 600 seconds (10 minutes).

Example timeline:

  • 06:59: Tag changed on EC2 instance.
  • 07:00: EventBridge triggers the workflow.
  • 07:10: Guardian is executed after 10-minute pause.
  1. Save the workflow

Save your final workflow to activate the automation.

Step 6: Validate and monitor the setup

Perform end-to-end validation by changing an EC2 tag again. Confirm the following:

  • The tag change event reaches Dynatrace.
  • The workflow triggers the guardian.
  • The guardian results appear in Dynatrace (e.g., heatmaps or relevant logs).

Run the following query in Dynatrace for additional monitoring:

fetch bizevents | filter event.type == "ec2.tag.change"

You should see log entries confirming the successful execution of your guardian process.

Achieve more with Site Reliability Guardian

In this blog, we highlighted the significant benefits of automating Site Reliability Guardian  triggers for Amazon EC2 changes. With automation, SRG helps engineering teams achieve efficiency, improved compliance, and cost optimization.

Learn more about Site Reliability Guardian in our documentation page.
Looking for more insights and support? Join the Automation Guild.

The post Enhance efficiency and compliance with automated AWS tag change triggers: A step-by-step guide appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/srg-aws-tag-changes/feed/ 0
Unified release management with Dynatrace: Empower DevOps and ITSM collaboration https://www.dynatrace.com/news/blog/unified-release-management-with-dynatrace/ https://www.dynatrace.com/news/blog/unified-release-management-with-dynatrace/#respond Tue, 14 Jan 2025 19:49:34 +0000 https://www.dynatrace.com/news/?p=67221 Distributed tracing best practices

Release management challenges with microservices Modern architecture often involves hundreds of microservices, each managed by its own CI/CD pipeline and often by different DevOps teams. While adding a release validation step to each pipeline is a best practice recommended by Dynatrace, implementing this across numerous pipelines can be resource-intensive. For organizations with hundreds of pipelines, […]

The post Unified release management with Dynatrace: Empower DevOps and ITSM collaboration appeared first on Dynatrace news.

]]>
Distributed tracing best practices

Release management challenges with microservices

Modern architecture often involves hundreds of microservices, each managed by its own CI/CD pipeline and often by different DevOps teams. While adding a release validation step to each pipeline is a best practice recommended by Dynatrace, implementing this across numerous pipelines can be resource-intensive. For organizations with hundreds of pipelines, achieving consistency requires careful coordination and planning to smoothly incorporate release validation.

Using Dynatrace’s centralized validation workflow, triggered via ITSM or GitOps, organizations can maintain consistent validation standards without individually modifying each pipeline. This flexible approach allows DevOps teams to leverage centralized control while enabling a gradual, manageable rollout of validation steps across pipelines as needed.

Unified approach with ITSM and CI/CD (including GitOps)

Dynatrace workflows provide a solution by integrating release validation with ITSM, GitOps, and CI/CD tools. Here’s how it works:

  1. Centralized validation via ITSM tools
    ITSM tools like ServiceNow and Jira provide centralized control over releases, incidents, and changes. With Dynatrace workflows, release managers can trigger validations without modifying each CI/CD pipeline, ensuring consistent quality standards across environments.
  1. GitOps for Kubernetes deployments
    GitOps tools, like ArgoCD, enable Kubernetes-based deployments by syncing cluster states with Git. This GitOps approach allows ArgoCD to trigger Dynatrace workflows upon deployment, running the same release validations managed by ITSM. Combining GitOps with ITSM ensures that all deployments—whether initiated by developers or release managers—adhere to organizational standards.
  1. Developer-focused CI/CD integration
    For teams using CI/CD platforms like Jenkins or GitLab, adding release validation steps provides immediate feedback, helping developers catch issues early. This developer-centric approach aligns with GitOps principles while still allowing ITSM-driven validation for broader oversight.

Unified approach with ITSM and CI/CD (including GitOps)

Defining Dynatrace composite SRGs and tag-based workflow implementation

In Dynatrace, Site Reliability Guardians (SRGs) provide automated validation checkpoints for each release. SRGs assess key metrics and thresholds, ensuring each deployment meets reliability standards before it progresses further. With SRGs, you can define specific service level objectives (SLO) like latency, error rates, or resource usage, tailoring validations to the needs of each service.

Composite SRGs for Targeted Validation and Notification

SRG supports tagging, allowing teams to group services by criteria such as environment, project, or criticality level. Using tags, you can configure Dynatrace workflows to execute SRGs selectively based on specific attributes, making release validation adaptable to different services.

For example:

Single-Service Validation: If a tag is set to “critical-service”, Dynatrace can trigger SRGs only for high-priority deployments, allowing focused validation where it’s needed most.

Multi-Service Validation: Broader tags like “production” or “release-group” can initiate SRG workflows across multiple related services, ensuring consistent standards for larger releases.

Implementing Workflows with Tag-Based Triggers

Workflows can leverage tags to determine which SRGs to execute, whether triggered from ITSM, GitOps, or CI/CD tools. This flexible setup means each deployment triggers only the necessary validations based on its tag profile, optimizing resources and reducing manual oversight.

Implementing Dynatrace Validation in a Phased Approach

Start by establishing a centralized workflow in ITSM to give Release Managers high-level control and visibility. As DevOps teams become familiar with the setup, expand to GitOps and CI/CD tools to support on-demand validation directly within development pipelines.

This phased approach balances centralized oversight for Release Managers with flexible, real-time validation for both DevOps engineers and developers, creating a comprehensive validation strategy across roles.

Conclusion

Integrating Dynatrace with ITSM and GitOps provides a unified, scalable release management strategy. By centralizing validation processes and automating checks across tools, organizations can improve release quality and responsiveness. Embrace this approach to empower both developers and release managers, achieving a seamless, high-confidence deployment workflow.

To help your team implement these solutions seamlessly, consider organizing structured workshops to cover essential steps like identifying critical services, setting up Composite Site Reliability Guardians with tags, and aligning workflows with existing CI/CD and GitOps practices. This approach ensures a smooth integration process and empowers your team to maintain high standards across all releases.

Ready to unify your release management with Dynatrace, ITSM, and GitOps? You and your team can experiment with all the functionality explored in this blog post in the Dynatrace Playground.

The post Unified release management with Dynatrace: Empower DevOps and ITSM collaboration appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/unified-release-management-with-dynatrace/feed/ 0
CrowdStrike update outage: Managing continuous delivery and deployment risk with Dynatrace https://www.dynatrace.com/news/blog/enhance-continuous-delivery-manage-deployment-risk/ https://www.dynatrace.com/news/blog/enhance-continuous-delivery-manage-deployment-risk/#respond Thu, 25 Jul 2024 21:43:26 +0000 https://www.dynatrace.com/news/?p=64984 Site Reliability Guardian, CrowdStrike outage

Resilient software update and delivery practices have become a key focus area in the aftermath of the CrowdStrike update outage. This blog is part of a series that explores how organizations can bolster IT practices at every stage of the software supply chain to ensure business resilience through any contingency.

The post CrowdStrike update outage: Managing continuous delivery and deployment risk with Dynatrace appeared first on Dynatrace news.

]]>
Site Reliability Guardian, CrowdStrike outage

After organizations struggled to recover from the CrowdStrike update outage on their Windows hosts in July, teams are now looking for ways to bolster the resiliency of their software update and continuous delivery practices.

The widespread impact of the CrowdStrike issue demonstrated how critical it is to avoid outages and failures during software deployments to maintain user trust and satisfaction. Traditional deployment techniques that roll out updates or patches directly into full production can present significant risks and lead to potential downtime.

Modern deployment techniques using progressive delivery—such as rolling, blue/green, canary, and feature flagging—offer a more controlled and gradual approach to releasing software. By implementing these strategies, organizations can minimize the impact of potential failures and ensure a smoother transition for users.

Eliminating the potential for outages may not be possible in all situations. That said, a unified observability and security platform such as Dynatrace can enhance modern deployment practices and enable teams to proactively monitor performance, validate changes, and best protect their downstream customers and end users from disruptions.

Bolster continuous delivery with software development lifecycle integrations

Software testing is critical, yet issues can still make it into production that negatively impact the customer experience. Dynatrace strives to protect its own customers through extensive automated test validation, complemented by staged rollouts to customer segments, which include quick feedback loops to determine whether the software is safe to continue deployments. Customers can adopt many modern and safe continuous delivery practices with Dynatrace in combination with their existing software pipelines.

By embedding Dynatrace observability into their own CI/CD pipeline, customers can prevent issues from ever leaving pre-production environments. This integration enhances the pre-release phase and plays a crucial role in the quick feedback cycle after deployment, allowing teams to identify issues immediately. Furthermore, the Dynatrace Site Reliability Guardian ensures service-level objectives (SLOs) and further validation thresholds are not violated both before and after deployment.

Progressive delivery methods to release higher-quality software, faster

The following common progressive delivery methods enable organizations to release software in a more controlled manner. 

Rolling deployments

In a rolling deployment, software is gradually rolled out, replacing previous software in a serial one-by-one fashion or in batch sets. Dynatrace can monitor production environments for performance degradations and outage events that may cause customers to lose access. The Dynatrace AutomationEngine, in conjunction with the existing deployment platforms, enables immediate and automatic software rollbacks to previous versions, limiting the blast radius for customers.

Blue/green deployments

A blue/green deployment strategy involves selecting a “blue” group to run the new software while the “green” group continues to run the previous version. The Davis AI engine immediately recognizes any anomalies, performance issues, or outages between the two groups. In case there is a degradation of service, Dynatrace picks it up and, embedded by the pipeline, can ensure all customers are rooted back to the proven, stable deployment.

Canary deployments

The Canary deployment strategy releases software to customers in incremental phases, gradually increasing the production load on the new deployment. Dynatrace enables customers to set quality measures or SLO targets for performance, outages, or other usage metrics to mitigate risk. By integrating AutomationEngine into the pipeline, customers can safely increase or decrease the load on canary deployments.

Feature flagging

Feature flagging is another popular technique used in progressive delivery that allows teams to roll out new features incrementally and with minimal risk. It supports A/B testing, canary releases, and quick rollbacks, ensuring smoother transitions and more controlled feature releases. The result is testing features in production with specific user segments, gathering feedback and making data-driven decisions on a broader rollout. Embedding feature flags in the codebase can make it easier to control the visibility and behavior of features in real time without redeploying or disrupting an application. This leads to enhanced agility, improves quality, and accelerates time-to-market, all while maintaining a seamless user experience.

As organizations implement progressive delivery strategies to enhance their release processes, integrating automated validation tools such as the Dynatrace Site Reliability Guardian (SRG) becomes essential for further optimizing these practices and ensuring robust performance monitoring

Site Reliability Guardian and AutomationEngine bolster continuous delivery tools and technologies

Dynatrace Site Reliability Guardian (SRG) integrates with continuous deployment practices and platforms to provide the following capabilities:

  1. Automated change impact analysis: SRG automates analyzing the impact of changes on service availability, performance, and capacity objectives across various systems before a deployment goes live or during any of the deployment strategies.
  2. Integration with CI/CD pipelines: Teams can integrate SRG into existing delivery pipelines including Jenkins, Github, GitLab, AWS, or Azure pipelines. This integration enables teams to validate releases automatically as part of their software development lifecycle before they are released to customers.
  3. Workflow automation: AutomationEngine and workflows automate the execution of guardians or problem remediation. This can be tied to specific events such as deployments or configuration changes, allowing for automated validation and response to changes.
  4. Service-level objectives (SLOs): SLOs enable site reliability engineers (SREs) to manage and track thresholds for critical services. The Davis AI engine proactively monitors these SLOs for degradations and acts before an SLO violation happens.
  5. Automated release validation: The platform supports automated release validation for security and quality gates to ensure that only high-quality code progresses through the delivery pipeline. This integration reduces the risk of deploying faulty code to production.

Embed Dynatrace into your release process and gain even more control

The Dynatrace platform is an essential tool for customers deploying software. It provides critical data that enables rapid, fact-based decisions about continuing or rolling back deployments. By offering real-time performance metrics and insights into application health, the Dynatrace platform empowers teams to assess the impact of changes quickly. This capability minimizes downtime and ensures potential issues are addressed before affecting end users, allowing organizations to navigate software deployment complexities and enhance reliability confidently.

Enhance continuous delivery quality and manage the risk of another CrowdStrike update outage

Incidents like the CrowdStrike update outage illustrate the importance of adopting modern progressive delivery strategies to enhance software reliability and customer satisfaction. By adopting approaches like rolling, blue/green, and canary deployments, teams can mitigate risks associated with large-scale releases and ensure a smoother user experience. Integrating Dynatrace into these processes provides invaluable insights and automated monitoring capabilities, allowing DevOps teams to detect issues early and respond swiftly.

As organizations continue to navigate the complexities of software delivery, prioritizing modern deployment techniques and leveraging robust monitoring solutions will be key to avoiding outages and failures. Embracing these practices will ultimately lead to more resilient applications, happier users, and a stronger competitive edge in the market.

Contact us to learn how you can bolster the resiliency of your progressive delivery techniques with AI-driven automation that can help you avoid the effects of an outage like CrowdStrike.

To learn more about the recent CrowdStrike update outage and explore more resources to help you maintain business resilience, check out the resource center, Business Resilience through CrowdStrike and Beyond.

The post CrowdStrike update outage: Managing continuous delivery and deployment risk with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enhance-continuous-delivery-manage-deployment-risk/feed/ 0
Auto-adaptive thresholds for AI-driven quality gating https://www.dynatrace.com/news/blog/auto-adaptive-thresholds-for-ai-driven-quality-gating/ https://www.dynatrace.com/news/blog/auto-adaptive-thresholds-for-ai-driven-quality-gating/#respond Tue, 04 Jun 2024 17:12:50 +0000 https://www.dynatrace.com/news/?p=64258 auto-adaptive thresholds

The Site Reliability Guardian automates the validation process for new software releases or changes. These validations involve setting specific performance, availability, or security objectives that must be met. For example, response time or failure rate thresholds can be used to define the desired state, warning range, or when an objective is violated. However, it can be difficult or even impossible to set these targets upfront for a new software component because it's unclear how the component will behave or what circumstances it will face.

The post Auto-adaptive thresholds for AI-driven quality gating appeared first on Dynatrace news.

]]>
auto-adaptive thresholds

The latest release of the Site Reliability Guardian incorporates assistance from Dynatrace Davis® AI to automatically derive appropriate threshold targets and adjust them over time to protect your quality improvements. This process, known as auto-adaptive thresholding, eliminates the need to define a static threshold upfront. Instead, it derives the suitable thresholds from previous validation results.

Build an umbrella for Development and Operations

In modern software engineering, the discipline of platform engineering delivers DevSecOps practices to developers to bridge the gaps between development, security, and operations and enhance the developer experience. A key element in platform engineering is the establishment of fast feedback cycles regarding the quality and security measures of new software releases. To provide automated feedback for developers, the concept of quality gates for static code analysis in continuous integration is widely adopted throughout the industry. However, this perspective differs in the continuous deployment practices of various organizations, where the feedback is either delayed or not returned to the developer.

While receiving no feedback on the quality or security of new features leaves developers uncertain about feature performance, delayed feedback also increases a developer’s cognitive load. The developer must pause their current engineering work to address the reported issue and consider the code changes they worked on a few days or weeks prior.

To reduce developers’ cognitive load by providing timely information, platform engineers must create tools that allow validations to be run in the early phases of development, with direct and fast feedback loops. Ideally, this should be a self-service offering that enables individual adoption by teams. While platform engineers can build and prepare the necessary infrastructure and templates for self-adoption, developers must still provide some customization. For example, the team must establish specific thresholds for desired service performance behavior. Setting these thresholds upfront can be challenging because the team might not know how a service will behave in its environment.

How we define auto-adaptive thresholds at Dynatrace

This blog post explores how Dynatrace leveraged the Site Reliability Guardian to establish a fast feedback loop for Davis AI model improvements. The conducted validations avoided regression within the models, and the outcomes were immediately fed back to the data science team when deviations were detected. While the data science team appreciated the quick insights and validations of their improvements, they initially struggled with the setup. Consequently, this blog post highlights the new capability of the Site Reliability Guardian to define auto-adaptive thresholds that tackle the challenge of configuring static thresholds and protect your quality and security investments with relative comparisons to previous validations.

Fast feedback cycles on model improvements

While the Site Reliability Guardian was originally designed to validate new software releases, Dynatrace has internally extended its application area to include validation of models for Davis AI.

The Dynatrace data science team continuously improves the machine learning models used by Davis AI, for example, by adding new features to forecasting or refining mathematical calculations. A single change can influence multiple models, as features are often used across several models. To ensure that model changes don’t lead to regressions, the data scientists set up Site Reliability Guardian, which is automatically triggered whenever a change is made in the codebase via CI/CD pipelines.

A series of models are continuously trained on Dynatrace tenants to effectively set objectives. The training times and other quality metrics, such as the RMSE (Root Mean Squared Error), SMAPE (Scaled Mean Absolute Percentage Error), and coverage probability, are monitored using Dynatrace. Our data scientists utilize metrics and events to store these quality metrics. However, other data formats, like logs, can also be employed. The quality metrics can then be easily queried using DQL and utilized for the objectives of a guardian. Validations are automatically triggered when a change is committed to the code base via the Dynatrace API. This helps data scientists quickly respond to recently introduced regressions. For instance, if an objective is violated, they’re immediately notified, for example, through a Slack channel.

Validation history
Figure 1. Validation history

One difficulty encountered when setting up objectives in the guardians was selecting an appropriate threshold for the quality metrics, as this is typically heavily dependent on the data.

Leverage Davis AI to quickly start validating

To address the challenge of defining a static threshold for an objective, the Site Reliability Guardian enables switching objective thresholds to auto-adaptive mode, as depicted in the screenshot below.

Activating an auto-adaptive threshold for the response time objective
Figure 2. Activating an auto-adaptive threshold for the response time objective

Davis AI controls auto-adaptive thresholds. It analyzes the next five validations to derive this objective’s proper warning and failure thresholds. Once the learning phase is complete, all subsequent validation results are fed into Davis AI to fine-tune the thresholds based on changed behavior.

Considering previous validation results, the latest validation is always relative to the past, protecting quality and security improvements. If, for example, recent performance improvements change a service’s response time from 200 ms to 175 ms, the auto-adaptive threshold is adjusted to reflect the new behavior. Nevertheless, the Site Reliability Guardian detects sudden behavior changes by reporting a warning or failure if response time returns to 200 ms or above.

Learning phase

Unless an objective has been validated five times, it’s still in the learning phase. During this phase, the measured values are informative, allowing observation of the objective’s development without affecting deployment or delivery processes. The Site Reliability Guardian denotes the learning phase of an objective with a loading symbol on the heatmap and in the objective details.

Representation of the learning phase of an auto-adaptive threshold
Figure 3. Representation of the learning phase of an auto-adaptive threshold

The warning and failure thresholds will be automatically set if sufficient validations are available to establish a solid baseline for an objective’s auto-adaptive thresholds. Consequently, the next objective validation will impact the overall validation result.

Auto-adaptive thresholds as code

To enhance the developer’s experience in adopting the Site Reliability Guardian in a self-service manner, the configuration for a guardian and its workflow can be provided in a configuration-as-code fashion. This enables the management of the configuration within the service’s code repository. Incorporating the new capability of auto-adaptive thresholds into configuration-as-code is as simple as adding the autoAdaptiveThresholdEnabled flag to an objective.

Configuration-as-code example for activating an auto-adaptive threshold
Figure 4. Configuration-as-code example for activating an auto-adaptive threshold

Before concluding, we wish to announce the change of the Site Reliability Guardian icon from purple to shiny golden. This new appearance enhances the icon’s geometry while preserving its core values. Therefore, the icon continues to feature the infinity loop as a symbol for the DevSecOps loop and the fast forward sign to expedite delivery while ensuring quality and security standards.

The new and shiny appearance of Site Reliability Guardian

Evolution of the Site Reliability Guardian icon
Figure 5. Evolution of the Site Reliability Guardian icon

What’s next

The new auto-adaptive thresholds capability is now available in Site Reliability Guardian. Open the app and switch from static to auto-adaptive thresholds for those objectives where Davis AI should derive the thresholds for you. For full details, see Dynatrace Documentation.

If you haven’t used Site Reliability Guardian yet, try it out in the Dynatrace Playground or watch the latest app spotlight recording.

The post Auto-adaptive thresholds for AI-driven quality gating appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/auto-adaptive-thresholds-for-ai-driven-quality-gating/feed/ 0
What are quality gates? How to use quality gates to deliver better software at speed and scale https://www.dynatrace.com/news/blog/what-are-quality-gates-how-to-use-quality-gates-with-dynatrace/ https://www.dynatrace.com/news/blog/what-are-quality-gates-how-to-use-quality-gates-with-dynatrace/#respond Wed, 21 Feb 2024 16:43:18 +0000 https://www.dynatrace.com/news/?p=62467 Site Reliability Guardian, CrowdStrike outage

In today’s fast-paced IT environments, speed matters. Organizations must ensure their digital infrastructure functions optimally and that they deliver software deployments and updates rapidly and consistently. But to meet demands and stay ahead of competitors, teams are under pressure to accelerate releases as much as possible. Verifying releases using systematic quality gates is a crucial […]

The post What are quality gates? How to use quality gates to deliver better software at speed and scale appeared first on Dynatrace news.

]]>
Site Reliability Guardian, CrowdStrike outage

In today’s fast-paced IT environments, speed matters. Organizations must ensure their digital infrastructure functions optimally and that they deliver software deployments and updates rapidly and consistently. But to meet demands and stay ahead of competitors, teams are under pressure to accelerate releases as much as possible. Verifying releases using systematic quality gates is a crucial practice to speed up the pace of deployments without sacrificing quality.

What are quality gates?

Quality gates are checkpoints that require deliverables to meet specific, measurable success criteria before progressing to the next development stage. They help foster confidence and consistency throughout the entire software development lifecycle (SDLC).

Organizations can customize quality gate criteria to validate technical service-level objectives (SLOs) and business goals, ensuring early detection and resolution of code deficiencies. Automating quality gates is ideal, as it minimizes manually checking and validating key metrics throughout the SDLC. Ultimately, quality gates safeguard code viability as it advances through the delivery pipeline.

Benefits of quality gates

Quality gates provide several advantages to organizations, including the following:

  • Optimized software performance: Quality gates assess code at different SDLC stages and ensure that only high-quality code progresses. This mechanism significantly boosts the likelihood of optimal functioning upon deployment. By actively monitoring metrics such as error rate, success rate, and CPU load, quality gates instill confidence in teams during software releases.
  • Fewer expensive fixes. Introducing well-defined quality gates at regular checkpoints prevents the escalation of issues as software advances through the pipeline. Early identification and resolution of potential deficiencies conserves resources and mitigates the need for expensive fixes later in the SDLC.
  • Continuous, informed improvement: Quality gates provide consistent feedback on key metrics. This steady input empowers teams to proactively address issues, which fosters continuous learning from the outset of the SDLC. This approach supports innovation, ambitious SLOs, DevOps scalability, and competitiveness.

Quality gates examples in Dynatrace

Quality gates hold much promise for organizations looking to release better software faster. But how do they function in practice? The following are specific examples that demonstrate quality gates in action:

Security gates

Security gates ensure code meets key security requirements defined by development and security stakeholders. In this example, we will focus on ensuring releases do not have any known vulnerabilities. In this context, before the gorouter service deploys, security gates defined in Dynatrace’s Site Reliability Guardian (SRG) validate whether the latest version of the software has introduced a vulnerability.

Security gate for gorouter

In this scenario, the gorouter service must meet specific key metrics relating to the following three security areas:

  • New vulnerabilities: The newly detected high-profile vulnerabilities in the environment. For each of these areas, the software receives a score from one to 10. Depending on the score thresholds for each SLO, the software will either “fail” or “pass” the gate.
  • Open vulnerability on process group: The total number of currently high-profile vulnerabilities related to a process group.
  • Open vulnerabilities: The total number of currently open high-profile vulnerabilities in the environment.
  • Vulnerability score: The highest vulnerability risk score for a process group.

To view a more in-depth breakdown of these key metrics and their thresholds, simply scroll down to the bottom:

Validated objectives

Validated objectives

For every new build, the SRG assigns a score for each of these SLO areas. Considering the above scores, this version of software would not pass through the security gates, and the deployment process for the gorouter service would be stalled. If even one of the key metrics receives a failing score —as it did on the open vulnerability on process group — the software cannot progress further. Adjustments must be made accordingly.

Quality gates after load/performance testing

Teams can use quality gates to evaluate performance metrics. Consider “Easytravel,” a customer-facing application used by a travel agency. Before a new version of the application is deployed, the software is subject to a series of load tests that evaluate capacity and performance under a series of simulated traffic and application demands.

Quality gates consider these metrics and determine if their values fulfill specific SLO requirements. The gates thus allow the application version to pass through if the software meets the criteria and halts the progression if not. Several tools can be used to collect metrics in load/performance testing. In this example, test results are gathered from K6, Gremlin, NeoLoad, and Dynatrace synthetics.

Easytravel

Easytravel

Per the figures above, SRG evaluates Easytravel metrics from all of these tools in one place. This way, the travel agency can easily streamline, organize, and consolidate their quality gates and metric evaluation process. The agency can also efficiently compare the newest version of Easytravel against previous versions of the software with regression testing facilitated by SRG.

The new version of the Easytravel app has passed the load/performance testing quality gates. Although there were a handful of “warnings” for the K6 metrics — this means that the metrics are close to the “failure” threshold — none of the metrics failed to meet the SLO entirely. As a result, this version of Easytravel passes through this set of quality gates and is one step closer to deployment.

Quality gates to validate the “four golden signals”

The “four golden signals” represent the most crucial metrics of a customer-facing system’s performance. These metrics are latency, traffic, errors, and saturation, all of which must be key considerations when curating user experience. Teams can use quality gates to track and evaluate such metrics. Let’s review how a quality gate can function for each of the four golden signals using the same Easytravel application. Below is a sample SRG dashboard for these signals:

Validation history

  • Latency
    • Latency refers to the amount of time that data takes to transfer from one point to another within a system. In the context of Easytravel, one can measure the speed at which a specific page of the application responds after a user clicks on it.
      • The passing threshold is anything below 50 ms.
        The warning threshold is 50-60 ms.
        The failure threshold is anything above 60ms.
    • In this scenario, the response time for the latest software version is 44 ms, thus, it passes through the latency quality gate.

Latency

Teams take a similar approach for the other three signals. In this example, unlike latency, the remaining three signals did not receive a “pass.” SLOs for saturation were not met, and thus, this specific build of Easytravel temporarily halts. In terms of errors and traffic, these SLOs did not receive a failing score but are close enough to the failure threshold to trigger a warning.

What’s next

Quality gates are only as effective as the tools that enable them. Dynatrace’s Site Reliability Guardian provides a platform for in-depth analysis and validation of service availability, performance, and capacity objectives throughout your entire digital environment. Equipped with these capabilities, Site Reliability Guardian can help you implement high-performing quality gates.

Interested in delivering better software faster? Start a 15-day free trial today.

The post What are quality gates? How to use quality gates to deliver better software at speed and scale appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/what-are-quality-gates-how-to-use-quality-gates-with-dynatrace/feed/ 0
Ensure safe and secure releases at scale by providing Golden Paths https://www.dynatrace.com/news/blog/ensure-safe-and-secure-releases-at-scale-by-providing-golden-paths/ https://www.dynatrace.com/news/blog/ensure-safe-and-secure-releases-at-scale-by-providing-golden-paths/#respond Tue, 14 Nov 2023 16:17:41 +0000 https://www.dynatrace.com/news/?p=60743 Unlocking the Power of Quality Gating by enabling Self-Service

Golden Paths for rapid product development Modern software development aims to streamline development and delivery processes to ensure fast releases to the market without violating quality and security standards. DevOps practices have been established in the last decade to accomplish this goal and deal with the dynamics of modern, cloud-native software architectures. To bring these […]

The post Ensure safe and secure releases at scale by providing Golden Paths appeared first on Dynatrace news.

]]>
Unlocking the Power of Quality Gating by enabling Self-Service

Golden Paths for rapid product development

Modern software development aims to streamline development and delivery processes to ensure fast releases to the market without violating quality and security standards. DevOps practices have been established in the last decade to accomplish this goal and deal with the dynamics of modern, cloud-native software architectures. To bring these practices to life within an organization at scale, the discipline of platform engineering has gained popularity. From a high-level point of view, platform engineering aims to:

  • Reduce the cognitive load on development teams.
  • Improve reliability and resiliency of products that rely on platform capabilities.
  • Accelerate product development and delivery by reusing and sharing platform tools and knowledge.
  • Reduce risk of security, regulatory, and functional issues in products and services.
  • Enable cost-effective and productive use of services.

While it takes multiple capabilities to achieve these goals, implementing “Golden Path” templates is a fundamental ingredient. The Cloud Native Computing Foundation defines a Golden Path as a “templated composition of well-integrated code and capabilities for rapid project development.” Simply put, a Golden Path is a self-service template for common tasks that allows for autonomy while providing guardrails that safeguard production environments.

Imagine that instead of development teams fending for themselves amidst a sea of tools and infrastructure, well-defined and enterprise-wide templates are provided for the development of all new product services. Such a template should contain a get-started tutorial, sample source-code framework, policy guardrails, CI/CD pipeline, infrastructure-as-code templates, and reference documentation. This approach helps you quickly integrate best practices within your organization and provides cloneable artifacts for rapid product development.

No developer is left in the dark to fend for themselves, Golden Paths light the way.

Ensure governance across your organization

While Golden Paths are key to bringing DevOps best practices to development teams, distributing them is challenging. This challenge is partially addressed by internal development platforms (IDP), which have been adopted by many, but not all, organizations.

Consequently, Dynatrace provides templates in the context they are needed—with or without the internal use of an IDP. This ensures governance across your organization with the proper templates in the right place. At the same time, this does not restrict you from managing them in your IDP, as explained below.

Leverage opinionated, self-service, and optional templates

The latest enhancements of the Site Reliability Guardian have introduced opinionated Golden Path templates. For now, they concentrate on Kubernetes, host, and security objectives but they will grow to cover even more reliability, performance, and security aspects. These templates enable development teams to instantiate a guardian for automating release validation in a self-service manner. Along the journey, monitored entities can be selected to provide the context to fetch the right data from Dynatrace Grail™.

Two-step approach to creating a guardian using a template.
A two-step approach to creating a guardian using a template.

After completing this two-step process, a ready-to-use guardian is created. It’s possible to add new objectives if needed or to tailor the objectives and their thresholds.

Fully functional and pre-configured guardian for a Kubernetes workload.
Fully functional and pre-configured guardian for a Kubernetes workload.

Allow for flexibility

Custom query variables are available to fine-tune guardian objectives and maintain flexibility in fetching data from Grail. An example of such a variable is the version number of a service. In many cases, you want to retrieve the logs, metrics, or traces of a particular version, which is unknown upfront. Consequently, the custom variable version can be defined in the DQL query, which retrieves its value before executing.

Custom query variables can be defined by starting with the $ sign followed by the variable name in the DQL query editor of a guardian objective. An overlay component then helps you add this variable and define the default value in the context of the guardian.

Leverage variables to parameterize the query of an objective.
Leverage variables to parameterize the query of an objective.

While a variable will fall back to its default if no value is provided, there are three ways of ingesting the value:

  1. When triggering a guardian validation from within the UI by selecting the Validate button, set a value for the configured variables using the Set variables option.
  2. The Site Reliability Guardian workflow action allows you to set a variable value. In this case, it is possible to consume data from the triggering event or a previous workflow action using Jinja expressions.
  3. If a workflow is used to automate the validation process, the workflow trigger event can provide automatically mapped properties to variables. Therefore, the properties must be within the execution_context property, as shown below. Given this example, the version number 0.1.0 is set where the variable $version is defined.

Code snippet to set variables

Since variables provide important context for validation, you can view validation-result values over time. The following screenshots depict the version number associated with the last 10 validation results.

Site Reliability Guardian Carts validation history in Dynatrace screenshot

Integrate with existing internal developer platforms

If you have an internal developer platform that serves as a central product engineering hub, you need Golden Path templates incorporated there so that you can scale out the platform for use by different teams. With Dynatrace, this can be achieved by extracting an instantiated template using Dynatrace Configuration as Code or cloning it from our public repository. This declarative format of a template, consisting of a guardian and workflow configuration, as depicted by the sample below, can then be added to the template repository managed and distributed by an IDP.

Dynatrace Configuration as Code sample

What’s next?

Site Reliability Guardian is built to ensure and maintain reliability and resiliency of products by leveraging automated release validations. Therefore, Golden Path templates in a self-service manner are offered for platform engineers to enable development teams to get started easily. The next enhancements of the Site Reliability Guardian will bring Davis® AI even closer to the Site Reliability Guardian by recommending relevant objectives and baselines for comparison.

  • The new functionality has already been released and is available for your use. If you’re using the Site Reliability Guardian already, update to the latest version. Otherwise, navigate to Dynatrace Hub and install it from there.
  • If the Site Reliability Guardian is new to you and you’re curious about it, go to Dynatrace Discovery and explore its capabilities in a playground environment.

We’d love to hear your feedback about Site Reliability GuardianLet us know your thoughts or share your ideas about further improvements.

The post Ensure safe and secure releases at scale by providing Golden Paths appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ensure-safe-and-secure-releases-at-scale-by-providing-golden-paths/feed/ 0
Implementing AWS Well-Architected pillars with automated workflows https://www.dynatrace.com/news/blog/implementing-aws-well-architected-pillars/ https://www.dynatrace.com/news/blog/implementing-aws-well-architected-pillars/#respond Wed, 13 Sep 2023 16:52:01 +0000 https://www.dynatrace.com/news/?p=59517 Dynatrace | AWS

If you use AWS cloud services to build and run your applications, you may be familiar with the AWS Well-Architected Framework. This is a set of best practices and guidelines that help you design and operate reliable, secure, efficient, cost-effective, and sustainable systems in the cloud. The framework comprises six pillars: Operational Excellence, Security, Reliability, […]

The post Implementing AWS Well-Architected pillars with automated workflows appeared first on Dynatrace news.

]]>
Dynatrace | AWS

If you use AWS cloud services to build and run your applications, you may be familiar with the AWS Well-Architected Framework. This is a set of best practices and guidelines that help you design and operate reliable, secure, efficient, cost-effective, and sustainable systems in the cloud. The framework comprises six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, and Sustainability.

Six AWS well architected pillars

But how can you ensure that your applications meet these pillars and deliver the best outcomes for your business? And how can you verify this performance consistently across a multicloud environment that also uses Microsoft Azure and Google Cloud Platform frameworks? Because Google offers its own Google Cloud Architecture Framework and Microsoft its Azure Well-Architected Framework, organizations that use a combination of these platforms triple the challenge of integrating their performance frameworks into a cohesive strategy.

This is where unified observability and Dynatrace Automations can help by leveraging causal AI and analytics to drive intelligent automation across your multicloud ecosystem. The Dynatrace platform approach to managing your cloud initiatives provides insights and answers to not just see what could go wrong but what could go right. For example, optimizing resource utilization for greater scale and lower cost and driving insights to increase adoption of cloud-native serverless services.

In this blog post, we’ll demonstrate how Dynatrace automation and the Dynatrace Site Reliability Guardian app can help you implement your applications according to all six AWS Well-Architected pillars by integrating them into your software development lifecycle (SDLC).

Dynatrace AutomationEngine workflows automate release validation using AWS Well-Architected pillars

With Dynatrace, you can create workflows that automate various tasks based on events, schedules or Davis problem triggers. Workflows are powered by a core platform technology of Dynatrace called the AutomationEngine. Using an interactive no/low code editor, you can create workflows or configure them as code. These workflows also utilize Davis®, the Dynatrace causal AI engine, and all your observability and security data across all platforms, in context, at scale, and in real-time.

One of the powerful workflows to leverage is continuous release validation. This process enables you to continuously evaluate software against predefined quality criteria and service level objectives (SLOs) in pre-production environments. You can also automate progressive delivery techniques such as canary releases, blue/green deployments, feature flags, and trigger rollbacks when necessary.

This workflow uses the Dynatrace Site Reliability Guardian application. The Site Reliability Guardian helps automate release validation based on SLOs and important signals that define the expected behavior of your applications in terms of availability, performance errors, throughput, latency, etc. The Site Reliability Guardian also helps keep your production environment safe and secure through automated change impact analysis.

But this workflow can also help you implement your applications according to each of the AWS Well-Architected pillars. Here’s an overview of how the Site Reliability Guardian can help you implement the six pillars of AWS Well-Architected.

AWS well-architected six pillars workflow
A Dynatrace Workflow that uses Dynatrace Site Reliability Guardian to implement the six AWS well-architected pillars

AWS Well-Architected pillar #1: Performance efficiency

The performance efficiency pillar focuses on using computing resources efficiently to meet system requirements, maintaining efficiency as demand changes, and evolving technologies.

A study by Amazon found that increasing page load time by just 100 milliseconds costs 1% in sales. Storing frequently accessed data in faster storage, usually in-memory caching, improves data retrieval speed and overall system performance. Beyond efficiency, validating performance thresholds is also crucial for revenues.

Once configured, the continuous release validation workflow powered by the Site Reliability Guardian can automatically do the following:

  • Validate if service response time, process CPU/memory usage, and so on, are satisfying SLOs
  • Stop promoting the release into production if the error rate in the logs is too high
  • Notify the SRE team using communication channels
  • Create a Jira ticket or an issue on your preferred Git repository if the release is violating the set thresholds for the performance SLOs
Pillar #1 performance efficiency of the AWS well architected pillars
The continuous release validation workflow powered by Dynatrace Site Reliability Guardian automatically verifies performance efficiency validation success and threshold violation cases

SLO examples for performance efficiency

The following examples show how to define an SLO for performance efficiency in the Site Reliability Guardian using Dynatrace Query Language (DQL).

Validate if response time is increasing under high load utilizing OpenTelemetry spans

fetch spans 
| filter endpoint.name == "/api/getProducts" 
| filter k8s.namespace.name == "catalog" 
| filter k8s.container.name == "product-service" 
| filter http.status_code == 200 
| summarize avg(duration) // in milliseconds
fetch spans result for AWS well architected pillar #1
* Please note that the Traces on Grail feature is currently in private preview, and the DQL syntax is subject to change.

Check if process CPU usage is in a valid range

timeseries val = avg(dt.process.cpu.usage) 
,filter in(dt.entity.process_group_instance, "PROCESS_GROUP_INSTANCE-ID") 
| fields avg = arrayAvg(val) // in percentage

CPU result

AWS Well-Architected pillar #2: Security

The security pillar focuses on protecting information system assets while delivering business value through risk assessment and mitigation strategies.

The continuous release validation workflow powered by Site Reliability Guardian can automatically do the following:

  • Check for vulnerabilities across all layers of your application stack in real-time, getting help from Dynatrace Davis Security Score as a validation metric
  • Block releases if they do not meet the security criteria
  • Notify the security team of the vulnerabilities in your application and create an issue/ticket to track the progress
Davis Security Score for AWS well architected pillar #2, security
Davis Security Score against third-party vulnerabilities

SLO examples for security

The following examples show how to define an SLO for security in the Site Reliability Guardian using DQL.

Runtime Vulnerability Analysis for a Process Group Instance – Davis Security Assessment Score

fetch events 
| filter event.kind == "SECURITY_EVENT" 
| filter event.type == "VULNERABILITY_STATE_REPORT_EVENT" 
| filter event.level == "ENTITY" 
| filter in("PROCESSGROUP_INSTANCE_ID",affected_entity.affected_processes.ids) 
| sort timestamp, direction:"descending" 
| summarize  
{  
status=takeFirst(vulnerability.resolution.status), 
score=takeFirst(vulnerability.davis_assessment.score), 
affected_processes=takeFirst(affected_entity.affected_processes.ids) 
}, 
by: {vulnerability.id, affected_entity.id} 
| filter status == "OPEN"  
| summarize maxScore=takeMax(score)

security score result

AWS Well-Architected pillar #3: Cost optimization

The cost optimization pillar focuses on avoiding unnecessary costs and understanding managing tradeoffs between cost capacity performance.

The continuous release validation workflow powered by the Site Reliability Guardian can automatically do the following:

  • Detect underutilized and/or overprovisioned resources in Kubernetes deployments considering the container limits and requests
  • Determine the non-Kubernetes-based applications that underutilize CPU, memory, and disk
  • Simultaneously validate if performance objectives are still in the acceptable range when you reduce the CPU, memory, and disk allocations
A graphic that shows the cost optimization without affecting the application performance for AWS well-architected pillar #3, cost performance
A graph that shows the cost optimization without affecting the application performance

SLO examples for cost optimization

The following examples show how to define an SLO for cost optimization in the Site Reliability Guardian using DQL.

Reduce CPU size and cost by checking CPU usage

To reduce CPU size and cost, check if CPU usage is below the SLO threshold. If so, test against the response time objective under the same Site Reliability Guardian. If both objectives pass, you have achieved your cost reduction on CPU size.

Screenshot of cost performance objective SLO in support of the AWS well-architected pillars

Here are the DQL queries from the image you can copy:

timeseries cpu = avg(dt.containers.cpu.usage_percent), 
filter: in(dt.containers.name, "CONTAINER-NAME") 
| fields avg = arrayAvg(cpu) // in percentage
fetch logs 
| filter k8s.container.name == "CONTAINER-NAME" 
| filter k8s.namespace.name == "CONTAINER-NAMESPACE" 
| filter matchesPhrase(content, "/api/uri/path") 
| parse content, "DATA '/api/uri/path' DATA 'rt:' SPACE? FLOAT:responsetime "  
| filter isNotNull(responsetime) 
| summarize avg(responsetime) // in milliseconds

Reduce disk size and re-validate

If the SLO specified below is not met, you can try reducing the size of the disk and then validating the same objective under the performance efficiency validation pillar. If the objective under the performance efficiency pillar is achieved, it indicates successful cost reduction for the disk size.

timeseries disk_used = avg(dt.host.disk.used.percent), 
filter: in(dt.entity.host,"HOST_ID") 
| fields avg = arrayAvg(disk_used) // in percentage

Disk used results

AWS Well-Architected pillar #4: Reliability

The reliability pillar focuses on ensuring a system can recover from infrastructure or service disruptions and dynamically acquire computing resources to meet demand and mitigate disruptions such as misconfigurations or transient network issues.

The continuous release validation workflow powered by the Site Reliability Guardian can automatically do the following:

  • Monitor the health of your applications across hybrid multicloud environments using Synthetic Monitoring and evaluate the results depending on your SLOs
  • Proactively identify potential availability failures before they impact users on production
  • Simulate failures in your AWS workloads using Fault Injection Simulator (FIS) and test how your applications handle scenarios such as instance termination, CPU stress, or network latency. SRG validates the status of the resiliency SLOs for the experiment period.
World map showing reliability metrics in support of AWS well architected pillar #4, reliability
Application availability validation across the world using Dynatrace Synthetic monitoring

SLO examples for reliability

The following examples show how to define an SLO for reliability in the Site Reliability Guardian using DQL.

Success Rate – Availability Validation with Synthetic Monitoring

fetch logs 
| filter log.source == "logs/requests" 
| parse content,"JSON:request" 
| fieldsAdd httpRequest = request[httpRequest] 
| fieldsAdd httpStatus = httpRequest[status] 
| fieldsAdd success = toLong(httpStatus < 400) 
| summarize successRate = sum(success)/count() * 100 // in percentage

logs request result

Number of Out of memory (OOM) kills of a container in the pod to be less than 5

timeseries oom_kills = avg(dt.kubernetes.container.oom_kills), 
filter: in(k8s.cluster.name,"CLUSTER-NAME") and in(k8s.namespace.name,"NAMESPACE-NAME") and in(k8s.workload.kind,"statefulset") and in (k8s.workload.name,"cassandra-workload-1") 
| fields sum = arraySum(oom_kills) // num of oom_kills

Reliability OOM kills result

AWS Well-Architected pillar #5: Operational excellence

The operational excellence pillar focuses on running and monitoring systems to deliver business value and continually improve supporting processes and procedures.

With continuous release validation workflow powered by the Site Reliability Guardian, you can:

  • Automatically verify service or application changes against key business metrics such as customer satisfaction score, user experience score, and Apdex rating
    Apdex rating in support of AWS well-architected pillar #5, operational excellence
  • Enhance collaboration with targeted notifications of relevant teams using the Ownership feature
  • Create an issue on your preferred Git repository to track and resolve the invalidated SLOs
  • Trigger remediation workflows based on events such as service degradations, performance bottlenecks, security vulnerabilities
  • Validate your CI/CD performance over time, considering the execution times, pipeline performance, failure rate, etc., which shows your operational efficiency in your software delivery pipeline.Screenshot of pipeline metrics in support of AWS well-architected pillar #5, operational excellence

SLO examples for operational excellence

The following examples show how to define an SLO for operational excellence in the Site Reliability Guardian.

Apdex rating validation of a web application

  1. Navigate to “Service-level objectives” and click on “Add new SLO” button
  2. Select “User experience” as a template. It will auto-generate the metric expression as the following:
    (100)*(builtin:apps.web.actionCount.category:filter(eq("Apdex category",SATISFIED)):splitBy())/(builtin:apps.web.actionCount.category:splitBy())
  3. Replace your application name in the entityName attribute:
    type("APPLICATION"),entityName("APPLICATION-NAME")
  4. Add a success criteria depending on your needs
    Success criteria for operational excellence
  5. Reference this SLO in your Site Reliability Guardian objective

AWS Well-Architected pillar #6: Sustainability

The sustainability pillar focuses on minimizing environmental impact and maximizing the social benefits of cloud computing.

The continuous release validation workflow powered by the Site Reliability Guardian can automatically do the following:

  • Measure and evaluate carbon footprint emissions associated with cloud usage
  • Leverage observability metrics to identify underutilized resources for reducing energy consumption and waste emissions
Sustainability dashboard supporting AWS well-architected pillar #6, Sustantainability
The Dynatrace Carbon Impact Dashboard evaluates the carbon impact of the resources in the cloud

SLO examples for sustainability

The following examples show how to define an SLO for sustainability in the Site Reliability Guardian using DQL.

Carbon emission total of the host running the application for the last 2 hours

fetch bizevents, from: -2h 
| filter event.type == "carbon.report" 
| filter dt.entity.host == "HOST-ID" 
| summarize toDouble(sum(emissions)), alias:total // total CO2e in grams

carbon emissions results supporting the AWS well-architected pillar #6, sustainability

Under-utilized memory resource validation

timeseries memory=avg(dt.containers.memory.usage_percent), by:dt.entity.host 
| filter dt.entity.host == "HOST-ID" 
| fields avg = arrayAvg(memory) // in percentage

memory resource validation for AWS well-architected pillar #6, sustainability

Implementing AWS Well-Architected Framework with Dynatrace: A practical guide

To help you get started with implementing the AWS Well-Architected Framework using Dynatrace, we’ve provided a sample workflow and SRGs in our official repository. This resource offers a step-by-step guide to quickly set up your validation tools and integrate them into your software development lifecycle.

Quick Implementation: Follow the instructions in our Dynatrace Configuration as Code Samples repository to deploy the sample workflow and SRGs. This will enable you to immediately begin validating your applications against the AWS Well-Architected pillars.

For more about how Site Reliability Guardian helps organizations automate change impact analysis, performance, and service level objectives, join us for the on-demand Observability Clinic, Site Reliability Guardian with DevSecOps activist Andreas Grabner.

The post Implementing AWS Well-Architected pillars with automated workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/implementing-aws-well-architected-pillars/feed/ 0
Automated observability, security, and reliability at scale https://www.dynatrace.com/news/blog/automated-observability-security-and-reliability-at-scale/ https://www.dynatrace.com/news/blog/automated-observability-security-and-reliability-at-scale/#respond Tue, 18 Jul 2023 15:54:24 +0000 https://www.dynatrace.com/news/?p=58664 Configuration at scale

Dynatrace automates setup and ensures observability, security, alerting, and remediations for newly developed software at any point in the software development lifecycle and in any number of environments.

The post Automated observability, security, and reliability at scale appeared first on Dynatrace news.

]]>
Configuration at scale

Dynatrace Configuration as Code enables complete automation of the Dynatrace platform’s configuration, ensuring that software is secure and reliable. With Configuration as Code, developers can manage their observability and security tasks with config files that can be developed alongside source code conveniently and at scale. Dynatrace offers Configuration as Code for the entire platform, covering all aspects, including app settings built for the AppEngine.

As software development grows more complex, managing components using an automated onboarding process becomes increasingly important. This is especially crucial in microservice architectures, where the number of components can be overwhelming. Furthermore, increasing the frequency of releases requires additional product lifecycle automation and setup.

While infrastructure has historically been treated as a bottleneck where proper scaling and compute power are applied to improve performance, these aspects are now typically addressed by hyperscalers that offer cloud-based infrastructure and infrastructure as a service. However, scaling up software development requires more tools along the software product lifecycle, which must be configured promptly and efficiently.

To handle this challenge, enterprises need to automate and streamline the onboarding and lifecycle of tool configurations in the software development processes, including aspects of observability, security, alerting, and remediation.

Efficient environment configuration at scale

One of software engineers’ most significant challenges is managing the numerous tools and technologies required for the software product lifecycle. Development teams must set up tailored configurations for each tool and component they’re responsible for. Developers can’t focus on software development while also managing the details of each tool’s configuration options, and tool administrators can’t scale to keep up with multiple development teams, numerous configurations, and the many environments they need to support. This is why it’s so essential to offer sharable configuration templates that can be easily reused and customized for specific teams, components, and environments.

Configuration as Code in Git repos, automatically applied by Dynatrace

Analogous to infrastructure as code, Configuration as Code, or “everything as code” is now essential for tackling software development challenges. With Dynatrace Configuration as Code, teams can meet these challenges while working with the IDEs, Git repos, and tools that they’re already familiar with. Configuration files allow for the automatic creation, update, and management of configurations for dashboards, synthetic monitors, alerts, SLOs, and security settings across multiple environments.
Configuration as Code in Git repos, automatically applied by Dynatrace

While developers edit files, a simple CLI command applies configurations to Dynatrace and, for example, automates the setup of a quality gate, including workflows and Site Reliability Guardians.

This can all be done safely and consistently in a repeatable manner. Configuration files can be reused, versioned, and shared across teams. Configuration as Code supports all the mechanisms and best practices of Git-based workflows, including pull requests, commit merging, and reviewer approval.

GitOps is a best-practice methodology for handling operation-relevant configurations that can be applied across the entire Dynatrace platform. With Dynatrace configuration files in Git repos, you gain:

  • Persistent configuration state available in files
    • Easy copy and paste
    • Easy templating
  • Established Git software development workflows
    • Repo branching
    • Change management
    • Reviews and approvals

From the developer perspective, you only need to fill out a form (YAML file) with key-value pairs (for example, to provide the thresholds for the API endpoint of a service) and an email address for escalation to get the benefit of the Dynatrace platform when you deliver a new service.

Enable self-service configuration at scale

You can easily set up observability, security, and automation by filling out a few key-value pairs in YAML format. Reduce the complexity of configuration down to only the parameters that vary within your organization, such as for different services and environments. Teams can quickly set the parameters they need to configure the Dynatrace platform to their specific requirements. We provide a CLI that is natively built for Dynatrace APIs and platform configurations, as well as apps built for our AppEngine that can be run and customized without third-party dependencies and with a lightweight setup. Additionally, we offer a Terraform provider if you already have a working Terraform-built environment in your organization.

Achieve automated observability and Site Reliability Engineering

To minimize any negative impact caused by problems such as security vulnerabilities, it’s crucial to respond and remediate the issues as quickly as possible. Dynatrace provides automation for detecting problems, and you can opt to automatically run a change-impact analysis report to proactively validate important objectives. This same mechanism can also be leveraged to validate the impact of new software releases on resources, logs, performance, reliability, or business measures. The screenshot below displays a workflow that listens for a deployment event of the easytrade service in the production stage.

Workflow that listens for a deployment event

The validation process is automated based on events that occur, while the objectives’ configuration, which is validated by the Site Reliability Guardian, is stored in a separate file. The screenshot below displays such a configuration. In summary, Configuration as Code enables the automatic execution of validations for a full set of configuration objectives and workflow definitions. Development teams can easily adopt this by providing key-value pairs in a YAML config file, such as the name of a service and the stage to listen for deployment events.

Key-value pairs in a YAML config file

Service-Level Objectives (SLOs) can provide additional support for the insights and outlook of services and components. Whether tracking internal, workload-centric indicators such as errors, duration, or saturation or focusing on the golden signals and other user-centric views such as availability, latency, traffic, or engagement, SLOs-as-code enables coherent and consistent monitoring throughout the environment at scale.

SLO tracking code

Proper notifications or escalations are automated based on ownership information. The Dynatrace platform supports ingesting ownership team information, as shown in the screenshots below. Dedicated configuration files are used to create teams and maintain relevant information, such as responsibilities and contact details, in a scalable and automated way. While this can be achieved through UIs and APIs, providing contact details and links to any supporting material when things go wrong in any environment is most conveniently done by a developer, as code, side-by-side with the software source code.

Configuration as Code side-by-side with software source code

Furthermore, by utilizing workflows, it’s easy to set up dynamically-queried ownership-team information from affected entities if a problem or a security vulnerability is detected. Leverage the full power of the Dynatrace AutomationEngine to automatically inform the right people and automatically resolve problems, all managed and configured easily via configuration as Code.

What’s next

Get started with Dynatrace Configuration as Code, natively built for the Dynatrace platform and third-party independent. You can read all about it in our Configuration as Code documentation. If Terraform is your tool of choice, please have a look at our Terraform provider documentation.

Stay tuned for more examples and easy-to-adopt automations provided in our public Github project.

The Developer’s Guide to Observability

Modern observability is no longer just an operations tool; it’s built for developers. When you bring observability into your IDE, pipelines, and AI driven development workflow, you can surface how code behaves in context across services, environments, and teams.

The post Automated observability, security, and reliability at scale appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automated-observability-security-and-reliability-at-scale/feed/ 0
Automated Change Impact Analysis with Site Reliability Guardian https://www.dynatrace.com/news/blog/site-reliability-guardian/ https://www.dynatrace.com/news/blog/site-reliability-guardian/#respond Thu, 16 Feb 2023 05:00:32 +0000 https://www.dynatrace.com/news/?p=56098 SLOs graphic

The Dynatrace® Site Reliability Guardian simplifies the adoption of DevOps and SRE best practices to ensure reliable, secure, and high-quality releases.

The post Automated Change Impact Analysis with Site Reliability Guardian appeared first on Dynatrace news.

]]>
SLOs graphic

Powered by Grail and the Dynatrace AutomationEngine, Site Reliability Guardian helps DevOps platform teams make better-informed release decisions by utilizing all the contextual observability and application security insights of the Dynatrace platform. The app also enables SREs to set and automate service-level objectives (SLOs) for critical services. Site Reliability Guardian further simplifies automation requirements by using entity and ownership information, alerting the right teams when regressions are detected. Dynatrace users can now unlock answer-driven automation to safely and securely release new features to customers while ensuring production SLOs are intact.

Streamline development and delivery processes

Nowadays, digital transformation strategies are executed by almost every organization across all industries. In the context of application development, digital transformation aims to streamline development and delivery processes to ensure quick delivery to the market with guaranteed high quality and security standards.

To achieve this, many organizations are adopting DevOps practices to provide developers with a delivery platform to release their applications and services autonomously and independently. While this empowers teams to frequently deliver new features, the overall business, security, and quality objectives must be maintained. This is where Site Reliability Engineering (SRE) practices are applied. SREs use Service-Level Indicators (SLI) to see the complete picture of service availability, latency, performance, and capacity across various systems, especially revenue-critical systems. Consequently, Service-Level objectives (SLO) are defined to enact countermeasures before the business is impacted.

Given the momentum of DevOps and SRE, digital transformation goals can be achieved when automation enables organizations to apply best practices rapidly and to keep pace with the scale of the organization and applications.

SLOs graphic

Automation is key for effective collaboration

While DevOps engineers rely on observability and automation to determine whether they can promote a new release into production, SREs look for answers when a new release has a positive or negative impact on application performance, affects user experience, or threatens other SLOs. SREs face ever more challenging situations as environment complexity increases, applications scale up, and organizations grow:

  • Growing dependency graphs result in blind spots and the inability to correlate performance metrics with user experience.
  • Siloed teams and multiple tools make it difficult to align on a single version of the truth for overall system health.

Other observability solutions don’t provide the required automation capabilities that would allow:

  • Automating and speeding up the SLO validation process and quickly reacting to regressions detected in application topology.
  • Informing the right people with the answers they need to implement targeted countermeasures.

Automate the validation of key objectives

Dynatrace evolves Cloud Automation release validation (powered by Keptn) into Site Reliability Guardian and natively enriches the Dynatrace platform by automating change impact analysis. For each service, a guardian can automatically validate service-level objectives, “golden signals,” or security vulnerabilities before and after any deployment or configuration changes. It detects regressions and deviations from previously observed behavior, including latency, traffic, error rates, saturation, security coverage, vulnerability risk levels, and memory consumption. Thus, Site Reliability Guardian supports DevOps and SREs in speeding up release delivery and improving release quality.

  • DevOps achieve safer and more secure releases by applying a gating mechanism that identifies release issues quickly and prevents poor-quality code from being promoted to production.
  • SREs leverage on-demand reliability validation by comparing observability data like service-level objectives against a dedicated release version from the past or as part of progressive delivery.

Get started

Get started with Site Reliability Guardian by selecting the components that you want to guard automatically. Relevant objectives, known as the four golden signals, are suggested by default:

  • Latency: the time it takes to service a request
  • Traffic: a measure of how much demand is placed on your system
  • Errors: the rate of failed requests
  • Saturation: a measure of a system’s remaining capacity, emphasizing the resources that are most constrained

The four golden signals are recommended for end-user-facing services because they have proven to be a good starting point for diving into SRE practices. Additionally, you can easily use any previously defined metrics and SLOs from your environments.

Site Reliability Guardian showing a failed validation
Site Reliability Guardian showing a failed validation due to an increase of error logs for a cart service.

Automation workflow runs validations and notifies responsible team members

With a solid set of objectives to validate, automatically generated workflows take care of fully automated validation for you. Whether triggered by a test result or a new release deployment, detected events work as a trigger to check the defined objectives and derive an overall status automatically. Based on the results, further actions are taken, like sending targeted notifications to the service and application owners. This is all available out-of-the-box with the default workflow template provided by Site Reliability Guardian.

Workflow leveraging Site Reliability Guardian
Workflow leveraging Site Reliability Guardian to make release decisions.

In addition, the workflow can be easily extended to any custom demands, for example, integration with tools that support your software product lifecycle. This includes executing tests, running Dynatrace Synthetic checks, or creating tickets. In this way, Dynatrace supports an ecosystem of tool integrations that unlock the full potential of DevOps and SRE with cloud-native delivery orchestration and remediation automation use cases.

What’s next?

Site Reliability Guardian is built for ensuring and maintaining production reliability and resilience. Further enhancements will focus on seamlessly integrating the solution into your existing pipelines to help you shift left and block poor-quality releases before they reach production. Davis will also assist Site Reliability Guardian in recommending relevant objectives and baselines for comparison.

  • For an early preview of this new functionality, please request Dynatrace platform preview access, which allows you to install the Site Reliability Guardian via Dynatrace Hub.
  • The Dynatrace Site Reliability Guardian will be available within the next 90 days. In the meantime, watch out for upcoming “Observability Clinics” and “Ask me anything” sessions covering the main topics of this blog post. You can either see a list of the next webinars on our website or follow us on LinkedIn to stay up-to-date with upcoming announcements and activities.

The post Automated Change Impact Analysis with Site Reliability Guardian appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/site-reliability-guardian/feed/ 0