cloud automation | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 11 Jun 2026 11:43:04 +0000 en hourly 1 Helping customers unlock the Power of Possible https://www.dynatrace.com/news/blog/helping-customers-unlock-the-power-of-possible/ https://www.dynatrace.com/news/blog/helping-customers-unlock-the-power-of-possible/#respond Tue, 29 Oct 2024 18:32:13 +0000 https://www.dynatrace.com/news/?p=66365 Abstract image depicting unlocking business potential with Dynatrace using power dashboarding

At Dynatrace, The Power of Possible is about uncovering insights that move business forward. Discover how our customers unlock business potential with Dynatrace.

The post Helping customers unlock the Power of Possible appeared first on Dynatrace news.

]]>
Abstract image depicting unlocking business potential with Dynatrace using power dashboarding

We are in the era of data explosion, hybrid and multicloud complexities, and AI growth. In this dynamic landscape, imagine understanding your digital environment every day—what’s working, what’s not, what may have an issue, and more importantly, how to solve it? Picture gaining insights into your business from the perspective of your users. What new possibilities would open up for your organization?

This isn’t about imagining what’s possible just because it’s possible. It’s about uncovering insights that move business forward. This is what companies like BT and TD Bank have achieved by leveraging Dynatrace.

What’s behind it all? The Dynatrace platform automatically captures and maps metrics, logs, traces, events, user experience data, and security signals into a single datastore, performing contextual analytics through a “power of three AI”—combining causal, predictive, and generative AI. Dynatrace analyzes billions of interconnected data points to deliver answers, not just data and dashboards sending signals without a path to resolution.

With over 2.5 quintillion bytes of data generated daily, managing this influx has far surpassed human capacity. Dynatrace transforms this unstructured data into a strategic advantage, processing it automatically—no manual tagging required.

Unlike traditional observability tools that merely bark signals through dashboards without real context, the Dynatrace AI-driven approach goes deeper. It delivers precise, actionable insights to help you understand what’s truly happening in complex environments and resolve issues in real time. It empowers teams to act proactively rather than reactively. And it enables executives to have unprecedented insight into how user experiences, applications and underlying infrastructure health can power their business.

Let’s explore how leading organizations have harnessed the power of end-to-end observability—to reduce costs, drive innovation and acceleration, and deliver exceptional experiences for their customers. You’ll see how a clear line of sight across your entire technology stack can be transformative and learn how to apply these lessons to your own business.

Breaking down complexity with unified observability

In today’s landscape, it’s imperative to know the health of your digital business. But many enterprises face escalating complexity across their digital environments, leading to visibility gaps and inefficiencies using their traditional toolsets. BT, the UK’s largest mobile and fixed broadband provider, faced this challenge when managing multiple monitoring tools across different teams. By consolidating their tools into Dynatrace, they were able to reduce outage times and digital incidents by 50%. This meant better service reliability, reduced costs, and less time spent on incident management—enabling their teams to focus on innovation.

Leveraging AI-driven automation to innovate faster

As its business expanded, TD Bank encountered challenges with frequent IT incidents and slow resolution times, which were creating inefficiencies in their operations. To improve this, they turned to Dynatrace for AI-driven automation to accelerate problem detection and resolution.

By automating root-cause analysis, TD Bank reduced incidents, speeding up resolution times and maintaining system reliability. The result? More time for teams to focus on developing new services and improving customer experience, all while keeping operational costs under control. This ability to innovate faster has given TD Bank a competitive edge in a complex market.

The Power of Possible: Uncovering insights that move business forward

These stories illustrate how leading organizations leverage Dynatrace to achieve real business impact. For BT, simplifying their observability strategy led to faster issue resolution and reduced costs. TD Bank’s AIOps capabilities meant more proactive service management, keeping customers satisfied.

At Dynatrace, we believe that anything is possible for our customers, and we are committed to helping them move beyond simply managing their digital businesses. Our goal is to empower customers with insights that enhance user experiences and drive business growth. As organizations like BT and TD Bank navigate the complexities of modern IT environments, we are humbled to be their trusted partner, helping to turn data into action and transform challenges into opportunities.

Ready to see how Dynatrace makes the impossible possible? Watch the replay of The Power of Possible to learn more.

Want to go deeper? Learn more about the innovations driving these outcomes, and build a case to understand the ROI Dynatrace could deliver for your organization:

Ready to explore how end-to-end observability can unlock new possibilities for your business?

Contact us today to get started.

The post Helping customers unlock the Power of Possible appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/helping-customers-unlock-the-power-of-possible/feed/ 0
Generative AI model observability, cloud modernization take center stage with partners at Dynatrace Perform 2024 https://www.dynatrace.com/news/blog/generative-ai-models-cloud-modernization-perform-2024/ https://www.dynatrace.com/news/blog/generative-ai-models-cloud-modernization-perform-2024/#respond Tue, 23 Jan 2024 20:27:29 +0000 https://www.dynatrace.com/news/?p=61676 Cost monitors

With our annual user conference, Dynatrace Perform 2024 rapidly approaching on January 29 through February 1, 2024, our teams, partners, and customers are buzzing with excitement and anticipation. Perform serves yearly as the marquis Dynatrace event to unveil new announcements, learn about new uses and best practices, and meet with peers and partners alike. At […]

The post Generative AI model observability, cloud modernization take center stage with partners at Dynatrace Perform 2024 appeared first on Dynatrace news.

]]>
Cost monitors

With our annual user conference, Dynatrace Perform 2024 rapidly approaching on January 29 through February 1, 2024, our teams, partners, and customers are buzzing with excitement and anticipation. Perform serves yearly as the marquis Dynatrace event to unveil new announcements, learn about new uses and best practices, and meet with peers and partners alike. At this year’s Perform, we are thrilled to have our three strategic cloud partners, Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP), returning as both sponsors and presenters to share their expertise about cloud modernization and observability of generative AI models.

Dynatrace unified observability and security is critical to not only keeping systems high performing and risk-free, but also to accelerating customer migration, adoption, and efficient usage of their cloud of choice. More so than ever before, organizations are investing in cloud migration and cloud modernization to lower total cost of ownership (TCO). These investments, in turn, extend to unifying observability with context and intelligence across increasingly dynamic and complex cloud environments, and ensuring that their cloud ecosystems are not only reliable and secure but also optimized to realize resource savings and accelerate release delivery.

At this year’s Perform, all three of our cloud partners will take the stage to speak about these benefits—and how to achieve them—in their expert-led sessions. Takeaways will be immediately applicable for both technologists and business stakeholders alike, helping attendees put into place the right observability and security practices with Dynatrace to further advance their cloud modernization journey (and make it a smooth one at that).

Read on to learn what you can look forward to hearing about from each of our cloud partners at Perform. If you’re unable to join us in Las Vegas, be sure to register to attend virtually—or view sessions on-demand afterward—so you don’t miss out!

Accelerating AWS migration and optimizing efficiency with Dynatrace

As many companies embrace cloud modernization and begin migrating to the AWS cloud (or continuing to move existing workloads), complexity can introduce uncertainty into the process. What can we move? What will the new architecture be? How can we ensure we see performance gains once migrated? These are the big questions that have slowed, or prevented, many teams from migrating.

In an upcoming partner session at Dynatrace Perform 2024, Mark Jaggers, AWS technical program manager, will detail the key steps of the migration process and showcase where Dynatrace is integral in not only answering these questions, but providing the technical ability and automation to migrate more quickly and confidently.

Session attendees will learn first-hand how Dynatrace natively integrates into the AWS Migration Hub to provide a full topology of on-prem workloads and dependencies in order to generate the ideal cloud-based architecture in the AWS cloud. Additionally, discover how Dynatrace is easily deployed on newly migrated workloads to get instant insights into performance, utilization, security, and efficiency post-migration. Finally, Mark will take attendees a step further to demonstrate how Dynatrace underpins the AWS Well-Architected pillars of cost optimization and operational excellence by helping enterprises to right-size AWS resources with utilization metrics and configuration for continuous efficiency in the cloud.

Learn more about Dynatrace and AWS in the whitepaper, Why modern, well-architected AWS clouds demand AI-powered observability.

Microsoft and Dynatrace solve cloud modernization complexities with generative AI models

In the cacophony of digital noise, where every buzzword promises to revolutionize businesses, how do organizations discern the transformative from the trivial? The struggle to prioritize digital transformation is real, and in this ever-evolving landscape, standing still is not an option. Innovation and cloud modernization aren’t luxuries; they’re the heartbeat of progress.

In this upcoming partner session at Dynatrace Perform 2024, Peter Laudati, Microsoft Cloud solution architect – GPS US, and Jay Gurbani, Dynatrace senior technical partner manager, will share how Microsoft and Dynatrace are helping enterprises solve the complexities introduced by cloud modernization and how organizations can use tools and innovations to unlock the true potential of the digital ecosystem. The session will explore leveraging AI for real-time business decisions and real-world applications.

Learn more about Dynatrace and Microsoft in the whitepaper, Why modern, well-architected Azure clouds demand AI-powered observability.

Enhancing generative AI models in real-world production settings with GCP and Dynatrace

In the ever-evolving landscape of artificial intelligence, the fusion of leading technologies can yield unparalleled results. This partner Perform session will delve into the exploration of Dynatrace, a leading observability and security platform, and Vertex AI Generative AI, a suite of tools within Google Cloud designed for constructing and deploying generative AI models. Attendees can anticipate an examination of the technical intricacies involved in both platforms, offering insights into the possibilities of both systems’ power together. The focus will be on empowering users with the knowledge of how monitoring, analyzing, and optimizing tools can help enhance generative AI models in real-world production settings.

Led by Merlin Yamssi, lead solutions consultant at Google Cloud’s AI/ML CoE partner engineering, and Mike Villiger, senior manager of technical alliances at Dynatrace, this partner session at Perform 2024 will explore Site Reliability Engineering (SRE) and the utilization of the Four Golden Signals for AI Observability. Attendees will gain valuable insight on how Dynatrace observability can complement key Google Cloud AI tools like Duet AI and Vertex AI through its Google Cloud integration and automated discovery processes. Learn more about enhancing system reliability by proactively detecting issues and understanding the capability of Dynatrace and Google Cloud in generative AI models and observability.

Learn more about Dynatrace and GCP from the ebook 5 Key Considerations for Monitoring Google Cloud.

From Las Vegas to the enterprise cloud

The learnings from Perform will be as vast and robust as the enterprise cloud itself and illuminate how Dynatrace is the (not so) secret weapon to accelerating your cloud modernization journey. Beyond the three breakout sessions from our cloud partners, as an attendee, you can also visit them in our expo hall and talk through unique challenges and use cases to further empower you to supercharge your own digital transformation in the cloud.

Don’t miss your chance to learn from and meet with these cloud powerhouses at Dynatrace Perform 2024. Register now to attend in person (or virtually). We’ll see you in Las Vegas and follow you to the cloud!

The post Generative AI model observability, cloud modernization take center stage with partners at Dynatrace Perform 2024 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/generative-ai-models-cloud-modernization-perform-2024/feed/ 0
Path to NoOps part 2: How infrastructure as code makes cloud automation attainable—and repeatable—at scale https://www.dynatrace.com/news/blog/infrastructure-as-code-for-cloud-automation/ https://www.dynatrace.com/news/blog/infrastructure-as-code-for-cloud-automation/#respond Tue, 29 Nov 2022 21:47:19 +0000 https://www.dynatrace.com/news/?p=53748 Dynatrace ensures continuous software quality by combining synthetic monitoring and automatic release validation, infrastructure as code, and cloud automation

NoOps may seem like a pipe dream. But our own experience at Dynatrace illustrates how NoOps can be a lot more attainable than you think by embracing an infrastructure-as-code mindset.

The post Path to NoOps part 2: How infrastructure as code makes cloud automation attainable—and repeatable—at scale appeared first on Dynatrace news.

]]>
Dynatrace ensures continuous software quality by combining synthetic monitoring and automatic release validation, infrastructure as code, and cloud automation

Infrastructure as code is a way to automate infrastructure provisioning and management. And it’s a crucial step toward achieving cloud automation on the path to NoOps.

In my previous blog post, Path to NoOps part 1: How modern AIOps brings NoOps within reach, I explored the aspirations of NoOps and how modern AIOps makes it possible. But how does it work in practice? Is it practical? And can other organizations realistically expect the same results? In this blog, I explore how Dynatrace has made cloud automation attainable—and repeatable—at scale by embracing the principles of infrastructure as code.

NoOps through modern AIOps: The Dynatrace story

As a torchbearer of modern AIOps, the Dynatrace’ AI engine, Davis®, provides a purpose-built AI platform for today’s web-scale modern cloud. Davis analyzes hundreds of billions of dependencies and auto-discovers billions of dynamic topology changes per second.
Mean time to repair (MTTR) includes the time teams need to detect, identify, fix, and verify issues. Without automation, these operations can span a wide range.

mean time to repair can span a long time without cloud automation

With its AI platform approach, however, Dynatrace eliminates mean time to detect (MTTD) and mean time to identify (MTTI) an issue’s root cause. Using AI, Dynatrace instantly and automatically identifies problems. Dynatrace can then automatically raise a service ticket and directly update the root cause in the service ticket. This automatic response eliminates time-consuming triage and reduces mean time to repair (MTTR) to just a few minutes, including fixing and verifying the issue.

cloud automation enables partial automation of MTTR processes

With this kind of reliable and automatic intelligence, teams can even automate fixing and verifying, thus attaining NoOps.

cloud automation makes it possible to fully automate MTTR for NoOps

Dynatrace IT itself has implemented a NoOps model for our own IT operations. With a skeleton staff of 7 on a on 24×7 schedule, we increased the number releases from 2 to 26 per year and reduced production bugs by 93% since 2014.

Dynatrace stats on its own NoOps capability using infrastructure as code and cloud automation
Dynatrace achieved its own IT transformation using infrastructure-as-code principles and Cloud Automation.

Hear the story of how Dynatrace achieved “NoOps” directly from our Chief Technology Officer, Bernd Greifeneder in “From 0 to NoOps in 80 days.”

But to fully integrate and automate our cloud-native continuous delivery and operations, we needed a control plane to automatically orchestrate most of the operations tasks. So we built one: The Dynatrace Cloud Automation control plane.

What is Dynatrace Cloud Automation?

Dynatrace Cloud Automation is an enterprise-grade control plane that extends intelligent observability, automation, and orchestration capabilities of the Dynatrace platform to DevOps pipelines. The goal of Cloud Automation is for development teams to build better software faster and operations to automate mundane repetitive tasks and focus on innovation.

This AI-driven control plane further abstracts away the complexity of underlying webhook-based integrations, enabling IT to assemble end-to-end processes that manage related interdependencies regardless of the underlying technology. As a result, IT teams can automate and manage processes across on-premises and cloud-based systems, or between multiple cloud services to prevent vendor lock-in.

Cloud Automation use cases

Let’s look at some scenarios where teams can use Cloud Automation in their NoOps journey.

Closed-loop remediation

Remediating an issue consists of the following stages:

  1. Identify the fix
  2. Apply the fix
  3. Validate the fix
  4. Update the service ticket
  5. Notify the relevant people

Individually, the tasks may take only a few minutes, but they add considerable overhead for the support staff. Many organizations execute scripts or runbooks manually to remediate trivial issues. These teams often hesitate to automate these runbooks because the triggering condition could result in a false positive. In other words, the condition could match even when the issue has disappeared or was not even there in the first place. As a result, teams must verify their data first because the triggering condition could be based on discovered data that is stale.

However, by advancing AIOps, Dynatrace considers dynamic CI relationships and dependencies instantly and automatically. Hence there are far fewer chances for false positives. Using Davis, Cloud Automation can trigger the right fix for an issue, validate the fix by running a synthetic test, update the service ticket, and notify stakeholders using communication channels—all in an automated way.

Davis AI automatically detects problems in incident response using cloud automation and infrastructure as code

The benefits of this closed-loop remediation include the following:

  • Freeing up your support staff from routine tasks
  • Lower MTTR
  • More controlled, consistent, and sustainable problem resolution
  • Transparency and scalability

Infrastructure-as-code

Infrastructure as code uses a declarative language to achieve the desired state as opposed to scripts that define a set of steps to execute. The purpose of infrastructure as code is to enable developers or operations teams to automatically manage, monitor, and provision resources, rather than manually configure discrete hardware devices and operating systems. Infrastructure as code is sometimes referred to as programmable or software-defined infrastructure. With the proliferation of infrastructure-as-code tools, operations teams can:

  • Deploy, configure, or tear down workloads into an instance in real-time
  • Ramp up or down resources in real-time based on workload requirements
  • Block or provision access to resources in real time based on threats or requests
  • Proactively manage web and mobile applications based on user experience or traffic

Embracing the concept of infrastructure as code has also helped to ease our own DevSecOps journey. Like our customers, we use Dynatrace to monitor thousands of environments and supporting applications, and we needed a way to streamline the configuration process. In response, Dynatrace introduced Monaco (Monitoring-as-code). Using Monaco, organizations can offer true application monitoring as a self-service to its users and applications. Teams can now onboarding hosts to the Dynatrace platform can be done in a few minutes instead of hours, reducing human dependencies.

Through its webhook-based integrations, Cloud Automation can further push configurations and deployments to a wide variety of tools using their infrastructure-as-code capabilities, through its webhook-based integrations.

SecOps

Dynatrace provides real-time, continuous surveillance of day 1 and day 2 vulnerabilities with full runtime context, no additional agents to install, no static scans, and no manual analysis for handling day 1 and day 2 vulnerabilities in real-time. The Dynatrace AI engine, Davis, provides intelligence and context to such detected events and helps to decide the remediation workflow automatically. These actions could be resetting a password, disabling a VM, blocking an IP address, or even patching a vulnerable application.

Using Dynatrace Cloud Automation, teams can easily mitigate data breaches and security violations.

Infrastructure as code and cloud automation pave the way to NoOps

While NoOps could probably be a never-ending journey, it is not just a passing cloud. It’s achievable by continuously innovating autonomic applications that require no manual intervention when an issue arises. With constant innovations based on our own path to NoOps, Dynatrace helps organizations on their own NoOps journeys. While it may not eliminate the requirement for Ops staff altogether, organizations can still benefit by greatly reducing their investment in operations and focusing on customer satisfaction, and improving application performance.

To learn more about how Dynatrace modern AIOps makes NoOps possible, join me on December 14, 2022, for our live event, the DevOps and SRE Virtual Workshop: Realize NoOps using Dynatrace Cloud Automation.

The post Path to NoOps part 2: How infrastructure as code makes cloud automation attainable—and repeatable—at scale appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/infrastructure-as-code-for-cloud-automation/feed/ 0
How OpenFeature — and feature flag standardization — enables high-quality continuous software delivery https://www.dynatrace.com/news/blog/openfeature-and-feature-flag-standardization/ https://www.dynatrace.com/news/blog/openfeature-and-feature-flag-standardization/#respond Mon, 24 Oct 2022 09:39:15 +0000 https://www.dynatrace.com/news/?p=53990 Dynatrace | OpenFeature

Earlier this year, Dynatrace announced its involvement in the open source feature flagging project OpenFeature that enables fast-paced, high-quality software development. Since then, the Cloud Native Computing Foundation (CNCF) has accepted OpenFeature as a sandbox project. With the support of many of the top feature flag companies and practitioners, OpenFeature has developed a vendor-neutral specification, […]

The post How OpenFeature — and feature flag standardization — enables high-quality continuous software delivery appeared first on Dynatrace news.

]]>
Dynatrace | OpenFeature

Earlier this year, Dynatrace announced its involvement in the open source feature flagging project OpenFeature that enables fast-paced, high-quality software development. Since then, the Cloud Native Computing Foundation (CNCF) has accepted OpenFeature as a sandbox project. With the support of many of the top feature flag companies and practitioners, OpenFeature has developed a vendor-neutral specification, and its software development kits (SDKs) for Java, JavaScript, .NET, and Go SDKs are now generally available as a 1.0 release.

Why do organizations need feature flags?

Organizations need to release software at a high velocity to stay competitive as the pace of business accelerates, but they can’t sacrifice software quality for speed. Agile companies that adopt continuous software delivery have increasingly enlisted feature flagging to enable more frequent code releases. Additionally, the open source and cloud-native communities have continued to develop open standards to ease communication between cloud-native tools and promote interoperability between vendors.

What is a feature flag?

In its simplest form, a feature flag is a toggle (switch) in an application that turns functionality on or off at runtime without deploying new code.

Feature flagging has been around for many years but has recently gained popularity, largely because companies need to innovate as quickly as possible. Feature flags support the rapid development of software features by allowing teams to decouple feature releases from deployments. Using feature flags, developers can test an unreleased feature by enabling access to it for a specific user, department, company, or any other desired unit. This allows new features to be safely tested in any environment, including production.

Feature flags example

As companies become more comfortable with feature flags, they’re using them for more than just rolling out new features. For example, teams use feature flagging for hypothesis-driven development and experimentation, which are now an essential part of the development lifecycle. Many feature flag tools and vendors support multivariant feature flags—companies can assign users to feature variants that are determined by any available factor (for example, company, geography, or timestamp) or pseudo-randomly. Teams can then measure the impacts and make data-driven decisions.

What is OpenFeature and why should organizations use it?

A common first step into feature flagging is typically a homegrown solution. However, custom-built tools require maintenance and support, which can be burdensome for teams and can limit the growth of these internal tools. Commercial solutions provide maintenance and support; however, both homegrown and commercial solutions can present interoperability challenges that complicate future requirements.

OpenFeature is an open source, vendor-agnostic feature-flagging application programming interface (API) that allows teams to start using feature flags quickly and confidently. It works with your preferred feature-flag management vendor or custom-built tool. As a result, teams can flexibly choose a feature flagging method that fits their current requirements while being able to switch easily to a different method if requirements change.

OpenFeature users can benefit from an ever-growing community of feature flag experts. The OpenFeature community pools the expertise of many top feature flag companies to develop open solutions to topics relevant to feature flagging, including an OpenTelemetry integration.

What does OpenFeature mean for observability?

Given the large number of vendors in the feature flag market, it’s impossible for observability vendors to support all vendor-provided and homegrown feature flag solutions natively. Instead of focusing on the feature level, most observability vendors treat each feature-flag toggle event (such as turning a feature on) like a new deployment or release. Teams then need to infer a feature’s impact by comparing metrics associated with requests before and after the toggle event.

Conversely, teams can integrate OpenFeature’s vendor-agnostic SDKs with any feature flag tool, easily enabling consistent observability support. By using our PurePath® technology for distributed tracing, Dynatrace can see all the feature flags that were evaluated for a given request and what values were returned. This allows teams to answer questions such as the following confidently.

  • What impact did a feature flag value have on a request?
  • Which services are using a particular feature flag?
  • Did a combination of feature flags cause unexpected behavior?

Teams can also enable feature flags for a subset of users (for example, for beta testing or A/B testing). Observing these cases as you would a traditional deployment can result in the impact of a new feature becoming lost in the noise. By analyzing feature flags at the level of distributed traces, Dynatrace can collect statistically significant metrics and compare a new feature’s behavior to that of the control. This allows teams to make better, data-driven decisions.

Dynatrace screenshot feature flag
Feature flag summary Dynatrace screenshot

What’s next for Dynatrace and OpenFeature?

OpenFeature provides the Dynatrace platform with a great opportunity to extend observability into feature flagging, allowing teams to monitor the deployment of a new feature, the results of an experiment, or a coordinated rollout. Combining this additional context with the already rich data within Dynatrace gives teams the confidence they need to safely release features without sacrificing velocity. With support for the Dynatrace Grail™ data lakehouse technology and Davis AI on the roadmap, now is the perfect time to start your journey with Dynatrace and OpenFeature.

Stop by the Dynatrace booth (#P24) to find out more about OpenFeature at KubeCon US (Detroit, MI) from October 26-28, 2022.

The post How OpenFeature — and feature flag standardization — enables high-quality continuous software delivery appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/openfeature-and-feature-flag-standardization/feed/ 0
Level up your production resiliency with automated problem remediation https://www.dynatrace.com/news/blog/dynatrace-automated-problem-remediation/ https://www.dynatrace.com/news/blog/dynatrace-automated-problem-remediation/#respond Wed, 20 Jul 2022 15:38:53 +0000 https://www.dynatrace.com/news/?p=52183 SLOs

Dynatrace Davis® AI detects problems, underlying root causes, business impact, and SLO impact across your full-stack production deployments. Dynatrace® Cloud Automation uses this information to trigger and execute remediation actions and observability-based validations. This blog post explains how Dynatrace Cloud Automation protects your production environments with automated remediation actions.

The post Level up your production resiliency with automated problem remediation appeared first on Dynatrace news.

]]>
SLOs

While development teams focus on innovation and quick release cycles, operations teams and Site Reliability Engineers (SREs) focus on stability. Of course, your customers expect both fast innovation and reliability.

Dynatrace Cloud Automation lets you track your most critical SLO). Further, it enables you to release with confidence by catching poor quality code before it reaches production. Now, we’re happy to extend our Cloud Automation offering with automated problem remediation.

Dynatrace Cloud Automation offering with automated problem remediation screenshot

The costs of downtime

Surveys suggest that a single hour of downtime can cost an organization from $1 million to over $5 million. For Fortune 1,000 companies, the average costs of unplanned application downtime per year are $1.25 billion to $2.5 billion.

Beyond extensive financial costs, broken SLOs lead to customer dissatisfaction, and customers might look elsewhere for similar services. A tarnished brand reputation might discourage prospects from trying your services, slow down your customer growth, and even hinder your talent recruiting.

Root cause-based automated problem remediation

What are the requirements for rapid problem remediation that prevents downtime? When looking at reports such as the DevOps Automation report 2021, it becomes clear that the most significant challenges during remediation are manual toil (lack of automation) as well as challenges related to communication, for example, reaching the right people, using the right runbooks, and ensuring that decisions are based on reliable data.

Dynatrace Davis AI detects problems, underlying root causes, business impact, and SLO impact across your full-stack production deployments. Dynatrace Cloud Automation uses this information to trigger and execute remediation actions and observability-based validations.

Based on data insights and customer research, we’ve identified the top five use cases for automated problem remediation:

  • Feature flag settings—Observe application and service behaviors, identify error-causing feature flags, and switch them accordingly to guarantee stable environments.
  • Process restarts (for example, JVM memory leaks)—Trigger a service restart or related actions for applications with underlying bug fixes that have been deprioritized or delayed.
  • Kubernetes resource adoption—Act on external, holistic, and customer-centric behavior observations—rather than on only internal parameters—and automatically roll out Custom Resource Definitions (CRDs) to designated environments.
  • Deployment and rollback—Trigger predefined rollback or roll-forward actions when a faulty deployment violates SLOs or decreases your error budget above the target.
  • Targeted notifications—Based on the auto-detected details of underlying root causes, keep your business and technical users, SREs, and Operations team updated regarding ongoing remediation actions and escalate if the situation requires higher visibility.

Take out the guesswork

Let’s look at an example scenario where automated problem remediation is applied. Consider that Davis AI has detected that turning on a particular feature flag leads to a failure rate increase and raises a problem. While this information is used to alert Operations or SREs, Dynatrace now allows you to link root causes to freely configurable runbooks. In this scenario, the runbook can turn off the failure-rate-increasing feature flag while checking the system’s health to see if the remediation action fixed the problem. Once the problem is fixed, Dynatrace Davis AI closes the problem. Otherwise, the next escalation level of the runbook is executed, for example, sending a notification to a human operator.

Dynatrace Cloud Automation offering with automated problem remediation screenshot
Fig 2. Dynatrace Cloud Automation turns off a faulty feature flag to prevent downtime.

What’s next

This blog post is the first in a series of publications centered around different problem remediation use cases and exemplary toolchain integrations. Check out further information in our SLO documentation.

Stay tuned for the next blog post in this series to learn how to extend problem remediation beyond the feature flag mechanism and level up your software delivery by integrating Cloud Automation into your existing DevOps toolchain. Then you can orchestrate the software development lifecycle and remediate issues automatically. Reach out to us today for a demo.

The post Level up your production resiliency with automated problem remediation appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-automated-problem-remediation/feed/ 0
Scale DevOps and SRE with open source Keptn https://www.dynatrace.com/news/blog/scale-devops-and-sre-with-open-source-keptn/ https://www.dynatrace.com/news/blog/scale-devops-and-sre-with-open-source-keptn/#respond Mon, 18 Apr 2022 19:28:21 +0000 https://www.dynatrace.com/news/?p=50026 Scale DevOps and SRE

Site reliability engineering (SRE) using DevOps requires high levels of automation to accommodate rapid testing and release schedules.

The post Scale DevOps and SRE with open source Keptn appeared first on Dynatrace news.

]]>
Scale DevOps and SRE

When it comes to site reliability engineering (SRE) initiatives adopting DevOps practices, developers and operations teams frequently find themselves at odds with one another. Developers want to write high-quality code and deploy it quickly. Operations teams want to make sure the system doesn’t break.

Observability seeks to find a happy medium between the two. Observability is a set of practices and technologies that helps IT teams understand what’s happening across complex environments so that teams can detect and resolve issues quickly, without disruption to users.

Andreas Grabner, DevOps Activist at Dynatrace, took to the virtual stage at the recent Dynatrace Perform conference to describe how the open source Keptn project automates the configuration of observability tools, dashboards, and alerting based on service-level objectives (SLOs).

Keptn: A reference implementation of Google’s SRE principles

Keptn is an open source control plane that enables cloud-native continuous delivery and automated operations. The goal is to accelerate innovation by eliminating the need for custom automation scripts and point-to-point tool integrations.

Dynatrace developed and released Keptn to open source in 2020. Software engineer Taras Tsugrii of Meta (formerly Facebook) paid Keptn a high compliment, saying it feels like a reference implementation of Google’s SRE principles, which are the search giant’s techniques for ensuring the integrity of its sites and services. Keptn’s global user base is developing, as is adoption within companies including Citrix, Amasol, ERT, and Kitopi. One particular use case for Austrian banking software developer Raiffeisen involves using Keptn to automate the production release and readiness validation of all its products using scoring metrics.

One of Keptn’s key advantages is that it automates SLO-driven, multi-stage delivery of apps and services. Why is automated orchestration critical?

Scale DevOps and SRE with open source Keptn

Too many SLOs create complexity for DevOps

DevOps and SRE engineers experience a lot of pressure to deliver applications faster and that adhere to standards like “the five nines” of availability, resulting in many new service level requirements. Developers also need to automate the release process to speed up deployment and reliability.

SLOs are a great way to define what software should do. But as teams adopt them more widely, they sometimes don’t break SLOs down into individual goals that make sense for development and release management.

By shifting SLOs left from operations back to development, teams can detect problems more quickly. SREs can then use SLOs for release quality checks, such as big bang, blue/green, and canary testing.

Bringing SLOs back into development also gives developers production-stage feedback on critical metrics that may later impact business SLOs. With instant feedback enabling teams to release clean software, developers can react faster and speed up the delivery of high-quality content.

With many pipelines to maintain, DevOps teams need automated orchestration. Standard automation techniques like scripting can be very powerful, but also quite complex.

Limits of scripting for DevOps and SRE

Classic automation has limits. Dieter Landenahuf, a senior ACE Engineer at Dynatrace, built Jenkins pipelines for new microservice architectures by creating templates and copying, pasting, and modifying them slightly. This created a classic “snowflake effect” because of the risk of code duplication: if something breaks, you need to fix it in multiple places. The process is error-prone, manual, and doesn’t scale.

Keptn reduces the complexity of pipelines while bringing automation into the delivery processes so that developers can focus on SREs.

Scale DevOps and SRE with open source Keptn

Keptn presents a declarative way to define orchestration so that developers can create automation sequences based on what should happen for delivery, remediation, and testing. Developers can define any sequence with chains of preferred tools. Keptn retains the individual tool configurations under version control.

Optimizations are based on SLOs, meaning Keptn decides whether to move forward with an orchestration based on SLOs. Keptn doesn’t replace the tools developers prefer: it connects them. Since Keptn is an open-source project, it will be easier to eventually replace tools because they all adhere to the same event standard. Teams can plug in new tools for testing, deployment, or even monitoring without having to modify many existing automation tools.

Keptn lets developers keep their favorite tools

Developers can use their existing tools to build artifacts. If they need to add more tasks — such as enforcing SLOs, automating testing, or connecting an observability platform — Keptn handles that for them. Keptn orchestrates all tools so that developers don’t have to. It simply reaches out to monitoring platforms like Dynatrace to extract the necessary SLOs. There is no need to write or maintain any customizations, nor figure out when the job is done and how to parse the results.

Keptn includes best practices that help developers choose which sequences to use. Ultimately, Keptn reduces code automation by 90% and makes every component SLO-driven. It’s based on open standards and is fully declarative. Everything is code and version-controlled in GitOps.

Dynatrace has an enterprise version of Keptn as part of its cloud software-as-a-service (SaaS). Here are some of the advantages of using the enterprise version:

  • Self-manages.
  • Authenticates using Dynatrace single sign-on.
  • Automatic upgrades.
  • Enables developers to plug in additional tools.

Developers can stay in production sequences. Every time the system creates a new artifact, Keptn triggers a new delivery sequence using tools based on events that the developer has subscribed to. Keptn keeps all the configuration information for every state along with Helm scripts and SLOs. Developers can also create different sequences, such as deployment using Monaco.

SLO-based release validation made easy

Release validation centers around SLO-based evaluations. To check quality, developers can simply trigger a release that evaluates performance against SLOs. The dashboard provides a heat map and every service level indicator (SLI) metric, criteria, and scoring.

There are two configuration options. The first is to configure as code using YAML and upload the configuration to a Git repository. The second, for Dynatrace users, is the Dashboard Link, which defines a dashboard with all the relevant SLOs and technical metrics. The dashboard becomes the basis for automated validation — ­­with all SLIs and SLOs stored in Dynatrace — and problems detected by Dynatrace’s Davis automated dependency analysis.

Scale DevOps and SRE with open source Keptn

Developers can also configure Slack and send notifications from Keptn cloud automation every time the system orchestrates a process. They can have multiple tools that subscribe to events, as well as a webhook service that lets them send events to an external tool.

Charting the course with Keptn

Developers who have been building many automation scripts for SRE should take a look at Keptn. The current version is 0.12, and the roadmap includes high-availability improvements, role-based access controls, and improved GitOps and auto-remediation. A Continuous Delivery Foundation special interest group is now standardizing events and the team is preparing a proposal for incubation in the Cloud Native Computing Foundation. Dynatrace’s hosted cloud version of Keptn, delivered through Cloud Automation, has planned support for a custom execution plane and custom integrations.

Check out Andreas’ presentation here along with other sessions from Perform 2022.

The post Scale DevOps and SRE with open source Keptn appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/scale-devops-and-sre-with-open-source-keptn/feed/ 0
How to start with SLOs to align Business, DevOps, and SREs https://www.dynatrace.com/news/blog/how-to-start-with-slos/ https://www.dynatrace.com/news/blog/how-to-start-with-slos/#respond Thu, 16 Dec 2021 12:24:20 +0000 https://www.dynatrace.com/news/?p=47361 Business analytics graphic

Service-level objectives (SLOs) are a great tool to align business goals with the technical goals that drive DevOps (Speed of Delivery) and Site Reliability Engineering (SRE) (Ensuring Production Resiliency). In a recent workshop I did with a global player in the financial market we used their new mobile banking app as a reference. The business […]

The post How to start with SLOs to align Business, DevOps, and SREs appeared first on Dynatrace news.

]]>
Business analytics graphic

Service-level objectives (SLOs) are a great tool to align business goals with the technical goals that drive DevOps (Speed of Delivery) and Site Reliability Engineering (SRE) (Ensuring Production Resiliency).

In a recent workshop I did with a global player in the financial market we used their new mobile banking app as a reference. The business said it wanted to increase the adoption of the new app vs the existing app. The team also wanted to increase the rating of the app which was – by analyzing the reviews – heavily impacted by crashes and slow responsiveness of the current versions. Based on these requirements and status we came up with the following proposal:

SLOs are a great way to align business goals with technical goals from DevOps and SRE
SLOs are a great way to align business goals with technical goals from DevOps and SRE

We first focused on the end-user because that’s what the business cares about. Therefore, the first three goals/objectives are all related to the mobile app itself including adoption, app store rating, and crashes. After that, we sat down and derived SLOs that are relevant for the backend services that have an impact on those end user-centric SLOs. We picked the Authentication Service as it’s the one first involved when a user opens the app. If that service is slow, failing, or not available at all it results in frustration mentioned in some of the comments on social media and the app store.

Introduction Objective Driven Development (ODD) for some Business SLOs

You could argue that App Adoption and App Rating are more goals than objectives as the values of today’s adoption and rating numbers are far below the proposed targets. But I would then argue that a goal is objective and therefore we should call it an SLO. The thing to keep in mind is that those SLOs will be RED (=violating the objective) in the beginning, and that you must keep working on it until they hopefully become GREEN (=meeting the objective). It’s the same concept as Test Driven Development (TDD) where you start with tests that will fail until you finish implementing the code so tests will succeed. Maybe we should call it ODD where we define an objective and work towards it.

I think it’s great to define those SLOs and put them on a dashboard you can give to your business. Instead of focusing on the error budget, which doesn’t make sense in this scenario, give them progress indicators so they can see whether you are moving towards or away from your goal. Below is an example of such a simple dashboard.

Mobile app rating is a good example for Objective Driven Development.
Mobile app rating is a good example of Objective Driven Development.

How to measure those end-user and service focused SLOs with Dynatrace

In the workshop, I also answered the question: How can we measure those metrics (=SLIs) that are behind our objectives? In Dynatrace that’s easy:

App Adoption Rate

Dynatrace’s Real User Monitoring (RUM) offering provides observability to every end-user that uses your mobile or web applications. For the adoption rate of the new vs the old app, we can simply take the number of unique users of the old app and divide it by the unique users of the new app to give you the rating.

App Rating

Dynatrace provides several ways to ingest data from external data sources. Whether its our Metrics Ingest API or building a Dynatrace Extension. This allows pulling or pushing mobile app ratings through the APIs that Google and Apple offer into Dynatrace. A recent blog from Wolfgang Beer discusses ingesting external data including blogs on SLOs to safeguard mobile app revenue.

Mobile Crashes

Dynatrace’s RUM for Mobile Apps provides crash analytics by default. Besides just a ratio, Dynatrace provides crash details including crash reports and stack traces which are great for developers to analyze and fix issues. For our SLO the only thing we need is the default Mobile Crash Rate metric.

Availability

For availability, I always propose to use Dynatrace Synthetic vs looking at real user traffic. Why? Because Synthetic tests are predictable and eliminate any seasonal behavior or impact of the end user’s environment (defect hardware, bad Wi-Fi, etc.). Dynatrace Synthetic allows us to test API endpoints as well as full end-user browser click paths. They are easy to set up and by default deliver an availability metric we can use for our SLOs.

Response time

Dynatrace automatically monitors the response time of all your services, either through using the Dynatrace OneAgent auto-instrumentation and ability to ingest OpenTelemetry data or by ingesting response time metrics from external sources such as Prometheus. It’s important you get response time metrics for the specific endpoints of a service, e.g., /authenticate. Also debate whether to exclude response time from certain requests, e.g., exclude requests from bogus requests or internal heartbeat requests. Dynatrace allows you to do all that, so you truly get the response time for the SLO you need.

Error rate

The error rate is similar to response time. Dynatrace has your request specific metrics including Error Rate, and allows you to configure what error really means – for example, do you want to include errors that are a result of a user making a typo in their username or not? Probably not. Dynatrace gives you that option, so you can get the exact Error Rate you need for your SLOs.

Creating an SLO dashboard for Business, DevOps, and SREs

Dynatrace’s SLO capability makes it easy to create the SLOs we just discussed. In my workshop, I showed typical SLO dashboards I build. I’m not just adding single or multiple SLOs on a dashboard; instead, for business and DevOps and SRE teams, I like to build a dashboard that shows the same SLO in different timeframes. The screenshot below is an example of this build of the dashboard with an explanation of what each SLO and the respective timeframe tells me:

Analyzing SLOs across multiple timeframes provide input for tactical (short term) and strategic (long term) decisions
Analyzing SLOs across multiple timeframes provide input for tactical (short term) and strategic (long term) decisions

The dashboard layout above is nothing I came up with. I saw this with other customers in the past where they explained how they analyze different timeframes to support tactical and strategical decisions. Simple yet powerful! And Dynatrace supports all this because SLOs can be viewed in different timeframes, either through our SLO tile including error budget or over time with our powerful charts.

From SLO reporting to ensuring business resiliency

What we have walked through so far is rather straightforward you may say. It’s just a bunch of metrics against some thresholds and then we put them on a dashboard in different colors and timeframes, and I agree! There is no real magic! It’s just simple reporting that allows you to react to a problem that’s already known based on those metrics (SLIs) you have specified as SLOs.

Dynatrace’s Davis AI takes SLO up a level; instead of just using it for reporting purposes, Dynatrace alerts on problems that will impact your SLOs before the SLO itself is impacted. This is possible through our automated problem detection and the fact that Dynatrace knows the dependencies of your SLOs to all backend supporting apps, services, and infrastructure. This capability was nicely explained in a recent blog.

Below is a high-level overview of what I often call “Minority report style pre-crime alerting”. It gives you time to act upon a problem before it becomes a real problem.

Dynatrace’s Pre-Crime Alerting on SLOs is what elevates regular SLO reporting to ensuring business resiliency
Dynatrace’s Pre-Crime Alerting on SLOs is what elevates regular SLO reporting to ensuring business resiliency

If you are thinking about SLOs, I hope this shows you that SLOs shouldn’t just be used for reporting or reacting to issues when they become apparent based on your error budget burn down. Modern observability coupled with deterministic AI – like Dynatrace offers – enables you to put out a flame before it becomes a fire. Explore how Dynatrace can help you create and monitor SLOs, track error budgets, and predict violations/problems before they occur, so teams can prevent failures and downtime.

Next step: Automating and scaling SLOs across from Ops to Dev

My workshops like with the mentioned customer in the financial sector where we discuss “How to start with SLOs” and “What SLOs are good for me” are just the start. The true power of SLOs comes when everyone in the value creation chain is actively working towards meeting SLOs. This requires a holistic approach to SLOs as part of the development lifecycle and not just being a siloed activity in production.

If you want to learn more, watch my latest Performance Clinic Automating SLOs from Ops to Dev with Dynatrace where I walk through creating your first SLOs, creating SLO dashboards, using SLOs as part of release validation, and ending up providing “SLOs as Code” to developers integrated into their development and delivery process.

If you have any further questions feel free to reach out to me via LinkedIn or Twitter or get in touch with one of my colleagues at Dynatrace whether this is your local Dynatrace account team, your CSM, Dynatrace One, or the Dynatrace community.

To learn more about how Dynatrace does SLOs, check out the on-demand performance clinic, Getting started with SLOs in Dynatrace.

The post How to start with SLOs to align Business, DevOps, and SREs appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-to-start-with-slos/feed/ 0
How to overcome the cloud observability wall https://www.dynatrace.com/news/blog/how-to-overcome-the-cloud-observability-wall/ https://www.dynatrace.com/news/blog/how-to-overcome-the-cloud-observability-wall/#respond Wed, 08 Dec 2021 22:36:46 +0000 https://www.dynatrace.com/news/?p=46664 cloud observability

As cloud environments become increasingly complex, legacy solutions can’t keep up with modern demands. As a result, companies run into the cloud complexity wall – also known as the cloud observability wall – as they struggle to manage modern applications and gain multicloud observability with outdated tools. But what exactly is this “wall,” and what […]

The post How to overcome the cloud observability wall appeared first on Dynatrace news.

]]>
cloud observability

As cloud environments become increasingly complex, legacy solutions can’t keep up with modern demands. As a result, companies run into the cloud complexity wall – also known as the cloud observability wall – as they struggle to manage modern applications and gain multicloud observability with outdated tools.

But what exactly is this “wall,” and what are the big-picture implications for your organization? Let’s explore this concept as we look at the best practices and solutions you should keep in mind to overcome the wall and keep up with today’s fast-paced and intricate cloud landscape.

What is the cloud observability wall?

Cloud applications are different from traditional monolithic applications – they are both ephemeral and dynamic. At any given time, the state of your application is undergoing rapid, automated changes in response to the environment. You may be using serverless functions like AWS Lambda, Azure Functions, or Google Cloud Functions, or a container management service, such as Kubernetes. Either way, you are spinning new resources up or down in response to the load, and a continuous delivery pipeline may be promoting a set of functions or containers to production.

These rapid changes — as well as the increasing volume and variety of data created — require a new approach to observability. Many customers try to use traditional tools to monitor and observe modern software stacks, but they struggle to deal with the dynamic and changing nature of cloud environments. As a result, they hit what the industry refers to as the cloud complexity wall – or cloud observability wall – and waste time, money, and resources trying to force their legacy toolsets to work in modern environments.

How observability works in a traditional environment

In contrast to modern software architecture, which uses distributed microservices, organizations historically structured their applications in a pattern known as “monolithic.” A monolithic software application has a few properties that are important to understand. Let’s break it down.

Centralized applications

Monolithic applications earned their name because their structure is a single running application, which often shares the same physical infrastructure. In a monolithic architecture, there is generally limited ability to evolve, upgrade, or enhance specific subsets of functionality without restarting or upgrading the entire application.

There are a few important details worth unpacking around monolithic observability as it relates to these qualities:

  1. The nature of a monolithic application using a single programming language can ensure all code uses the exact same logging standards, location, and internal diagnostics. Just as the code is monolithic, so is the logging.
  2. When an application runs on a single large computing element, a single operating system can monitor every aspect of the system. Modern operating systems provide capabilities to observe and report various metrics about the applications running.
  3. The last aspect is the centralization of compute. As the entire application shares the same computing environment, it collects all logs in the same location, and developers can gain insight from a single storage area.

This centralization means all aspects of the system can share underlying hardware, are generally written in the same programming language, and the operating system level monitoring and diagnostic tools can help developers understand the entire state of the system.

Dynamic applications with ephemeral services

Modern cloud-native architectures leverage a completely different development paradigm compared to monolithic applications. The core of a microservice design pattern aims to make each discrete subset of system functionality into its own self-contained unit, known as a microservice. Each microservice, running as a discrete, completely self-contained, stateless application, runs inside a container or serverless function that shares no underlying operating system with any other microservice.

The components of partitioned applications generally communicate over a network call. This boundary is language agnostic, which means the service is compatible with all other microservices written in any language, so long as the network interfaces remain the same. As it relates to observability, logging practices don’t require singular technical enforcement because services don’t share code across microservice boundaries.

Another aspect of microservices is how the service itself relates to the underlying hardware. Serverless functions typically run on hyperscale clouds and so there’s no hardware to manage. Containers and container managers, such as Kubernetes, allow the hardware to be abstracted away from the application. In both cases, microservices are in a constant, ephemeral state of transition, scaling up and down in response to the environment. Because the state of the system is always in flux, there is no centralized log location that one can readily use to gain observability.

But it’s important to note, it’s not the case that traditional observability tools are bad; they’re simply not the right tools for multicloud observability.

Observability challenges of multicloud environments

Multicloud microservices-based environments bring a new set of challenges, many of which are around the velocity and volume of data generated. In many cloud-native deployments, there can be hundreds or thousands of containers and serverless functions running at any given time. With this volume of data, not only do traditional tools break down due to technical differences, but the sheer volume of data generated is also orders of magnitude greater.

Furthermore, the large data quantities generated, on top of the increasing need to sift through millions of events to uncover actionable patterns and unexpected discrepancies to optimize application performance and availability, is much more than a single person – or even a single legacy observability system – can manage to gain any insight from. As a result, organizations are looking to AI as the modern observability capability to address these challenges, using advanced AI algorithms to understand and make sense of all the discrete events in a cloud-native microservices environment is absolutely critical for your application.

Overcoming the cloud complexity wall with Dynatrace

Dynatrace provides a unique Software Intelligence Platform built for multicloud observability. Incorporating AI, automation, and front-end monitoring, the Dynatrace Platform provides complete end-to-end visibility for modern applications enabling companies to optimize application performance and reliability while accelerating innovation.

As organizations steadily replace monolithic applications with emerging cloud-native solutions, modern cloud observability tools will provide the vital backbone. They have become the requisite tools needed to help companies overcome the cloud complexity wall and accelerate application performance and innovation.

5 challenges to achieving observability at scale

Learn how Dynatrace can help you make sense of an increasingly changing cloud landscape.

The post How to overcome the cloud observability wall appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-to-overcome-the-cloud-observability-wall/feed/ 0
Automatic connection of logs and traces accelerates AI-driven cloud analytics https://www.dynatrace.com/news/blog/automatic-connection-of-logs-and-traces-accelerates-ai-driven-cloud-analytics/ https://www.dynatrace.com/news/blog/automatic-connection-of-logs-and-traces-accelerates-ai-driven-cloud-analytics/#respond Thu, 18 Nov 2021 12:50:13 +0000 https://www.dynatrace.com/news/?p=47229 Observability graphic

As digital transformation continues to accelerate and enterprises modernize with the adoption of cloud-native architectures, the number of interconnected components and microservices is exploding. Logs are a critical ingredient in managing and optimizing these application environments. Dynatrace now unifies log monitoring with its patented PurePath technology for distributed tracing and code-level analysis. Logs are now automatically connected to distributed traces for faster analysis and optimization of cloud-native and hybrid applications.

The post Automatic connection of logs and traces accelerates AI-driven cloud analytics appeared first on Dynatrace news.

]]>
Observability graphic

Customers expect enterprises to deliver increasingly better, faster, and more reliable digital experiences. Cloud-native observability is a prerequisite for companies that need to meet these expectations. Observability enables a holistic approach to automation and BizDevOps collaboration for the optimization of applications and business outcomes.

Logs are a crucial component in the mix that help BizDevOps teams understand the full story of what’s happening in a system. Logs include critical information that can’t be found elsewhere, like details on transactions, processes, users, and environment changes.

A key element of effectively leveraging observability is analyzing telemetry data in context. Being able to cut through the noise, with all the relevant logs at hand, dramatically reduces the time it takes to get actionable insights into the optimization and troubleshooting of workloads.

Without automation, this contextualization is hardly feasible, especially in large and dynamic environments. Modern heterogeneous stacks consist of countless interconnected and ephemeral components and microservices. Log entries related to individual transactions can be spread across multiple microservices or serverless workloads. Manual and configuration-heavy approaches to putting telemetry data into context and connecting metrics, traces, and logs simply don’t scale.

Automatically connect logs and distributed traces at scale

With PurePath® distributed tracing and analysis technology at the code level, Dynatrace already provides the deepest possible insights into every transaction. Starting with user interactions, PurePath technology automatically collects all code execution details, executed database statements, critical transaction-based metrics, and topology information end-to-end.

By unifying log analytics with PurePath tracing, Dynatrace is now able to automatically connect monitored logs with PurePath distributed traces. This provides a holistic view, advanced analytics, and AI-powered answers for cloud optimization and troubleshooting.

Automatic contextualization of log data works out-of-the-box for popular languages like Java, .NET, Node.js, Go, and PHP, as well as for NGiNX and Apache Web servers. Unlike other approaches in the market, Dynatrace allows you to apply this new functionality broadly via central activation. This automated approach avoids any manual configuration of tracers or agents and the need to restart processes.

In addition, Dynatrace offers an open-source approach to the contextualization of log entries and distributed traces as well via OpenTelemetry.

PurePath traces provide a transaction-centric view across all telemetry data

You can instantly investigate logs related to individual transactions on the new Logs tab in the PurePath view. This instantly reveals additional context.

The example below includes analysis of a payment issue. It shows that a call to the payment provider was declined because the credit card verification failed. From the call perspective, you can easily see the related entry in the code-level information to understand where in your code this specific log entry was really created.

Error trace with expanded logs Dynatrace screenshot

Seamlessly switch context and analyze individual spans, transactions, or entire workloads

Dynatrace makes it easy to view the log lines related to individual spans or a broader view that covers transactions end-to-end or even entire workloads.

Uniquely, Dynatrace also provides connections to the processes that handle each call.

Log viewer Dynatrace screenshot

In the screenshot above, you can see that a single trace created 64 different log entries. The top entry, marked with status ERROR is critical to the analysis of this issue. 

This seamless user journey is also available from the log viewer side. You can easily get from individual log lines to a transaction-centric view for additional context and analysis.

How to get started

Starting with OneAgent version 1.231, you can activate our OneAgent code modules for Java, .Net, Go, Node.js, PHP, NGiNX, or Apache Web server to automatically enrich logs with trace context without any manual code or configuration change on your workload. Just go to Settings > Server-side service monitoring > Deep Monitoring > New OneAgent features. This ensures that trace IDs are automatically added to log lines for transaction-based analytics.

Structured log entries that are ingested via the Generic log ingestion of Log Monitoring V2 will show up in related PurePath traces starting with Dynatrace version 1.232.

Within the next 90 days, all transaction-related logs will show up in PurePath view after activation of this new functionality.

To find out more, see:

Log Monitoring Classic

Log Monitoring Classic is available for SaaS and managed deployments. For the latest Dynatrace log monitoring offering on SaaS, upgrade to Log Management and Analytics.

New to Dynatrace?

If so, start your free trial today!

The post Automatic connection of logs and traces accelerates AI-driven cloud analytics appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automatic-connection-of-logs-and-traces-accelerates-ai-driven-cloud-analytics/feed/ 0
How to automate version aware distributed trace analysis https://www.dynatrace.com/news/blog/how-to-automate-version-aware-distributed-trace-analysis/ https://www.dynatrace.com/news/blog/how-to-automate-version-aware-distributed-trace-analysis/#respond Mon, 13 Sep 2021 09:17:52 +0000 https://www.dynatrace.com/news/?p=46075 Version Aware Diagnostics Dashboard

Distributed Traces are “the source of truth” for developers and architects as they capture the true end-to-end execution path for each individual request processed by your applications and services. When running unit or API tests in a local dev environment, IDE (Integrated Development Environment) and test tool integrations with observability platforms make it easy for […]

The post How to automate version aware distributed trace analysis appeared first on Dynatrace news.

]]>
Version Aware Diagnostics Dashboard

Distributed Traces are “the source of truth” for developers and architects as they capture the true end-to-end execution path for each individual request processed by your applications and services. When running unit or API tests in a local dev environment, IDE (Integrated Development Environment) and test tool integrations with observability platforms make it easy for engineers to run distributed trace analysis on the distributed traces they just generated. Answering questions like “Does my code correctly call the new backend service version for my specific use case?” becomes easy as the outgoing call with request and response details will be shown in the distributed trace.

The following screenshot shows a distributed trace in Dynatrace with detailed version information of every service call involved. In the rest of the blog, you will learn more about how to capture version information and use it to answer your version-specific questions:

Dynatrace gives developers full version and request details of every service involved in end-2-end distributed trace
Dynatrace gives developers full version and requests details of every service involved in end-2-end distributed trace

If you’re not already capturing distributed traces in your development environments, I hope this blog enables you to make the case for it.

From a handful to millions of distributed traces requires automation

When distributed traces are captured in shared testing or production environments, where hundreds of services/microservices are deployed in one or multiple versions, you potentially end up with millions of captured traces in a short time span. If you then figure out how to automate the analysis of those traces you can empower DevOps and SRE teams as they need answers to questions such as:

  • “Which versions of our services are currently processing our critical transactions?”
  • “How does an overloaded backend service impact the SLOs of the frontend service?”
  • “Which frontend services are responsible for the changed traffic behavior on the backend services?”
  • “Is there a different behavior between two versions of a service? If so – shall we stop the rollout into production?”

I personally keep hearing those questions more frequently these days, which is somewhat worrying. Many members in our Dynatrace community are moving towards k8s and microservices which allows DevOps and SREs to leverage zero-downtime update strategies (also known as Progressive Delivery), such as Blue/Green, Rolling Updates, Canary Deployments or Feature Flags more easily. To answer version-specific questions like the one above it’s not only necessary to capture version information on every distributed trace but you must also automate the analysis as no one can dig through millions of traces manually to end up with answers that lead to better delivery and release decisions.

The good news is that Dynatrace PurePath, our leading automated distributed trace technology for the past 15+ years, is version aware by default meaning that it automatically captures the version information on every PurePath and provides automated analysis options to answer those version specific questions. But it’s not just the raw data that matters – it’s what Dynatrace does with the data, which I’ll delve into now. , which I’ll delve into now.

To learn more about Dynatrace’s version aware analysis capabilities I invited Thomas Rothschaedl, Product Manager at Dynatrace, to my latest Performance Clinic where he explained how version aware PurePaths are captured, how Dynatrace provides real-time release overview, and how Dynatrace provides automated answers to DevOps & SRE based on version-aware PurePath data.

While I encourage you to watch the full 30 minutes recording on YouTube or Dynatrace University I captured the key learnings in the remainder of this blog:

Dynatrace real-time release and version overview

Thomas started of reminded me about the Dynatrace Releases screen, which gives our users a live overview of all deployed releases in every monitored environment, even providing release lifecycle events (deployment, tests, quality gate, promote, rollouts, rollbacks, problems, etc.) as well as direct access to any open development or support tickets:

Dynatrace provides a real-time version-aware release overview answering critical questions for DevOps, SREs and Release Managers
Dynatrace provides a real-time version-aware release overview answering critical questions for DevOps, SREs, and Release Managers

If you’d like to learn more about the releases overview make sure to watch my Performance Clinic on Risk-Free Delivery with Dynatrace Cloud Automation Release Management.

Analyzing rolling version updates through multi-dimensional analysis

At Dynatrace we’re proud to use Dynatrace on Dynatrace which allowed Thomas to show an internal example of analyzing rolling software updates. The screenshot below shows a 72-hour analysis window of the rolling updates we do in our end-to-end testing environment. Here, we continuously roll out the latest Dynatrace versions that come out of our build system. The environment is constantly under load and it’s therefore great to see how the rolling update is truly and smoothly updating from one version to the next:

Dynatrace version-aware PurePath analysis automates the validation of successful rolling updates or canary deployments
Dynatrace version-aware PurePath analysis automates the validation of successful rolling updates or canary deployments

Analyzing request count (=throughput) is just one option, as you can see in the next section.

Automatic regression detection across versions

While the above example nicely shows the rolling updates that happened during constant load on the system, it also only focuses on throughput. What’s more interesting is if you switch to a different metric that can be extracted from PurePaths such as “Number of Exceptions Thrown”, “Number of Database Calls Made”, “Number of Database Rows Fetched”, or “Time Spent in I/O”.

Thomas demonstrated this in the performance clinic I mentioned above, where he wanted to know if any of our builds introduced a regression that caused more exceptions to be thrown. Exceptions are by default captured by Dynatrace OneAgent as part of the version aware PurePath. As you can see from the below screenshot, he immediately found a regression that was introduced in one of the builds that were rolled out earlier that day:

Automatically detect regressions introduced with a particular version by focusing on different metrics extracted from version-aware PurePaths
Automatically detect regressions introduced with a particular version by focusing on different metrics extracted from version-aware PurePaths

This is the true power of analyzing a massive amount of PurePaths in an automated way. No one would be able to identify those problems quickly by manually digging through millions of PurePaths. That’s why Dynatrace’s automation is valued by our users as it finds these issues automatically without any manual effort. But there’s more than what Thomas showed us.

Diagnostics, metrics, dashboards, and alerting

In the 30 minutes I had, Thomas walked me through the use cases he additionally demonstrated how to:

  • Drill to the offending line of code of the exception regression
  • Create metrics, put them on a dashboard and roll those out across all teams
  • Get alerted on version-specific anomalies
Dynatrace dashboard including host, service and version specific request metrics. Easy to share between teams
Dynatrace dashboard including host, service, and version-specific request metrics. Easy to share between teams

If you want to see all these demos, then check out the Performance Clinic recording. The live demo piece starts at the 15:35 timestamp.

Make better release decisions through version aware distributed traces

Whether you use Dynatrace or any other tool to capture distributed traces, it should be clear that you must make sure to capture version information on each trace and have a way to analyze large volumes of distributed traces to make better release and deployment decisions. If you don’t yet have a distributed tracing option, or if your current tooling doesn’t support what Thomas has shown in his demo, then feel free to sign up for a Dynatrace trial and try it out yourself.

To end this, I’d like to say THANK YOU Thomas for your great preparation of the Performance Clinic content. You did an amazing job in showing the value of the latest capabilities in Dynatrace. I also want to say THANK YOU to Dynatrace engineering, which has not only built a great platform but also uses Dynatrace on Dynatrace and with that, makes our demos and storytelling even easier.

The post How to automate version aware distributed trace analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-to-automate-version-aware-distributed-trace-analysis/feed/ 0
How an AIOps platform can shift left–and why it should https://www.dynatrace.com/news/blog/how-an-aiops-platform-can-shift-left/ https://www.dynatrace.com/news/blog/how-an-aiops-platform-can-shift-left/#respond Tue, 24 Aug 2021 07:51:28 +0000 https://www.dynatrace.com/news/?p=45872 Benefits of AIOps transform business operations.

As organizations layer more technologies into their DevOps toolchains, an observability-based AIOps platform that can shift left is a good strategy.

The post How an AIOps platform can shift left–and why it should appeared first on Dynatrace news.

]]>
Benefits of AIOps transform business operations.

Since the term artificial intelligence for IT operations (AIOps) was coined by Gartner in 2016, organizations have considered it a good strategy to adopt an AIOps platform or AIOps tools. These tools can help manage and automate anomaly detection and incident response for IT operations in production environments.

But as organizations adopt CI/CD practices and layer in a growing array of cloud-native solutions and open-source technologies to their DevOps toolchains, it’s becoming clear that AIOps can unlock value along the entire digital value chain.

Data in modern cloud-native environments is a continuum—from software development through service delivery all the way to customer interactions. Everything that happens provides telemetry to help teams discover root causes, inform decisions, and automate processes.

In this increasingly integrated landscape, organizations benefit not just from an AIOps platform, but an all-in-one observability, deterministic AI, and analytics platform—a software intelligence platform—to continuously automate, analyze, predict, and remediate IT issues while navigating their digital acceleration journey.

AI applied to cloud ops

As the size and complexity of distributed cloud computing systems continue to grow, typical AIOps approaches to monitoring, diagnosing, and repairing software are not scaling. Traditional AIOps relies on correlating data in order to reduce alerts, which is slow and inaccurate and does little to identify root causes.

An intelligent AIOps platform with end-to-end observability can leverage AI-based algorithms and real-time data analytics to automate triaging, response, and remediation for common IT issues, including unexpected downtime, system latency, or determining why a Kubernetes pod was terminated.

An integrated platform that includes AIOps, observability, and analytics can consume and analyze the increasing volume of cloud data to automate and optimize these routine monitoring and management tasks. Such an observability-based AIOps platform can also provide advanced insight for IT and DevOps, while reducing mean time to resolution (MTTR) and speeding up mean time to discovery (MTTD). With end-to-end visibility into multicloud environments, an intelligent AIOps platform with advanced analytics enables faster innovation, higher quality, more efficiency, and ultimately, better business outcomes.

In addition to driving enterprise automation, there are six key capabilities an all-in-one AIOps platform approach delivers:

  1. Alert management
    • Replaces monitoring tool alert storms with accurate, reliable root-cause analysis.
    • Eliminates up to 90% of false alarms and reduces noise with deterministic AI fault tree analysis.
    • Observes, analyzes, and enables automated response in near-real time.
  2. Automation
    • Contextualizes and processes large volumes of operational data.
    • Uses this high-fidelity, context-rich collected data to create real-time topology and service flow maps.
    • Provides analysis and AI-powered insights across the application lifecycle.
    • Continuously discovers changes to environments, apps, and services.
  3. Incident prioritization and routing
    • Delivers relevant insights to the right people at the right time, providing precise answers with root-cause determination, prioritized by business impact.
  4. Event causation
    • Uses causation-based AI to point directly to the root cause and impact of a failed test run, an application slowdown, or system outage, or to drive decisions about whether to release a piece of software.
  5. Predictive analytics
    • Monitors the entire technology stack end-to-end to predict and prevent future disruptions before they occur.
  6. Auto-remediation
    • Automates anomaly detection, problem notification, and self-remediation with full-stack monitoring and integration with workflow automation platforms, such as ServiceNow and Jenkins.

Why AIOps needs to “shift-left”

With the volume of data increasing, and the demand for services rocketing upward, the need for AIOps is no longer limited to IT operations. As DevOps and SRE practices mature, pre-production workflows need AIOps capabilities just as acutely.

A typical continuous integration/continuous delivery (CI/CD) pipeline follows the following sequence:

  1. Source — creating source code
  2. Build — compiling the application
  3. Test — testing code for functionality
  4. Release — pushing code to the repository
  5. Deploy — moving code to production

The term “shift-left” refers to the practice of performing a task at an earlier stage of development before it goes to production, such as automated testing at the source phase instead of when code is ready to be released.

Shift-left applied to AIOps integrates AI into the full DevOps lifecycle, including data ingestion, building code, and testing for enhanced software quality and deeper root-cause analysis before code is deployed to production. A software intelligence platform that includes end-to-end observability in its approach to AIOps delivers continuous alert and incident management, automatically observes and identifies anomalies in CI/CD pipelines, and prevents issues from reaching the production stage, resulting in more efficient builds and quicker, higher-quality releases of new versions of software.

Shifting AIOps left means development teams can easily leverage production service-level objectives (SLOs) as criteria for building quality gates into acceptance testing earlier in the development cycle. It also means teams can initiate auto-remediation for CI/CD workflows by integrating with software configuration and deployment management technologies, such as Chef, Puppet, and Ansible.

Faster time-to-value with an intelligent AIOps platform

An AIOps platform based on continuous discovery and end-to-end observability can detect anomalies before they affect the CI/CD pipeline or impact customer experience and SLOs. It can automate validation processes by using SLO-based quality gates with events, tags, and APIs integrating seamlessly with existing CI/CD workflows — for faster automated deployments and time-to-value.

AI-enhanced alerting and escalation using advanced algorithms based on deterministic fault-tree analysis can automatically route incidents to the appropriate team, empowering them with metadata and context that results in accurate, reliable, and precise root-cause analysis. If automated processes are unable to address a slowdown or outage, DevOps is then given a clear path to remediation, eliminating time spent on “problem triage,” which drives faster innovation and better quality.

A shift-left AIOPs platform approach mitigates the cost of IT downtime

Every CIO and CFO knows IT downtime is costly — potentially adding up to thousands of dollars per minute or more depending on the organization’s size and reach of services. But what may be as significant is the impact downtime and performance issues can have on already pushed-to-the-limit DevOps and IT teams. They are under increasing pressure to maintain system reliability and prevent outages of highly complex, distributed multicloud operations. The constant context switching required to hop from issue to issue also derails focus on mission-critical concerns and distracts teams from their core functions.

By shifting AIOps left using production-based performance criteria as a quality gate earlier in the development cycle, teams can release more resilient software. If an issue arises anywhere in the DevOps workflow, an integrated AIOps platform with all-in-one observability, deterministic AI, and analytics can automatically remediate issues and optimize performance based on system health and user demands — preventing or minimizing the duration of outages and reducing costs. It also combats IT teams’ top stressors by reducing the need for manual intervention in IT operations and DevOps workflows, helping to prevent burnout and costly employee churn.

Shift AIOps left with an integrated software intelligence platform

Teams undergoing digital transformation are discovering that shifting AIOps left into DevOps and SRE workflows can accelerate and increase the effectiveness of their DevOps and SRE initiatives.

To help developers pinpoint problems and automate more processes during the development, test, and delivery phases of DevOps, Dynatrace seamlessly integrates AIOps into the CI/CD pipeline, bringing fault-tree analysis to pre-production workflows. Shifting AIOps left enables developers to discover and auto-remediate issues in pre-production so they can optimize processes and deliver higher quality code to production.

Using the Cloud Automation control plane—powered by Keptn, an open-source technology for cloud-native application life-cycle orchestration—Dynatrace provides release analysis, version awareness, and SLO-based quality gates so teams can automate releases at all stages of the DevOps pipeline. By integrating with DevOps tools like Chef, Puppet and Ansible, Dynatrace can execute closed-loop remediation workflows or orchestrate ITSM tools to trigger incident management workflows.

To learn more about how Dynatrace approaches AIOps, see the eBook: AIOps Done Right.

Dynatrace was also named a leader in AIOps in the Forrester Wave. Read the report here.

Developing an AIOps strategy for cloud observability

Download our free eBook to learn the best practices for developing an AIOps strategy that drives efficiency, innovation, and better business outcomes.

The post How an AIOps platform can shift left–and why it should appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-an-aiops-platform-can-shift-left/feed/ 0
A three-step implementation guide to answer-driven SLO-based release validation https://www.dynatrace.com/news/blog/a-three-step-implementation-guide-to-answer-driven-slo-based-release-validation/ https://www.dynatrace.com/news/blog/a-three-step-implementation-guide-to-answer-driven-slo-based-release-validation/#respond Fri, 16 Jul 2021 14:53:27 +0000 https://www.dynatrace.com/news/?p=45314 SLO-based Quality Gates is the latest capability we added to the Dynatrace Cloud Automation solution proudly powered by Keptn

The Dynatrace Software Intelligence Platform already comes with release analysis, version awareness, and Service Level Objective (SLO) support as part of the Dynatrace Cloud Automation solution, helping DevOps and SRE teams automate the delivery and operational decisions. This week my colleague Michael Winkler announced the general availability of Cloud Automation quality gates, a new capability […]

The post A three-step implementation guide to answer-driven SLO-based release validation appeared first on Dynatrace news.

]]>
SLO-based Quality Gates is the latest capability we added to the Dynatrace Cloud Automation solution proudly powered by Keptn

The Dynatrace Software Intelligence Platform already comes with release analysis, version awareness, and Service Level Objective (SLO) support as part of the Dynatrace Cloud Automation solution, helping DevOps and SRE teams automate the delivery and operational decisions. This week my colleague Michael Winkler announced the general availability of Cloud Automation quality gates, a new capability that aims to provide answer-driven release validation as part of your delivery process.

In this blog, I want to guide you through the steps of implementing release quality gates in your delivery automation or as we often call it “Shift-Left SLOs with Quality Gates”. We have seen users who joined our preview program “speed up their release validation by 90%”. And it’s not just the release validation. Many of our users are performance engineers using Cloud Automation Quality Gates to automate the analysis of their performance and load tests – saving hours of analysis time for each test they run. Don’t just take my word for it, hear from Mike Kobush, Performance Engineer at NAIC, who gave us a tour of his performance test automation using Dynatrace Cloud Automation.

The innovative technology powering these recent and future use cases comes from Keptn, our recognized CNCF open-source project. Here is a good overview of how Dynatrace users are now benefiting from this stream of innovation:

SLO-based Quality Gates is the latest capability we added to the Dynatrace Cloud Automation solution proudly powered by Keptn
SLO-based Quality Gates is the latest capability we added to the Dynatrace Cloud Automation solution proudly powered by Keptn

If you want more information on Keptn, then I suggest joining the Keptn Open Source community and help us drive innovation that benefits both the open-source project and Dynatrace Cloud Automation Solution. For an overview of Keptn, I can also recommend watching our cdCON talk Behind the Scenes of Keptn: Event-Driven Delivery & Ops Orchestration.

Ready to create your first release validation automation? Whether it’s used to automate your test analysis or fully integrated into your delivery pipeline, this blog will guide you through it!

Let’s Start: Answer driven automation – your full video tutorial

If this is blog is TLDR (Too Long Didn’t Read) I have good news for you. Simply watch the following video tutorial on YouTube where I cover the following topics:

  • 0:00 – Intro to Dynatrace Cloud Automation Use Cases
  • 01:19 – Introducing Shift-Left SLO Quality Gates
  • 03:24 – Pre-Requisites
  • 05:00 – Quick Start – Connect Cloud Automation
  • 07:58 – Quick Start – Automatic Quality Gates
  • 18:48 – GitOps – Codify Quality Gates
  • 29:10 – GitOps – Dashboard based Automation
  • 37:13 – Integrate with Delivery
  • 40:08 – Expand to more use cases



Answer driven automation tutorial

Watch my tutorial and learn everything from getting started via GitOps to integrate with your CI/CD

As the video alone shows you every step in detail, including live demos, I will just give you a high-level overview and the outcomes of the individual sections:

Pre-requisite: Cloud Automation SaaS Tenant

If you’re a Dynatrace user and want to leverage automated SLO-based quality gates, then you need to request a Cloud Automation SaaS Tenant through your Dynatrace representative. This could be your account team, Dynatrace One, or support.

While I think every user will benefit from this new feature make sure you meet the following criteria to really reap the benefits:

  1. Dynatrace requirements
  2. You have automated tests as part of delivery, monitored by Dynatrace, and you want to automatically validate to speed up your delivery pipeline (Lead Time)
  3. You run load tests monitored with Dynatrace and you want to automatically validate to eliminate the manual analysis effort

Quick Start: Answer-driven quality gates in 5 minutes

The first goal is to get a default automated quality gate for one of your services monitored with Dynatrace. It should be a service you regularly run tests against so we can automate the analysis of the data captured during the test runs with Cloud Automation Quality Gates.

It only takes two steps, as you can see in the QuickStart section of my video tutorial:

  1. Connect your Cloud Automation tenant with your Dynatrace environment
  2. Tag a service with keptn_managed and keptn_service:<servicename>
Quick Start: All it takes is connect and tag your services. This enables automatic quality gates
Quick Start: All it takes is to connect and tag your services. This enables automatic quality gates

In the quick start, we also talk about how to trigger an evaluation of a timeframe. This can either be done through the Keptn CLI (Command Line Interface) or through the Keptn API. The screenshot below demonstrates how to trigger an evaluation for 30 minutes, adding additional metadata to the evaluation and how it then looks like in the Cloud Automation UI:

Evaluations can be triggered through the CLI or API and therefore easy to integrate with your other DevOps tooling
Evaluations can be triggered through the CLI or API and therefore easy to integrate with your other DevOps tooling

GitOps: Cloud automation as code

To achieve true automation, we can’t rely on configurations done manually through UIs. The goal is to store and manage all configurations in a version control system such as Git. That’s what Cloud Automation does: SLIs, SLOs, Monitoring Definition, Automation Sequences, Tests, Deployment Definitions, etc. are stored in a Git repository that Cloud Automation holds internally by default.

To harness the power of this Git repo, we can connect it to a Git repo residing in your Git service of choice, e.g., GitHub, GitLab, Azure, Bitbucket, and more. This enables you and your development teams to modify configuration files used by Cloud Automation through their regular way of work, for example adding new SLOs through a Pull Request process.

For a more detailed explanation on how to set up an upstream, as well as how to then modify SLIs or SLOs and see them reflected in your next evaluation run as shown in the image below, give the GitOps section on Codify Quality Gates of my tutorial a watch.

GitOps: integrate Cloud Automation configuration into your development processes through the upstream git capability
GitOps: integrate Cloud Automation configuration into your development processes through the upstream git capability

A very popular way to define quality gates is through Dynatrace dashboards. My video tutorial demonstrates the steps of using SLO dashboards, which ultimately are generating SLI and SLO YAML files that end up in the Git repository. This means that even this approach is fully GitOps compliant!

Dashboard based Automation: easy to manage and fully GitOps compliant
Dashboard based Automation: easy to manage and fully GitOps compliant

If you want to learn more about the dashboard approach make sure to check out my YouTube Tutorial on Building SLO-based Quality Gates in 5 Minutes.

Integrate with delivery

Knowing how to trigger Dynatrace Cloud Automation Quality Gates through the CLI and the API makes us ready to integrate it with your existing DevOps delivery tools such as Jenkins, GitLab, Harness, Tekton, Azure DevOps, Bamboo, or others. The integration allows you to bring automated release validation into your existing pipelines – using the result of Cloud Automation Quality Gates to make automated decisions on whether to stop or continue your delivery process.

Watch the Integrate with Jenkins and other CI/CD tools section of my YouTube tutorial to learn more about how you can integrate this in your environment:

Integrate: Use existing libraries & extensions or simply call the API to integrate Cloud Automation with your DevOps tools
Integrate: Use existing libraries & extensions or simply call the API to integrate Cloud Automation with your DevOps tools

If you need help integrating this with your tools, feel free to reach out to your Dynatrace team or get in contact with your Autonomous Cloud Enablement Practice team which has helped many of our customers bring cloud automation into their DevOps tools.

Expand to more use cases

Automated release validation through quality gates is just the start. But these use cases are to be enabled by Cloud Automation in the not-so-distant future:

  • Monitoring as Code (Dynatrace Monaco)
    • Automatically configure Dynatrace, e.g., dashboards, management zones, metrics
  • Performance as a Self-Service (JMeter, Neotys, Locust,)
    • Automatically execute and analyze performance tests based on SLOs
  • Progressive Delivery (Helm, Argo, Tekton, Harness, GitLab …)
    • Automate multi-stage delivery with canaries and feature flags
  • Auto Remediation (Ansible, ServiceNow, Generic Executor, …)
    • Automatically respond to problem tickets by executing remediation actions

I cover some of the details in the Expand section of my YouTube tutorial:

Expand: Quality gates is just the start. Bring the power of automation into your environment
Expand: Quality gates are just the start. Bring the power of automation into your environment

If you have specific use cases or want to learn more about what’s possible and coming up, make sure to contact your Dynatrace team.

Friends don’t let friends build their own cloud automation

It’s as simple as the headline: “Friends don’t let friends build their own cloud automation!”.

Your challenge as a DevOps or SRE engineer in the coming months will be to implement many of the use cases listed under our Cloud Automation solution. We have already brought you, release analysis, version awareness, and SLO. Now we bring you Quality Gates to automate release validation and coming up are those relevant expansions as we move towards a more cloud-native delivery and operations model; automating performance and resiliency engineering, progressive delivery, and auto-remediation. Stay tuned, stay connected, stay healthy!

The post A three-step implementation guide to answer-driven SLO-based release validation appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/a-three-step-implementation-guide-to-answer-driven-slo-based-release-validation/feed/ 0
Dynatrace AI predicts SLO violations and pinpoints root causes proactively https://www.dynatrace.com/news/blog/dynatrace-ai-predicts-slo-violations-and-pinpoints-root-causes-proactively/ https://www.dynatrace.com/news/blog/dynatrace-ai-predicts-slo-violations-and-pinpoints-root-causes-proactively/#respond Mon, 28 Jun 2021 23:07:44 +0000 https://www.dynatrace.com/news/?p=45137 SLOs

Dynatrace enables Site Reliability Engineering (SRE) teams to proactively ensure the highest service quality levels. Davis, the Dynatrace AI engine, identifies potential contributors to SLO violations in real time, before thresholds are breached. Dynatrace pinpoints the root causes of problems and their impact on SLOs.

The post Dynatrace AI predicts SLO violations and pinpoints root causes proactively appeared first on Dynatrace news.

]]>
SLOs

In modern service landscapes, Service-level-objectives (SLOs) are the chosen methodology of Site Reliability Engineering (SRE) teams for ensuring the high quality of delivery of their digital services. There is however a major challenge faced by many SRE teams: how to catch relevant degradations early, before a long-term SLO shows an unhealthy state.

Are you still “reacting to bad numbers”?

SLOs with an observation period of, for example, one week, are of course not overly affected by short-lived outliers. However, such observation periods come with a disadvantage: incidents can pile up and there is a delay between those incidents and the corresponding health metrics ultimately dropping low enough to trigger a warning.

Teams who are primarily reactive in their approach therefore use SLOs to decide when the state of a system has become so bad that it requires intervention.

Some SRE teams counter this by defining the same SLOs for different observation periods to reduce the reaction times in case of incidents. Many teams use three different levels of observation periods, one for strategic decisions, one for tactical decisions, and one short period for catching incidents. These redundancies can of course create additional efforts and complexity.

Error budgets and the tracking of their burn rates offer a much better approach, however without extensive manual effort, this approach still leaves two questions open:

  • How can I detect anomalies early, before they impact your SLOs?
  • To facilitate fast remediation, how can I quickly identify the root causes of emerging issues that have massive potential SLO impact

Most monitoring tools offer only a single SLO metric. However, watching a single SLO health metric and error budget drop doesn’t provide much in the way of answers; it only confirms the obvious—that your SLO is unhealthy. In the best case scenario, to answer the above questions, you need experts to conduct manual investigation and interpret the data for you. In the worst case scenario, the nature of today’s dynamic and heterogenous environments renders such manual investigation impossible.

Dynatrace proactively helps Site Reliability Engineers keep their SLOs healthy

Dynatrace Davis, our AI-engine, offers a unique feature that overcomes the fundamental challenge of reacting quickly enough, even within strategic observation periods. Davis notifies you when any of your SLOs are at risk, before any metrics turn red.

This works out-of-the-box because Dynatrace understands how all your application and infrastructure components depend on each other. In this way, Davis can link defined SLOs to those anomalies that present potential negative impact.

Davis AI predicts if future SLO health is at risk

Let’s look at an example where an SLO was defined for the stability of a frontend service that shows a perfect 100% SLO health status:

Dynatrace screenshot SLO status

Notice in the above SLO tile that Davis has displayed a critical error indicator to inform the SRE team about an ongoing incident within the service topology that the SLO covers. Even though the SLO still shows perfect 100% health, Davis AI is proactively predicting that the future SLO health is at risk.

Dynatrace AI pinpoints the root causes of SLO-impacting incidents

Further, a single click on this tile displays all active incidents along with the potential negative impact on the future health of the SLO.

Dynatrace screenshot What's the root-cause

A drill-down from an unhealthy SLO takes you to a filtered feed of detected problems that are the root causes of these incidents. This precise AI-assisted identification of root causes saves valuable time for SRE and DevOps teams during critical service outages, instead of just showing a single, isolated health metric.

Get up and running in under a minute with SLO templates

Service-level-objectives consist of carefully selected Service-level-indicators (SLIs) which provide a quantitative measure of each aspect of the service level. Typically, an SRE team spends a good amount of time selecting the best indictor metrics for their given services, which then leads to well-defined SLOs that reflect the service quality.

The Google Site Reliability Engineering page is a great read for understanding and embracing the idea of defining SLOs for reliable global IT services.

However, getting started with SLOs in Dynatrace is even easier.

We offer a collection of best-practice SLO definitions for various use cases beyond the observability domain; simply choose one of the predefined SLO templates that Dynatrace provides out-of-the-box.

For example, you can measure the quality of service of your mobile app offering. Dynatrace offers a mobile crash-free users SLO template that you can use to create a best-practice SLO for measuring the reliability and stability of your mobile apps.

Dynatrace screenshot Add new SLO

Once you’ve defined your business-critical SLOs, you’re all set. Davis will then automatically analyze your SLOs continuously and provide a truly proactive approach to SLOs.

Get started with SLOs

Proactive SLO impact analysis is available with the release of Dynatrace version 1.220. If you’re new to Dynatrace, you can experience the magic yourself by starting a Dynatrace free trial.

Learn more

If you want to learn more about SLIs/SLOs, here are a few resources that we recommend:

You can check out the session recording to find all the details.

The post Dynatrace AI predicts SLO violations and pinpoints root causes proactively appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-ai-predicts-slo-violations-and-pinpoints-root-causes-proactively/feed/ 0