Artificial intelligence | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 11 Jun 2026 11:43:04 +0000 en hourly 1 Causal AI use cases for modern observability that can transform any business https://www.dynatrace.com/news/blog/causal-ai-use-cases-for-modern-observability/ https://www.dynatrace.com/news/blog/causal-ai-use-cases-for-modern-observability/#respond Mon, 22 Jan 2024 18:56:22 +0000 https://www.dynatrace.com/news/?p=61642 Causal AI use cases for modern observability; exploratory data analytics

Artificial intelligence adoption is on the rise everywhere—throughout industries and in businesses of all sizes. And while generative AI was much hyped in 2023, the deterministic nature of causal AI—which determines the precise root cause of an issue—is a key foundational requirement to get reliable decisions and recommendations from generative AI technologies. Further, not every […]

The post Causal AI use cases for modern observability that can transform any business appeared first on Dynatrace news.

]]>
Causal AI use cases for modern observability; exploratory data analytics

Artificial intelligence adoption is on the rise everywhere—throughout industries and in businesses of all sizes. And while generative AI was much hyped in 2023, the deterministic nature of causal AI—which determines the precise root cause of an issue—is a key foundational requirement to get reliable decisions and recommendations from generative AI technologies.

Further, not every business uses AI in the same way or for the same reasons. So, it’s important for organizations to choose the AI type that best meets their needs. While predictive AI relies on machine learning algorithms that find correlations in data, causal AI aims to determine the precise underlying mechanisms that drive events and outcomes. As a result, causal AI use cases are key to enabling organizations to identify the root cause of problems and determine remediation.

Making the case for causal AI

Most AI today uses machine learning models like neural networks that find correlations and make predictions based on them. However, correlation does not imply causation. So, these models are limited in their ability to explain why outputs occurred or to make reliable decisions in new situations. They’re essentially informed guesses or likelihoods of outcomes. The growing recognition of these limitations is driving increased interest and research into causal AI use cases.

Causal AI use cases for modern observability

Integrating causal AI into observability systems can significantly advance organizations’ understanding of their environments. Traditional monitoring tools can alert organizations to issues, but causal AI can precisely identify the root cause of operational and quality issues. This facilitates quicker and more effective problem solving, reducing downtime and improving reliability through intelligent automation.

More generally, causal AI can contribute to explainable and fair AI systems. That’s important as regulatory scrutiny and demands for responsible AI are growing. According to a recent Dynatrace survey of 1,300 CIOs, CTOs, and other senior technology leaders, 98% of technology leaders are concerned that generative AI could be susceptible to unintentional bias, error, and misinformation. AI systems’ ability to explain the reasons for their recommendations grounded in causal AI could go a long way in resolving these trust issues.

Take causal AI to the next level with a composite approach

The benefits of causal AI are obvious, as it determines the exact underlying causes and effects of a digital system’s events or behaviors based on the system’s topology. The same cannot be said for predictive AI, which makes predictions about future events based on data patterns, and generative AI, which uses training data to create content that reflects its users’ natural language queries. But nothing is perfect, and each AI type has specific capabilities and limitations.

That’s where Dynatrace can help. Dynatrace takes a composite approach, called hypermodal AI, which combines causal, predictive, and generative AI to drive fast, precise, and trustworthy answers and automation.

This hypermodal approach features the following:

Automated root-cause analysis. Dynatrace automated root-cause analysis uses causal AI to rapidly pinpoint the source issues behind user experience, application, and infrastructure performance problems before they result in outages. Through dependency mapping, Dynatrace causal AI can contextualize and explain incident alerts, saving teams substantial time compared with manual troubleshooting across complex, modern IT environments.

Automated root-cause analysis with Davis CoPilot

Intelligent alert prioritization. By determining the likely business effects of service issues using causal AI, Dynatrace automatically prioritizes issues and alerts the relevant teams while allowing auto-remediation on routine alerts. This reduces alert fatigue and speeds up the restoration of critical systems through auto-remediation.

Failure prediction. Dynatrace uses causality graphs and analysis of the sequence of events to determine how chains of dependent application or infrastructure events will potentially lead to slowdowns, failures, and outages. By predicting failure risk, Dynatrace enables pre-emptive changes such as resource autoscaling, traffic shifting, or preventative rollbacks of bad code deployment ahead of time.

Forecasting with Davis CoPilot

Resource optimization. Dynatrace uses the causal relationships between events across user experience, application, and infrastructure layers and ties them to business KPIs to optimize dynamic policy decisions for cloud resource or container scaling and cloud cost optimization to meet performance and efficiency goals even in highly dynamic and complex environments.

Automated remediation. For well-defined remediation processes, teams can automate remediation tasks, such as server restarts, spinning up new nodes, code rollbacks, and configuration changes based on the determination of causal AI—all without manual intervention in many cases.

For more information on where AI is heading this year and why taking a composite AI approach is critical to organizational success, check out our recent research report, “The state of AI 2024.”

The post Causal AI use cases for modern observability that can transform any business appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/causal-ai-use-cases-for-modern-observability/feed/ 0
Measuring the importance of data quality to causal AI success https://www.dynatrace.com/news/blog/the-importance-of-data-quality-to-causal-ai/ https://www.dynatrace.com/news/blog/the-importance-of-data-quality-to-causal-ai/#respond Thu, 04 Jan 2024 19:12:59 +0000 https://www.dynatrace.com/news/?p=61445 Generative AI poised to have an impact by automating software development. And why AI projects fail

Causal AI can accurately pinpoint why an event occurred, but the effectiveness of AI depends on high-quality data. Discover common data quality challenges, how to improve data quality, and more.

The post Measuring the importance of data quality to causal AI success appeared first on Dynatrace news.

]]>
Generative AI poised to have an impact by automating software development. And why AI projects fail

Traditional analytics and AI systems rely on statistical models to correlate events with possible causes. While this approach can be effective if the model is trained with a large amount of data, even in the best-case scenarios, it amounts to an informed guess, rather than a certainty. That’s where causal AI can help.

Causal AI is a different approach that goes beyond event correlations to understand the underlying reasons for trends and patterns. It uses fault-tree analysis to identify the component events that cause outcomes at a higher level. Causal AI is particularly effective in observability. It removes much of the guesswork of untangling complex system issues and establishes with certainty why a problem occurred.

Causal AI applies a deterministic approach to anomaly detection and root-cause analysis that yields precise, continuous, and actionable insights in real time. But to be successful, data quality is critical. High-quality data creates the foundation for credible insights organizations can use to make sound decisions.

In what follows, we’ll discuss how to assess data quality, common data quality challenges, how to overcome them, and more.

Key considerations for assessing data quality

Assessing data quality requires organizations to consider several key factors, including the following:

Accuracy. Teams need to ensure the data is accurate and correctly represents real-world scenarios. Additionally, it’s important to consider all variables.

Completeness. Is any information missing from the data set? Omissions can create wrong conclusions and contribute to bias.

Consistency. Ensure there are no discrepancies in the data. Contradictory or inconsistent data confuses AI models and increases the risk of errors.

Timeliness. The data should be up-to-date and relevant to the current context. Timeliness is a critical factor in AI for IT operations (AIOps). Because IT systems change often, AI models trained only on historical data struggle to diagnose novel events. Causal AI requires real-time updates to the training model.

Relevancy. The data needs to be appropriate for the questions asked. In AIOps, this means providing the model with the full range of logs, events, metrics, and traces needed to understand the inner workings of a complex system.

How can organizations improve data quality?

Improving data quality is a strategic process that involves all organizational members who create and use data. It starts with implementing data governance practices, which set standards and policies for data use and management in areas such as quality, security, compliance, storage, stewardship, and integration.

Data stewardship is an increasingly important factor in data quality. It ensures the data people and departments generate and maintain is clean, consistent, and complete. Data mesh is a popular new concept that encourages the people who create data to treat it as a product to be managed like any other product. But it suffers from limitations such as multiple copies of data. High-quality operational data in a central data lakehouse that is available for instant analytics is often teams’ preferred way to get consistent and accurate answers and insights.

Data-cleaning tools and methods are needed to identify and fix errors. Additionally, teams should perform continuous audits to evaluate data against benchmarks and implement best practices for ensuring data quality.

Common data quality challenges to consider

Organizations may encounter numerous barriers to ensuring data quality. For starters, the sheer amount of data can make management daunting. Modern, cloud-native architectures have many moving parts, and identifying them all is a daunting task with human effort alone. Modern observability solutions that automatically and instantly detect all IT assets in an environment — applications, containers, services, processes, and infrastructure — can save time.

Fragmented and siloed data storage can create inconsistencies and redundancies. Stakeholders need to put aside ownership issues and agree to share information about the systems they oversee, including success factors and critical metrics.

Another common impediment is manual data tagging and handling, an error-prone process that teams should minimize. Observability solutions automate much of the task of identifying the variables that go into application performance and availability. Human involvement should be limited to verifying the features or attributes machine learning algorithms use to make predictions or decisions.

Improving data quality management using causal AI

Causal AI can be a powerful tool for improving systems management, observability, and troubleshooting. It can highlight inconsistencies or outliers in data sets that indicate anomalies and pinpoint the root causes. It also enables an AIOps approach with proactive visibility that helps companies improve operational efficiency and reduce false-positive alerts by 95%, according to a Forrester Consulting report.

Causal AI informs better data governance policies by providing insight into how to improve data quality. It improves time management and event prioritization by helping developers, administrators, and site reliability engineers identify the alerts that matter most. Identifying issues before an application or service outage occurs can reduce costs. IT teams can focus on strategic initiatives to drive business success, rather than firefighting — and it accelerates digital transformation through automation and self-maintaining systems.

Unleash the power of causal AI

Dynatrace provides an AI-powered, automated IT performance monitoring platform with advanced observability and analytics capabilities. It enables real-time health and performance tracking, intelligent anomaly detection, data quality controls, and automated issue resolution.

By accurately assessing, managing, and continuously improving data quality, organizations can use causal AI to its full potential. Platforms such as Dynatrace help ensure that data quality rises to the standard required for effective causal analysis.

Learn more about how to make the most of your immense data — and store it — with this free guide, “Data insights get an upgrade with data lakehouse architecture.”

The post Measuring the importance of data quality to causal AI success appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-importance-of-data-quality-to-causal-ai/feed/ 0
The state of AI in 2024: Overcoming adoption challenges to unlock organizational success https://www.dynatrace.com/news/blog/state-of-ai-in-2024/ https://www.dynatrace.com/news/blog/state-of-ai-in-2024/#respond Tue, 12 Dec 2023 18:45:03 +0000 https://www.dynatrace.com/news/?p=61083 How generative AI is fueling IT operations modernization

In the "State of AI" report, respondents outlined the benefits and challenges of AI. They also indicated how to overcome AI challenges with a ‘composite’ approach in which teams combine multiple types of AI to generate accurate and trustworthy answers.

The post The state of AI in 2024: Overcoming adoption challenges to unlock organizational success appeared first on Dynatrace news.

]]>
How generative AI is fueling IT operations modernization

Artificial intelligence (AI) has revolutionized the business and IT landscape. And now, it has become integral to organizations’ efforts to drive efficiency and improve productivity.

In fact, according to the recent Dynatrace survey, “The state of AI 2024,” the majority of technology leaders (83%) say AI has become mandatory. However, most organizations are still in relatively uncharted territory with their AI adoption strategies. Alongside the numerous benefits, these organizations need to manage the increased risks the technology brings.

Looking into the future of AI in 2024, the report explores these challenges and highlights how technology leaders can overcome them to drive fast, precise, and trustworthy answers and automation.

AI investment is accelerating

The report indicates that organizations are already recognizing the vast potential of AI. Their plans to increase investment in these technologies over the next 12 months show no signs of slowing. For example, nearly two-thirds (61%) of technology leaders say they will increase investment in AI over the next 12 months to speed software development.

As they continue on this path, organizations expect other benefits, from enabling business users to easily customize dashboards (54%) to building interactive queries for analytics (48%). This means AI will affect not only IT and back-office support functions but also front-line staff in customer-facing roles.

AI is essential to taming multicloud complexity

One area where organizations see significant potential for AI is in helping to reduce the complexity of their modern cloud environments. Eighty-seven percent of technology leaders say AI-powered issue prevention and remediation are critical to managing multicloud complexity.

Organizations will increase AI investment over the next 12 months to tackle this complexity by delivering predictable, trustworthy, and precise answers in real time. For example, 73% of technology leaders are investing in AI to generate insight from observability, security, and business events data.

This means greater productivity for individual teams. DevOps teams, for example, can focus on driving innovation instead of grinding through manual jobs. According to “The state of AI” report, nearly three-quarters of IT operations, development, and security teams plan to use AI to become more proactive in executing their work.

Technology leaders also expect AI to become critical to the success of core DevOps use cases, including the following:

  • threat detection, investigation, and response (82%);
  • automating complex operations tasks (63%); and
  • eliminating false alerts and the manual effort of validating code deployments (58%).

Minimizing AI risk is an urgent priority

Alongside the clear advantages of AI, the report indicates that there are challenges for AI adoption. In the wake of a significant hype cycle following the 2022 launch of ChatGPT, a chatbot based on generative AI, most technology leaders are concerned that generative AI could be susceptible to unintentional bias, error, and misinformation.

To address this, DevOps teams need to find ways to easily engineer AI prompts that contain detailed context and precision. In doing so, they can achieve meaningful, AI-generated responses that users can trust and avoid inaccurate or inconsistent statements.

But it’s not just the accuracy of AI-generated answers that’s a concern. Organizations must also be mindful of the potential security and compliance risks.

The report indicates that 95% of technology leaders are concerned that using generative AI to create code could result in data leakage as well as improper or illegal use of intellectual property.

Organizations need sufficient guardrails to manage the data that AI models ingest. Otherwise, employees could accidentally expose sensitive information. This need will drive demand for AI platforms that are purpose-built with security and privacy requirements in mind.

AI will have a widespread impact

Delving beyond the impact on IT, the report shows that AI is set to improve workforce satisfaction throughout the organization. Nontechnical workers can make informed, data-driven decisions with easier access to analytics through natural language queries and virtual assistants.

As a result, the burden on DevOps teams will ease, as the pressure to deliver on the business’ needs for data-driven insights will no longer depend solely on DevOps.

To realize these benefits, organizations must get their AI strategy right. Technology leaders can lay the groundwork for success by recognizing that not all AI is created equal. More complex use cases, such as writing software code and resolving security vulnerabilities, require a combination of AI types and different data sources, such as observability, security, and business events.

The report identifies this “composite AI” approach — where the precision of causal AI meets the forecasting capabilities of predictive AI to provide essential context for generative AI prompts — is essential for organizational success in 2024.

To take a closer look at what technology leaders around the world are saying, read more in “The state of AI 2024.”

The post The state of AI in 2024: Overcoming adoption challenges to unlock organizational success appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/state-of-ai-in-2024/feed/ 0
What is causal AI? Why this deterministic AI approach is critical to business success https://www.dynatrace.com/news/blog/what-is-causal-ai-deterministic-ai/ https://www.dynatrace.com/news/blog/what-is-causal-ai-deterministic-ai/#respond Tue, 25 Jul 2023 01:51:40 +0000 https://www.dynatrace.com/news/?p=58787 The keys to responsible AI and the importance of trusted AI

Today's organizations need to go beyond a traditional, correlation-driven approach to identify the underlying causes and effects of an event or behavior and drive better DevOps automation. Enter causal AI.

The post What is causal AI? Why this deterministic AI approach is critical to business success appeared first on Dynatrace news.

]]>
The keys to responsible AI and the importance of trusted AI

Today’s organizations need to solve increasingly complex human problems, making advancements in artificial intelligence (AI) more important than ever. Conventional data science approaches and analytics platforms can predict the correlation between an event and possible sources. But they often fall short when it comes to understanding why an event occurred. That’s where causal AI, also referred to as deterministic AI, makes a crucial difference.

In what follows, we’ll discuss causal AI, how it works, and how it compares to other types of artificial intelligence. We’ll also discuss why it’s essential for business success in the age of generative AI.

What is causal AI?

Causal AI is an artificial intelligence technique used to determine the exact underlying causes and effects of events or behaviors. Unlike correlation-based machine learning, which calculates probabilities based on statistics, causal AI uses fault-tree analysis to determine system-level failures based on component-level failures. With this systematic, top-down approach, causal AI and modern deterministic AIOps provide a determinative basis for automatic anomaly detection, root-cause analysis, security risk ranking, and business impact assessment.

Causal AI draws on supporting data, such as relationships, dependencies, and other context among network entities and events. With this context, causal AI determines the precise root cause of an issue. This approach helps teams to develop effective models or interventions for change while also predicting their potential effectiveness. It can increase confidence in business and IT decision making by clearly connecting events to an intended or unintended outcome.

The deterministic quality of causal AI can also form the foundation for reliable recommendations from emerging generative AI technologies.

Why is causal AI important?

Most AIOps approaches use predictive analytics that apply algorithms and machine learning to historical data to predict future outcomes. Such an outcome could be a CPU spike that progresses into a system failure. Predictive analysis helps an organization manage resources and improve incident response times.

This blind spot between the underlying cause and resulting effect can lead to unwanted bias and poor decision making. Predictive analysis can observe an event and predict an outcome will occur, but it can’t show that the outcome occurred because of the event. In other words, correlation doesn’t equal causation.

Causal AI, on the other hand, identifies the underlying cause of an event and its precise relationship to the outcome. Organizations can use causal AI frameworks and algorithms to ask questions and gain a deeper understanding of their CloudOps, DevOps, and SecOps use cases. For instance, these questions can include the following:

  • Why aren’t customers completing their transactions?
  • What’s causing customer churn?
  • Why is this application sluggish at certain times of the day?

Additionally, the deterministic AI approach of causal AI can determine the cause-and-effect relationship of events from a combination of metrics, traces, and log data, as well as user behavior data and other details. Thus, teams can resolve incidents immediately to prevent disruptions in service and keep an organization in compliance with service-level agreements.

Correlation AI vs. causal AI: Weighing the differences

Deterministic AI vs. statistical correlation-based AI

Correlation-based machine learning models predict outcomes from statistical relationships and are useful in many scenarios. For example, facial recognition, personal shopping, and predictive maintenance.

However, the shortcomings of correlation-based AI become evident when teams need to determine how an action would affect an outcome. While predictive models can identify the likelihood of certain positive or negative events happening, they’re unable to explain how they arrived at that forecast. They’re also unable to identify the underlying factors and cause-and-effect relationships.

Correlation-based AI and causal AI have a few additional differences, including the following:

Correlation-based AI Causal AI
Correlation-based AI relies on statistics to provide assumptions about what’s happening. Causal AI can clearly trace and explain exactly what’s happening at every step based on specific contextual data.
Correlation-based AI is probabilistic and requires humans to verify the accuracy of results. Causal AI is fact-based and thus can do automated analyses.
Correlation-based AI can make only predictions with limited ability to explain an event. Causal AI, on the other hand, provides details on how it arrived at a conclusion.
Correlation-based AI needs to be checked for bias due to the limitations of various data, algorithms, or sampling. Causal AI, however, relies on actual data and not training data and is therefore not prone to bias issues.
Correlation-based AI may be completely off base in novel situations. Causal AI can adapt to new situations and find unknown unknowns.

How does causal AI work?

Causal AI essentially works in two steps. First, it collects information and discovers problems within the data set. Then, it looks for causal relationships that help explain those issues using a plan devised from the collected data.

To better understand how causal AI works, it’s important to understand fault-tree analysis—a data-driven, fault-tree methodology used for causality analysis. Fault-tree analysis uses boolean logic to explore system-level failures. It’s a top-down approach used to identify the component-level failure, or basic event, that caused the system-level failure, or top event.

Causal AI that uses fault-tree analysis works the following way:

  1. Defines the scope of the system and what’s considered a failure.
  2. Defines top-level faults and the analysis starting point with details of the failure.
  3. Identifies precipitating events that could cause the top-level fault to occur, whether alone or with multiple concurring events.
  4. Finds the root causes of each precipitating event and event sequence.
  5. Analyzes the fault tree by looking for the events that lead to failure or are most likely to fail.

With the certainty of this systematic approach, teams can gain insight into ways to mitigate paths to failure and support system improvements, and automate resolutions.

Applying causal AI to your organization

Dynatrace Davis® AI offers continuous causal analysis to the code level that maps and understands the relationships between all of an organization’s networks, applications, and services. Using fault-tree analysis, this causal AI approach seamlessly combines topological context with metric data to quickly identify observability signals for any behavior of interest. The analysis provides insights into every entity a problem affects, enabling developers to solve problems without having to reproduce errors.

With its deterministic AI approach, causal AI provides the perfect basis for automating responses and supplying facts for reliable generative AI recommendations.

To learn more, join us for the free Dynatrace observability clinic with a live Q&A: “Observability Clinic: Leverage Davis AI to analyze your system before things break.”

The post What is causal AI? Why this deterministic AI approach is critical to business success appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/what-is-causal-ai-deterministic-ai/feed/ 0
IT automation: An AIOps guide https://www.dynatrace.com/news/blog/an-aiops-guide/ https://www.dynatrace.com/news/blog/an-aiops-guide/#respond Thu, 05 Jan 2023 20:09:36 +0000 https://www.dynatrace.com/news/?p=55517 AIOps eBook, women in leadership

AIOps is an IT approach that applies artificial intelligence to IT operations, bringing process efficiencies to organizations. Learn more about the benefits of AIOps in this AIOps guide.

The post IT automation: An AIOps guide appeared first on Dynatrace news.

]]>
AIOps eBook, women in leadership

As the globe strides into 2023 — with rapid change and macroeconomic uncertainty looming — organizations want tools and technologies that enable them to become more efficient, reduce costs, and innovate more.

These are precisely the business goals of AIOps: an IT approach that applies artificial intelligence (AI) to IT operations, bringing process efficiencies. AIOps, together with an observability platform, can enable organizations to precisely identify the root cause of cloud application performance and security issues.

AIOps also enables organizations to remediate these issues as they occur — and, in some cases, can prevent problems before they occur and disrupt operations.

According to Dynatrace data, 71% of CIOs say the explosion of data produced by cloud-native technology stacks is beyond human ability to manage.

That may be why numerous organizations have invested in AIOps technology. According to recent data, the AIOps market is worth some $17 billion per year. Growing investment reflects organizations’ need to have greater visibility into incidents within their IT environments. Further, Gartner expects the proportion of large companies that use AIOps and digital experience monitoring tools to monitor apps and infrastructure will increase from 5% in 2018 to 30% in 2023.

In this AIOps guide, we outline what AIOps is, the benefits of AIOps for organizations, and how they are using AIOps and automation.

What is AIOps?

AIOps is an IT approach that uses artificial intelligence to automate IT operations (ITOps), such as event correlation, anomaly detection, and root-cause analysis. A modern approach to AIOps serves the full software delivery lifecycle. It addresses the volume, velocity, and variety of data in complex multicloud environments with advanced AI techniques to provide precise answers and intelligent automation.

In its most basic form, IT automation executes scripts or processes on a schedule or in response to particular events. A unified approach to automating IT processes using a tightly integrated automation platform is essential to avoid poorly integrated, siloed services. An IT automation strategy should start by breaking down the workflow, the types of operations they will perform, and how teams can best monitor and optimize them in production.

Observability graphic What is AIOps? An insider’s guide to AI for IT ops — and beyond – blog

Many organizations are turning to AI for ITOps in place of manual processes. But what is AIOps, and how can it support your organization?

Dynatrace ensures continuous software quality by combining synthetic monitoring and automatic release validation, infrastructure as code, and cloud automation Developing an AIOps strategy for cloud observability – eBook

Facing increasing IT complexity, customer demand, and security issues, organizations are embracing AIOps to automate how they develop software.

Automation graphic What is IT automation? – blog

IT automation is crucial to address the challenges of maintaining complex services. See how IT automation helps teams reduce manual tasks and more.

What are the benefits of AIOps for organizations?

As organizations automate processes in their complex multicloud environments, they can gain several benefits. These environments are rapidly changing and sometimes ephemeral. IT teams can’t keep up with routine tasks of cloud management. As they fall behind, environments can suffer from application outages, increased costs, and frustrated customers. Ultimately, IT automation can deliver consistency, efficiency, and better business outcomes for modern enterprises.

Automating IT practices offers enterprises faster data centers and cloud operations, as well as increased flexibility and accuracy. Additionally, automating routine IT tasks eliminates the human element — and the potential mistakes that come with it.​

Benefits of AIOps transform business operations. AIOps and digital transformation modernize BT – blog

AIOps and digital transformation go hand in hand. Discover how companies like BT have enlisted AIOps as they evolve their service portfolio.

AIOps graphic AIOps done right – eBook

AIOps enables autonomous operations and boosts innovation, but you need to know how to implement it correctly. Learn the keys to AIOps success.

Database observability graphic Power boundless observability, security, and business analytics with Grail – resource center

Discover the benefits of the Dynatrace data lakehouse technology, Grail, as well as the importance of a data lakehouse, how it can bring insights to life, and more.

AIOps use cases: How have organizations used AIOps?

Today, many organizations strive for digital transformation. As organizations digitize and modernize, they can in turn drive innovation, increase revenue, and create operational efficiencies. But many also lack a vision for digital transformation or the means to execute on that vision through technology. Enter AIOps and digital transformation, which can help organizations actualize their goals.

For successful organizations engaged in digital transformation, AIOps and digital transformation go hand in hand.

Here are some stories of how users take to cloud modernization, digital transformation, and workflow automation with AIOps.

AIOps graphic Applying real-world AIOps use cases to your operations – blog

AIOps offers myriad benefits, including increased automation. Find out how to apply AIOps use cases to address real-world operations issues.

Digital Experience Customer panel: Innovating with advanced AIOps – video

View the on-demand panel discussion of how AIOps is powering a new era of automation and innovation in the modern cloud.

Monaco tools How Park ‘N Fly innovates with IT automation, AIOps, and observability – blog

For Park ‘N Fly, positive customer experience via technology is the backbone of the business. That’s where AIOps and observability come in.

What is hyperscale computing? How Park ‘N Fly eliminated silos and improved customer experience with Dynatrace cloud monitoring – blog

Park ‘N Fly relies on successfully integrating its booking system with its custom-built kiosks located at its off-airport parking lots. See how an observability platform and AIOps help.

Serverless computing, multi-cloud, multicloud observability Creating a seamless end-user experience with an AIOps platform approach to DEM – blog

Digital experience management helps companies deliver seamless, end-user experiences. Learn why an AIOps platform approach to DEM is the key.

The post IT automation: An AIOps guide appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/an-aiops-guide/feed/ 0
Max Tegmark on artificial intelligence: The ultimate technology for game-changers https://www.dynatrace.com/news/blog/max-tegmark-on-artificial-intelligence/ https://www.dynatrace.com/news/blog/max-tegmark-on-artificial-intelligence/#respond Mon, 14 Feb 2022 22:31:23 +0000 https://www.dynatrace.com/news/?p=48627 artificial intelligence, AI, Max Tegmark

MIT physics professor and Future of Life Institute co-founder Max Tegmark shares his big thoughts on the big possibilities of AI to change human innovation.

The post Max Tegmark on artificial intelligence: The ultimate technology for game-changers appeared first on Dynatrace news.

]]>
artificial intelligence, AI, Max Tegmark

“Think big. Really big. Cosmically big.” When it comes to artificial intelligence, MIT physics professor and futurist Max Tegmark thinks in terms of 13.8 billion years of cosmic history and the potential of the human race to influence the next 13.8 billion.

In his keynote address at Dynatrace Perform 2022, Tegmark set the stage for AI as the ultimate technology for game-changers. And more importantly, the role of humans in commanding the power in our grasp.

“When we use technology wisely, we can accomplish things our ancestors could only dream of,” Tegmark says. Through the frame of technological accomplishments in the past half-century, Tegmark laid out the possibilities for AI to transform life on earth. Using AI, humankind is already accelerating the capacity to bring forth life-saving technologies, such as diagnosing cancer and solving the protein-folding problem for biomedical research.

“The technology we’re developing is giving life the opportunity to flourish,” Tegmark says. “Not just for the next election cycle, but for billions of years.”

How far will artificial intelligence go?

Max Tegmark defines artificial intelligence simply as the “ability to accomplish complex goals.” The more complex the goals, the more intelligence they call for. There’s no law of physics that precludes artificial general intelligence (AGI), or the ability for technology to learn and accomplish anything a human can. Polls show that most AI researchers expect AGI within decades.

artificial intelligence, AI, Max Tegmark

But if a technology can learn like a human through recursive self-improvement, does that mean AI will leave humanity in the dust? Will self-learning technologies create a superintelligence that far exceeds human capacity? And if so, are we doomed or saved?

To answer these questions, Tegmark suggests it’s a matter of perspective. Through human ingenuity, the tech industry has improved computational ability many millions of times since computers were invented. And engineers have extracted only a minute fraction of the energy that is theoretically possible from energy sources. That includes known sources, like gasoline and coal, or what people can possibly extract from other sources.

Max Tegmark sees the enormous benefits of AI as long as humans cultivate the wisdom we need to minimize risks.

Winning the wisdom race with artificial intelligence

“I’m confident we can have an inspiring future with high tech, but it’s going to require winning the wisdom race,” Tegmark says. “The race between the growing power of the technology and the wisdom with which we manage it.”

In the analog world, people learn by making mistakes. If you try something and it fails (or someone dies in a car crash), you adjust the approach (and invent seat belts). But at the scale of AGI, a reactive trial-and-error approach can be costly and potentially catastrophic. Instead, you can begin to proactively predict what could go wrong and apply safety engineering principles.

To help win the wisdom race, Tegmark and four colleagues co-founded the Future of Life Institute, designed to keep powerful technologies going in the right direction. Isaac Asimov’s three laws of robotics were too limited, so Tegmark and his colleagues developed the 23 Asilomar AI Principles, a set of practical and ethical guidelines for developing and applying artificial intelligence. More than 1,000 researchers and scientists worldwide have adopted and signed these principles.

Aligning AI’s goals with our own

“Any science can be used as a new way of harming people or a new way of helping people,” Tegmark says. To illustrate, he shares three of the 23 Asilomar Principles:

  • Avoid a destabilizing arms race in lethal autonomous weapons. We shouldn’t allow AI algorithms to decide to kill people.
  • Mitigate AI-fueled inequality. We should share the great wealth artificial intelligence helps produce so everyone is better off.
  • Invest in AI safety research. This effort can make systems robust, secure, and trustworthy.

artificial intelligence, AI, Max Tegmark

AGI safety requires what Max Tegmark calls “AI alignment.”

“The biggest threat from AGI is not that it’s going to turn evil, like in some silly movie,” Tegmark says. “The worry is it’s going to turn really competent and accomplish goals that aren’t aligned with our goals.” For example, one way to look at the extinction of the West African black rhino is that humans’ goals weren’t aligned with the rhinos’ goals.

So humanity doesn’t go the way of those rhinos, we must design AI to understand, adopt, and retain our goals. “This way we can steer AGI to accomplish our goals for an inspiring future,” Tegmark explains.

Envision an amazing future, not a dystopic one

As any captain of industry knows, a positive vision is essential for business success. Once you know where you want to go, then you can identify the problems and potential pitfalls. Instead of imagining a dystopic future, we should envision an amazing future. The United Nations’ 17 Sustainable Development Goals, for example, provide a roadmap to a future in which humanity thrives.

artificial intelligence, AI, Max Tegmark

“These are challenging and noble goals adopted by nearly every country on earth,” Tegmark says. “Artificial intelligence can help us attain these sustainability goals better and faster. As we continue toward AGI and beyond, let’s not just aim toward them by 2030, let’s accomplish all of them and raise our ambition to go beyond them.”

As individuals, we have an important role to figure out how to steer artificial intelligence and make these changes happen.

As a company, Dynatrace and its causation-based Davis AI are building this future for customers by delivering the vision of a world where software works perfectly.

For the 30,000 AI game-changers and technologists attending Dynatrace Perform across the world, Professor Tegmark gave an assignment. “Be proactive. Think in advance about how to steer technology and where you want to go with it. We will be the masters of our own destiny by actually building it.”

The post Max Tegmark on artificial intelligence: The ultimate technology for game-changers appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/max-tegmark-on-artificial-intelligence/feed/ 0
How an AIOps platform can shift left–and why it should https://www.dynatrace.com/news/blog/how-an-aiops-platform-can-shift-left/ https://www.dynatrace.com/news/blog/how-an-aiops-platform-can-shift-left/#respond Tue, 24 Aug 2021 07:51:28 +0000 https://www.dynatrace.com/news/?p=45872 Benefits of AIOps transform business operations.

As organizations layer more technologies into their DevOps toolchains, an observability-based AIOps platform that can shift left is a good strategy.

The post How an AIOps platform can shift left–and why it should appeared first on Dynatrace news.

]]>
Benefits of AIOps transform business operations.

Since the term artificial intelligence for IT operations (AIOps) was coined by Gartner in 2016, organizations have considered it a good strategy to adopt an AIOps platform or AIOps tools. These tools can help manage and automate anomaly detection and incident response for IT operations in production environments.

But as organizations adopt CI/CD practices and layer in a growing array of cloud-native solutions and open-source technologies to their DevOps toolchains, it’s becoming clear that AIOps can unlock value along the entire digital value chain.

Data in modern cloud-native environments is a continuum—from software development through service delivery all the way to customer interactions. Everything that happens provides telemetry to help teams discover root causes, inform decisions, and automate processes.

In this increasingly integrated landscape, organizations benefit not just from an AIOps platform, but an all-in-one observability, deterministic AI, and analytics platform—a software intelligence platform—to continuously automate, analyze, predict, and remediate IT issues while navigating their digital acceleration journey.

AI applied to cloud ops

As the size and complexity of distributed cloud computing systems continue to grow, typical AIOps approaches to monitoring, diagnosing, and repairing software are not scaling. Traditional AIOps relies on correlating data in order to reduce alerts, which is slow and inaccurate and does little to identify root causes.

An intelligent AIOps platform with end-to-end observability can leverage AI-based algorithms and real-time data analytics to automate triaging, response, and remediation for common IT issues, including unexpected downtime, system latency, or determining why a Kubernetes pod was terminated.

An integrated platform that includes AIOps, observability, and analytics can consume and analyze the increasing volume of cloud data to automate and optimize these routine monitoring and management tasks. Such an observability-based AIOps platform can also provide advanced insight for IT and DevOps, while reducing mean time to resolution (MTTR) and speeding up mean time to discovery (MTTD). With end-to-end visibility into multicloud environments, an intelligent AIOps platform with advanced analytics enables faster innovation, higher quality, more efficiency, and ultimately, better business outcomes.

In addition to driving enterprise automation, there are six key capabilities an all-in-one AIOps platform approach delivers:

  1. Alert management
    • Replaces monitoring tool alert storms with accurate, reliable root-cause analysis.
    • Eliminates up to 90% of false alarms and reduces noise with deterministic AI fault tree analysis.
    • Observes, analyzes, and enables automated response in near-real time.
  2. Automation
    • Contextualizes and processes large volumes of operational data.
    • Uses this high-fidelity, context-rich collected data to create real-time topology and service flow maps.
    • Provides analysis and AI-powered insights across the application lifecycle.
    • Continuously discovers changes to environments, apps, and services.
  3. Incident prioritization and routing
    • Delivers relevant insights to the right people at the right time, providing precise answers with root-cause determination, prioritized by business impact.
  4. Event causation
    • Uses causation-based AI to point directly to the root cause and impact of a failed test run, an application slowdown, or system outage, or to drive decisions about whether to release a piece of software.
  5. Predictive analytics
    • Monitors the entire technology stack end-to-end to predict and prevent future disruptions before they occur.
  6. Auto-remediation
    • Automates anomaly detection, problem notification, and self-remediation with full-stack monitoring and integration with workflow automation platforms, such as ServiceNow and Jenkins.

Why AIOps needs to “shift-left”

With the volume of data increasing, and the demand for services rocketing upward, the need for AIOps is no longer limited to IT operations. As DevOps and SRE practices mature, pre-production workflows need AIOps capabilities just as acutely.

A typical continuous integration/continuous delivery (CI/CD) pipeline follows the following sequence:

  1. Source — creating source code
  2. Build — compiling the application
  3. Test — testing code for functionality
  4. Release — pushing code to the repository
  5. Deploy — moving code to production

The term “shift-left” refers to the practice of performing a task at an earlier stage of development before it goes to production, such as automated testing at the source phase instead of when code is ready to be released.

Shift-left applied to AIOps integrates AI into the full DevOps lifecycle, including data ingestion, building code, and testing for enhanced software quality and deeper root-cause analysis before code is deployed to production. A software intelligence platform that includes end-to-end observability in its approach to AIOps delivers continuous alert and incident management, automatically observes and identifies anomalies in CI/CD pipelines, and prevents issues from reaching the production stage, resulting in more efficient builds and quicker, higher-quality releases of new versions of software.

Shifting AIOps left means development teams can easily leverage production service-level objectives (SLOs) as criteria for building quality gates into acceptance testing earlier in the development cycle. It also means teams can initiate auto-remediation for CI/CD workflows by integrating with software configuration and deployment management technologies, such as Chef, Puppet, and Ansible.

Faster time-to-value with an intelligent AIOps platform

An AIOps platform based on continuous discovery and end-to-end observability can detect anomalies before they affect the CI/CD pipeline or impact customer experience and SLOs. It can automate validation processes by using SLO-based quality gates with events, tags, and APIs integrating seamlessly with existing CI/CD workflows — for faster automated deployments and time-to-value.

AI-enhanced alerting and escalation using advanced algorithms based on deterministic fault-tree analysis can automatically route incidents to the appropriate team, empowering them with metadata and context that results in accurate, reliable, and precise root-cause analysis. If automated processes are unable to address a slowdown or outage, DevOps is then given a clear path to remediation, eliminating time spent on “problem triage,” which drives faster innovation and better quality.

A shift-left AIOPs platform approach mitigates the cost of IT downtime

Every CIO and CFO knows IT downtime is costly — potentially adding up to thousands of dollars per minute or more depending on the organization’s size and reach of services. But what may be as significant is the impact downtime and performance issues can have on already pushed-to-the-limit DevOps and IT teams. They are under increasing pressure to maintain system reliability and prevent outages of highly complex, distributed multicloud operations. The constant context switching required to hop from issue to issue also derails focus on mission-critical concerns and distracts teams from their core functions.

By shifting AIOps left using production-based performance criteria as a quality gate earlier in the development cycle, teams can release more resilient software. If an issue arises anywhere in the DevOps workflow, an integrated AIOps platform with all-in-one observability, deterministic AI, and analytics can automatically remediate issues and optimize performance based on system health and user demands — preventing or minimizing the duration of outages and reducing costs. It also combats IT teams’ top stressors by reducing the need for manual intervention in IT operations and DevOps workflows, helping to prevent burnout and costly employee churn.

Shift AIOps left with an integrated software intelligence platform

Teams undergoing digital transformation are discovering that shifting AIOps left into DevOps and SRE workflows can accelerate and increase the effectiveness of their DevOps and SRE initiatives.

To help developers pinpoint problems and automate more processes during the development, test, and delivery phases of DevOps, Dynatrace seamlessly integrates AIOps into the CI/CD pipeline, bringing fault-tree analysis to pre-production workflows. Shifting AIOps left enables developers to discover and auto-remediate issues in pre-production so they can optimize processes and deliver higher quality code to production.

Using the Cloud Automation control plane—powered by Keptn, an open-source technology for cloud-native application life-cycle orchestration—Dynatrace provides release analysis, version awareness, and SLO-based quality gates so teams can automate releases at all stages of the DevOps pipeline. By integrating with DevOps tools like Chef, Puppet and Ansible, Dynatrace can execute closed-loop remediation workflows or orchestrate ITSM tools to trigger incident management workflows.

To learn more about how Dynatrace approaches AIOps, see the eBook: AIOps Done Right.

Dynatrace was also named a leader in AIOps in the Forrester Wave. Read the report here.

Developing an AIOps strategy for cloud observability

Download our free eBook to learn the best practices for developing an AIOps strategy that drives efficiency, innovation, and better business outcomes.

The post How an AIOps platform can shift left–and why it should appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-an-aiops-platform-can-shift-left/feed/ 0
Auto-Mitigation with Dynatrace AI – or shall we call it Self-Healing? https://www.dynatrace.com/news/blog/auto-mitigation-with-dynatrace-ai-or-shall-we-call-it-self-healing/ https://www.dynatrace.com/news/blog/auto-mitigation-with-dynatrace-ai-or-shall-we-call-it-self-healing/#respond Tue, 07 Nov 2017 14:02:56 +0000 https://www.dynatrace.com/blog/?p=21638 Cloud-Microservices

After our AI-Driven DevOps webinar with Anil from Verizon Enterprise I got into a debate with my colleague Dave Anderson on how to call the auto mitigation approach we discussed during the webinar: Is it “auto mitigation”, “auto remediation” or shall we be bold and call it “self-healing”? Instead of getting caught up on terminology […]

The post Auto-Mitigation with Dynatrace AI – or shall we call it Self-Healing? appeared first on Dynatrace news.

]]>
Cloud-Microservices

After our AI-Driven DevOps webinar with Anil from Verizon Enterprise I got into a debate with my colleague Dave Anderson on how to call the auto mitigation approach we discussed during the webinar: Is it “auto mitigation”, “auto remediation” or shall we be bold and call it “self-healing”?

Instead of getting caught up on terminology I thought to write this blog post and explain our thoughts and let you decide on what you think we should be calling it.

Here is the animated slide from our webinar that started the discussion:

AI Detected Problem details allow us to build smarter automated mitigation actions. No need to wake up engineers at 2AM every time a problem happens.
AI Detected Problem details allow us to build smarter automated mitigation actions. No need to wake up engineers at 2AM every time a problem happens.

Our point of the webinar was that with the Dynatrace AI (Artificial Intelligence) analyzed data, we can trigger and build much smarter Auto Mitigation actions. Here are our thoughts to explain the slide:

  • Escalate at 2AM? Dynatrace auto detects the problem and how many end users and service endpoints are impacted and this translates directly to the severity of our escalation process.
  • Auto Mitigate! Dynatrace is aware of all important events across all entities involved in the problem. (e.g: network connection issue, critical log message after a configuration change, CPU exhaustion) This allows us to write smarter auto-mitigation steps to address the root cause and not the symptom of the problem!
  • Update Dev Ticket! If the mitigation actions work, we can automatically update the Jira ticket about the executed actions and in the daily stand-up, developers can discuss what happened last night.
  • Mark Bad Commits! If the mitigation actions didn’t solve the problem, we still have the option to rollback and mark the responsible Pull Request as BAD and Detailed Analysis can be done in the post-mortem retrospective!
  • Escalate as last resort! If rolling back doesn’t solve the situation, it’s time to definitely escalate – even at 2AM!

Auto Mitigation Implementation with AWS Lambda

Inspired by the work that my colleague Alois Reitbauer did around Auto-Mitigation in the last couple of months, we sat down with our easyTravel Demo Team. easyTravel – in case you’ve never heard of it – is our #1 application we use to demo the capabilities of Dynatrace Fullstack Monitoring and the Dynatrace AI. It is also available for anyone to download and install.

easyTravel comes with many different components, services and some built-in problem patterns, that can be enabled or scheduled, on demand or via REST. easyTravel was also recently enhanced to scale up and down, individual dockerized components, such as the backend service.

Rafal Psciuk, Team Lead in our Gdansk office, thought about good auto-remediation use cases, just in case something goes wrong with easyTravel. He implemented two use cases which he recently demoed to me and I found it just “AWSome.” 😊 With the following explanation, I hope it inspires you to combine Dynatrace AI detected problems with your automation tools to implement auto-mitigation, auto-remediation or self-healing (or whatever you want to call it 😊).

As we host easyTravel in AWS, Rafal decided to leverage AWS Lambda to automate mitigation. He wrote Lambda functions Dynatrace triggers when a problem gets detected. The function then analyzes the actual problem, all its correlated events and issues course correcting actions, depending on the actual root cause of the problem. These actions can range from restarting processes, scaling up or down docker containers, rerouting traffic, running cleanup or database scripts.

Here is the schematic overview of what happens when Dynatrace AI detects a problem:

An AI Detected Problem triggers the Lambda Mitigation Call and passes all Problem details for smarter mitigation actions to fix the problem.
An AI Detected Problem triggers the Lambda Mitigation Call and passes all Problem details for smarter mitigation actions to fix the problem.

Process Crash: Restart Process

The first use case is a simple, but very common use case. Dynatrace AI detected a JavaScript error rate increased to 94%, impacting 713 real user actions per minute. Dynatrace AI detected the root cause of this error to be a crash of a CouchDB process which is used by Tomcat (Application Server) to serve all dynamic page requests:

Dynatrace automatically tells us impact and root cause thanks to data from OneAgent, Smartscape and AI-based Anomaly Detection detection
Dynatrace automatically tells us impact and root cause thanks to data from OneAgent, Smartscape and AI-based Anomaly Detection detection

One would normally not suspect the correlation of a spike in JavaScript Errors to a CoucheDB process crash. Thanks to Dynatrace OneAgent (FullStack Data), Smartscape (Automated Dependency Model) and the AI engine (Anomaly Detection) the root cause was automatically detected.

We can leverage this information to our advantage and build smarter mitigation scripts. Here are some additional thoughts on this problem pattern:

#1 – Restart Process: If we see a crash of a back-end process that is impacting our end users or any type of SLA (Service Level Agreement) we have to try to restart that process. We do have all the information which process is impacted and where that process normally runs.

#2 – Prevent Future Process Crashes: Besides knowing which process, on which machine crashed, Dynatrace also captures any critical log message, deployment or configuration change event that preceded the crash. (e.g: an update to CouchDB was deployed 5 mins prior or there was a connection pool configuration change leading to too many incoming connections causing the crash). Knowing more about what caused the crash allows us to not just restart the process and see it crash again soon after, but it also allows us to solve the problem that caused it to crash in the first place.

Our AWS Lambda Mitigation Function

Here are parts of the AWS Lambda function that Rafal implemented. First, he parses the incoming problem details and then processes the event based on whether it matches the problem pattern he is handling:

exports.handler = function(event, context, callback) {
  var parsedEvent = parseEvent(event);
  if (!isCorrectProblem(parsedEvent)) {
       callback(null, { "Not a valid problem": "true" });
  } else {
    processEvent(parsedEvent, callback);
  }
};

function parseEvent(event) {
  console.log('Loading event');
  console.log(event);
  var parser = new EventParser(event);

 var parsedEvent = {
    hasApplication: parser.hasApplication('www.easytravel.com'),
    hasProcess: parser.hasProcess('CouchDB_ET'),
    isOpen: parser.isOpen(),
  problemId: parser.getProblemId()
};

  console.log("isOpen: " + parsedEvent.isOpen + " hasApplication: " + parsedEvent.hasApplication + " hasProcess: " + parsedEvent.hasProcess);

  return parsedEvent;
}

If he needs more data from Dynatrace, he simply queries the Dynatrace REST API to access deployment events, timeseries or Smartscape information. Based on the information in the problem details and event history reaches out to the easyTravel Orchestration Engine to restart CouchDB.

function processEvent(parsedEvent, done) {
  if (parsedEvent.isOpen) {
    console.log("Resolving problem");
    disablePlugins(parsedEvent, done);
  } else {
    console.log("Send resolved comment");
    addProblemResolvedComment(parsedEvent.problemId, done);
  }
}

function disablePlugins(parsedEvent, done) {
  var error = {
    couchErr: null,
    javaScriptErr: null
  };

  disablePlugin(couchDBPlugin, function(couchErr) {
    error.couchErr = couchErr;
    disablePlugin(javascriptErrorsPlugin, function(javaScriptErr) {
      error.javaScriptErr = javaScriptErr;
      disablePluginCallback(parsedEvent, error, done);
    });
  });
}

The Lambda function also puts a comment on the Dynatrace Problem to indicate when the function started to mitigate the problem and when it eventually solved the problem! All of this is neatly documented when looking at the Dynatrace Problem history!

The Lambda Mitigation Code not only fixes the problem the AI detected, it also protocols each step along the way back to the Dynatrace Problem ticket.
The Lambda Mitigation Code not only fixes the problem the AI detected, it also protocols each step along the way back to the Dynatrace Problem ticket.

The key point in this use case is that we can quickly react to a situation where a critical process crashes. We know which process it is and we also know which impact it currently has.

Slow Database impact Microservices: Increase Microservice Capacity

The second use case is one we keep seeing more frequently with architectures leveraging microservices that share resources such as a back-end service or database.

The login page is experiencing a 140% increased load time caused a slowdown of the backend database which is used by 5 different service instances!
The login page is experiencing a 140% increased load time caused a slowdown of the backend database which is used by 5 different service instances!

The first observation is that the database is the problem and therefore, it has to be fixed. A closer look at all the details Dynatrace captured, shows us that the slow down came from a high number of UPDATE statements executed by one of the 5 services that share this database. It was basically a batch job that somebody had triggered, impacting all other services that use that database.

All other services had to wait much longer for database responses, which meant that their threads were blocked and couldn’t accept newly incoming requests fast enough.

The increase in throughput and response time for UPDATE queries caused an impact on all sorts of SQL queries executed by 5 different services.
The increase in throughput and response time for UPDATE queries caused an impact on all sorts of SQL queries executed by 5 different services.

Knowing these details about the database activity and impact, we have different options (and probably more than these listed) to mitigate this problem:

#1 Throttle Batch Service: Depending on how important the batch job is, we could give it lower priority and slow down its execution of SQL updates.

#2 Temporarily Increase or deploy Cache Layer for Read Only Use Cases: While these UPDATEs run, switch your front-end services to access a cached data set instead of hitting the database. And change that route once updates are completed. While this means that users may see slightly outdated data for a couple of minutes, it prevents the performance impact.

#3 Scale front-end services: If service requests are slowed down due to a slow database, it means you may run into capacity issues, leading to not only very slow requests, but rejected requests. In that case, we could simply scale up instances of the front-end services until the database is back to normal speed. The same approach is valid if we would see an increase in end user requests – we can simply scale that tier and dependent tiers to handle the new load patterns.

Our AWS Lambda Mitigation Function

In this use case, Rafal implemented an extended version of the mitigation function that executes remediation actions based on the details that Dynatrace is pushing as part of the Problem Notification Integration. Rafal’s Lambda function additionally queried the Dynatrace Problem and Timeseries API, to learn more about traffic patterns on both front-end, as well as other services that access that problematic database. He was automating what a human would do as well by pulling in more information to make smarter decisions!

Rafal ended up implementing the mitigation function to temporarily scale up the number of front-end microservice instances that supported the login page. That opened up more worker threads for incoming requests from people hitting the login page. This mitigated the problem, as users could access the login page without a major performance impact or even error pages.

Here is the Problem Evolution replay of this use case – showing how the Dynatrace AI not only detects the problem, but also detects when it was mitigated / solved:

Problem Evolution showing how the slowdown in the database impacted the MicroJourneyService until we scaled up by adding more instances
Problem Evolution showing how the slowdown in the database impacted the MicroJourneyService until we scaled up by adding more instances.

The Lambda function also puts a more detailed comment on the Dynatrace Problem to indicate WHICH situation was detected (high database traffic) and how it solved the problem!

“Self Healing” function not only documented remediating action but also more details about what more details it found in the problem details that lead to that specific scaling action
“Self Healing” function not only documented remediating action, but also more details about what more details it found in the problem details that lead to that specific scaling action.

The key revelation of this use case is that just treating this as a database issue would probably result in suboptimal responses. (e.g: fixing the problem on the database level and not on the service that causes the problem). With the AI, we can execute much better and automated remediation actions.

Auto-Mitigation, Auto-Remediation or Self-Healing! What is it?

Reading my own blog (before publishing), shows me that even I used all three terms throughout the blog 😊 – well – whatever we call it. I believe it is the way we have to think about automating problem resolution. Gone are the days where we are always getting people on bridge calls, spending time to find out who is to blame and how to actually fix the problem. We now have better data (OneAgent), automatic dependencies (Smartscape), anomaly detection (baselining, machine learning and domain knowledge) and a REST API that allows us to automatically react to the actual root cause. This will reduce MTTR (Mean Time To Repair) drastically!

The post Auto-Mitigation with Dynatrace AI – or shall we call it Self-Healing? appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/auto-mitigation-with-dynatrace-ai-or-shall-we-call-it-self-healing/feed/ 0