problem analysis | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Tue, 23 Jun 2026 07:00:25 +0000 en hourly 1 Davis CoPilot expands: Get answers and insights across the Dynatrace platform https://www.dynatrace.com/news/blog/davis-copilot-expands-get-answers-and-insights-across-the-dynatrace-platform/ https://www.dynatrace.com/news/blog/davis-copilot-expands-get-answers-and-insights-across-the-dynatrace-platform/#respond Tue, 04 Feb 2025 16:00:17 +0000 https://www.dynatrace.com/news/?p=67510 Davis CoPilot

We’re excited to announce that Davis CoPilot Chat is now available across the Dynatrace platform. Davis CoPilot™, launched in October 2024 to support Dynatrace users with access to their data, now extends across the platform, streamlining user onboarding and providing comprehensive support and contextual insights from various Dynatrace® Apps. With the new Davis CoPilot conversational […]

The post Davis CoPilot expands: Get answers and insights across the Dynatrace platform appeared first on Dynatrace news.

]]>
Davis CoPilot


Update: We’ve launched Dynatrace Assist, our next-generation AI chat that goes far beyond answering questions.
Dynatrace Assist is the evolution of Davis CoPilot®.

We’re excited to announce that Davis CoPilot Chat is now available across the Dynatrace platform. Davis CoPilot™, launched in October 2024 to support Dynatrace users with access to their data, now extends across the platform, streamlining user onboarding and providing comprehensive support and contextual insights from various Dynatrace® Apps. With the new Davis CoPilot conversational interface, users can leverage natural language to quickly get answers to their questions, making it easier than ever for users to interact with Dynatrace.

Intuitive access to information boosts team productivity

We understand that taking advantage of the numerous features and functionalities offered by platforms like Dynatrace can be challenging. To help you navigate this and boost your efficiency, we’re excited to announce that Davis CoPilot Chat is now generally available (GA). This new feature provides information and guidance exactly when and where you need it, making your Dynatrace experience smoother and more efficient.

Davis CoPilot can be accessed anytime directly from the Dock.

Davis CoPilot leverages the power of generative AI to answer your questions through a globally accessible chat interface. We’re proud to say that Davis CoPilot is multilingual: you can ask questions and get answers in many different languages, including French, Spanish, German, Portuguese, Chinese, Japanese, and, of course, English. Davis CoPilot provides immediate, accurate responses, eliminating the need for extensive searches and reducing dependency on support channels. This makes knowledge more readily available and boosts productivity and user experience for both new and experienced users.

Davis CoPilot Chat follows our recent announcement of the general availability of Quick Analysis in Notebooks and Dashboards, which makes data accessible to technical and non-technical users alike. This means you can interact with data stored in the Dynatrace Grail™ data lakehouse just by using natural language.

Simplify onboarding and quickly find what you’re looking for with Davis CoPilot

You can start using the Davis CoPilot conversational interface immediately. Simply enable Davis CoPilot and assign the relevant user permissions, and the Davis CoPilot button will appear in the Dock.

Start a new conversation with Davis CoPilot Chat by selecting it in the Dock or by pressing CTRL/CMD + I and entering your question.

Davis CoPilot is great for guiding new and occasional users
Figure 2. Davis CoPilot is great for guiding new and occasional users

New users can quickly get up to speed with Dynatrace by asking Davis CoPilot for help with basic commands, setup instructions, and troubleshooting tips. This reduces the learning curve and enables new users to become productive faster. The conversational interface provides step-by-step guidance, making the onboarding process smoother and more efficient.

If you’re already familiar with Dynatrace, you can rely on Davis CoPilot to provide detailed explanations for a wide range of expert questions related to exploring new use cases, advanced configuration topics, and building custom apps.

Here are some examples of questions you can ask Davis CoPilot:

  • Onboarding: How do we start sending OpenTelemetry data to Dynatrace?
  • Understanding Dynatrace: What is the difference between an event and a problem in Dynatrace?
  • Exploring Dynatrace solutions: How can we comply with the Digital Operational Resilience Act (DORA) using Dynatrace?
  • Configuring your environment: How do I set up an alert based on an anomaly detector?
  • Developing custom apps: How can I import external table data and visualize it using the Dynatrace App Toolkit?

Get contextual assistance at the press of a button

Davis CoPilot seamlessly integrates into our use-case-specific Dynatrace Apps, offering you contextual insights and guidance at the press of a button. While we plan to release additional contextual app integrations in the coming months, several will be available a few weeks after launch, allowing Davis CoPilot to provide you with insights into:

  • Kubernetes warning signals
  • Individual problem details and the relationships between problems
  • Database performance optimization

Simplify Kubernetes: Davis CoPilot decodes warning signals

Understanding the background and root cause of warnings often requires in-depth subject matter expertise. That’s why we integrated Davis CoPilot into Kubernetes. Instead of manually looking up error messages, Davis CoPilot translates warning signals into clear, understandable language. In addition, Davis CoPilot offers a list of typical root causes and related remediation steps. This way, newcomers can quickly become proficient, and experts can elevate their expertise to hero status.

Davis CoPilot provides contextual guidance for Kubernetes warning signals
Figure 3. Davis CoPilot provides contextual guidance for Kubernetes warning signals

Problems demystified: Davis CoPilot provides insights into root causes

In Problems, Davis CoPilot provides clear summaries of problems, their root causes, and the suggested remediation steps. Davis CoPilot explains individual issues in clear language from the problem details page and can perform a comparative analysis when multiple problems are selected from the list view. This helps you identify common root causes and propose corrective steps without relying on a team of experts and waiting for hours for critical insights. If you want to learn more, have a look at Wolfgang Beer’s latest blog post and learn more about recent advancements in the Problems app.

Davis CoPilot explains problems in clear language
Figure 4. Davis CoPilot explains problems in clear language

Optimize database performance: Understand query execution plans

Query execution plans provide detailed information on how a database will execute an SQL query. While these provide the raw data on how to improve query performance and reduce resource consumption, they require expert knowledge to read and interpret. Now, in Databases, Davis CoPilot can provide natural language explanations of execution plans, breakdowns of relevant details, and recommendations on how to improve statement performance. This gives non-expert database users, such as developers, the knowledge they need to optimize their application performance and database utilization.

Davis CoPilot explains query execution plans
Figure 5. Davis CoPilot explains query execution plans

Empower your teams with Davis CoPilot today

The launch of Davis CoPilot Chat marks the second milestone of our journey. We’re committed to continuously enhancing the assistant’s capabilities with upcoming features, including query explanations, workflow actions, and troubleshooting guides.

Get started with Davis CoPilot today and transform how you and your teams interact with Dynatrace:

Thanks for joining us on this exciting journey. We look forward to your feedback and to seeing how Davis CoPilot helps your teams achieve their goals.

Davis CoPilot Chat, as well as the Dynatrace Apps integrations mentioned in this blog post, will be available starting with the release of Dynatrace SaaS version 1.307.

The post Davis CoPilot expands: Get answers and insights across the Dynatrace platform appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/davis-copilot-expands-get-answers-and-insights-across-the-dynatrace-platform/feed/ 0
Advancing AIOps: Preventive operations powered by Davis AI https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/ https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/#respond Tue, 04 Feb 2025 16:00:06 +0000 https://www.dynatrace.com/news/?p=67673 Davis AI alerts

The 2024 CrowdStrike incident demonstrated our societal vulnerabilities to IT outages. A faulty software update caused widespread issues, impacting critical services globally, including airlines, banks, hospitals, and public safety systems. Despite recent advancements such as containers, Kubernetes, and platform engineering, it’s evident that managing enterprise software services has become increasingly complex. IT operations must be prepared to quickly address and mitigate disruptions, ensuring business continuity and minimizing damage.

The post Advancing AIOps: Preventive operations powered by Davis AI appeared first on Dynatrace news.

]]>
Davis AI alerts

AI, especially AIOps, has emerged as a pivotal solution, promising to avoid downtime. The 2024 State of AI Report highlights this trend, with 89% of technology leaders anticipating that AI will significantly enhance incident response by learning to automate and optimize various tasks, such as performance monitoring and workload scheduling.

Blue screens of death at LGA airport due to the July 2024 CrowdStrike outage. (Source: Wikimedia Commons.)
Figure 1. Blue screens of death at LGA airport due to the July 2024 CrowdStrike outage. (Source: Wikimedia Commons.)

AIOps can identify and address potential issues before they become major incidents by learning from history and analyzing large amounts of data in real time. This approach improves operational efficiency and resilience, though it’s not without flaws. The complexity of IT environments and the changing nature of threats necessitate human oversight and ongoing adjustment of AIOps systems to handle unforeseen challenges and ensure optimal performance. Additionally, predictions based on historical data are reactive, solely relying on past information to anticipate future events, and can’t prevent all new or emerging issues. This limitation highlights the importance of continuous innovation and adaptation in IT operations and AIOps strategies.

“The shift from reactive to preventive operations represents the next evolution in AIOps.”
Bernd Greifeneder, CTO Dynatrace

When Dynatrace set out with Davis® AI over 10 years ago, pioneering AI-driven operations, we focused initially on problem identification before moving on to problem remediation. The next milestone in enhancing the capabilities of Davis AI—another pioneering step forward in AI-driven operations—is outright problem prevention. In this blog post, we explain how the unique combination of causal, predictive, and generative AI—augmented by the latest Davis AI advancements—is transforming how Dynatrace customers manage and optimize their IT infrastructure.

Automatic root cause detection

Modern, complex, and distributed environments generate a substantial number of events. This necessitates additional requirements such as minimizing the total number of issues, eliminating false positives, and conducting accurate root cause analysis.

Dynatrace has a longstanding reputation for accurately analyzing root causes and identifying related events. While other methods typically rely on mere correlation and historical data analysis, we’ve further enhanced our capabilities by implementing causational analysis, which leverages contextual information automatically gathered during data ingestion and processing in addition to historical data analysis. This is achieved using Dynatrace Grail™, our causational data lakehouse, which unifies all data in an always-up-to-date topology model. By applying causal AI to incoming data in real time, Davis instantly learns and continuously adapts to new information. This facilitates more precise root cause analysis and anomaly detection, including identifying seasonal anomalies and establishing auto-adaptive thresholds.

Root cause analysis with the Problems app
Figure 2. Root cause analysis with the Problems app

When applying this Davis root cause detection within our own IT environment, Davis effectively filters out over 99.9% of incoming data noise, condensing hundreds of thousands of daily system events into no more than four or five incidents that require attention from our IT operations team.

These algorithms are not limited to monitoring IT environments. At our February 2025 Dynatrace Perform session on exploratory analytics with AI-driven insights, the Performance Engineering Lead of XXXLutz—one of the world’s largest furniture retailers operating more than 370 stores across Europe—explains how XXXLutz utilizes Davis AI to proactively identify critical order drops, allowing them to respond quickly and effectively to changing market conditions and ensuring that their business remains agile and responsive to the needs of their customers.

Problem journey and reactive remediation

At the core of Dynatrace problem remediation stands the Problems app—an optimized view into opinionated insights, details, and context of each detected issue—for Operations, SREs, and developers. It filters billions of log lines, including the topology of each incident and its affected entities, for efficient problem triaging and troubleshooting, resulting in a 56% faster mean time to repair (MTTR) for critical incidents.

With the latest release, we drive this further by improving the automatic connection of relevant log and trace data for further drill down, presenting the full context of an issue in a single view. This provides comprehensive visibility into even complex architectures, simplifying the process of examining relevant details and addressing code-level issues, reducing 100 clicks and manual filtering to a single click with no loss of context.

Comparative analysis of multiple problems with Davis CoPilot
Figure 3. Comparative analysis of multiple problems with Davis CoPilot

By utilizing Davis CoPilot™, you can conduct comparative analyses of multiple issues, obtain natural language summaries of individual problems, and receive contextual recommendations along with specific remediation steps.

You can also link troubleshooting guides created in Notebooks to remediated issues, thereby building an intelligent knowledge base. Davis automatically connects additional documents as well as stored workflows. So the next time a similar problem arises, Davis brings up related guides, enabling teams to learn from previous experiences and reducing the risk of knowledge loss.

Harness your collective knowledge by connecting troubleshooting guides
Figure 4. Harness your collective knowledge by connecting troubleshooting guides

Please refer to our recent blog posts for more information on utilizing Problems for AI-driven insights and the latest Davis CoPilot advancements.

Automating the remediation

While obtaining comprehensive insights is beneficial, true transformation occurs through the use of tools that automatically execute remediation steps. To implement these “AI-driven operations,” it’s essential to forecast future requirements, including capacity demands, potential system failures, and security incidents.

Traditional forecasting engines typically depend on historical data, stored in metrics. In contrast, Davis AI generates real-time predictions, facilitating proactive operations. This capability is due to Davis’s ability to process raw data, such as logs, for forecasting, leveraging Grail to execute previously unattainable queries.

Consider the following scenario: You begin by retrieving and analyzing logs to identify relevant values for automation. Once this task is complete, you proceed to your pipelining tool to configure ingestion rules that extract these values into metrics and then wait several weeks for your prediction engine to generate alerts that can serve as triggers for your workflows.

However, when utilizing Dynatrace with its integrated anomaly detection and forecasting capabilities, you gain the advantage of schema-less data analysis and the ability to process any raw data into time series in real time. This significantly reduces the time required to establish AIOps workflows from several weeks to less than 30 minutes.

Preventive operations

The complexity of modern software environments makes it challenging to determine a service’s reliability solely through testing. It’s impractical to emulate scenarios such as generating a million tickets to assess performance capabilities. This necessitates real-time insights and operations rather than reactive problem-solving or raising alerts to notify personnel.

Preventive operations address this need by enabling proactive corrective actions before issues arise, akin to predictive maintenance. AI-supported anomaly detection identifies parameters that deviate from the norm, allowing for automatic configuration adjustment to mitigate potential problems preemptively.

Dynatrace offers the only unified, AI-powered platform for all data, all teams, and all possibilities.
Figure 5. Dynatrace offers the only unified, AI-powered platform for all data, all teams, and all possibilities.

Davis CoPilot combines the “power of three”:

  • Davis causal AI for identifying anomalies and root cause analysis
  • Davis predictive AI for precise forecasting and determining when to take action
  • Generative AI capabilities that perform actions beyond simply sending notifications or restarting services

In this way, Dynatrace extends AIOps beyond traditional IT operations tasks and addresses complex scenarios, including security use cases such as threat observability. Consider the following real-world example:

At Dynatrace, we log all failed login attempts. We can predict potential threats when abnormal patterns are identified and raise a security event by utilizing seasonal baselining. The subsequent workflow involves checking the IP address and generating a threat score. Upon reaching a certain threshold, a new ruleset is automatically added to the web application firewall. This entire process is fully automated, running before a problem even occurs, significantly reducing the response time from over an hour to a fraction of a second.

In another instance, automatic log pattern analysis crawling our application logs decreased the number of bugs in the production environment by 15% and freed up time previously spent on log analysis and triaging (in pre-prod), equivalent to 17 full-time employees. Consequently, these 17 developers can now dedicate their efforts to adding more value to Dynatrace.

Summary

The State of AI report states that over 88% of technology leaders anticipate AI will enhance incident responses and improve their teams’ ability to predict and proactively resolve service-affecting issues.

With Dynatrace, organizations are prepared to evolve their ITOps and SRE departments from troubleshooting to prevention, getting proactive with forecasting, and utilizing generative AI instead of purely focusing on history-focused root cause analysis.

Start your preventive operations journey with smart automation and auto-remediation that prevents larger issues.

Are you interested in gaining more insights?

The post Advancing AIOps: Preventive operations powered by Davis AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/feed/ 0
Transform your operations with Davis AI root cause analysis https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/ https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/#respond Tue, 08 Oct 2024 19:03:46 +0000 https://www.dynatrace.com/news/?p=66044 root cause analysis

Complexity is ever-increasing in today’s fast-paced world of software deployments and cloud infrastructure. This is why Davis® AI root cause analysis is an indispensable tool for Operations, Site Reliability, and DevOps teams.

The post Transform your operations with Davis AI root cause analysis appeared first on Dynatrace news.

]]>
root cause analysis

Without AI-assisted observability tooling, the productivity of operations teams drops, leading to a dramatic increase in Mean Time to Repair (MTTR) and a significant rise in the personnel needed to manage critical incidents. In an era dominated by automated, code-driven software deployments through Kubernetes and cloud services, human operators simply can’t keep up without intelligent observability and root cause analysis tools.

Modern observability has evolved from simple metric telemetry monitoring to encompass a wide range of data, including logs, traces, events, alerts, and resource attributes. Dynatrace Root Cause Analysis (RCA) seamlessly integrates all this information, providing crucial analysis to remediate incidents in real time.

Problem feed for fast triage and remediation of AI-detected problems.
Figure 1. Problem feed for fast triage and remediation of AI-detected problems.

By offering root cause analysis on top of the highly flexible Grail™ data lakehouse, Dynatrace empowers SRE and operations teams to further reduce MTTR. Direct access to the underlying data allows the automatic RCA analysis to eliminate data silos and to dive deep into every aspect of the collected incident data.

Unlike generic DIY query frontends, the Dynatrace Problems app is a tailor-made solution for efficiently supporting operations use cases. This approach ensures that your operation teams have all the tools they need to manage modern software deployments.

Transform your operations today with the new Problems app and stay ahead in the ever-evolving software and cloud infrastructure landscape.

Rapid response to critical incidents

Operations teams can quickly focus on incoming Davis AI-detected and -analyzed problems by referring to the problems feed.

The problem feed is designed to prioritize active issues, ensuring they always appear at the top, regardless of how long they’ve been ongoing. This default sorting strategy, which uses time as a secondary criterion, guarantees that Operations teams never overlook an active problem, no matter which primary filter is applied.

You can focus on your domain using the filter bar at the top, the quick filters on the side, or both. The chart feature allows for quick analysis of problem peaks at specific times.

Operations teams will appreciate the ability to sort problems by duration and the number of affected entities. This aids in assessing Davis-detected root causes and prioritizing remediation efforts. The native multi-select feature lets users open a filtered group of problems simultaneously, facilitating quick comparisons and detailed analysis.

Streamline deployment insights with AI-generated summaries

Every second counts during wide-scale incidents affecting large parts of your production systems. This is why precisely showing the root cause ultimately helps to speed up problem resolution.

You can multi-select a cohort of active problems, select Show detail, and review all critical problem details, including preview charts and event details, without losing the context of your problem feed.

The new problem experience transparently displays all the available details, with prominently displayed root-cause markers to precisely guide your attention.

In the realm of cloud infrastructure management, having a clear and concise view of your deployment’s health is crucial. Our dedicated deployment perspective offers just that, showcasing the hierarchy of affected and related infrastructure components. The root cause of any issue is prominently marked with a root-cause badge, making it easy to identify and address problems swiftly.

This perspective not only highlights the affected cloud regions but also provides a quick summary of the Kubernetes context where your workloads encountered failures. Gone are the days of clicking and navigating through multiple dashboards. Instead, you receive an AI-generated summary as an affected deployment architecture diagram.

This diagram, akin to a UML (Unified Modeling Language) deployment diagram, offers a familiar representation for software architects, ensuring they can quickly grasp the situation and take necessary actions. By streamlining the visualization of deployment issues, we empower teams to resolve problems more efficiently and maintain optimal performance.

To save time, the root-cause component is preselected, and all the details of the root cause are displayed on the right, along with charts showing the detected breaches from learned normal behavior.

You can review each individual finding on all problem-affected entities by selecting the individual deployment components or by switching to the detailed event perspective, which shows all the single events that the root cause analysis collected into a single problem.

Confirm the AI-detected root cause and review the deployment context.
Figure 2. Confirm the AI-detected root cause and review the deployment context.

In addition to using markers for swift root cause analysis, operations teams often seek to attach valuable remediation hints and playbooks for familiar scenarios.

By implementing a flexible event tagging mechanism, event sources and detectors can be easily customized to include additional custom event properties. This allows for markdown-formatted event description text that can contain remediation links, as illustrated in the screenshot below.

Root cause remediation hints as markdown links
Figure 3: Root cause remediation hints as markdown links

The Dynatrace Semantic Dictionary helps identify the semantics of well-known event properties and provides convenient platform intents. For instance, entity links (dt.entity.*) or links to the responsible settings entry (dt.settings.object_id) that detected and opened an event can be included. These settings links save valuable time when adjusting detection sensitivity for thresholds or baselines. Additionally, the event setting property can be utilized in a DQL query to create a table of the top-triggering configurations or to automate settings changes using an automation workflow.

Quick access to incident logs

The seamless integration of logs powered by Dynatrace Grail™ data lakehouse with Davis AI root cause analysis is a game changer for modern operation teams, as it offers a quick summary of all incident-relevant logs.

The Dynatrace root cause engine already combines all incident-relevant information to recommend log queries, which saves a lot of navigation time and completely eliminates the need to manually identify complex log filters.

A single click on the Problem details log perspective immediately surfaces all relevant logs related to the given incident, as shown below.

Failure rate increase logs
Figure 4.
100 errors and warnings of failure rate logs
Figure 5.

Within this view the Operations team can further refine the query or adapt the filters and open a notebook to persist the log findings for critical post-mortem documentation purposes.

Root cause analysis in a user-focused context

Most modern application stacks are deployed through Kubernetes, making it essential for operations teams to focus on Kubernetes clusters, cloud resources, and workloads of critical services.

Since operations engineers prefer not to switch contexts, a consistent root-cause experience is provided regardless of where the user journey begins.

Whether you start your remediation journey within the Infrastructure & Operations app or the Kubernetes app, you receive the same root-cause information without needing to navigate between different apps. This seamless embedding of root-cause information into the current context saves valuable time during incident remediation.

Root cause shown in context of the Infrastructure & Operations context.
Figure 6. The root cause is shown in the context of Infrastructure & Operations.
CPU throttling root cause shown in Kubernetes context.
Figure 7. CPU throttling root cause shown in Kubernetes context.

Notify and automate to speed up remediation

The Problems app features a global problem indicator that is always visible within the Dock to capture your attention. This indicator shows whether there are active problems within the environment. You can personalize this number by selecting and saving a problem filter within the problem feed, as demonstrated below. The saved default filter is then automatically applied to the global problem indicator, reducing the number of active problems for the user.

Select Alerting (bell icon) to set up alerts related to filtered problems and configure email addresses for notification recipients.

The email payload and the use of an email address for notifications are preset, allowing for a personalized notification setup, as shown below.

Save the personal default filter and set up email notifications.
Figure 8. Save the personal default filter and set up email notifications.
Find the global problem indicator in the Dock.
Figure 9. Find the global problem indicator in the Dock.

You can take a further step towards answer-driven automation and use the detected Davis problem event to trigger workflow automation. Automatically remediate an issue using our no-code workflow actions for collaboration (for example, Slack, Microsoft Teams, ServiceNow, Pagerduty) and remediation (for example, AWS, Red Hat Ansible, Kubernetes).

The introduction of a filterable global problem indicator ensures that Operations teams remain focused on active problems within the environment, even while exploring data in Notebooks or Dashboards.

In future updates, the Problems app will support multiple named filters and introduce Segments as the primary method for using and sharing numerous predefined filters among operations teams.

Outlook

The newly released Problems app enhances transparency by providing detailed AI-detected root-cause information. It also offers convenient deployment and architectural visualizations, along with a log perspective, to help operations teams reduce Mean Time to Repair (MTTR).

In future updates, we aim to support the ability to acknowledge and label incoming problems, improving team coordination. Additionally, plans include a visual representation of the application map, direct propagation of information such as application IDs into the problem feed, and support for segments to filter the problem feed.

Summary

For over a decade, Dynatrace has been at the forefront of integrating AI into incident analysis, particularly through Davis root cause analysis.

Davis is now essential for Operations, Site Reliability, and DevOps teams, helping them to navigate the complexities of modern software deployments and cloud infrastructure.

Without Davis, the productivity of these teams would plummet, leading to longer Mean Time to Repair (MTTR) and increased staffing needs to handle critical incidents.

In today’s automated deployments and cloud services, traditional observability tools fall short, unable to keep pace with the intelligence needed for effective root cause analysis.

Modern observability encompasses various data sources, from metrics to logs and events, requiring intelligent tools like Davis to seamlessly integrate and analyze this information in real time. By providing Davis on top of the flexible Grail data lakehouse, Dynatrace empowers teams to swiftly reduce MTTR by accessing and previewing incident data comprehensively.

The Davis Problems app streamlines triage, allowing teams to swiftly focus on AI-detected issues. Its intuitive interface simplifies problem resolution.

Try out the new Problems app in the Dynatrace Playground.

The post Transform your operations with Davis AI root cause analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/feed/ 0
Davis AI: Your personal interactive troubleshooting assistant  https://www.dynatrace.com/news/blog/davis-ai-your-personal-interactive-troubleshooting-assistant/ https://www.dynatrace.com/news/blog/davis-ai-your-personal-interactive-troubleshooting-assistant/#respond Wed, 24 May 2023 08:59:02 +0000 https://www.dynatrace.com/news/?p=57846 business observability

When Dynatrace started reinventing cloud-native service tracing and observability ten years ago, it was already clear that human operators were overwhelmed with traditional monitoring systems' massive raw data inflow. Besides being unable to watch that amount of telemetry data on dashboards, classic operations teams were also blown away by the sheer number of alerts they received 24/7 from hundreds of different monitoring tools. 

The post Davis AI: Your personal interactive troubleshooting assistant  appeared first on Dynatrace news.

]]>
business observability

Update: We’ve expanded AI-powered dashboarding with Dynatrace Intelligence, delivering smarter insights, forecasting, and anomaly detection across the Dynatrace platform.
Dynatrace Intelligence is the evolution of Davis AI®, improving how users interact with and understand observability data.

With the introduction of Davis® root-cause detection, Dynatrace reduced the amount of single-alert spam that arises when large-scale incidences occur. Instead of immediately firing off an alert for all raw events, the Davis root-cause engine follows each violating service’s causal relationships. By automatically following the causal direction of the topology between services and their underlying infrastructure, Davis collects all raw events that belong to the same root cause and then notifies you by raising a problem.

With interactive problem mode, Dynatrace introduces a new, powerful troubleshooting assistant. This blog post explains how Davis can help reduce your MTTR (mean time to resolve) using interactive user guidance that retains context when drilling deeper into problem analysis.

Davis problem analysis
Select any entry in the side panel to navigate to the corresponding metric, in context.

Faster remediation through precise root cause analysis

Once Davis identifies a problem, a Problem overview page is created, which shows a comprehensive management summary of what happened (impact) and the root cause of the problem. DevOps teams use this page to quickly identify and remediate unexpected incidences.

Usually, the journey doesn’t stop here. When the DevOps team has finished their work, software experts must investigate the underlying software stack. They need to analyze all relevant information that Davis found along the deployment stack to avoid such problems in the future. When navigating to the underlying service—identified as the root cause—the problem detail page opens with retained problem context, which includes:

  • Date and time of the current problem, so you don’t need to manually adapt the date and time on each page in the analysis journey.
  • A side panel that interactively informs you about all problem-related information for the relevant service.
  • Davis highlights all relevant problem information on each page you navigate to.

The screenshot below shows how Davis interactively guides you by highlighting all the relevant information with red and yellow markers (on the left side) while showing a list of AI root-cause findings in the side panel on the right (if the Davis side panel is closed, an icon is displayed on the right-hand panel so you can re-open it).

AI root-cause findings

Davis highlighting detected problems in side panel

Optimize your software stack using Davis interactive problem mode

Watch out for red and yellow markers in the navigation section headers—these indicate that Davis has found information related to the problem.

The red marker highlights events and their duration, whereas the yellow marker indicates metric anomalies where suspicious metric change points were found during the problem analysis. The yellow metric change points highlight a point in time, while the red markers represent event durations.

If you select one of the markers (either directly or via the side panel), you can view additional information, such as the timeframe and duration.

Davis AI change point and event markers

Davis AI change point (in yellow on the left) and event duration (in red on the right) markers

Meeting SLO requirements

In addition to providing context to detected problems, Davis also supports you when spikes are detected in connected SLOs (Service Level Objectives). Via the dedicated SLO button in the top bar, service-level objectives relating to the selected service can be reviewed immediately without losing context.

Spikes can easily be investigated by selecting a timeframe and clicking Analyze. Davis instantly collects all connected signals and provides relevant, contextual information. Watch the following video for examples of how the interactive problem mode helps identify SLO-relevant issues.

Davis SLO analysis
Review related Service Level Objectives (SLOs)

Summary

Davis problem detection and root cause analysis is essential for modern AIOps (Artificial Intelligence for IT Operations) and DevOps to minimize the MTTR. Real-time insights are crucial for quickly triaging unexpected incidents and remediating them in a timely manner.

Davis interactive problem mode guides you through all the detailed problem-related information and marks problems visually to make them easier to understand. It also seamlessly integrates user-defined SLOs, including leveraging Davis AI for analyzing SLO degradations, which saves precious time during critical incidents. You no longer need to leave the context of your page when using the side panel for navigational help to dig through all relevant findings and SLOs discovered during root cause analysis.

We’re, of course, highly interested in your feedback! We encourage you to try the interactive problem mode and share your feedback and product ideas via the Dynatrace Community. Every message we receive helps us to continuously improve the Dynatrace platform.

The post Davis AI: Your personal interactive troubleshooting assistant  appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/davis-ai-your-personal-interactive-troubleshooting-assistant/feed/ 0
Automate prioritization of quality improvements with Dynatrace SLO violation prediction and problem analysis https://www.dynatrace.com/news/blog/automate-prioritization-of-quality-improvements-with-dynatrace-slo-violation-prediction-and-problem-analysis/ https://www.dynatrace.com/news/blog/automate-prioritization-of-quality-improvements-with-dynatrace-slo-violation-prediction-and-problem-analysis/#respond Wed, 24 Aug 2022 17:38:21 +0000 https://www.dynatrace.com/news/?p=52855 SLOs graphic

Does either of the following situations sound familiar to you? You have plans to enjoy an upcoming vacation, but production problems in your area of responsibility prevent you from taking time away. Or, you’ve been asked to ensure the customer experience of a production system that you have no familiarity with—you’re now responsible for a […]

The post Automate prioritization of quality improvements with Dynatrace SLO violation prediction and problem analysis appeared first on Dynatrace news.

]]>
SLOs graphic

Does either of the following situations sound familiar to you? You have plans to enjoy an upcoming vacation, but production problems in your area of responsibility prevent you from taking time away. Or, you’ve been asked to ensure the customer experience of a production system that you have no familiarity with—you’re now responsible for a system for which you don’t know the structure of the underlying services.

How can you know where best to invest time and money into quality improvements in such situations? Which known issues have the most significant impact on your customer experience and business success? Without automatic problem prioritization, you might easily misallocate your resources to low-priority issues.

This blog post explains how you can effectively use your business-critical metrics as service level objectives (SLOs). The problems that Dynatrace identifies in your systems are automatically linked to critical SLOs and their related error budget and burndown rate.

The error budget and burn rate provide site reliability engineers (SREs) with the information they need to take action before their end users are affected. Such information dramatically improves your chances of avoiding war-room meetings with other stakeholders.

Error budgets and burndown rates

Dynatrace provides alerts on high error budget burn rates that predict when an error budget will be depleted if no action is taken. Analysis of detected problems includes root cause analysis for quick problem remediation and the assurance that your SLO targets are met. While the SLO status of a critical metric might be okay (displayed in green) or at a warning level (indicated in yellow), the error budget might be consumed quickly. In such situations, Dynatrace Davis® AI detected problems show the identified root cause of the problem in addition to a call to action to mitigate the SLO violation before it affects your users or your error budget. This way, you can stop the consumption of your error budget and focus on the right problem at the right time.

Example SLO dashboard tiles
Figure 1: The SLO dashboard tiles provide all the information you need: The red arrows show the error budget trend going down, and the red warning icons indicate that Davis AI has detected problems that impact these SLOs. Select these warning icons to view the related problem descriptions, which include root cause analysis and call-to-action details that you can use to fix SLO-impacting problems.

Error budgets as a tool for prioritizing investments

An error budget can be understood as a metric value that equates to an acceptable rate of technical errors that can occur in a system before the errors affect end user experience. In essence, error budgets tell you when investments into quality improvements are worth the effort.

To succeed, organizations must put their customers’ needs front and center. By utilizing an error budget, SREs can measure where customer satisfaction is at risk due to a high burn rate.

Dynatrace supports SREs in their need to:

  • see which SLO error budgets are burning down to effectively prioritize work on problems that impact those SLOs.
  • see the trend of SLO status/error budgets to predict if and when an SLO will exhaust its error budget.
  • receive alerts for high burn rates so that mediation efforts can be planned.
  • get support finding the root cause of a high error budget burn rate so that the problem can be fixed before the error budget is depleted.

Different approaches to investment prioritization

It’s vital to distinguish between reactive work that’s based on alerting (“fire fighting”) and proactive planning for investments into quality and automation. Both approaches require prioritization. The reactive approach results in a low mean time to repair rate (MTTR—one of the DORA metrics) for newly discovered problems that impact SLOs and a high error budget burn rate. With the proactive alerting-based approach, SREs must identify and implement solutions for mitigating depleting error budgets. Conversely, intelligent prioritization of investments into quality improvements and automation better ensures SLOs. In many cases, multiple problems contribute to a depleted error budget and SREs must manually investigate all related problems to prioritize investments into quality.

Dynatrace provides solutions for both the proactive approach and the reactive approach to prioritizing quality investments.

The reactive approach

In reactive fire-fighting style prioritization based on identified root causes, Dynatrace Davis® AI identifies problems and shows you the number of potentially impacted SLOs. You can link directly to the impacted SLOs from the problem page (Figure 2 below). This way, prioritizing work on one problem over the other is easy. SREs can set up alerts for high error budget burn rates so that they can react quickly to impacted SLOs before those error budgets are depleted. Dynatrace Davis AI presents the root cause analysis for each detected problem so that SREs can define action items that will improve error budget burn rates and avoid any SLO breaches with minimal mean time to repair.

The Problems view shows the count of affected SLOs and crosslinks to those SLOs
Figure 2: The Problems view shows the count of affected SLOs and crosslinks to those SLOs.

The proactive approach

With the proactive approach to investment prioritization for quality improvements and automation, problems related to SLOs show all affected SLO error budget burn rates and depleted error budgets (Figure 3 below). They also show a count of all problems that affect each SLO. With this information, SREs gain an overview of all problems that contribute to each depleted error budget. This sort of crucial insight is invaluable for planning future quality improvements and implementing automated problem remediation.

This SLO overview shows an error budget burn rate icon and enables you to create alerts for high error budget burn rates
Figure 3: This SLO overview shows an error budget burn rate icon and enables you to create alerts for high error budget burn rates.

While the status of the first SLO shown in Figure 3 is still okay, the error budget has already been consumed. This is why SREs need to receive such alerts before error budgets are depleted. As the second SLO is already in a bad state, besides fire fighting, the investigation of the seven related problems will help the SRE to understand the history of previous problems and common root causes. The SRE can then determine exactly where quality improvements or remediation automation is needed most.

What’s next?

Find out how easy it is to set up your first SLOs and then automate and scale SLO practices in your organization.

Explore how fully automated remediation of problems can help to keep your SLOs in good shape. Utilizing on-demand synthetic tests and release validation in Dynatrace provides you with continuous assurance of the status of your SLOs—all with a single solution.

Dynatrace is happy to provide you with a demo or proof of concept for Cloud Automation. We also offer a free Dynatrace trial if you want to get started directly!

The post Automate prioritization of quality improvements with Dynatrace SLO violation prediction and problem analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automate-prioritization-of-quality-improvements-with-dynatrace-slo-violation-prediction-and-problem-analysis/feed/ 0
Identify issues immediately with actionable metrics and context in the Dynatrace Problem view https://www.dynatrace.com/news/blog/identify-issues-immediately-with-actionable-metrics-and-context-in-dynatrace-problem-view/ https://www.dynatrace.com/news/blog/identify-issues-immediately-with-actionable-metrics-and-context-in-dynatrace-problem-view/#respond Fri, 03 Jun 2022 15:13:27 +0000 https://www.dynatrace.com/news/?p=51283 Dynatrace Community graphic

Dynatrace introduces new enhancements to application problem analysis that allow I&O teams to identify and prioritize issues such as crashes and errors in the problem view. These enhancements add enriched context that help explain not just the “what” but also the “why” behind the issues so that DevOps teams can take corrective action based on reliable AI-driven answers rather than basic monitoring data.

The post Identify issues immediately with actionable metrics and context in the Dynatrace Problem view appeared first on Dynatrace news.

]]>
Dynatrace Community graphic

Whenever a performance problem is flagged, Infrastructure and Operations (I&O) practitioners strive to resolve the issue as soon as possible by identifying the root cause, understanding the impact, obtaining the relevant details, and fixing the issue within the shortest possible timeframe—the meantime to resolution (MTTR). But this is often not as intuitively simple as it should be in other solutions where DevOps teams must click through a series of screens and dashboards to get to the root cause. This results in delays, frustration amongst team members, and lost conversions.

So, whenever your end users’ digital experience is bogged down by a problem, whether it’s the result of a synthetic monitor (browser as well as HTTP), mobile app monitoring, or web monitoring, your teams need to see the most pertinent information about the impact and the root cause at a glance. Often, raised problems are the result of custom settings with fixed thresholds or the creation of custom events for alerting. In large enterprise environments, it’s often difficult to determine who configured such settings and thresholds.

Leverage AI assistance to deliver better customer experience

At the heart of Dynatrace Digital Experience Monitoring (DEM) is Davis, the state-of-the-art AI engine that accurately prioritizes the severity of each detected performance anomaly in terms of its potential impact on real users and business KPIs. Without any configuration or the need for a data scientist, Davis provides instant and automatic answers to degradations in service, anomalies in behavior, and impact on user experience so that I&O teams can chart a clear course of action to resolve issues. To facilitate this further, we’ve introduced new information in problem details when Digital Experience Monitoring issues are raised.

Problem details have been further enhanced so that practitioners can confidently rely on summarized intelligence to ensure a better user experience. You can swiftly determine why an alert was raised or understand how a custom performance threshold that was set up previously by another person in the organization is related to a performance/slowdown issue.

  • For DEM problems, the business impact analysis shows the number of users that are potentially impacted by a problem. This section has been improved to show you the ratio of the number of users affected by the problem compared to the total number of users using the application during the problem timeframe.
  • The impact section is now enriched with the list of affected applications, services, and other entities affected by the problem for increased productivity and data-driven decisions.
  • The root cause analysis section now contains links to custom events for alerting and manual performance thresholds. This ensures greater agility and reduces the time to resolution.

How you can leverage the enhanced intelligence

Here are five use cases where you can benefit the most from the new information in DEM problem details:

DEM problem information use cases

  • Problem prioritization: The ratio of affected users to observed users for web and mobile problems is clearly shown. This ensures that I&O practitioners are better able to understand the impact of a problem. With this information, you can prioritize problems that have the highest number of affected users, drill down to affected user sessions, and understand the impact of a problem.

Business impact analysis Screenshot Dynatrace

  • Mobile crash increase troubleshooting: Your teams are short on time. In case of a spike in the crash rate for your app, you can go directly to the crash overview. With this update, your teams save time and can quickly access the crashes to analyze the root cause. The type of breached baseline (auto-detected baseline or fixed manual threshold) is also available as additional information in the crash rate increase section.

Problem detail Screenshot Dynatrace

  • Root cause and settings in enterprise environments: With a large user base, you face a mammoth task in identifying which custom events trigger certain alerts. With this update, any I&O practitioner who comes across a custom alert in problem details can identify exactly which custom event was manually set up for alerting. This is particularly handy in enterprise environments where only a select few people can create custom events for alerting.

Problem root cause Screenshot Dynatrace

  • Synthetic problem troubleshooting: Looking for the root cause of a failing synthetic monitor can be tricky, especially when the tested application is not fully monitored. To facilitate the troubleshooting of synthetic monitors (HTTP as well as browser monitors), we’ve added more actionable data directly in problem details (for example, direct links to recent failing executions, monitor settings, and monitor results pages filtered by the problem duration). You can also see more information like timestamps for recent configuration changes, which, in some cases, can be the cause of a synthetic monitor’s failure. Bringing this information into problem details saves you time in troubleshooting synthetic monitoring problems.

Synthetic problem troubleshooting Screenshot Dynatrace

  • Web error-rate increase troubleshooting: Looking for the errors that need to be assigned and subsequently fixed is time-consuming. Hence, the top errors can now be seen at a glance, and you can directly access the multidimensional analysis page by selecting the Analyze errors button. This enables you to quickly see the details of an error and spend more time solving the problem rather than looking for it.

Web error-rate increase troubleshooting Screenshot Dynatrace

How to get started

If you’re interested in seeing this in action, the good news is that most of this intelligence and analysis has been available since Dynatrace version 1.231; the root cause details and affected user ratios are available since Dynatrace version 1.238.

New to Dynatrace?

To learn more about how Dynatrace can help optimize your user experiences across mobile, web, IoT, and APIs, visit Dynatrace Digital Experience Monitoring (DEM) or sign up for a 15-day free trial.

The post Identify issues immediately with actionable metrics and context in the Dynatrace Problem view appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/identify-issues-immediately-with-actionable-metrics-and-context-in-dynatrace-problem-view/feed/ 0