Davis AI | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Tue, 23 Jun 2026 07:06:15 +0000 en hourly 1 Dynatrace observability is now a Kiro power https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/ https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/#respond Fri, 12 Jun 2026 21:12:32 +0000 https://www.dynatrace.com/news/?p=74536

In this blog, we'll introduce the Kiro power for Dynatrace, show what it unlocks for developers, and walk you through how to get it up and running.

The post Dynatrace observability is now a Kiro power appeared first on Dynatrace news.

]]>

What is the Kiro power for Dynatrace?

The Kiro power for Dynatrace delivers live observability data, root cause analysis, and remediation suggestions directly into the Kiro IDE, with no JSON editing or manual MCP setup.

Kiro is an AI-powered IDE that helps developers move from idea to working code through spec-driven development and an agentic assistant. To make the assistant genuinely useful in unfamiliar domains, Kiro recently introduced powers: curated, partner-validated bundles of MCP servers, steering files, and best practices that install with a single click and load on demand when a relevant task comes up. Install a power, and Kiro’s agent gains specialized expertise the moment you need it.

For Dynatrace customers already working in Kiro, it’s the shortest path yet from code to production insight. For developers new to Dynatrace, it’s a one-click way to ground Kiro’s reasoning in real facts from your environment, not guesses.

Why this matters for developers

Developers have historically been one step removed from production. When something breaks after deployment, the path to figuring out what went wrong usually runs through a Site Reliability Engineering (SRE) or operations team, and AI coding assistants can’t automatically and reliably remediate issues in software they’re unfamiliar with. Agents that can write code are guessing about how their code behaves in production unless they have access to real telemetry data.

The Dynatrace Kiro power for Dynatrace closes this gap through Dynatrace Intelligence, the agentic operations system at the core of the Dynatrace platform. Kiro’s answers are grounded in deterministic, causal AI and real-time production data, not probabilistic guesses.

When a developer starts a task by writing a prompt, Kiro evaluates the conversation, identifies the relevant power using keywords, and dynamically activates power. Kiro then loads Dynatrace MCP tools and power instructions, providing skills to investigate problems, query live observability data, surface root causes, and even execute and verify remediations.
Figure 1. When a developer starts a task by writing a prompt, Kiro evaluates the conversation, identifies the relevant power using keywords, and dynamically activates the power. Kiro then loads Dynatrace MCP tools and the power instructions, providing the skills needed to investigate problems, query live observability data, surface root causes, and even execute and verify remediations.

With the tools provided by the Kiro power, developers can:

  • Investigate live incidents and get root cause analysis directly in Kiro chat
  • Query metrics, logs, and traces from production using natural language
  • Surface security vulnerabilities affecting the code they’re working on
  • Get remediation suggestions grounded in what’s actually happening in their environment

“Using Kiro powers for Dynatrace has been a total game-changer in the observability space. Deep-dive root cause analysis of complex system issues that once required lengthy manual intervention now happens in seconds, giving us unprecedented speed and confidence.”

Mike Kobush, Sr. Software Performance Engineer, NAIC

How to install the Kiro power for Dynatrace

Getting started takes only a few steps. Once installed, the Kiro power activates automatically when Kiro detects a relevant task. Mention an incident, a slow service, or anything that needs production context, and the Dynatrace tools and guidance will load in Kiro chat.

Prerequisites

  • A Dynatrace account. If you don’t already have one, you can start a free 15-day trial.
  • Kiro installed on your system.

Prepare the Dynatrace connection

First, create a Dynatrace Platform Token, which Kiro will use to authenticate. Then add the required permissions for the Dynatrace MCP server.

Install the Kiro power

The power can be installed from either the Kiro IDE or the Kiro powers website. For this walkthrough, we’ll use the IDE.

  1. Launch the Kiro IDE.
  2. Select the Ghosty icon with the lightning bolt to open the powers panel.
  3. Select Dynatrace Observability from the Recommended
  4. Select Install. The power is registered with placeholder values for the Dynatrace URL and token. Therefore, Kiro will show an error message that the MCP server can’t be reached.
  5. To complete the configuration, select Open Settings and replace the placeholders with your environment details.

Configure your tenant and token

In the settings file, replace the two placeholders:

Placeholder Replace with
YOUR_DT_URL https://TENANT_ID.apps.dynatrace.com/platform-reserved/mcp-gateway/v0.1/servers/dynatrace-mcp/mcp. Replace TENANT_ID with your Dynatrace environment ID (visible in your environment URL, for example https://<ENVIRONMENT_ID>.apps.dynatrace.com/ui).
YOUR_BEARER_TOKEN The Dynatrace platform token you created earlier (for example, dt0s16.XXXXX).

Start asking questions

Open a new chat in Kiro and start interacting with your Dynatrace environment using natural language. Query active problems or security vulnerabilities, request a root cause analysis to identify critical issues in production, or pull related logs and traces, all without leaving the IDE.

See it in action

The short demo below walks through installing the Kiro power for Dynatrace, verifying the connection, and running a first query against your environment to list the top 10 vulnerabilities detected by Dynatrace.

Installing and activating the Kiro power for Dynatrace (video)
Figure 2. Installing and activating the Kiro power for Dynatrace (video)

Get started with the Kiro power for Dynatrace

Kiro powers transform what used to be a stitching exercise (MCP servers here, steering files there, custom instructions somewhere else) into one single, ready-to-use bundle. The Kiro power for Dynatrace applies the same idea to observability: live production insight, causal root cause analysis, and remediation grounded in real telemetry, all available the moment a developer needs them.

The result is a tighter loop between writing code and understanding how it behaves in production. Less waiting for diagnostic data from someone else. Less guesswork from an AI assistant operating without context. And, more time spent on the work that actually matters.

Ready to try it? The Kiro Power for Dynatrace is publicly available: install it from kiro.dev or the Kiro IDE and start asking your environment questions.

Using Kiro and the Kiro power for Dynatrace root cause analysis (video)
Figure 3. Using Kiro and the Kiro power for Dynatrace root cause analysis (video)
Experience the Kiro power for Dynatrace for yourself.

The post Dynatrace observability is now a Kiro power appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/feed/ 0
Write the future: Create your own agentic workflows https://www.dynatrace.com/news/blog/write-the-future-create-your-own-agentic-workflows/ https://www.dynatrace.com/news/blog/write-the-future-create-your-own-agentic-workflows/#respond Thu, 08 Jan 2026 08:00:10 +0000 https://www.dynatrace.com/news/?p=72347 Agentic workflows with Davis CoPilot

Imagine commissioning le Carré and Fleming to build your perfect undercover agent: quietly embedded in the system you’re watching. You hand in your mission brief, which includes the target, objective, and behaviors to track. Your agent observes without drawing attention, reporting insights back to you. On cue, the information flow you’ve carefully orchestrated turns signals […]

The post Write the future: Create your own agentic workflows appeared first on Dynatrace news.

]]>
Agentic workflows with Davis CoPilot

Imagine commissioning le Carré and Fleming to build your perfect undercover agent: quietly embedded in the system you’re watching. You hand in your mission brief, which includes the target, objective, and behaviors to track. Your agent observes without drawing attention, reporting insights back to you. On cue, the information flow you’ve carefully orchestrated turns signals into actionable intelligence that helps pre-empt risk.

Dynatrace doesn’t write spy fiction. However, even better, Dynatrace now lets you write your own smart agentic workflows that deliver intelligent reports and react to changes in your environment based on your objectives.

Adding generative AI to your workflow

Using the power of gen AI, Davis CoPilot® transforms your workflows into agentic instructions. Davis CoPilot lets you explore data using conversational language, translating complex data and queries into summaries, and provides intelligent recommendations across Dynatrace.

Integrated into Dynatrace Workflows, Davis CoPilot is your toolkit for building smart automations, bringing the power of generative AI into your mission-critical workflows.

Build conversational automation that adjusts to live data based on your instructions, sending summaries of your crash logs directly to Slack
Figure 1. Build conversational automation that adjusts to live data based on your instructions, sending summaries of your crash logs directly to Slack.

By embedding Davis CoPilot in your workflows, you can associate any automation with any number of conversational automations. Your workflows can even perform actions autonomously when combined with precise Davis® AI forecasting, for instance, scaling resources based on forecast demands.

In real time, these workflows monitor live data, summarize critical issues, and identify remediation paths or emerging threats. When scheduled, these smart workflows help you outsource routine tasks, such as alerting stakeholders of costly queries.

Let’s look at some examples of how these smart workflows can help you in your daily work.

Build agentic workflows that respond to critical events

Proactive guiding through complex problem remediation

Let’s assume you want to build an automation that cuts through alert noise and analyzes a problem as it occurs, guiding you through the remediation. When a new problem is detected, Davis CoPilot extracts the problem details, summarizes the situation, and provides tailored remediation guidance. By embedding it into a smart workflow, you can select your preferred automation to automatically syndicate this information, populate a ticket in ServiceNow or Jira, or post it to a dedicated Slack channel.

See how you can set up a workflow automation that automatically sends summaries and remediation guidance when a new problem is detected.

Monitor emerging threats to help you orchestrate a response

Next, you can build an agentic workflow that helps you monitor emerging threats and assess their risk to your environment as vulnerabilities are detected in your tenant. In plain language, you instruct your agent to extract IOCs, query security events in your environment, and correlate them with observability data in your environment. Information provided by the external threat feed is automatically matched against the live context in your tenant. The agent has now collected all the necessary information and provides a reliable risk assessment, along with a plan to orchestrate a response, directly in your Slack channel, ensuring around-the-clock visibility and a rapid response.

With a single workflow, you can monitor emerging security events as they occur, understand their impact, and determine the next steps.
Figure 2. With a single workflow, you can monitor emerging security events as they occur, understand their impact, and determine the next steps.
Example of a tailored and contextual analysis delivered to Slack as the issue arises
Figure 3. Example of a tailored and contextual analysis delivered to Slack as the issue arises

Write the future: Build an agentic workflow that autonomously auto-scales your resources

Dynatrace helps you build agents that reason autonomously. The key is to deliver data as precise as Dynatrace forecast capabilities. In this example, we linked the power of Davis AI to forecast demand, with generative AI and GitHub automations. Davis AI predicts the number of resources the hyperscaler infrastructure will need based on forecasted demand. When Davis AI notices a scaling need, Davis CoPilot interprets the data and autonomously edits the manifest using the GitHub automation. Giving you one end-to-end workflow that automatically scales resources up or down based on forecasted needs. To see this in action, watch how this workflow autonomously edits a manifest based on Davis AI suggestions to auto-scale a Kubernetes cluster.

Schedule agentic workflows to optimize routine tasks

Do you feel like sleeping in a little later? Maybe stretching your lunch break a little longer? Running that extra hill without sacrificing your productivity? Scheduling Davis CoPilot into your smart workflow is a great way to automate recurring tasks and save time.

Build an automation that predicts resource consumption

A recurring challenge for SREs is analyzing the full environment to predict future bottlenecks or over-resourcing and continuously translating the data to update stakeholders. Even with great observability in place, you need to ensure that you interpret the data and make timely decisions to inform future provisioning.

By combining Davis AI forecasting automation with Davis CoPilot, you can build an agent that answers key questions, such as which workloads are most resource-intensive, which resources show the most variance, and which require frequent scaling. This automation is capable of highly reliable forecasts, even when data points are limited. The automation interprets the data and emails actionable recommendations directly to you and anyone else who needs to stay informed.

To see this in action, watch the section of this video that explores predicting resource consumption.

Smart workflows that optimize query costs

Scheduling tasks can even help you keep costs lean and efficient. For admins or budget owners, staying within financial limits while maintaining performance is a constant challenge. In this example, we built a smart workflow that identifies the top 20 most expensive queries from the last 24 hours. Davis CoPilot analyzes each query and sends optimization recommendations directly to the query authors via email.

Smart workflow leveraging Davis CoPilot to recommend query optimizations tailored to your tenant
Figure 4. Smart workflow leveraging Davis CoPilot to recommend query optimizations tailored to your tenant
Example of an optimization suggestion delivered to the inbox of the query author
Figure 5. Example of an optimization suggestion delivered to the inbox of the query author

To implement this yourself, tailored to the most expensive queries executed on your tenant, go to our documentation

Conclusion: Adapt your workflows to any stage of your automation journey

These are just a few examples; the applications for it are endless. We’ve designed this workflow action to cater to your organization’s automation appetite. You may want to transform how you keep business stakeholders informed about what’s happening in your environment, leveraging Dynatrace’s highly accurate insights, which are translated into plain language and actionable next steps.

Alternatively, you may be ready to transition towards autonomous operations, where automation not only supports but also acts in a controlled and reliable manner. Davis CoPilot embedded into your workflows opens the door to your agentic journey.

Start your agentic journey and join the Davis CoPilot for Workflows Preview

Davis CoPilot for Workflows is available as a Preview. Sign up now and see how generative intelligence embedded into your workflows transforms your automation. Today, it helps you react faster, optimize more effectively, and collaborate seamlessly. Tomorrow, it will go even further: anticipating needs, orchestrating actions, and enabling truly autonomous reasoning.

Gain efficiency and have your agentic workflows do the work for you!

The post Write the future: Create your own agentic workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/write-the-future-create-your-own-agentic-workflows/feed/ 0
Real-time insights: Leverage Dynatrace observability capabilities within Kiro powered by AWS https://www.dynatrace.com/news/blog/real-time-insights-leverage-dynatrace-observability-capabilities-within-amazon-kiro/ https://www.dynatrace.com/news/blog/real-time-insights-leverage-dynatrace-observability-capabilities-within-amazon-kiro/#respond Mon, 24 Nov 2025 19:42:17 +0000 https://www.dynatrace.com/news/?p=72036 Amazon Q Developer CLI and Dynatrace

In today’s cloud-native environments, having real-time observability data at your fingertips is crucial. By integrating Kiro powered by AWS with Dynatrace, you can leverage powerful AI-assisted monitoring and troubleshooting capabilities directly in your development workflow. Kiro—which recently reached general availability—helps developers by bringing structure to AI coding with spec-driven development. When a developer needs to fix […]

The post Real-time insights: Leverage Dynatrace observability capabilities within Kiro powered by AWS appeared first on Dynatrace news.

]]>
Amazon Q Developer CLI and Dynatrace

In today’s cloud-native environments, having real-time observability data at your fingertips is crucial. By integrating Kiro powered by AWS with Dynatrace, you can leverage powerful AI-assisted monitoring and troubleshooting capabilities directly in your development workflow.

Kiro—which recently reached general availability—helps developers by bringing structure to AI coding with spec-driven development. When a developer needs to fix an issue, investigate an error, or optimize resource usage, it’s crucial they can analyze what happened just before the issue occurred and delve deeper into the infrastructure utilization of your applications in your cloud or container environment.

Unlock development productivity with live production insights

Developers typically face restricted access to production environments, being fully dependent on site reliability engineers (SREs) or operations teams to detect and report issues post-deployment, and provide them with the necessary information to fix an issue. This segmented workflow can result in delayed problem identification and resolution, an increased risk of failures in production, and reduced efficiency throughout the development lifecycle.

By connecting Dynatrace with Kiro, developers can access real-time insights from production environments, gain contextual information down to the root cause of an incident, and receive remediation proposals—all within their Kiro environment.

Figure 1: Dynatrace Agentic AI ecosystem for developers
Figure 1. Dynatrace Agentic AI ecosystem for developers

Kiro has a built-in Model Context Protocol (MCP) client that can be used to extend its capabilities to communicate securely and flexibly with external data sources and tools such as Dynatrace.

Let’s dig deeper into how to leverage this capability and provide Dynatrace’s unique insights to your development teams.

Step-by-step integration guide

Prerequisites

  • You’ll need a Dynatrace account. If you don’t already have one, you can start a free 15-day trial.
  • Kiro must be installed on your system.
  • You must have basic familiarity with AWS services and the Dynatrace platform.

Prepare integration with Dynatrace

First, you need to create a Dynatrace Platform Token, which is used to define Kiro access, and then add the required permissions for the Dynatrace MCP server.

Configure Kiro MCP Settings

The Kiro MCP configuration is managed through a JSON file. The interface supports two levels of configuration:

  • User-level: ~/.kiro/settings/mcp.json applies to all workspaces
  • Workspace-level: .kiro/settings/mcp.json is specific to the current workspace

You can apply the configuration using two different methods:

Method 1: Open the command palette (use Cmd + Shift + P on Mac or Ctrl + Shift + P on Windows/Linux), search for MCP and select Kiro: Open workspace MCP config (JSON) or Kiro: Open user MCP config (JSON), depending on whether or not you want to configure the settings for the workspace or user level.

Method 2: Alternatively, you can use the Kiro Panel. Open Kiro and select the Kiro ghost icon to open the left-side panel. Locate the MCP SERVERS section, select  Open MCP Config, and then start configuring the connection for the Dynatrace MCP Server.

Dynatrace specific settings

Note: Only add one of the following configurations, depending on whether you want to use the remote MCP server or the local MCP server. You can’t use both at the same time.

Using the remote MCP server

Use the following configuration. Replace $TENANT_ID with your Dynatrace environment ID. (You can find your environment ID in the URL of your Dynatrace environment — for example, https://<ENVIRONMENT_id>.apps.dynatrace.com/ui.) Then, replace $DT_PLATFORM_TOKEN with the ID of the Dynatrace platform token you created previously (for example, dt0s16.XXXXX).

{ 
"mcpServers": 
  { 
    "dynatrace": {
      "type": "http",
      "url": "https://$TENANT_ID.apps.dynatrace.com/platform-reserved/mcp-gateway/v0.1/servers/dynatrace-mcp/mcp",
      "headers": {
        "Authorization": "Bearer $DT_PLATFORM_TOKEN"
      },
      "tools": ["*"]
      }
  }
}

Connect the local MCP server

The configuration for the local Dynatrace MCP server can be added to the Kiro IDE using one-click installation or by following the manual configuration as shown below. Don’t forget to replace $TENANT_ID with the ID of your tenant.

{
  "mcpServers": {
    "dynatrace-mcp-server": {
      "command": "npx",
      "args": ["-y", "@dynatrace-oss/dynatrace-mcp-server@latest"],
      "env": {
        "DT_ENVIRONMENT": "https://$TENANT_ID.apps.dynatrace.com"
      }
    }
  }
}


Figure 2. Add Dynatrace via one-click installation (video)
Figure 2. Add Dynatrace via one-click installation (video)

Verify the integration

Once configured, you can use the Kiro chat to interact with Dynatrace through natural language conversations. Simply tell Kiro what you need, whether it’s investigating a critical incident, gaining insights into metrics, logs, or traces from your application, analyzing dependencies, or setting up automated alerts.

In the screenshot below, you can see in the lower left which capabilities are provided by the Dynatrace MCP Server. Beyond the standardized actions, such as listing active vulnerabilities or problems, querying data stored in Dynatrace, or creating a workflow, you can also interact with Davis CoPilot®, the Dynatrace natural language assistant.

Figure 3: Amazon Kiro with an established connection to Dynatrace.
Figure 3. Kiro with an established connection to Dynatrace.

Conclusion

This integration isn’t just another feature; it’s a fundamental shift in how Dynatrace integrates with your development workflow. It brings together the power of Kiro’s AI capabilities with the Dynatrace unified observability platform, allowing developers to access critical monitoring data and gain a real-time understanding of their production environments via natural language interaction.

Spend less time context switching and more time creating value for your customers. Start today and benefit from real-time insights, precise root cause analysis based on causal understanding or improved troubleshooting capabilities, and enhanced development workflows.

Explore how Dynatrace can integrate seamlessly into your development landscape using our remote MCP Server. If you’re interested in learning more about Kiro, have a look at their launch blog post or visit the documentation and dig deeper into how to connect with MCP Servers.

Gain efficiency by empowering Kiro with insights from Dynatrace.

The post Real-time insights: Leverage Dynatrace observability capabilities within Kiro powered by AWS appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/real-time-insights-leverage-dynatrace-observability-capabilities-within-amazon-kiro/feed/ 0
Boost cloud reliability: Dynatrace and Azure SRE Agent unite for autonomous operations https://www.dynatrace.com/news/blog/boost-cloud-reliability-dynatrace-and-azure-sre-agent-unite-for-autonomous-operations/ https://www.dynatrace.com/news/blog/boost-cloud-reliability-dynatrace-and-azure-sre-agent-unite-for-autonomous-operations/#respond Wed, 19 Nov 2025 17:19:49 +0000 https://www.dynatrace.com/news/?p=71938 Dynatrace and Azure SRE Agent

The integration of Dynatrace with Microsoft Azure SRE Agent establishes a new benchmark for cloud operations by leveraging AI-based root cause analysis and real-time production insights, alongside a comprehensive understanding of complex, large-scale IT environments. You can leverage the combined strengths of Dynatrace and Microsoft, enabling teams to resolve complex problems in large-scale IT environments […]

The post Boost cloud reliability: Dynatrace and Azure SRE Agent unite for autonomous operations appeared first on Dynatrace news.

]]>
Dynatrace and Azure SRE Agent

The integration of Dynatrace with Microsoft Azure SRE Agent establishes a new benchmark for cloud operations by leveraging AI-based root cause analysis and real-time production insights, alongside a comprehensive understanding of complex, large-scale IT environments. You can leverage the combined strengths of Dynatrace and Microsoft, enabling teams to resolve complex problems in large-scale IT environments more quickly and efficiently, and automate incident remediation, moving one step closer to driving autonomous operations across their complex environments.

In today’s cloud-first world, reliability isn’t just a goal; it’s a competitive advantage. As more services move online and LLM-powered assistants evolve into autonomous agents, maintaining the reliability, scalability, and cost-efficiency of critical systems becomes essential.

That’s why Dynatrace and Microsoft teamed up to integrate Dynatrace® AI-powered observability with the Azure SRE Agent. This collaboration allows site reliability engineers (SREs) to ensure seamless operations while proactively planning for future scalability and reliability requirements.

Transform your incident management through the combined capabilities of Azure SRE Agent and Dynatrace AI

Azure SRE Agent, introduced earlier this year, provides SREs and developers with the tools they need to increase the speed and efficiency of incident responses, diagnostics, and collaboration, allowing them to resolve problems quickly.

Automate monitoring of cloud environments
Figure 1. Automate monitoring of cloud environments

Seamlessly integrated with incident management tools such as ServiceNow, as well as the developer ecosystem, represented by GitHub Copilot or Azure DevOps, the agent runs in the background 24/7, learning and monitoring the health and performance of your cloud environment.

As a reliability assistant, Azure SRE Agent supports teams by efficiently diagnosing and resolving production issues. You can ask the agent questions in natural language, easily access clear and concise problem summaries, and coordinate incident workflows with integrated human-in-the-loop approvals.

Dynatrace enhances Azure SRE Agent’s troubleshooting and automation capabilities with advanced observability insights. By mapping topology, data, and business context, Dynatrace gains a comprehensive understanding and delivers production-accurate visibility across your entire IT system. This visibility feeds Dynatrace deterministic AI, allowing precise root-cause identification and impact analysis. All these insights are now seamlessly supplied to the Azure SRE Agent, equipping it with real-time production context and reliable root cause analysis.

This allows your teams to move beyond simply receiving alerts; teams are now provided with AI that acts, guides safe mitigations, and accelerates resolution within Azure-native workflows.

Gain efficiency across every stage of the incident lifecycle

Using the Model Context Protocol (MCP), the Azure SRE Agent is securely connected with Dynatrace. Whether a team member uses the agent to ask questions in plain natural language, or the agent interacts with Dynatrace directly—sharing insights, asking for real-time observability data, or root cause analysis, together with remediation steps—the close collaboration supports use cases across every stage of incident management, allowing you to:

  • Cut MTTR by automating routine runbooks and diagnostics, with safe, approved mitigation actions based on full context.
  • Reduce security risk by triaging vulnerabilities faster with production evidence, triggering guided fixes, and validating outcomes.
  • Accelerate delivery with contextual GitHub issues and PRs that include root cause, blast radius, and tests, minimizing issue reproduction time and rework.
  • Improve fix accuracy by correlating Azure and Dynatrace telemetry for precise root-cause and impact analysis.
  • Prevent incidents before they happen using real-time signals and historical trends to stop regressions and reduce toil.

Illustrating the value: Proactively detect and remediate security vulnerabilities

Let’s take a look at a concrete example, which we presented at Microsoft Ignite. Imagine you run a Java-based payroll app on Azure, and a new security warning (CVE) appears. Every second matters now, and there’s no room for error: you need the issue fixed quickly, without lots of back-and-forth between teams.

Schematic illustration – proactive vulnerability remediation with Dynatrace, Azure SRE Agent and GitHub
Figure 2: Schematic illustration – proactive vulnerability remediation with Dynatrace, Azure SRE Agent, and GitHub
  • Once the vulnerability is detected, Dynatrace automatically identifies the library that caused the vulnerability, opens a GitHub issue containing all relevant information, such as which parts of your app are affected, and informs Azure SRE Agent.
  • The SRE agent reviews the GitHub issue and requests additional information from Dynatrace via the MCP server, such as the number of users affected, how often it happens, which endpoints are involved, or which customers might be affected, to assess the scope and impact of the vulnerability.
Azure SRE automatically creates a GitHub issue with all the details.
Figure 3. Azure SRE automatically creates a GitHub issue with all the details.
  • After gathering all necessary details, the SRE Agent synthesizes the information and creates a new GitHub issue, assigning it to GitHub Copilot for remediation.
  • GitHub Copilot then takes action by updating the configuration and code in the GitHub repository to resolve the vulnerability automatically.
  • The pull request not only includes the necessary version changes but also includes documentation, highlighting all findings as well as how the issue was remediated, along with unit tests, to prevent the issue from recurring.

Demo of Azure SRE Agent thumbnail

Try the power of Agentic AI for incident resolution

Dynatrace delivers deep, causation-based insights into your live systems, now seamlessly integrated with Azure SRE Agent to elevate your incident management. With this integration, you can unlock:

  • Smarter detection and remediation: Deep contextual observability from Dynatrace, correlated with Azure telemetry, enhances issue identification and resolution across complex environments.
  • Automated operations: Routine runbook actions and diagnostic workflows can be automated, reducing mean time to repair and freeing teams to focus on innovation.
  • Proactive reliability: Continuous analysis of real-time and historical data identifies leading indicators of failure, allowing teams to prevent incidents before they impact customers.

Azure customers can now access Azure SRE Agent directly in the Azure portal. To connect Dynatrace with the agent and learn how to set up Dynatrace MCP Server, see Dynatrace Documentation.

The post Boost cloud reliability: Dynatrace and Azure SRE Agent unite for autonomous operations appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/boost-cloud-reliability-dynatrace-and-azure-sre-agent-unite-for-autonomous-operations/feed/ 0
Enhance the impact of Dynatrace Davis CoPilot with built-in observability https://www.dynatrace.com/news/blog/enhance-the-impact-of-dynatrace-davis-copilot-with-built-in-observability/ https://www.dynatrace.com/news/blog/enhance-the-impact-of-dynatrace-davis-copilot-with-built-in-observability/#respond Fri, 07 Nov 2025 18:10:53 +0000 https://www.dynatrace.com/news/?p=71730 Dynatrace Davis CoPilot

Ninety-five percent of Generative AI projects fail to deliver measurable value, and leaders are under mounting pressure to demonstrate that their AI investments are effective. Achieving this requires clear visibility into how and where AI is used, and the outcomes it’s driving. Dynatrace is setting a standard for observability across the AI stack, and we’re […]

The post Enhance the impact of Dynatrace Davis CoPilot with built-in observability appeared first on Dynatrace news.

]]>
Dynatrace Davis CoPilot
Update: We’ve launched Dynatrace Assist, our next-generation AI chat that goes far beyond answering questions.
Dynatrace Assist is the evolution of Davis CoPilot®.

Ninety-five percent of Generative AI projects fail to deliver measurable value, and leaders are under mounting pressure to demonstrate that their AI investments are effective. Achieving this requires clear visibility into how and where AI is used, and the outcomes it’s driving.

Dynatrace is setting a standard for observability across the AI stack, and we’re extending that same level of insight to our own AI tools. The new Davis CoPilot® Feature Adoption Dashboard utilizes the same telemetry that Dynatrace teams rely on to improve product quality. Assess effectiveness and optimize how Davis CoPilot supports productivity and decision-making, to move from experimentation to sustainable results.

Davis CoPilot, the Dynatrace platform’s LLM-powered assistant, helps teams work faster by leveraging the full context of their data on Dynatrace to deliver precise and actionable answers. The result is less time spent searching for information or onboarding users, and more time extracting the maximum value from the Dynatrace platform and achieving measurable outcomes.

IT and central team leaders typically offer Davis CoPilot to their users, with specific goals in mind that align with their organization’s broader AI strategy. Initiatives like these typically aim to achieve three key objectives:

  • Adoption and engagement: Ensure AI becomes an integral part of routine workflows, so value can scale across teams.
  • Productivity gains: Reduce manual effort, increase speed to insight, and improve the quality of outcomes.
  • Demonstrable business value: Connect usage to measurable results, such as reduced operational costs, faster incident resolution, or improved service levels.

Understand how AI is used and how it delivers value

To make these objectives measurable, you need visibility into how AI is adopted by your users, the purposes it serves, and whether it delivers the intended value. Only then can you identify where improvements are needed. The ready-made Davis CoPilot Feature Adoption Dashboard delivers this visibility out of the box, showing how Dynatrace generative AI features are used across your organization. Based on the provided metrics and insights, administrators and central teams can make data-driven adjustments.

Customers who opt in to Davis CoPilot can find the Feature Adoption Dashboard in the “Ready-made” category.
Figure 1. Customers who opt in to Davis CoPilot can find the Feature Adoption Dashboard in the “Ready-made” category.

Know how frequently and for what purpose Davis CoPilot is used, in real time

Gain real-time visibility into when and how regularly teams are using Davis CoPilot in their workflows. The dashboard highlights active engagement, query activity, and usage trends across your organization, helping you understand where Davis CoPilot delivers the most value and where additional enablement may be needed.

By analyzing usage patterns, you can identify high-performing teams, monitor overall adoption progress, and ensure employees are using Davis CoPilot effectively to achieve meaningful outcomes.

For a deeper analysis, break Davis CoPilot usage down further by skill:

  • Chat: Analyze chat interactions and workflow actions (currently in private preview), showing how users engage with Davis CoPilot to ask questions, troubleshoot issues, and automate routine tasks.
  • Natural language querying: Tracks how users convert everyday language into Dynatrace Query Language (DQL) commands, supporting faster data exploration for both technical and non-technical users.
  • Explain DQL queries: Shows how users rely on Davis CoPilot to interpret and summarize complex queries, making it easier to understand and build on existing work.
  • Document search: Tracks how users engage with AI-driven document retrieval for accelerated troubleshooting in the Problems app.
Get insights into AI usage and interaction success rates, split by AI skill.
Figure 2. Get insights into AI usage and interaction success rates, split by AI skill.

Track user experience and satisfaction

To determine whether Davis CoPilot delivers value, it’s important to measure not only usage but also the quality of user interactions and outcomes. The dashboard tracks execution times and success rates to demonstrate how well Davis CoPilot performs in real-world scenarios. This makes it easier to identify technical issues such as invalid DQL generation or prompts blocked by guardrails and content filters.

On the Failed NL2DQL interaction details tile, try out Open with... > Davis CoPilot on the response column to understand why the generated DQL is considered invalid.
Figure 3. On the Failed NL2DQL interaction details tile, try out Open with… > Davis CoPilot on the response column to understand why the generated DQL is considered invalid.

Additional user feedback adds context to these signals. Thumbs-up and thumbs-down reactions help indicate where users achieve the desired outcome and where they run into problems. When negative feedback clusters around similar prompts or skills, administrators can examine the failed prompts, identify common failure modes, and understand the conditions that lead to them. This supports targeted follow-up actions, such as improving internal guidance for AI usage, reinforcing enablement for specific teams, and surfacing actionable improvement requests to Dynatrace.

For example, if multiple users struggle with natural language queries for Kubernetes data, admins can review the failed prompts, provide best practices, and verify that these measures lead to higher success rates over time. You can even consider enriching your data by adding common synonyms with OpenPipeline. Nequi shared their story at Perform 2025.

Together, operational metrics and contextual feedback help organizations to quickly identify friction points and take concrete steps to improve user outcomes and overall satisfaction.

Get detailed insights on user satisfaction.
Figure 4. Get detailed insights on user satisfaction.

Optimize performance of AI-generated insights

The dashboard also provides transparency into the queries executed through Davis CoPilot, including query counts and the volume of data scanned. This helps you better understand the resource and cost impact of AI-generated insights across your environment. This level of visibility is critical, as many AI initiatives stall because teams lack the insight to understand the operational impact of increased usage.

By identifying data-intensive queries early, you can optimize performance, control cost exposure, and avoid unexpected resource spikes that can undermine confidence in scaling AI. Capabilities such as segment filtering or organizing data into dedicated buckets allow you to fine-tune data access based on organizational needs. This gives you the ability not only to monitor AI activity but also to adjust and govern it responsibly, ensuring Davis CoPilot remains efficient, controlled, and aligned with your broader business objectives.

Understand the number of executed queries and the scanned data volume.
Figure 5. Understand the number of executed queries and the volume of scanned data.

Make use of the full potential of Davis CoPilot

The Davis CoPilot Feature Adoption Dashboard equips you with the insights needed to scale Davis CoPilot responsibly, maximizing productivity gains while maintaining control. With clear visibility into usage, success rates, and operational impact, you can build a stronger foundation for continued AI expansion.

The dashboard is instantly available in the environments of Dynatrace customers who have enabled Davis CoPilot.

The post Enhance the impact of Dynatrace Davis CoPilot with built-in observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enhance-the-impact-of-dynatrace-davis-copilot-with-built-in-observability/feed/ 0
Unlocking productivity and trust: Dynatrace observability in NVIDIA AI Factory https://www.dynatrace.com/news/blog/unlocking-productivity-and-trust-dynatrace-observability-in-nvidia-ai-factory-environments/ https://www.dynatrace.com/news/blog/unlocking-productivity-and-trust-dynatrace-observability-in-nvidia-ai-factory-environments/#respond Tue, 28 Oct 2025 18:30:05 +0000 https://www.dynatrace.com/news/?p=71582 Davis CoPilot for NVIDIA

The NVIDIA Enterprise AI Factory addresses the rapidly evolving needs for AI infrastructure to support the rise of agentic AI. Since its launch, customers have leveraged this validated design to build agents by following structured methodology and recommended frameworks, which simplifies deployment and configuration while facilitating the implementation of AI factories in both on-premises and […]

The post Unlocking productivity and trust: Dynatrace observability in NVIDIA AI Factory appeared first on Dynatrace news.

]]>
Davis CoPilot for NVIDIA

The NVIDIA Enterprise AI Factory addresses the rapidly evolving needs for AI infrastructure to support the rise of agentic AI. Since its launch, customers have leveraged this validated design to build agents by following structured methodology and recommended frameworks, which simplifies deployment and configuration while facilitating the implementation of AI factories in both on-premises and hybrid cloud environments.

Dynatrace has been an integral part of this initiative. Dynatrace full-stack AI and LLM observability helps organizations move forward with confidence in building their AI and agentic AI initiatives.

Observable AI: Turn a black box into a glass box to build confidence

With the publication of comprehensive guidelines, it’s simpler than ever for Dynatrace customers to set up and start monitoring their full-stack NVIDIA enterprise AI infrastructure, including its key tiers and components. Covering the infrastructure layer from GPUs to Kubernetes, NVIDIA NIM microservices, NVIDIA NeMo, and other technologies up to the application layer, Dynatrace observability enables customers to confidently run and operate complex AI workflows on NVIDIA infrastructure.

NVIDIA Enterprise AI Factory for Agents including components covered by ecosystem partners (such as Observability). Picture taken from NVIDIA Enterprise AI Factory - Design Guide White Paper
Figure 1: NVIDIA Enterprise AI Factory for Agents, including components covered by ecosystem partners (such as Observability). Picture taken from NVIDIA Enterprise AI Factory – Design Guide White Paper

In parallel, Dynatrace has worked to significantly advance our AI and LLM observability offering by introducing the following:

Dynatrace AI Observability
Figure 2: Dynatrace AI Observability

These improvements address challenges such as missing observability insights, scale, sovereignty, and trust. This empowers organizations to operationalize AI by building trust and monitoring guardrails; providing analytics capabilities to detect user-facing issues; helping SREs and AI-native engineers maintain performance, reliability, and security; and reducing cost across the agentic, AI, and LLM stack.

Privacy and security lead the way to scaling AI with confidence

AI is delivering significant productivity improvements, with 66% of senior executives reporting positive trends in productivity, according to PwC’s AI Agent Survey. This momentum is driving the demand to manage AI expenditures, enhance the decision-making quality of agents, and optimize development through visibility into AI components’ behavior in production environments — from pilot projects to full-scale operations.

However, sensitive data considerations and strict compliance requirements often impede progress, preventing organizations from fully realizing the benefits of AI adoption. As enterprises prioritize data privacy, regulatory compliance, and data sovereignty, there is an increasing need for high-performance NVIDIA AI infrastructure alongside frameworks designed to preserve control, trust, and autonomy in AI development.

In a recent blog on sovereign AI, NVIDIA shares strategies for nations and enterprises to develop AI factories that uphold local governance, security, and cultural values. Combining such factories with the Dynatrace advanced observability solution enables organizations to operationalize AI at scale — building secure and scalable agents, deployed on premises or in hybrid environments.

From privacy needs to public-sector requirements: NVIDIA AI Factory for Government

At NVIDIA GTC Washington, D.C. today, NVIDIA AI Factory for Government was announced, in support of the needs for regulated environments to drive AI initiatives. The U.S. Office of Management and Budget’s decision to establish scorecards for agencies’ AI maturity and management is in line with a 2024 Gartner Research forecast that more than 60% of government organizations will be prioritizing their investments in business automation by 2026 — up from 35% in 2022. The NVIDIA AI Factory for Government is a full-stack, end-to-end reference design that brings the power of reasoning AI to federal organizations. It helps organizations unlock productivity gains just like it does for enterprises, from service delivery to threat detection and day-to-day operations.

Built on the experience of deploying internal AI factories, the reference design offers guidance for deploying agentic AI, physical AI, and high-performance computing workloads on premises and in hybrid cloud environments, while meeting the compliance needs of federal and other secure organizations. The NVIDIA AI Factory for Government reference design includes NVIDIA Blackwell accelerated computing and NVIDIA networking, NVIDIA-Certified Systems, NVIDIA AI Enterprise software, NVIDIA Nemotron open models, and third-party software from AI leaders, all validated by NVIDIA.

Dynatrace delivers trusted observability and automation for regulated environments

Dynatrace has always been committed to supporting the public sector and other industries with regulatory requirements by providing customers with capabilities to control data flow through its lifecycle and manage sensitive data from ingestion to deletion, as well as global deployment options to meet data residency requirements, configurable retention times for different data types and use cases, unique encryption keys for customer’s stored data, and more.

Our dedication is reflected in customers’ success stories from regulated industries, as well as a growing list of global and local certifications, such as ISO 27001, SOC 2 Type II, CSA STAR 2, ENS, Tisax, and others. Find out more about our certifications and supported compliance frameworks in our Trust Center. For organizations also navigating evolving sovereignty requirements, our approach to digital sovereignty demonstrates how Dynatrace combines technical innovation with policy alignment to deliver trusted solutions globally.

Benefit from full-stack observability for end-to-end validated design

Dynatrace observability with the NVIDIA AI Factory for Government reference design enables organizations to accelerate the deployments of their AI agents and applications for federal and enterprise environments, and benefit from real-time, AI-powered insights.

These benefits range from improved scalability and performance to reduced complexity and total cost of ownership by simplifying processes, mitigating deployment risks to improved data security and compliance.

Visit the Dynatrace Playground to experience the possibilities of AI and LLM observability, and discover how Dynatrace is accelerating enterprise AI at scale.

Dynatrace and the Dynatrace logo are trademarks of the Dynatrace, Inc. group of companies. All other trademarks are the property of their respective owners.

The post Unlocking productivity and trust: Dynatrace observability in NVIDIA AI Factory appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/unlocking-productivity-and-trust-dynatrace-observability-in-nvidia-ai-factory-environments/feed/ 0
Dynatrace and Atlassian deliver agentic AI that transforms end-to-end incident management https://www.dynatrace.com/news/blog/dynatrace-and-atlassian-delivering-agentic-ai-that-transforms-your-end-to-end-incident-management/ https://www.dynatrace.com/news/blog/dynatrace-and-atlassian-delivering-agentic-ai-that-transforms-your-end-to-end-incident-management/#respond Wed, 08 Oct 2025 05:45:57 +0000 https://www.dynatrace.com/news/?p=71166 Dynatrace and Atlassian

When incidents occur, engineers and incident managers often lack the production visibility they need to fully understand the underlying issues and act quickly. Most tickets fail to include details about severity, impact, or next steps. This forces teams to waste time jumping between tools and manually stitching data together, delaying recovery and driving up costs.

The post Dynatrace and Atlassian deliver agentic AI that transforms end-to-end incident management appeared first on Dynatrace news.

]]>
Dynatrace and Atlassian

The new Dynatrace integration with Atlassian solves this by embedding real-time production insights directly into incident management processes. Teams gain instant visibility into what’s happening, who’s impacted, and the actions required to resolve issues faster — all without the need to switch tools.

Dynatrace uniquely detects problems in real time by understanding topology, data context, and dependencies across your entire digital ecosystem. Incidents are automatically tied to underlying root causes, giving teams a complete, production-accurate, “live” picture of problem details, severity, and impact.

Dynatrace insights are now accessible in Jira Service Management through human-readable summaries generated by Atlassian Rovo. By bringing production context directly into Jira, Confluence, and Jira Service Management, you’ll accelerate response times and significantly reduce mean time to resolution (MTTR).

At Dynatrace, context is our mantra, sitting at the core of everything we do. This means more than just data enrichment: Every piece of data is automatically contextualized, and dependencies are mapped to reveal the full picture. However, context also means delivering the right data exactly when and where you need it. To do just that, Dynatrace is bringing these insights directly into Atlassian. This is not just limited to IT service management (ITSM). You can get access to contextualized insights directly within an IDE as described in our latest blog post about the new Dynatrace  MCP Server.

Diagnose faster with context from production at your fingertips

Most incident tickets land on an engineer’s desk with little more than a timestamp, a vague description, or a user complaint. They rarely reveal the severity of the issue, which systems are affected, or what might be causing it. This lack of context in an ITSM workflow forces teams to spend unnecessary time digging through monitoring dashboards, chasing logs, or switching between tools just to piece together the basics of the problem.

Instead of getting frustrated, you can now instantly ask the Rovo Ops agent to identify anomalies that occurred around the incident timeframe. The agent queries Dynatrace via our MCP Server and returns the findings directly in the same browser window.

Get problem insights from Dynatrace directly delivered in the ticket context.
Figure 1. Get problem insights from Dynatrace directly delivered in the ticket context.

Having contextual details and alerts available directly in the ticket context means you gain immediate clarity into health, what’s wrong, the impact, and the evidence. This leads to faster diagnosis and quicker recovery, while also reducing unnecessary escalations of already-known or related issues, ensuring internal resources aren’t tied up with redundant work.

Remediate smarter with AI-driven root-cause analysis and automation

Once an incident is identified, the Rovo Ops agent utilizes Dynatrace production insights, which accelerate triage and root-cause analysis for incident managers, pinpointing the actual root cause in real time and delivering a higher level of insight and accuracy.

Rovo can now pull in Dynatrace Causal AI insights, including the precise root cause and blast radius of the issue, and combines these with Jira Service Management incident and change history. With Dynatrace contextual intelligence, Rovo delivers fact-based, AI-generated problem summaries and clear remediation recommendations, outperforming the guesswork of pure GenAI approaches.

From this point, just follow the remediation recommendation and trigger a suggested automation action in Jira Service Management, or ask follow-up questions for clarification.

Perform contextual analytics with follow-up questions
Figure 2. Perform contextual analytics with follow-up questions

Learn for the future with automated post-incident reviews

The job isn’t finished after an incident is mitigated and marked resolved in Jira Service Management, as you still need to capture what happened and determine how to prevent its recurrence. Instead of spending hours on manual write-ups, Rovo automatically triggers the post-incident review (PIR) process.

In the auto-generated PIR, Rovo surfaces all of the relevant details and history, from the root cause to detected anomalies, all of which are enriched by Dynatrace AI-driven insights. This provides a complete, time-ordered view of the incident, which is combined with Jira Service Management context attributes like assignees, tags, outage duration, and related change logs. With this context, the agent generates a draft PIR. Inside the PIR, you’ll find monitoring charts showing the status before, during, and after the incident, a clear summary of the cause, and a pre-filled prevention plan. All that’s left for you to do is review, refine, and finalize the PIR.

The automatically documented PIRs act as built-in retrospectives, helping teams continuously mature their operations. They also feed insights back into Rovo to sharpen its future recommendations.

Transform how you work, beyond incident management

These are just a few examples of what’s now possible through the extended Dynatrace + Atlassian integration. We’ll continue to explore deeper integrations to make your troubleshooting journey even more efficient in the future.

Imagine directly following up on investigations from within Rovo, with seamless drill-downs into Dynatrace® Apps, or surfacing related post-mortem information and runbooks stored in Jira or Confluence to SREs when investigating an issue in Dynatrace.

And the potential impact goes well beyond incident management. By bringing reliable, real-time production truth into daily workflow and connecting that truth directly to business outcomes, more teams and roles can fundamentally transform the way they work, harnessing the full power of agentic AI.

  • Get instant release validation: Developers can query Rovo for pre- and post-deployment failure rates, SLOs, and outcome metrics, allowing them to release with confidence, roll back faster when needed, and validate hypotheses with real data.
  • Make decisions based on outcomes: Product managers can ask Rovo or Davis CoPilot® to analyze the impact of a new feature or release by investigating KPI shifts such as user engagement or a drop in check-outs.
  • Speed up triage based on business impact: Support engineers working on Jira tickets see Dynatrace insights related to the root cause, blast radius, affected applications, and services. These insights are enriched with further details on user and business impact, allowing engineers to perform instant impact analysis before assigning tickets.
  • Run smarter daily stand-ups: Development teams receive ready-made summaries, including exceptions, user analysis, and deployment reports from the last 24 hours, providing relevant insights into what’s actually happening in production.

Start benefiting from deeper integrations with Dynatrace as your trusted foundation for agentic AI

Dynatrace delivers a deep, causation-based understanding of your live digital systems, providing the precise, reliable insights that enterprises can trust as a foundation for agentic AI.

Ready to see how Dynatrace and Atlassian work together and benefit from adopting agentic AI concepts? Then dig deeper into the new possibilities using our remote MCP Server and experience how real-time production context makes your operations more efficient.

See our documentation to learn more about how to connect the Dynatrace MCP Server.

Gain efficiency by empowering your AI agents with insights from Dynatrace.

The post Dynatrace and Atlassian deliver agentic AI that transforms end-to-end incident management appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-and-atlassian-delivering-agentic-ai-that-transforms-your-end-to-end-incident-management/feed/ 0
Sky-high developer productivity with Dynatrace MCP and GitHub Copilot https://www.dynatrace.com/news/blog/sky-high-developer-productivity-with-dynatrace-mcp-and-github-copilot/ https://www.dynatrace.com/news/blog/sky-high-developer-productivity-with-dynatrace-mcp-and-github-copilot/#respond Fri, 03 Oct 2025 17:12:09 +0000 https://www.dynatrace.com/news/?p=71251 Agentic AI

Unless you’ve been living under a rock this past year, you’ve heard the buzz about how AI is changing the day-to-day lives of developers, whether it’s using LLMs for vibe coding or adopting agentic AI concepts to improve productivity. In particular, the introduction of Model Context Protocol (MCP) as a standard for connecting AI agents […]

The post Sky-high developer productivity with Dynatrace MCP and GitHub Copilot appeared first on Dynatrace news.

]]>
Agentic AI

Unless you’ve been living under a rock this past year, you’ve heard the buzz about how AI is changing the day-to-day lives of developers, whether it’s using LLMs for vibe coding or adopting agentic AI concepts to improve productivity. In particular, the introduction of Model Context Protocol (MCP) as a standard for connecting AI agents with other agents or tools has led to exponential growth in the number of solutions connecting AI coding assistants with APIs, databases, and more. Last month, GitHub, publisher of the prominent AI coding assistant GitHub Copilot, announced its new MCP registry as a place where developers can find links to MCP servers. This registry solves the challenge of numerous registries, repos, or community threads listing the same MCP servers.

We’re excited to share that the GitHub MCP registry now includes Dynatrace MCP, so developers can integrate Dynatrace observability and security analysis directly into their workflows.

In this blog post, we’ll explore how developers can use Dynatrace MCP together with GitHub Copilot to streamline troubleshooting, enhance security, and boost productivity—without ever leaving their IDEs.

Need to troubleshoot an issue? Dynatrace MCP has the answers

One of the biggest challenges in troubleshooting and observability is knowing where to look for missing data when issues arise in production. It might sound odd, but when you consider the sheer number of tools, applications, environments, and layers of source code developers need to navigate, it’s not at all clear where they should look for quick answers. No wonder onboarding time is such a major factor in engineering productivity.

Your day as a developer might start with a complaint about something not working correctly. You receive a Jira ticket with a cryptic error message that includes a link to a service health dashboard or a notification in the problem list. You start rummaging through the logs and dashboards to find answers outside your IDE, with little context. This greatly increases your cognitive load and delays the problem’s resolution.

Developers who use Dynatrace, though, don’t have to manually dig through logs and dashboards. They get clear summaries, root cause information, and all relevant data in the context of the affected service. Thanks to the updated problem flow, it’s easy for them to identify the root cause and remediate it.

Dynatrace identifies the likely root cause, performing failure analysis in the context of the affected service.
Figure 1. Dynatrace identifies the likely root cause, performing failure analysis in the context of the affected service.

Context is key

Now imagine a scenario in which your DevOps team uses MCP as their standard for integrating external tools and insights. As a developer, your instinct is to jump straight into the source code, so why not share as much context as possible right there? Utilizing GitHub Copilot with the Dynatrace MCP integration in place, you get all relevant data from your production environment and quickly isolate the problem, giving you the ability to:

  • Get information on the root cause
  • Query related logs, metrics, and traces
  • Get real-time data from all environments, including production
  • Leverage additional gathered insights, such as CPU and memory profiling, from the affected service
  • Get further context using query patterns, such as “group logs by customer impact”

And the best is, you don’t need to know how or where to look—Dynatrace handles all this for you automatically. And there’s more: you can even ask Dynatrace for remediation suggestions.

Dynatrace identifies the problem root cause as an arithmetic exception, all via LLM and the MCP protocol in the developer VS Code IDE. (video)
Figure 2. Dynatrace identifies the problem root cause as an arithmetic exception, all via LLM and the MCP protocol in the developer VS Code IDE. (video)

What’s happening behind the scenes

When you type a natural language prompt into the GitHub Copilot chat, GitHub’s MCP client establishes a connection with the Dynatrace MCP server, which then connects to Dynatrace. The LLM converts the prompt into a context-aware call and ultimately transforms it into a DQL query that it executes on Dynatrace Grail® data lakehouse, which retrieves the necessary data.

Simplified communication flow.
Figure 3. Simplified communication flow.

Need to fix a security issue? Dynatrace MCP has the answers

By integrating security testing practices earlier in the Software Development Lifecycle, the Shift-left principle has brought security responsibilities into the developer’s world, and they’re here to stay. Developers often find themselves reacting to alerts from SREs, manually auditing code, or relying on static analysis tools that frequently miss context-specific vulnerabilities.

You can instantly access insights like:

  • Vulnerabilities in your code, including open source and third party vulnerabilities
  • Recommendations for fixing an issue
  • Proactive checks tailored to your current coding context

All of this happens without pulling developers away from their flow. MCP allows for a smarter, more proactive approach to security—one that’s embedded in their development process.

Dynatrace shares details of known vulnerabilities related to the affected component/service.
Figure 4. Dynatrace shares details of known vulnerabilities related to the affected component/service.
Dynatrace assesses the query pattern and suggests a fix.
Figure 5. Dynatrace assesses the query pattern and suggests a fix.

Need to create new code? Dynatrace is here to help

Writing new code with Dynatrace allows developers to look ahead proactively. By typing natural-language, conversational questions like “Have I bumped into CPU limits?” or “What is my CPU usage? Is it too high?” developers can identify potential bottlenecks and performance issues before they become real problems, ultimately delivering higher quality code.

Developing new code or optimizing a service in an existing app brings even more complexity. Such work involves identifying inefficient API usage, reducing unnecessary load, improving performance, and assessing the potential impact of the new code so as to minimize deployment issues.

Need to verify recent CICD builds? Dynatrace helps you shift left

As a developer, you want to catch build issues early—before they snowball into deployment delays or production incidents. With Dynatrace MCP, you can type questions like:

  • “What failed in the last build?”
  • “Are there any performance regressions tied to this commit?”
  • “Did this deployment introduce any anomalies?”

By integrating Dynatrace into your CI/CD pipeline, you gain real-time visibility into build health, test coverage, and deployment impact. This means faster feedback loops, fewer surprises, and higher delivery quality. Dynatrace helps you to shift from reactive debugging to proactive delivery assurance—all through natural language interactions.

Get started with Dynatrace MCP

The Dynatrace MCP server is available as a community-supported open source project. To familiarize yourself with all it can do, visit the Dynatrace MCP project. in our GitHub repository. There you’ll find all the necessary documentation to guide you through the setup process and explain all the available capabilities.

Explore how developers use Dynatrace MCP with GitHub Copilot to streamline troubleshooting, enhance security, and boost productivity without ever leaving their IDEs.

The post Sky-high developer productivity with Dynatrace MCP and GitHub Copilot appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/sky-high-developer-productivity-with-dynatrace-mcp-and-github-copilot/feed/ 0
Let the problem guide you: Smart remediation with Dynatrace https://www.dynatrace.com/news/blog/let-the-problem-guide-you-smart-remediation-with-dynatrace/ https://www.dynatrace.com/news/blog/let-the-problem-guide-you-smart-remediation-with-dynatrace/#respond Wed, 13 Aug 2025 15:38:50 +0000 https://www.dynatrace.com/news/?p=70376 Problem alert dashboard

Imagine a friend pings you, asking to meet them at a new restaurant whose location you don’t know. Would you simply step outside and start walking, hoping that you’ll somehow find the restaurant by luck? More likely, you’ll open a navigation app on your mobile device, search for the address, and set your route. A critical incident in your IT stack is no different. You shouldn’t have to guess your way through the dark alleys of logs, traces, and warnings, asking passersby for assistance, just to be bounced from one guess to the next.
This is where Dynatrace comes in: curating what’s most relevant, guiding engineers to the next best steps, and confidently guiding them from alert to resolution.

The post Let the problem guide you: Smart remediation with Dynatrace appeared first on Dynatrace news.

]]>
Problem alert dashboard

At Dynatrace, we’ve long led the way in root cause analysis, automatically detecting when something’s wrong and identifying why. These AI-powered insights form the intelligent backbone of how Dynatrace guides engineers through complex, entangled problems, turning a confusing cityscape of data into a focused remediation journey.

In this blog post, we follow Omar, a fictitious SRE working on Astroshop, the Dynatrace OpenTelemetry demo application. Omar works with Sophie, a developer, to resolve a critical failure.

Dynatrace guides them, and any team, through every step of the incident lifecycle, with speed, precision, and confidence. Here’s how:

  • Problem diagnosisDynatrace automatically packages the failures into a single correlated problem.
  • AlertingOmar receives a rich, actionable notification, right where he works, with root cause analysis (RCA), ownership, and deep links already in place.
  • TriageThe problem details page serves as a “triage and remediation command center,” showing Omar what’s impacted, who’s affected, and which issues to prioritize for remediation.
  • HandoffFrom within the problem context, a guided Jira ticket populated with RCA insights, logs, and deep links can be created with one click.
  • Guided investigationsProblems view lists all incidents related to the context of the problem.
  • Remediation – Using Live Debugger and real-time metrics, Sophie fixes the issue and sees immediate confirmation that the fix worked.
  • Reinforcement and learning for the future – The fix is automatically documented in a troubleshooting guide, turning today’s incident into tomorrow’s intelligence.

Problem diagnosis: multiple incidents and a stream of signals tied into one cohesive problem

The Astroshop online store recently began experiencing errors when customers attempted to pay using American Express credit cards. Error messages prevent users from completing their purchases. Dynatrace seasonal baselining detected this service health issue, reduced alert noise by clustering all related events into a single problem, and tied all the relevant telemetry (logs, metrics, and traces) to the affected users, infrastructure, and SLOs.

By determining the ownership of the service, based on the Kubernetes label, the problem is automatically routed to the right team: Omar’s SRE team.

Figure 1. The Problems app clusters related events into one cohesive problem.
Figure 1. The Problems app clusters related events into one cohesive problem.

Contextual alerting

Meanwhile, Omar, the on-call SRE, receives the notification directly in the context where Omar and his team work, Slack. The notification reads, Failure rate increase in payment service.

Figure 2. Receive alert notifications in the tools where you work.
Figure 2. Receive alert notifications in the tools where you work.

Notifications like these can, of course, also show up in JIRA, ServiceNow, or other tools, depending on where your team works and wants to be notified. The notification provides a link to the full context, where Omar can begin remediation straight away. The notification identifies the root cause and service owners. It also provides directly accessible opinionated drilldowns, removing the need to open different tools. Omar starts his work by investigating the context and responding to the incident.

Triage: understand what, who, and how bad

Selecting the View problem button in the Slack notification brings Omar to the related Problem page, his triage and remediation command center.

At a glance, Omar can answer the following questions: How bad is this problem? Who and what is affected? How long has the problem existed? And, what is my best course of action?

Figure 3. Problem details: triage key impact and root cause at a glance.
Figure 3. Problem details: triage key impact and root cause at a glance.

First question: How bad is it, and who’s affected?

From the Problem Incident header, Omar immediately understands that the issue is blocking revenue: 600 users failed to check out their purchase. Furthermore, he sees that there is one frontend and six services are affected, and the error puts five SLOs at risk. Clicking one of the tiles allows the engineer to drill down into further details, offering clear “what to look at next” calls to action.

While unfamiliar with the error, Omar understands that this issue is severe and must be prioritized. Before taking action, he wants to understand the current situation compared to the normal state, which leads to his next question.

Second question: How long has this been going on, and how many checkouts are failing?

The event chart makes it clear: the incident has been ongoing for thirty minutes. Omar sees exactly how the event has evolved over time. The baseline-aware metrics allow him to distinguish between normal fluctuations and true anomalies. Here, the current behavior deviates from normal, with a significant increase in errors, evidently deviating from the baseline, where no errors are expected over that same period, around the same time of the day.

Meaning, for the last thirty minutes, many more people were unable to complete their purchases than usual. The issue is ongoing. In the blink of an eye, he has validated the severity of the incident.

Third question: What’s broken?

The deployment view on the left provides a visual breakdown of the failure by component and cluster. The root cause engine correlated all contributing events and factors, and pinpointed the exact service at fault: the culprit, a Kubernetes workload associated with the “payment service.”

Fourth question: Who can fix it?

This is now easy. The root cause, clearly labeled within the affected infrastructure, clarifies ownership. Omar sees that Sophie, the developer from Team Finance, owns the failing service.

Swiftly guided through the incident, Omar can confidently hand over the issue to the right person.

Handoff: effortless and precise handoff to the responsible team

Omar kicks off a handoff directly from the problem details page within Dynatrace. With a single click, Omar creates a Jira Issue containing all relevant information.

The Jira ticket is populated with the problem identifiers, making it easy to understand the affected services and components at stake without toggling between tools. No ping-pong escalations, just clear, confident delegation with minimal overhead.

Sophie, the owning developer, receives the Jira notification and opens the ticket to find everything she needs: a brief summary with direct links to the problem view, failure analysis, and Live Debugger.

Guided investigation in problem mode

While guided through the next best course of action, Sophie chooses to analyze the failure further using the Services app. This opens problem mode, a visual assistive layer inside the Services app, scoped to the incident’s impact window. It pre-filters all logs, traces, and warnings based on:

  • The scope of the incident
  • The services and infrastructure involved
  • Event severity and user impact
Figure 4. Problem mode in the Services app: Investigate only what’s relevant in the automatically scoped context.
Figure 4. Problem mode in the Services app: Investigate only what’s relevant in the automatically scoped context.

This view automatically prioritizes what matters, highlighting the signals that explain why the failure occurred, not just what broke. Instead of wrestling with contextless, stale, and unrelated logs, Sophie sees error logs and telemetry tied to the exact timeframe of the checkout errors. It cuts to the relevant error logs linked to the failure, including the exact error message. Sophie sees immediately that this is a bug in the code, specifically a JavaScript error at line 73.

Figure 5. Drill down to the root cause on the code level.
Figure 5. Drill down to the root cause on the code level.

Remediation: fix rapidly, validate instantly

Sophie has enough context to fix the failure. She opens Dynatrace Live Debugger, navigates to where the error logs occurred, and sets a non-breaking breakpoint at line 22, where she knows the error originates. Once Live Debugger captures a snapshot, Sophie is able to quickly identify the culprit: a mismatched character, American-Express vs American_Express, is causing the checkout service to bounce.

Figure 6. Start debugging at the exact line where the error occurs with live production data in Live Debugger.
Figure 6. Start debugging at the exact line where the error occurs with live production data in Live Debugger.

Sophie patches the issue and deploys the fix. A couple of minutes later, when having a look at the business metrics dashboard, she validates the remediation action, observing in real-time that checkout failures are gone, and revenue is flowing again.

Document the issue for future reinforcement and learning

Once the fix is deployed, Sophie and Omar go to Troubleshooting in the problem details. Together, they jot down details of the errors and their cause: the key mismatch between American-Express and American_Express, and how it was resolved. Sophie adds the link to the dashboard she used for validation, next to the incident details, automatically prefilled by Davis® AI.

Figure 7. Davis AI automatically recommends relevant documents with actionable troubleshooting steps for this problem.
Figure 7. Davis AI automatically recommends relevant documents with actionable troubleshooting steps for this problem.

Once this is done, Dynatrace automatically

  • Indexes the incident’s metadata and telemetry
  • Prefills it with Incident details linked to the RCA
  • Links it to similar incidents using graph-based AI and vector search

Weeks or months later, when a similar checkout spike appears, Dynatrace remediation intelligence automatically surfaces the relevant troubleshooting guide, providing necessary context and remediation guidance.

Rather than starting from scratch in the future, users benefit from a living knowledge base that not only accelerates problem resolution by reducing repeated investigation but turns every fix into team intelligence.

Try out the Dynatrace problem-remediation journey today

Dynatrace doesn’t just help you detect problems; it guides your teams calmly and confidently from >notification to resolution. Start now, eliminating repetitive war rooms, and try it out yourself:

Explore the problem remediation journey in the Dynatrace Playground.

The post Let the problem guide you: Smart remediation with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/let-the-problem-guide-you-smart-remediation-with-dynatrace/feed/ 0
Remediation intelligence: Accelerate MTTR with AI-powered context and knowledge https://www.dynatrace.com/news/blog/remediation-intelligence-accelerate-mttr-with-ai-powered-context-and-knowledge/ https://www.dynatrace.com/news/blog/remediation-intelligence-accelerate-mttr-with-ai-powered-context-and-knowledge/#respond Wed, 13 Aug 2025 15:38:30 +0000 https://www.dynatrace.com/news/?p=70391 Remediation intelligence

There’s a hidden barrier to deep understanding of your revenue, performance, and bottom line that has nothing to do with tools or telemetry. It’s organizational knowledge, the remediation know-how that’s scattered across hundreds of documents, private notebooks, dashboards, and in the minds of your engineers. This implicit, loosely documented knowledge is invisible to machines and rarely timely for humans. So, when a high-priority incident hits, this knowledge gap fills your war rooms with engineers tasked with solving problems they don’t own.

The post Remediation intelligence: Accelerate MTTR with AI-powered context and knowledge appeared first on Dynatrace news.

]]>
Remediation intelligence

Delivering reliable, business-critical applications to production is more complex than ever. The growing complexity and granularity of modern software systems and the trend to shifting more and more responsibilities to development teams (shift left) lead to increased pressure on your development teams.

In fact, traditional development teams now have a wider set of responsibilities; they must be highly skilled and educated in a multitude of domains. The responsibilities of these teams now range from specification, planning, testing, risk assessment, cost estimates, and deployment, to load testing, UI testing, integration testing, and on-call responsibilities for their software services.

While development teams need to be literate in all new technology stacks, cloud resources, and quality assurance methods, they also need to work with numerous tools.

The invisible bottleneck in your remediation process

The whole shift-left trend has pushed operational responsibility closer to development, expanding workloads to include on-call rotations, more frequent deployments, and growing expectations around uptime. When incidents occur, dozens of engineers are dragged into war rooms to perform analysis of the underlying root causes.

Reducing the Mean Time to Repair (MTTR) is essential to business continuity and customer satisfaction. Without access to the right knowledge at the right time, even skilled teams lose momentum. In high-pressure situations caused by critical incidents, it becomes more important than ever to ensure that all relevant information is shared with every role involved. This requires the most automated and intelligent methods available for effectively collecting, analyzing, and distributing information.

The on-call engineer’s journey

Take Omar, an SRE; an automated voice jolts him awake to summon him into a war room in the middle of the night after a routine update caused a spike in failed requests for a cloud-based payment service. It’s a P1 incident. With each passing minute, merchants are losing value, support requests surge, and customers are complaining. Dozens of caffeinated engineers are already in the war room.

Logs point to timeouts, but this is just a symptom; the real problem is somewhere else. Reading every message and document would take hours, so Omar scans for summaries, key findings, and any mention of his team’s services. Several hypotheses have already been tested, and one points to a potential issue in a backend service Omar’s team owns. Meanwhile, customer complaints are beginning to surface from other time zones. The payment service is failing, and customer success managers are growing increasingly anxious.

And while a similar outage has happened before, Omar is not able to find any documentation or insights into how to remediate the issue.

Why organizational knowledge doesn’t scale

When remediation history lives in documents, scattered across teams, formats, and platforms, engineers waste time searching instead of solving. Even well-documented incidents don’t prevent recurrence if they’re disconnected from future incidents. Even centralized platforms like Backstage don’t help if they can’t surface the right guidance at the right time.

Without a way to systematically identify, reuse, and scale this knowledge, it remains reactive. That’s not just inefficient, it’s a blocker to building intelligent automation and truly preventative operations. This is the hidden obstacle, silently inflating your MTTR, buried knowledge that costs time, delays response, and drains focus from what really matters. This is what remediation intelligence solves.

Introducing remediation intelligence

Dynatrace has a long history of providing DevOps teams with AI-driven tools for anomaly detection, root cause identification, and incident impact assessment in complex application environments. Over the past decade, it has contributed to reducing mean time to resolution (MTTR) by learning application behavior and analyzing dependencies in real time.

Building on this foundation, Dynatrace launched remediation intelligence, which adds an additional element to the incident response process from alert through resolution. It assists engineers during remediation by combining Davis® AI root cause and impact analysis with input from global community knowledge and internal expertise. It integrates data such as logs, metrics, traces, and topological context into a single view and offers support for documenting post-incident reviews.

Figure 1. The problems page displays all important information, allowing you to directly access all incident-relevant error logs.
Figure 1. The problems page displays all important information, allowing you to directly access all incident-relevant error logs.

Close the knowledge gap:  Embedded troubleshooting knowledge

What truly sets Dynatrace remediation intelligence apart is its ability to proactively surface relevant internal knowledge at the moment it’s needed most. It adds an AI-guided assistive layer to the Problems app that brings implicit, organizational knowledge directly into the flow of incident response. Once a problem is detected, Davis AI scans the historical data, surfacing past remediation playbooks, troubleshooting dashboards, and notebooks that were used to resolve similar issues.

With troubleshooting guides, we introduce a context-aware guidance system, built on Davis AI, that connects current incidents with prior resolution paths. It makes organizational knowledge queryable, remediation patterns reusable, and every responder effective, even when they’re solving an unfamiliar issue.

Figure 2. Review related documents from similar past incidents.
Figure 2. Review related documents from similar past incidents.

Remediation intelligence surfaces the most relevant remediation insights

When an incident occurs, Dynatrace excels at automatically analyzing and surfacing technical insights. It collects and organizes all relevant signals—logs, metrics, traces, and topology—into a single coherent problem. No fragmented alerts. No disconnected symptoms. Just one structured, AI-curated incident view. In parallel, Dynatrace AI scans all documents marked as troubleshooting-relevant. Using advanced semantic search and vector embeddings, it ranks and surfaces the most relevant past incidents, dashboards, notes, and postmortems, based on their similarity to the current problem. This is not just keyword matching; Dynatrace understands patterns, failure modes, and system relationships, surfacing ranked, high-similarity incidents in the problem view.

Figure 3. Example of a troubleshooting guide
Figure 3. Example of a troubleshooting guide

From observability to trusted automation: The power of context-aware AI

The future of resilient, self-healing systems lies in the seamless integration of observability, AI, and organizational knowledge. When these elements come together, they form the foundation for trusted automation—a system that not only reacts to incidents but learns from them, adapts, and eventually prevents them altogether.

At the heart of this vision is context. Effective auto-remediation depends on the ability to precisely identify the root cause of an issue and understand its broader impact across the application stack. But automation doesn’t stop at detection. By capturing and integrating the remediation strategies used by engineers, Dynatrace builds a living knowledge base. This organizational knowledge, when combined with AI-driven root cause analysis, allows the system to replicate proven remediation paths and suggest next steps with increasing accuracy.

With all relevant data and insights unified in a single platform, engineers gain a single pane of glass view into their systems. This not only streamlines manual remediation efforts but also lays the groundwork for flexible, context-aware auto-remediation. Each incident you resolve fuels the knowledge. Over time, as the system learns, it evolves from reactive automation to proactive incident prevention—anticipating issues before they escalate and taking preemptive action.

This is the vision Dynatrace is delivering: a future where engineers can trust automation not just to respond, but to understand, learn, and improve—turning every incident into a step toward greater system intelligence and reliability.

Empower your teams: turn hard-won operational insights into scalable remediation power

Dynatrace remediation intelligence is ready to work for you today. To start benefiting, opt into Davis CoPilot®, the Dynatrace generative AI assistant. Once turned on, you’ll need to configure Davis CoPilot to learn from a curated set of Dynatrace documents—specifically, Notebooks and Dashboards that are either created directly from detected problems or clearly labeled with the prefix [TSG] in their titles (short for Troubleshooting Guide).

Davis CoPilot will analyze and learn from your team’s historical remediation efforts, capturing valuable insights and strategies, allowing it to proactively suggest relevant documentation and guidance when similar incidents are detected in the future, helping your engineers respond faster and more effectively.

Figure 4. Turn on Davis CoPilot in Settings.
Figure 4. Turn on Davis CoPilot in Settings.

All data uploaded to Dynatrace remains strictly private. All data and remediation insights remain within your tenant. Davis CoPilot treats your documents as strictly confidential and never shares or transfers this information outside your Dynatrace environment. Your team’s knowledge stays private, secure, and entirely under your control, while still powering smarter, more context-aware automation.

Learn more about document suggestions and Dynatrace remediation intelligence in our documentation, and read about discovering relevant troubleshooting guides and how to create new ones.

Don’t let organizational knowledge stay buried. Make it actionable. Make it scalable.

The post Remediation intelligence: Accelerate MTTR with AI-powered context and knowledge appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/remediation-intelligence-accelerate-mttr-with-ai-powered-context-and-knowledge/feed/ 0
Insights into your Azure DevOps pipelines https://www.dynatrace.com/news/blog/insights-into-your-azure-devops-pipelines/ https://www.dynatrace.com/news/blog/insights-into-your-azure-devops-pipelines/#respond Tue, 11 Mar 2025 08:17:36 +0000 https://www.dynatrace.com/news/?p=68202 Insights into Azure DevOps pipelines

In this blog post, we demonstrate how Dynatrace® observability in CI/CD pipelines provides detailed insights that help developers debug faster and improve code quality.

The post Insights into your Azure DevOps pipelines appeared first on Dynatrace news.

]]>
Insights into Azure DevOps pipelines

In this blog post, we guide you through configuring a project to visualize real-time CI/CD build and release data for your Azure DevOps pipelines. The resulting visibility enhances collaboration with transparent data, increases productivity by automating monitoring tasks, and enables teams to detect issues proactively. With insights into your Azure DevOps pipelines, you can continuously improve processes, ensuring smoother and more reliable deployments over time.

A notebook version of the blog is also available on the Dynatrace GitHub repository. Download the AzureDevOps - Dynatrace Integration.json file and upload it to your Notebooks app in Dynatrace.

Prerequisites

Before we begin, ensure you have the following:

  • Access to your Dynatrace Tenant and permission to create tokens.
  • Your tenant ID, which can be found in your environment URL: https://<YOUR_TENANT_ID>.live.dynatrace.com/.
  • A token with Ingest Logs v2 scope.
  • A token with Write Settings scope.
ADO token properties example
Figure 1: Create a token with log ingest and write settings scope

Create webhooks in Azure DevOps

First, we need to create two service hooks subscriptions in Azure DevOps: one for Builds Completed and one for Release Deployment Completed.

  1. Navigate to https://{orgName}/{project_name}/_settings/serviceHooks.
  2. During the configuration, do not apply any filters.
  3. In the settings page of the subscription, fill in the following fields:
    • URL: https://<YOUR_TENANT_ID>.live.dynatrace.com/api/v2/logs/ingest
    • HTTP Headers: Authorization: Api-token <YOUR_LOG_INGEST_TOKEN>
    • Ensure the text above is copied exactly, replacing only the token.
    • Change “Messages to send” and “Detailed Messages to send” to Text.
Project settings showing URL, HTTP headers, and Messages details
Figure 2: Set up service hook subscriptions in Azure DevOps

Add a new Grail logs bucket

Follow these steps to create a new logs bucket in Dynatrace. Note: This step is optional. You’re welcome to use an existing Grail logs bucket.

  1. Open the Storage Management app in your tenant: Select CTRL/CMD + K and enter Storage.
  2. Create a new bucket by selecting + in the top right corner.
  3. Name the bucket azure_devops_logs.
  4. Set the retention time as desired.
  5. Set the bucket type to logs.

Configure OpenPipeline with log processing rules

Using OpenPipeline™, you can easily define a rule that routes all relevant log lines you created in the previous step into the bucket. This process also performs the necessary processing steps, such as renaming fields or transforming log data into dedicated events.

  1. Open the OpenPipeline app and select Logs in the left pane.
  2. Select the Pipelines tab and create a new one by selecting + Pipeline.
  3. Name the new pipeline AzureDevOps.
  4. Go to Dynamic Routing and create the following new rule:
matchesPhrase(eventType,"ms.vss-release.deployment-completed-event") OR matchesPhrase(eventType,"build.complete")
  1. From the Pipeline dropdown, select “AzureDevOps”.
  2. Return to your pipelines and open “AzureDevOps”.
  3. Under the Storage section, add a new processor > Bucket assignment, set a name and select the azure_devops_logs bucket from the dropdown.
Pipeline screen showing AzureDevOpsLogs bucket assignment attributes
Figure 3: Define the bucket assignment within OpenPipeline
  1. Next, open the Processing tab and define a couple of rules.
    1. First, create a new Rename Fields rule and call it “Rename Build Fields” where the left value is the new field name and the right value is the existing field name.
1. resource.buildNumber: buildNumber
2. resource.result: result
    1. Second, create another Rename Fields rule called “Rename Release Fields” where the left value is the new field name and the right value is the existing field name.
1. resource.stageName: stageName  
2. resource.project.name: projectName 
3. resource.deployment.release.name: releaseName 
4. releaseStatus:resource.environment.status: releaseStatus

Use the following sample data to verify your rules are working as expected.

Release event sample data:

{ 
       "timestamp": "2024-11-11T15:14:51.104000000-05:00", 
       "loglevel": "NONE", 
       "status": "NONE", 
       "createdDate": "2024-11-11T20:14:50.6300269Z", 
       "detailedMessage.text": "Deployment of release Release-946 on stage Staging succeeded. Time to deploy: 00:14:14.", 
       "dt.auth.origin": "dt0c01.YFMJ6LUO43SFFDW2SF7EW5YZ", 
       "eventType": "ms.vss-release.deployment-completed-event", 
       "id": "1862ab11-c0d4-451a-8b9b-0dfe7f517297", 
       "message.text": "Deployment of release Release-946 on stage Staging succeeded.", 
       "resource.environment.status": "succeeded", 
       "resource.project.name": "devlove-alpha", 
       "resource.stageName": "Staging", 
       "resource.deployment.release.name": "Release-946" 
    }
OpenPipeline screen showing rename field values with sample data
Figure 4: Testing the processing rule

Create an Azure DevOps dashboard and visualize log data

Now that we have ingested the logs coming from AzureDevOps, let’s visualize the data to get better and quicker insights into our CICD pipelines.

  1. Go to the AzureDevOps Git Repository and download the AzureDevOps Dashboard (on Logs).json file.
  2. Within Dynatrace, open the Dashboards app and select Upload at the top left corner.
  3. Upload the JSON file to start visualizing your Azure DevOps data.
Azure DevOps screen showing 85 releases succeeded and 89 releases failed.
Figure 5: Ready made Dashboard to monitor ingested ADO logs.

Extract “release” and “build” events

Taking this a step further, we can convert the ingested log data to SDLC (Software Development Lifecycle) events in case we detect a new release or build and discard the related log line afterward.

This helps reduce the number of stored log data and supports further platform engineering use cases, such as calculating DORA metrics, automating development processes, or observing the health of your engineering pipeline.

Disclaimer: The following instructions extract Business events from log data. Once supported by OpenPipeline we propose to extract Software Development Lifecycle Events (SDLC events), which are the preferred way of storing the extracted information.

Note: You may need to request a Log Content Length (MaxContentLength_Bytes) increase depending on how many steps your Release Events have. The integration generates ingest costs for logs and metrics (according to your rate card), depending on how many build/release events you ingest.

  1. In Dynatrace, open OpenPipeline:
    • Go to OpenPipeline > Logs > Pipelines > AzureDevOps.
    • Navigate to the Data Extraction section.
  2. Next, create a Business Event Processor rule, for all build events, using the following parameters:
    • Name: Build Result
    • Matching condition: matchesPhrase(eventType,"build.complete")
    • Event type: field name eventType
    • Event provider: Change to Static String: AzureDevOps
    • Fields to extract: result, buildNumber, resource.reason
OpenPipeline screen of Azure DevOps data extraction properties build properties
Figure 6: Create a new “build” event
  1. Finally, we use another Business Event Processor rule, to extract release events:
    • Name: Release Result
    • Matching condition: matchesPhrase(eventType,"ms.vss-release.deployment-completed-event")
    • Event type: field name eventType
    • Event provider: Change to Static String: AzureDevOps
    • Fields to extract: stageName, projectName, releaseName, releaseStatus, resource.deplyoment.startedOn, resource.deployment.completedOn
OpenPipeline screen showing Azure DevOps data extraction properties release results
Figure 7: Create a new “release” event

Extract Davis® AI events (optional)

Besides transforming the log line into SDLC events, you can also extract Davis events in case something goes wrong. These events can be used to create an alert or trigger a (remediation) workflow.

  1. First, add an event if a build has failed. Add a new “Davis event” processor using the following data:
    • Name: Build Complete Failed
    • Matching condition: matchesPhrase(eventType,"build.complete”) AND result== “build”
    • Event description: Unable to generate build {buildNumber}

Note: You can change the event.type in case you want to increase the severity level.

OpenPipeline screen showing Azure DevOps data extraction properties for Build Complete Failed
Figure 8: Create a new Davis event for failed builds.
  1. Next, extract a Davis event if the deployment is rejected:
    • Name: Release deployment rejected
    • Matching condition: matchesPhrase(eventType,"ms.vss-release.deployment-completed-event”) AND releaseStatus== “rejected”
    • Event description: Unable to deploy release {releaseName}
OpenPipeline screen for Azure DevOps showing data extraction properties for release rejected
Figure 9: Create a new Davis event for rejected deployments.

Discard logs (optional)

Now that we have successfully converted the log event into a SDLC (Business) and a Davis event, you can disable the storage assignment rule. This will reduce the amount of data stored in Grail and help you saving some money (and consequently speeding up your log queries). Go to OpenPipeline > Logs > Pipelines > AzureDevOps.

  1. Open the OpenPipeline app and select the pipeline we created before.
  2. Select the tab Storage and change the matching condition to false.

Analyzing data

After deleting the logs events we need to adapt our dashboard to consume business instead of log data. To speed up things you can upload a ready-to-use dashboard into your environment:

  1. Go to the AzureDevOps Git Repository and download the AzureDevOps Dashboard (on BizEvents).json file.
  2. Within Dynatrace, open the Dashboards, select Upload at the top left corner, and select the JSON file.

Additionally, we recommend that you create a segment to filter all of your monitored entities across different apps.

  1. Go to the Segments app and create a new segment by selecting + in the top right corner.
  2. Rename the segment to AzureDevOps.
  3. Select + Business events and add the following filter: event.provider = AzureDevOps
  4. Select + Logs and add the following filter: dt.system.bucket = azure_devops_logs
  5. Select Preview to validate the filters and Save once you’re done.
AzureDevOps screen showing variables for business events and logs
Figure 10: Dynamically filter your data across apps using segments

What’s next

By following these steps, you’ll be able to seamlessly integrate Azure DevOps with Dynatrace, enabling efficient log management and insightful data visualization.

Happy monitoring! 🚀

If you don’t already have Dynatrace, you can try this yourself in the Dynatrace Playground sandbox environment.

The post Insights into your Azure DevOps pipelines appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/insights-into-your-azure-devops-pipelines/feed/ 0
Powerful exploratory analytics for AI-driven insights https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/ https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/#respond Tue, 04 Feb 2025 16:00:42 +0000 https://www.dynatrace.com/news/?p=67543 Problem alert dashboard

The Dynatrace platform empowers Operations, SRE, and DevOps teams to maintain high software quality, security, and reliability, allowing organizations to innovate and scale confidently. By leveraging Davis® AI with enhanced predictive analytics and automated workflows, Dynatrace simplifies issue detection and resolution, reduces MTTR, and enables proactive incident prevention.

The post Powerful exploratory analytics for AI-driven insights appeared first on Dynatrace news.

]]>
Problem alert dashboard


Deploying and safeguarding software services has become increasingly complex despite numerous innovations, such as containers, Kubernetes, and platform engineering. Recent global IT outages, such as the CrowdStrike incident, remind us how dependent society is on software that works perfectly.

Organizations must balance many factors to stay competitive.
Figure 1. Organizations must balance many factors to stay competitive.

Organizations strive to strike a delicate balance between cost, time to market, and innovation. This challenge is more pressing than ever as businesses seek to stay competitive while ensuring their software remains robust and secure.

This necessitates a comprehensive platform that empowers enterprises to understand IT and software within the broader context of their business operations, giving them confidence that their software and IT infrastructure are reliable.

Scale with confidence: Leverage AI for instant insights and preventive operations

Using Dynatrace, Operations, SRE, and DevOps teams can scale efficiently while maintaining software quality and ensuring security and reliability. Its AI-driven exploratory analytics help organizations navigate modern software deployment complexities, quickly identify issues before they arise, shorten remediation journeys, and enable preventive operations.

We’ve added numerous enhancements to our platform, leveraging advanced AI and automation for smarter software observability.

In this blog post, we show you how to

  • Get AI-driven insights directly on your operations dashboards
  • Improve MTTR with AI-assisted problem analysis and logs and traces in context
  • Leverage Gen AI through Davis CoPilot to get insights into root causes
  • Automate remediation of AI-detected problems with simple workflows
  • Adopt Preventive Operations with AI forecasting and automated action

Get AI-driven insights directly on your operations dashboards

A high-level, customizable view of your data is crucial in modern software operations. Dynatrace Dashboards, powered by Grail™ data lakehouse and Davis® AI, offer precisely that. They provide a comprehensive overview, seamlessly integrating health and problem-related information into a single view. You can chart your topology across data silos alongside all alerts, events, and problems using honeycomb tiles, which offer convenient drill-downs into the problem-debugging user flow.

Dynatrace ensures that context is seamlessly integrated into the platform, thus simplifying complexity for you as a user when analyzing issues and allowing you to focus on what truly matters. AI-driven analytics transform data analysis, making it faster and easier to uncover insights and act. This approach not only improves user experiences, it ensures that critical insights are accessible to both experts and novices. By simplifying remediation journeys and extending features to more user groups, Dynatrace enables results across all teams.

The new Problems dashboard, including rich honeycomb visualization, helps you focus on what’s important, turning technical data into a visual story.
Figure 2. The new Problems dashboard, including rich honeycomb visualization, helps you focus on what’s important, turning technical data into a visual story.

When a truly important issue stands out, the next step is refinement. With a few clicks, you can segment and filter your data to focus on specific applications, assignment groups, or regions. Directly mapping and surfacing ownership information within data segments accelerates incident assignment notifications and triggers automatic remediations.

Utilize the comprehensive filter functionality to update your dashboards dynamically.
Figure 3. Utilize the comprehensive filter functionality to update your dashboards dynamically.

If you see an issue or need to look closely at a specific application where an issue was identified, simply select the element to be seamlessly directed to the Problems app. There, you can dig deeper while continuing to focus on your selected segment. This tight integration, following a golden thread of insights, ensures that you’re more productive. To experience the possibilities of AI-empowered dashboards, try our example dashboard on the Dynatrace Playground.

Improve MTTR with AI-assisted problem analysis, logs, and traces in context

The Problems app delivers opinionated AI-assisted problem analysis optimized for Operations and Site Reliability Engineers (SREs) and developers. According to IDC, guiding users visually and automatically surfacing all critical details enables a 56% faster mean time to repair (MTTR) for critical incidents.

When a large-scale incident occurs, follow the red flag that Davis AI uses to identify the root cause, pinpoint all relevant details, and visually reproduce the details in charts, highlighting the affected deployment.

Analyze the root cause in the Problems app.
Figure 4. Analyze the root cause in the Problems app.

Besides identifying the root cause, Davis AI also automatically connects all relevant log lines. Logs are invaluable for identifying further insights and detecting fundamental flaws, such as process crashes or exceptions. With a single click in Problems, all incident logs are surfaced automatically. But we don’t stop there, Dynatrace also seamlessly integrates relevant trace data, offering full visibility into even complex, microservices-based architectures.

By providing these end-to-end insights, Dynatrace and Davis AI empower SREs, developers, and architects to quickly dive deep into an incident’s details, including all relevant logs and traces. Using this context, they can effectively focus on fixing and remediating code-level issues, significantly improving MTTR, and ensuring that critical incidents are resolved swiftly and efficiently.

Leverage GenAI via Davis CoPilot for insights into root causes

Dynatrace offers precision tools for domain experts to solve complex problems and dig deeper into their data. While product owners often focus on the intricate technical details of an incident, they often prefer a quick summary of what happened and what caused it. The soon-to-be-globally available Davis CoPilot™ bridges this gap by summarizing problems and their root causes and suggesting remediation steps based on these insights.

You’re not limited to one problem; Davis CoPilot can simultaneously analyze multiple problems, draw conclusions about their relationships, identify the common root cause, and propose corrective steps. Instead of relying on a team of experts and waiting hours for insights, Davis CoPilot helps you identify similarities and draw relevant conclusions independently and efficiently.

The use of generative AI adds significant value by augmenting Dynatrace-detected technical root causes with knowledge from the global tech community. Generative AI can access and synthesize vast amounts of information from various sources, providing a broader context and deeper insights. This ensures that your teams benefit from the latest advancements and solutions, enhancing their ability to resolve issues effectively and efficiently.


Dynatrace Problems App - Explain Problems video

Gain a better understanding of root causes with Davis CoPilot
Figure 5. Gain a better understanding of root causes with Davis CoPilot

Automate remediation of AI-detected problems with simple workflows

To automatically remediate Davis AI-detected problems, Dynatrace leverages powerful Workflows. Dynatrace workflows can be triggered by any problem or alerting event, automating domain-specific tasks to take remedial actions.

For example, workflows can scale up capacity to adapt to demand or automatically restart a service in case of a crash. With a large catalog of available workflow actions, you can react efficiently to AI-detected problems, reducing mean time to repair (MTTR) by automatically remediating issues.

But you can do much more with it: The recently introduced Simple Workflows, which are included in your Dynatrace subscription with no extra cost, offer greater flexibility and power than standard notifications. You can use the same mechanisms and trigger types to notify your developer team via Slack, create a JIRA issue, or send a PagerDuty alert.

This ensures that your operations, SRE, and DevOps teams can focus on more strategic tasks while the system handles routine problem resolutions. Automation enhances operational efficiency and ensures that your systems remain robust and reliable, even in the face of unexpected issues.

Easily set up automated remediation with the new Simple Workflows.
Figure 6. Easily set up automated remediation with the new Simple Workflows.

Adopt Preventive Operations with AI forecasting and automated action

Going beyond reactive problem detection, analysis, and remediation, Dynatrace can also leverage predictive AI to anticipate and avoid critical situations before they occur. Using Davis AI forecast, you can easily predict future capacity demands. Combining this knowledge with workflows allows you to take proactive measures to ensure system stability and performance.

Let’s have a look at a concrete example:

It’s easy to predict key indicators of your application, such as order levels or service request counts. Once load and demand rise and Davis AI identifies a potential future issue in your infrastructure setup, Davis CoPilot can automatically generate an updated Kubernetes configuration script for you and automatically upscale the environment to meet future demand. This ensures that your system scales appropriately to handle the anticipated demand, preventing incidents before they occur and eliminating the need to generate a problem.

That’s what we call Preventive Operations. Instead of sending an alert and notifying people, Dynatrace simply fixes the issue. According to Gartner’s Analytics Maturity Model, using predictive AI can significantly reduce the likelihood of incidents by taking preemptive action and remediation.

Start using Davis AI to analyze your environments and predict and address potential issues in advance. This will empower your teams to avoid potential problems and ensure a smooth, uninterrupted user experience.

Initiate automated, corrective action before an issue occurs
Figure 7. Initiate automated, corrective action before an issue occurs.

Tackle business challenges with confidence

Ensure your software runs securely and reliably with Dynatrace and Davis AI.

Dynatrace and Davis AI support you by running your software securely and reliably. This includes advanced root cause analysis, deep insights into detected issues, and corrective actions—whether manual or automatic—to prevent outages before they occur.

Get started

For more information, have a look at our documentation or explore the available resources on the Dynatrace Playground to experience some of these enhancements first-hand:

The post Powerful exploratory analytics for AI-driven insights appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/feed/ 0
Davis CoPilot expands: Get answers and insights across the Dynatrace platform https://www.dynatrace.com/news/blog/davis-copilot-expands-get-answers-and-insights-across-the-dynatrace-platform/ https://www.dynatrace.com/news/blog/davis-copilot-expands-get-answers-and-insights-across-the-dynatrace-platform/#respond Tue, 04 Feb 2025 16:00:17 +0000 https://www.dynatrace.com/news/?p=67510 Davis CoPilot

We’re excited to announce that Davis CoPilot Chat is now available across the Dynatrace platform. Davis CoPilot™, launched in October 2024 to support Dynatrace users with access to their data, now extends across the platform, streamlining user onboarding and providing comprehensive support and contextual insights from various Dynatrace® Apps. With the new Davis CoPilot conversational […]

The post Davis CoPilot expands: Get answers and insights across the Dynatrace platform appeared first on Dynatrace news.

]]>
Davis CoPilot


Update: We’ve launched Dynatrace Assist, our next-generation AI chat that goes far beyond answering questions.
Dynatrace Assist is the evolution of Davis CoPilot®.

We’re excited to announce that Davis CoPilot Chat is now available across the Dynatrace platform. Davis CoPilot™, launched in October 2024 to support Dynatrace users with access to their data, now extends across the platform, streamlining user onboarding and providing comprehensive support and contextual insights from various Dynatrace® Apps. With the new Davis CoPilot conversational interface, users can leverage natural language to quickly get answers to their questions, making it easier than ever for users to interact with Dynatrace.

Intuitive access to information boosts team productivity

We understand that taking advantage of the numerous features and functionalities offered by platforms like Dynatrace can be challenging. To help you navigate this and boost your efficiency, we’re excited to announce that Davis CoPilot Chat is now generally available (GA). This new feature provides information and guidance exactly when and where you need it, making your Dynatrace experience smoother and more efficient.

Davis CoPilot can be accessed anytime directly from the Dock.

Davis CoPilot leverages the power of generative AI to answer your questions through a globally accessible chat interface. We’re proud to say that Davis CoPilot is multilingual: you can ask questions and get answers in many different languages, including French, Spanish, German, Portuguese, Chinese, Japanese, and, of course, English. Davis CoPilot provides immediate, accurate responses, eliminating the need for extensive searches and reducing dependency on support channels. This makes knowledge more readily available and boosts productivity and user experience for both new and experienced users.

Davis CoPilot Chat follows our recent announcement of the general availability of Quick Analysis in Notebooks and Dashboards, which makes data accessible to technical and non-technical users alike. This means you can interact with data stored in the Dynatrace Grail™ data lakehouse just by using natural language.

Simplify onboarding and quickly find what you’re looking for with Davis CoPilot

You can start using the Davis CoPilot conversational interface immediately. Simply enable Davis CoPilot and assign the relevant user permissions, and the Davis CoPilot button will appear in the Dock.

Start a new conversation with Davis CoPilot Chat by selecting it in the Dock or by pressing CTRL/CMD + I and entering your question.

Davis CoPilot is great for guiding new and occasional users
Figure 2. Davis CoPilot is great for guiding new and occasional users

New users can quickly get up to speed with Dynatrace by asking Davis CoPilot for help with basic commands, setup instructions, and troubleshooting tips. This reduces the learning curve and enables new users to become productive faster. The conversational interface provides step-by-step guidance, making the onboarding process smoother and more efficient.

If you’re already familiar with Dynatrace, you can rely on Davis CoPilot to provide detailed explanations for a wide range of expert questions related to exploring new use cases, advanced configuration topics, and building custom apps.

Here are some examples of questions you can ask Davis CoPilot:

  • Onboarding: How do we start sending OpenTelemetry data to Dynatrace?
  • Understanding Dynatrace: What is the difference between an event and a problem in Dynatrace?
  • Exploring Dynatrace solutions: How can we comply with the Digital Operational Resilience Act (DORA) using Dynatrace?
  • Configuring your environment: How do I set up an alert based on an anomaly detector?
  • Developing custom apps: How can I import external table data and visualize it using the Dynatrace App Toolkit?

Get contextual assistance at the press of a button

Davis CoPilot seamlessly integrates into our use-case-specific Dynatrace Apps, offering you contextual insights and guidance at the press of a button. While we plan to release additional contextual app integrations in the coming months, several will be available a few weeks after launch, allowing Davis CoPilot to provide you with insights into:

  • Kubernetes warning signals
  • Individual problem details and the relationships between problems
  • Database performance optimization

Simplify Kubernetes: Davis CoPilot decodes warning signals

Understanding the background and root cause of warnings often requires in-depth subject matter expertise. That’s why we integrated Davis CoPilot into Kubernetes. Instead of manually looking up error messages, Davis CoPilot translates warning signals into clear, understandable language. In addition, Davis CoPilot offers a list of typical root causes and related remediation steps. This way, newcomers can quickly become proficient, and experts can elevate their expertise to hero status.

Davis CoPilot provides contextual guidance for Kubernetes warning signals
Figure 3. Davis CoPilot provides contextual guidance for Kubernetes warning signals

Problems demystified: Davis CoPilot provides insights into root causes

In Problems, Davis CoPilot provides clear summaries of problems, their root causes, and the suggested remediation steps. Davis CoPilot explains individual issues in clear language from the problem details page and can perform a comparative analysis when multiple problems are selected from the list view. This helps you identify common root causes and propose corrective steps without relying on a team of experts and waiting for hours for critical insights. If you want to learn more, have a look at Wolfgang Beer’s latest blog post and learn more about recent advancements in the Problems app.

Davis CoPilot explains problems in clear language
Figure 4. Davis CoPilot explains problems in clear language

Optimize database performance: Understand query execution plans

Query execution plans provide detailed information on how a database will execute an SQL query. While these provide the raw data on how to improve query performance and reduce resource consumption, they require expert knowledge to read and interpret. Now, in Databases, Davis CoPilot can provide natural language explanations of execution plans, breakdowns of relevant details, and recommendations on how to improve statement performance. This gives non-expert database users, such as developers, the knowledge they need to optimize their application performance and database utilization.

Davis CoPilot explains query execution plans
Figure 5. Davis CoPilot explains query execution plans

Empower your teams with Davis CoPilot today

The launch of Davis CoPilot Chat marks the second milestone of our journey. We’re committed to continuously enhancing the assistant’s capabilities with upcoming features, including query explanations, workflow actions, and troubleshooting guides.

Get started with Davis CoPilot today and transform how you and your teams interact with Dynatrace:

Thanks for joining us on this exciting journey. We look forward to your feedback and to seeing how Davis CoPilot helps your teams achieve their goals.

Davis CoPilot Chat, as well as the Dynatrace Apps integrations mentioned in this blog post, will be available starting with the release of Dynatrace SaaS version 1.307.

The post Davis CoPilot expands: Get answers and insights across the Dynatrace platform appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/davis-copilot-expands-get-answers-and-insights-across-the-dynatrace-platform/feed/ 0
Advancing AIOps: Preventive operations powered by Davis AI https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/ https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/#respond Tue, 04 Feb 2025 16:00:06 +0000 https://www.dynatrace.com/news/?p=67673 Davis AI alerts

The 2024 CrowdStrike incident demonstrated our societal vulnerabilities to IT outages. A faulty software update caused widespread issues, impacting critical services globally, including airlines, banks, hospitals, and public safety systems. Despite recent advancements such as containers, Kubernetes, and platform engineering, it’s evident that managing enterprise software services has become increasingly complex. IT operations must be prepared to quickly address and mitigate disruptions, ensuring business continuity and minimizing damage.

The post Advancing AIOps: Preventive operations powered by Davis AI appeared first on Dynatrace news.

]]>
Davis AI alerts

AI, especially AIOps, has emerged as a pivotal solution, promising to avoid downtime. The 2024 State of AI Report highlights this trend, with 89% of technology leaders anticipating that AI will significantly enhance incident response by learning to automate and optimize various tasks, such as performance monitoring and workload scheduling.

Blue screens of death at LGA airport due to the July 2024 CrowdStrike outage. (Source: Wikimedia Commons.)
Figure 1. Blue screens of death at LGA airport due to the July 2024 CrowdStrike outage. (Source: Wikimedia Commons.)

AIOps can identify and address potential issues before they become major incidents by learning from history and analyzing large amounts of data in real time. This approach improves operational efficiency and resilience, though it’s not without flaws. The complexity of IT environments and the changing nature of threats necessitate human oversight and ongoing adjustment of AIOps systems to handle unforeseen challenges and ensure optimal performance. Additionally, predictions based on historical data are reactive, solely relying on past information to anticipate future events, and can’t prevent all new or emerging issues. This limitation highlights the importance of continuous innovation and adaptation in IT operations and AIOps strategies.

“The shift from reactive to preventive operations represents the next evolution in AIOps.”
Bernd Greifeneder, CTO Dynatrace

When Dynatrace set out with Davis® AI over 10 years ago, pioneering AI-driven operations, we focused initially on problem identification before moving on to problem remediation. The next milestone in enhancing the capabilities of Davis AI—another pioneering step forward in AI-driven operations—is outright problem prevention. In this blog post, we explain how the unique combination of causal, predictive, and generative AI—augmented by the latest Davis AI advancements—is transforming how Dynatrace customers manage and optimize their IT infrastructure.

Automatic root cause detection

Modern, complex, and distributed environments generate a substantial number of events. This necessitates additional requirements such as minimizing the total number of issues, eliminating false positives, and conducting accurate root cause analysis.

Dynatrace has a longstanding reputation for accurately analyzing root causes and identifying related events. While other methods typically rely on mere correlation and historical data analysis, we’ve further enhanced our capabilities by implementing causational analysis, which leverages contextual information automatically gathered during data ingestion and processing in addition to historical data analysis. This is achieved using Dynatrace Grail™, our causational data lakehouse, which unifies all data in an always-up-to-date topology model. By applying causal AI to incoming data in real time, Davis instantly learns and continuously adapts to new information. This facilitates more precise root cause analysis and anomaly detection, including identifying seasonal anomalies and establishing auto-adaptive thresholds.

Root cause analysis with the Problems app
Figure 2. Root cause analysis with the Problems app

When applying this Davis root cause detection within our own IT environment, Davis effectively filters out over 99.9% of incoming data noise, condensing hundreds of thousands of daily system events into no more than four or five incidents that require attention from our IT operations team.

These algorithms are not limited to monitoring IT environments. At our February 2025 Dynatrace Perform session on exploratory analytics with AI-driven insights, the Performance Engineering Lead of XXXLutz—one of the world’s largest furniture retailers operating more than 370 stores across Europe—explains how XXXLutz utilizes Davis AI to proactively identify critical order drops, allowing them to respond quickly and effectively to changing market conditions and ensuring that their business remains agile and responsive to the needs of their customers.

Problem journey and reactive remediation

At the core of Dynatrace problem remediation stands the Problems app—an optimized view into opinionated insights, details, and context of each detected issue—for Operations, SREs, and developers. It filters billions of log lines, including the topology of each incident and its affected entities, for efficient problem triaging and troubleshooting, resulting in a 56% faster mean time to repair (MTTR) for critical incidents.

With the latest release, we drive this further by improving the automatic connection of relevant log and trace data for further drill down, presenting the full context of an issue in a single view. This provides comprehensive visibility into even complex architectures, simplifying the process of examining relevant details and addressing code-level issues, reducing 100 clicks and manual filtering to a single click with no loss of context.

Comparative analysis of multiple problems with Davis CoPilot
Figure 3. Comparative analysis of multiple problems with Davis CoPilot

By utilizing Davis CoPilot™, you can conduct comparative analyses of multiple issues, obtain natural language summaries of individual problems, and receive contextual recommendations along with specific remediation steps.

You can also link troubleshooting guides created in Notebooks to remediated issues, thereby building an intelligent knowledge base. Davis automatically connects additional documents as well as stored workflows. So the next time a similar problem arises, Davis brings up related guides, enabling teams to learn from previous experiences and reducing the risk of knowledge loss.

Harness your collective knowledge by connecting troubleshooting guides
Figure 4. Harness your collective knowledge by connecting troubleshooting guides

Please refer to our recent blog posts for more information on utilizing Problems for AI-driven insights and the latest Davis CoPilot advancements.

Automating the remediation

While obtaining comprehensive insights is beneficial, true transformation occurs through the use of tools that automatically execute remediation steps. To implement these “AI-driven operations,” it’s essential to forecast future requirements, including capacity demands, potential system failures, and security incidents.

Traditional forecasting engines typically depend on historical data, stored in metrics. In contrast, Davis AI generates real-time predictions, facilitating proactive operations. This capability is due to Davis’s ability to process raw data, such as logs, for forecasting, leveraging Grail to execute previously unattainable queries.

Consider the following scenario: You begin by retrieving and analyzing logs to identify relevant values for automation. Once this task is complete, you proceed to your pipelining tool to configure ingestion rules that extract these values into metrics and then wait several weeks for your prediction engine to generate alerts that can serve as triggers for your workflows.

However, when utilizing Dynatrace with its integrated anomaly detection and forecasting capabilities, you gain the advantage of schema-less data analysis and the ability to process any raw data into time series in real time. This significantly reduces the time required to establish AIOps workflows from several weeks to less than 30 minutes.

Preventive operations

The complexity of modern software environments makes it challenging to determine a service’s reliability solely through testing. It’s impractical to emulate scenarios such as generating a million tickets to assess performance capabilities. This necessitates real-time insights and operations rather than reactive problem-solving or raising alerts to notify personnel.

Preventive operations address this need by enabling proactive corrective actions before issues arise, akin to predictive maintenance. AI-supported anomaly detection identifies parameters that deviate from the norm, allowing for automatic configuration adjustment to mitigate potential problems preemptively.

Dynatrace offers the only unified, AI-powered platform for all data, all teams, and all possibilities.
Figure 5. Dynatrace offers the only unified, AI-powered platform for all data, all teams, and all possibilities.

Davis CoPilot combines the “power of three”:

  • Davis causal AI for identifying anomalies and root cause analysis
  • Davis predictive AI for precise forecasting and determining when to take action
  • Generative AI capabilities that perform actions beyond simply sending notifications or restarting services

In this way, Dynatrace extends AIOps beyond traditional IT operations tasks and addresses complex scenarios, including security use cases such as threat observability. Consider the following real-world example:

At Dynatrace, we log all failed login attempts. We can predict potential threats when abnormal patterns are identified and raise a security event by utilizing seasonal baselining. The subsequent workflow involves checking the IP address and generating a threat score. Upon reaching a certain threshold, a new ruleset is automatically added to the web application firewall. This entire process is fully automated, running before a problem even occurs, significantly reducing the response time from over an hour to a fraction of a second.

In another instance, automatic log pattern analysis crawling our application logs decreased the number of bugs in the production environment by 15% and freed up time previously spent on log analysis and triaging (in pre-prod), equivalent to 17 full-time employees. Consequently, these 17 developers can now dedicate their efforts to adding more value to Dynatrace.

Summary

The State of AI report states that over 88% of technology leaders anticipate AI will enhance incident responses and improve their teams’ ability to predict and proactively resolve service-affecting issues.

With Dynatrace, organizations are prepared to evolve their ITOps and SRE departments from troubleshooting to prevention, getting proactive with forecasting, and utilizing generative AI instead of purely focusing on history-focused root cause analysis.

Start your preventive operations journey with smart automation and auto-remediation that prevents larger issues.

Are you interested in gaining more insights?

The post Advancing AIOps: Preventive operations powered by Davis AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/feed/ 0
Better dashboarding with Dynatrace Davis AI: Instant meaningful insights https://www.dynatrace.com/news/blog/better-dashboarding-with-dynatrace-davis-ai/ https://www.dynatrace.com/news/blog/better-dashboarding-with-dynatrace-davis-ai/#respond Tue, 21 Jan 2025 21:38:07 +0000 https://www.dynatrace.com/news/?p=67370 abstract image showing connected dots and waves representing MCP best practices for agentic AI

Discover the value of Davis® AI when working with dashboards for observability, security, or business use cases. Quickly spot anomalies by activating Davis AI on any numeric time series chart data. Stay ahead with visual, AI-powered forecasting, or get new insights into your data with just a few clicks by leveraging Davis CoPilot™.

The post Better dashboarding with Dynatrace Davis AI: Instant meaningful insights appeared first on Dynatrace news.

]]>
abstract image showing connected dots and waves representing MCP best practices for agentic AI

Ensuring smooth operations is no small feat, whether you’re in charge of application performance, IT infrastructure, or business processes. Chances are, you’re a seasoned expert who visualizes meticulously identified key metrics across several sophisticated charts. Your trained eye can interpret them at a glance, a skill that sets you apart.

However, your responsibilities might change or expand, and you need to work with unfamiliar data sets. The market is saturated with tools for building eye-catching dashboards, but ultimately, it comes down to interpreting the presented information. This is where Davis AI for exploratory analytics can make all the difference.

Activate Davis AI to analyze charts within seconds
Figure 1. Activate Davis AI to analyze charts within seconds

Davis AI can help you expand your dashboards and dive deeper into your available data to extract additional information. Our customers value the nearly unlimited possibilities for querying and joining data on the Dynatrace platform, with the option of instant, real-time visualization of query results. Whether you’re an expert or an occasional user, our recently launched Davis CoPilot will enable you to get instant results without the need to write complex queries yourself. Have a look at our recent Davis CoPilot blog post for more information and practical use cases.

If you’ve already created your dashboards, now is the time to use Davis AI to identify anomalies or predict future trends without restricting use cases.

Leverage Davis AI for anomaly detection and instant insights

“My chart shows a peak at 8:00 AM. Do I need to investigate this further?” You might be regularly confronted with this or similar questions. Davis AI machine learning capabilities will help you identify actual anomalies within seconds, enabling you to focus resources on issues that matter.

Based on your requirements, you can select one of three approaches for Davis AI anomaly detection directly from any time series chart:

  • Auto-Adaptive Threshold: This dynamic, machine-learning-driven approach automatically adjusts reference thresholds based on a rolling seven-day analysis, continuously adapting to changes in metric behavior over time. For example, if you’re monitoring network traffic and the average over the past 7 days is 500 Mbps, the threshold will adapt to this baseline. An anomaly will be identified if traffic suddenly drops below 200 Mbps or above 800 Mbps, helping you identify unusual spikes or drops.
  • Seasonal Baseline: Ideal for metrics with predictable seasonal patterns, this option leverages Davis AI to create a confidence band based on historical data, accounting for expected variations. For instance, in a web shop, sales might vary by day of the week. Using a seasonal baseline, you can monitor sales performance based on the past fourteen days. An anomaly is identified if sales on a Friday are significantly lower than on previous Fridays, indicating a potential issue.
  • Static Threshold: This approach defines a fixed threshold suitable for well-known processes or when specific threshold values are critical. For example, if you have an SLA guaranteeing 95% uptime, you can set a static threshold to alert you whenever uptime drops below this value, ensuring you meet your service commitments.

Davis AI is particularly powerful because it can be applied to any numeric time series chart independently of data source or use case.

The following example will monitor an end-to-end order flow utilizing business events displayed on a Dynatrace dashboard. By leveraging Davis AI anomaly detection, we can identify potentially fraudulent behavior by activating anomaly detection on the Average order size chart. As shown in the chart below on the lower left, most values fall within the band of acceptable response time (highlighted in green), with only one spike occurring at 5:00 AM. Since this spike was outside the expected range, an anomaly was identified.

Apply Davis AI anomaly detection to detect fraudulent behavior in a business process
Figure 2. Apply Davis AI anomaly detection to detect fraudulent behavior in a business process
  • Application Observability: Identify unexpected error rate increases in application performance, helping pinpoint and resolve issues quickly.
  • Digital Experience Management: Monitor user interaction patterns to spot anomalies in website or app performance that could affect user experience, such as slow page load times.
  • FinOps: Track irregularities in cloud spending or resource usage, enabling cost optimization and preventing budget overruns.

Davis AI forecast analysis predicts future numeric values of any time series. It can even process external datasets or the results of any data query if it can be displayed as a numeric time series, such as occurrences over time.

The forecast is created instantly, even for large data sets, and updates dynamically whenever filter settings are changed.

In application performance management, acting with foresight is paramount. Maintaining reliability and scalability requires a good grasp of resource management; predicting future demands helps prevent resource shortages, avoid over-provisioning, and maintain cost efficiency.

On this SRE dashboard, we utilize Davis AI to forecast and visualize future resource utilization:

SRE dashboard monitoring the four golden signals and forecasting resource utilization
Figure 3. SRE dashboard monitoring the four golden signals and forecasting resource utilization

Other potential applications for forecasting include:

  • Kubernetes: Forecasting helps dynamically scale Kubernetes clusters by predicting future resource needs. This ensures optimal resource utilization and cost efficiency. Forecasting can identify potential anomalies in node performance, helping to prevent issues before they impact the system.
  • Business: Using information on past order volumes, businesses can predict future sales trends, helping to manage inventory levels and effectively plan marketing strategies.

AIOps: Utilize Davis AI to predict and prevent

Utilizing the Dynatrace AutomationEngine, Davis AI forecasting capabilities can even trigger automated actions. One of our customers’ SRE teams needed to increase disk space to avoid ongoing over- and under-provisioning, which was time-consuming and annoying. Now, with Davis AI forecasting capabilities, the target disk size is predicted automatically, and an automated task for disk resizing is triggered when necessary.

If you want to further explore the possibilities for prediction and prevention management with Dashboards, have a look at our example dashboard in the Dynatrace Playground.

Prevent incidents through predictive maintenance and capacity management
Figure 4. Prevent incidents through predictive maintenance and capacity management

Experience Davis AI in action

To experience the possibilities of Davis AI, look at this short introduction video by Andreas Grabner:
How to chart and forecast any data point

To explore the depth of functionality of Dynatrace Dashboards yourself and get first-hand experience, try out the app in the Dynatrace Playground.

The post Better dashboarding with Dynatrace Davis AI: Instant meaningful insights appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/better-dashboarding-with-dynatrace-davis-ai/feed/ 0
Announcing General Availability of Davis CoPilot: Your new AI assistant https://www.dynatrace.com/news/blog/announcing-general-availability-of-davis-copilot-your-new-ai-assistant/ https://www.dynatrace.com/news/blog/announcing-general-availability-of-davis-copilot-your-new-ai-assistant/#respond Thu, 10 Oct 2024 14:18:13 +0000 https://www.dynatrace.com/news/?p=66106 Davis CoPilot icon

We're excited to announce the general availability of Davis CoPilot™, our groundbreaking generative AI assistant crafted to transform your data interaction experience with Dynatrace. Leveraging advanced large language models, Davis CoPilot converts your conversational prompts into accurate Dynatrace Query Language (DQL) commands, facilitating smooth and intuitive data analysis for both beginners and seasoned professionals.

The post Announcing General Availability of Davis CoPilot: Your new AI assistant appeared first on Dynatrace news.

]]>
Davis CoPilot icon

Update: We’ve launched Dynatrace Assist, our next-generation AI chat that goes far beyond answering questions.
Dynatrace Assist is the evolution of Davis CoPilot®.

Deal with data overload in the enterprise

In today’s rapidly evolving digital landscape, enterprises are inundated with vast amounts of data. Extracting meaningful insights from this data is crucial for staying competitive. However, traditional data analysis techniques can be time-consuming and demand specialized expertise, limiting how quickly and easily insights can be obtained.

Empower deep data analysis with natural language queries

Davis CoPilot enhances efficiency and productivity by seamlessly integrating generative AI throughout the Dynatrace platform. This feature allows you to effortlessly gain insights and generate queries without needing to learn new syntax or manage complex commands. Consequently, Dynatrace becomes accessible to a broader audience, including non-technical users and those who don’t work with Dynatrace on a daily basis, and empowers teams to make faster data-driven decisions.

Examples of generated queries
Figure 1. Examples of generated queries

Empower users with intuitive data access—without compromising security

At Dynatrace, we recognize the complexities associated with data environments. DQL, the query language employed to analyze data stored in Dynatrace Grail™ data lakehouse, offers remarkable versatility and power, serving as an essential tool for experienced users seeking to fully harness Grail’s capabilities. Davis CoPilot simplifies the data querying process for both professionals and beginners by enabling interactions through natural language. This democratizes data access, allowing all users to generate valuable insights swiftly and effortlessly. Consequently, the data analysis process is accelerated, empowering teams to make informed, data-driven decisions with increased speed and precision.

At Dynatrace, we prioritize the protection of your data. Our solutions are engineered to be secure, reliable, and entirely transparent. Davis CoPilot guarantees that your confidential information is never at risk of being leaked or disclosed across environments, as we ensure continuous protection of your prompts and data. Furthermore, there is no automatic model training or fine-tuning based on your usage, ensuring that your data is employed strictly for its intended purpose—to generate DQL and provide swift insights. This steadfast dedication to security and transparency enables you to use our tools confidently, trusting that your data is well-protected. Look at our documentation to get more insights into the privacy and security aspects of Davis CoPilot.

Get started with quick analysis in Notebooks and Dashboards

Davis CoPilot allows you to perform rapid data analysis in Notebooks and Dashboards by translating natural language prompts into Dynatrace Query Language (DQL). The results are automatically executed and returned, making complex data analysis more accessible than ever before.

Simply create a new notebook or dashboard, then select + Add > Davis CoPilot. Enter your prompt (or try one of our suggestions), and select Run. Davis CoPilot will generate and auto-execute the DQL so you can go from question to data insights in seconds. If you’d rather refine your query before executing it, open the dropdown list next to the run button and select Generate DQL only (this feature is currently only available in Notebooks).

Davis CoPilot video

Environment-aware queries unlock full data-context awareness

Davis CoPilot is much more than an AI tool that helps you create queries. Davis CoPilot knows the context of your data, which results in more precise answers using a feature called environment-aware queries.

Having environment-aware queries configured allows Davis CoPilot to identify unique data fields and custom metrics in your environment. You can now run more complex analyses and get better results by crafting more accurate queries that identify and reference relevant entities, events, spans, and metrics straight from your environment. And, of course, we do this without putting you or your data at risk. This functionality is opt-in, and you have full control over which data tables and buckets are accessible to Davis CoPilot. Let’s look at some examples:

If you’re an application owner tracking travel bookings for new trips on a travel website, you’ll likely need to track:

  • profit made on each booking  (as a business event)
  • applicable discounts (as a business event)
  • length of time it takes customers to complete a booking (as a custom metric)

With this in mind, you might give Davis CoPilot the following command: “Show me the average revenue and price reduction for new trips over the last month.”

If you have environment-aware queries configured, the following DQL will be generated automatically, and you’ll get the relevant results you’re looking for.

fetch bizevents , from:now() – 30d 
| filter event.type == “new trip” 
| makeTimeseries interval:1h, {profit= avg(profit), discount= avg(discount)

With environment-aware queries configured, Davis CoPilot infers that “revenue” refers to the profit field and “price reduction” refers to the discount field, even though your prompt doesn’t use the correct field names. However, if you don’t have environment-aware queries configured, Davis CoPilot can’t identify all relevant fields. For example, the following incorrect DQL will be generated if the same conversational command is issued when environment-aware queries are not configured. In such cases, you won’t get any results since the fields mentioned in the command don’t exist in your environment.

fetch bizevents, from:now() – 30d 
| filter event.type ==  “new trip”
| makeTimeseries interval:1h, {avg_revenue = avg(revenue), 
  avg_price_reduction = avg(price_reduction)

Alternatively, you might ask Davis CoPilot the following: “On average, how long does it take customers to book new trips?” If you have environment-aware queries enabled, the following DQL will be generated, and you’ll get the relevant results you need.

timeseries avg(new_trip_booking_duration)

Conversely, if you don’t have environment-aware queries configured, you’ll likely receive an error message because Davis CoPilot can’t correctly map your question to your custom metric key. In this case, Davis CoPilot can’t generate a valid DQL query since it won’t be able to find a matching built-in metric.

User permissions are enforced both with and without environment-aware queries, ensuring that Davis CoPilot provides relevant responses that comply with individual data-access rights. Environment-aware queries truly unlock the power of Grail for everyone in your organization.

What’s next for Davis CoPilot

This is just the beginning of our new AI assistant journey. We’re committed to making Davis CoPilot even better, and we’ve got some fantastic features coming your way, from query explanations to problem insights, document generators, and more.

We value your feedback and are continuously working to enhance our product. Want to share your thoughts? You can share your learnings directly from the Davis CoPilot interface. Your feedback helps us refine the functionality and better meet your needs. You can also request to participate in ongoing or upcoming Preview programs. Get in touch with your Dynatrace account manager if you’re interested.

Get started today and embrace the future of data analytics

The launch of Davis CoPilot marks a significant advancement in data analysis capabilities. If you have a Dynatrace Platform Subscription, Davis CoPilot is available for you with the release of Dynatrace SaaS version 1.301. If you have a classic license, Davis CoPilot is available for you with the release of Dynatrace SaaS version 1.304.

Empower your team with the ability to effortlessly transform natural language prompts into actionable insights. Activate Davis CoPilot in your Dynatrace environment today and explore how it can transform your data analysis workflows.

For more information and to get started, please visit our documentation. Thank you for being part of this exciting journey with us. We look forward to your feedback and seeing how Davis CoPilot helps you achieve your goals.

Ready to try out Davis CoPilot yourself?

The post Announcing General Availability of Davis CoPilot: Your new AI assistant appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/announcing-general-availability-of-davis-copilot-your-new-ai-assistant/feed/ 0
Transform your operations with Davis AI root cause analysis https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/ https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/#respond Tue, 08 Oct 2024 19:03:46 +0000 https://www.dynatrace.com/news/?p=66044 root cause analysis

Complexity is ever-increasing in today’s fast-paced world of software deployments and cloud infrastructure. This is why Davis® AI root cause analysis is an indispensable tool for Operations, Site Reliability, and DevOps teams.

The post Transform your operations with Davis AI root cause analysis appeared first on Dynatrace news.

]]>
root cause analysis

Without AI-assisted observability tooling, the productivity of operations teams drops, leading to a dramatic increase in Mean Time to Repair (MTTR) and a significant rise in the personnel needed to manage critical incidents. In an era dominated by automated, code-driven software deployments through Kubernetes and cloud services, human operators simply can’t keep up without intelligent observability and root cause analysis tools.

Modern observability has evolved from simple metric telemetry monitoring to encompass a wide range of data, including logs, traces, events, alerts, and resource attributes. Dynatrace Root Cause Analysis (RCA) seamlessly integrates all this information, providing crucial analysis to remediate incidents in real time.

Problem feed for fast triage and remediation of AI-detected problems.
Figure 1. Problem feed for fast triage and remediation of AI-detected problems.

By offering root cause analysis on top of the highly flexible Grail™ data lakehouse, Dynatrace empowers SRE and operations teams to further reduce MTTR. Direct access to the underlying data allows the automatic RCA analysis to eliminate data silos and to dive deep into every aspect of the collected incident data.

Unlike generic DIY query frontends, the Dynatrace Problems app is a tailor-made solution for efficiently supporting operations use cases. This approach ensures that your operation teams have all the tools they need to manage modern software deployments.

Transform your operations today with the new Problems app and stay ahead in the ever-evolving software and cloud infrastructure landscape.

Rapid response to critical incidents

Operations teams can quickly focus on incoming Davis AI-detected and -analyzed problems by referring to the problems feed.

The problem feed is designed to prioritize active issues, ensuring they always appear at the top, regardless of how long they’ve been ongoing. This default sorting strategy, which uses time as a secondary criterion, guarantees that Operations teams never overlook an active problem, no matter which primary filter is applied.

You can focus on your domain using the filter bar at the top, the quick filters on the side, or both. The chart feature allows for quick analysis of problem peaks at specific times.

Operations teams will appreciate the ability to sort problems by duration and the number of affected entities. This aids in assessing Davis-detected root causes and prioritizing remediation efforts. The native multi-select feature lets users open a filtered group of problems simultaneously, facilitating quick comparisons and detailed analysis.

Streamline deployment insights with AI-generated summaries

Every second counts during wide-scale incidents affecting large parts of your production systems. This is why precisely showing the root cause ultimately helps to speed up problem resolution.

You can multi-select a cohort of active problems, select Show detail, and review all critical problem details, including preview charts and event details, without losing the context of your problem feed.

The new problem experience transparently displays all the available details, with prominently displayed root-cause markers to precisely guide your attention.

In the realm of cloud infrastructure management, having a clear and concise view of your deployment’s health is crucial. Our dedicated deployment perspective offers just that, showcasing the hierarchy of affected and related infrastructure components. The root cause of any issue is prominently marked with a root-cause badge, making it easy to identify and address problems swiftly.

This perspective not only highlights the affected cloud regions but also provides a quick summary of the Kubernetes context where your workloads encountered failures. Gone are the days of clicking and navigating through multiple dashboards. Instead, you receive an AI-generated summary as an affected deployment architecture diagram.

This diagram, akin to a UML (Unified Modeling Language) deployment diagram, offers a familiar representation for software architects, ensuring they can quickly grasp the situation and take necessary actions. By streamlining the visualization of deployment issues, we empower teams to resolve problems more efficiently and maintain optimal performance.

To save time, the root-cause component is preselected, and all the details of the root cause are displayed on the right, along with charts showing the detected breaches from learned normal behavior.

You can review each individual finding on all problem-affected entities by selecting the individual deployment components or by switching to the detailed event perspective, which shows all the single events that the root cause analysis collected into a single problem.

Confirm the AI-detected root cause and review the deployment context.
Figure 2. Confirm the AI-detected root cause and review the deployment context.

In addition to using markers for swift root cause analysis, operations teams often seek to attach valuable remediation hints and playbooks for familiar scenarios.

By implementing a flexible event tagging mechanism, event sources and detectors can be easily customized to include additional custom event properties. This allows for markdown-formatted event description text that can contain remediation links, as illustrated in the screenshot below.

Root cause remediation hints as markdown links
Figure 3: Root cause remediation hints as markdown links

The Dynatrace Semantic Dictionary helps identify the semantics of well-known event properties and provides convenient platform intents. For instance, entity links (dt.entity.*) or links to the responsible settings entry (dt.settings.object_id) that detected and opened an event can be included. These settings links save valuable time when adjusting detection sensitivity for thresholds or baselines. Additionally, the event setting property can be utilized in a DQL query to create a table of the top-triggering configurations or to automate settings changes using an automation workflow.

Quick access to incident logs

The seamless integration of logs powered by Dynatrace Grail™ data lakehouse with Davis AI root cause analysis is a game changer for modern operation teams, as it offers a quick summary of all incident-relevant logs.

The Dynatrace root cause engine already combines all incident-relevant information to recommend log queries, which saves a lot of navigation time and completely eliminates the need to manually identify complex log filters.

A single click on the Problem details log perspective immediately surfaces all relevant logs related to the given incident, as shown below.

Failure rate increase logs
Figure 4.
100 errors and warnings of failure rate logs
Figure 5.

Within this view the Operations team can further refine the query or adapt the filters and open a notebook to persist the log findings for critical post-mortem documentation purposes.

Root cause analysis in a user-focused context

Most modern application stacks are deployed through Kubernetes, making it essential for operations teams to focus on Kubernetes clusters, cloud resources, and workloads of critical services.

Since operations engineers prefer not to switch contexts, a consistent root-cause experience is provided regardless of where the user journey begins.

Whether you start your remediation journey within the Infrastructure & Operations app or the Kubernetes app, you receive the same root-cause information without needing to navigate between different apps. This seamless embedding of root-cause information into the current context saves valuable time during incident remediation.

Root cause shown in context of the Infrastructure & Operations context.
Figure 6. The root cause is shown in the context of Infrastructure & Operations.
CPU throttling root cause shown in Kubernetes context.
Figure 7. CPU throttling root cause shown in Kubernetes context.

Notify and automate to speed up remediation

The Problems app features a global problem indicator that is always visible within the Dock to capture your attention. This indicator shows whether there are active problems within the environment. You can personalize this number by selecting and saving a problem filter within the problem feed, as demonstrated below. The saved default filter is then automatically applied to the global problem indicator, reducing the number of active problems for the user.

Select Alerting (bell icon) to set up alerts related to filtered problems and configure email addresses for notification recipients.

The email payload and the use of an email address for notifications are preset, allowing for a personalized notification setup, as shown below.

Save the personal default filter and set up email notifications.
Figure 8. Save the personal default filter and set up email notifications.
Find the global problem indicator in the Dock.
Figure 9. Find the global problem indicator in the Dock.

You can take a further step towards answer-driven automation and use the detected Davis problem event to trigger workflow automation. Automatically remediate an issue using our no-code workflow actions for collaboration (for example, Slack, Microsoft Teams, ServiceNow, Pagerduty) and remediation (for example, AWS, Red Hat Ansible, Kubernetes).

The introduction of a filterable global problem indicator ensures that Operations teams remain focused on active problems within the environment, even while exploring data in Notebooks or Dashboards.

In future updates, the Problems app will support multiple named filters and introduce Segments as the primary method for using and sharing numerous predefined filters among operations teams.

Outlook

The newly released Problems app enhances transparency by providing detailed AI-detected root-cause information. It also offers convenient deployment and architectural visualizations, along with a log perspective, to help operations teams reduce Mean Time to Repair (MTTR).

In future updates, we aim to support the ability to acknowledge and label incoming problems, improving team coordination. Additionally, plans include a visual representation of the application map, direct propagation of information such as application IDs into the problem feed, and support for segments to filter the problem feed.

Summary

For over a decade, Dynatrace has been at the forefront of integrating AI into incident analysis, particularly through Davis root cause analysis.

Davis is now essential for Operations, Site Reliability, and DevOps teams, helping them to navigate the complexities of modern software deployments and cloud infrastructure.

Without Davis, the productivity of these teams would plummet, leading to longer Mean Time to Repair (MTTR) and increased staffing needs to handle critical incidents.

In today’s automated deployments and cloud services, traditional observability tools fall short, unable to keep pace with the intelligence needed for effective root cause analysis.

Modern observability encompasses various data sources, from metrics to logs and events, requiring intelligent tools like Davis to seamlessly integrate and analyze this information in real time. By providing Davis on top of the flexible Grail data lakehouse, Dynatrace empowers teams to swiftly reduce MTTR by accessing and previewing incident data comprehensively.

The Davis Problems app streamlines triage, allowing teams to swiftly focus on AI-detected issues. Its intuitive interface simplifies problem resolution.

Try out the new Problems app in the Dynatrace Playground.

The post Transform your operations with Davis AI root cause analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/feed/ 0
Learn how to create a Davis AI anomaly detector on Grail https://www.dynatrace.com/news/blog/create-a-davis-ai-anomaly-detector-on-grail/ https://www.dynatrace.com/news/blog/create-a-davis-ai-anomaly-detector-on-grail/#respond Tue, 11 Jun 2024 19:22:55 +0000 https://www.dynatrace.com/news/?p=64325 Davis AI Fetch logs

From working with Dynatrace Notebooks, you know that exploratory analytics are crucial for uncovering the narratives within your organization’s data. By leveraging visual data analytics and collaboration input from development, security, and business teams, such insights become transparent, enabling immediate understanding and action on the implications for your business. Further, it’s essential to take automated actions to proactively use anomaly detection to determine if your business is at risk. Such anomaly detection should be implemented in straightforward steps, as described in this blog post.

The post Learn how to create a Davis AI anomaly detector on Grail appeared first on Dynatrace news.

]]>
Davis AI Fetch logs
Update: We’ve enhanced anomaly detection on Grail with Dynatrace Intelligence, enabling smarter, AI-powered insights and automated actions across the Dynatrace platform.
Dynatrace Intelligence builds on Davis AI®, advancing how teams detect, analyze, and respond to anomalies in their data.

Dynatrace Grail™ data lakehouse provides contextual analytics across unified observability, security, and business data. It allows you to query and combine data anytime using the Dynatrace Query Language (DQL). This enables exploratory data analysis and the ability to collaborate visually on the results with your colleagues.

Anomaly detection in Notebooks

You likely encounter “why” questions in your daily work. Why did we have an outage? Why did the system behave differently? Why did I receive an alert? These questions can be effectively investigated in Dynatrace Notebooks, where you can easily compile the necessary data and break it down into a time series. However, in the time series example below, we must determine whether the number of access attempts to our example Travel Mobile app is normal or abnormal.

Figure 1: Generated time series based on access logs in Notebooks
Figure 1: Generated time series based on access logs in Notebooks

In many cases, it’s evident, based on your past experiences looking at time series data, whether or not something is an anomaly. But how can you automate your expertise? Such automation could ensure that you and your colleagues don’t have to manually monitor time series to identify whether or not they include anomalies.

Davis® AI provides such automated anomaly detection out of the box. Still, your business requires the flexibility of Davis AI to detect anomalies based on your specific requirements, for example, to automatically generate a Davis problem based on a detected anomaly. For this purpose, we provide the Davis AI Analyzer, which allows you to select a specific analyzer. Three anomaly detection analyzers are available, each equipped with unique mechanisms to detect anomalies in your data that significantly deviate from the norm.

One unique feature of the Davis AI Analyzer is that it works on any time series, regardless of its origin—whether generated with makeTimeseries from events, business events, logs, or other sources or the joining of different time series. As you can see in the screenshot below, Davis AI Analyzer gains the full power of DQL, making Davis anomaly detection even more flexible and stronger than ever. This power can be easily experienced by selecting the desired Davis anomaly detection analyzer in Notebooks or Dashboards.

Figure 2: Using the seasonal baseline anomaly detection analyzer in Notebooks.
Figure 2: Using the seasonal baseline anomaly detection analyzer in Notebooks.

By selecting the seasonal baseline analyzer, Davis AI recognizes that the number of attempted accesses to the app in this example doesn’t deviate from the norm based on the past data during the same period. The time series falls within the seasonal green confidence band. A potential alert would be visually simulated if the time series fell outside this band.

This anomaly detector observes the number of attempted accesses per minute and triggers an event when anomalies are detected. You can create a similar Davis anomaly detector in a few simple steps.

Automate your experience with Davis Anomaly Detection

In Notebooks, select open with and choose Davis Anomaly Detection; all settings required for creating an anomaly detector will be carried over.

Create a new anomaly detector in Davis Anomaly Detection.
Figure 3: Create a new anomaly detector in Davis Anomaly Detection.

The new anomaly detector is created in four steps; the first two steps are carried over automatically from Notebooks. Let’s start with the most straightforward step, Get started, where you define a title for your anomaly detector and a description for the configuration.

The next two steps, as mentioned, have already been prefilled from Notebooks. In the Configure your query step, you’ll find the DQL query you predefined, and in the Customize parameters step, you’ll find your selected anomaly detection analyzer. The last significant step, the Create an event template step, remains. Here, you can define the template for your event and describe all essential information for the subsequent process.

Define the description and properties in the event template.
Figure 4: Define the description and properties in the event template.

What makes this template exceptional is that you can use {placeholder} hints to add additional context to the text about the event. For example, the value of the violation or the source entity where the anomaly was detected. This ensures that all essential information about the event is immediately visible to the Site Reliability Engineer (SRE). After completing all four steps, we can create the Davis anomaly detector by selecting Create. The anomaly detector will automatically monitor your defined time series every minute and trigger your specified event upon detection of an anomaly.

The new anomaly detector is now listed in Davis Anomaly Detection. Here, you’ll find all anomaly detector configurations, and you can filter them according to your specific criteria. Additionally, you can expand this table with extra information about the configurations, such as when the anomaly detectors were last modified.

Overview of anomaly detectors available within Davis Anomaly Detection.
Figure 5: Overview of anomaly detectors available within Davis Anomaly Detection.

Of course, you always have the option to reopen an anomaly detector directly in Notebooks, where all configuration settings are carried over. You also have the option to display a preview of your anomaly detector directly in Davis Anomaly Detection.

Figure 6: Visualize your custom anomaly detectors in Notebooks without leaving Davis Anomaly Detection.
Figure 6: Visualize your custom anomaly detectors in Notebooks without leaving Davis Anomaly Detection.

The exciting challenge is finding answers to your everyday “why” questions using Grail and DQL analytics capabilities. If the answer is successfully identified in a time series and you want to automate the result with anomaly detection, this can be done in just a few steps. We recommend you explore the new Davis Anomaly Detection analyzer in Notebooks; we’re confident you’ll quickly discover its many uses.

Try out Davis Anomaly Detection

Want to know more? Check out the following video, in which Andreas Grabner and I collaborated on a new episode of the Dynatrace Observability Clinic. Here, we share a live introduction to Anomaly Detection based on DQL.

We also recommend watching the exciting use case for Anomaly Detection and the 5 Pillars of Data Observability.

What’s next

Davis Anomaly Detection is automatically enabled for all Dynatrace SaaS environments with the release of Dynatrace version 1.291. No effort is needed from your side. We’re, of course, highly interested in your feedback. So, please head to the Dynatrace Community and share your suggestions and product ideas to help us continuously improve Dynatrace Anomaly Detection.

Are you interested in learning more? In Dynatrace Documentation, you can learn more about Davis Anomaly Detection and how to use anomaly detection within Notebooks, or look at our Playground, where you can explore practical examples of how to utilize Davis AI Analyzer in your anomaly detection.

See examples of using Davis AI to detect anomalies. Visit Dynatrace Playground.

The post Learn how to create a Davis AI anomaly detector on Grail appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/create-a-davis-ai-anomaly-detector-on-grail/feed/ 0
Auto-adaptive thresholds for AI-driven quality gating https://www.dynatrace.com/news/blog/auto-adaptive-thresholds-for-ai-driven-quality-gating/ https://www.dynatrace.com/news/blog/auto-adaptive-thresholds-for-ai-driven-quality-gating/#respond Tue, 04 Jun 2024 17:12:50 +0000 https://www.dynatrace.com/news/?p=64258 auto-adaptive thresholds

The Site Reliability Guardian automates the validation process for new software releases or changes. These validations involve setting specific performance, availability, or security objectives that must be met. For example, response time or failure rate thresholds can be used to define the desired state, warning range, or when an objective is violated. However, it can be difficult or even impossible to set these targets upfront for a new software component because it's unclear how the component will behave or what circumstances it will face.

The post Auto-adaptive thresholds for AI-driven quality gating appeared first on Dynatrace news.

]]>
auto-adaptive thresholds

The latest release of the Site Reliability Guardian incorporates assistance from Dynatrace Davis® AI to automatically derive appropriate threshold targets and adjust them over time to protect your quality improvements. This process, known as auto-adaptive thresholding, eliminates the need to define a static threshold upfront. Instead, it derives the suitable thresholds from previous validation results.

Build an umbrella for Development and Operations

In modern software engineering, the discipline of platform engineering delivers DevSecOps practices to developers to bridge the gaps between development, security, and operations and enhance the developer experience. A key element in platform engineering is the establishment of fast feedback cycles regarding the quality and security measures of new software releases. To provide automated feedback for developers, the concept of quality gates for static code analysis in continuous integration is widely adopted throughout the industry. However, this perspective differs in the continuous deployment practices of various organizations, where the feedback is either delayed or not returned to the developer.

While receiving no feedback on the quality or security of new features leaves developers uncertain about feature performance, delayed feedback also increases a developer’s cognitive load. The developer must pause their current engineering work to address the reported issue and consider the code changes they worked on a few days or weeks prior.

To reduce developers’ cognitive load by providing timely information, platform engineers must create tools that allow validations to be run in the early phases of development, with direct and fast feedback loops. Ideally, this should be a self-service offering that enables individual adoption by teams. While platform engineers can build and prepare the necessary infrastructure and templates for self-adoption, developers must still provide some customization. For example, the team must establish specific thresholds for desired service performance behavior. Setting these thresholds upfront can be challenging because the team might not know how a service will behave in its environment.

How we define auto-adaptive thresholds at Dynatrace

This blog post explores how Dynatrace leveraged the Site Reliability Guardian to establish a fast feedback loop for Davis AI model improvements. The conducted validations avoided regression within the models, and the outcomes were immediately fed back to the data science team when deviations were detected. While the data science team appreciated the quick insights and validations of their improvements, they initially struggled with the setup. Consequently, this blog post highlights the new capability of the Site Reliability Guardian to define auto-adaptive thresholds that tackle the challenge of configuring static thresholds and protect your quality and security investments with relative comparisons to previous validations.

Fast feedback cycles on model improvements

While the Site Reliability Guardian was originally designed to validate new software releases, Dynatrace has internally extended its application area to include validation of models for Davis AI.

The Dynatrace data science team continuously improves the machine learning models used by Davis AI, for example, by adding new features to forecasting or refining mathematical calculations. A single change can influence multiple models, as features are often used across several models. To ensure that model changes don’t lead to regressions, the data scientists set up Site Reliability Guardian, which is automatically triggered whenever a change is made in the codebase via CI/CD pipelines.

A series of models are continuously trained on Dynatrace tenants to effectively set objectives. The training times and other quality metrics, such as the RMSE (Root Mean Squared Error), SMAPE (Scaled Mean Absolute Percentage Error), and coverage probability, are monitored using Dynatrace. Our data scientists utilize metrics and events to store these quality metrics. However, other data formats, like logs, can also be employed. The quality metrics can then be easily queried using DQL and utilized for the objectives of a guardian. Validations are automatically triggered when a change is committed to the code base via the Dynatrace API. This helps data scientists quickly respond to recently introduced regressions. For instance, if an objective is violated, they’re immediately notified, for example, through a Slack channel.

Validation history
Figure 1. Validation history

One difficulty encountered when setting up objectives in the guardians was selecting an appropriate threshold for the quality metrics, as this is typically heavily dependent on the data.

Leverage Davis AI to quickly start validating

To address the challenge of defining a static threshold for an objective, the Site Reliability Guardian enables switching objective thresholds to auto-adaptive mode, as depicted in the screenshot below.

Activating an auto-adaptive threshold for the response time objective
Figure 2. Activating an auto-adaptive threshold for the response time objective

Davis AI controls auto-adaptive thresholds. It analyzes the next five validations to derive this objective’s proper warning and failure thresholds. Once the learning phase is complete, all subsequent validation results are fed into Davis AI to fine-tune the thresholds based on changed behavior.

Considering previous validation results, the latest validation is always relative to the past, protecting quality and security improvements. If, for example, recent performance improvements change a service’s response time from 200 ms to 175 ms, the auto-adaptive threshold is adjusted to reflect the new behavior. Nevertheless, the Site Reliability Guardian detects sudden behavior changes by reporting a warning or failure if response time returns to 200 ms or above.

Learning phase

Unless an objective has been validated five times, it’s still in the learning phase. During this phase, the measured values are informative, allowing observation of the objective’s development without affecting deployment or delivery processes. The Site Reliability Guardian denotes the learning phase of an objective with a loading symbol on the heatmap and in the objective details.

Representation of the learning phase of an auto-adaptive threshold
Figure 3. Representation of the learning phase of an auto-adaptive threshold

The warning and failure thresholds will be automatically set if sufficient validations are available to establish a solid baseline for an objective’s auto-adaptive thresholds. Consequently, the next objective validation will impact the overall validation result.

Auto-adaptive thresholds as code

To enhance the developer’s experience in adopting the Site Reliability Guardian in a self-service manner, the configuration for a guardian and its workflow can be provided in a configuration-as-code fashion. This enables the management of the configuration within the service’s code repository. Incorporating the new capability of auto-adaptive thresholds into configuration-as-code is as simple as adding the autoAdaptiveThresholdEnabled flag to an objective.

Configuration-as-code example for activating an auto-adaptive threshold
Figure 4. Configuration-as-code example for activating an auto-adaptive threshold

Before concluding, we wish to announce the change of the Site Reliability Guardian icon from purple to shiny golden. This new appearance enhances the icon’s geometry while preserving its core values. Therefore, the icon continues to feature the infinity loop as a symbol for the DevSecOps loop and the fast forward sign to expedite delivery while ensuring quality and security standards.

The new and shiny appearance of Site Reliability Guardian

Evolution of the Site Reliability Guardian icon
Figure 5. Evolution of the Site Reliability Guardian icon

What’s next

The new auto-adaptive thresholds capability is now available in Site Reliability Guardian. Open the app and switch from static to auto-adaptive thresholds for those objectives where Davis AI should derive the thresholds for you. For full details, see Dynatrace Documentation.

If you haven’t used Site Reliability Guardian yet, try it out in the Dynatrace Playground or watch the latest app spotlight recording.

The post Auto-adaptive thresholds for AI-driven quality gating appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/auto-adaptive-thresholds-for-ai-driven-quality-gating/feed/ 0
Unified observability delivers deeper insights with AI-driven analytics and automation https://www.dynatrace.com/news/blog/ai-driven-analytics-and-automation-for-unified-observability/ https://www.dynatrace.com/news/blog/ai-driven-analytics-and-automation-for-unified-observability/#respond Mon, 12 Feb 2024 21:23:39 +0000 https://www.dynatrace.com/news/?p=62396 Perform 2024: Make waves

Today’s organizations flock to multicloud environments for myriad reasons, including increased scalability, agility, and performance. However, these environments can drown enterprises in data, forcing them to adopt multiple tools and services to manage and secure it. This fragmented approach adds complexity and opens the door to security vulnerabilities. In fact, according to recent Dynatrace research, […]

The post Unified observability delivers deeper insights with AI-driven analytics and automation appeared first on Dynatrace news.

]]>
Perform 2024: Make waves

Today’s organizations flock to multicloud environments for myriad reasons, including increased scalability, agility, and performance. However, these environments can drown enterprises in data, forcing them to adopt multiple tools and services to manage and secure it. This fragmented approach adds complexity and opens the door to security vulnerabilities.

In fact, according to recent Dynatrace research, 85% of technology leaders say the number of tools, platforms, dashboards, and applications they use adds to the complexity of managing a multicloud environment. Further, 84% of technology leaders say multicloud complexity makes it harder to protect applications from security vulnerabilities and attacks.

With unified observability and security, organizations can protect their data and avoid tool sprawl with a single platform that delivers AI-driven analytics and intelligent automation.

During a Dynatrace Perform 2024 breakout session, Dynatrace colleagues Bipin Singh, product marketing director, and Markie Duby, principal solutions engineer, showed how organizations can bring together observability, security, and business data from cloud-native and multicloud environments with Dynatrace.

Update: We’ve expanded Dynatrace Intelligence, extending AI-powered insights across the Dynatrace platform. Dynatrace Intelligence is the evolution of Davis AI®, delivering deeper observability and more actionable intelligence.

The secret sauce of unified observability

Observability enables teams to measure a system’s state based on the data it generates. A unified observability approach takes it a step further, enabling teams to monitor and secure their full stack on an AI-powered data platform.

With the Davis AI engine, Grail data lakehouse, and Smartscape topology visualization at its core, the Dynatrace unified observability and security platform provides AI-driven analytics and automation capabilities.

An overview of the Dynatrace unified observability and security platform.
An overview of the Dynatrace unified observability and security platform.

“Grail handles data storage, data management, and processes data at massive speed, scale, and cost efficiency,” Singh said. “Smartscape contextualizes your entire environment and builds a real-time topology map that’s dynamic and stays up to date as your environment changes. And the Davis AI engine is continuously watching your environment and evaluating the emerging situation, automatically detecting problems, creating automated root-cause analysis for you and business impact analysis for prioritization.”

The importance of hypermodal AI to unified observability

Artificial intelligence is a critical aspect of a unified observability strategy. In fact, according to the recent Dynatrace report, “The state of AI 2024,” 83% of technology leaders say AI has become mandatory to keep up with the growing complexity of multicloud environments.

The Davis AI engine uses a hypermodal approach to bring together causal, predictive, and generative AI. Causal AI determines the underlying causes and effects of issues based on the system’s topology. Predictive AI, meanwhile, makes predictions about future events based on patterns from historical data. And generative AI, termed Davis CoPilot, creates queries, notebooks, and dashboards to simplify analytics, and provides workflow and automation recommendations.

By bringing together these AI types, organizations receive generative AI recommendations based on the precise context from predictive and causal AI. This coactive AI approach enables organizations to spend more time on innovation by simplifying and automating routine tasks.

A breakdown of how Grail, Smartscape, and Davis work together in the Dynatrace unified observability and security platform.
A breakdown of how Grail, Smartscape, and Davis work together.

How Davis tackles root cause for AI-driven analytics

Duby discussed how Dynatrace OneAgent, Smartscape, and Davis work together to take information from many different layers in a full stack to provide root-cause analysis.

“When [Davis is] going through and detecting anomalies within your environment, it’s using data both from the underlying interdependency link as well as that end-to-end trace from the end user all the way back in order to do things like root-cause analysis,” Duby said. “We’re using that causal AI to determine what is actually the underlying root cause.”

A visual representation of what Davis uses for its own analysis in the Dynatrace unified observability and security platform.
A visual representation of what Davis uses for its own analysis.

Davis enables users to go deeper into the details of the underlying processes running on a particular host. The hypermodal AI engine shows what’s happening in a system down to the data coming in, while presenting the information in context.

“It’s one thing to have the data; it’s another thing to have it in context,” Duby continued. “For performance, for security analytics, you have to have the data in context. You need to understand how these different pieces interact with each other and how those pieces are actually coming through.”

Once Davis has gathered all the necessary information throughout the different layers of the stack, it can determine what’s changing, what’s breaking, where the issues are, and how to resolve them. Additionally, it helps users prioritize which issues need immediate attention by providing the necessary context.

“[Davis is] looking at the business context—not just the IT, not just the individual metrics, but understanding the whole picture,” Duby said.

How Davis CoPilot takes AI further and promotes collaboration

A significant piece of the Dynatrace hypermodal AI approach is Davis CoPilot, the generative AI part of the hypermodal engine. Davis CoPilot enables users to create queries, dashboards, and notebooks using natural language input, while offering coding suggestions for workflow automation. Additionally, it simplifies the processes of onboarding, configuring, and adopting the Dynatrace unified observability and security platform.

“This is Davis CoPilot. This is your helper to make sure you can actually go in and take advantage of all of that underlying data,” Duby said. “So, you have the analytics and the performance tracking that Dynatrace is doing, and you also have the ability to build it for your own custom use case.”

A preview of Davis CoPilot returning results from a natural-language input.
A preview of Davis CoPilot returning results from a natural-language input.

In addition to creating queries, dashboards, and notebooks using natural language, Davis CoPilot enables users to share these notebooks with other team members across the organization, boosting collaboration efforts.

“[Davis CoPilot will] start building out queries. And now I can take this information that I just got back, and I can share this notebook with my colleague. And now they have the exact same information,” Duby continued. “I can reuse the same report next month and see if it changed. … This functionality allows me to collaborate with my team. It allows me to run with a new idea and see what comes back. And then I can take this information, and I can build on top of it to do more advanced analytics for my different teams.”

For an in-depth demonstration of using the Dynatrace unified observability and security platform, watch our on-demand session, “AI-driven analytics and automation for unified observability and security.” And for more coverage from Perform 2024, check out our guide.

The post Unified observability delivers deeper insights with AI-driven analytics and automation appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ai-driven-analytics-and-automation-for-unified-observability/feed/ 0
Introducing Dynatrace built-in data observability on Davis AI and Grail https://www.dynatrace.com/news/blog/introducing-dynatrace-built-in-data-observability-on-davis-ai-and-grail/ https://www.dynatrace.com/news/blog/introducing-dynatrace-built-in-data-observability-on-davis-ai-and-grail/#respond Wed, 31 Jan 2024 17:00:22 +0000 https://www.dynatrace.com/news/?p=61558 Database observability graphic

“Great! I have ingested important custom data into Dynatrace, critical to running my applications and making accurate business decisions… but can I trust the accuracy and reliability?” Welcome to the world of data observability. The Dynatrace open platform is well-positioned to take advantage of the exponential increase in data generation. However, coupled with the increase […]

The post Introducing Dynatrace built-in data observability on Davis AI and Grail appeared first on Dynatrace news.

]]>
Database observability graphic

“Great! I have ingested important custom data into Dynatrace, critical to running my applications and making accurate business decisions… but can I trust the accuracy and reliability?”

Welcome to the world of data observability.

The Dynatrace open platform is well-positioned to take advantage of the exponential increase in data generation. However, coupled with the increase of external data sources that can now be ingested, there are new challenges in data management that need to be addressed.

 “Every year, poor data quality costs organizations an average $12.9 million”
– Gartner

Data observability is a practice that helps organizations understand the full lifecycle of data, from ingestion to storage and usage, to ensure data health and reliability. Data observability involves monitoring and managing the internal state of data systems to gain insight into the data pipeline, understand how data evolves, and identify any issues that could compromise data integrity or reliability. At its core, data observability is about ensuring the availability, reliability, and quality of data.

Data observability is crucial to analytics and automation, as business decisions and actions depend on data quality. In the age of AI, data observability has become foundational and complementary to AI observability, data quality being essential for training and testing AI models.

Dynatrace now addresses many of the issues customers experience around the health, quality, freshness, and general usefulness of data that is externally sourced into Dynatrace Grail™, allowing them to make better-informed decisions and optimize their efforts for digital transformation and data-driven operations.

The rise of data observability in DevOps

Data forms the foundation of decision-making processes in companies across the globe. Data is the foundation upon which strategies are built, directions are chosen, and innovations are pursued. Consequently, the importance of continuously observing data quality, and ensuring its reliability, is paramount. Surveys from our recent Automation Pulse Report underscore this sentiment: 57% of C-level executives say the absence of data observability and data flow analysis makes it difficult to drive automation in a compliant way. This not only underscores the universal significance of data, it also hints at its pivotal role within DevOps. For DevOps teams that inform deployment strategies, optimize processes, and drive continuous improvement, the integrity and timeliness of data are of significant importance.

As organizations scale and accelerate their digital transformation journeys, a major hurdle to proper DevOps adoption is the trustworthiness of the massive volume of data coming from various sources, much of which goes into data silos such as log management tools, SIEM solutions, and others.

The rise of data observability needs is where Dynatrace capabilities around Grail, analytics, and Davis® AI are in an outstanding and unmatched position to deliver the currently missing value to the market: a leading and single solution for all data observability analytics needs. This reduces the demand for further data flow analysis tools and clears any hurdles to making data useable for DevOps automation use cases.

Davis AI, Grail, and data observability

By grouping common data observability issues into industry-standard pillars, we can provide tangible examples and showcase current capabilities. The five pillars we focus on are freshness, volume, distribution, schema, and lineage.

Freshness: Timeliness of data

In an ideal ecosystem, actionable data should be as recent as possible, supported by learnings from accurate, historical data. Observing the freshness of data helps to ensure that decisions are based on the most recent and relevant information.

Scenario: Due to an undetected configuration issue, a flight status system from a popular airline had been buffering data for the last two hours before sending it on in one batch. Downstream dashboards and system automations were using outdated data, leading to incorrect statuses of flights in reports.

Solution: After setting up data ingestion into Grail, Dynatrace Query Language (DQL) is used to add a freshness field (Figure 1) which is calculated from the delta between when the signal was written and when it was ingested. This freshness measurement can then be used by out-of-the-box Dynatrace anomaly detection to actively alert on abnormal changes within the data ingest latency to ensure the expected freshness of all the data records. Furthermore, the new Alert on missing data feature in the Anomaly Detector panel can be used to trigger notifications when data is not coming in as expected after being baselined.

Value: The possibility of alerting on data freshness issues, based on a learned baseline through Davis AI, allows for faster time-to-detect where there are seemingly no infrastructure issues. Normally this would have left an issue undetected for much longer, providing a false sense of security, eventually leading to a much bigger customer and monetary impact for the organization.

Use of Dynatrace Notebook to track when a flight status table was last updated.
Figure 1. Use of Dynatrace Notebook to track when a flight status table was last updated.

Volume: Quantity of data generated or processed within a given timeframe

Unexpected increases or drops in the volume of data are often a good indication of an undetected issue.

Scenario: For many B2B SaaS companies, the number of reported customers is an important metric. It heavily influences downstream reports, and dashboards, shaping decisions from daily operations to strategic monthly reviews. In this scenario, a manually triggered run of a production pipeline had the unintended consequence of duplicating the reported customer metric. If left unchecked, this misrepresentation of a single KPI could lead to misguided decision-making processes through multiple layers of the organization.

Solution: Like the freshness example, Dynatrace can monitor the record count over time. Once a DQL query has been set up, it can be used in an automation workflow (Figure 2) where scheduling, prediction, comparison to actual value, and, finally, alerting are all taken care of to enable a fully flexible way to detect anomalies in data volume.

Value: KPIs and metrics such as the number of reported customers are central to an organization’s business and strategic processes. Any issues here will result in a loss of trust in the data, and, if left undetected, they will eventually lead to monetary impact, including loss of reputation for an organization.

Using Dynatrace Workflows to alert on data volume anomalies
Figure 2 Using Dynatrace Workflows to alert on data volume anomalies

Distribution: The statistical spread or ranges of data

The distribution of data is essential in identifying patterns, outliers, or anomalies in the data. Deviation from the expected distribution can signal an issue in data collection or processing.

Scenario: A financial institution processes millions of transactions daily, ranging from credit card purchases and mortgage payments to interbank transfers and ATM withdrawals. An erroneous change in the database system leads to a subset of the data being categorized incorrectly. After several days, the fraud detection system starts triggering on a frequent basis, and liquidity management dashboards begin showing questionable values.

Solution: Baselining and raising alerts on anomalies are core capabilities of Davis AI. After setting up ingestion for the data that you want to monitor, it’s simple to use Dynatrace full AI capabilities to observe and alert on any anomalies in the data. In the example above, ingesting the number of transactions as business events, anomaly detection could be based on this to proactively alert and trigger mitigation activities.

Value: While variations are expected in financial trends, anomalies should be auto-detected, and manual detection should not be relied on. Earlier detection of these issues will keep the fallout as low as possible.

Schema: Structure and relationships of data between entities

Observing the schema can help identify and flag unanticipated changes, such as the addition of new fields or deletion of existing fields.

Scenario: An externally connected database system made an update that inadvertently dropped the account_id column in the customers table. The automated data pipeline propagated these changes, leading to downstream reports, dashboards, and applications breaking as the previous field reference is now missing.

Solution: Using the DQL FieldsSummary command, we can keep track of the number of distinct field keys within a given family of data records. Once confirmed in a notebook, the number of field keys can be used in an automated workflow to continuously monitor the count and write it back to a new metric (Figure 3). Once the new metric is established, out-of-the-box Dynatrace anomaly detection can be used to alert on either a static threshold or a learned baseline.

Value: Observing incoming data Schemas, and thus placing expectations on what the external data should look like and must contain, allows for pro-active alerting and mitigation of issues long before they can lead to widespread business impact such as broken reports, dashboards, or further analytics on top of the data.

Keeping track of the field count in a new metric (data.observability.fields) using Workflows and Typescript.
Figure 3. Keeping track of the field count in a new metric (data.observability.fields) using Workflows and Typescript.

Lineage: Journey of data through a system

Data lineage provides insights into where the data came from (upstream) and what is impacted (downstream). It plays a crucial role in root cause analysis as well as informing impacted systems about an issue as quickly as possible.

Scenario: The hourly_consumption table was deprecated and removed by an overzealous database administrator as there were no known downstream consumers of this data, breaking a monthly integration check used for consumption reporting for shareholders.

Solution: In the future, Dynatrace Smartscape® could be used, which already builds a dependency graph, to enable a data lineage view. This would enable faster root cause analysis of any data-related problems, as well as allow for easy notification of downstream consumers who would be impacted.

Value: A proper understanding of the source of the data, as well as where it is used, helps drive down time-to-alert and time-to-repair. Time-to-alert is achieved by quickly and automatically alerting those who are impacted by a data issue by quickly understanding downstream consumers of the data, while time-to-repair informs on the source of where the data originated from, to quickly drill down into those systems.

Data availability: A prerequisite

You could implement the most contemporary, accurate, and useful data observability solution possible, but what good will it be if all the data simply does not arrive as expected? Broken pipelines or missing data sources would mean that there is simply no data to observe and that data may never arrive, forever lost.

A truly valuable data observability solution should be able to alert on data issues as early in the process as possible. This requires monitoring of the upstream infrastructure, applications, or platform supporting those data streams. This is where the power of Dynatrace end-to-end observability comes into play. Dynatrace can leverage existing Infrastructure Monitoring and Application Observability solutions to surface problems that can affect later data observability workstreams—long before a traditional data observability solution would pick up the issue.

Leverage the power of Dynatrace and Davis AI—now and into the Future

Anomaly Detection

Anomaly detection is grounded in the idea of baselining typical patterns of ingested data, designed to alert where a change or deviation from the norm is observed. These patterns typically go beyond simple flat or trend lines, often exhibiting complex seasonal behaviors, such as business hours or weekly patterns related to the industry. Dynatrace is particularly strong in this area: Davis predictive AI has been enriched over the years with a set of advanced machine learning (ML) algorithms optimized for time-series observability datasets to cope with these challenges. Davis AI anomaly detection, leveraging these ML algorithms, can already be used on the results of DQL queries. (Embedding ML algorithms into DQL as functions is on the Dynatrace platform roadmap.)

Considering the examples and solutions provided above, anomaly detection plays a pivotal role in numerous data observability use cases and can be harnessed to effectively address these challenges.

Triage and resolution of a data incident

Triaging requires an ability to identify the root cause of a data incident, which is particularly challenging as an organization scales up the volume and speed of data ingest typical of an enterprise environment. It’s easy to see how Davis causal AI problem detection could be extended in the future to identify root-cause data observability issues.

Depending on the incident, there might be different paths to resolution. One acceptable path could be full auto-remediation, whereby Dynatrace AutomationEngine could be triggered, scripts executed, permissions granted, security checked, and data corrected. A second path might require Jira tickets to be created and human intervention through an approval process. A data problem alert could be used as the event allowing for multiple methods to notify the correct data owners, stewards, governors, or data teams.

An incident requires not only resolution but also understanding and alerting upstream data providers and downstream data subscribers to the potential impact. Dynatrace is strong on the observability of data pipelines ingesting data into Grail and consumers of Grail data, although this is an area that will be enhanced and improved in the future product roadmap.

Prevention of future incidents

Not all data quality incidents can be prevented, especially because ELT/ETL data pipelines typically tend to grow over time and span many different heterogeneous collectors that have different ownerships. There are, however, mitigation techniques you can use, for example:

  • Health tracking of key datasets or streams over time—alerting on anomalies
  • Monitoring standard query results and changes over time
  • Well-designed, data-focused dashboards for monitoring
  • Auto remediation where appropriate with built-in audit logging
  • Forensic abilities for ad-hoc data analysis

Summary

Dynatrace is uniquely positioned to provide even more value by extending our world-class observability platform into the data observability realm. To achieve this, we leverage Infrastructure Monitoring and Application Observability for early warnings on data pipeline issues and use DQL, Workflows, and Grail for data observability—all enabled by our best-in-class Davis AI engine.

Ensuring the quality and reliability of underlying data is more crucial than ever now that many organizations are deploying Generative AI models. Data observability is becoming a mandatory part of business analytics, automation, and AI. Davis AI and data observability together uniquely ensure the quality and reliability of data at the level of hypermodal AI—predictive, causal, and generative.

You can now monitor sources and incoming data pipelines for freshness, volume, distribution, lineage, and availability issues early on without added noise and in a central location, the Dynatrace platform. This gives your teams additional confidence over data quality, saves time, prevents inaccurate analyses and automation outcomes, leads to more trustworthy AI models, and supports efforts to consolidate or reduce the number of IT tools they rely on.

Ready to get started with Dynatrace data observability? For complete details, best practices, and detailed use cases, see Dynatrace data observability documentation.

The post Introducing Dynatrace built-in data observability on Davis AI and Grail appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/introducing-dynatrace-built-in-data-observability-on-davis-ai-and-grail/feed/ 0
Enhanced AI model observability with Dynatrace and Traceloop OpenLLMetry https://www.dynatrace.com/news/blog/enhanced-ai-model-observability-with-dynatrace-and-traceloop-openllmetry/ https://www.dynatrace.com/news/blog/enhanced-ai-model-observability-with-dynatrace-and-traceloop-openllmetry/#respond Mon, 04 Dec 2023 18:32:01 +0000 https://www.dynatrace.com/news/?p=60953 Enhancing AI model observability

In the rapidly evolving landscape of artificial intelligence, ensuring your AI model’s optimal performance, reliability, security, and user trust is paramount. This blog post explores how combining the Dynatrace full stack observability platform and Traceloop's OpenLLMetry OpenTelemetry SDK can seamlessly provide comprehensive insights into Large Language Models (LLMs) in production environments. Observing AI models enables you to make informed decisions, optimize performance, and ensure compliance with emerging AI regulations.

The post Enhanced AI model observability with Dynatrace and Traceloop OpenLLMetry appeared first on Dynatrace news.

]]>
Enhancing AI model observability

“Engineers today lack an easy way to track the tokens and prompt usage of their LLM applications in production. By using OpenLLMetry and Dynatrace, anyone can get complete visibility into their system, including gen-AI parts with 5 minutes of work.”

Nir Gazit, CEO and Co-Founder Traceloop

Why AI model observability matters

The adoption of LLMs has surged across various industries, particularly since the introduction of OpenAI’s GPT model. While these models yield impressive results, the challenge of maintaining their operation within defined boundaries has increased.

AI model observability plays a crucial role in achieving this by addressing these key aspects:

  1. Model performance and reliability: Evaluating the model’s ability to provide accurate and timely responses, ensuring stability, and assessing domain-specific semantic accuracy.
  2. Resource consumption: Observing computational resource availability and saturation, whether deployed in cloud-native environments like Kubernetes or CPU-enabled servers.
  3. Data quality and drift: Monitoring the quality and characteristics of training and runtime data to detect significant changes that might impact model accuracy.
  4. Explainability and interpretability: Providing information on model versions, parameters, and deployment schedules, which is essential for interpreting and understanding model answers.
  5. Security and compliance: Actively preventing security threats at both the application and model levels to ensure responsible and compliant AI usage.

The challenge of AI model observability

One challenge in AI model observability is the diverse tooling landscape required to gain critical insights. OpenTelemetry has become a standard for collecting traces, metrics, and logs. However, seamless support for various SDKs and AI model frameworks, such as LangChain and Pinecone, remains essential.

Combining Dynatrace with Traceloop’s OpenLLMetry addresses the heterogeneity challenge by supporting a range of popular LLMs, prompt engineering, and chaining frameworks. OpenLLMetry, an open source SDK built on OpenTelemetry, offers standardized data collection for AI Model observability.

How OpenLLMetry works

OpenLLMetry supports AI model observability by capturing and normalizing key performance indicators (KPIs) from diverse AI frameworks. Utilizing an additional OpenTelemetry SDK layer, this data seamlessly flows into the Dynatrace environment, offering advanced analytics and a holistic view of the AI deployment stack.

Given the prevalence of Python in AI model development, OpenTelemetry serves as a robust standard for collecting observability data, including traces, metrics, and logs. While OpenTelemetry’s auto-instrumentation provides valuable insights into spans and basic resource attributes, it falls short in capturing specific KPIs crucial for AI models, such as model name, version, prompt and completion tokens, and temperature parameters.

OpenLLMetry bridges this gap by supporting popular AI frameworks like OpenAI, HuggingFace, Pinecone, and LangChain. Standardizing the collection of essential model KPIs through OpenTelemetry ensures comprehensive observability. The open source OpenLLMetry SDK, built atop OpenTelemetry, enables thorough insights into your Large Language Model (LLM) applications.

As the collected data seamlessly integrates with your Dynatrace environment, you can analyze LLM metrics, spans, and logs in the context of all traces and code-level information. Maintained under the Apache 2.0 license by Traceloop, OpenLLMetry is a valuable asset for product owners, providing a transparent view of AI model performance.

The diagram below illustrates how OpenLLMetry captures and transmits AI model KPIs to your Dynatrace environment, empowering your business with unparalleled insights into your AI deployment landscape.

Enhancing AI model observability

Dynatrace OneAgent® is perfectly capable of automatically injecting and tracing code-level information for many technologies, such as Java, .NET, Golang, and NodeJS. However, Python models are trickier.

In the Dynatrace web UI, you can track your AI model in real time, examine its model attributes, and assess the reliability and latency of each specific LangChain task, as demonstrated below.

LangChain task distributed traces in Dynatrace screenshot

The captured span by Traceloop automatically displays vital details, including the mode utilized by our LangChain model gpt-3-5-turbo, the model’s invocation with a temperature parameter of 0.7, and the utilization of 53 completion tokens for this individual request.

LangChain task distributed traces in Dynatrace screenshot

With the growth of AI, maintaining transparency is essential

Observing AI models like Large Language Models (LLMs) in production is crucial for enhancing performance, reliability, security, and user trust. This includes the monitoring of AI-related costs to ensure they remain within acceptable margins. The Dynatrace platform, coupled with Traceloop’s OpenLLMetry OpenTelemetry SDK, offers comprehensive visibility from model inception to completion.

As AI adoption grows, maintaining transparency is essential for regulatory compliance. While Dynatrace automates tracing for various technologies, Python-based AI models require OpenTelemetry. OpenLLMetry bridges this gap, supporting popular AI frameworks and vendors to ensure standardized data collection. OpenLLMetry provides an open source SDK for LLM observability, seamlessly integrating with Dynatrace for in-depth analysis.

References

The post Enhanced AI model observability with Dynatrace and Traceloop OpenLLMetry appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enhanced-ai-model-observability-with-dynatrace-and-traceloop-openllmetry/feed/ 0
Automate predictive capacity management with Davis AI for Workflows https://www.dynatrace.com/news/blog/automate-predictive-capacity-management-with-davis-ai-for-workflows/ https://www.dynatrace.com/news/blog/automate-predictive-capacity-management-with-davis-ai-for-workflows/#respond Tue, 11 Jul 2023 20:19:17 +0000 https://www.dynatrace.com/news/?p=58574 predictive capacity management

>> Scroll down to see predictive capacity management in action (14-second video) Our recent blog post, Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI, showed how Dynatrace Notebooks are used to predict the future behavior of time series data stored in Grail™. This follow-up post introduces Davis® AI for […]

The post Automate predictive capacity management with Davis AI for Workflows appeared first on Dynatrace news.

]]>
predictive capacity management
>> Scroll down to see predictive capacity management in action (14-second video)

Our recent blog post, Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI, showed how Dynatrace Notebooks are used to predict the future behavior of time series data stored in Grail™. This follow-up post introduces Davis® AI for Workflows, showing you how to fully automate prediction and remediation of your future capacity demands. The anticipation of future capacity demands makes it possible to completely avoid critical outages by notifying you days in advance, well before incidents arise.

Predictive capacity management starts within a Dynatrace Notebook, where the operations team explores important capacity indicators, such as the percentage of free disks, as shown below.

Figure 1. Example forecast of remaining disk capacity with upper/lower bounds and an anticipated value.
Figure 1. Example forecast of remaining disk capacity with upper/lower bounds and an anticipated value.

After exploring and selecting the most important capacity indicators for your environment, a workflow triggers forecast reporting at regular intervals. The example workflow below is triggered every Monday at 8:00 AM to provide a capacity report for all the disks that will likely run out of space within the next week.

Figure 2. Over of the predict disk capacity workflow
Figure 2. Predict disk capacity workflow

Define the forecast

The workflow uses the Davis for Workflows action to automatically trigger a forecast for a selected set of disks. The forecast operation is selected within the Davis action, and a DQL query is used to specify the set of disks and the capacity indicator metric that should be predicted. Note that you can use any time series data you can fetch from Grail using DQL within the forecast action.

While this example uses the metric dt.host.disk.free, you can choose any kind of capacity metric, such as host CPU, memory, or network load—you can even extract a metric value from a given log line.

The forecast is trained on a relative timeframe (for example, the last seven days) which is specified in the configured DQL query. The DQL query example below trains forecasting on a relative timeframe of the last seven days:

timeseries avg(dt.host.disk.free), by:{dt.entity.host, dt.entity.disk}, bins: 120, from:now()-7d, to:now()

The configuration below shows that a forecast horizon of 100 data points is requested, which means that 100 additional predicted points will expand the initially fetched 120 data bins of the source DQL query. This predicts one week into the future.

Figure 3. Detail of the forecasting workflow step
Figure 3. Detail of the forecasting workflow step

The prediction action returns all its forecasted time series lines, which can include hundreds or even thousands of individual disk predictions.

Evaluate the forecast results

Within the following TypeScript action, each disk prediction is tested against a threshold to determine if the disk will run out of space in the next week. The TypeScript code snippet below is responsible for checking for threshold violations and for preparing all the violations in a result object for subsequent actions to follow up on:

Figure 4. Evaluating the results with a custom TypeScript action
Figure 4. Evaluating the results with a custom TypeScript action

The TypeScript action returns a custom object that uses a Boolean flag (violation) to tell the follow-up actions about violations and an array of all the violation details (violations).

const predictionSummary = { violation: false, violations: new Array<Record<string, string>>() };

Tip: Download the TypeScript template from our documentation.

Trigger remediation actions

A collection of remediation actions can be used to follow up on predicted capacity shortages. In this example, two parallel actions are defined. One action sends out an email notification; the other raises a Davis problem for each violating disk. All remediation actions use the Boolean violation flag of the previous workflow action to avoid invocations when there are no violations.

Here you can see the invocation condition used in the follow-up actions that control the invocation.

Figure 5. Conditional execution
Figure 5. Conditional execution

Raise events in case of disk capacity shortage!

A TypeScript remediation action is used to iterate through all the predicted disk shortages and to raise individual alarm events. Each alarm event has custom event properties that can be used to deliver further details about the situation and to further identify the disk or host.

Figure 6. Create an alarm event for predicted shortages.
Figure 6. Create an alarm event for predicted shortages.

Tip: Download the TypeScript template from our documentation.

Review all Davis-predicted capacity problems

Navigating to the Davis problems feed, the operations team can review all the predicted disk capacity shortages. Remember, raising events and problems is an optional remediation step that can be skipped entirely by directly sending emails or Slack messages to the responsible teams.

The creation of alerting events within this workflow example highlights the flexibility and power of the Dynatrace AutomationEngine combined with the analytical capabilities of Davis AI and Grail.

Figure 7. List of events created by the workflow.
Figure 7. List of events created by the workflow.

Summary

The combination of Davis AI forecasts with Dynatrace AutomationEngine and Grail opens the door for many valuable use cases—anticipative management of capacity being the most prominent of these. Predicting future capacity shortages for thousands of disks or hosts allows operations teams to anticipate critical situations weeks before incidents occur. The flexibility and power of the Dynatrace AutomationEngine allow operations teams to react to detected shortages flexibly and to customize and implement their remediation flows.

You can install Davis® for Workflows via the Dynatrace Hub. As a starting point for implementing your own anticipative capacity management workflow, you can download all the TypeScript code used in this example from our documentation:

For full details, see Davis AI analysis in workflows documentation.

Predictive capacity management in action (14-second video)

The post Automate predictive capacity management with Davis AI for Workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automate-predictive-capacity-management-with-davis-ai-for-workflows/feed/ 0
Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI https://www.dynatrace.com/news/blog/dynatrace-automatically-monitors-openai-chatgpt-for-companies-that-deliver-reliable-cost-effective-services-powered-by-generative-ai/ https://www.dynatrace.com/news/blog/dynatrace-automatically-monitors-openai-chatgpt-for-companies-that-deliver-reliable-cost-effective-services-powered-by-generative-ai/#respond Wed, 07 Jun 2023 17:07:42 +0000 https://www.dynatrace.com/news/?p=58130 Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI

This blog post looks at how Dynatrace automatically collects OpenAI/GPT model requests and charts them within Dynatrace, as well as how abnormal service behavior can be used to identify slowdowns in OpenAI/GPT requests as the root cause of large-scale issues. Both functionalities have been part of the Dynatrace platform for a couple of years already, and so have withstood the challenges of customer usage.

The post Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI appeared first on Dynatrace news.

]]>
Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI

AI observability is becoming imperative as businesses in all sectors are introducing novel approaches to innovate with generative AI in their domains. Advanced AI applications using OpenAI services don’t just forward user input to OpenAI models; they also require client-side pre- and post-processing. A typical design pattern is the use of a semantic search over a domain-specific knowledge base, like internal documentation, to provide the required context in the prompt. This is achieved by using OpenAI services to compute numerical representations of text data that ease the computation of text similarity, called “embeddings,” for the documents as well as for the user input.

Furthermore, tools like LangChain leverage large language models (LLM) as one of their basic building blocks for creating AI agents (think of AI agents as APIs that perform a series of chat interactions that target a desired outcome) which perform complex and potentially large queries against an LLM like GPT-4. They then connect to third-party services such as online calculators, web search, or flight status information to combine real-time information with the power of an LLM.

One of the crucial success factors for delivering cost-efficient and high-quality AI-agent services following the approach described above is using AI observability to closely observe their cost, latency, and reliability.

Dynatrace enables enterprises to automatically collect, visualize, and alert on OpenAI API request consumption, latency, and stability information in combination with all other services that are used to build AI applications. This includes OpenAI as well as Azure OpenAI services, such as GPT-3, Codex, DALL-E, or ChatGPT.

AI observability example: OpenAI token consumption

Our example dashboard below visualizes OpenAI token consumption. It shows critical SLOs for latency and availability, as well as the most important OpenAI generative AI service metrics, such as response time, error count, and the overall number of requests.

AI observability dashboard showing OpenAI service health and performance
With these latency, reliability, and cost measurements in place, your operations team can now define their own OpenAI dashboards and SLOs.

Dynatrace OneAgent® discovers, observes, and protects access to OpenAI automatically, with no manual configuration, revealing the full context of used technologies, service interaction topology, security-vulnerability analysis, and the observability of all metrics, traces, logs, and business events in real time.

How Dynatrace traces OpenAI model requests

Let’s use a simple NodeJS example service to show how Dynatrace OneAgent automatically traces OpenAI model requests. OpenAI offers an official NodeJS language binding that allows the direct integration of a model request by adding the following lines of code to your own NodeJS AI application:

const { Configuration, OpenAIApi } = require("openai");

const configuration = new Configuration({

apiKey: process.env.OPENAI_API_KEY

});

const openai = new OpenAIApi(configuration);

const response = await openai.createCompletion({

model: "text-davinci-003",

prompt: "Say hello!",

temperature: 0,

max_tokens: 10,

});

Once the AI application is started on a OneAgent-monitored server, the application is automatically detected, and the traces and metrics for all outgoing requests are collected. OneAgent automatic injection of monitoring and tracing code works not only for the NodeJS language binding but also when using the raw HTTPS request in NodeJS. While OpenAI offers official language bindings only for Python and NodeJS, there is a long list of community-provided language bindings.

OneAgent can automatically monitor all C#, .NET, Java, Go, and NodeJS bindings. However, we recommend following the OpenTelemetry approach to monitoring Python with Dynatrace.

The screenshot below shows the traces that OneAgent collects, along with all the latency and reliability measurements for each of the outgoing GPT model requests.

Traces that OneAgent collects, along with all the latency and reliability measurements for each of the outgoing GPT model requests in Dynatrace screenshot

Dynatrace further refines the OpenAI calls by automatically splitting specific services for the OpenAI domain, as shown below.

General Settings for OpenAI calls in Dynatrace screenshot

Once this is done, the Dynatrace Service Flow shows the flow of your requests, starting with your NodeJS service and calling the OpenAI model, as shown below.

AI observability service flow for conversastionService

As shown in the example above, Dynatrace OneAgent automatically collects all latency and reliability-related information along with all the traces showing how your OpenAI requests traverse your service graph.

The seamless tracing of OpenAI model requests allows operators to identify behavioral patterns within their AI service landscape and to understand the typical load situation of their infrastructure.

This AI observability knowledge is essential for further optimizing the performance and cost of services.

By adding some lines of manual instrumentation to a NodeJS service, cost-related measurements are also picked up by OneAgent, collecting the number of OpenAI conversational tokens used.

Observing OpenAI request cost

Each request to an OpenAI model, such as text-davinci-003, gpt-3.5-turbo, or GPT-4 reports back how many tokens were used for the request prompt (the length of your text question) and how many tokens the model generated as a response.

OpenAI customers are billed based on the total number of tokens consumed by all the requests they make. By extracting these token measurements from the returning payload and reporting them through Dynatrace OneAgent, users can observe token consumption across all OpenAI-enhanced services in their monitoring environment.

Here is the instrumentation used to extract the token count from the OpenAI response and to report the three measurements to the local OneAgent:

function report_metric(openai_response) {

var post_data = "openai.promt_token_count,model=" + openai_response.model + " " + openai_response.usage.prompt_tokens + "\n";

post_data += "openai.completion_token_count,model=" + openai_response.model + " " + openai_response.usage.completion_tokens + "\n";

post_data += "openai.total_token_count,model=" + openai_response.model + " " + openai_response.usage.total_tokens + "\n";

console.log(post_data);

var post_options = {

host: 'localhost',

port: '14499',

path: '/metrics/ingest',

method: 'POST',

headers: {

'Content-Type': 'text/plain',

'Content-Length': Buffer.byteLength(post_data)

}

};

var metric_req = http.request(post_options, (resp) => {}).on("error", (err) => { console.log(err); });

metric_req.write(post_data);

metric_req.end();

}

After adding these lines to your NodeJS service, three new OpenAI token consumption metrics are available in Dynatrace, as shown below.

OpenAI token consumption metrics available in Dynatrace screenshot

Davis AI automatically detects ChatGPT as the root-cause

One of the superb features of Dynatrace is Davis® AI, which automatically learns the typical behavior of monitored services. Once an abnormal slowdown or increase of errors is detected, Davis AI triggers root cause analysis to identify the cause.

Our simple example of a NodeJS service entirely depends on the ChatGPT model response. So, whenever the latency of the model response degrades or the model request returns an error, Davis AI automatically detects it.

In the example below, Davis AI automatically reported a slowdown of the NodeJS prompt service and correctly detected the OpenAI generative service as the root cause of the slowdown.

Davis AI automatically reportes a slowdown of the NodeJS prompt service and correctly detected the OpenAI generative service as the root cause of the slowdown

The Davis problem details page shows all affected services for which the OpenAI generative service was the root cause of the slowdown, along with the ripple effects of the slowdown.

The problem details also list all Service Level Objectives that were negatively impacted by the slowdown.

List of all Service Level Objectives that were negatively impacted by a slowdown

AI observability with Dynatrace brings peace of mind when using OpenAI models

The massive popularity of generative AI cloud services, such as OpenAI’s GPT-4 model, is forcing companies to rethink and redesign their existing service landscapes. Integrating generative AI into traditional service landscapes comes with all kinds of uncertainties. Using AI observability from Dynatrace to observe OpenAI cloud services helps you gain cost transparency and ensure the operational health of your AI-enhanced services.

Also, full transparency and observability of AI services will play a significant role in upcoming AI regulations at a national level and for risk assessments within your own company.

For further details, you can view the full source of the NodeJS service on GitHub.

The post Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-automatically-monitors-openai-chatgpt-for-companies-that-deliver-reliable-cost-effective-services-powered-by-generative-ai/feed/ 0