troubleshooting | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Tue, 23 Jun 2026 07:00:25 +0000 en hourly 1 Orchestrate multicloud AI agents for autonomous incident resolution https://www.dynatrace.com/news/blog/orchestrate-multicloud-ai-agents-for-autonomous-incident-resolution/ https://www.dynatrace.com/news/blog/orchestrate-multicloud-ai-agents-for-autonomous-incident-resolution/#respond Mon, 15 Jun 2026 20:11:14 +0000 https://www.dynatrace.com/news/?p=74557 Observability data

Cloud SRE Agents is a Dynatrace app that orchestrates AWS®, Azure®, and Google® AI agents for automated investigation and resolution assistance for incidents across multicloud environments. Cloud SRE Agents routes identified issues based on configurable rules, centralizes its findings, and provides a single audit trail for autonomous operations.

The post Orchestrate multicloud AI agents for autonomous incident resolution appeared first on Dynatrace news.

]]>
Observability data

Organizations are evolving from human-driven operations to supervised autonomous operations, where AI investigates, recommends, and remediates, and humans stay in control of what matters most. A big part of delivering on that vision is working with the agents that customers already run in their cloud environments.

Harness the power of hyperscale agents

Each hyperscaler has AI agents that automatically investigate and help resolve production incidents using native cloud telemetry and tools. They act like embedded site reliability engineers, analyzing issues and recommending or executing remediation steps without waiting for a human to start the process.

AWS DevOps Agent provides investigation and remediation in AWS using native tooling. An Azure SRE Agent specializes in investigating and remediating Azure issues. And Google Gemini Cloud Assist is for incident analysis across Google Cloud Platform (GCP).

Over the past year, we’ve published how Dynatrace supercharges each of these cloud agents individually. When an issue occurs, Dynatrace Intelligence combines causal, predictive, and agentic AI using the Smartscape dependency graph to automatically link related symptoms and root causes across the environment into one unified problem card.

When Dynatrace integrates with the AWS DevOps Agent, dependency-aware root cause analysis combines with AWS frontier-agent capabilities, and joint customers report up to 70% reductions in mean time to resolution. When Azure SRE Agent connects with Dynatrace, deterministic, causation-based AI flows directly into Azure-native remediation workflows, cutting the back-and-forth between teams. And with Google Gemini Cloud Assist, Dynatrace delivers the same production context layer to GCP-hosted incidents: precise root cause, full topology, real business impact.

Problem detected by Dynatrace Intelligence, investigated and remediated by AWS DevOps Agent (see documentation in the right-hand panel)
Figure 1. Problem detected by Dynatrace Intelligence, investigated and remediated by AWS DevOps Agent (see documentation in the right-hand panel)

From integrations to intelligent orchestration

Many enterprises run workloads across AWS, Azure, and Google Cloud simultaneously, and managing three separate integrations with separate routing logic and separate cost controls is its own operational tax. Cloud SRE Agents provides a single orchestration layer that routes problems to specific hyperscaler agents based on configurable profiles to see everything happening across all three cloud agents.

The Cloud SRE Agents app writes findings back to Dynatrace, and provides your team with measurable visibility into autonomous actions.

The Overview tab's interactive graph shows a live view of problems and their activity status, grouped by related SRE agent.
Figure 2. The Overview tab’s interactive graph shows a live view of problems and their activity status, grouped by related SRE agent.

How Cloud SRE Agents works

When Dynatrace Intelligence detects a problem and identifies the root cause, Cloud SRE Agents calls dedicated cloud-native agents from AWS, Azure, and Google Cloud to retrieve deeper insights from the sources that only they can reach: CloudTrail history, Azure subscription policy, GCP project IAM, recent deployments, and native runbooks. These agents run in parallel, gathering evidence as soon as the problem is detected. Their findings, and, where applicable, the recommended remediation path, are displayed in the same Dynatrace problem view that the on-call SRE is already using in their day-to-day workflow.

One view. No tab-switching. The work starts without you.

Three workflows do the orchestration in the background:

  • Investigate evaluates your Interaction Profiles and dispatches matching problems to the right agents in parallel.
  • Periodic Tasks polls each cloud provider for completion, detects stalled or timed-out investigations, and writes findings back as problem annotations.
  • Event Handlers normalize the cloud-provider event stream so every action correlates back to its originating problem, end to end.

Cloud SRE Agents has the insights and intelligence to decide which agent gets which problem, tracks each run to completion, and brings the answers back together in a single view. The Overview tab provides a real-time, interactive network graph of problems, agents, and activities. The replay view allows the user to step back in time and get an overview of what has happened when, as well as the status of each investigation.

Replay functionality in the Cloud SRE Agents Overview
Figure 3. Replay functionality in the Cloud SRE Agents Overview

Intelligent routing with Interaction Profiles

In agentic operations, routing rules make the difference between turning autonomous systems loose on every alert and pointing them precisely where they earn their keep. Interaction Profiles are how you express routing judgment in Cloud SRE Agents. Each profile pairs a set of conditions with the agent or agents that should handle the problems flagged by the profile, and evaluates the conditions whenever Dynatrace Intelligence detects a problem.

The conditions you can write are deliberately broad. You can route by the cloud account, subscription, or project an incident touches; by problem category (availability, error, slowdown, resource contention); by affected entity type (a Kubernetes cluster, a database, a Lambda function); by tag, label, or any custom attribute carried in the problem record. Conditions combine with AND/OR logic and nest as deeply as you need, keeping real production routing policy inside the app rather than spilling into custom workflows or scripts.

Three ways teams put it to work

Route problems to the right cloud, automatically

A spike in Lambda error rates belongs to AWS DevOps Agent. An Azure App Service degradation calls for Azure SRE Agent. A Pub/Sub latency issue lands with Gemini Cloud Assist. In a multicloud estate, none of those decisions should fall to a human at 2:00 AM. A profile filtered by AWS Account ID, Azure Subscription ID, or GCP Project ID, then narrowed by resource type or tag, settles the routing question once. Every matching problem is automatically routed to the right specialist with the right cloud-native context.

Optimize spend with budget-aware routing

Cloud AI agents do work, and that work has a cost. Cloud SRE Agents lets you set a Monthly Duration Budget per agent and gate dispatch on it via a Has Available Budget filter: once the budget is exhausted, new investigations either stop (in strict enforcement mode) or proceed with a logged warning. The duration figure itself is a proxy, derived from Dynatrace event timestamps rather than the cloud provider’s clock, which makes it useful as a circuit breaker and directional signal, not a substitute for AWS, Azure, or GCP usage reports. The governance value is what matters: you decide how much autonomous investigation you’re willing to underwrite each month, and the system holds the line.

Tier autonomous investigation by problem type and entity

Not every Dynatrace problem warrants an autonomous investigation. Problem Category filters let you dispatch agents only to the problem categories that warrant it, for example, availability or error problems that require immediate action, rather than slowdowns or custom alerts where human triage might still be the right call. Layer on Entity Type filters, and you can further focus on specific infrastructure tiers (hosts, services, process groups, Kubernetes clusters). The result is a tiered model: high-severity issues receive immediate autonomous investigation, lower-severity signals queue for human review, and your team controls the threshold.

Governance that makes autonomous work measurable

Agentic operations earn trust when teams can see what the agents did, why, and whether it worked. Cloud SRE Agents treats that as a first-class concern, with two views built for the two audiences who care about it.

The Activity tab is the audit trail. Every investigation and mitigation appears as a card on a unified timeline; expand any card to see the agent’s full findings, the evidence it pulled, and the action it took or recommended. Each response can be rated Good, OK, or Bad, building a quality signal grounded in what your team actually saw rather than what the system predicted. When a single problem triggers work across multiple agents, those activities roll up to a single status (in progress, done, or stalled), so you always know where things stand without having to reconstruct the run from individual records.

Activity tab showing an expanded investigation card with agent findings and rating control.
Figure 4. Activity tab showing an expanded investigation card with agent findings and rating control.

The Statistics tab is where autonomous operations become a number you can show to a leadership team: problems handled, mitigations executed, average investigation time, MTTR and MTTI trends, success rates, and satisfaction scores broken down by agent. The same view doubles as a directional cost lens, since agent working time is the dominant driver on the cloud side of the bill. Treat the number as a trend signal and a circuit-breaker input, not a billing record (reconcile against AWS, Azure, and GCP usage reports for exact spend), and it makes the case for expanding agentic coverage with evidence rather than anecdote.

The Statistics tab shows key metrics and per-agent insights across a selected time range.
Figure 5. The Statistics tab shows key metrics and per-agent insights across a selected time range.

Why production context multiplies the value

What changes Cloud SRE Agents from a smart dispatcher into something more is what Dynatrace Intelligence contributes before an agent ever begins its analysis. Dynatrace delivers deterministic, causation-based root cause analysis grounded in Dynatrace’s Smartscape real-time dependency mapping, alongside business impact assessment and correlated telemetry. That context shapes the entire direction of the investigation. A cloud agent arriving with that foundation starts from “this specific service on this specific host is the root cause, and here’s the customer impact” rather than “something is wrong somewhere in this account.”

The numbers reflect it. According to AWS, organizations using the AWS DevOps Agent with Dynatrace see up to a 75% reduction in mean time to resolution.

Western Governors University, which runs a fully online learning environment for 200,000 students, uses AWS DevOps Agent with Dynatrace to automate cross-system correlation that previously required manual effort across multiple tools. At a larger scale, United Airlines transports more than 500,000 passengers daily across a hybrid environment that includes more than 500 AWS accounts, 20,000 Lambda functions, and 38,000 OneAgent deployments.

The team’s description of the before and after status is direct: previously, multiple tools with overlapping functions created gaps and black boxes during troubleshooting. With AWS DevOps Agent and Dynatrace, Dynatrace identifies the responsible layer, the agent investigates and provides resolution steps, and everything surfaces in a single Dynatrace view. No 3:00 AM tool-switching required.

Get started

For a closer look at the individual integrations, read the posts on AWS DevOps Agent and Dynatrace and Azure SRE Agent and Dynatrace, or see how Dynatrace Intelligence powers autonomous operations. To put your cloud agents to work today, install Cloud SRE Agents from the Dynatrace Hub. Cloud SRE Agents is currently available as a community-supported app.

Harness the power of your hyperscaler agents

The post Orchestrate multicloud AI agents for autonomous incident resolution appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/orchestrate-multicloud-ai-agents-for-autonomous-incident-resolution/feed/ 0
Dynatrace observability is now a Kiro power https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/ https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/#respond Fri, 12 Jun 2026 21:12:32 +0000 https://www.dynatrace.com/news/?p=74536

In this blog, we'll introduce the Kiro power for Dynatrace, show what it unlocks for developers, and walk you through how to get it up and running.

The post Dynatrace observability is now a Kiro power appeared first on Dynatrace news.

]]>

What is the Kiro power for Dynatrace?

The Kiro power for Dynatrace delivers live observability data, root cause analysis, and remediation suggestions directly into the Kiro IDE, with no JSON editing or manual MCP setup.

Kiro is an AI-powered IDE that helps developers move from idea to working code through spec-driven development and an agentic assistant. To make the assistant genuinely useful in unfamiliar domains, Kiro recently introduced powers: curated, partner-validated bundles of MCP servers, steering files, and best practices that install with a single click and load on demand when a relevant task comes up. Install a power, and Kiro’s agent gains specialized expertise the moment you need it.

For Dynatrace customers already working in Kiro, it’s the shortest path yet from code to production insight. For developers new to Dynatrace, it’s a one-click way to ground Kiro’s reasoning in real facts from your environment, not guesses.

Why this matters for developers

Developers have historically been one step removed from production. When something breaks after deployment, the path to figuring out what went wrong usually runs through a Site Reliability Engineering (SRE) or operations team, and AI coding assistants can’t automatically and reliably remediate issues in software they’re unfamiliar with. Agents that can write code are guessing about how their code behaves in production unless they have access to real telemetry data.

The Dynatrace Kiro power for Dynatrace closes this gap through Dynatrace Intelligence, the agentic operations system at the core of the Dynatrace platform. Kiro’s answers are grounded in deterministic, causal AI and real-time production data, not probabilistic guesses.

When a developer starts a task by writing a prompt, Kiro evaluates the conversation, identifies the relevant power using keywords, and dynamically activates power. Kiro then loads Dynatrace MCP tools and power instructions, providing skills to investigate problems, query live observability data, surface root causes, and even execute and verify remediations.
Figure 1. When a developer starts a task by writing a prompt, Kiro evaluates the conversation, identifies the relevant power using keywords, and dynamically activates the power. Kiro then loads Dynatrace MCP tools and the power instructions, providing the skills needed to investigate problems, query live observability data, surface root causes, and even execute and verify remediations.

With the tools provided by the Kiro power, developers can:

  • Investigate live incidents and get root cause analysis directly in Kiro chat
  • Query metrics, logs, and traces from production using natural language
  • Surface security vulnerabilities affecting the code they’re working on
  • Get remediation suggestions grounded in what’s actually happening in their environment

“Using Kiro powers for Dynatrace has been a total game-changer in the observability space. Deep-dive root cause analysis of complex system issues that once required lengthy manual intervention now happens in seconds, giving us unprecedented speed and confidence.”

Mike Kobush, Sr. Software Performance Engineer, NAIC

How to install the Kiro power for Dynatrace

Getting started takes only a few steps. Once installed, the Kiro power activates automatically when Kiro detects a relevant task. Mention an incident, a slow service, or anything that needs production context, and the Dynatrace tools and guidance will load in Kiro chat.

Prerequisites

  • A Dynatrace account. If you don’t already have one, you can start a free 15-day trial.
  • Kiro installed on your system.

Prepare the Dynatrace connection

First, create a Dynatrace Platform Token, which Kiro will use to authenticate. Then add the required permissions for the Dynatrace MCP server.

Install the Kiro power

The power can be installed from either the Kiro IDE or the Kiro powers website. For this walkthrough, we’ll use the IDE.

  1. Launch the Kiro IDE.
  2. Select the Ghosty icon with the lightning bolt to open the powers panel.
  3. Select Dynatrace Observability from the Recommended
  4. Select Install. The power is registered with placeholder values for the Dynatrace URL and token. Therefore, Kiro will show an error message that the MCP server can’t be reached.
  5. To complete the configuration, select Open Settings and replace the placeholders with your environment details.

Configure your tenant and token

In the settings file, replace the two placeholders:

Placeholder Replace with
YOUR_DT_URL https://TENANT_ID.apps.dynatrace.com/platform-reserved/mcp-gateway/v0.1/servers/dynatrace-mcp/mcp. Replace TENANT_ID with your Dynatrace environment ID (visible in your environment URL, for example https://<ENVIRONMENT_ID>.apps.dynatrace.com/ui).
YOUR_BEARER_TOKEN The Dynatrace platform token you created earlier (for example, dt0s16.XXXXX).

Start asking questions

Open a new chat in Kiro and start interacting with your Dynatrace environment using natural language. Query active problems or security vulnerabilities, request a root cause analysis to identify critical issues in production, or pull related logs and traces, all without leaving the IDE.

See it in action

The short demo below walks through installing the Kiro power for Dynatrace, verifying the connection, and running a first query against your environment to list the top 10 vulnerabilities detected by Dynatrace.

Installing and activating the Kiro power for Dynatrace (video)
Figure 2. Installing and activating the Kiro power for Dynatrace (video)

Get started with the Kiro power for Dynatrace

Kiro powers transform what used to be a stitching exercise (MCP servers here, steering files there, custom instructions somewhere else) into one single, ready-to-use bundle. The Kiro power for Dynatrace applies the same idea to observability: live production insight, causal root cause analysis, and remediation grounded in real telemetry, all available the moment a developer needs them.

The result is a tighter loop between writing code and understanding how it behaves in production. Less waiting for diagnostic data from someone else. Less guesswork from an AI assistant operating without context. And, more time spent on the work that actually matters.

Ready to try it? The Kiro Power for Dynatrace is publicly available: install it from kiro.dev or the Kiro IDE and start asking your environment questions.

Using Kiro and the Kiro power for Dynatrace root cause analysis (video)
Figure 3. Using Kiro and the Kiro power for Dynatrace root cause analysis (video)
Experience the Kiro power for Dynatrace for yourself.

The post Dynatrace observability is now a Kiro power appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/feed/ 0
Bring real-time production insights into Claude Code with the Dynatrace MCP Server https://www.dynatrace.com/news/blog/bring-real-time-production-insights-into-claude-code-with-the-dynatrace-mcp-server/ https://www.dynatrace.com/news/blog/bring-real-time-production-insights-into-claude-code-with-the-dynatrace-mcp-server/#respond Mon, 30 Mar 2026 18:36:42 +0000 https://www.dynatrace.com/news/?p=73599 Dynatrace and Claude logos

Get immediate production visibility inside Claude Code, the next-gen AI coding assistant from Anthropic. The Dynatrace MCP server can now be used as a ready-to-use connector for Claude Code, Cowork, and Chat. Connect in minutes to query logs, traces, and problems; conduct live debugging from your terminal, and validate AI-first workflows with real data. Model […]

The post Bring real-time production insights into Claude Code with the Dynatrace MCP Server appeared first on Dynatrace news.

]]>
Dynatrace and Claude logos

Get immediate production visibility inside Claude Code, the next-gen AI coding assistant from Anthropic. The Dynatrace MCP server can now be used as a ready-to-use connector for Claude Code, Cowork, and Chat. Connect in minutes to query logs, traces, and problems; conduct live debugging from your terminal, and validate AI-first workflows with real data.

Model Context Protocol (MCP) is the standard for connecting AI assistants to live tools and data. The Dynatrace MCP server brings your full observability and security context into every Claude session. This means fewer context switches, less time hunting through dashboards, and better answers that are grounded in what’s actually happening in your environment.

How to connect Claude Code with Dynatrace

Getting started is straightforward. Open Claude, go to Connectors, and search for “Dynatrace.” Install the Dynatrace MCP Server connector, follow the setup steps, and you’re connected.

Figure 1. Dynatrace MCP Server connector setup in Claude
Figure 1. Dynatrace MCP Server connector setup in Claude

How Dynatrace MCP Server and Claude give you visibility into your production data

Whether you’re investigating a production issue, reviewing a deployment, or working through a security vulnerability, you can ask Claude in plain language and get answers backed by your actual Dynatrace production data. This is not just documentation summaries or generic guidance, but a real window into production.

Troubleshoot without leaving Claude Code

Let’s say you receive a Jira ticket with details of an error. Instead of switching to dashboards and digging through logs to find out what went wrong, you can now ask Claude. Dynatrace pulls all root-cause information, related logs, metrics, traces, and CPU and memory profiling data from your production environment directly into the session. You can query these details by impact, filter by service, and request remediation suggestions, all without knowing in advance where to look or how to write the DQL query.

Figure 2. Claude gets the problem description and production data via the Dynatrace MCP Server.
Figure 2. Claude gets the problem description and production data via the Dynatrace MCP Server.

This same workflow applies when you’re checking for vulnerabilities in your running workloads, verifying whether a recent deployment introduced regressions, or proactively catching performance issues before they hit production.

You can now also use dtctl, the open source CLI for the Dynatrace platform, in Claude Code alongside the Dynatrace MCP server to manage dashboards, run workflows, and execute DQL from your terminal.

Ready to get started? Install Dynatrace MCP Server in Claude Code today.

The post Bring real-time production insights into Claude Code with the Dynatrace MCP Server appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/bring-real-time-production-insights-into-claude-code-with-the-dynatrace-mcp-server/feed/ 0
Deep troubleshooting and crash intelligence in the new Dynatrace RUM experience https://www.dynatrace.com/news/blog/deep-troubleshooting-and-crash-intelligence-in-the-new-dynatrace-rum-experience/ https://www.dynatrace.com/news/blog/deep-troubleshooting-and-crash-intelligence-in-the-new-dynatrace-rum-experience/#respond Fri, 20 Feb 2026 17:08:10 +0000 https://www.dynatrace.com/news/?p=73116 Troubleshooting RUM Experience

Mobile teams have spent years stitching together crash logs, SDK outputs, and guesswork to understand why their apps freeze, crash, or feel slow. Dynatrace has removed that friction with a mobile troubleshooting experience that unifies performance and stability into streamlined end-to-end workflows. This brings together app insights, view‑level performance, crashes, and application not responding (ANR) errors with full contextual enrichment and adds automatic React Native symbolication through built‑in source‑map management.

The post Deep troubleshooting and crash intelligence in the new Dynatrace RUM experience appeared first on Dynatrace news.

]]>
Troubleshooting RUM Experience

With crash intelligence from the new Dynatrace RUM experience, practitioners get a dedicated mobile health view that instantly highlights regressions and stability issues, while two ready‑made dashboards surface performance and troubleshooting insights with zero setup. Where traditional tools stop at isolated crash reporting, Dynatrace connects every error and crash signal—across native, hybrid, and cross‑platform apps. The result is a single, intelligent experience that helps teams ship faster, fix issues sooner, and deliver trustworthy mobile apps.

Why mobile troubleshooting has fallen behind

In the web world, metrics like Core Web Vitals provide a common language for performance. In mobile, no such standard exists, despite mobile apps being far more sensitive to performance regressions than desktop applications. There’s an even bigger challenge: while browser‑based issues can be fixed instantly with a quick deployment, mobile bugs require a full app‑store release and rely on users updating their devices; this makes mobile bugs far more difficult and slower to reverse. It also makes proactive observability even more critical for mobile than for web. Furthermore, today’s mobile apps span hundreds of device types and unpredictable network conditions, and frameworks like React Native or Flutter must function reliably across a range of hardware and operating‑system versions. These layers make traditional “crash-only” analytics insufficient.

While a slow app start can cost you a new user before they’ve even reached your home page, a poorly performing view can stall your conversion funnel. And a frozen view can quietly destroy user trust without leaving any clear trace in your logs. Other monitoring tools often treat these problems in isolation: a crash report here, a view-analytics SDK there, an app-start metric captured in another tool… But teams need unified context, not fragmented data on glass.

A dedicated mobile performance experience

The new Dynatrace RUM experience solves this problem by bringing together performance and stability in a coherent picture rooted in real user journeys, not isolated events. This combined visibility is what sets Dynatrace apart. The foundation of the new RUM experience lies in a new performance analysis capability that reflects how mobile apps actually work. Apps are now represented through mobile “views”—the mobile app pages that users interact with as they move through the RUM experience. Dynatrace automatically detects views for native iOS and Android apps, with manual instrumentation available for frameworks like Flutter and React Native. Performance signals for each view tell a story, including, for developers, where the errors or exceptions occurred.

Error Explorer Instrumentation in Dynatrace screenshot

These metrics are stored as Dynatrace Grail®-native data with all relevant dimensions, enabling rich analysis across devices, OS versions, app versions, and frontends. These metrics are fed directly into Experience Vitals, where App Owners and PMs can drill down from a high-level performance view into detailed metrics for each view. Select any specific view to reveal its performance characteristics, error footprint, and context across the user journey.

App health intelligence on the basis of crashes, ANRs, and error context

Crashes have always been a pain point for mobile developers, but ANRs are often worse. Many tools treat ANRs as an afterthought, offering incomplete or inconsistent visibility. Error Inspector, the dedicated app for frontend troubleshooting on the Dynatrace platform, now captures ANRs on both iOS and Android, as well as crash reporting for Android NDK. Error Inspector reports errors as structured user events with precisely defined semantics. Previously, Dynatrace practitioners could add properties to automatically monitored telemetry data, turning previously opaque frontend degradations into immediately diagnosable issues. Similarly, in the new RUM experience, user events triggered by mobile apps are automatically enriched with view context, device details, and session information to identify exactly which users, contexts, and conditions triggered a crash, thereby dramatically accelerating root‑cause identification and troubleshooting.

Error Explorer in Dynatrace screenshot

Server-side processing applies efficient grouping logic to provide a clear, deduplicated view of what’s happening, enabling meaningful dashboarding and reliable future alerting. With Dynatrace Experience Vitals, SREs gain insights into app performance and stability. Insights are elevated to the same first-class status as performance, ensuring teams see both sides of the experience.

Crashes and ANRs flow naturally to Error Inspector, where they’re prioritized by severity—crashes have higher severity than errors, and all ANRs are highlighted. For developers, this means less noise and faster diagnosis. For business leaders, it means clarity on what matters most, making prioritization a breeze.

Introducing React Native symbolication

Cross-platform frameworks create powerful opportunities but introduce complexity when things go wrong. Stack traces often combine JavaScript and native code, leaving teams with unhelpful, obfuscated output like hashes. To solve this, Dynatrace now supports full React Native source map uploading and management, with automatic symbolication of React Native crashes directly in Error Inspector.

Error Explorer in Dynatrace screenshot

Teams can upload symbol files through standard methods and manage them via a new interface. Unlike competing tools, which only partially support hybrid stack symbolication/deobfuscation or require manual post-processing, symbolicated crashes appear automatically in Dynatrace, making debugging dramatically more efficient.

For developers working with React Native, this eliminates the friction of switching tools and patching together build artifacts.

Error Explorer in Dynatrace

A new mobile health overview

Mobile developers often collaborate with their SRE counterparts on observability teams, who frequently face a simple yet critical question from App Owners and PMs: “Is my mobile app healthy right now?”

The answer to this question is provided by Experience Vitals. The mobile health overview shows whether RUM is enabled on your monitored devices, highlights performance indicators such as app start thresholds, and surfaces stability concerns such as crash and ANR rates. Crashes and errors link directly to Error Inspector, creating a smooth transition from error awareness to analysis.

Error Explorer in Dynatrace screenshot

This focus on clarity allows decision makers to grasp app health at a glance while giving developers frictionless access to the underlying data.

Good insights require the correct setup

To help teams get started faster and avoid misconfiguration, Dynatrace introduces improved instrumentation guides for iOS and Android. These guides add necessary context to ensure new customers receive the exact data they expect. The mobile settings experience has been redesigned to make it clear how UI selections correspond to code snippet changes, aligning it more closely with what teams experience in agentless web flows. This consistency matters, especially for practitioners managing multiple platforms on the enterprise level.

To accelerate adoption and make insights immediately usable, Dynatrace is introducing two ready-made dashboards:

  • App start health dashboard, which includes app start analysis and high-level performance trends.
  • Mobile health dashboard, designed for investigating crashes and errors, providing stability insights without requiring configuration.

App Health dashboard in Dynatrace screenshot

These dashboards incorporate the new metrics, dimensions, and logic introduced across the release, giving teams instant, actionable visibility into their mobile apps. Decision-makers get the reports they need, developers get the views that help them move faster. And everyone benefits from data that’s consistent, contextual, and unified.

App Health dashboard in Dynatrace screenshot

Raising the bar for mobile troubleshooting

Most mobile monitoring solutions emphasize either performance or stability, but rarely both in a unified, contextual way. Many tools require heavy manual instrumentation, fragmented SDKs, or separate troubleshooting workflows for native and cross-platform apps. Crash grouping is inconsistent, ANR support varies widely, symbolication often requires external tools, and virtually no other mobile monitoring solution combines all this with a platform that also ties into backend traces and business impact.

Mobile performance and stability in Dynatrace

Dynatrace delivers something fundamentally and wonderfully different: a single, Grail-powered, full stack view of mobile performance and stability that blends performance metrics, view insights, app-start intelligence, ANR/crash analytics, and hybrid framework support—all in one place, all contextualized by real user journeys and backend traces. This is not just crash reporting. It’s not just performance monitoring. It’s a complete solution for understanding the mobile experience, from the first view to the last tap.

With these enhancements, Dynatrace sets a new standard for mobile observability. Mobile developers gain a more intuitive debugging experience. Product Managers get clear KPIs. App owners receive actionable health insights. And organizations as a whole gain the ability to deliver reliable, performant, and engaging mobile experiences. In a digital world increasingly shaped by mobile interactions, Dynatrace provides the visibility teams need—and the intelligence mobile troubleshooting has been waiting for.

Get started

To instrument your mobile apps and start collecting the full breadth of performance and stability data, please visit our mobile instrumentation documentation. And, if you want to explore how Dynatrace helps you debug crashes, ANRs, and React Native issues with full context, open Error Inspector to see these new capabilities in action.

The post Deep troubleshooting and crash intelligence in the new Dynatrace RUM experience appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/deep-troubleshooting-and-crash-intelligence-in-the-new-dynatrace-rum-experience/feed/ 0
Unprecedented insights into frontend user experience with Dynatrace Real User Monitoring https://www.dynatrace.com/news/blog/unprecedented-insights-into-frontend-user-experience-with-dynatrace-real-user-monitoring/ https://www.dynatrace.com/news/blog/unprecedented-insights-into-frontend-user-experience-with-dynatrace-real-user-monitoring/#respond Wed, 28 Jan 2026 16:55:06 +0000 https://www.dynatrace.com/news/?p=72345 User Experience Monitoring

The next-generation Dynatrace® Real User Monitoring (RUM) experience, powered by Grail® and DQL, delivers unified, actionable insights across modern web and mobile applications. By extending cloud native observability to the frontend, Dynatrace unifies user experience data with backend intelligence, enabling teams to optimize performance, shorten troubleshooting cycles, and improve business outcomes, while preserving strict enterprise-grade privacy standards.

The post Unprecedented insights into frontend user experience with Dynatrace Real User Monitoring appeared first on Dynatrace news.

]]>
User Experience Monitoring

Why RUM matters more than ever

As the frontend becomes the primary point of customer engagement, real user monitoring is a critical link between IT performance and business outcomes. The ways users interact with web and mobile apps have evolved, driven by single-page applications (SPAs), asynchronous rendering, and nonlinear navigation. These changes have increased the demand for RUM solutions that can keep pace with modern architectures and deliver actionable insights.

Teams face growing challenges:

  • Limited visibility into user journeys across SPAs and mobile applications
  • Fragmented analysis between frontend telemetry and backend execution
  • Slow troubleshooting when context is missing or spread across tools
  • Increased pressure to capture meaningful data while maintaining strict privacy standards

Modern frameworks and libraries introduce new patterns and constraints, making it harder to understand user behavior and optimize experiences using legacy tools. Teams require solutions that not only keep pace with evolving architectures but also empower them to act quickly and confidently on user experience data.

Reinventing RUM for modern application experiences

The new Dynatrace RUM experience was purpose-built to address these challenges and futureproof RUM use cases. By leveraging unified, end-to-end observability with deep business context, Dynatrace introduces a streamlined, modern approach to understanding frontend performance and user behavior. Advanced analytics and intuitive, purpose-built workflows are layered on a foundation that aligns with how today’s teams operate.

With this new RUM experience, teams can:

  • Understand modern application behavior with new frontend signals: Gain deeper visibility into how users interact with single-page and mobile applications, leveraging new data models and behavioral signals.
  • Improve user experience with out-of-the-box insights and intuitive practitioner workflows: Access immediate, actionable insights through purpose-built apps and dashboards designed for common practitioner tasks.
  • Extend insights with advanced analytics and built-in privacy controls: Perform deeper analysis, maintain responsible long-term data retention, and ensure privacy and compliance with flexible data access and built-in controls.

Together, these capabilities empower teams to optimize performance, resolve complaints, analyze user behavior, and accelerate DevOps workflows—collaborating more effectively and acting faster on real user impact.

DevOps workflows dashboard

Understand modern application behavior with new frontend signals

The ability to capture and analyze new frontend signals that reflect how modern applications behave in production gives teams a more accurate view of how users move through single-page and mobile applications—without additional instrumentation.

Industry standards such as Core Web Vitals, along with mobile-specific performance signals coupled with enhancements for troubleshooting: application not responding (ANR) and symbolication, provide a consistent way to assess application health and troubleshoot issues across platforms. The new RUM data model captures events that reflect how modern applications behave in production, including:

  • Background requests: Network calls initiated by an application without direct user interaction, typically to fetch or update data in the background to maintain app functionality and user experience.
  • User interactions: Events such as clicks, taps, scrolls, or keypresses that reflect how users engage with an application and provide behavioral insights without necessarily invoking a backend request.
  • Soft navigation: As modern web applications often don’t perform a full-page load when navigating, soft navigation helps analyze the user’s path through the application from one view to the next.

Improve user experience faster with out-of-the-box insights and intuitive practitioner workflows

The new RUM experience delivers immediate value through redesigned, task-oriented experiences built around common practitioner workflows. Out-of-the-box apps and dashboards make RUM data easier to understand and act on from the start.

Pre-built dashboards surface industry-standard performance metrics such as Core Web Vitals and key mobile signals, giving teams a clear and consistent view of frontend health. Unified dashboards bring frontend and backend data into a single view, making it easier to trace request errors to their underlying exceptions, monitor customer experience over time, and understand how performance issues impact business outcomes.

Purpose-built apps include:

Experience Vitals: Quickly identify which requests or assets are contributing to slowdowns, using redesigned waterfall analysis and industry standards like Core Web Vitals. DevOps teams can define custom alerts and integrate them into existing workflows.

Experience Vitals Dynatrace

Error Inspector: Group and prioritize errors, presenting critical context in one place. Session details are shown alongside stack traces, making it easier to assess impact and resolve problems quickly. Integration with Jira accelerates issue tracking and accountability.

Error Inspector in Dynatrace

Users & Sessions: Ground investigations in real user session data. Support teams can quickly find and analyze individual sessions, reducing guesswork and time to resolution. Product owners can identify friction points, validate reported issues, and improve user journeys based on actual behavior.

User session metrics in Dynatrace

These capabilities help teams:

  • Optimize frontend availability and performance
  • Troubleshoot issues faster with clear, end-to-end context
  • Resolve customer complaints using real session data

“UWM is deeply committed to site reliability engineering to deliver the best possible client experience. With the new RUM experience, we gain unprecedented, actionable insights into user interactions, empowering our SRE teams to optimize performance.”

Gavin Wallis, Observability Architect, United Wholesale Mortgage

Extend insights with advanced analytics and built-in privacy controls

While out-of-the-box experiences address the most common workflows, some scenarios require a deeper dive. The new RUM experience provides flexible access to all frontend data through a unified data model, enabling teams to explore real user behavior with greater context and precision.

  • Custom dashboards and collaborative Notebooks: Correlate frontend behavior with backend performance, analyze experience trends over time, and connect user experience directly to business outcomes such as conversion or revenue impact.
  • Dynatrace Query Language (DQL): Explore and connect data, with Dynatrace Intelligence available to help with query generation and data interpretation.
  • Privacy and compliance: Built-in privacy and permission controls support responsible data use. Extended data retention (public preview) allows deep investigations, compliance support, and historical trend analysis.

Getting started with the new RUM experience

The new RUM experience is now generally available, with all apps and features available by default on DPS SaaS tenants with version 330 or higher. You can allow data ingestion to the Dynatrace 3rd-generation platform for individual frontends or whole environments. (Classic data capture isn’t impacted.)

Real User Experience in Dynatrace

Dynatrace RUM setup guides

  • For web application setup, follow these steps. If you use automatic injection, the new configuration will be applied within 5 minutes. If you insert the RUM JavaScript manually, you might need to update the snippet, depending on the snippet format you’re using.
  • For mobile app setup, follow these steps to set up the new RUM experience.
  • For web frontends at the environment level, follow these steps.

The post Unprecedented insights into frontend user experience with Dynatrace Real User Monitoring appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/unprecedented-insights-into-frontend-user-experience-with-dynatrace-real-user-monitoring/feed/ 0
Integration with AWS DevOps Agent: Autonomous investigations powered by production context https://www.dynatrace.com/news/blog/integration-with-aws-devops-agent-autonomous-investigations-powered-by-production-context/ https://www.dynatrace.com/news/blog/integration-with-aws-devops-agent-autonomous-investigations-powered-by-production-context/#respond Thu, 15 Jan 2026 16:39:34 +0000 https://www.dynatrace.com/news/?p=72432 AWS icon and agentic AI

The integration of Dynatrace with AWS DevOps Agent delivers a powerful combination for autonomous incident response, pairing Dynatrace’s AI-powered root cause analysis and real-time production context with AWS’s new frontier agent capabilities. Together, the two platforms bring complementary strengths that accelerate investigations, reduce handoffs and “war room ping-pong,” and ultimately cut time and cost. Teams running AWS applications can investigate incidents more quickly, identify root causes with precision, and move closer to achieving truly autonomous cloud operations.

The post Integration with AWS DevOps Agent: Autonomous investigations powered by production context appeared first on Dynatrace news.

]]>
AWS icon and agentic AI

March 31, 2026 update

Today we congratulate AWS on the general availability of AWS DevOps Agent. This marks an important step forward in how teams operate and innovate in the cloud, moving closer to systems that can investigate and respond with minimal human intervention.

At Dynatrace, we are proud to have collaborated with AWS on this initiative from the beginning. Together, we have worked to bring observability, AI, and automation closer together to help customers simplify operations and resolve incidents faster.
Our joint customers are already seeing measurable value, including up to a 70 percent reduction in mean time to resolution, as teams move from reactive troubleshooting to more intelligent and automated workflows.

This work reflects a broader shift toward agentic operations. We are continuing to deepen our collaboration with AWS across DevOps Agent and other AI services as this space evolves. This foundation sets the stage for how the Dynatrace platform and AWS DevOps Agent integration works in practice.

How AWS DevOps Agent and Dynatrace complement each other to resolve incidents faster

It’s late at night, you’re on call, and an alert fires for an AWS application. You need to assess the severity, understand the impact, and quickly notify the relevant teams. Until now, that potentially meant toggling between Dynatrace and the AWS Console to piece together the full picture. With the AWS DevOps Agent and Dynatrace integration, you instantly have all the information you need at every stage of remediation.

AWS DevOps Agent represents a new class of frontier agents: AI that works autonomously for hours or days, investigating incidents without constant human intervention. Dynatrace provides causal and predictive AI that pinpoints the root cause of issues and anticipates problems before they escalate. Together, they create something neither can deliver alone: end-to-end incident resolution that spans from early warning through root cause to remediation.

When AWS announced the DevOps Agent at re:Invent last December, they showcased this integration as a key use case, demonstrating how autonomous investigation becomes dramatically more effective when powered by Dynatrace precise, topology-aware production context. The agent doesn’t just correlate signals; it understands what those signals mean for your business.

Experience topology-aware root cause analysis with guided mitigation

Consider a typical CRM stack: a React frontend on S3 and CloudFront, an ALB routing to Lambda-hosted Python services implementing the CRM business logic, backed by an RDS PostgreSQL. During normal operations, the responses take ~1 ms, but suddenly those degrade to 1 s+. Dynatrace instantly detects the problem with all relevant context, including business impact, and automatically triggers the AWS DevOps Agent to initiate further investigation.

Dynatrace and the AWS DevOps Agent work hand in hand to analyze and mitigate the problem
Figure 1: Dynatrace and the AWS DevOps Agent work hand in hand to analyze and mitigate the problem.

  • Root cause analysis with causal AI: Dynatrace detects response-time degradation and automatically gathers relevant context, including potentially AWS resources causing the issue, such as Lambda, RDS, and ALB.
  • Seamless collaboration with AWS DevOps Agent: Dynatrace triggers the AWS DevOps Agent and passes full runtime context, allowing it to trace the execution path from symptom to failing component.
  • Pinpoint the root cause: The AWS DevOps Agent analyzes underlying RDS logs, identifies DROP INDEX commands that correlate with slowdown events, and surfaces the findings directly in Dynatrace, without tool switching. The commands are traced to an input error by a database administrator.
  • Recommend and stage a fix: the agent provides a clear diagnosis, recommended remediation steps, and proposed action that’s ready for human approval.
  • Prevent recurrence: The agent suggests proactive monitoring of database logs for similar commands to prevent future incidents.

This always-on, on-call workflow accelerates triage, allows topology-aware root cause analysis, guides mitigation, and adds preventative recommendations. It works across a broad set of AWS services, including AWS Lambda function errors, Amazon EKS container failures, Amazon VPC connectivity issues, and more.

Ready to try it out yourself?

With the Dynatrace integration into AWS DevOps agents, you get:

  • Fewer handoffs and clearer ownership with a single investigation narrative (no bouncing between teams/tools)
  • Less manual correlation as Dynatrace supplies topology, dependencies, and traces as a ready-to-use production context
  • Faster “why” analysis as AWS DevOps Agent correlates AWS telemetry with change/deployment history and proposes mitigations
  • A more repeatable incident response, including prevention recommendations to reduce repeats

For more details on the preview and how to try it yourself, have a look at this hands-on walkthrough on the AWS Cloud Operations Blog. You can also refer to AWS documentation for instructions on connecting Dynatrace and the AWS DevOps Agent.

For more news on Dynatrace and AWS, have a look at this recent blog post.

The post Integration with AWS DevOps Agent: Autonomous investigations powered by production context appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/integration-with-aws-devops-agent-autonomous-investigations-powered-by-production-context/feed/ 0
Real-time insights: Leverage Dynatrace observability capabilities within Kiro powered by AWS https://www.dynatrace.com/news/blog/real-time-insights-leverage-dynatrace-observability-capabilities-within-amazon-kiro/ https://www.dynatrace.com/news/blog/real-time-insights-leverage-dynatrace-observability-capabilities-within-amazon-kiro/#respond Mon, 24 Nov 2025 19:42:17 +0000 https://www.dynatrace.com/news/?p=72036 Amazon Q Developer CLI and Dynatrace

In today’s cloud-native environments, having real-time observability data at your fingertips is crucial. By integrating Kiro powered by AWS with Dynatrace, you can leverage powerful AI-assisted monitoring and troubleshooting capabilities directly in your development workflow. Kiro—which recently reached general availability—helps developers by bringing structure to AI coding with spec-driven development. When a developer needs to fix […]

The post Real-time insights: Leverage Dynatrace observability capabilities within Kiro powered by AWS appeared first on Dynatrace news.

]]>
Amazon Q Developer CLI and Dynatrace

In today’s cloud-native environments, having real-time observability data at your fingertips is crucial. By integrating Kiro powered by AWS with Dynatrace, you can leverage powerful AI-assisted monitoring and troubleshooting capabilities directly in your development workflow.

Kiro—which recently reached general availability—helps developers by bringing structure to AI coding with spec-driven development. When a developer needs to fix an issue, investigate an error, or optimize resource usage, it’s crucial they can analyze what happened just before the issue occurred and delve deeper into the infrastructure utilization of your applications in your cloud or container environment.

Unlock development productivity with live production insights

Developers typically face restricted access to production environments, being fully dependent on site reliability engineers (SREs) or operations teams to detect and report issues post-deployment, and provide them with the necessary information to fix an issue. This segmented workflow can result in delayed problem identification and resolution, an increased risk of failures in production, and reduced efficiency throughout the development lifecycle.

By connecting Dynatrace with Kiro, developers can access real-time insights from production environments, gain contextual information down to the root cause of an incident, and receive remediation proposals—all within their Kiro environment.

Figure 1: Dynatrace Agentic AI ecosystem for developers
Figure 1. Dynatrace Agentic AI ecosystem for developers

Kiro has a built-in Model Context Protocol (MCP) client that can be used to extend its capabilities to communicate securely and flexibly with external data sources and tools such as Dynatrace.

Let’s dig deeper into how to leverage this capability and provide Dynatrace’s unique insights to your development teams.

Step-by-step integration guide

Prerequisites

  • You’ll need a Dynatrace account. If you don’t already have one, you can start a free 15-day trial.
  • Kiro must be installed on your system.
  • You must have basic familiarity with AWS services and the Dynatrace platform.

Prepare integration with Dynatrace

First, you need to create a Dynatrace Platform Token, which is used to define Kiro access, and then add the required permissions for the Dynatrace MCP server.

Configure Kiro MCP Settings

The Kiro MCP configuration is managed through a JSON file. The interface supports two levels of configuration:

  • User-level: ~/.kiro/settings/mcp.json applies to all workspaces
  • Workspace-level: .kiro/settings/mcp.json is specific to the current workspace

You can apply the configuration using two different methods:

Method 1: Open the command palette (use Cmd + Shift + P on Mac or Ctrl + Shift + P on Windows/Linux), search for MCP and select Kiro: Open workspace MCP config (JSON) or Kiro: Open user MCP config (JSON), depending on whether or not you want to configure the settings for the workspace or user level.

Method 2: Alternatively, you can use the Kiro Panel. Open Kiro and select the Kiro ghost icon to open the left-side panel. Locate the MCP SERVERS section, select  Open MCP Config, and then start configuring the connection for the Dynatrace MCP Server.

Dynatrace specific settings

Note: Only add one of the following configurations, depending on whether you want to use the remote MCP server or the local MCP server. You can’t use both at the same time.

Using the remote MCP server

Use the following configuration. Replace $TENANT_ID with your Dynatrace environment ID. (You can find your environment ID in the URL of your Dynatrace environment — for example, https://<ENVIRONMENT_id>.apps.dynatrace.com/ui.) Then, replace $DT_PLATFORM_TOKEN with the ID of the Dynatrace platform token you created previously (for example, dt0s16.XXXXX).

{ 
"mcpServers": 
  { 
    "dynatrace": {
      "type": "http",
      "url": "https://$TENANT_ID.apps.dynatrace.com/platform-reserved/mcp-gateway/v0.1/servers/dynatrace-mcp/mcp",
      "headers": {
        "Authorization": "Bearer $DT_PLATFORM_TOKEN"
      },
      "tools": ["*"]
      }
  }
}

Connect the local MCP server

The configuration for the local Dynatrace MCP server can be added to the Kiro IDE using one-click installation or by following the manual configuration as shown below. Don’t forget to replace $TENANT_ID with the ID of your tenant.

{
  "mcpServers": {
    "dynatrace-mcp-server": {
      "command": "npx",
      "args": ["-y", "@dynatrace-oss/dynatrace-mcp-server@latest"],
      "env": {
        "DT_ENVIRONMENT": "https://$TENANT_ID.apps.dynatrace.com"
      }
    }
  }
}


Figure 2. Add Dynatrace via one-click installation (video)
Figure 2. Add Dynatrace via one-click installation (video)

Verify the integration

Once configured, you can use the Kiro chat to interact with Dynatrace through natural language conversations. Simply tell Kiro what you need, whether it’s investigating a critical incident, gaining insights into metrics, logs, or traces from your application, analyzing dependencies, or setting up automated alerts.

In the screenshot below, you can see in the lower left which capabilities are provided by the Dynatrace MCP Server. Beyond the standardized actions, such as listing active vulnerabilities or problems, querying data stored in Dynatrace, or creating a workflow, you can also interact with Davis CoPilot®, the Dynatrace natural language assistant.

Figure 3: Amazon Kiro with an established connection to Dynatrace.
Figure 3. Kiro with an established connection to Dynatrace.

Conclusion

This integration isn’t just another feature; it’s a fundamental shift in how Dynatrace integrates with your development workflow. It brings together the power of Kiro’s AI capabilities with the Dynatrace unified observability platform, allowing developers to access critical monitoring data and gain a real-time understanding of their production environments via natural language interaction.

Spend less time context switching and more time creating value for your customers. Start today and benefit from real-time insights, precise root cause analysis based on causal understanding or improved troubleshooting capabilities, and enhanced development workflows.

Explore how Dynatrace can integrate seamlessly into your development landscape using our remote MCP Server. If you’re interested in learning more about Kiro, have a look at their launch blog post or visit the documentation and dig deeper into how to connect with MCP Servers.

Gain efficiency by empowering Kiro with insights from Dynatrace.

The post Real-time insights: Leverage Dynatrace observability capabilities within Kiro powered by AWS appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/real-time-insights-leverage-dynatrace-observability-capabilities-within-amazon-kiro/feed/ 0
Boost cloud reliability: Dynatrace and Azure SRE Agent unite for autonomous operations https://www.dynatrace.com/news/blog/boost-cloud-reliability-dynatrace-and-azure-sre-agent-unite-for-autonomous-operations/ https://www.dynatrace.com/news/blog/boost-cloud-reliability-dynatrace-and-azure-sre-agent-unite-for-autonomous-operations/#respond Wed, 19 Nov 2025 17:19:49 +0000 https://www.dynatrace.com/news/?p=71938 Dynatrace and Azure SRE Agent

The integration of Dynatrace with Microsoft Azure SRE Agent establishes a new benchmark for cloud operations by leveraging AI-based root cause analysis and real-time production insights, alongside a comprehensive understanding of complex, large-scale IT environments. You can leverage the combined strengths of Dynatrace and Microsoft, enabling teams to resolve complex problems in large-scale IT environments […]

The post Boost cloud reliability: Dynatrace and Azure SRE Agent unite for autonomous operations appeared first on Dynatrace news.

]]>
Dynatrace and Azure SRE Agent

The integration of Dynatrace with Microsoft Azure SRE Agent establishes a new benchmark for cloud operations by leveraging AI-based root cause analysis and real-time production insights, alongside a comprehensive understanding of complex, large-scale IT environments. You can leverage the combined strengths of Dynatrace and Microsoft, enabling teams to resolve complex problems in large-scale IT environments more quickly and efficiently, and automate incident remediation, moving one step closer to driving autonomous operations across their complex environments.

In today’s cloud-first world, reliability isn’t just a goal; it’s a competitive advantage. As more services move online and LLM-powered assistants evolve into autonomous agents, maintaining the reliability, scalability, and cost-efficiency of critical systems becomes essential.

That’s why Dynatrace and Microsoft teamed up to integrate Dynatrace® AI-powered observability with the Azure SRE Agent. This collaboration allows site reliability engineers (SREs) to ensure seamless operations while proactively planning for future scalability and reliability requirements.

Transform your incident management through the combined capabilities of Azure SRE Agent and Dynatrace AI

Azure SRE Agent, introduced earlier this year, provides SREs and developers with the tools they need to increase the speed and efficiency of incident responses, diagnostics, and collaboration, allowing them to resolve problems quickly.

Automate monitoring of cloud environments
Figure 1. Automate monitoring of cloud environments

Seamlessly integrated with incident management tools such as ServiceNow, as well as the developer ecosystem, represented by GitHub Copilot or Azure DevOps, the agent runs in the background 24/7, learning and monitoring the health and performance of your cloud environment.

As a reliability assistant, Azure SRE Agent supports teams by efficiently diagnosing and resolving production issues. You can ask the agent questions in natural language, easily access clear and concise problem summaries, and coordinate incident workflows with integrated human-in-the-loop approvals.

Dynatrace enhances Azure SRE Agent’s troubleshooting and automation capabilities with advanced observability insights. By mapping topology, data, and business context, Dynatrace gains a comprehensive understanding and delivers production-accurate visibility across your entire IT system. This visibility feeds Dynatrace deterministic AI, allowing precise root-cause identification and impact analysis. All these insights are now seamlessly supplied to the Azure SRE Agent, equipping it with real-time production context and reliable root cause analysis.

This allows your teams to move beyond simply receiving alerts; teams are now provided with AI that acts, guides safe mitigations, and accelerates resolution within Azure-native workflows.

Gain efficiency across every stage of the incident lifecycle

Using the Model Context Protocol (MCP), the Azure SRE Agent is securely connected with Dynatrace. Whether a team member uses the agent to ask questions in plain natural language, or the agent interacts with Dynatrace directly—sharing insights, asking for real-time observability data, or root cause analysis, together with remediation steps—the close collaboration supports use cases across every stage of incident management, allowing you to:

  • Cut MTTR by automating routine runbooks and diagnostics, with safe, approved mitigation actions based on full context.
  • Reduce security risk by triaging vulnerabilities faster with production evidence, triggering guided fixes, and validating outcomes.
  • Accelerate delivery with contextual GitHub issues and PRs that include root cause, blast radius, and tests, minimizing issue reproduction time and rework.
  • Improve fix accuracy by correlating Azure and Dynatrace telemetry for precise root-cause and impact analysis.
  • Prevent incidents before they happen using real-time signals and historical trends to stop regressions and reduce toil.

Illustrating the value: Proactively detect and remediate security vulnerabilities

Let’s take a look at a concrete example, which we presented at Microsoft Ignite. Imagine you run a Java-based payroll app on Azure, and a new security warning (CVE) appears. Every second matters now, and there’s no room for error: you need the issue fixed quickly, without lots of back-and-forth between teams.

Schematic illustration – proactive vulnerability remediation with Dynatrace, Azure SRE Agent and GitHub
Figure 2: Schematic illustration – proactive vulnerability remediation with Dynatrace, Azure SRE Agent, and GitHub
  • Once the vulnerability is detected, Dynatrace automatically identifies the library that caused the vulnerability, opens a GitHub issue containing all relevant information, such as which parts of your app are affected, and informs Azure SRE Agent.
  • The SRE agent reviews the GitHub issue and requests additional information from Dynatrace via the MCP server, such as the number of users affected, how often it happens, which endpoints are involved, or which customers might be affected, to assess the scope and impact of the vulnerability.
Azure SRE automatically creates a GitHub issue with all the details.
Figure 3. Azure SRE automatically creates a GitHub issue with all the details.
  • After gathering all necessary details, the SRE Agent synthesizes the information and creates a new GitHub issue, assigning it to GitHub Copilot for remediation.
  • GitHub Copilot then takes action by updating the configuration and code in the GitHub repository to resolve the vulnerability automatically.
  • The pull request not only includes the necessary version changes but also includes documentation, highlighting all findings as well as how the issue was remediated, along with unit tests, to prevent the issue from recurring.

Demo of Azure SRE Agent thumbnail

Try the power of Agentic AI for incident resolution

Dynatrace delivers deep, causation-based insights into your live systems, now seamlessly integrated with Azure SRE Agent to elevate your incident management. With this integration, you can unlock:

  • Smarter detection and remediation: Deep contextual observability from Dynatrace, correlated with Azure telemetry, enhances issue identification and resolution across complex environments.
  • Automated operations: Routine runbook actions and diagnostic workflows can be automated, reducing mean time to repair and freeing teams to focus on innovation.
  • Proactive reliability: Continuous analysis of real-time and historical data identifies leading indicators of failure, allowing teams to prevent incidents before they impact customers.

Azure customers can now access Azure SRE Agent directly in the Azure portal. To connect Dynatrace with the agent and learn how to set up Dynatrace MCP Server, see Dynatrace Documentation.

The post Boost cloud reliability: Dynatrace and Azure SRE Agent unite for autonomous operations appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/boost-cloud-reliability-dynatrace-and-azure-sre-agent-unite-for-autonomous-operations/feed/ 0
From code to cloud: Dynatrace launches first GitHub custom agent, consolidating observability for developers https://www.dynatrace.com/news/blog/from-code-to-cloud-dynatrace-launches-first-github-custom-agent-consolidating-observability-for-developers/ https://www.dynatrace.com/news/blog/from-code-to-cloud-dynatrace-launches-first-github-custom-agent-consolidating-observability-for-developers/#respond Tue, 18 Nov 2025 17:43:25 +0000 https://www.dynatrace.com/news/?p=71917 Dynatrace observability and Security Agent

Much has been said about the distributed and ephemeral nature of today’s software development landscape. In recent years, the rise of AI and agentic ecosystems has further transformed the way developers work. Yet, despite these advancements, developers still struggle with scattered tech stacks, fragmented tools, and, until now, fragmented insights.

The post From code to cloud: Dynatrace launches first GitHub custom agent, consolidating observability for developers appeared first on Dynatrace news.

]]>
Dynatrace observability and Security Agent

With Dynatrace, developers can rest assured that everything they need—from development to deployed infrastructure—is available with full context at every touchpoint. Traditionally, developers have had limited visibility into production systems, relying on operations teams to surface issues after deployment. This disjointed approach often leads to delays in identifying and resolving problems, increased risk of production failures, and inefficiencies in the development process

With GitHub’s recent announcement of custom agents, Dynatrace is proud to launch its first custom agent, introducing a powerful new entry point for Dynatrace insights directly within GitHub repositories, where developers can act on them immediately. This integration ensures that developers no longer need to leave their familiar GitHub environment to access critical observability and security data. Whether it’s triaging production errors, validating deployments, or responding to security vulnerabilities

This isn’t just another way to consume observability and security insights from Dynatrace—this is about achieving true end-to-end visibility, from code to production, by combining the capabilities of Dynatrace with GitHub Copilot and its custom agents.

In this blog post, we’ll explore the most prominent use cases offered by this integration and demonstrate how deploying these two tools in tandem can unlock new levels of efficiency, reliability, and security for your software development lifecycle.

How it works

Dynatrace and Git solutions support opposite ends of the Software Delivery Lifecycle: Git hosts your code repository while Dynatrace monitors your production infrastructure. Agentic AI and the MCP protocol allow bridging of this gap, essentially allowing AI assistants to verify their code in production.

Dynatrace custom agent concept
Figure 1. Dynatrace custom agent concept

Setting up the connection with Dynatrace is simple

We recommend using your local MCP server for prototyping and testing things out. First, follow these instructions to create a platform token with the necessary scopes. Once this is done, install the MCP server in your GitHub repository: you’ll find the configuration in Settings > CoPilot > Coding Agent (an example can be found in the repository linked above).

Finally, you need to set up the custom agent in your repository under .github/agents/Dynatrace.md. For your convenience, we’ve provided an example file that you can simply drag and drop. Update your MCP server’s credentials, and you’re all set.

Tip: Before promotion to production, we recommend switching from your local MCP Server to our remote MCP Server.

Use cases: Unlock the power of production insights

Dynatrace’s custom agent for GitHub Copilot unlocks the following core use cases, each designed to enhance developer productivity and ensure operational excellence:

Incident Response and Root Cause Analysis

Developers can investigate and resolve production failures directly within GitHub by leveraging Dynatrace Davis® AI for real-time problem detection, backend stack trace analysis, and business impact assessment. This provides higher precision in less time, without requiring changes to workflows at other locations.

How to find the Dynatrace agent
Figure 2. How to find the Dynatrace agent

Deployment impact analysis

The agent validates deployments by comparing pre- and post-deployment metrics, detecting anomalies, and providing data-driven health assessments to ensure smooth rollouts.

Production error triage

Dynatrace systematically monitors and categorizes production errors. Now, with a custom agent, we can prioritize remediation efforts, reduce error backlogs, and improve application reliability directly in GitHub Copilot.

A 404 error is investigated and resolved
Figure 3. A 404 error is investigated and resolved

Performance regression detection

Dynatrace continuously monitors application performance baselines, ensuring that any degradation in latency, throughput, or error rates is flagged before it impacts your end-users or breaches service-level objectives (SLOs).

Release validation and health checks

Acting as an automated quality gate, the agent can validate releases, monitor post-deployment stabilization, and ensure compliance with SLOs.

Security vulnerability response and compliance monitoring

Dynatrace integrates security workflows into GitHub, proactively identifying vulnerabilities, mapping them to compliance frameworks, and providing prioritized remediation recommendations

To infinity and beyond

Dynatrace’s custom agent is more than just a tool for many use cases—it’s a powerful toolkit that helps you unlock the full potential of your existing data to uncover the insights you need, precisely when you need them. The use cases we’ve highlighted here are just the beginning. The true power of this integration lies in the dynamic nature of AI, enabling you to search, query, and connect insights across a vast network of agents and platforms. While the journey may take you through uncharted territories of interconnected ecosystems, and it may sound a bit scary at first, there is huge value (and fun!) in turning complexity into clarity and productivity.

Ready to consolidate observability and get answers faster?

Have you already seen the recently launched GitHub custom agents, and are you interested in trying out Dynatrace’s custom agents? Then familiarize yourself with our remote MCP Server and learn how to integrate it into your development landscape.

Experience how real-time production context makes your organization more efficient.

The post From code to cloud: Dynatrace launches first GitHub custom agent, consolidating observability for developers appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/from-code-to-cloud-dynatrace-launches-first-github-custom-agent-consolidating-observability-for-developers/feed/ 0
Hands-free vulnerability remediation with Dynatrace and GitHub Copilot https://www.dynatrace.com/news/blog/dynatrace-mcp-server-and-github-copilot-coding-agent/ https://www.dynatrace.com/news/blog/dynatrace-mcp-server-and-github-copilot-coding-agent/#respond Tue, 28 Oct 2025 10:32:53 +0000 https://www.dynatrace.com/news/?p=71564 Hand Free remediation - GitHub Universe

In today’s fast-paced development environments, security vulnerabilities can be a major bottleneck, slowing down releases and increasing risks. Traditional approaches to addressing vulnerabilities often involve manual triaging and prioritization, high development efforts, and downtime, which can hinder developer productivity, delay critical fixes, and risk production apps.

The post Hands-free vulnerability remediation with Dynatrace and GitHub Copilot appeared first on Dynatrace news.

]]>
Hand Free remediation - GitHub Universe

The value of automation and agentic AI

The Dynatrace® AI-powered observability and security platform, with its remote Model Context Protocol (MCP) server, revolutionizes this process by integrating observability context into automation and agentic AI-driven workflows. By connecting Dynatrace with GitHub Copilot coding agent, organizations can achieve prioritized and automated vulnerability remediation that not only streamlines security but also maintains system performance and developer efficiency.

Both developers and site reliability engineers (SREs) benefit from this agentic AI collaboration, bringing actionable runtime insights from Dynatrace directly into GitHub and efficiently automating vulnerability remediation.

Automated security remediation with smart runtime verification use case

GitHub Dependabot proactively alerts developers when known vulnerabilities are detected in their projects’ dependencies, helping teams stay secure and compliant. As organizations and projects scale, prioritizing and addressing the alerts becomes a challenge: deciding which vulnerabilities matter most in a given context and streamlining the path to remediation.

Dependabot alerts page
Figure 1. Dependabot alerts page

Let’s dig deeper into two different scenarios, where Dynatrace can help to improve remediation by automating developer tasks and prioritizing work based on impact.

To optimize remediation efforts and minimize release delays, it’s essential to prioritize vulnerabilities based on their actual impact on production applications.

Dynatrace providing context for automating the remediation of GitHub Dependabot alerts

In a typical environment, when streamlining vulnerability remediation, you might utilize GitHub Actions workflows to regularly poll Dependabot. Once a new alert is detected, the workflow creates a new issue and assigns it to the GitHub Copilot coding agent.

To fully understand the problem and its impact, GitHub Copilot coding agent queries Dynatrace—using the remote MCP server—to get additional runtime data. Dynatrace confirms that the vulnerable library is loaded, and that its own Runtime Vulnerability Analytics (RVA) identifies the same issue and verifies the impact by highlighting if the vulnerable function is used in live environments and if it’s exploitable.

Typical high-level architecture for vulnerability remediation workflow.
Figure 2. Typical high-level architecture for vulnerability remediation workflow.

Equipped with the additional context, the coding agent generates a code-level fix. To remediate and prevent insecure code, the fix is added as a pull request to the GitHub repository, awaiting developer review to ensure human oversight.

Once the change is approved and the fix deployed, Dynatrace continuously monitors the environment, verifying the remediation and ensuring the issue is resolved without introducing new problems.

Automate and orchestrate security findings from GitHub Dependabot with Dynatrace Workflows

Unifying and contextualizing vulnerability findings across different tools helps apply prioritization and centralize automation for real-time runtime validation and fix deployment, supporting SREs to minimize potential disruptions to their services.

With the Dynatrace integration for GitHub Advanced Security, you can continuously forward GitHub Dependabot alerts to Dynatrace. Once ingested, Dynatrace uses its capabilities for further analysis, including an RVA verification that provides a contextual understanding of the potential impact on the monitored environment.

Enhanced vulnerability remediation architecture with Dependabot alerts.
Figure 3. Enhanced vulnerability remediation architecture with Dependabot alerts.

The workflow creates a new GitHub issue, including a comprehensive summary of alerts with their confirmation status.

GitHub Copilot Coding Agent automatically picks up the issue and submits the proposed fix for verified alerts as a pull request for a developer to review. This ensures the remediation process remains transparent and allows for human oversight before deployment. Once the pull request is approved and merged, the alert is remediated and the code is secured.

Automated end-to-end security remediation

These two scenarios exemplify how organizations can shift from reactive to proactive security management, automating repetitive tasks and providing actionable insights for developers.

The relationship between Dynatrace and GitHub demonstrates the power of agentic AI in modern software development, where runtime data drives decision-making and empowers coding agents to apply fixes to code environments—all in a standardized way, applying enterprise guardrails.

Reduce your mean time to resolution (MTTR), enhance developer productivity, and ensure robust system security by automating security remediation—all without sacrificing performance or uptime.

The benefits of the coding agent are further amplified with the GitHub announcement of introducing custom agents. Dynatrace has just launched its first custom agent, which seamlessly integrates Dynatrace’s observability and security capabilities into GitHub Copilot, empowering teams to maintain operational excellence, ensure application reliability, and uphold security compliance across the entire software development lifecycle (SDLC).

This allows teams to streamline incident response, perform root cause analysis, validate deployments, triage production errors, and many more use cases—all directly within their GitHub repository workflows.

By delivering real-time insights and actionable data from production environments, Dynatrace extends GitHub’s reach into production, effectively closing the SDLC. Stay tuned for our upcoming blog, where we’ll dive deeper into the powerful capabilities of Dynatrace’s custom agent for GitHub Copilot.

Ready to transform your security workflows?

Explore how Dynatrace can integrate seamlessly into your development landscape using our remote MCP Server and read through our walkthrough documentation.

Sign up for the preview and experience how real-time production context makes your organization more efficient.

The post Hands-free vulnerability remediation with Dynatrace and GitHub Copilot appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-mcp-server-and-github-copilot-coding-agent/feed/ 0
Sky-high developer productivity with Dynatrace MCP and GitHub Copilot https://www.dynatrace.com/news/blog/sky-high-developer-productivity-with-dynatrace-mcp-and-github-copilot/ https://www.dynatrace.com/news/blog/sky-high-developer-productivity-with-dynatrace-mcp-and-github-copilot/#respond Fri, 03 Oct 2025 17:12:09 +0000 https://www.dynatrace.com/news/?p=71251 Agentic AI

Unless you’ve been living under a rock this past year, you’ve heard the buzz about how AI is changing the day-to-day lives of developers, whether it’s using LLMs for vibe coding or adopting agentic AI concepts to improve productivity. In particular, the introduction of Model Context Protocol (MCP) as a standard for connecting AI agents […]

The post Sky-high developer productivity with Dynatrace MCP and GitHub Copilot appeared first on Dynatrace news.

]]>
Agentic AI

Unless you’ve been living under a rock this past year, you’ve heard the buzz about how AI is changing the day-to-day lives of developers, whether it’s using LLMs for vibe coding or adopting agentic AI concepts to improve productivity. In particular, the introduction of Model Context Protocol (MCP) as a standard for connecting AI agents with other agents or tools has led to exponential growth in the number of solutions connecting AI coding assistants with APIs, databases, and more. Last month, GitHub, publisher of the prominent AI coding assistant GitHub Copilot, announced its new MCP registry as a place where developers can find links to MCP servers. This registry solves the challenge of numerous registries, repos, or community threads listing the same MCP servers.

We’re excited to share that the GitHub MCP registry now includes Dynatrace MCP, so developers can integrate Dynatrace observability and security analysis directly into their workflows.

In this blog post, we’ll explore how developers can use Dynatrace MCP together with GitHub Copilot to streamline troubleshooting, enhance security, and boost productivity—without ever leaving their IDEs.

Need to troubleshoot an issue? Dynatrace MCP has the answers

One of the biggest challenges in troubleshooting and observability is knowing where to look for missing data when issues arise in production. It might sound odd, but when you consider the sheer number of tools, applications, environments, and layers of source code developers need to navigate, it’s not at all clear where they should look for quick answers. No wonder onboarding time is such a major factor in engineering productivity.

Your day as a developer might start with a complaint about something not working correctly. You receive a Jira ticket with a cryptic error message that includes a link to a service health dashboard or a notification in the problem list. You start rummaging through the logs and dashboards to find answers outside your IDE, with little context. This greatly increases your cognitive load and delays the problem’s resolution.

Developers who use Dynatrace, though, don’t have to manually dig through logs and dashboards. They get clear summaries, root cause information, and all relevant data in the context of the affected service. Thanks to the updated problem flow, it’s easy for them to identify the root cause and remediate it.

Dynatrace identifies the likely root cause, performing failure analysis in the context of the affected service.
Figure 1. Dynatrace identifies the likely root cause, performing failure analysis in the context of the affected service.

Context is key

Now imagine a scenario in which your DevOps team uses MCP as their standard for integrating external tools and insights. As a developer, your instinct is to jump straight into the source code, so why not share as much context as possible right there? Utilizing GitHub Copilot with the Dynatrace MCP integration in place, you get all relevant data from your production environment and quickly isolate the problem, giving you the ability to:

  • Get information on the root cause
  • Query related logs, metrics, and traces
  • Get real-time data from all environments, including production
  • Leverage additional gathered insights, such as CPU and memory profiling, from the affected service
  • Get further context using query patterns, such as “group logs by customer impact”

And the best is, you don’t need to know how or where to look—Dynatrace handles all this for you automatically. And there’s more: you can even ask Dynatrace for remediation suggestions.

Dynatrace identifies the problem root cause as an arithmetic exception, all via LLM and the MCP protocol in the developer VS Code IDE. (video)
Figure 2. Dynatrace identifies the problem root cause as an arithmetic exception, all via LLM and the MCP protocol in the developer VS Code IDE. (video)

What’s happening behind the scenes

When you type a natural language prompt into the GitHub Copilot chat, GitHub’s MCP client establishes a connection with the Dynatrace MCP server, which then connects to Dynatrace. The LLM converts the prompt into a context-aware call and ultimately transforms it into a DQL query that it executes on Dynatrace Grail® data lakehouse, which retrieves the necessary data.

Simplified communication flow.
Figure 3. Simplified communication flow.

Need to fix a security issue? Dynatrace MCP has the answers

By integrating security testing practices earlier in the Software Development Lifecycle, the Shift-left principle has brought security responsibilities into the developer’s world, and they’re here to stay. Developers often find themselves reacting to alerts from SREs, manually auditing code, or relying on static analysis tools that frequently miss context-specific vulnerabilities.

You can instantly access insights like:

  • Vulnerabilities in your code, including open source and third party vulnerabilities
  • Recommendations for fixing an issue
  • Proactive checks tailored to your current coding context

All of this happens without pulling developers away from their flow. MCP allows for a smarter, more proactive approach to security—one that’s embedded in their development process.

Dynatrace shares details of known vulnerabilities related to the affected component/service.
Figure 4. Dynatrace shares details of known vulnerabilities related to the affected component/service.
Dynatrace assesses the query pattern and suggests a fix.
Figure 5. Dynatrace assesses the query pattern and suggests a fix.

Need to create new code? Dynatrace is here to help

Writing new code with Dynatrace allows developers to look ahead proactively. By typing natural-language, conversational questions like “Have I bumped into CPU limits?” or “What is my CPU usage? Is it too high?” developers can identify potential bottlenecks and performance issues before they become real problems, ultimately delivering higher quality code.

Developing new code or optimizing a service in an existing app brings even more complexity. Such work involves identifying inefficient API usage, reducing unnecessary load, improving performance, and assessing the potential impact of the new code so as to minimize deployment issues.

Need to verify recent CICD builds? Dynatrace helps you shift left

As a developer, you want to catch build issues early—before they snowball into deployment delays or production incidents. With Dynatrace MCP, you can type questions like:

  • “What failed in the last build?”
  • “Are there any performance regressions tied to this commit?”
  • “Did this deployment introduce any anomalies?”

By integrating Dynatrace into your CI/CD pipeline, you gain real-time visibility into build health, test coverage, and deployment impact. This means faster feedback loops, fewer surprises, and higher delivery quality. Dynatrace helps you to shift from reactive debugging to proactive delivery assurance—all through natural language interactions.

Get started with Dynatrace MCP

The Dynatrace MCP server is available as a community-supported open source project. To familiarize yourself with all it can do, visit the Dynatrace MCP project. in our GitHub repository. There you’ll find all the necessary documentation to guide you through the setup process and explain all the available capabilities.

Explore how developers use Dynatrace MCP with GitHub Copilot to streamline troubleshooting, enhance security, and boost productivity without ever leaving their IDEs.

The post Sky-high developer productivity with Dynatrace MCP and GitHub Copilot appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/sky-high-developer-productivity-with-dynatrace-mcp-and-github-copilot/feed/ 0
Let the problem guide you: Smart remediation with Dynatrace https://www.dynatrace.com/news/blog/let-the-problem-guide-you-smart-remediation-with-dynatrace/ https://www.dynatrace.com/news/blog/let-the-problem-guide-you-smart-remediation-with-dynatrace/#respond Wed, 13 Aug 2025 15:38:50 +0000 https://www.dynatrace.com/news/?p=70376 Problem alert dashboard

Imagine a friend pings you, asking to meet them at a new restaurant whose location you don’t know. Would you simply step outside and start walking, hoping that you’ll somehow find the restaurant by luck? More likely, you’ll open a navigation app on your mobile device, search for the address, and set your route. A critical incident in your IT stack is no different. You shouldn’t have to guess your way through the dark alleys of logs, traces, and warnings, asking passersby for assistance, just to be bounced from one guess to the next.
This is where Dynatrace comes in: curating what’s most relevant, guiding engineers to the next best steps, and confidently guiding them from alert to resolution.

The post Let the problem guide you: Smart remediation with Dynatrace appeared first on Dynatrace news.

]]>
Problem alert dashboard

At Dynatrace, we’ve long led the way in root cause analysis, automatically detecting when something’s wrong and identifying why. These AI-powered insights form the intelligent backbone of how Dynatrace guides engineers through complex, entangled problems, turning a confusing cityscape of data into a focused remediation journey.

In this blog post, we follow Omar, a fictitious SRE working on Astroshop, the Dynatrace OpenTelemetry demo application. Omar works with Sophie, a developer, to resolve a critical failure.

Dynatrace guides them, and any team, through every step of the incident lifecycle, with speed, precision, and confidence. Here’s how:

  • Problem diagnosisDynatrace automatically packages the failures into a single correlated problem.
  • AlertingOmar receives a rich, actionable notification, right where he works, with root cause analysis (RCA), ownership, and deep links already in place.
  • TriageThe problem details page serves as a “triage and remediation command center,” showing Omar what’s impacted, who’s affected, and which issues to prioritize for remediation.
  • HandoffFrom within the problem context, a guided Jira ticket populated with RCA insights, logs, and deep links can be created with one click.
  • Guided investigationsProblems view lists all incidents related to the context of the problem.
  • Remediation – Using Live Debugger and real-time metrics, Sophie fixes the issue and sees immediate confirmation that the fix worked.
  • Reinforcement and learning for the future – The fix is automatically documented in a troubleshooting guide, turning today’s incident into tomorrow’s intelligence.

Problem diagnosis: multiple incidents and a stream of signals tied into one cohesive problem

The Astroshop online store recently began experiencing errors when customers attempted to pay using American Express credit cards. Error messages prevent users from completing their purchases. Dynatrace seasonal baselining detected this service health issue, reduced alert noise by clustering all related events into a single problem, and tied all the relevant telemetry (logs, metrics, and traces) to the affected users, infrastructure, and SLOs.

By determining the ownership of the service, based on the Kubernetes label, the problem is automatically routed to the right team: Omar’s SRE team.

Figure 1. The Problems app clusters related events into one cohesive problem.
Figure 1. The Problems app clusters related events into one cohesive problem.

Contextual alerting

Meanwhile, Omar, the on-call SRE, receives the notification directly in the context where Omar and his team work, Slack. The notification reads, Failure rate increase in payment service.

Figure 2. Receive alert notifications in the tools where you work.
Figure 2. Receive alert notifications in the tools where you work.

Notifications like these can, of course, also show up in JIRA, ServiceNow, or other tools, depending on where your team works and wants to be notified. The notification provides a link to the full context, where Omar can begin remediation straight away. The notification identifies the root cause and service owners. It also provides directly accessible opinionated drilldowns, removing the need to open different tools. Omar starts his work by investigating the context and responding to the incident.

Triage: understand what, who, and how bad

Selecting the View problem button in the Slack notification brings Omar to the related Problem page, his triage and remediation command center.

At a glance, Omar can answer the following questions: How bad is this problem? Who and what is affected? How long has the problem existed? And, what is my best course of action?

Figure 3. Problem details: triage key impact and root cause at a glance.
Figure 3. Problem details: triage key impact and root cause at a glance.

First question: How bad is it, and who’s affected?

From the Problem Incident header, Omar immediately understands that the issue is blocking revenue: 600 users failed to check out their purchase. Furthermore, he sees that there is one frontend and six services are affected, and the error puts five SLOs at risk. Clicking one of the tiles allows the engineer to drill down into further details, offering clear “what to look at next” calls to action.

While unfamiliar with the error, Omar understands that this issue is severe and must be prioritized. Before taking action, he wants to understand the current situation compared to the normal state, which leads to his next question.

Second question: How long has this been going on, and how many checkouts are failing?

The event chart makes it clear: the incident has been ongoing for thirty minutes. Omar sees exactly how the event has evolved over time. The baseline-aware metrics allow him to distinguish between normal fluctuations and true anomalies. Here, the current behavior deviates from normal, with a significant increase in errors, evidently deviating from the baseline, where no errors are expected over that same period, around the same time of the day.

Meaning, for the last thirty minutes, many more people were unable to complete their purchases than usual. The issue is ongoing. In the blink of an eye, he has validated the severity of the incident.

Third question: What’s broken?

The deployment view on the left provides a visual breakdown of the failure by component and cluster. The root cause engine correlated all contributing events and factors, and pinpointed the exact service at fault: the culprit, a Kubernetes workload associated with the “payment service.”

Fourth question: Who can fix it?

This is now easy. The root cause, clearly labeled within the affected infrastructure, clarifies ownership. Omar sees that Sophie, the developer from Team Finance, owns the failing service.

Swiftly guided through the incident, Omar can confidently hand over the issue to the right person.

Handoff: effortless and precise handoff to the responsible team

Omar kicks off a handoff directly from the problem details page within Dynatrace. With a single click, Omar creates a Jira Issue containing all relevant information.

The Jira ticket is populated with the problem identifiers, making it easy to understand the affected services and components at stake without toggling between tools. No ping-pong escalations, just clear, confident delegation with minimal overhead.

Sophie, the owning developer, receives the Jira notification and opens the ticket to find everything she needs: a brief summary with direct links to the problem view, failure analysis, and Live Debugger.

Guided investigation in problem mode

While guided through the next best course of action, Sophie chooses to analyze the failure further using the Services app. This opens problem mode, a visual assistive layer inside the Services app, scoped to the incident’s impact window. It pre-filters all logs, traces, and warnings based on:

  • The scope of the incident
  • The services and infrastructure involved
  • Event severity and user impact
Figure 4. Problem mode in the Services app: Investigate only what’s relevant in the automatically scoped context.
Figure 4. Problem mode in the Services app: Investigate only what’s relevant in the automatically scoped context.

This view automatically prioritizes what matters, highlighting the signals that explain why the failure occurred, not just what broke. Instead of wrestling with contextless, stale, and unrelated logs, Sophie sees error logs and telemetry tied to the exact timeframe of the checkout errors. It cuts to the relevant error logs linked to the failure, including the exact error message. Sophie sees immediately that this is a bug in the code, specifically a JavaScript error at line 73.

Figure 5. Drill down to the root cause on the code level.
Figure 5. Drill down to the root cause on the code level.

Remediation: fix rapidly, validate instantly

Sophie has enough context to fix the failure. She opens Dynatrace Live Debugger, navigates to where the error logs occurred, and sets a non-breaking breakpoint at line 22, where she knows the error originates. Once Live Debugger captures a snapshot, Sophie is able to quickly identify the culprit: a mismatched character, American-Express vs American_Express, is causing the checkout service to bounce.

Figure 6. Start debugging at the exact line where the error occurs with live production data in Live Debugger.
Figure 6. Start debugging at the exact line where the error occurs with live production data in Live Debugger.

Sophie patches the issue and deploys the fix. A couple of minutes later, when having a look at the business metrics dashboard, she validates the remediation action, observing in real-time that checkout failures are gone, and revenue is flowing again.

Document the issue for future reinforcement and learning

Once the fix is deployed, Sophie and Omar go to Troubleshooting in the problem details. Together, they jot down details of the errors and their cause: the key mismatch between American-Express and American_Express, and how it was resolved. Sophie adds the link to the dashboard she used for validation, next to the incident details, automatically prefilled by Davis® AI.

Figure 7. Davis AI automatically recommends relevant documents with actionable troubleshooting steps for this problem.
Figure 7. Davis AI automatically recommends relevant documents with actionable troubleshooting steps for this problem.

Once this is done, Dynatrace automatically

  • Indexes the incident’s metadata and telemetry
  • Prefills it with Incident details linked to the RCA
  • Links it to similar incidents using graph-based AI and vector search

Weeks or months later, when a similar checkout spike appears, Dynatrace remediation intelligence automatically surfaces the relevant troubleshooting guide, providing necessary context and remediation guidance.

Rather than starting from scratch in the future, users benefit from a living knowledge base that not only accelerates problem resolution by reducing repeated investigation but turns every fix into team intelligence.

Try out the Dynatrace problem-remediation journey today

Dynatrace doesn’t just help you detect problems; it guides your teams calmly and confidently from >notification to resolution. Start now, eliminating repetitive war rooms, and try it out yourself:

Explore the problem remediation journey in the Dynatrace Playground.

The post Let the problem guide you: Smart remediation with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/let-the-problem-guide-you-smart-remediation-with-dynatrace/feed/ 0
Remediation intelligence: Accelerate MTTR with AI-powered context and knowledge https://www.dynatrace.com/news/blog/remediation-intelligence-accelerate-mttr-with-ai-powered-context-and-knowledge/ https://www.dynatrace.com/news/blog/remediation-intelligence-accelerate-mttr-with-ai-powered-context-and-knowledge/#respond Wed, 13 Aug 2025 15:38:30 +0000 https://www.dynatrace.com/news/?p=70391 Remediation intelligence

There’s a hidden barrier to deep understanding of your revenue, performance, and bottom line that has nothing to do with tools or telemetry. It’s organizational knowledge, the remediation know-how that’s scattered across hundreds of documents, private notebooks, dashboards, and in the minds of your engineers. This implicit, loosely documented knowledge is invisible to machines and rarely timely for humans. So, when a high-priority incident hits, this knowledge gap fills your war rooms with engineers tasked with solving problems they don’t own.

The post Remediation intelligence: Accelerate MTTR with AI-powered context and knowledge appeared first on Dynatrace news.

]]>
Remediation intelligence

Delivering reliable, business-critical applications to production is more complex than ever. The growing complexity and granularity of modern software systems and the trend to shifting more and more responsibilities to development teams (shift left) lead to increased pressure on your development teams.

In fact, traditional development teams now have a wider set of responsibilities; they must be highly skilled and educated in a multitude of domains. The responsibilities of these teams now range from specification, planning, testing, risk assessment, cost estimates, and deployment, to load testing, UI testing, integration testing, and on-call responsibilities for their software services.

While development teams need to be literate in all new technology stacks, cloud resources, and quality assurance methods, they also need to work with numerous tools.

The invisible bottleneck in your remediation process

The whole shift-left trend has pushed operational responsibility closer to development, expanding workloads to include on-call rotations, more frequent deployments, and growing expectations around uptime. When incidents occur, dozens of engineers are dragged into war rooms to perform analysis of the underlying root causes.

Reducing the Mean Time to Repair (MTTR) is essential to business continuity and customer satisfaction. Without access to the right knowledge at the right time, even skilled teams lose momentum. In high-pressure situations caused by critical incidents, it becomes more important than ever to ensure that all relevant information is shared with every role involved. This requires the most automated and intelligent methods available for effectively collecting, analyzing, and distributing information.

The on-call engineer’s journey

Take Omar, an SRE; an automated voice jolts him awake to summon him into a war room in the middle of the night after a routine update caused a spike in failed requests for a cloud-based payment service. It’s a P1 incident. With each passing minute, merchants are losing value, support requests surge, and customers are complaining. Dozens of caffeinated engineers are already in the war room.

Logs point to timeouts, but this is just a symptom; the real problem is somewhere else. Reading every message and document would take hours, so Omar scans for summaries, key findings, and any mention of his team’s services. Several hypotheses have already been tested, and one points to a potential issue in a backend service Omar’s team owns. Meanwhile, customer complaints are beginning to surface from other time zones. The payment service is failing, and customer success managers are growing increasingly anxious.

And while a similar outage has happened before, Omar is not able to find any documentation or insights into how to remediate the issue.

Why organizational knowledge doesn’t scale

When remediation history lives in documents, scattered across teams, formats, and platforms, engineers waste time searching instead of solving. Even well-documented incidents don’t prevent recurrence if they’re disconnected from future incidents. Even centralized platforms like Backstage don’t help if they can’t surface the right guidance at the right time.

Without a way to systematically identify, reuse, and scale this knowledge, it remains reactive. That’s not just inefficient, it’s a blocker to building intelligent automation and truly preventative operations. This is the hidden obstacle, silently inflating your MTTR, buried knowledge that costs time, delays response, and drains focus from what really matters. This is what remediation intelligence solves.

Introducing remediation intelligence

Dynatrace has a long history of providing DevOps teams with AI-driven tools for anomaly detection, root cause identification, and incident impact assessment in complex application environments. Over the past decade, it has contributed to reducing mean time to resolution (MTTR) by learning application behavior and analyzing dependencies in real time.

Building on this foundation, Dynatrace launched remediation intelligence, which adds an additional element to the incident response process from alert through resolution. It assists engineers during remediation by combining Davis® AI root cause and impact analysis with input from global community knowledge and internal expertise. It integrates data such as logs, metrics, traces, and topological context into a single view and offers support for documenting post-incident reviews.

Figure 1. The problems page displays all important information, allowing you to directly access all incident-relevant error logs.
Figure 1. The problems page displays all important information, allowing you to directly access all incident-relevant error logs.

Close the knowledge gap:  Embedded troubleshooting knowledge

What truly sets Dynatrace remediation intelligence apart is its ability to proactively surface relevant internal knowledge at the moment it’s needed most. It adds an AI-guided assistive layer to the Problems app that brings implicit, organizational knowledge directly into the flow of incident response. Once a problem is detected, Davis AI scans the historical data, surfacing past remediation playbooks, troubleshooting dashboards, and notebooks that were used to resolve similar issues.

With troubleshooting guides, we introduce a context-aware guidance system, built on Davis AI, that connects current incidents with prior resolution paths. It makes organizational knowledge queryable, remediation patterns reusable, and every responder effective, even when they’re solving an unfamiliar issue.

Figure 2. Review related documents from similar past incidents.
Figure 2. Review related documents from similar past incidents.

Remediation intelligence surfaces the most relevant remediation insights

When an incident occurs, Dynatrace excels at automatically analyzing and surfacing technical insights. It collects and organizes all relevant signals—logs, metrics, traces, and topology—into a single coherent problem. No fragmented alerts. No disconnected symptoms. Just one structured, AI-curated incident view. In parallel, Dynatrace AI scans all documents marked as troubleshooting-relevant. Using advanced semantic search and vector embeddings, it ranks and surfaces the most relevant past incidents, dashboards, notes, and postmortems, based on their similarity to the current problem. This is not just keyword matching; Dynatrace understands patterns, failure modes, and system relationships, surfacing ranked, high-similarity incidents in the problem view.

Figure 3. Example of a troubleshooting guide
Figure 3. Example of a troubleshooting guide

From observability to trusted automation: The power of context-aware AI

The future of resilient, self-healing systems lies in the seamless integration of observability, AI, and organizational knowledge. When these elements come together, they form the foundation for trusted automation—a system that not only reacts to incidents but learns from them, adapts, and eventually prevents them altogether.

At the heart of this vision is context. Effective auto-remediation depends on the ability to precisely identify the root cause of an issue and understand its broader impact across the application stack. But automation doesn’t stop at detection. By capturing and integrating the remediation strategies used by engineers, Dynatrace builds a living knowledge base. This organizational knowledge, when combined with AI-driven root cause analysis, allows the system to replicate proven remediation paths and suggest next steps with increasing accuracy.

With all relevant data and insights unified in a single platform, engineers gain a single pane of glass view into their systems. This not only streamlines manual remediation efforts but also lays the groundwork for flexible, context-aware auto-remediation. Each incident you resolve fuels the knowledge. Over time, as the system learns, it evolves from reactive automation to proactive incident prevention—anticipating issues before they escalate and taking preemptive action.

This is the vision Dynatrace is delivering: a future where engineers can trust automation not just to respond, but to understand, learn, and improve—turning every incident into a step toward greater system intelligence and reliability.

Empower your teams: turn hard-won operational insights into scalable remediation power

Dynatrace remediation intelligence is ready to work for you today. To start benefiting, opt into Davis CoPilot®, the Dynatrace generative AI assistant. Once turned on, you’ll need to configure Davis CoPilot to learn from a curated set of Dynatrace documents—specifically, Notebooks and Dashboards that are either created directly from detected problems or clearly labeled with the prefix [TSG] in their titles (short for Troubleshooting Guide).

Davis CoPilot will analyze and learn from your team’s historical remediation efforts, capturing valuable insights and strategies, allowing it to proactively suggest relevant documentation and guidance when similar incidents are detected in the future, helping your engineers respond faster and more effectively.

Figure 4. Turn on Davis CoPilot in Settings.
Figure 4. Turn on Davis CoPilot in Settings.

All data uploaded to Dynatrace remains strictly private. All data and remediation insights remain within your tenant. Davis CoPilot treats your documents as strictly confidential and never shares or transfers this information outside your Dynatrace environment. Your team’s knowledge stays private, secure, and entirely under your control, while still powering smarter, more context-aware automation.

Learn more about document suggestions and Dynatrace remediation intelligence in our documentation, and read about discovering relevant troubleshooting guides and how to create new ones.

Don’t let organizational knowledge stay buried. Make it actionable. Make it scalable.

The post Remediation intelligence: Accelerate MTTR with AI-powered context and knowledge appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/remediation-intelligence-accelerate-mttr-with-ai-powered-context-and-knowledge/feed/ 0
Powerful exploratory analytics for AI-driven insights https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/ https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/#respond Tue, 04 Feb 2025 16:00:42 +0000 https://www.dynatrace.com/news/?p=67543 Problem alert dashboard

The Dynatrace platform empowers Operations, SRE, and DevOps teams to maintain high software quality, security, and reliability, allowing organizations to innovate and scale confidently. By leveraging Davis® AI with enhanced predictive analytics and automated workflows, Dynatrace simplifies issue detection and resolution, reduces MTTR, and enables proactive incident prevention.

The post Powerful exploratory analytics for AI-driven insights appeared first on Dynatrace news.

]]>
Problem alert dashboard


Deploying and safeguarding software services has become increasingly complex despite numerous innovations, such as containers, Kubernetes, and platform engineering. Recent global IT outages, such as the CrowdStrike incident, remind us how dependent society is on software that works perfectly.

Organizations must balance many factors to stay competitive.
Figure 1. Organizations must balance many factors to stay competitive.

Organizations strive to strike a delicate balance between cost, time to market, and innovation. This challenge is more pressing than ever as businesses seek to stay competitive while ensuring their software remains robust and secure.

This necessitates a comprehensive platform that empowers enterprises to understand IT and software within the broader context of their business operations, giving them confidence that their software and IT infrastructure are reliable.

Scale with confidence: Leverage AI for instant insights and preventive operations

Using Dynatrace, Operations, SRE, and DevOps teams can scale efficiently while maintaining software quality and ensuring security and reliability. Its AI-driven exploratory analytics help organizations navigate modern software deployment complexities, quickly identify issues before they arise, shorten remediation journeys, and enable preventive operations.

We’ve added numerous enhancements to our platform, leveraging advanced AI and automation for smarter software observability.

In this blog post, we show you how to

  • Get AI-driven insights directly on your operations dashboards
  • Improve MTTR with AI-assisted problem analysis and logs and traces in context
  • Leverage Gen AI through Davis CoPilot to get insights into root causes
  • Automate remediation of AI-detected problems with simple workflows
  • Adopt Preventive Operations with AI forecasting and automated action

Get AI-driven insights directly on your operations dashboards

A high-level, customizable view of your data is crucial in modern software operations. Dynatrace Dashboards, powered by Grail™ data lakehouse and Davis® AI, offer precisely that. They provide a comprehensive overview, seamlessly integrating health and problem-related information into a single view. You can chart your topology across data silos alongside all alerts, events, and problems using honeycomb tiles, which offer convenient drill-downs into the problem-debugging user flow.

Dynatrace ensures that context is seamlessly integrated into the platform, thus simplifying complexity for you as a user when analyzing issues and allowing you to focus on what truly matters. AI-driven analytics transform data analysis, making it faster and easier to uncover insights and act. This approach not only improves user experiences, it ensures that critical insights are accessible to both experts and novices. By simplifying remediation journeys and extending features to more user groups, Dynatrace enables results across all teams.

The new Problems dashboard, including rich honeycomb visualization, helps you focus on what’s important, turning technical data into a visual story.
Figure 2. The new Problems dashboard, including rich honeycomb visualization, helps you focus on what’s important, turning technical data into a visual story.

When a truly important issue stands out, the next step is refinement. With a few clicks, you can segment and filter your data to focus on specific applications, assignment groups, or regions. Directly mapping and surfacing ownership information within data segments accelerates incident assignment notifications and triggers automatic remediations.

Utilize the comprehensive filter functionality to update your dashboards dynamically.
Figure 3. Utilize the comprehensive filter functionality to update your dashboards dynamically.

If you see an issue or need to look closely at a specific application where an issue was identified, simply select the element to be seamlessly directed to the Problems app. There, you can dig deeper while continuing to focus on your selected segment. This tight integration, following a golden thread of insights, ensures that you’re more productive. To experience the possibilities of AI-empowered dashboards, try our example dashboard on the Dynatrace Playground.

Improve MTTR with AI-assisted problem analysis, logs, and traces in context

The Problems app delivers opinionated AI-assisted problem analysis optimized for Operations and Site Reliability Engineers (SREs) and developers. According to IDC, guiding users visually and automatically surfacing all critical details enables a 56% faster mean time to repair (MTTR) for critical incidents.

When a large-scale incident occurs, follow the red flag that Davis AI uses to identify the root cause, pinpoint all relevant details, and visually reproduce the details in charts, highlighting the affected deployment.

Analyze the root cause in the Problems app.
Figure 4. Analyze the root cause in the Problems app.

Besides identifying the root cause, Davis AI also automatically connects all relevant log lines. Logs are invaluable for identifying further insights and detecting fundamental flaws, such as process crashes or exceptions. With a single click in Problems, all incident logs are surfaced automatically. But we don’t stop there, Dynatrace also seamlessly integrates relevant trace data, offering full visibility into even complex, microservices-based architectures.

By providing these end-to-end insights, Dynatrace and Davis AI empower SREs, developers, and architects to quickly dive deep into an incident’s details, including all relevant logs and traces. Using this context, they can effectively focus on fixing and remediating code-level issues, significantly improving MTTR, and ensuring that critical incidents are resolved swiftly and efficiently.

Leverage GenAI via Davis CoPilot for insights into root causes

Dynatrace offers precision tools for domain experts to solve complex problems and dig deeper into their data. While product owners often focus on the intricate technical details of an incident, they often prefer a quick summary of what happened and what caused it. The soon-to-be-globally available Davis CoPilot™ bridges this gap by summarizing problems and their root causes and suggesting remediation steps based on these insights.

You’re not limited to one problem; Davis CoPilot can simultaneously analyze multiple problems, draw conclusions about their relationships, identify the common root cause, and propose corrective steps. Instead of relying on a team of experts and waiting hours for insights, Davis CoPilot helps you identify similarities and draw relevant conclusions independently and efficiently.

The use of generative AI adds significant value by augmenting Dynatrace-detected technical root causes with knowledge from the global tech community. Generative AI can access and synthesize vast amounts of information from various sources, providing a broader context and deeper insights. This ensures that your teams benefit from the latest advancements and solutions, enhancing their ability to resolve issues effectively and efficiently.


Dynatrace Problems App - Explain Problems video

Gain a better understanding of root causes with Davis CoPilot
Figure 5. Gain a better understanding of root causes with Davis CoPilot

Automate remediation of AI-detected problems with simple workflows

To automatically remediate Davis AI-detected problems, Dynatrace leverages powerful Workflows. Dynatrace workflows can be triggered by any problem or alerting event, automating domain-specific tasks to take remedial actions.

For example, workflows can scale up capacity to adapt to demand or automatically restart a service in case of a crash. With a large catalog of available workflow actions, you can react efficiently to AI-detected problems, reducing mean time to repair (MTTR) by automatically remediating issues.

But you can do much more with it: The recently introduced Simple Workflows, which are included in your Dynatrace subscription with no extra cost, offer greater flexibility and power than standard notifications. You can use the same mechanisms and trigger types to notify your developer team via Slack, create a JIRA issue, or send a PagerDuty alert.

This ensures that your operations, SRE, and DevOps teams can focus on more strategic tasks while the system handles routine problem resolutions. Automation enhances operational efficiency and ensures that your systems remain robust and reliable, even in the face of unexpected issues.

Easily set up automated remediation with the new Simple Workflows.
Figure 6. Easily set up automated remediation with the new Simple Workflows.

Adopt Preventive Operations with AI forecasting and automated action

Going beyond reactive problem detection, analysis, and remediation, Dynatrace can also leverage predictive AI to anticipate and avoid critical situations before they occur. Using Davis AI forecast, you can easily predict future capacity demands. Combining this knowledge with workflows allows you to take proactive measures to ensure system stability and performance.

Let’s have a look at a concrete example:

It’s easy to predict key indicators of your application, such as order levels or service request counts. Once load and demand rise and Davis AI identifies a potential future issue in your infrastructure setup, Davis CoPilot can automatically generate an updated Kubernetes configuration script for you and automatically upscale the environment to meet future demand. This ensures that your system scales appropriately to handle the anticipated demand, preventing incidents before they occur and eliminating the need to generate a problem.

That’s what we call Preventive Operations. Instead of sending an alert and notifying people, Dynatrace simply fixes the issue. According to Gartner’s Analytics Maturity Model, using predictive AI can significantly reduce the likelihood of incidents by taking preemptive action and remediation.

Start using Davis AI to analyze your environments and predict and address potential issues in advance. This will empower your teams to avoid potential problems and ensure a smooth, uninterrupted user experience.

Initiate automated, corrective action before an issue occurs
Figure 7. Initiate automated, corrective action before an issue occurs.

Tackle business challenges with confidence

Ensure your software runs securely and reliably with Dynatrace and Davis AI.

Dynatrace and Davis AI support you by running your software securely and reliably. This includes advanced root cause analysis, deep insights into detected issues, and corrective actions—whether manual or automatic—to prevent outages before they occur.

Get started

For more information, have a look at our documentation or explore the available resources on the Dynatrace Playground to experience some of these enhancements first-hand:

The post Powerful exploratory analytics for AI-driven insights appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/feed/ 0
Advancing AIOps: Preventive operations powered by Davis AI https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/ https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/#respond Tue, 04 Feb 2025 16:00:06 +0000 https://www.dynatrace.com/news/?p=67673 Davis AI alerts

The 2024 CrowdStrike incident demonstrated our societal vulnerabilities to IT outages. A faulty software update caused widespread issues, impacting critical services globally, including airlines, banks, hospitals, and public safety systems. Despite recent advancements such as containers, Kubernetes, and platform engineering, it’s evident that managing enterprise software services has become increasingly complex. IT operations must be prepared to quickly address and mitigate disruptions, ensuring business continuity and minimizing damage.

The post Advancing AIOps: Preventive operations powered by Davis AI appeared first on Dynatrace news.

]]>
Davis AI alerts

AI, especially AIOps, has emerged as a pivotal solution, promising to avoid downtime. The 2024 State of AI Report highlights this trend, with 89% of technology leaders anticipating that AI will significantly enhance incident response by learning to automate and optimize various tasks, such as performance monitoring and workload scheduling.

Blue screens of death at LGA airport due to the July 2024 CrowdStrike outage. (Source: Wikimedia Commons.)
Figure 1. Blue screens of death at LGA airport due to the July 2024 CrowdStrike outage. (Source: Wikimedia Commons.)

AIOps can identify and address potential issues before they become major incidents by learning from history and analyzing large amounts of data in real time. This approach improves operational efficiency and resilience, though it’s not without flaws. The complexity of IT environments and the changing nature of threats necessitate human oversight and ongoing adjustment of AIOps systems to handle unforeseen challenges and ensure optimal performance. Additionally, predictions based on historical data are reactive, solely relying on past information to anticipate future events, and can’t prevent all new or emerging issues. This limitation highlights the importance of continuous innovation and adaptation in IT operations and AIOps strategies.

“The shift from reactive to preventive operations represents the next evolution in AIOps.”
Bernd Greifeneder, CTO Dynatrace

When Dynatrace set out with Davis® AI over 10 years ago, pioneering AI-driven operations, we focused initially on problem identification before moving on to problem remediation. The next milestone in enhancing the capabilities of Davis AI—another pioneering step forward in AI-driven operations—is outright problem prevention. In this blog post, we explain how the unique combination of causal, predictive, and generative AI—augmented by the latest Davis AI advancements—is transforming how Dynatrace customers manage and optimize their IT infrastructure.

Automatic root cause detection

Modern, complex, and distributed environments generate a substantial number of events. This necessitates additional requirements such as minimizing the total number of issues, eliminating false positives, and conducting accurate root cause analysis.

Dynatrace has a longstanding reputation for accurately analyzing root causes and identifying related events. While other methods typically rely on mere correlation and historical data analysis, we’ve further enhanced our capabilities by implementing causational analysis, which leverages contextual information automatically gathered during data ingestion and processing in addition to historical data analysis. This is achieved using Dynatrace Grail™, our causational data lakehouse, which unifies all data in an always-up-to-date topology model. By applying causal AI to incoming data in real time, Davis instantly learns and continuously adapts to new information. This facilitates more precise root cause analysis and anomaly detection, including identifying seasonal anomalies and establishing auto-adaptive thresholds.

Root cause analysis with the Problems app
Figure 2. Root cause analysis with the Problems app

When applying this Davis root cause detection within our own IT environment, Davis effectively filters out over 99.9% of incoming data noise, condensing hundreds of thousands of daily system events into no more than four or five incidents that require attention from our IT operations team.

These algorithms are not limited to monitoring IT environments. At our February 2025 Dynatrace Perform session on exploratory analytics with AI-driven insights, the Performance Engineering Lead of XXXLutz—one of the world’s largest furniture retailers operating more than 370 stores across Europe—explains how XXXLutz utilizes Davis AI to proactively identify critical order drops, allowing them to respond quickly and effectively to changing market conditions and ensuring that their business remains agile and responsive to the needs of their customers.

Problem journey and reactive remediation

At the core of Dynatrace problem remediation stands the Problems app—an optimized view into opinionated insights, details, and context of each detected issue—for Operations, SREs, and developers. It filters billions of log lines, including the topology of each incident and its affected entities, for efficient problem triaging and troubleshooting, resulting in a 56% faster mean time to repair (MTTR) for critical incidents.

With the latest release, we drive this further by improving the automatic connection of relevant log and trace data for further drill down, presenting the full context of an issue in a single view. This provides comprehensive visibility into even complex architectures, simplifying the process of examining relevant details and addressing code-level issues, reducing 100 clicks and manual filtering to a single click with no loss of context.

Comparative analysis of multiple problems with Davis CoPilot
Figure 3. Comparative analysis of multiple problems with Davis CoPilot

By utilizing Davis CoPilot™, you can conduct comparative analyses of multiple issues, obtain natural language summaries of individual problems, and receive contextual recommendations along with specific remediation steps.

You can also link troubleshooting guides created in Notebooks to remediated issues, thereby building an intelligent knowledge base. Davis automatically connects additional documents as well as stored workflows. So the next time a similar problem arises, Davis brings up related guides, enabling teams to learn from previous experiences and reducing the risk of knowledge loss.

Harness your collective knowledge by connecting troubleshooting guides
Figure 4. Harness your collective knowledge by connecting troubleshooting guides

Please refer to our recent blog posts for more information on utilizing Problems for AI-driven insights and the latest Davis CoPilot advancements.

Automating the remediation

While obtaining comprehensive insights is beneficial, true transformation occurs through the use of tools that automatically execute remediation steps. To implement these “AI-driven operations,” it’s essential to forecast future requirements, including capacity demands, potential system failures, and security incidents.

Traditional forecasting engines typically depend on historical data, stored in metrics. In contrast, Davis AI generates real-time predictions, facilitating proactive operations. This capability is due to Davis’s ability to process raw data, such as logs, for forecasting, leveraging Grail to execute previously unattainable queries.

Consider the following scenario: You begin by retrieving and analyzing logs to identify relevant values for automation. Once this task is complete, you proceed to your pipelining tool to configure ingestion rules that extract these values into metrics and then wait several weeks for your prediction engine to generate alerts that can serve as triggers for your workflows.

However, when utilizing Dynatrace with its integrated anomaly detection and forecasting capabilities, you gain the advantage of schema-less data analysis and the ability to process any raw data into time series in real time. This significantly reduces the time required to establish AIOps workflows from several weeks to less than 30 minutes.

Preventive operations

The complexity of modern software environments makes it challenging to determine a service’s reliability solely through testing. It’s impractical to emulate scenarios such as generating a million tickets to assess performance capabilities. This necessitates real-time insights and operations rather than reactive problem-solving or raising alerts to notify personnel.

Preventive operations address this need by enabling proactive corrective actions before issues arise, akin to predictive maintenance. AI-supported anomaly detection identifies parameters that deviate from the norm, allowing for automatic configuration adjustment to mitigate potential problems preemptively.

Dynatrace offers the only unified, AI-powered platform for all data, all teams, and all possibilities.
Figure 5. Dynatrace offers the only unified, AI-powered platform for all data, all teams, and all possibilities.

Davis CoPilot combines the “power of three”:

  • Davis causal AI for identifying anomalies and root cause analysis
  • Davis predictive AI for precise forecasting and determining when to take action
  • Generative AI capabilities that perform actions beyond simply sending notifications or restarting services

In this way, Dynatrace extends AIOps beyond traditional IT operations tasks and addresses complex scenarios, including security use cases such as threat observability. Consider the following real-world example:

At Dynatrace, we log all failed login attempts. We can predict potential threats when abnormal patterns are identified and raise a security event by utilizing seasonal baselining. The subsequent workflow involves checking the IP address and generating a threat score. Upon reaching a certain threshold, a new ruleset is automatically added to the web application firewall. This entire process is fully automated, running before a problem even occurs, significantly reducing the response time from over an hour to a fraction of a second.

In another instance, automatic log pattern analysis crawling our application logs decreased the number of bugs in the production environment by 15% and freed up time previously spent on log analysis and triaging (in pre-prod), equivalent to 17 full-time employees. Consequently, these 17 developers can now dedicate their efforts to adding more value to Dynatrace.

Summary

The State of AI report states that over 88% of technology leaders anticipate AI will enhance incident responses and improve their teams’ ability to predict and proactively resolve service-affecting issues.

With Dynatrace, organizations are prepared to evolve their ITOps and SRE departments from troubleshooting to prevention, getting proactive with forecasting, and utilizing generative AI instead of purely focusing on history-focused root cause analysis.

Start your preventive operations journey with smart automation and auto-remediation that prevents larger issues.

Are you interested in gaining more insights?

The post Advancing AIOps: Preventive operations powered by Davis AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/feed/ 0
Transform your operations with Davis AI root cause analysis https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/ https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/#respond Tue, 08 Oct 2024 19:03:46 +0000 https://www.dynatrace.com/news/?p=66044 root cause analysis

Complexity is ever-increasing in today’s fast-paced world of software deployments and cloud infrastructure. This is why Davis® AI root cause analysis is an indispensable tool for Operations, Site Reliability, and DevOps teams.

The post Transform your operations with Davis AI root cause analysis appeared first on Dynatrace news.

]]>
root cause analysis

Without AI-assisted observability tooling, the productivity of operations teams drops, leading to a dramatic increase in Mean Time to Repair (MTTR) and a significant rise in the personnel needed to manage critical incidents. In an era dominated by automated, code-driven software deployments through Kubernetes and cloud services, human operators simply can’t keep up without intelligent observability and root cause analysis tools.

Modern observability has evolved from simple metric telemetry monitoring to encompass a wide range of data, including logs, traces, events, alerts, and resource attributes. Dynatrace Root Cause Analysis (RCA) seamlessly integrates all this information, providing crucial analysis to remediate incidents in real time.

Problem feed for fast triage and remediation of AI-detected problems.
Figure 1. Problem feed for fast triage and remediation of AI-detected problems.

By offering root cause analysis on top of the highly flexible Grail™ data lakehouse, Dynatrace empowers SRE and operations teams to further reduce MTTR. Direct access to the underlying data allows the automatic RCA analysis to eliminate data silos and to dive deep into every aspect of the collected incident data.

Unlike generic DIY query frontends, the Dynatrace Problems app is a tailor-made solution for efficiently supporting operations use cases. This approach ensures that your operation teams have all the tools they need to manage modern software deployments.

Transform your operations today with the new Problems app and stay ahead in the ever-evolving software and cloud infrastructure landscape.

Rapid response to critical incidents

Operations teams can quickly focus on incoming Davis AI-detected and -analyzed problems by referring to the problems feed.

The problem feed is designed to prioritize active issues, ensuring they always appear at the top, regardless of how long they’ve been ongoing. This default sorting strategy, which uses time as a secondary criterion, guarantees that Operations teams never overlook an active problem, no matter which primary filter is applied.

You can focus on your domain using the filter bar at the top, the quick filters on the side, or both. The chart feature allows for quick analysis of problem peaks at specific times.

Operations teams will appreciate the ability to sort problems by duration and the number of affected entities. This aids in assessing Davis-detected root causes and prioritizing remediation efforts. The native multi-select feature lets users open a filtered group of problems simultaneously, facilitating quick comparisons and detailed analysis.

Streamline deployment insights with AI-generated summaries

Every second counts during wide-scale incidents affecting large parts of your production systems. This is why precisely showing the root cause ultimately helps to speed up problem resolution.

You can multi-select a cohort of active problems, select Show detail, and review all critical problem details, including preview charts and event details, without losing the context of your problem feed.

The new problem experience transparently displays all the available details, with prominently displayed root-cause markers to precisely guide your attention.

In the realm of cloud infrastructure management, having a clear and concise view of your deployment’s health is crucial. Our dedicated deployment perspective offers just that, showcasing the hierarchy of affected and related infrastructure components. The root cause of any issue is prominently marked with a root-cause badge, making it easy to identify and address problems swiftly.

This perspective not only highlights the affected cloud regions but also provides a quick summary of the Kubernetes context where your workloads encountered failures. Gone are the days of clicking and navigating through multiple dashboards. Instead, you receive an AI-generated summary as an affected deployment architecture diagram.

This diagram, akin to a UML (Unified Modeling Language) deployment diagram, offers a familiar representation for software architects, ensuring they can quickly grasp the situation and take necessary actions. By streamlining the visualization of deployment issues, we empower teams to resolve problems more efficiently and maintain optimal performance.

To save time, the root-cause component is preselected, and all the details of the root cause are displayed on the right, along with charts showing the detected breaches from learned normal behavior.

You can review each individual finding on all problem-affected entities by selecting the individual deployment components or by switching to the detailed event perspective, which shows all the single events that the root cause analysis collected into a single problem.

Confirm the AI-detected root cause and review the deployment context.
Figure 2. Confirm the AI-detected root cause and review the deployment context.

In addition to using markers for swift root cause analysis, operations teams often seek to attach valuable remediation hints and playbooks for familiar scenarios.

By implementing a flexible event tagging mechanism, event sources and detectors can be easily customized to include additional custom event properties. This allows for markdown-formatted event description text that can contain remediation links, as illustrated in the screenshot below.

Root cause remediation hints as markdown links
Figure 3: Root cause remediation hints as markdown links

The Dynatrace Semantic Dictionary helps identify the semantics of well-known event properties and provides convenient platform intents. For instance, entity links (dt.entity.*) or links to the responsible settings entry (dt.settings.object_id) that detected and opened an event can be included. These settings links save valuable time when adjusting detection sensitivity for thresholds or baselines. Additionally, the event setting property can be utilized in a DQL query to create a table of the top-triggering configurations or to automate settings changes using an automation workflow.

Quick access to incident logs

The seamless integration of logs powered by Dynatrace Grail™ data lakehouse with Davis AI root cause analysis is a game changer for modern operation teams, as it offers a quick summary of all incident-relevant logs.

The Dynatrace root cause engine already combines all incident-relevant information to recommend log queries, which saves a lot of navigation time and completely eliminates the need to manually identify complex log filters.

A single click on the Problem details log perspective immediately surfaces all relevant logs related to the given incident, as shown below.

Failure rate increase logs
Figure 4.
100 errors and warnings of failure rate logs
Figure 5.

Within this view the Operations team can further refine the query or adapt the filters and open a notebook to persist the log findings for critical post-mortem documentation purposes.

Root cause analysis in a user-focused context

Most modern application stacks are deployed through Kubernetes, making it essential for operations teams to focus on Kubernetes clusters, cloud resources, and workloads of critical services.

Since operations engineers prefer not to switch contexts, a consistent root-cause experience is provided regardless of where the user journey begins.

Whether you start your remediation journey within the Infrastructure & Operations app or the Kubernetes app, you receive the same root-cause information without needing to navigate between different apps. This seamless embedding of root-cause information into the current context saves valuable time during incident remediation.

Root cause shown in context of the Infrastructure & Operations context.
Figure 6. The root cause is shown in the context of Infrastructure & Operations.
CPU throttling root cause shown in Kubernetes context.
Figure 7. CPU throttling root cause shown in Kubernetes context.

Notify and automate to speed up remediation

The Problems app features a global problem indicator that is always visible within the Dock to capture your attention. This indicator shows whether there are active problems within the environment. You can personalize this number by selecting and saving a problem filter within the problem feed, as demonstrated below. The saved default filter is then automatically applied to the global problem indicator, reducing the number of active problems for the user.

Select Alerting (bell icon) to set up alerts related to filtered problems and configure email addresses for notification recipients.

The email payload and the use of an email address for notifications are preset, allowing for a personalized notification setup, as shown below.

Save the personal default filter and set up email notifications.
Figure 8. Save the personal default filter and set up email notifications.
Find the global problem indicator in the Dock.
Figure 9. Find the global problem indicator in the Dock.

You can take a further step towards answer-driven automation and use the detected Davis problem event to trigger workflow automation. Automatically remediate an issue using our no-code workflow actions for collaboration (for example, Slack, Microsoft Teams, ServiceNow, Pagerduty) and remediation (for example, AWS, Red Hat Ansible, Kubernetes).

The introduction of a filterable global problem indicator ensures that Operations teams remain focused on active problems within the environment, even while exploring data in Notebooks or Dashboards.

In future updates, the Problems app will support multiple named filters and introduce Segments as the primary method for using and sharing numerous predefined filters among operations teams.

Outlook

The newly released Problems app enhances transparency by providing detailed AI-detected root-cause information. It also offers convenient deployment and architectural visualizations, along with a log perspective, to help operations teams reduce Mean Time to Repair (MTTR).

In future updates, we aim to support the ability to acknowledge and label incoming problems, improving team coordination. Additionally, plans include a visual representation of the application map, direct propagation of information such as application IDs into the problem feed, and support for segments to filter the problem feed.

Summary

For over a decade, Dynatrace has been at the forefront of integrating AI into incident analysis, particularly through Davis root cause analysis.

Davis is now essential for Operations, Site Reliability, and DevOps teams, helping them to navigate the complexities of modern software deployments and cloud infrastructure.

Without Davis, the productivity of these teams would plummet, leading to longer Mean Time to Repair (MTTR) and increased staffing needs to handle critical incidents.

In today’s automated deployments and cloud services, traditional observability tools fall short, unable to keep pace with the intelligence needed for effective root cause analysis.

Modern observability encompasses various data sources, from metrics to logs and events, requiring intelligent tools like Davis to seamlessly integrate and analyze this information in real time. By providing Davis on top of the flexible Grail data lakehouse, Dynatrace empowers teams to swiftly reduce MTTR by accessing and previewing incident data comprehensively.

The Davis Problems app streamlines triage, allowing teams to swiftly focus on AI-detected issues. Its intuitive interface simplifies problem resolution.

Try out the new Problems app in the Dynatrace Playground.

The post Transform your operations with Davis AI root cause analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/feed/ 0
The Dynatrace troubleshooting community: Experts tips and tricks to get the most out of Dynatrace https://www.dynatrace.com/news/blog/dynatrace-troubleshooting-community/ https://www.dynatrace.com/news/blog/dynatrace-troubleshooting-community/#respond Wed, 24 Jul 2024 08:00:33 +0000 https://www.dynatrace.com/news/?p=64913 Dynatrace troubleshooting

Get expert tips and tricks about integrating and monitoring your favorite technologies on the Dynatrace platform. The Dynatrace Community is the perfect place to find answers.

The post The Dynatrace troubleshooting community: Experts tips and tricks to get the most out of Dynatrace appeared first on Dynatrace news.

]]>
Dynatrace troubleshooting

Something not working as expected? Not sure what an error message means? At Dynatrace, we’re committed to helping you get the most out of observability and security to keep your software working perfectly. The perfect place to start is the Dynatrace troubleshooting community forum.

What is the Dynatrace troubleshooting community?

The Dynatrace troubleshooting community is a website that hosts articles written by Dynatrace experts with quick answers to common issues.

If you’re unsure about something, it’s easy to engage Dynatrace support directly. Either open a chat right on the platform or open a support ticket. But you can also find self-service guidance and troubleshooting tips and tricks in the Troubleshooting Community. The forum hosts technical experts throughout Dynatrace.

Offering 24/7 self-service, the Dynatrace Troubleshooting community forum is the perfect place to start your Dynatrace support journey. You’ll find fantastic articles that can help you resolve issues and connect to like-minded practitioners to assist along the way.

Dynatrace engineering teams now contributing directly

For the first time, the Dynatrace Troubleshooting community forum is officially enabling our Engineering (R&D) teams to create customer-facing content in collaboration with our support team and community. With this addition, we’re bringing the best possible knowledge directly to you.

Experts from our R&D teams, Customer Success group, and Support organization share their expertise and experience through these troubleshooting articles to provide you with help and guidance if something doesn’t work out as planned.

For example, the Kubernetes and SSO R&D teams recently collaborated with the support team to provide Dynatrace tips and troubleshooting articles for a professional deep dive into specific issues and resolutions such as the following:

Dynatrace troubleshooting and expertise 24×7

Questions are welcome—the article owner, contributor, or one of our Community super-users, whom we affectionately call DynaMights, will answer them.

We want to offer our customers the best possible self-service options for handling issues. We want searchers to find the abundance of know-how and expertise that welcomes all in to learn and grow with us.

While we have only highlighted a few articles here, we’re always releasing more, and more fantastic information is available to help you perfect your software. We’ll continue to innovate and give out the best knowledge possible.

If you can’t locate the answer to your issue, leave us your feedback so we can improve.

Visit the Dynatrace Troubleshooting Community Forum today, we would love to hear your feedback on the articles!

The post The Dynatrace troubleshooting community: Experts tips and tricks to get the most out of Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-troubleshooting-community/feed/ 0
Davis AI: Your personal interactive troubleshooting assistant  https://www.dynatrace.com/news/blog/davis-ai-your-personal-interactive-troubleshooting-assistant/ https://www.dynatrace.com/news/blog/davis-ai-your-personal-interactive-troubleshooting-assistant/#respond Wed, 24 May 2023 08:59:02 +0000 https://www.dynatrace.com/news/?p=57846 business observability

When Dynatrace started reinventing cloud-native service tracing and observability ten years ago, it was already clear that human operators were overwhelmed with traditional monitoring systems' massive raw data inflow. Besides being unable to watch that amount of telemetry data on dashboards, classic operations teams were also blown away by the sheer number of alerts they received 24/7 from hundreds of different monitoring tools. 

The post Davis AI: Your personal interactive troubleshooting assistant  appeared first on Dynatrace news.

]]>
business observability

Update: We’ve expanded AI-powered dashboarding with Dynatrace Intelligence, delivering smarter insights, forecasting, and anomaly detection across the Dynatrace platform.
Dynatrace Intelligence is the evolution of Davis AI®, improving how users interact with and understand observability data.

With the introduction of Davis® root-cause detection, Dynatrace reduced the amount of single-alert spam that arises when large-scale incidences occur. Instead of immediately firing off an alert for all raw events, the Davis root-cause engine follows each violating service’s causal relationships. By automatically following the causal direction of the topology between services and their underlying infrastructure, Davis collects all raw events that belong to the same root cause and then notifies you by raising a problem.

With interactive problem mode, Dynatrace introduces a new, powerful troubleshooting assistant. This blog post explains how Davis can help reduce your MTTR (mean time to resolve) using interactive user guidance that retains context when drilling deeper into problem analysis.

Davis problem analysis
Select any entry in the side panel to navigate to the corresponding metric, in context.

Faster remediation through precise root cause analysis

Once Davis identifies a problem, a Problem overview page is created, which shows a comprehensive management summary of what happened (impact) and the root cause of the problem. DevOps teams use this page to quickly identify and remediate unexpected incidences.

Usually, the journey doesn’t stop here. When the DevOps team has finished their work, software experts must investigate the underlying software stack. They need to analyze all relevant information that Davis found along the deployment stack to avoid such problems in the future. When navigating to the underlying service—identified as the root cause—the problem detail page opens with retained problem context, which includes:

  • Date and time of the current problem, so you don’t need to manually adapt the date and time on each page in the analysis journey.
  • A side panel that interactively informs you about all problem-related information for the relevant service.
  • Davis highlights all relevant problem information on each page you navigate to.

The screenshot below shows how Davis interactively guides you by highlighting all the relevant information with red and yellow markers (on the left side) while showing a list of AI root-cause findings in the side panel on the right (if the Davis side panel is closed, an icon is displayed on the right-hand panel so you can re-open it).

AI root-cause findings

Davis highlighting detected problems in side panel

Optimize your software stack using Davis interactive problem mode

Watch out for red and yellow markers in the navigation section headers—these indicate that Davis has found information related to the problem.

The red marker highlights events and their duration, whereas the yellow marker indicates metric anomalies where suspicious metric change points were found during the problem analysis. The yellow metric change points highlight a point in time, while the red markers represent event durations.

If you select one of the markers (either directly or via the side panel), you can view additional information, such as the timeframe and duration.

Davis AI change point and event markers

Davis AI change point (in yellow on the left) and event duration (in red on the right) markers

Meeting SLO requirements

In addition to providing context to detected problems, Davis also supports you when spikes are detected in connected SLOs (Service Level Objectives). Via the dedicated SLO button in the top bar, service-level objectives relating to the selected service can be reviewed immediately without losing context.

Spikes can easily be investigated by selecting a timeframe and clicking Analyze. Davis instantly collects all connected signals and provides relevant, contextual information. Watch the following video for examples of how the interactive problem mode helps identify SLO-relevant issues.

Davis SLO analysis
Review related Service Level Objectives (SLOs)

Summary

Davis problem detection and root cause analysis is essential for modern AIOps (Artificial Intelligence for IT Operations) and DevOps to minimize the MTTR. Real-time insights are crucial for quickly triaging unexpected incidents and remediating them in a timely manner.

Davis interactive problem mode guides you through all the detailed problem-related information and marks problems visually to make them easier to understand. It also seamlessly integrates user-defined SLOs, including leveraging Davis AI for analyzing SLO degradations, which saves precious time during critical incidents. You no longer need to leave the context of your page when using the side panel for navigational help to dig through all relevant findings and SLOs discovered during root cause analysis.

We’re, of course, highly interested in your feedback! We encourage you to try the interactive problem mode and share your feedback and product ideas via the Dynatrace Community. Every message we receive helps us to continuously improve the Dynatrace platform.

The post Davis AI: Your personal interactive troubleshooting assistant  appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/davis-ai-your-personal-interactive-troubleshooting-assistant/feed/ 0
Find and analyze your web frontend errors faster https://www.dynatrace.com/news/blog/find-and-analyze-your-web-frontend-errors-faster/ https://www.dynatrace.com/news/blog/find-and-analyze-your-web-frontend-errors-faster/#respond Thu, 05 Nov 2020 20:56:47 +0000 https://www.dynatrace.com/news/?p=40683 Web frontend errors

Besides out-of-the-box error count charting for Dynatrace RUM in version 1.206, you can now also leverage granular alerting for individual errors, improved RUM UI flows during error analysis, and extended error information.

The post Find and analyze your web frontend errors faster appeared first on Dynatrace news.

]]>
Web frontend errors

In a world of accelerating deployment cycles, understanding and fixing web front-end errors quickly can make or break your business. For example, developers know that a JavaScript error can ruin the user experience of your web application, so you want to find such errors quickly, analyze them, and ensure that they don’t occur again. Dynatrace Real User Monitoring (RUM) allows you to be proactive, reduce the number of front-end errors, and massively improve user experience.

Whether you check errors based on Davis®-reported problem alerts that you receive or are proactively exploring them in the Dynatrace web UI, your problem analysis should be based on an intuitive workflow that is consistent and explicit. To provide you with such a solution, we’ve further enhanced Dynatrace Real User Monitoring.

In August, as a first step, we announced extended Davis AI awareness of HTTP and custom errors with automatic monitoring of your JavaScript, HTTP, and custom errors. Now, with Dynatrace version 1.206, we also enable you to alert on individual HTTP and custom errors . Additionally, we’ve made error counts available consistently across the product and aided charting of error counts per type on your dashboards. We’ve improved the RUM web UI with additional links and flows to switch between Dynatrace entities, better labels and names for consistency, and extended error information.

Catch errors faster, create tailored front-end error alerts with Davis, and more

Most monitoring solutions available today try to distinguish themselves by providing you with “the best” dashboards, lists, filters, or information to help you find and analyze your front-end errors. While all these elements are certainly important, simply crossing them off a “must-have” list isn’t enough! At Dynatrace, we believe that a truly valuable solution automatically detects the front-end errors that matter to you and provides guidance and comprehensive explanations as to where and why they occurred. Let’s look at two concrete scenarios to illustrate this:     

  • As an application owner, you frequently need to report on and improve the error counts of your applications. But you aren’t only interested in error counts per application; you’re probably more focused on ensuring that a specific geolocation or user group doesn’t encounter any errors. Now, using Dynatrace error count metrics, you can tailor the Dynatrace Davis AI to only alert you of errors that matter to you and track those errors over time to optimize your application.
  • As a developer, part of your regular performance-improvement routine likely includes exploring and catching errors. For example, as you go through our aggregated waterfall charts to analyze user action performance, you can now see errors from the edge to the core so that you can immediately fix them.

With other new improvements, you can:

Get Dynatrace Davis alerts on your error count metrics

We’re excited to announce that with Dynatrace version 1.206, we now provide error counts in the Multidimensional analysis for RUM, which you can now use with calculated metrics. So now you can create your own HTTP or custom error counts for a specific error and use them to set up Davis alerting.

For example, you can make sure that your employees have the required permissions to access certain resources in your internal web application by setting up alerts for HTTP 403 errors (see the image below). Or you can answer questions like “How many of my loyalty customers in Japan run into unavailable resources” by setting up HTTP 404 error alerts that leverage session and action properties to segment your traffic.

Note that we’re working on providing you with the ability to create custom counts and alerts for JavaScript errors too!

Create error count metrics to use them for custom error alerting

Easily visualize error counts per type by pinning custom charts to your dashboard

Besides making error counts available for calculated metrics, we’ve also provided some out-of-the-box application-level error counts that you can use in custom charting. We’ve made it easy for you to pick and chart the correct time series. Simply use the two-click approach shown in the screenshots below to pin error counts per type onto your dashboards.

It’s as simple as selecting Create custom chart on the Multidimensional analysis page for errors and then selecting Pin to dashboard on the custom chart. If you need to change the default settings, you can adjust them on the custom chart.

Create a custom chart from multidimensional analysis of RUM errors

Custom chart for error count, showing metric settings

See errors from the edge to the core

At Dynatrace, we’re convinced that automation isn’t just a matter of problem detection but should also be reflected in the user interface and product experience. So we’ve added hints and links at just those places that we believe will aid your error analysis by enabling you to draw the right conclusions and find true insights.   

  • Instantly see which user actions and sessions are affected by JavaScript errors

You can now see information for JavaScript errors that was previously only available for HTTP and custom errors. For example, we’ve added the ability to navigate directly from the JavaScript error details page to the affected user actions and user sessions.

Directly jump to affected user actions and user sessions from JavaScript error details

  • Analyze individual occurrences with full context

This improvement covers the use case where you likely start off by looking into how to improve the performance of a certain user action and then find errors in an aggregated waterfall—now the Errors finding displays a link and guidance per error type to help you further analyze such errors.

On aggregated waterfalls, follow links per error type to analyze individual occurrences with full context

These links automatically take you to the individual waterfall instances filtered by the error type you selected (for example, JavaScript errors) on the aggregated waterfall. We now clearly mark the user action instances that are affected by errors. The new Error count column shows the errors per instance so that you can better prioritize which instances to look at first.

Use the Error count column to prioritize what to look at first

Select an instance and select the Errors finding to see the individual errors.

Use the Errors finding to display individual errors

  • Compare user action instances to their aggregates to check “normal” behavior

Another improvement that’s very useful for troubleshooting and performance optimization is the ability to jump from a single user action instance in the waterfall chart (the micro level) to aggregated information on the user action details page (the macro level). This is helpful when you want to compare timings and information for an individual user action with aggregate information to check what’s considered normal.

User action instance waterfall

Run better-informed error analysis by accessing the summary of all occurrences

To empower you to run better-informed, high-level error analysis and segmentation, we now show a generic Errors column in all tables and charts. With extended Davis awareness of HTTP and custom errors in RUM, this column now sums up all error occurrences of all error types. You can now use this complete error information when prioritizing errors for analysis and breaking them down by dimensions such as geolocation or user type.

New Errors column in the user type card

New Errors column in the world map and in geolocation breakdowns

In addition to all application-related pages in the Dynatrace web UI, we’ve also introduced consistent error counts per type and per user action in USQL, which means that you can now use USQL to check for the user actions that have the most JavaScript, HTTP, or custom errors.

Note: If you’re already working with the existing error counts in USQL or the Session export API, please review the details below that explain how the counts will change and what you might need to do.

Leverage consistent error counts per type in USQL

Capture every failed image and every error page by using extended error information

  • Easily catch failed images

Now you can also use extended error information to find failed images. These are indicated with the prefix HTTP <unknown>: in the error name on the Multidimensional analysis page for errors. (Note that error names will be changed to use the prefix “Failed image:…” in a future release.)

Failed image in HTTP error list

Error details for a failed image

Failed image in a waterfall chart

In another improvement, we’ve added the HTTP status code on the HTTP error details page so that you don’t need to look these up.

HTTP status code on error details page

  • Capture every single error page

Remember the traditional HTTP 404 error pages that featured cats, dogs, or even Lego figures pulling out an electrical plug? While such pages can be fun, we’ve seen more than a few HTTP error pages that aren’t well designed (think of “Product not found” pages that return an HTTP response code of 200 and are therefore not captured as HTTP errors by Dynatrace RUM).

We’ve now introduced the ability to mark such pages as “true” error pages by using our dtrum JavaScript API. All such pages are then counted as additional HTTP errors and are also highlighted in waterfall charts. The following illustration shows a page that was marked as an error page using the command dtrum.markAsErrorPage(501, "error").

Mark your pages with an HTTP status code of 200 as error pages

Analyze the error type you’re most interested in using the simplified Error type filter

When you select an error type or user type using the filters in the upper-right corner of the Multidimensional analysis page for errors, your selections are automatically propagated to the Detail analysis below.

Filters propagated on the Multidimensional analysis page for errors

Previously, you could filter displayed user actions per error type by selecting, for example, Has JavaScript errors. As the underlying user action could also have other associated error types, this turned out to be confusing. So we’ve changed the logic and the name of the filter. The new Error type filter now does exactly what you’d expect—it only displays the error count of the error type that you select.

Use the simplified Error type filter to look up an error type of interest

Update on error counts in USQL and Session exports

With Dynatrace version 1.204, we’ve introduced the following three new error counts in USQL and Session export for every user action:

  • javaScriptErrorCount
  • requestErrorCount (already configured for the name change of HTTP errors to Request errors. See the What’s next section below for details)
  • customErrorCount
  • useraction.failedXHRRequests—Considers only failed XHR calls made in your end user’s browser, which could possibly can be the flip side of your server-side HTTP errors (i.e., httpRequestWithErrors).

These error counts are fully consistent across Dynatrace. Going forward, the existing totalErrorCount for every session will be the sum of these three new error counts.

In turn, we’ll deprecate these existing error counts starting with version 1.212.

  • useraction.errorCount—Only includes JavaScript errors.
  • useraction.httpRequestsWithErrors—Considers only server-side errors.
  • useraction.failedImages—Will be part of the new requestErrorCount.

What should I do now?

If you currently use error counts that will be deprecated in USQL or Session export, please use the following replacements:

  • Use javaScriptErrorCount instead of useraction.errorCount.
  • Use requestErrorCount instead of useraction.httpRequestsWithErrors and useraction.failedXHRRequests.

What happens after version 1.212?

  • USQL and the Session export API result will no longer serve the deprecated counts.
  • USQL dashboards using the deprecated counts can’t be displayed, and you’ll need to adapt the underlying queries.
  • User session pages will be based on the new counts.

What’s next

Further improvements that we’re already working on include:

  • Updated session pages that consider and display the new error counts per user action.
  • CSP violations for Chrome browsers—these new errors will be part of a new, more generic error type called “Request errors,” into which HTTP errors and Failed images will be merged.
  • Filter and search for individual JavaScript errors.
  • The ability to use individual JavaScript errors for calculated metrics, charting, and creating individual error alerts.

Questions?

We’d love to hear your feedback. Please share your feedback with us at Dynatrace Community and feel free to mention any other improvements you believe will help you in better exploring and analyzing your web front-end errors.

Or, if you’re new to Dynatrace, start your free trial now to explore all our performance and troubleshooting capabilities for your front-end applications.

The post Find and analyze your web frontend errors faster appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/find-and-analyze-your-web-frontend-errors-faster/feed/ 0