Observability Archives | Dynatrace news https://www.dynatrace.com/news/category/observability/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 16 Jul 2026 13:35:42 +0000 en hourly 1 Seeing the core clearly: Dynatrace mainframe monitoring meets IBM z17 https://www.dynatrace.com/news/blog/seeing-the-core-clearly-dynatrace-mainframe-monitoring-meets-ibm-z17/ https://www.dynatrace.com/news/blog/seeing-the-core-clearly-dynatrace-mainframe-monitoring-meets-ibm-z17/#respond Thu, 16 Jul 2026 13:35:42 +0000 https://www.dynatrace.com/news/?p=74813 Connecting Logs and Traces related content

Enterprise computing just got a major upgrade. With the launch of IBM z17™, organizations running mission-critical workloads now have access to a platform purpose-built for AI at scale, end-to-end automation, and advanced security capabilities. For the enterprises that run the world’s most important transactions in banking, insurance, retail, and government, this is a significant step […]

The post Seeing the core clearly: Dynatrace mainframe monitoring meets IBM z17 appeared first on Dynatrace news.

]]>
Connecting Logs and Traces related content

Enterprise computing just got a major upgrade. With the launch of IBM z17™, organizations running mission-critical workloads now have access to a platform purpose-built for AI at scale, end-to-end automation, and advanced security capabilities. For the enterprises that run the world’s most important transactions in banking, insurance, retail, and government, this is a significant step forward.

A more powerful platform also means more complexity to manage. More AI workloads in production, more transactions per second, and more pressure on the teams responsible for keeping it all running.

That’s where Dynatrace comes in. As a trusted IBM partner, Dynatrace brings AI-powered, full-stack observability to the mainframe, giving teams the visibility and intelligence they need to operate IBM z17 environments with confidence.

IBM z17™: A new era for enterprise AI and automation

IBM z17 is designed for the next era of enterprise computing, bringing AI, automation, and security closer to the mission-critical transactions that run the business.

With the IBM Telum II™ processor and IBM Spyre™ Accelerator, organizations can run AI inferencing directly within transaction flows, helping support real-time decisions without moving data off-platform.

But as AI-powered workloads become more embedded in core systems, teams need more than infrastructure performance. They need real-time visibility into how those workloads behave across the full hybrid environment.

The observability challenge on the mainframe

As organizations accelerate AI adoption and automation on IBM z17, how do teams know it’s all working?

Mission-critical mainframe workloads have always been difficult to observe. Legacy monitoring tools were built for a different era, designed for periodic sampling rather than real-time intelligence. They tell you something went wrong after the fact, not why, and not how to fix it.

Bringing AI inference into live transaction flows raises the stakes. Unexpected model behavior, workload spikes, or latency issues can affect critical services, so teams need to understand what is happening and respond in seconds, not hours.

Dynatrace mainframe monitoring is built specifically to close that gap.

Dynatrace mainframe monitoring: Full-stack visibility for the core

Dynatrace mainframe monitoring gives enterprises deep, real-time visibility into their IBM Z environments, unified with the rest of their technology stack in a single platform, helping to reduce silos, minimize manual correlation, and support more informed decision making.

Here is what that means in practice:

  • Real-time performance monitoring. Dynatrace continuously monitors CPU utilization, response times, transaction throughput, and resource consumption on IBM Z, surfacing anomalies when they occur rather than after a batch job completes.
  • AI-powered root cause analysis. Dynatrace Intelligence, an agentic operations system, can analyze billions of dependencies across mainframe, cloud, and hybrid environments. When something goes wrong, Dynatrace Intelligence pinpoints the root cause, reducing mean time to resolution and freeing teams from manual investigation.
  • End-to-end transaction tracing. Dynatrace traces transactions across the full hybrid stack, from the web front end through microservices, APIs, and into the mainframe core. Teams get a complete picture of how workloads behave across every tier.
  • Unified observability across hybrid cloud. Whether workloads run on IBM z17, in a public cloud, or in containers on OpenShift, Dynatrace brings it all into a single view. No more switching between tools or reconciling data from disconnected systems.
  • Security and compliance insights. Dynatrace can surface runtime vulnerabilities, detect unusual behavior, and generate compliance-relevant telemetry to extend the security capabilities built into IBM z17.

IBM z17 and Dynatrace: From infrastructure power to operational intelligence

IBM z17 delivers the infrastructure power: AI at the core, intelligent automation, and ironclad security. Dynatrace delivers the operational intelligence: real-time visibility, automated answers, and end-to-end context. Together, they give teams what they need to harness that power with confidence.

In practice, that means organizations can:

  • Accelerate AI adoption on IBM z17 with the observability to validate model performance and catch issues before they impact customers.
  • Automate IT operations from the mainframe to the cloud with AI-driven insights that eliminate manual toil.
  • Better maintain the reliability, security, and compliance that mission-critical environments demand, with enhanced transparency across every workload.

IBM z17 raises the ceiling for what enterprise infrastructure can do. Dynatrace enables organizations to see everything happening across it and act on operational intelligence in real time.

See what’s possible

If your organization runs workloads on IBM Z, now is the moment to ensure you have the observability to match the platform’s capabilities. IBM z17 is built for the future of enterprise computing. Dynatrace is built to help you operate it.

Continue exploring Dynatrace mainframe monitoring.

The post Seeing the core clearly: Dynatrace mainframe monitoring meets IBM z17 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/seeing-the-core-clearly-dynatrace-mainframe-monitoring-meets-ibm-z17/feed/ 0
Dynatrace named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms for the 16th consecutive time https://www.dynatrace.com/news/blog/2026-gartner-magic-quadrant-observability-platforms/ https://www.dynatrace.com/news/blog/2026-gartner-magic-quadrant-observability-platforms/#respond Wed, 15 Jul 2026 16:45:57 +0000 https://www.dynatrace.com/news/?p=74637 GartnerMQ-2026

​Observability began in a world where software was more predictable. You could usually see what went wrong and where. AI changes that. Now a system can be technically healthy and still produce a bad answer, break a policy, or quietly burn money at scale.​ That shift is forcing observability to evolve quickly. It must account […]

The post Dynatrace named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms for the 16th consecutive time appeared first on Dynatrace news.

]]>
GartnerMQ-2026

​Observability began in a world where software was more predictable. You could usually see what went wrong and where. AI changes that. Now a system can be technically healthy and still produce a bad answer, break a policy, or quietly burn money at scale.​

That shift is forcing observability to evolve quickly. It must account for behavior, judgment, cost, and risk, not just uptime. That’s why we’re proud to share that Dynatrace has been named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms, marking the 16th time Dynatrace has been recognized as a Leader in this report. We think Dynatrace continues to deliver consistent business value for our customers at scale as technology evolves.​

We believe this recognition reflects a simple reality: Teams need a unified view across applications, infrastructure, cloud environments, and AI systems in one place, with the context to get answers—not guesses—about what is happening, why it matters, and what to do next.

The AI era demands end-to-end visibility​

Observability was already under pressure. Cloud-native architectures, distributed systems, and constantly changing environments made it harder to follow cause and effect across the stack. AI raises the stakes.​

AI systems do not fail like traditional software. They can look healthy at the system level while producing low-quality outputs, violating policy, exposing sensitive data, or quietly consuming far more resources than expected. Those problems often sit outside the view of conventional monitoring tools. And when observability is fragmented across separate products, teams lose the context they need to understand how AI behavior connects to infrastructure, applications, and business outcomes.​

Organizations need a complete, connected view from GPU to business outcome. That is what makes it possible to adopt AI safely, operate it confidently, and avoid trading speed for risk.

The Dynatrace difference: A unified platform, powered by AI and built for AI​

Dynatrace approaches observability with a unified platform, not a collection of fragmented tools. Built on Grail®, Smartscape®, and Dynatrace Intelligence — with integrations into the tools, clouds, and AI agents your teams already rely on — the platform brings together a unified data foundation, real-time contextual understanding, and AI-powered intelligence to help teams understand complex systems and act with confidence.​

That matters for two reasons:​

  • Dynatrace is powered by AI. Dynatrace applies deterministic AI to deliver precise, trustworthy answers across applications, infrastructure, and cloud environments. Dynatrace agentic AI can then act on those answers, helping teams move faster, resolve issues sooner, and execute at scale with confidence. ​
  • Dynatrace is built for AI. AI is now becoming part of the software stack itself, and it introduces new failure modes that traditional observability cannot fully explain. Dynatrace gives teams visibility into what their AI is actually doing — not just how the system is performing — with insight into performance, cost, quality, and compliance, all on the same platform and with no additional tooling required.​

Together, these capabilities help organizations move beyond isolated dashboards and alerts. They make it possible to observe, analyze, and automate across modern environments with the context required to keep AI systems reliable, governed, and aligned to business goals.​

We believe this recognition reflects where the market is going

The observability market is changing quickly as organizations invest in AI-powered applications, modernize technology stacks, and look for ways to reduce operational complexity. In this environment, platform depth, unified data, context, and AI matter more than ever.​

​We feel Dynatrace’s continued recognition as a Leader reflects the strength of this approach: A platform that helps customers unify and contextualize data across complex environments, transform it into actionable answers, and support intelligent automation at scale. We think sixteen times as a Leader also speaks to consistent delivery through wave after wave of technology change.​

Read the full Gartner® report​

We’re proud of this recognition, and grateful to the customers, partners, and teams that continue to push observability forward with us.​

Read the 2026 Gartner® Magic Quadrant™ for Observability Platforms to learn more about why Dynatrace was recognized as a Leader and how we think the category continues to evolve in the AI era.

Access the 2026 Gartner® Magic Quadrant™ for Observability Platforms report.

FAQ

What does it mean that Dynatrace was named a Leader in the Gartner® Magic Quadrant™ for Observability Platforms?

The Gartner Magic Quadrant evaluates vendors based on Completeness of Vision and Ability to Execute. We believe Dynatrace’s position as a Leader reflects the strength of our unified observability platform and our ability to help customers manage modern complexity at scale in the AI-era.

Why does observability need to change in the AI era?

AI introduces new kinds of operational risk. A system can appear healthy while still producing poor outputs, violating guardrails, or increasing cost. Teams need observability that can connect AI behavior to the rest of the environment and provide context across performance, cost, quality, and compliance. 

What makes the Dynatrace approach different?

Dynatrace combines a unified data foundation, real-time topology and context, and AI-powered intelligence in one platform. That lets teams move from fragmented signals to precise answers and intelligent action, without relying on separate tools to understand AI systems. Dynatrace offers a combination of:

– A unified platform with Grail® lakehouse, exabyte-scale data foundation
– Business-aware insights with Smartscape® real-time topology and contextual understanding
– Answers, not guesses and governed agentic automation with Dynatrace Intelligence

Gartner Disclaimer

Gartner, Magic Quadrant for Observability Platforms, Padraig Byrne, Martin Caren, D.B. Cummings, Neil Young, 13 July 2026

Gartner does not endorse any vendor, product or service depicted in its research publications and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner research publications consist of the opinions of Gartner’s Research & Advisory organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this research, including any warranties of merchantability or fitness for a particular purpose.

Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates

Dynatrace was recognized as Compuware from 2010-2014.

The post Dynatrace named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms for the 16th consecutive time appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/2026-gartner-magic-quadrant-observability-platforms/feed/ 0
2026 Gartner® Magic Quadrant™ for Observability Platforms https://www.dynatrace.com/gartner-magic-quadrant-for-observability-platforms/ Mon, 13 Jul 2026 14:23:25 +0000 https://www.dynatrace.com/news/?post_type=analyst-reports&p=65241 The post 2026 Gartner® Magic Quadrant™ for Observability Platforms appeared first on Dynatrace news.

]]>
The post 2026 Gartner® Magic Quadrant™ for Observability Platforms appeared first on Dynatrace news.

]]>
Dynatrace Release Radar 06.26 https://www.dynatrace.com/news/blog/dynatrace-release-radar-06-26/ https://www.dynatrace.com/news/blog/dynatrace-release-radar-06-26/#respond Thu, 09 Jul 2026 16:52:43 +0000 https://www.dynatrace.com/news/?p=74748 Release Radar

This series covers recent Dynatrace releases and updates, focusing on what’s new, what’s changed, and how these recent enhancements can benefit you and your organization. Each post covers newly available capabilities and where to explore them.

The post Dynatrace Release Radar 06.26 appeared first on Dynatrace news.

]]>
Release Radar

If you want to see them in action, head over to our Release Radar launchpad on the Dynatrace Playground.

Smartscape gets a unified topology view and ad-hoc filters

In a significant Smartscape update, a new All topology view shows every relationship for a given node in a single graph: the infrastructure stack, communication flows, and relationships such as monitoring, load balancing, routing, and API dependencies. Where the existing Vertical and Horizontal views each focus on a subset of relationships, the All view provides a more comprehensive view of relationships, from any node in any app across the platform.

Two changes make these views faster and more focused:

  • The AWS and Kubernetes views now use flat layouts instead of nested ones, bringing the same relevance-based edge fetching and priority-driven node loading used elsewhere in Smartscape for more consistent visibility across your cloud landscape.
  • New ad-hoc node and edge filters let you narrow any view by node type, cloud and infrastructure labels, team ownership, environment, and other properties.

Filters work alongside segments and are saved in the URL, so you can bookmark and share a focused view with segment, timeframe, and filters all preserved.

Ad-hoc filters narrow a Smartscape view by team ownership and environment, highlighting matching nodes and preserving the filter state in the URL.
Ad-hoc filters narrow a Smartscape view by team ownership and environment, highlighting matching nodes and preserving the filter state in the URL.

Press Ctrl+F (Cmd+F on Mac) in any Smartscape view to find nodes by name or ID. Matching nodes are highlighted in the graph, and the legend is narrowed to matching entity groups.

For broader context on Smartscape, see The new Dynatrace Smartscape improves operational efficiency across clouds, Kubernetes, infrastructure, and more.

AI Observability gains LLM evaluation, OpenInference, and Python instrumentation

AI applications can fail without obvious indicators — returning responses at normal speed with no errors, while delivering answers that are inaccurate, unsafe, or inconsistent. Traditional performance monitoring misses this entirely.

dt-evals is a new open source CLI that closes that gap. It pulls live gen_ai.* spans directly from your Dynatrace environment. Built-in evaluators use an LLM judge to score real production interactions for faithfulness, hallucination, relevance, toxicity, bias, PII leakage, prompt injection, and drift. The judge writes structured results back to Dynatrace as business events. Evaluation scores sit alongside latency and error metrics in the same dashboards. These scores can trigger alert workflows and gate CI/CD releases based on quality thresholds the same way that performance metrics do. For the thinking behind this approach, see Evaluate LLM and agent quality in Dynatrace AI Observability and LLM evaluations as a foundation for trustworthy agentic AI systems.

Evaluation quality scores, pass rates, and drift trends from dt-evals running alongside model latency and token usage — turning AI quality into the same kind of operational signal as performance.
Evaluation quality scores, pass rates, and drift trends from dt-evals running alongside model latency and token usage — turning AI quality into the same kind of operational signal as performance.

Dynatrace OneAgent now automatically instruments Python applications that use AWS Bedrock, OpenAI, Azure OpenAI, and LangChain. Dynatrace captures distributed traces, logs, and AI-related telemetry for supported model interactions — provider, operation, model, duration, token usage, and prompt and completion metadata where available. To capture prompt and completion content, go to OneAgent features and turn on Python OpenAI prompt capture.

The same visibility extends to teams using OpenInference with OpenTelemetry (OTel). Dynatrace ingests OpenInference traces and normalizes them to the same gen_ai.* attribute schema — covering model usage, token consumption, prompts, completions, agents, tools, embeddings, and guardrails — so OTel-instrumented applications get consistent telemetry without switching instrumentation frameworks.

As AI adoption grows, evaluation and instrumentation together turn AI services into observable, governable assets rather than black boxes.

Logs gains pattern analysis, Kubernetes insights, and in-context traces

Log analysis gets three meaningful upgrades.

Log pattern analysis (Preview) lets you aggregate query results in Logs into patterns that cluster similar logs together. You can focus quickly on recurring errors, reduce thousands of similar logs to a handful of patterns, recognize the changing parts of a pattern (and their datatypes), and reuse the generated Dynatrace Pattern Language (DPL) for other queries or in OpenPipeline.

Log pattern analysis grouping thousands of similar entries into a handful of patterns, with dynamic segments highlighted and DPL ready to reuse.
Log pattern analysis grouping thousands of similar entries into a handful of patterns, with dynamic segments highlighted and DPL ready to reuse.

In-context trace details mean that when you investigate a log entry with trace context, you can open the associated trace directly inside Logs. A waterfall icon signals that you stay in context rather than navigating away to Distributed Tracing.

Log insights in ready-made Kubernetes dashboards provide built-in log analytics for clusters, namespace workloads, namespace pods, and node pods. Error log counts appear alongside health metrics, with log level distribution and severity trends below. Direct links to the Logs app ensure that a deeper investigation is only one click away.

Faster service investigation with the Services Explorer Preview

The Services app now includes a visual service map that overlays performance and health indicators on service-to-service relationships and messaging flows. It’s the fastest way to understand blast radius during an incident, providing a single view of topology context, performance signals, and bottlenecks without switching views.

The Services Explorer service map overlaying performance indicators on service-to-service relationships to pinpoint blast radius during an incident.
The Services Explorer service map overlays performance indicators on service-to-service relationships to pinpoint the blast radius during an incident.

You can also filter services directly by primary Grail fields such as k8s.cluster.name, k8s.namespace.name, aws.region, and azure.location — the same attributes that power segments across Dynatrace. Both capabilities are available in the Explorer Preview view and open for feedback before general availability; see the Community post for details.

New security integrations and a Kubernetes security tab

Threat Observability expands its ingestion options with new integrations. Dynatrace now integrates with Checkmarx for software composition analysis and container security findings, and adds CrowdStrike and Kyverno integrations — pulling detection findings and Kubernetes policy compliance data into Dynatrace as security events. For Kyverno, see Ingest Kyverno compliance findings.

Kubernetes monitoring also gets a dedicated security tab (Kubernetes app version 1.42.0+) that replaces the Vulnerability tab in the Explorer, bringing security context into the same place teams already investigate cluster health.

The Security tab surfacing vulnerability, detection, and misconfiguration findings alongside Kubernetes cluster health — without leaving the monitoring context.
The Security tab surfacing vulnerability, detection, and misconfiguration findings alongside Kubernetes cluster health — without leaving the monitoring context.

Runtime Vulnerability Analytics now has a native interface, replacing the legacy management-zone-based monitoring rules with a single consolidated workflow.

One change worth flagging for security teams: ingested security.events must now carry a timestamp within −1h/+10min, tightened from the previous −24h/+10min window. Events with older timestamps are dropped, so please review any pipelines that backfill security events.

Performance, drilldowns, and navigation improvements

Improved discovery of ready-made dashboards. Ready-made dashboards deliver instant insights without requiring complex queries. Finding, installing, configuring, and customizing them is now more straightforward — so new users get value faster and experienced users can build confidently on best-practice templates.

The Hub discovery workflow guides you from platform search to installable ready-made dashboards.
The Hub discovery workflow guides you from platform search to installable, ready-made dashboards.

Contents tab added to all extension apps. All extension apps in Dynatrace Hub now include a Contents tab that surfaces the extension’s ready-made dashboards, so you can quickly go from installation to insights.

Session Replay has two improvements:

  • Full-screen mode is now available, removing viewport constraints during playback.
  • Navigating to a session through Error Inspector now opens Session Replay directly in context, keeping the investigation continuous.

Cleaner Smartscape topology. Inactive Synthetic Locations no longer appear in Smartscape, keeping topology views focused on what’s live.

Smartscape navigation for database tables and indexes. Direct navigation intents let you jump from a database node to its table or index detail view in one click.

Filters stay with you. Automated filtering suggestions scope correctly to OR and AND conditions across all apps. Filter state, search terms, and highlights survive page reloads. HTTP Status Filter selections persist through navigation steps in Distributed Tracing.

DQL durations support decimals. Duration literals (h, m, s, ms, us, ns) now accept decimal numbers — for example, 0.5h or .2m. Note, however, that this doesn’t apply to calendar durations.

More headroom in Distributed Tracing. The log viewer no longer caps at 1,000 entries, with full deduplication across trace and span IDs. Span scan limits are configurable from settings (default 5,000, up to 10,000). Field naming is also cleaned up — Smartscape fields drop the redundant prefix, and classic ME fields are clearly labeled.

More allowlist entries for external requests. You can now add up to 100 allowlist entries, double the previous limit of 50, with existing entries preserved across all environments.

Affected entity names enriched in problem records. A new affected_entity_names array is now populated alongside the existing affected_entity_ids and affected_entity_types arrays, index-aligned across all three.

The Problems feed displaying affected entity names alongside IDs, enabling notification workflows and integrations to reference entities without a separate lookup.
The Problems feed displays affected entity names alongside IDs, enabling notification workflows and integrations to reference entities without a separate lookup.

This brings the 3rd-gen platform to parity with classic problem notifications and enables notification workflows and external integrations to reference entity names without additional lookup. The Problems app v1.27 reached General Availability on June 29.

Proactive Cost Intelligence across your entire stack

Dynatrace now makes it easier to understand costs, act before they spike, and optimize with less effort. New Optimize documentation walks Dynatrace Platform Subscription (DPS) customers through the full journey from understanding to optimizing costs, aligned with the FinOps Foundation framework.

Dynatrace Assist surfaces the root cause of a cost spike directly from billing usage events, without requiring specialist knowledge.
Dynatrace Assist surfaces the root cause of a cost spike directly from billing usage events, without requiring specialist knowledge.

The bigger shift is that Dynatrace Assist can now do the cost analysis work for you, designed to reduce the need for specialist expertise. You can ask it to:

  • Understand spikes — “I received a notification that costs have increased. Can you find anything notable?” returns the root cause along with a full drilldown into your billing_usage
  • Predict costs — “Based on my log ingest usage over the last 90 days, can you predict my usage for the next 30 days?” returns a capability-level forecast based on actual consumption, useful when onboarding new teams.
  • Optimize usage — “Are there any log queries duplicated by multiple users?” surfaces overlapping queries with concrete suggestions to improve them.

For more on building cost discipline into your observability practice, see Driving your FinOps strategy with observability best practices.

Why these changes matter

Taken together, the June releases make everyday investigation work feel less fragmented. You get more context in the places where teams already troubleshoot: a fuller Smartscape view, AI quality signals alongside performance data, log patterns that identify root causes faster, service maps for incident response, and security and cost insights that are easier to act on without switching tools or relying on specialists.

These are the kinds of changes that add up across a week of real work.

Check out all these updates in action on our Release Radar launchpad.

The post Dynatrace Release Radar 06.26 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-release-radar-06-26/feed/ 0
Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/ https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/#respond Thu, 02 Jul 2026 23:42:51 +0000 https://www.dynatrace.com/news/?p=74651 NVIDIA and Dynatrace

Enterprise AI is rapidly evolving from standalone models to agentic AI systems, where multiple AI agents collaborate to gather information, reason across data sources, and generate complex outputs. These systems unlock powerful new capabilities, but they also introduce significant operational challenges. Organizations must be able to observe, govern, and optimize AI agents, models, and infrastructure in real […]

The post Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q appeared first on Dynatrace news.

]]>
NVIDIA and Dynatrace

Enterprise AI is rapidly evolving from standalone models to agentic AI systems, where multiple AI agents collaborate to gather information, reason across data sources, and generate complex outputs. These systems unlock powerful new capabilities, but they also introduce significant operational challenges. Organizations must be able to observe, govern, and optimize AI agents, models, and infrastructure in real time.

Dynatrace helps support this need by providing broad visibility across key layers of the AI stack—from agent orchestration and model inference to GPU infrastructure and enterprise applications. With Dynatrace, teams can monitor AI workflows, understand model behavior, optimize costs, and help improve reliability as agentic systems scale.

Why agentic AI needs full-stack observability

As organizations build GPU-accelerated platforms for AI training and inference, understanding system behavior becomes increasingly complex, with bottlenecks potentially occurring anywhere – from GPU utilization, model latency, token consumption, and downstream service dependencies.

Dynatrace connects these layers through full-stack AI observability, designed to help teams monitor model performance, trace multi-agent workflows, track GPU and infrastructure utilization, detect bottlenecks across AI pipelines, and potentially accelerate troubleshooting with AI-powered root cause analysis.

This unified visibility helps organizations run AI workloads with the same reliability, efficiency, and operational confidence expected from modern enterprise systems.

This unified visibility helps organizations operate AI workloads with improved visibility and operational confidence. By integrating with NVIDIA AI–Q Blueprint and the NVIDIA Agent Toolkit, Dynatrace enriches agent reasoning with high-quality operational telemetry while at the same time helping teams govern and identify opportunities to optimize costs.

How Dynatrace addresses Agentic AI

Dynatrace is designed to assist your team with monitoring infrastructure usage and model behavior and detecting pipeline bottlenecks and token consumption while improving reliability by accelerating troubleshooting and root cause analysis. It also provides a unified view of AI workflows from agent to model down to the infrastructure, allowing organizations to support responsible AI operations, manage cost, improve performance and support agentic workflows at scale.

Every agentic deployment is customized with different agents, tools, models, and data pipelines; therefore, observability is an important capability for understanding how these systems behave in production. The complexity arises as agents interact with multiple enterprise data sources, including:

  • internal datasets
  • external web and knowledge repositories
  • proprietary research systems
  • models served through NVIDIA NIM and Nemotron

Dynatrace can serve as operational data source for AI agents that may help improve the quality of generated insights and enable more informed decision-making. With flexible integration across customized AI-Q implementations, this architecture also lays out the groundwork for automated analysis, research, and decision making.

How Dynatrace integrates NVIDIA AI-Q

By combining NVIDIA’s AI-Q Blueprint with Dynatrace AI observability, organizations gain the transparency and operational intelligence needed to govern, optimize, and scale complex AI systems.

Dynatrace integrates into AI-Q environments in two ways.

1. Observability and cost intelligence for Agentic AI workflows

The NVIDIA Agent Toolkit generates lightweight OpenTelemetry traces that Dynatrace ingests to visualize agent workflows and model interactions.

Dynatrace automatically maps the underlying infrastructure supporting AIQ deployments including NVIDIA NIM and Nemotron microservices and enriches telemetry with AI-specific signals such as:

  • token usage
  • inference latency
  • model metadata
  • GPU utilization

This provides comprehensive visibility across key components including:

  • AI models and inference workloads
  • agent orchestration pipelines
  • GPU and infrastructure resources
  • enterprise data interactions

With these insights, teams can quickly detect performance bottlenecks across agent pipelines, monitor GPU utilization and overall infrastructure health, and identify inefficient model usage. This visibility can help organizations identify cost optimization opportunities associated with AI workloads. Together, these capabilities position observability as important components for building reliable and scalable AI systems.

2. Dynatrace as a high-quality data source for AI agents

Dynatrace can also serve as an operational intelligence source for AI agents.

Through Model Context Protocol (MCP) integrations, Dynatrace exposes telemetry that agents can use in their reasoning workflows, including:

  • infrastructure performance metrics
  • operational incidents and problems
  • deployment and reliability trends
  • system behavior and resource consumption

This allows AI agents to incorporate real-time operational insights into their decision-making. Instead of relying solely on external data, agents gain contextual awareness of enterprise systems, which may support more informed outputs Dynatrace ingests NVIDIA Agent Toolkit OpenTelemetry traces, model telemetry, and infra metrics exposing operational context via MCP.

Together, these technologies create a powerful foundation for deploying deep research in the enterprise as reflected in the picture below.

Dynatrace AI Observability - NVIDIA
Figure 1: Dynatrace providing AI Observability for NVIDIA AI-Q

AI-Q use cases

The following are illustrative examples of what becomes possible when AI-Q-based research agents incorporate Dynatrace operational data and insights into their reasoning workflows. While NVIDIA AI-Q is a reference framework rather than a formal certified Dynatrace integration, these scenarios show how agentic research systems could use Dynatrace AI observability to generate richer analysis, identify patterns, and support more informed decisions.

Infrastructure migration analysis

AI agents combine Dynatrace operational telemetry such as performance trends, incidents, and deployment velocity with infrastructure and cloud cost data to evaluate platform migration scenarios (for example, OpenShift to AKS). The system produces data-driven recommendations with quantified tradeoffs to support strategic decisions.

Large-scale incident analysis

By analyzing thousands of historical problems, AI agents can identify recurring patterns, understand infrastructure behavior, and correlate technical issues with business KPIs. This enables deep operational insights and long-form analysis that would be difficult and time-consuming for humans to produce.

AI cost governance and optimization

Enterprises can use observability data from Dynatrace to analyze token consumption, model usage, and inefficient data interactions across AI workloads. Agents can identify patterns and suggest potential optimizations such as more efficient models or improved workflows.

Software delivery and reliability insights

DevOps and SRE teams can use agentic analysis to correlate deployments with incidents, assess build quality trends, forecast reliability risks, and identify engineering priorities—using Dynatrace as the trusted operational data source.

Get started today

The post Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/feed/ 0
Beyond LLM-as-a-judge: Establishing LLM evaluations as a foundation for trustworthy agentic AI systems https://www.dynatrace.com/news/blog/llm-evaluations-as-a-foundation-for-trustworthy-agentic-ai-systems/ https://www.dynatrace.com/news/blog/llm-evaluations-as-a-foundation-for-trustworthy-agentic-ai-systems/#respond Fri, 26 Jun 2026 15:54:20 +0000 https://www.dynatrace.com/news/?p=74677

Large language models and agents are rapidly transforming how organizations build software, automate workflows, and interact with data. From copilots to autonomous agents, AI-powered systems are increasingly responsible for answering questions, generating code, and supporting operational decisions. But as organizations move from experimentation to production, measuring performance reliably is no longer optional; this is where LLM evaluations become essential.

The post Beyond LLM-as-a-judge: Establishing LLM evaluations as a foundation for trustworthy agentic AI systems appeared first on Dynatrace news.

]]>

This is the second post in our series on LLM evaluations. In the companion post, Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals, we showed you how to run online evaluations against real GenAI prompt traces and bring quality scores into Dynatrace AI Observability alongside latency, cost, and errors. This post steps back to the fundamentals: what evaluations are, how they work, and the methods teams use to measure AI quality.

Just as traditional software relies on testing frameworks to ensure reliability, AI systems require robust evaluation frameworks to measure the quality, accuracy, and safety of model outputs. Evals are the primary mechanism by which teams build trust in, iterate on, and responsibly deploy AI systems. Without them, organizations may risk deploying systems that produce unreliable answers, hallucinate facts, or quietly degrade in performance over time.

Key takeaways

  • Evaluations are how teams move from “the LLM feels right” to “we can prove the LLM works.”
  • There is no single best evaluation method. The right approach depends on what you’re measuring and why.
  • LLM-as-a-Judge is one powerful tool within the broader evaluation ecosystem, not synonymous with evals as a whole.
  • Online and offline evaluations serve complementary roles: offline for development, online for production monitoring.
  • A mature evaluation strategy combines code-based, model-based, and human-based methods.
  • Evals should be treated as living artifacts — maintained, versioned, and improved over time like any other engineering asset.

Why LLM evaluation is fundamentally different from traditional testing

Traditional software produces deterministic outputs — the same input consistently returns the same result, making pass/fail testing straightforward. LLMs are probabilistic systems: the same prompt can produce different responses depending on context, temperature, and model behavior. This variability makes conventional testing methods insufficient.

Instead of verifying a single correct output, teams must evaluate across multiple dimensions simultaneously:

  • Correctness— does the response answer the question accurately?
  • Relevance — is the output aligned with the user’s intent?
  • Faithfulness — is the response grounded in source data, not invented?
  • Safety and bias — does the output comply with organizational policies?

This transforms evaluation from simple pass/fail checks into continuous measurement of AI quality.

Prompt stream with evaluation results shown in AI Observability app
Figure 1. Prompt stream with evaluation results shown in AI Observability app

The hallucination problem

The most well-known consequence of probabilistic generation is hallucination — when a model produces plausible-sounding but factually incorrect information. This happens because LLMs predict likely word sequences rather than verify facts, which enables powerful reasoning but introduces serious risk in enterprise environments where accuracy is critical.

Addressing this requires evaluation frameworks that track signals like factual accuracy, semantic similarity, groundedness in source data, and consistency across responses. These metrics transform subjective quality judgments into measurable, improvable signals.

What is an LLM evaluation?

An LLM evaluation is a systematic process of testing a model or AI-powered system to determine whether it meets a defined standard of quality. That standard could be factual accuracy, helpfulness, safety, tone, latency, cost-efficiency, or any other measurable dimension that matters to the application.

Evaluations translate vague product goals (“the assistant should be helpful and safe”) into concrete, repeatable measurements. They allow teams to:

  • Catch regressions when a model is updated, or a prompt is changed.
  • Compare candidates — different models, prompt versions, or retrieval strategies — objectively.
  • Build accountability by producing evidence that a system behaves as intended.
  • Accelerate iteration by giving developers fast, structured feedback loops.

Evals exist on a spectrum of formality, from a small hand-curated test set run locally, to a large, automated pipeline running thousands of test cases in CI/CD on every deployment.

AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability
Figure 2. AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability

How do LLM evaluations operate?

At their core, evaluations follow a consistent pattern regardless of their complexity:

  1. Define the task and success criteria. What should the LLM model do, and how will you know when it does it correctly? This is the hardest and most important step.
  2. Assemble a dataset. A set of inputs (prompts, user messages, documents) paired with expected outputs or grading rubrics. Datasets can be human-curated, synthetically generated, or sampled from production traffic.
  3. Run inference. Pass the inputs through the system under test and collect outputs.
  4. Score the outputs. Apply a scoring method — a function, a model, or a human — to assess how well each output meets the success criteria.
  5. Aggregate and analyze. Roll up scores into metrics (accuracy, pass rate, average score), visualize distributions, and compare against baselines or previous runs.
  6. Act on results. Use the findings to accept or reject a change, file a bug, update a prompt, or trigger retraining.

This loop can run manually during development, automatically in CI/CD pipelines, or continuously against live production traffic.

What’s the difference between LLM evaluations and LLM-as-a-Judge?

This is one of the most common points of confusion in the space.

LLM evaluations are the broader discipline — the full process described above. They encompass everything from how you define success to how you collect test data to how you score outputs to how you act on results.

LLM-as-a-Judge is one specific scoring method that can be used within an evaluation pipeline. It involves using a language model (often a strong general-purpose model like GPT-5 or Claude Sonnet 4.6) to automatically assess the quality of another model’s outputs.

Think of it this way: evaluations are the framework, and LLM-as-a-Judge is one type of grader you can plug into that framework — alongside code-based graders, human graders, or embedding-based similarity checks.

  • LLM-as-a-judge handles open-ended, subjective dimensions (tone, creativity, helpfulness) that are hard to capture in code.
  • It scales to large datasets without human effort.
  • It can be surprisingly well-calibrated when prompts and rubrics are carefully designed.

Limitations

  • Inherent biases of the LLM model used to judge (verbosity bias, position bias, self-preference).
  • Requires prompt engineering and validation to ensure the judge is grading what you intend.
  • Adds cost and latency to the evaluation pipeline.
  • Not appropriate for tasks with clear ground-truth answers where code-based checks suffice.

Code-based evaluations

Code-based evaluations use deterministic functions — written in Python or any language — to score model outputs. No secondary LLM model is involved.

How it works

You write a function that takes the model output as input and returns a score. The function might check for exact string matches, run regex patterns, execute generated code and test it, parse JSON and validate its structure, call an external API to verify a fact, or compare numerical results.

Common patterns

  • Exact match — does the output equal the expected answer?
  • Contains / regex match — does the output include a required phrase or follow a required format?
  • Execution-based — for code generation tasks, run the output and check whether tests pass.
  • Structured output validation — parse JSON/XML outputs and verify schema and values.
  • Tool call verification — for agentic tasks, did the model call the right tool with the right parameters?

Strengths

  • Fully deterministic and reproducible.
  • Fast and cheap to run at scale.
  • Easy to understand, debug, and audit.
  • No dependence on a secondary model’s judgment.

Limitations

  • Cannot handle open-ended or subjective quality dimensions.
  • Requires knowing the exact expected output or a verifiable property of the output.
  • Brittle for tasks where there are many valid correct outputs (for example, summarization, creative writing).

Code-based LLM evals are the first tool to reach for whenever a task has a clear, verifiable answer. They form the backbone of any reliable eval suite.

Online vs. offline evaluations

These two modes are not competing approaches — they’re complementary phases of a complete evaluation strategy.

Offline evaluations

Offline evals run against a static, pre-collected dataset before a system reaches production. They’re the evaluation equivalent of unit and integration tests in software development.

  • When: During development, before deploying a new model, prompt, or retrieval change.
  • Dataset: Curated, labeled, or synthetically generated. Often maintained in version control.
  • Latency: Can run in batch; speed is less critical.
  • Use cases: Regression testing, model comparison, prompt optimization, safety red-teaming, fine-tune evaluation.

Key advantage: Full control over the test distribution and ground-truth labels.

Key limitation: The dataset may not reflect real user behavior or the long tail of production inputs.

Online Evaluations

Online evals run against live production traffic in real time or near real time. They observe what is actually happening when real users interact with the system.

  • When: Continuously, in production.
  • Dataset: Real user inputs — unlabeled, unpredictable, and representative.
  • Latency: Must be fast or asynchronous to avoid slowing down user-facing requests.
  • Use cases: Production monitoring, anomaly detection, drift detection, A/B testing, continuous quality assurance.

Key advantage: Captures real-world usage patterns, prompts, and failure modes from production traffic, giving teams the most representative signal for monitoring AI quality over time.

Key limitation: No pre-defined labels; scoring must rely on heuristics, implicit signals (thumbs up/down, re-prompts), or async LLM-as-a-Judge pipelines.

Get started with LLM evaluations today

The field of LLM evals is evolving rapidly. As enterprises deploy increasingly autonomous AI systems, evaluation can play an important role in improving AI accuracy, reliability, and safety.

Here are the trends worth watching and investing in:

  1. Evaluation-driven development. Treat evals as a first-class engineering artifact. Write eval cases before building features, maintain them in version control, and integrate them into CI/CD pipelines — mirroring test-driven development practices from software engineering.
  2. Agentic and multi-step evaluation. As AI systems move from single-turn Q&A to multi-step agents that use tools and maintain state, evaluations must evolve to assess full trajectories rather than just individual outputs. This includes evaluating tool use, planning quality, error recovery, and task completion over long horizons.
  3. Adversarial and safety evals. Red-teaming — probing a system for failures, biases, and unsafe behaviors — is becoming a standard part of the eval lifecycle, especially as regulatory requirements around AI safety mature.
  4. Human-in-the-loop calibration. Even automated eval pipelines benefit from periodic human review to catch drift in what the judge model or scoring function is measuring. Building lightweight human-annotation workflows alongside automated evaluations yields a more reliable signal over time.
  5. Standardization and benchmarking. The industry is moving toward shared benchmarks and eval frameworks (for example, HELM, MMLU, LMSYS Chatbot Arena, OpenAI Evals) that allow apples-to-apples comparisons across models. Building internal evals that complement these public benchmarks will be an increasingly important capability for any team deploying LLMs.
  6. Cost-aware evaluation. As evals scale, cost becomes a real constraint. Emerging approaches include training lightweight specialized judge models, using embedding-based similarity as a cheap first filter, and intelligently sampling which examples need expensive LLM-as-a-Judge scoring.

Organizations that invest early in robust evaluation frameworks and combine them with AI observability will be positioned to scale AI safely across their operations.

Ready to put this into practice?

See our companion blog post, Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals, to learn how dt-evals lets you run LLM-as-a-judge evaluations on real GenAI traces and turn AI quality into a queryable, trendable, and alertable signal inside Dynatrace AI Observability.

Because in the end, AI systems are only as trustworthy as the processes used to evaluate them.

The post Beyond LLM-as-a-judge: Establishing LLM evaluations as a foundation for trustworthy agentic AI systems appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/llm-evaluations-as-a-foundation-for-trustworthy-agentic-ai-systems/feed/ 0
Log management for AI workloads: How to bring your logs and telemetry plan into the AI-first century https://www.dynatrace.com/news/blog/2026-log-management-for-ai-workloads-action-plan/ https://www.dynatrace.com/news/blog/2026-log-management-for-ai-workloads-action-plan/#respond Wed, 24 Jun 2026 12:19:44 +0000 https://www.dynatrace.com/news/?p=74644 2026 Logs Report action plan blog

AI is stretching the boundaries of traditional log management. More data without context slows insight, increases risk, and stalls AI progress. Teams need to rethink how they capture, process, and use telemetry from ingest to analysis. This action plan outlines how to unify telemetry, optimize pipelines, and turn data into real‑time, trusted intelligence teams can […]

The post Log management for AI workloads: How to bring your logs and telemetry plan into the AI-first century appeared first on Dynatrace news.

]]>
2026 Logs Report action plan blog

AI is stretching the boundaries of traditional log management. More data without context slows insight, increases risk, and stalls AI progress. Teams need to rethink how they capture, process, and use telemetry from ingest to analysis. This action plan outlines how to unify telemetry, optimize pipelines, and turn data into real‑time, trusted intelligence teams can use to scale AI operations with confidence.

AI workloads aren’t just increasing telemetry—they’re exposing the limits of how teams capture, store, and use it. Traditional log management was built for predictable systems and finite telemetry. AI systems break those assumptions. The result: more telemetry, less context, and increasing cost pressures to stay operational. The shift is how to manage logs better—and changing how teams capture, process, and make logs telemetry available across the entire lifecycle.

Here’s how that shift looks in practice:

Traditional log management AI-ready log management
Indexing and re-indexing, schema-first, archiving and rehydrating  Schema-on-read, always hydrated, always queryable 
Tool-specific context  Unified telemetry context 
Reactive troubleshooting  Preventive operations 

The State of Log Management 2026 research report—based on a global survey of 450 senior IT leaders—examines how AI is reshaping log economics, instrumentation, and observability strategies, and what shifts technical leaders must make to telemetry capture, storage, and management to support and scale agentic AI projects.

5 actions to kickstart your new log management plan

  • Create a single source of truth for AI systems by centralizing all telemetry in a unified, continuously queryable context layer and platform.
  • Establish causation across AI systems by automatically unifying logs with metrics, traces, and lifecycle context—not relying on logs alone.
  • Control costs without losing visibility by optimizing telemetry before ingest and eliminating indexing, archiving, and rehydration dependencies.
  • Standardize and govern telemetry at ingest to ensure data quality, compliance, and real-time usability at AI scale.
  • Enable preventive AI operations by turning contextual telemetry into real-time insight and automated remediation.

Why should teams unify telemetry on a single observability platform for AI workloads?

Unifying telemetry reduces manual correlation, preserves context, and keeps logs, metrics, and traces continuously queryable as telemetry scale increases.

AI workloads are exacerbating an existing problem by fragmenting even more telemetry across tools just as systems require more context. Teams now use an average of seven log tools, forcing manual correlation that doesn’t scale.

  • Unify all telemetry—logs, metrics, traces, security signals, user behavior, business events—into a single, continuously queryable context layer where telemetry is correlated automatically.
  • Enrich telemetry at ingest starting at the edge with shared technical and business context to explain system behavior as dependencies multiply.
  • Democratize access using intuitive querying so more teams can validate AI behavior and act faster with confidence.

How do logs and traces work together for reliable and explainable autonomous operations?

Logs don’t explain AI behavior independently. Understanding comes from unifying logs with traces and other telemetry signals optimized throughout the telemetry lifecycle.

Autonomous systems demand deterministic signals that explain what happened, why it happened, and how to respond—something logs alone can’t fully provide.

  • Instrument logs to capture AI‑specific details at every inference layer to preserve the exact sequence of events.
  • Correlate logs with traces automatically to establish causation and pinpoint root causes.
  • Automate remediation using continuously enriched telemetry to enable reliable, explainable autonomous operations at scale.

How can teams optimize log management costs without sacrificing insight?

Teams can manage costs by retaining high‑value telemetry without rigid schemas, indexing overhead, or rehydration delays that limit analysis.

Managing log costs involves data strategy, not just storage. Logs consume nearly half of observability budgets, yet even after reducing volume by filtering, masking, and aggregating, 50% of organizations don’t collect or discard an average of 86% of logs specifically to manage costs, and 74% say indexing and rehydration costs are barriers to value.

  • Ingest and retain telemetry without rigid schemas or indexes, eliminating the need to predict questions in advance.
  • Store exabytes of data in one queryable layer, avoiding cold archives and rehydration costs and delays.
  • Analyze telemetry in full context to reduce waste and maximize business value from AI‑generated data.

What changes to instrumentation and ingest should teams make to support AI workloads?

Teams must optimize telemetry before ingest—standardizing instrumentation and automating parsing and configurations—so data remains high-quality, contextual, and continuously queryable at AI scale.

Fragmented instrumentation and brittle ingest pipelines slow insight and delay AI projects from reaching production. 85% of organizations struggle to ingest logs at AI scale, and 80% say turning telemetry into insight delays AI initiatives.

  • Standardize instrumentation across logs, traces, metrics, and other telemetry signals to maintain context and reduce downstream correlation.
  • Streamline ingestion in real time by automating parsing, configurations, and enrichment to retain only high‑value, compliant data.
  • Sustain telemetry at scale with an always‑queryable data layer that supports real‑time analytics and automation.

Why are preventive operations critical to AI-native environments?

As AI workloads increase telemetry volume and autonomous operations, teams need detailed intelligence about what’s happening in AI output to predictively detect early signals of unexpected results.

Because reactive troubleshooting can’t keep up with autonomous systems, teams need real‑time, contextual telemetry to detect drift and prevent failures early. 84% say customer trust in AI depends on their ability to use log analytics to predict and prevent problems.

  • Correlate logs with end‑to‑end traces automatically to create a reliable understanding of AI behavior before failures escalate.
  • Analyze AI telemetry in real time and full context to detect early signs of drift or degradation.
  • Automate response to reduce risk and scale AI‑driven operations safely.

Upleveling log management to advance trustworthy agentic AI

Expanding AI workloads demand more from log management—an approach built on unified observability, open, optimized ingest at massive scale, and real‑time analytics without rigid schemas, indexing overhead, or rehydration delays. Logs remain the accountability anchor, but trust emerges only when all telemetry signals come together in context.

Get the State of Log Management 2026 report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI operations.

FAQ: Log management action plan for AI workloads

Why do AI workloads require a different log management approach?

AI workloads are variable and generate significantly more telemetry, which demands explainability, reliability, and cost control capabilities that traditional log architectures weren’t designed to support.

What is the first step teams should take to modernize log management for AI?

Unify logs, metrics, and traces on a single observability platform so telemetry is always available in context and doesn’t require manual correlation.

How can organizations reduce log management costs without losing insight?

By starting observability at the edge and retaining high‑value telemetry without rigid schemas, indexes, or rehydration delays, teams avoid discarding data while controlling cost.

Why aren’t logs alone enough to support autonomous operations?

Logs show what happened and why, but traces show how and what’s affected; together they explain AI behavior and enable reliable, automated remediation.

What enables preventive operations in AI‑native environments?

Real‑time analysis of contextual telemetry that detects early signs of drift or degradation and triggers automated guardrails before failures escalate.

The post Log management for AI workloads: How to bring your logs and telemetry plan into the AI-first century appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/2026-log-management-for-ai-workloads-action-plan/feed/ 0
How AI workloads are changing what logs must deliver, forcing a new strategy https://www.dynatrace.com/news/blog/log-management-research-key-findings/ https://www.dynatrace.com/news/blog/log-management-research-key-findings/#respond Wed, 17 Jun 2026 12:34:03 +0000 https://www.dynatrace.com/news/?p=73713 Future of log management: 2026 research shows AI demands of logs

AI workloads are redefining what logs must deliver—and exposing where traditional approaches fall short. New research reveals how rising scale, cost pressures, and missing context are reshaping log management. The path forward is clear: unify logs with traces and other telemetry in context to turn fragmented signals into reliable insight and build the foundation for trusted, scalable AI operations.

The post How AI workloads are changing what logs must deliver, forcing a new strategy appeared first on Dynatrace news.

]]>
Future of log management: 2026 research shows AI demands of logs

New Dynatrace research reveals how logs are becoming the accountability anchor for AI systems and why cost-driven trade-offs (sampling, cold storage, and discarding data) increase operational risk. In fact, with legacy log management tools now consuming 45% of observability budgets, 67% say the costs of these tools outweigh their value, indicating a need to modernize quickly (source: Dynatrace State of Log Management 2026 research).

3 takeaways for engineering and business leaders

  • Scale is accelerating. Log and telemetry volume increased 93% on average with 1 in 5 organizations seeing growth above 150%.
  • Costs are breaking traditional models. Existing log management tools consume 45% of observability budgets with respondents estimating an average annual spend near $2.5M.
  • Context with traces is the path to AI trust. 73% say logs reveal only part of what’s happening with AI workloads, and 70% rank traces as a top source for evaluating AI performance and behavior. But 65% rely heavily on logs because they don’t have easy visibility into other telemetry signals.

Logs capture the precise details of events from cloud, AI, and infrastructure. As AI workloads and agentic systems (AI systems that act autonomously) make up an increasing proportion of technology stacks, logs serve as a crucial shared language for humans and agents to reason and troubleshoot system state and autonomously build and optimize infrastructure and applications.

Accountability depends on placing high-fidelity log telemetry in context: connecting it with traces, metrics, security events, user behavior, and business signals to understand AI behavior and guide remediation at scale.

How is AI breaking the economics of traditional log management?

AI workloads break the financials of traditional logging approaches by driving massive telemetry growth that forces many teams to not collect or discard data to avoid runaway costs.

Over the past year, AI workloads triggered a 93% average increase in log and telemetry volume, with one in five experiencing expansions above 150%. At the same time, teams rely on an average of seven different log and telemetry tools, forcing manual correlation that doesn’t scale.

The financial impact is just as stark. According to the report, existing log management tools now consume 45% of observability budgets, with average annual spend among survey respondents nearing an estimated $2.5M per organization. To contain costs, many teams limit ingestion, sample telemetry, or push logs into cold storage, losing crucial context and increasing security, compliance, and operational risk. In fact, 67% say the cost of existing log management tools now outweighs their value. And even after filtering, masking, and aggregating to reduce volume, the limitations of traditional logging tools force teams into imprecise compromises, resulting in half of organizations not collecting or discarding 86% of logs specifically to manage costs.

Traditional vs. AI-native log management

What changes

Traditional log management
(cost-first)

AI-native log management
(unified observability)

How teams find answers  Manual stitching slows analysis; teams spend 58% of analysis time correlating telemetry  Automated ingestion, enrichment, and correlation reduces manual work and accelerates time to answers 
AI trust and validation  Logs alone are incomplete; 73% say logs reveal only part of what’s happening in AI workloads  Context builds trust: logs + traces (top-ranked by 70%) + metrics + events show behavior and causality 
Readiness for what’s next  79% worry current ingest/storage won’t meet future needs; instrumentation lags AI requirements  Updated instrumentation starting at the edge and open, automated processing at scale (supported by 81%) for accelerated AI innovation 
Data strategy  Many teams control spend by reducing ingest using sampling, cold storage, or discarding data  Retain high-fidelity telemetry at scale without rehydration/indexing friction 
Tooling approach  Fragmented tooling; teams use an average of seven tools, forcing manual correlation  One real-time observability context layer that unifies logs with metrics, traces, security events, and business signals 
Operational impact  Blind spots and risks grow as half of teams don’t collect or discard 86% of logs  Answers, not guesses; more context preserved for faster diagnosis and safer automation 

What must change in instrumentation and ingestion for autonomous systems?

To build trust in AI workloads and advance the business value of autonomous decision-making, optimizing and streamlining telemetry must start before ingest and be open and automated at a massive scale.

Traditional logging architectures weren’t built for AI‑driven scale or autonomy.

  • As AI workloads proliferate, 79% of technical leaders worry their current ingest and storage approaches won’t meet future needs.
  • 80% say they must update instrumentation to support new AI‑specific metrics.

The operational toll is significant. Teams spend 58% of their analysis time stitching together logs, metrics, and traces before extracting insight, which slows decisions and delays AI projects from moving into production.

To break this bottleneck, 81% of organizations say log ingestion and processing must be open and automated at massive scale, enabling real‑time analysis without rigid schemas, indexing overhead, or rehydration delays.

Why do logs need an observability ecosystem to build AI trust?

While logs are a crucial component of AI observability, they need context to tell the whole story.

  • 73% of organizations say logs reveal only part of what’s happening in AI workloads
  • Teams spend 58% of analysis time stitching telemetry together
  • 72% say standalone log management tools are obsolete—AI workloads demand a platform approach that combines all types of telemetry in one place to accelerate time to answers

As a result, most organizations rely on logs alongside other telemetry signals—especially traces, which 70% rank as the top source for evaluating AI performance and behavior.

Trust in AI develops when analysis is based on high-fidelity telemetry in context:

  • Logs provide the fact basis
  • Traces expose flow and causality
  • Metrics quantify performance
  • Security events surface risk
  • Business signals connect system behavior to outcomes

Unified observability turns this telemetry into a coherent narrative—enabling teams to validate AI behavior, assess impact radius, guide remediation, and move from reactive troubleshooting to preventive operations.

What does AI-native log management look like with unified observability?

Unified observability transforms log management for the AI era by optimizing telemetry instrumentation and ingest, unifying telemetry with context, and eliminating crippling cost-cutting measures.

With exponential increase in telemetry volumes due to AI workloads, the path forward can’t just be shrinking log ingest to manage costs. AI innovation depends on leaning into telemetry volume with the right strategy and capabilities. The report points to key actions leaders can take to adopt AI-native log management practices.

  • Centralize all telemetry (logs, traces, metrics, events) in a unified, continuously queryable context layer to eliminate silos and scale AI visibility.
  • Automatically correlate logs with traces and lifecycle context to establish causation, understand AI behavior end to end, and enable reliable autonomous operations.
  • Control log costs before ingest while retaining full-fidelity telemetry in exabyte-capacity storage with no rigid schemas, indexes, cold archives, or rehydration.
  • Standardize instrumentation and optimize ingestion to capture high-value, governed telemetry that tracks agent actions, reasoning traces, and lifecycle events.
  • Enable preventive operations by detecting early signals and automating remediation to reduce risk, strengthen reliability, and safely scale AI projects.

Fragmented vs Unified log management

The goal is reliable operations and AI accountability at scale. Together, logs, traces, metrics, and events in context can enable preventive operations, reliable autonomy, and confident decision‑making as AI systems move from pilots into production. Organizations that build this unified foundation will be best positioned to scale AI without sacrificing trust.

Download the State of Log Management 2026 report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

FAQ: State of Log management 2026

Why is log management getting harder in the AI era?

Because AI workloads are driving rapid telemetry growth—organizations saw a 93% average increase in log and telemetry volume in the past year.

Why do log management costs feel out of control?

Existing log management tools consume 45% of observability budgets, and average annual spend is nearing $2.5M per organization.

Do teams still see value in log management at today’s price?

Not consistently. 67% of respondents say costs of existing log management tools now outweigh their value.

Why do teams discard so many logs?

Cost pressure. Using existing tools, 50% of organizations don’t collect or discard 86% of their logs on average, often using sampling or limiting ingestion specifically to reduce spend.

What’s the biggest operational bottleneck with today’s tooling?

Manual correlation. Teams spend 58% of analysis time stitching together logs, metrics, and traces before they can extract insight.

Why aren’t standalone log tools enough for AI workloads?

Because logs alone rarely tell the whole story. 72% say standalone log management tools are obsolete, and 73% say logs reveal only part of what’s happening in AI workloads.

What telemetry signal do teams rely on most to evaluate AI behavior?

Traces—70% rank them as the top source for evaluating AI performance and behavior beyond logs alone.

The post How AI workloads are changing what logs must deliver, forcing a new strategy appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/log-management-research-key-findings/feed/ 0
Dynatrace Release Radar 05.26 https://www.dynatrace.com/news/blog/dynatrace-release-radar-05-26/ https://www.dynatrace.com/news/blog/dynatrace-release-radar-05-26/#respond Tue, 16 Jun 2026 18:49:01 +0000 https://www.dynatrace.com/news/?p=74546 Release Radar

This series covers recent Dynatrace releases and updates, focusing on what’s new, what’s changed, and how these recent enhancements can benefit you and your organization. Each post covers newly available capabilities and points you toward where to explore them.

The post Dynatrace Release Radar 05.26 appeared first on Dynatrace news.

]]>
Release Radar

This edition of Release Radar covers the Dynatrace releases from May, 2026. Here are the changes that should matter right away to practitioners:

  • AI tooling
  • AI coding agent monitoring
  • Jira integration
  • Dashboard productivity
  • Pipeline grouping

To see them in action, head over to our release radar launchpad on the Dynatrace Playground.

Dynatrace Assist: At your side with more context

Dynatrace Assist gets four upgrades in sprint 1.338 that deepen its usefulness during active investigations.

Side-by-side mode puts the chat interface in a collapsible panel alongside your current view — dashboard, notebook, or any other app page stays visible while you work with Assist. Chat and investigate at the same time without losing your place.

Reference files and skills give Assist access to a curated knowledge base built on Dynatrace documentation and product expertise, following the Anthropic Claude Agent Skills format. Responses are designed to draw on structured Dynatrace knowledge, so answers are grounded in how the platform actually works.

Anthropic Claude Sonnet 4.6 as the foundation model can help bring stronger multi-step reasoning and improved performance on complex, multi-tool investigations — the same model used in the latest Anthropic API and Claude Code.

A purpose-built NL2DQL model makes natural-language-to-DQL generation more accurate. The capability now uses a fine-tuned foundation model based on Llama 3.1 8B, trained specifically on Dynatrace query patterns. Write a question in plain language and get a working DQL query with fewer iterations.

Figure 1. Dynatrace Assist now works side by side with your current view, with stronger reasoning, grounded reference skills, and more accurate natural-language-to-DQL generation.
Dynatrace Assist now works side by side with your current view, with stronger reasoning, grounded reference skills, and more accurate natural-language-to-DQL generation.

AI coding agents get unified monitoring

Dynatrace now provides observability for five major AI coding agents: Claude Code, Google Gemini CLI, OpenAI Codex CLI, OpenCode, and GitHub Copilot SDK.

As  your team adopts multiple AI agents in parallel, you need shared visibility into what they cost, how they behave, and what they produce. This release gives platform teams, engineering leaders, and security teams a  unified observability across all five agents, all built on OpenTelemetry.

What teams get across the supported agents:

  • Adoption and token tracking — session counts, token consumption, and cost trends across agents and teams.
  • Tool behavior visibility — which tools each agent calls, how often, and where runs slow down or fail.
  • Production context in the IDE — engineers can query live Dynatrace data through the Dynatrace MCP Server without leaving their coding environment.
  • Engineering outcome correlation — connect agent activity to downstream delivery signals like commits and pull requests (available for Claude Code).

Pre-configured dashboards are available for each agent. For the full breakdown by agent and setup details, see Dynatrace expands AI coding agent monitoring.

Figure 2. Dynatrace brings unified observability to five major AI coding agents, helping teams track adoption, token usage, tool behavior, and cost across environments.
Dynatrace brings unified observability to five major AI coding agents, helping teams track adoption, token usage, tool behavior, and cost across environments.

Investigate production problems without leaving Jira

The Dynatrace MCP Server now integrates with Atlassian Rovo, bringing observability context directly into Jira and JSM tickets.

When an incident or issue is open in Jira or JSM, Rovo can now call Dynatrace tools in natural language — querying metrics, traces, logs, and topology — and post the results as ticket comments. Root cause analysis and dependency mapping happen inside the ticket, so engineers stay in context instead of switching between platforms.

Key points for practitioners:

  • No context switching — investigate and document findings without leaving Jira or JSM.
  • Per-user OAuth 2.1 — every Dynatrace call runs as the requesting user, with a full audit trail across both platforms.
  • Admin-controlled tool exposure — administrators choose which Dynatrace tools Rovo can access.
  • Included with Dynatrace SaaS — no additional cost, and adding users doesn’t change the pricing.

It’s designed for fast setup: authenticate via the Rovo admin UI, select the tools to expose, and the integration is live across Jira, JSM, and Confluence.

For the full walkthrough, see Dynatrace MCP Server for Atlassian Rovo.

Figure 3. With the Dynatrace MCP Server for Atlassian Rovo, teams can investigate incidents and add observability findings directly inside Jira and JSM tickets. (Video)
With the Dynatrace MCP Server for Atlassian Rovo, teams can investigate incidents and add observability findings directly inside Jira and JSM tickets. (Video)

Dashboards: build faster, navigate better

Releases 1.338 and 1.339 bring a focused set of dashboard improvements that add up across a day of analysis work.

Ready-made tiles and sections expand the dashboard and notebook library with pre-configured visualizations and built-in drill-downs to other Dynatrace apps. Browse or search the tile library to find components that are ready to use, and customize from there rather than starting from a blank canvas.

Direct JSON editing lets power users open and edit the full dashboard definition as JSON from the dashboard Actions menu. The format matches the Dashboard API, so configuration changes, bulk tile edits, and version-controlled workflows are all faster in the editor than in the visual UI.

URL-driven variables make it possible to configure hidden dashboard variables through URL parameters, enabling pre-configured views to be linked directly with filters already applied — useful for sharing context-specific dashboards with specific teams or stakeholders.

Launcher link reordering lets users drag links between sections in the launcher, making personal navigation layouts easy to maintain.

Active tile tab persistence keeps the last-active editing tab visible when switching between tiles, so the configuration state is preserved while navigating across a dashboard.

Figure 4. New dashboard enhancements — including ready-made tiles, direct JSON editing, and URL-driven variables — make it faster to build, refine, and share analysis views.
New dashboard enhancements — including ready-made tiles, direct JSON editing, and URL-driven variables — make it faster to build, refine, and share analysis views.

Pipeline groups get a configuration UI

Pipeline groups, which let central teams enforce shared policies across multiple OpenPipeline pipelines, now have a dedicated configuration interface in Early Access (sprint 1.339).

Platform teams can now configure group-level policies in the UI, including cost allocation and sensitive data scanning. Pipeline teams still control their own parsing and extraction logic, so central governance doesn’t come at the cost of local flexibility. No direct API access required.

This UI makes the feature accessible to a wider set of platform operators and can help reduce setup costs for organizations running large, multi-team pipeline environments.

For background on pipeline groups and the governance model they enable, see Pipeline Groups in Dynatrace OpenPipeline.

Figure 5. Pipeline groups now include a dedicated configuration interface in Early Access, making shared governance policies easier to manage across OpenPipeline pipelines.
Pipeline groups now include a dedicated configuration interface in Early Access, making shared governance policies easier to manage across OpenPipeline pipelines.

Latest UX improvements

Investigation workflows in logs get sharper. Join the Log pattern analysis preview and see how selected log patterns now open a dedicated deep-dive panel showing behavior over time and associated records — inspect pattern-level context without losing the overview, and drill into individual log records from the same panel without navigating away. Log attributes open in full-screen mode for reading long or nested JSON in place. When Logs is opened from a contextual link in a dashboard or alert, the query runs automatically — no extra click required.

Navigation, filtering, and performance also improve across several surfaces. In the Session List, a tooltip on the Session Replay icon shows replay availability — full, partial, or none — so you can triage which sessions have usable replay before opening them. The Services explorer gains dynamic tag filters for Kubernetes namespace annotations and labels, expanding filter coverage for k8s-heavy environments. Infrastructure inventory in large network environments now loads in 3–5 seconds in typical environments, with status and reachability columns loading progressively so the view is immediately usable at scale. Press Shift + ? anywhere in the platform to open the keyboard shortcut reference.

Recent UX updates improve investigation workflows across logs, session replay, services, and infrastructure, helping teams move faster with less context switching.
Recent UX updates improve investigation workflows across logs, session replay, services, and infrastructure, helping teams move faster with less context switching.

Why these changes matter

The May releases extend capabilities across the workflows that practitioners use most. AI tooling stays present during investigations. Coding agents become observable assets, not black boxes. Dashboards get faster to build and easier to manage. Cost data supports annual planning. And pipeline governance reaches more teams through a UI.

These are the kinds of changes that add up across a week of real work.

Check out the updates in action on our Release Radar Launchpad.

The post Dynatrace Release Radar 05.26 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-release-radar-05-26/feed/ 0
Orchestrate multicloud AI agents for autonomous incident resolution https://www.dynatrace.com/news/blog/orchestrate-multicloud-ai-agents-for-autonomous-incident-resolution/ https://www.dynatrace.com/news/blog/orchestrate-multicloud-ai-agents-for-autonomous-incident-resolution/#respond Mon, 15 Jun 2026 20:11:14 +0000 https://www.dynatrace.com/news/?p=74557 Observability data

Cloud SRE Agents is a Dynatrace app that orchestrates AWS®, Azure®, and Google® AI agents for automated investigation and resolution assistance for incidents across multicloud environments. Cloud SRE Agents routes identified issues based on configurable rules, centralizes its findings, and provides a single audit trail for autonomous operations.

The post Orchestrate multicloud AI agents for autonomous incident resolution appeared first on Dynatrace news.

]]>
Observability data

Organizations are evolving from human-driven operations to supervised autonomous operations, where AI investigates, recommends, and remediates, and humans stay in control of what matters most. A big part of delivering on that vision is working with the agents that customers already run in their cloud environments.

Harness the power of hyperscale agents

Each hyperscaler has AI agents that automatically investigate and help resolve production incidents using native cloud telemetry and tools. They act like embedded site reliability engineers, analyzing issues and recommending or executing remediation steps without waiting for a human to start the process.

AWS DevOps Agent provides investigation and remediation in AWS using native tooling. An Azure SRE Agent specializes in investigating and remediating Azure issues. And Google Gemini Cloud Assist is for incident analysis across Google Cloud Platform (GCP).

Over the past year, we’ve published how Dynatrace supercharges each of these cloud agents individually. When an issue occurs, Dynatrace Intelligence combines causal, predictive, and agentic AI using the Smartscape dependency graph to automatically link related symptoms and root causes across the environment into one unified problem card.

When Dynatrace integrates with the AWS DevOps Agent, dependency-aware root cause analysis combines with AWS frontier-agent capabilities, and joint customers report up to 70% reductions in mean time to resolution. When Azure SRE Agent connects with Dynatrace, deterministic, causation-based AI flows directly into Azure-native remediation workflows, cutting the back-and-forth between teams. And with Google Gemini Cloud Assist, Dynatrace delivers the same production context layer to GCP-hosted incidents: precise root cause, full topology, real business impact.

Problem detected by Dynatrace Intelligence, investigated and remediated by AWS DevOps Agent (see documentation in the right-hand panel)
Figure 1. Problem detected by Dynatrace Intelligence, investigated and remediated by AWS DevOps Agent (see documentation in the right-hand panel)

From integrations to intelligent orchestration

Many enterprises run workloads across AWS, Azure, and Google Cloud simultaneously, and managing three separate integrations with separate routing logic and separate cost controls is its own operational tax. Cloud SRE Agents provides a single orchestration layer that routes problems to specific hyperscaler agents based on configurable profiles to see everything happening across all three cloud agents.

The Cloud SRE Agents app writes findings back to Dynatrace, and provides your team with measurable visibility into autonomous actions.

The Overview tab's interactive graph shows a live view of problems and their activity status, grouped by related SRE agent.
Figure 2. The Overview tab’s interactive graph shows a live view of problems and their activity status, grouped by related SRE agent.

How Cloud SRE Agents works

When Dynatrace Intelligence detects a problem and identifies the root cause, Cloud SRE Agents calls dedicated cloud-native agents from AWS, Azure, and Google Cloud to retrieve deeper insights from the sources that only they can reach: CloudTrail history, Azure subscription policy, GCP project IAM, recent deployments, and native runbooks. These agents run in parallel, gathering evidence as soon as the problem is detected. Their findings, and, where applicable, the recommended remediation path, are displayed in the same Dynatrace problem view that the on-call SRE is already using in their day-to-day workflow.

One view. No tab-switching. The work starts without you.

Three workflows do the orchestration in the background:

  • Investigate evaluates your Interaction Profiles and dispatches matching problems to the right agents in parallel.
  • Periodic Tasks polls each cloud provider for completion, detects stalled or timed-out investigations, and writes findings back as problem annotations.
  • Event Handlers normalize the cloud-provider event stream so every action correlates back to its originating problem, end to end.

Cloud SRE Agents has the insights and intelligence to decide which agent gets which problem, tracks each run to completion, and brings the answers back together in a single view. The Overview tab provides a real-time, interactive network graph of problems, agents, and activities. The replay view allows the user to step back in time and get an overview of what has happened when, as well as the status of each investigation.

Replay functionality in the Cloud SRE Agents Overview
Figure 3. Replay functionality in the Cloud SRE Agents Overview

Intelligent routing with Interaction Profiles

In agentic operations, routing rules make the difference between turning autonomous systems loose on every alert and pointing them precisely where they earn their keep. Interaction Profiles are how you express routing judgment in Cloud SRE Agents. Each profile pairs a set of conditions with the agent or agents that should handle the problems flagged by the profile, and evaluates the conditions whenever Dynatrace Intelligence detects a problem.

The conditions you can write are deliberately broad. You can route by the cloud account, subscription, or project an incident touches; by problem category (availability, error, slowdown, resource contention); by affected entity type (a Kubernetes cluster, a database, a Lambda function); by tag, label, or any custom attribute carried in the problem record. Conditions combine with AND/OR logic and nest as deeply as you need, keeping real production routing policy inside the app rather than spilling into custom workflows or scripts.

Three ways teams put it to work

Route problems to the right cloud, automatically

A spike in Lambda error rates belongs to AWS DevOps Agent. An Azure App Service degradation calls for Azure SRE Agent. A Pub/Sub latency issue lands with Gemini Cloud Assist. In a multicloud estate, none of those decisions should fall to a human at 2:00 AM. A profile filtered by AWS Account ID, Azure Subscription ID, or GCP Project ID, then narrowed by resource type or tag, settles the routing question once. Every matching problem is automatically routed to the right specialist with the right cloud-native context.

Optimize spend with budget-aware routing

Cloud AI agents do work, and that work has a cost. Cloud SRE Agents lets you set a Monthly Duration Budget per agent and gate dispatch on it via a Has Available Budget filter: once the budget is exhausted, new investigations either stop (in strict enforcement mode) or proceed with a logged warning. The duration figure itself is a proxy, derived from Dynatrace event timestamps rather than the cloud provider’s clock, which makes it useful as a circuit breaker and directional signal, not a substitute for AWS, Azure, or GCP usage reports. The governance value is what matters: you decide how much autonomous investigation you’re willing to underwrite each month, and the system holds the line.

Tier autonomous investigation by problem type and entity

Not every Dynatrace problem warrants an autonomous investigation. Problem Category filters let you dispatch agents only to the problem categories that warrant it, for example, availability or error problems that require immediate action, rather than slowdowns or custom alerts where human triage might still be the right call. Layer on Entity Type filters, and you can further focus on specific infrastructure tiers (hosts, services, process groups, Kubernetes clusters). The result is a tiered model: high-severity issues receive immediate autonomous investigation, lower-severity signals queue for human review, and your team controls the threshold.

Governance that makes autonomous work measurable

Agentic operations earn trust when teams can see what the agents did, why, and whether it worked. Cloud SRE Agents treats that as a first-class concern, with two views built for the two audiences who care about it.

The Activity tab is the audit trail. Every investigation and mitigation appears as a card on a unified timeline; expand any card to see the agent’s full findings, the evidence it pulled, and the action it took or recommended. Each response can be rated Good, OK, or Bad, building a quality signal grounded in what your team actually saw rather than what the system predicted. When a single problem triggers work across multiple agents, those activities roll up to a single status (in progress, done, or stalled), so you always know where things stand without having to reconstruct the run from individual records.

Activity tab showing an expanded investigation card with agent findings and rating control.
Figure 4. Activity tab showing an expanded investigation card with agent findings and rating control.

The Statistics tab is where autonomous operations become a number you can show to a leadership team: problems handled, mitigations executed, average investigation time, MTTR and MTTI trends, success rates, and satisfaction scores broken down by agent. The same view doubles as a directional cost lens, since agent working time is the dominant driver on the cloud side of the bill. Treat the number as a trend signal and a circuit-breaker input, not a billing record (reconcile against AWS, Azure, and GCP usage reports for exact spend), and it makes the case for expanding agentic coverage with evidence rather than anecdote.

The Statistics tab shows key metrics and per-agent insights across a selected time range.
Figure 5. The Statistics tab shows key metrics and per-agent insights across a selected time range.

Why production context multiplies the value

What changes Cloud SRE Agents from a smart dispatcher into something more is what Dynatrace Intelligence contributes before an agent ever begins its analysis. Dynatrace delivers deterministic, causation-based root cause analysis grounded in Dynatrace’s Smartscape real-time dependency mapping, alongside business impact assessment and correlated telemetry. That context shapes the entire direction of the investigation. A cloud agent arriving with that foundation starts from “this specific service on this specific host is the root cause, and here’s the customer impact” rather than “something is wrong somewhere in this account.”

The numbers reflect it. According to AWS, organizations using the AWS DevOps Agent with Dynatrace see up to a 75% reduction in mean time to resolution.

Western Governors University, which runs a fully online learning environment for 200,000 students, uses AWS DevOps Agent with Dynatrace to automate cross-system correlation that previously required manual effort across multiple tools. At a larger scale, United Airlines transports more than 500,000 passengers daily across a hybrid environment that includes more than 500 AWS accounts, 20,000 Lambda functions, and 38,000 OneAgent deployments.

The team’s description of the before and after status is direct: previously, multiple tools with overlapping functions created gaps and black boxes during troubleshooting. With AWS DevOps Agent and Dynatrace, Dynatrace identifies the responsible layer, the agent investigates and provides resolution steps, and everything surfaces in a single Dynatrace view. No 3:00 AM tool-switching required.

Get started

For a closer look at the individual integrations, read the posts on AWS DevOps Agent and Dynatrace and Azure SRE Agent and Dynatrace, or see how Dynatrace Intelligence powers autonomous operations. To put your cloud agents to work today, install Cloud SRE Agents from the Dynatrace Hub. Cloud SRE Agents is currently available as a community-supported app.

Harness the power of your hyperscaler agents

The post Orchestrate multicloud AI agents for autonomous incident resolution appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/orchestrate-multicloud-ai-agents-for-autonomous-incident-resolution/feed/ 0
Dynatrace observability is now a Kiro power https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/ https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/#respond Fri, 12 Jun 2026 21:12:32 +0000 https://www.dynatrace.com/news/?p=74536

In this blog, we'll introduce the Kiro power for Dynatrace, show what it unlocks for developers, and walk you through how to get it up and running.

The post Dynatrace observability is now a Kiro power appeared first on Dynatrace news.

]]>

What is the Kiro power for Dynatrace?

The Kiro power for Dynatrace delivers live observability data, root cause analysis, and remediation suggestions directly into the Kiro IDE, with no JSON editing or manual MCP setup.

Kiro is an AI-powered IDE that helps developers move from idea to working code through spec-driven development and an agentic assistant. To make the assistant genuinely useful in unfamiliar domains, Kiro recently introduced powers: curated, partner-validated bundles of MCP servers, steering files, and best practices that install with a single click and load on demand when a relevant task comes up. Install a power, and Kiro’s agent gains specialized expertise the moment you need it.

For Dynatrace customers already working in Kiro, it’s the shortest path yet from code to production insight. For developers new to Dynatrace, it’s a one-click way to ground Kiro’s reasoning in real facts from your environment, not guesses.

Why this matters for developers

Developers have historically been one step removed from production. When something breaks after deployment, the path to figuring out what went wrong usually runs through a Site Reliability Engineering (SRE) or operations team, and AI coding assistants can’t automatically and reliably remediate issues in software they’re unfamiliar with. Agents that can write code are guessing about how their code behaves in production unless they have access to real telemetry data.

The Dynatrace Kiro power for Dynatrace closes this gap through Dynatrace Intelligence, the agentic operations system at the core of the Dynatrace platform. Kiro’s answers are grounded in deterministic, causal AI and real-time production data, not probabilistic guesses.

When a developer starts a task by writing a prompt, Kiro evaluates the conversation, identifies the relevant power using keywords, and dynamically activates power. Kiro then loads Dynatrace MCP tools and power instructions, providing skills to investigate problems, query live observability data, surface root causes, and even execute and verify remediations.
Figure 1. When a developer starts a task by writing a prompt, Kiro evaluates the conversation, identifies the relevant power using keywords, and dynamically activates the power. Kiro then loads Dynatrace MCP tools and the power instructions, providing the skills needed to investigate problems, query live observability data, surface root causes, and even execute and verify remediations.

With the tools provided by the Kiro power, developers can:

  • Investigate live incidents and get root cause analysis directly in Kiro chat
  • Query metrics, logs, and traces from production using natural language
  • Surface security vulnerabilities affecting the code they’re working on
  • Get remediation suggestions grounded in what’s actually happening in their environment

“Using Kiro powers for Dynatrace has been a total game-changer in the observability space. Deep-dive root cause analysis of complex system issues that once required lengthy manual intervention now happens in seconds, giving us unprecedented speed and confidence.”

Mike Kobush, Sr. Software Performance Engineer, NAIC

How to install the Kiro power for Dynatrace

Getting started takes only a few steps. Once installed, the Kiro power activates automatically when Kiro detects a relevant task. Mention an incident, a slow service, or anything that needs production context, and the Dynatrace tools and guidance will load in Kiro chat.

Prerequisites

  • A Dynatrace account. If you don’t already have one, you can start a free 15-day trial.
  • Kiro installed on your system.

Prepare the Dynatrace connection

First, create a Dynatrace Platform Token, which Kiro will use to authenticate. Then add the required permissions for the Dynatrace MCP server.

Install the Kiro power

The power can be installed from either the Kiro IDE or the Kiro powers website. For this walkthrough, we’ll use the IDE.

  1. Launch the Kiro IDE.
  2. Select the Ghosty icon with the lightning bolt to open the powers panel.
  3. Select Dynatrace Observability from the Recommended
  4. Select Install. The power is registered with placeholder values for the Dynatrace URL and token. Therefore, Kiro will show an error message that the MCP server can’t be reached.
  5. To complete the configuration, select Open Settings and replace the placeholders with your environment details.

Configure your tenant and token

In the settings file, replace the two placeholders:

Placeholder Replace with
YOUR_DT_URL https://TENANT_ID.apps.dynatrace.com/platform-reserved/mcp-gateway/v0.1/servers/dynatrace-mcp/mcp. Replace TENANT_ID with your Dynatrace environment ID (visible in your environment URL, for example https://<ENVIRONMENT_ID>.apps.dynatrace.com/ui).
YOUR_BEARER_TOKEN The Dynatrace platform token you created earlier (for example, dt0s16.XXXXX).

Start asking questions

Open a new chat in Kiro and start interacting with your Dynatrace environment using natural language. Query active problems or security vulnerabilities, request a root cause analysis to identify critical issues in production, or pull related logs and traces, all without leaving the IDE.

See it in action

The short demo below walks through installing the Kiro power for Dynatrace, verifying the connection, and running a first query against your environment to list the top 10 vulnerabilities detected by Dynatrace.

Installing and activating the Kiro power for Dynatrace (video)
Figure 2. Installing and activating the Kiro power for Dynatrace (video)

Get started with the Kiro power for Dynatrace

Kiro powers transform what used to be a stitching exercise (MCP servers here, steering files there, custom instructions somewhere else) into one single, ready-to-use bundle. The Kiro power for Dynatrace applies the same idea to observability: live production insight, causal root cause analysis, and remediation grounded in real telemetry, all available the moment a developer needs them.

The result is a tighter loop between writing code and understanding how it behaves in production. Less waiting for diagnostic data from someone else. Less guesswork from an AI assistant operating without context. And, more time spent on the work that actually matters.

Ready to try it? The Kiro Power for Dynatrace is publicly available: install it from kiro.dev or the Kiro IDE and start asking your environment questions.

Using Kiro and the Kiro power for Dynatrace root cause analysis (video)
Figure 3. Using Kiro and the Kiro power for Dynatrace root cause analysis (video)
Experience the Kiro power for Dynatrace for yourself.

The post Dynatrace observability is now a Kiro power appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/feed/ 0
From reactive to proactive: How NAIC embedded AI‑powered observability directly into the IDE https://www.dynatrace.com/news/blog/how-naic-embedded-ai-powered-observability-directly-into-the-ide/ https://www.dynatrace.com/news/blog/how-naic-embedded-ai-powered-observability-directly-into-the-ide/#respond Fri, 12 Jun 2026 17:54:29 +0000 https://www.dynatrace.com/news/?p=74532 Achieving enhanced observability for Alibaba Cloud in multi-cloud environments with Dynatrace

Every developer knows the feeling: You’re in your IDE when something breaks. Error rates spike, alerts fire, and suddenly you’re out of the flow. Michael Kobush, Performance Engineer III at the National Association of Insurance Commissioners (NAIC®), wanted to eliminate the gap between development and runtime. Instead of switching tools or waiting on SRE support, […]

The post From reactive to proactive: How NAIC embedded AI‑powered observability directly into the IDE appeared first on Dynatrace news.

]]>
Achieving enhanced observability for Alibaba Cloud in multi-cloud environments with Dynatrace

Every developer knows the feeling: You’re in your IDE when something breaks. Error rates spike, alerts fire, and suddenly you’re out of the flow. Michael Kobush, Performance Engineer III at the National Association of Insurance Commissioners (NAIC®), wanted to eliminate the gap between development and runtime. Instead of switching tools or waiting on SRE support, NAIC set out to bring production insight directly into the developer workflow.

Let’s take a look at how NAIC embedded real-time observability directly into their development workflow and reduced investigation time to a few minutes.

The problem: Context switching kills developer productivity

Developers lose time the moment they leave their IDE, jumping between views of logs, metrics, and traces simply to understand what has changed.

For NAIC, this friction was slowing down their teams. Developers didn’t have access to production context, creating a dependency on SRE teams whenever investigations were needed. An analysis that should have taken minutes routinely took 45 minutes to an hour. Root-cause identification required manual correlation across multiple systems, a process that was neither scalable nor sustainable.

At Dynatrace Perform 2026, Kobush demonstrated how his team uses Kiro and Dynatrace at NAIC: Real-time observability in your IDE: How NAIC uses Kiro powers to drive developer productivity

The solution: Kiro powers and intelligent observability

Kiro is AWS’s agentic AI-powered IDE that takes a spec-driven approach to software development by turning natural language prompts into structured requirements, architecture designs, and implementation tasks to carry code from prototype to production.

NAIC installed the Dynatrace power for Kiro, one of Kiro’s installable powers that
dynamically connect domain-specific tools and context to the agent. Once connected, Kiro gives developers and AI agents access to Dynatrace data and insights, helping them pinpoint root causes and receive remediation recommendations directly in their workflow.

No switching between tools. No waiting on another team.

The aha moment: Root-cause analysis in minutes, not hours

The first prompt NAIC ran after connecting Kiro to Dynatrace set the tone for everything that followed. Kobush typed a single line into Kiro: “Tell me about problem P-18576.”

Within 30 seconds, Kiro returned a full problem summary with details and recommendations, pulling everything from Dynatrace automatically. Then, he pushed further: “Give me a really deep dive root-cause analysis of what happened.”

In under two minutes, Kiro returned a full root-cause analysis correlating telemetry, infrastructure signals, historical incidents, and the current problem from Dynatrace into a structured response that included:

  • An executive summary
  • Detailed problem context
  • Infrastructure analysis
  • Technical root-cause analysis
  • Remediation strategies
  • Conclusions and next steps

A preliminary assessment that would have previously taken 45 minutes to an hour was now done in minutes. More importantly, it wasn’t just faster; it gave the team a clear, connected view of how services, infrastructure, and dependencies contributed to the issue.

Beyond root cause: Automation across the entire workflow

What makes this more than just a faster diagnostic tool is how NAIC extended Kiro’s capabilities to automate the full incident response workflow.

Using Kiro’s steering files feature, NAIC configured Kiro to automatically generate a structured Markdown file whenever a root-cause analysis was completed. That file includes:

  • Relevant DQL queries used during the investigation
  • Direct links to the Dynatrace dashboards and data sources that surfaced the issue
  • A clear summary of findings

With a Targetprocess MCP also connected, Kiro can take that analysis and populate a ticket directly, automatically loading all relevant context and sending it to the development team. For NAIC, this means the handoff from investigation to remediation is essentially hands-off. This level of automation doesn’t just save time; it creates consistent, repeatable workflows with built-in guardrails. Every incident gets the same structured, data-rich documentation, regardless of who’s investigating it or when.

This isn’t just about faster incident response. It changes how teams build and release software—giving developers immediate feedback on how their changes behave in real environments.

Proactive alerting: Catching problems before they crash

Root-cause analysis after the fact is valuable. With observability embedded directly into the workflow, teams can detect issues earlier in development and respond faster in production, closing the gap between building and operating software.

After noticing that a specific process had crashed, Kobush asked Kiro to set up an alerting profile that would trigger both before the crash, based on stress signals visible in the logs, and at the point of the crash. Kiro analyzed historical log data, identified pre-crash indicators, and built the alert profile automatically.

The result: NAIC’s team now receives early warning signals before a process fails, giving engineers time to intervene rather than react.

This shift from reactive to proactive operations is central to what the Dynatrace and AWS partnership enables. When observability data is embedded in the developer workflow rather than siloed in a separate platform, the entire engineering organization is better equipped to prevent incidents, not just resolve them.

Debugging a sneaky production bug

Perhaps the most telling story from NAIC’s experience with Kiro occurred during a routine error-rate investigation.

An application error rate had increased unexpectedly. Kobush asked Kiro to investigate. Two minutes later, Kiro identified the culprit: A developer had left debug code in the development environment, and it had made its way into production. Every time a user triggered that code path, it threw errors.

When Kobush sent the Markdown report to the developer, the response was immediate: “How did you find that? I’ve been looking for that.”

Kiro leveraged correlated logs, traces, systems context, and historical behavior from Dynatrace to pinpoint exactly where the issue originated.

Start embedding observability into your development workflow

NAIC’s experience highlights a broader shift: When developers, AI assistants, and systems all operate from the same runtime context, debugging becomes faster, releases become safer, and teams spend less time chasing issues and more time building.

The broader message from Kobush is simple: “I’m not a developer. I have a degree in biology and a minor in chemistry… But this, to me, is a game changer in the observability space. I can do things in seconds that would take me hours.”

The productivity gap between observability data and developer action is a solvable problem.

For DevOps engineers, SREs, and platform teams looking to accelerate incident resolution, reduce context switching, and move from reactive troubleshooting to proactive operations, the Dynatrace and Kiro integration offers a practical, immediately actionable path forward.

For developers, this means fewer interruptions, faster answers, and the ability to stay in flow, even when issues arise.

For more information on how Dynatrace and AWS work together, and to access integration best practices, read our guide, Master AI Observability, or come and see us at an AWS Summit near you.

The post From reactive to proactive: How NAIC embedded AI‑powered observability directly into the IDE appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-naic-embedded-ai-powered-observability-directly-into-the-ide/feed/ 0
Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals https://www.dynatrace.com/news/blog/evaluate-llm-and-agent-quality-in-dynatrace-ai-observability/ https://www.dynatrace.com/news/blog/evaluate-llm-and-agent-quality-in-dynatrace-ai-observability/#respond Thu, 11 Jun 2026 19:44:47 +0000 https://www.dynatrace.com/news/?p=74476

AI applications fail in ways that differ from traditional software. They can return responses quickly, with no errors, and still deliver answers that are inaccurate, ungrounded, unsafe, or unusable. That's why AI quality can't be treated as a side project.

The post Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals appeared first on Dynatrace news.

]]>


For AI systems, reliability is defined by response quality, factual grounding, data security, and usability — and those signals need to live alongside the same observability data teams already trust to monitor performance and availability.

When evaluation scores are isolated in notebooks, spreadsheets, standalone tools, or CI logs, they’re hard to operationalize. By bringing AI quality metrics into Dynatrace AI Observability—next to latency, cost, errors, traces, and user behavior—teams can connect poor responses and hallucinations directly to the prompts, models, retrieval contexts, tool calls, services, and traces that produced them.

What is dt-evals?

dt-evals is an open source CLI for evaluating LLM and agent quality from real GenAI traces, agentic interactions. Teams can run online evaluations against live or recent interactions, score outputs with an LLM judge, and send structured results back to Dynatrace AI Observability so quality becomes visible, queryable, trendable, and actionable.

A minor prompt edit, model change, or retrieval update to an AI application can improve one behavior while quietly breaking another. The challenge to tracking down where and why these systems break is that evaluation results are often maintained outside the operational workflow, making it difficult to connect a low score to the exact trace, prompt, model version, retrieval context, tool call, or service that produced the unwanted behavior.

Dynatrace AI Observability closes this loop. With dt-evals and the AI Observability Evaluation Preview teams can pull recent gen_ai.*  spans, score real interactions with an LLM judge, and write structured evaluation results back as business events. These scores can be viewed with the originating trace, queried for custom analysis, trended in dashboards, and used to trigger alerts or workflow-driven remediation.

A failing faithfulness score is no longer just a number in a report. It’s now an operational signal.

What are LLM evaluations?

An LLM evaluation system scores an AI response against a range of quality and safety dimensions. Common examples include whether the answer is relevant to the question, faithful to the provided context, free of hallucinations, safe for users, complete enough to be useful, and resistant to prompt-injection attempts.

LLM evaluations are typically applied in two modes:

Offline evaluations run before release against a fixed test set or curated trace dataset. These are used to compare a proposed prompt, model, retriever, or agent-tool change against a known baseline before shipping. For example, replay 500 representative support questions in CI and block the release if faithfulness drops below the configured threshold.

Online evaluations run after deployment against sampled production or user traffic. Use online evaluations to detect regressions caused by live inputs, changing retrieval results, tool behavior, traffic mix, or model drift. For example, evaluate 10% of support-agent traces from the last hour and alert the team if hallucination failures exceed the configured window.

With dt-evals, you can run evaluations from the command line, use them in CI/CD, or schedule them to detect quality regressions autonomously after deployment as a post-processing quality gate for your AI agents and LLM output.

AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability
Figure 1. AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability

Run evaluations from the command line

dt-evals is an open source evaluation toolkit for teams that want to bring their own data, judge provider, and evaluation logic while keeping traces, scores, dashboards, and alerts connected.

Install the CLI:

npm install -g @dynatrace-oss/dt-evals

Or run it directly with npx:

npx @dynatrace-oss/dt-evals <command>

A typical first run has three steps:

  1. Configure your environment and judge provider (Bring Your Own AI API key):

dt-evals configure

  1. Verify your local setup and connection:

dt-evals doctor

  1. Run evaluations on recent GenAI traces:

dt-evals run --since 1h --sample 10

This command evaluates traces from the last hour and samples 10% of them. In other words, dt-evals evaluates roughly one out of every ten matching traces, including the prompt and completion messages associated with each selected trace.

During configuration, you provide the connection to your Dynatrace environment and the credentials for the LLM judge provider you want to use. dt-evals does not require teams to send evaluations through a fixed provider. You bring your own judge credentials and control where evaluation execution happens.

For CI/CD use cases, run in CI mode:

dt-evals run --since 6h –ci

In CI mode, dt-evals emits machine-readable output and can fail the pipeline when a configured threshold is breached. This makes quality checks part of the same delivery process used for prompt changes, model upgrades, retrieval updates, and agent releases.

dt-evals in action
Video 1. dt-evals in action

Bring your own LLM judge provider

The “LLM-as-judge” evaluation approach involves using an AI model to score another model or agent response. The LLM judge needs to come from an AI provider your team trusts and has approved for the type of data being evaluated.

dt-evals supports common LLM judge AI models and inference providers, including OpenAI, Anthropic, Google/Vertex/Gemini, AWS Bedrock, and Azure OpenAI. Depending on the package and configuration path you use. Teams provide their own credentials, choose the judge model, and can tune execution settings such as thresholds and concurrency.

This matters for both governance and cost control. Teams can decide which LLM provider is assigned to evaluate which traffic, how many judge calls run in parallel, and where evaluation results are stored.

Which quality and safety dimensions are evaluated by dt-evals?

dt-evals supports built-in LLM-as-Judge evaluators for a range of quality and safety dimensions, including:

  • Relevance: Does the response answer the user’s question?
  • Faithfulness: Is the response supported by the provided context?
  • Hallucination: Does the response invent facts that are not present in the available context?
  • Answer completeness: Does the response fully address the user’s request?
  • Context relevance: Is the retrieved or supplied context useful for answering the question?
  • Factual accuracy: Does the response match an expected or known-correct answer?
  • Summarization quality: Does the summary preserve the important information?
  • Conciseness: Is the response direct with no unnecessary detail?
  • Fluency: Is the response clear and readable?
  • Toxicity: Does the response contain harmful or abusive content?
  • Bias: Does the response show unfair or inappropriate bias?
  • PII leakage: Does the response expose sensitive personal information?
  • Prompt injection: Did the input or response show signs of instruction manipulation?
  • User frustration: Does the interaction suggest the user is blocked or dissatisfied?
  • Drift: Are scores changing meaningfully compared with prior behavior?

A Retrieval Augmented Generation (RAG) application might focus on faithfulness, hallucination, context relevance, and answer completeness. A customer-facing support agent might focus on relevance, fluency, bias, toxicity, and prompt-injection risk. An internal assistant might add custom checks for tone, policy compliance, or whether the answer includes required next steps.

Add custom evaluations

Built-in metrics are useful, but most production AI systems also need checks that are specific to the business, domain, or workflow.

Custom evaluations let teams define their own judge prompts, scoring rules, labels, and thresholds. For example, a support team can create a custom evaluator that checks whether an answer includes a required troubleshooting step before recommending escalation. A financial services team can check whether responses include the required disclaimers. A platform team can check whether an agent uses the correct tool before answering.

The critical point is that custom evaluators run through the same pipeline as built-in evaluators. They can produce the same structured results, appear alongside other scores, and be used in dashboards, alerts, and release checks.

A typical configuration defines the target service, judge provider, sampling strategy, enabled metrics, and thresholds:

schemaVersion: 1
name: support-agent-prod

dynatrace:
  environmentUrl: https://your-env.apps.dynatrace.com
  platformToken: dt0s16.xxxxx

judge:
  provider: openai
  model: gpt-5.5

scope:
  service: support-agent
  since: 1h
  sampling:
    strategy: random
    percent: 10

metrics:
  enabled:
    - faithfulness
    - hallucination
    - relevance
    - drift

alerts:
  thresholds:
    faithfulness: 0.7
    relevance: 0.7

Evaluation results in the AI Observability app

Evaluation results appear directly in the AI Observability app, so teams don’t have to jump between a trace view, an eval report, and a separate dashboard to understand what happened.

In the Prompts view, teams can filter for prompts with evaluation scores and inspect row-level verdicts. Score badges such as relevance, fluency, bias, faithfulness, or toxicity make response quality easy to scan without opening every trace.

This is useful when triaging a regression. Instead of starting with a generic failure count, teams can quickly see which prompts failed, which evaluator failed them, and whether the issue is isolated or widespread.

AI Observability App Prompts stream with evaluation results
Figure 2: AI Observability App Prompts stream with evaluation results

From an individual prompt or trace, the Evaluations tab shows run-level details, including the evaluation name, score, provider, judge model, method, and supporting metadata.

That detail matters because a failed score is only useful if teams can explain it. Engineers and evaluation owners can move from a low score to the exact prompt, response, trace, model, evaluator, and rationale that produced it.

Prompt detail view with evaluation results and trace context in Dynatrace AI Observability
Figure 3:  Prompt detail view with evaluation results and trace context in Dynatrace AI Observability

Query, trend, and alert on evaluation scores

Because dt-evals writes results back as structured events, evaluation scores can be analyzed with the rest of your telemetry.

Teams can ask questions such as:

  • Which evaluator has the lowest average score?
  • Which services are producing the most failed evaluations?
  • Did quality drop after a model or prompt change?
  • Are hallucinations increasing over time?
  • Is quality improving at the cost of latency or token usage?

For example, to get average score by evaluator you could write this query:

fetch bizevents
| filter event.type == "gen_ai.evaluation.result"
| summarize avg_score = avg(gen_ai.evaluation.score.value),
    by: { gen_ai.evaluation.name }
| sort avg_score asc 
Querying failed evaluations by service and evaluator in Dynatrace AI Observability
Figure 4: Querying failed evaluations by service and evaluator in Dynatrace AI Observability

Failed evaluations by service and metric can be determined with this query:

fetch bizevents
| filter event.type == "gen_ai.evaluation.result"
| filter gen_ai.evaluation.score.label == "fail"
| summarize failures = count(),
    by: { dt.service.name, gen_ai.evaluation.name }
| sort failures desc 
Average evaluation scores by evaluator in Dynatrace AI Observability
Figure 5: Average evaluation scores by evaluator in Dynatrace AI Observability

Trending is where evaluation data becomes more useful than a point-in-time report. A single failed score can show an issue. A trend can show whether quality is drifting slowly, whether a release caused a sudden drop, or whether a fix actually improved behavior over time.

On the AI Evaluation & LLM App Performance dashboard, teams can track trends in quality score, pass rate, failed evaluations, drift detections, evaluator health, run cadence, and pass/fail volume over time.

AI Evaluation &amp; LLM App Performance dashboard
Video 2: AI Evaluation & LLM App Performance dashboard

How to turn quality regressions into alerts

Evaluation results can also drive alerts. For example, a support agent team may want to notify the AI team when hallucinations appear in production, or when faithfulness drops for more than a few minutes.

name: support-agent-prod

alerts:
  notifications:
    - name: hallucination-detected
      metric: hallucination
      condition: count > 0
      window: 5m
      channel:
        type: slack
        connection: ai-observability-slack
        channel: "#ai-alerts"

    - name: faithfulness-regression
      metric: faithfulness
      condition: fail_rate > 10%
      window: 15m
      channel:
        type: email
        connection: ai-team-email
        to: [ai-team@example.com] 

Deploy the alerts with:

dt-evals alerts list ./support-agent-prod.yaml
dt-evals alerts apply ./support-agent-prod.yaml

Once applied, Dynatrace runs these checks continuously as Workflows. If hallucinations appear in the last five minutes, the team gets a Slack alert. If more than 10% of faithfulness checks fail over 15 minutes, the AI team receives an email. This turns LLM quality from something teams inspect manually into something Dynatrace can monitor and route automatically.

For continuous alerting, evaluation runs need to happen continuously or on a schedule. You can run dt-evals in CI for release checks (see our example here), run it manually during investigation, or deploy a scheduled runner for ongoing production evaluation. An alert is only as fresh as the evaluation results it carries.

Once configured, quality signals can be routed to the teams that need to act. If hallucinations appear in the last five minutes, the team can receive a Slack alert. If more than 10% of faithfulness checks fail over 15 minutes, the AI team can receive an email. This turns LLM quality from something teams inspect manually into something they can monitor and route automatically.

Close the loop in the AI software delivery lifecycle

Evaluation gates are most useful when they meet developers where they already work. Because dt-evals writes evaluation results back into the observability data layer, those results are not limited to dashboards or post-release reviews. They can be queried, inspected, and acted on from development workflows, CI/CD pipelines, and agentic coding environments.

For example, a team can run dt-evals after any change to a prompt, model, retriever, or agent tool, and then use dtctl (Dynatrace CLI tool for AI Agents) to query the resulting evaluation data, inspect related traces, review dashboards, or validate whether a release threshold was met. In an AI-assisted workflow, tools such as Claude Code, Cursor, GitHub Copilot, or an internal agent harness can leverage MCP or CLI access to bring that same observability context into the developer’s daily workflow.

That closes the loop of the AI software delivery lifecycle: teams can evaluate behavior, control rollout decisions, remediate regressions, and feed production learning back into the next development cycle. Quality signals are no longer in a separate report; they’ve become a part of how AI software is built, shipped, and operated.

Bring evaluations into the release process

Evaluation support is not just for inspection after something breaks. It can also help prevent regressions before they reach users.

Overview of a typical release workflow for an AI app with dt-evals
Figure 6: Overview of a typical release workflow for an AI app with dt-evals

This makes AI quality part of the release process. Teams can gate changes based on relevance, faithfulness, hallucination risk, prompt-injection risk, toxicity, or custom metrics, rather than relying solely on latency and error rate.

What makes this meaningful is that it’s the same pipeline teams already run. AI quality gates sit alongside the latency, error rate, and SLO gates teams have been using for years. There’s no second CI system, no second platform to learn, no second dashboard to monitor. Quality becomes one more dimension of the release decision, gated the same way performance is gated, by the same platform, in the same pipeline.

Coming next

The current experience makes evaluation results visible and actionable inside Dynatrace AI Observability. Next, the focus is on making evaluation workflows easier to run at scale and easier to compare across changes.

Planned improvements include targeted and bulk trace evaluations, custom evaluation libraries, evaluator versioning and lineage, baseline comparisons, experiment views, native quality gates, and deeper visibility into online evaluations.

These capabilities will help teams compare prompt and model variants, understand quality versus cost and latency tradeoffs, and detect sustained quality regressions before they affect more users.

Start today

To get started, check out the Git repository. You’ll need:

  • Node.js 20 or later
  • A Dynatrace environment with GenAI spans and the AI Observability app installed
  • Credentials for the judge provider you want to use
  • A service, trace sample, or CI workflow you want to evaluate

Then, install the CLI:

npm install -g @dynatrace-oss/dt-evals

Configure your service and judge provider:

dt-evals configure

Run your first evaluation:

dt-evals run --since 1h --sample 10

With Dynatrace AI Observability and dt-evals, teams can bring LLM and agent evaluations into the operational loop, where they can trace behavior, score outputs, trend results, alert on regressions, and gate releases before silent failures reach production.

The post Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/evaluate-llm-and-agent-quality-in-dynatrace-ai-observability/feed/ 0
Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/ https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/#respond Mon, 08 Jun 2026 15:03:09 +0000 https://www.dynatrace.com/news/?p=74422 Blog OTP Observability for Agentic AI

Agentic AI is breaking the mold of what organizations need from observability. Fragmented, correlation-dependent observability platforms are no longer “good enough.” Enterprises with dynamic, hybrid environments require observability that provides real-time, precise answers, so AI agents can prevent problems, automate workflows, and deliver better, more secure software.

The post Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI appeared first on Dynatrace news.

]]>
Blog OTP Observability for Agentic AI

As more agentic AI projects come online, the observability market is abuzz with familiar promises: tool consolidation, AI-powered insights, and faster remediation through smarter tools. On the surface, this sounds like progress. But beneath the excitement, many discussions are framed around the wrong question.

The real issue isn’t about how to adopt autonomous operations; it’s about ensuring AI agents are operating reliably and resolving problems without introducing new ones. When evaluating new observability solutions, the question should be:

Can this observability solution accurately analyze complex, dynamic telemetry in context so AI agents can act autonomously with trust, precision, and reliability?

As systems become increasingly agent driven, observability is crossing a structural boundary. Approaches designed for environments where only humans decide and act must adapt to a world where agents increasingly operate autonomously with human oversight, while keeping organizations informed.

Rethinking observability for the agentic age

Observability platforms were initially intended to support engineers in delivering reliable applications, services, and infrastructure to users, and alert them in the event of a problem. Dashboards, alerts, and correlation helped teams investigate incidents, piece together what happened, diagnose issues, decide on next steps, and resolve the problem. This model worked when changes were pushed manually.

The assumption was that more data, better correlation, and cleaner interfaces will lead to increased visibility and improved operational decision making.

Agentic AI systems break that assumption.

With faster release cycles and AI-generated code, manual investigations can no longer keep pace. Moreover, observability platforms must now provide actionable insights to both humans and AI agents.

As agents begin operating as autonomous participants in software environments by triggering mitigations, scaling infrastructure, and optimizing behavior in real time, observability can no longer function solely as a human interface. It must also provide AI systems with a reliable, contextual fact basis that agents can act on programmatically. Machines can’t rely on dashboards and alerts. They require a deterministic foundation of unified, real-time data that delivers accurate, context-rich answers at exabyte scale.

Agentic systems break the mold of “good enough”

Many observability platforms layer probabilistic AI on top of siloed data. They use LLMs to correlate signals and rank likely causes—but they can’t always determine correctness.

“Probabilistic” means that the same input will generate a different output based on a probability distribution of predefined outputs, delivering a different answer when the same problem occurs. This approach is also prone to hallucinations, requiring additional human validation, which can increase operational overhead and token costs, delay resolution of business-critical issues, and divert resources from strategic initiatives.

Enterprise-grade observability must now answer: Is this insight reliable enough for autonomous action?

AI built on siloed data is inherently unreliable. Autonomous systems depend on deterministic, contextual, and trustworthy data to act reliably.

“Deterministic” means that the same input always results in the same output by using factual data to trace the exact causal changes that created the issue. When agentic AI systems act on business-critical applications, the cost of being “mostly right” becomes operationally unacceptable.

This is where a subtle but critical divide appears in the market. Aggregating signals and correlating anomalies can surface patterns. Patterns alone are not a solid basis for decisions, and without deterministic understanding, AI systems inherit that uncertainty and can propagate it downstream.

To drive reliable enterprise autonomous operations, AI agents require a unified, AI-powered observability platform that can analyze exabytes of data in real time and across models to pinpoint root cause, delivering actionable answers in context of what’s affected and its business impact.

From correlated guesses to deterministic answers

This shift in the demands of observability hinges on a clear distinction:

  • Probabilistic AI correlates signals that happened around the same time and therefore appear related, pulling information from fragmented data stores to propose a likely root cause.
  • Deterministic AI uses causal analysis to pinpoint what happened and why, recommend remediation actions, and identify business impact.

Probabilistic AI is intended to narrow the search space and direct engineers toward potential resolution, but it still requires interpretation.

Deterministic AI establishes sequence, dependency, and impact, enabling systems to decide safely without waiting for humans to connect the dots.

Auto‑remediation, auto-prevention, and auto-optimization all depend on this leap. A platform that unifies telemetry only at the UI layer may deliver data and potential root cause, but it can’t compensate for fragmented understanding and missing context underneath. When context is pieced together after the fact, confidence is never guaranteed.

You can’t automate what you don’t precisely understand.

Context driven observability as the control plane for AI

In an autonomous enterprise, observability doesn’t sit beside execution; it’s embedded within it. This integration requires that teams adopt a new mindset toward observability architecture.

Because more AI workloads are happening at the source, telemetry must be optimized and streamlined before ingest, not after the fact, from the edge to the back end. Data access must be unified, context-aware, and always-hydrated on a massive scale. Answers must be explicit, not implicit, and they must be informed by automatic, real-time dependency mapping.

Likewise, intelligence must combine deterministic and agentic AI—not as add‑ons, but as a single reasoning system from ingest to execution.

In this model:

  • AI agents can become the primary consumers of observability data.
  • Humans can shift toward strategy, architecture, oversight, and exception handling.
  • Observability evolves from a reactive lens into a control plane for autonomous operations.

Observability purpose-built for autonomous operations ensures successful agentic AI initiatives

This moment represents an architectural transition, not just an incremental upgrade cycle. Correlation-dependent observability that uses probabilistic AI can be extended, augmented, and rebranded, but it will always carry the limitations of approximation and human validation.

The next era belongs to an observability platform that’s built for machine understanding from the start: a unified, context driven architecture that delivers deterministic answers at machine speed, precision, and scale.

Do you want more data or better decisions? Learn why enterprises are switching to Dynatrace.

The post Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/feed/ 0
The rise of business observability https://www.dynatrace.com/news/blog/the-rise-of-business-observability/ https://www.dynatrace.com/news/blog/the-rise-of-business-observability/#respond Sun, 07 Jun 2026 14:10:28 +0000 https://www.dynatrace.com/news/?p=73757 Business Process analytics

Every organization runs on data. But for most leaders, the challenge is clarity. Traditional dashboards and reports often arrive too late, and system-level metrics don’t explain the business consequences of technical events. When a payment API slows by 200 milliseconds, what does that mean for daily revenue? When user sessions drop, is it a performance […]

The post The rise of business observability appeared first on Dynatrace news.

]]>
Business Process analytics

Every organization runs on data. But for most leaders, the challenge is clarity. Traditional dashboards and reports often arrive too late, and system-level metrics don’t explain the business consequences of technical events.

When a payment API slows by 200 milliseconds, what does that mean for daily revenue? When user sessions drop, is it a performance issue or a recent pricing change? These are the questions business observability is designed to answer.

Business observability connects technical telemetry to business outcomes, creating a shared understanding of how digital performance drives, or hinders, organizational results. It’s less about adding more data, and more about connecting the right data in real time to the decisions that matter.

Importantly, business observability is broader than observing individual business processes. While end‑to‑end workflows like “order to fulfillment” or “claim to payout” are one expression of business observability, the discipline also encompasses digital experience, customer behavior, revenue impact, operational efficiency, and risk. Business observability focuses on understanding how technology performance influences business outcomes across the organization and not just how a single process executes.

Moving beyond traditional monitoring

Traditional monitoring focuses on uptime, latency, and error rates. These are essential for engineers but often disconnected from business context. Business observability closes that gap by linking system health directly to business performance.

That connection – the ability to translate technical telemetry into business insight – is what distinguishes business observability from conventional monitoring.

What makes business observability different

Business observability builds on traditional monitoring by expanding its scope from system health to business outcomes. It introduces three essential capabilities that operate across customer experience, revenue‑generating transactions, and end‑to‑end processes:

  • Business context integration: IT metrics are linked with business KPIs, translating system behavior into measurable impact.
  • End-to-end business process visibility: Processes like “loan approval to disbursement,” “claim to payout,” or “application to onboarding” are monitored across applications, infrastructure, and external systems.
  • Real-time decision support: Business observability surfaces business implications in real time, allowing teams to respond to issues before they escalate.

Together, these principles give organizations a live understanding of how every digital interaction affects outcomes, from conversions and revenue to efficiency and satisfaction.

Where business observability gets its data

Business observability relies on unifying multiple data streams into a single, contextual view of performance:

  • Business events: Structured, contextualized data that represents key business outcomes, such as completed checkouts, submitted claims, bookings, or failed transactions.
  • Logs: Time-stamped records of system activity that show what happened and when, often containing critical technical and business data.
  • Real user monitoring (RUM): Continuous insight into how users actually experience applications and digital services across devices, channels, and geographies.
  • External business tools: Data from tools such as ERP, CRM, billing, and payment platforms that provide commercial, operational, and customer context.
  • Instrumentation and agents: Technologies that collect telemetry and business-relevant data from applications and services, ranging from manual instrumentation to automated approaches that observe activity as it flows through systems with minimal configuration effort.
  • OpenTelemetry: A standardized instrumentation framework that enables consistent collection of metrics, logs, and traces across hybrid and cloud-native environments.

When correlated, these data sources bridge the gap between technical operations and business results, providing a comprehensive view of how technology supports, or disrupts, performance.

How organizations put business observability to work

Modern organizations apply business observability in three key ways. Organizations may pursue these use cases through a variety of approaches, ranging from manually instrumented metrics and custom reporting to more automated, integrated platforms that reduce effort and time to insight.

  1. Drive real-time decisions with IT context: When issues arise, teams and business leaders can see the business impact immediately. A sudden dip in conversions, for example, can be traced to a misconfigured API or third-party service outage, enabling immediate, targeted response.
  2. Track and optimize business processes: Complex workflows—like “order to fulfillment” or “quote to claim”—are mapped from end to end. Teams can pinpoint where time, cost, or customer satisfaction are being lost and address bottlenecks before they affect outcomes.
  3. Accelerate sustainability and reduce costs: By correlating business events with resource usage, organizations can identify inefficiencies in automation, cloud consumption, or scheduling that inflate cost and carbon impact.

Business observability, in action

Each of these examples shows how business observability does more than surface anomalies. It provides the context to act on them. By connecting business events, telemetry, and user experience data, organizations move from simply detecting issues to understanding their impact, cause, and resolution path in real time.

In practice, achieving this level of insight often depends on how business and technical data are captured and correlated. While some organizations rely on custom instrumentation, manual analysis, or post‑incident reporting, others use more automated approaches that make it possible to detect impact, trace root cause, and act in real time.

Financial services

A large financial institution monitored loan application volume as a business event rather than relying solely on system health metrics. When completed applications began declining, traditional monitoring showed no infrastructure failures. Business observability revealed that timeouts from a third-party credit scoring service were affecting only new applicants following a recent integration change. By correlating business events with external service performance, the bank isolated the issue quickly and rolled back the configuration—preventing lost loan volume and downstream compliance exposure.

Payments

A global payment services provider processes billions of transactions annually and needs to understand not just whether systems are available, but how performance impacts transaction success and revenue. By connecting transaction latency and failure rates directly to payment outcomes, the organization can quantify the financial impact of technical issues in real time. This shared visibility allows both internal teams and customers to see how payment flows are performing and proactively optimize transaction speed and reliability.

Travel and hospitality

A large travel platform aggregates booking data across partners, regions, and channels. Business observability enables teams to track quotes, bookings, and completed reservations as business events, segmented by partner and geography. When booking volume dips, teams can immediately determine whether the cause is a partner integration issue, a regional performance problem, or a downstream service slowdown—allowing rapid response to protect revenue across the ecosystem.

Retail and consumer services

A national restaurant chain observed high abandonment rates during online reservation flows. Rather than treating this as a generic user experience issue, business observability correlated real user behavior with backend availability and booking outcomes. This insight enabled automated recovery workflows that re-engaged customers who abandoned reservations due to technical or availability issues—recovering lost bookings without manual intervention.

Aviation and transportation

An international airline struggled with fragmented visibility across booking, pricing, and fulfillment systems. By modeling bookings as end-to-end business processes, business observability provided real-time insight into how technical issues affected customer bookings and operational teams. This shared context improved collaboration between IT, call centers, and operations, reducing response times and improving passenger experience during disruptions.

Insurance and regulated industries

An insurer tracked claims submissions as business events and noticed rising exception rates on mobile channels. Business observability revealed that document uploads from newer devices exceeded a backend file-size limit introduced during a recent update. Because business events, logs, and user session data were correlated in real time, teams deployed a same-day fix and notified affected customers—preventing claim backlogs and demonstrating operational transparency to regulators.

Manufacturing and public sector

Organizations running complex, multi-step production or licensing processes use business observability to monitor each step as it moves across applications, infrastructure, and external systems. In manufacturing and government services alike, correlating process steps with technical events allows teams to identify bottlenecks that delay outcomes—such as throttled payment validation or downstream capacity limits—and resolve them without disrupting customer- or citizen-facing services.

Enabling data-driven executive insight

For executives, business observability shifts technology from a cost center to a source of executive intelligence, providing leaders with real‑time visibility into revenue, customer experience, operational performance, and risk. It enables:

  • Real-time business health monitoring: Live dashboards show key performance indicators such as order volume, claim processing times, or fulfillment success rates.
  • Anomaly detection: Advanced analytics identify deviations in both technical and business metrics before they escalate.
  • Impact analysis: Teams can immediately quantify how issues affect revenue, engagement, or satisfaction, and prioritize based on real business value.
  • Trend analysis: Historical and real-time data combine to forecast outcomes and guide strategic decisions.

Improving collaboration between business and IT

One of the most transformative aspect of business observability is the way it unifies language and priorities across teams.

  • IT teams can prioritize work by business impact instead of technical urgency.
  • Business stakeholders gain visibility into technical dependencies that influence performance.
  • Cross-functional teams align around shared outcomes rather than isolated metrics.

This shared context strengthens trust and speeds response, particularly during digital transformation, where both technology performance and customer experience are constantly evolving.

Laying the groundwork for business observability

While business observability may be expressed through dashboards, process views, or executive metrics, its success depends less on how data is visualized and more on how consistently business outcomes are connected to technical signals across teams. Adopting business observability effectively requires thoughtful preparation:

  • Data integration: Pull from diverse sources – applications, infrastructure, and business systems – to form a cohesive view.
  • Shared metrics: Define how business KPIs map to technical signals.
  • Organizational alignment: Ensure teams are trained and incentivized to act on shared insights.
  • Platform scalability: Choose an observability solution that supports hybrid, cloud, and partner ecosystems as data volume grows.

The road ahead

Business observability represents a shift from reactive monitoring to outcome-driven intelligence. Rather than focusing solely on system health or individual process performance, it enables organizations to understand in real time how technology influences revenue, customer experience, operational efficiency, and strategic decision‑making. For executives, it means real-time visibility into how technology influences outcomes. For IT teams, it means prioritizing based on business value. For organizations, it means a unified, data-driven way to make decisions with confidence.

In practice, many organizations attempt to reach this level of insight through manually instrumented metrics, custom dashboards, and offline analysis – approaches that require ongoing effort and often delay understanding when it matters most. Dynatrace Business Observability takes a different approach, capturing business events from multiple sources and automatically correlating them in real time with full‑stack telemetry. This delivers the context leaders need to act decisively, without the manual overhead of traditional approaches.

Take Dynatrace for a spin

FAQs: Business observability

What is business observability in simple terms?

Business observability is the ability to understand how technical performance affects business outcomes in real time. It connects IT telemetry, such as logs, traces, and metrics, with business data such as transactions, claims, or orders, so teams can see both the cause and the consequence of an issue in one view.

How is business observability different from traditional monitoring?

Traditional monitoring reports on system health (uptime, latency, or resource use) without showing the business impact. Business observability goes further by linking these technical metrics with key performance indicators (KPIs), providing the context to understand how technical changes influence revenue, customer experience, and operational efficiency.

What types of data does business observability rely on?

Business observability combines multiple data sources, including:

  • Business data from transactions or processes
  • Logs and traces from applications and infrastructure
  • Real user monitoring (RUM) for end-user experience
  • Data from external business tools like CRM, ERP, or payment systems
  • Instrumentation frameworks such as OneAgent and OpenTelemetry for consistent telemetry across environments

Who benefits most from business observability?

Executives gain real-time visibility into how technology affects performance and outcomes. IT and engineering teams gain business context to prioritize fixes based on impact. Together, these perspectives drive faster decision-making and closer alignment between technology operations and business goals.

What are common use cases for business observability?

Typical use cases include:

  • Detecting and resolving process slowdowns in finance, healthcare, or logistics workflows
  • Tracking and optimizing customer journeys and digital transactions
  • Measuring the business impact of new releases or integrations
  • Identifying inefficiencies that increase costs or carbon footprint

How does business observability support sustainability and cost efficiency?

By correlating business events with resource consumption, business observability helps identify where automation, compute, or storage are overused. This enables teams to optimize cloud spend, reduce energy consumption, and track sustainability metrics alongside operational performance.

What challenges do organizations face when implementing business observability?

Key challenges include integrating data from multiple systems, defining the right shared KPIs between business and IT, and ensuring teams are trained to act on insights collaboratively. Success depends on cross-functional alignment as much as on technology.

Is business observability only relevant for large enterprises?

No. Any organization that relies on digital processes – from mid-sized financial firms to healthcare networks or government agencies – can benefit. The ability to link technical performance to business results is valuable at any scale

The post The rise of business observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-rise-of-business-observability/feed/ 0
Port and Dynatrace: One-prompt incident triage with the Dynatrace MCP Server https://www.dynatrace.com/news/blog/port-and-dynatrace-one-prompt-incident-triage/ https://www.dynatrace.com/news/blog/port-and-dynatrace-one-prompt-incident-triage/#respond Fri, 05 Jun 2026 12:28:09 +0000 https://www.dynatrace.com/news/?p=74412 Agent graphic

The Dynatrace MCP Server is now available in Port via Port MCP Connectors. A single OAuth flow connects it in minutes. Set up Port AI to communicate with Dynatrace, GitHub, Slack, and your service catalog in a single conversation, correlating production signals, code, and ownership in a single agent run rather than three browser tabs […]

The post Port and Dynatrace: One-prompt incident triage with the Dynatrace MCP Server appeared first on Dynatrace news.

]]>
Agent graphic

The Dynatrace MCP Server is now available in Port via Port MCP Connectors. A single OAuth flow connects it in minutes. Set up Port AI to communicate with Dynatrace, GitHub, Slack, and your service catalog in a single conversation, correlating production signals, code, and ownership in a single agent run rather than three browser tabs and a manual handoff. For incident triage, this turns a multi-tool investigation into a single prompt: just ask Port AI what’s wrong; you’ll get details about the failing service, the error signature, the file and function, and the suspect commit.

Ground every Port AI conversation in live production data from Dynatrace

Port is an agentic engineering platform that platform teams use to organize their software development lifecycle. It gives software and DevOps teams a central place where engineers can find system information, take action on it, and route work across their connected tools, without waiting on IT or operations.

Port AI is the assistant that queries the catalog using natural language. Through Port MCP Connectors, the same chat also reaches external systems, such as Dynatrace, GitHub, and Slack.

Dynatrace complements Port by providing context-rich observability and security insights, right where you need them:

What Dynatrace brings What Port brings
Live observability signal across logs, traces, and metrics Service ownership and team responsibility
Dependency topology between affected services On-call rotation and escalation paths
Open problems with root cause already identified Recent deploys and commit history per service
Security vulnerabilities and exposures detected in running services, with severity and affected entities Remediation ownership and the team that’s accountable for the fix

The result: Team members ask a question in the Port AI chat, where they’re already working, and get back complete answers that no single tool could produce on its own.

Complete triage run with Port AI [VIDEO]
Figure 1. Complete triage run with Port AI [VIDEO]

Incident triage from a single Port AI prompt

The Dynatrace MCP Server provides Port AI with a set of tools it can call during any conversation. For incident triage, the most relevant needs are:

  • Query production data. Logs, traces, metrics, and events from across the Dynatrace tenant returned in structured form.
  • List open problems. An overview of all active problems on the tenant.
  • Get problem details. Root cause, causal chain, and affected entities for a specific problem.
  • Get troubleshooting guidance. Relevant troubleshooting guides matched to a problem description.

The full toolset also covers security findings, entity and topology lookup, query generation, forecasting, and more. See the Dynatrace Hub for the complete list. In combination with Port AI, these capabilities turn the Port AI chat into a single place to ask production questions.

The example below walks through the triage of an incident affecting broker_service, a fictional service. The same investigation, done manually without this integration, starts in Dynatrace (where the failing service shows up in seconds), then jumps to GitHub to scan recent commits, to Slack to confirm ownership, and back to a doc to write up the summary. With this Port AI integration, those steps run in a single agent run, with the Dynatrace signal at the center of the chain.

The scenario begins when broker_service degrades and lands as a new incident, INC-1003. In Port AI chat, an SRE asks, “Help me understand the root cause of INC-1003.”

Port correlates signals from Dynatrace, GitHub, and Slack in a single agent run.
Figure 2. Port correlates signals from Dynatrace, GitHub, and Slack in a single agent run.

Port AI loads the ai-incident-triage skill and runs the following steps:

  1. Resolve the incident in Port’s catalog. Port AI looks up the incident entity and pulls the affected service identifier.
  2. Query Dynatrace. The Dynatrace MCP Server queries logs and traces for broker_service. The response carries the first-seen failure timestamp and the top error signature.
  3. Find the suspect commit in GitHub. Port AI passes the failure window to the GitHub MCP Server, locates the failing function, and lists the commits to that file. One commit aligns with the first-seen timestamp.
  4. Return a structured triage summary. Port AI returns the failing service, the error signature, the file and function, and the suspect commit.
  5. Post to Slack and close the loop. The Slack MCP Server posts the same summary to #incident-updates (Figure 2). The incident entity records a triaged_at timestamp through a Port self-service action.
The triage summary is posted to the Slack #incident-updates channel.
Figure 3. The triage summary is posted to the Slack #incident-updates channel.

Within a single agent run, the engineer receives a triage summary in Slack that already includes the live Dynatrace signal.

The same Dynatrace integration allows many more use cases, such as deployment correlation, on-call summaries, and postmortem drafts, each built as a Port AI skill.

Security triage works just as easily: ask Port AI about a vulnerability; the Dynatrace MCP Server returns the affected running services, severity, and exposed entities, while Port resolves ownership and routes the fix to the accountable team.

Roll it out across teams, govern centrally

The integration is designed for organization-wide rollout. Admins maintain central control over which tools are exposed and who can access them, while each query remains scoped to the user’s existing permissions.

  • Per-tool selection. Admins choose which Dynatrace tools Port AI can call across the organization. Sensitive tools can be scoped to selected groups.
  • Per-user authentication. Each user authenticates to Dynatrace through OAuth. Queries return only the data that their existing Dynatrace permissions already allow.
  • Audit trail on both sides. Every Port AI call and every Dynatrace MCP call names the same person, with no stitching required between platforms.
  • Scales without new workflows. The same per-user model that works for a pilot team works for hundreds of developers. No separate access-request workflow needed.

Get started: connect Port with Dynatrace

The Dynatrace MCP Server connects to Port through a single OAuth flow. Setup takes a few minutes and is done once by an admin.

  • Add Dynatrace as a data source. In Port, go to Data Sources > + Data source > MCP Servers. Select Dynatrace, fill in the connector details, and select Connect to authenticate with your Dynatrace tenant.
  • Expose the tools you want Port AI to use. Under Allowed Tools, add the Dynatrace capabilities you want available to your organization (querying production data, listing problems, getting problem details, finding troubleshooting guidance). Select Publish.
  • Register the incident triage skill. Add the ai-incident-triage skill from Port’s skill library and point it at the incident entities in your catalog.

Once published, the integration is available to every authenticated Port user in your organization. Each user authenticates to Dynatrace individually through OAuth on first use.

For full setup walkthroughs, see Port documentation

Port MCP Connectors documentation: connector setup and admin configuration

Triage incidents with AI: full skill walkthrough

The post Port and Dynatrace: One-prompt incident triage with the Dynatrace MCP Server appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/port-and-dynatrace-one-prompt-incident-triage/feed/ 0
AI agents are redefining software development—but they’re flying blind without observability https://www.dynatrace.com/news/blog/ai-agents-are-redefining-software-development-but-theyre-flying-blind-without-observability/ https://www.dynatrace.com/news/blog/ai-agents-are-redefining-software-development-but-theyre-flying-blind-without-observability/#respond Thu, 28 May 2026 17:09:42 +0000 https://www.dynatrace.com/news/?p=74210 AI agents are redefining software development

Imagine a team of AI agents building, deploying, and running software at machine speed—yet unable to see what’s happening in production. This is the new reality for enterprise technology leaders. As one Fortune 500 CTO told us, “Speed is now the primary driver of innovation, forcing organizations to rethink processes, compliance, and roles; it’s a […]

The post AI agents are redefining software development—but they’re flying blind without observability appeared first on Dynatrace news.

]]>
AI agents are redefining software development

Imagine a team of AI agents building, deploying, and running software at machine speed—yet unable to see what’s happening in production. This is the new reality for enterprise technology leaders. As one Fortune 500 CTO told us, “Speed is now the primary driver of innovation, forcing organizations to rethink processes, compliance, and roles; it’s a necessity for innovation teams.”

Observability—real-time visibility into how software behaves in production—has become the critical enabler for both human-led and agent-led teams. Without it, AI agents are powerful but blind.


Key executive insights

  1. Software production is being redefined by AI agents. This transformation is a structural shift, not a trend.
  2. The world is bimodal again. Human-led and agent-led environments coexist.
  3. AI agents are powerful but blind. Without rich context from production, they cannot deliver reliably.
  4. Observability is a crucial enabler to gradually transform from human-led to agent-led operations. Observability is what allows organizations to industrialize software delivery with confidence.
  5. The new KPI for agent-led teams is the percentage of human intervention required. The lower the number, the better the AI is working.

The market reality: A bimodal world

Organizations are accelerating AI adoption not because it is trendy, but because it is existential. Companies that fail to transform risk being outpaced by competitors that can deliver software faster, cheaper, and at higher quality. CTOs and CIOs are making statements like “speed over compliance” not out of recklessness, but because they recognize that without radical acceleration, their businesses face disruption.

At the frontier of this shift is a fundamentally new way of building software: AI-first development. In these environments, 100% of coding, testing, deployment, operations, bug fixing, and optimization are performed by AI agents. The human role shifts to specification, goal setting, supervision, and correction. Intellectual property moves from the code to the specification—code becomes a generated artifact, not the source of truth. With a complete, well-architected spec, agents can fully rebuild the software from it again.

This creates a bimodal operating environment:

  • Human-led teams—the majority today—are existing operations, SREs, and developers augmenting their workflows with AI. They follow the traditional SDLC, increasingly supported by AI agents that auto-prevent, auto-remediate, and auto-optimize, which reduces manual effort and achieves more with the same resources.
  • Agent-led teams—growing fast—are innovation groups operating in full AI development life cycle (AIDLC) mode. Swarms of AI agents build, deploy, and run software end-to-end. Humans write specifications and intent, not code. For these teams, the KPI is no longer “how many story points were solved?” but “what percentage of human intervention is required?”

Observability enables a reliable transition to autonomous operations

In the early 2010s, a similar bimodal pattern emerged with cloud: one team running thousands of servers on-premises, another in stealth mode on AWS. The pattern is repeating now with AI.

Why not switch everything to agent-led right away? Because existing systems follow processes, compliance, and technology stacks that can’t be immediately automated in an AI-first way. Moreover, it’s too risky to move all business-critical systems simultaneously. The safer path: start with an innovation team, build less critical applications first, and only when those are successful and trusted, begin migrating more of the business-critical services.

New foundation models that arrived in early 2026 have accelerated the path to fully autonomous operations, making agent-led teams realistic at small scale today, with large scale within sight. These systems focus on AI-first software generation first, with a clear goal to eventually master operational challenges (resilience, performance, scale, security) entirely with agents as well.

Observability plays a critical role not only in making both modes work reliably, but also in enabling the transformation from the first mode to the second. The context observability provides—understanding existing system behavior, dependencies, and requirements—is exactly what agents need to create the reliable and scalable software. Observability is what makes both modes work, and it is the critical bridge between them.

The core problem: AI agents are blind

AI agents can code, deploy, refactor, and operate software faster than humans ever could. But there is one thing AI cannot do without help: AI has no awareness of what happens in production. It’s blind to the real world: without real-time feedback from running software—in development and production —agents make decisions without context and without understanding their consequences. They operate at speed, but without sight.

77% of IT teams still lack full visibility across hybrid environments (IBM Institute for Business Value, 2025). If you can’t see it, you can’t scale it. Observability is not optional for AI-first operations, it’s a prerequisite.

The Dynatrace response: Real-time observability for both worlds

Dynatrace addresses both sides of this bimodal reality: a complementary response to the two speeds at which enterprises now operate.

For human-led teams: Autonomous operations at scale

This year, Dynatrace launched Dynatrace Intelligence: a full agentic operations system that orchestrates dozens of agents that auto-prevent, auto-remediate, and auto-optimize across site reliability, development, and application security. These AI agents deliver the following value in production:

  • SRE Agent: Kubernetes troubleshooting, infrastructure optimization, and automated incident resolution – reducing mean time to resolution at scale.
  • Developer Agent: Surfaces production context during deployment, validates changes, and prevents issues before they reach customers.
  • Security Agent: Identifies vulnerabilities, triages threats, and accelerates security response, all in real time.

The deterministic foundation underneath: what separates Dynatrace agents from others is its deterministic foundation: real-time, full-stack, and cross-model root-cause analysis, anomaly detection, and forecasting, all grounded by data in a unified, purpose-built data lakehouse that delivers accurate, contextual answers from exabytes of information. This is not AI that guesses; it’s AI that reasons from facts. Benchmarks from internal testing and observed customer use cases: 12× higher success rate in SRE use cases, 3× faster problem resolution, 2.5× lower token cost.

Ecosystem integrations that extend intelligence beyond the platform: Dynatrace Intelligence extends into third-party tools to drive autonomous actions across development, SRE, and ITOps workflows.

For agent-led teams: develop and run software reliably

Dynatrace enables AI-first teams to let swarms of agents to build and run software reliably, providing real-world awareness from observability, run-time context across development, security, and operations, and self-optimization toward SLAs, cost, and resilience. Key capabilities include:

  • Agentic observability: Closed-loop autonomous operations where observability agents coordinate with coding and deployment agents to self-heal.
  • AI and cloud observability: Full-stack visibility across cloud infrastructure and AI workloads, covering resilience, performance, security, user experience, and LLM evaluations to assess the quality and reliability of agent outputs, helping identify potential inaccuracies, hallucinations, or risks.
  • AI data lakehouse (Grail): Real-time context engine that provides long-term memory for agent decisions—sub-second, API-native, at an exabyte scale.

The goal: a closed loop where agents detect issues, resolve them, and ship the fix—autonomously, 24/7.

Dynatrace is on the same bimodal journey – our entire business runs on Dynatrace Intelligence in human-led mode, with agents taking over more tasks continuously, while our AI-first offering and new services are built and operated entirely by agent swarms, using our own observability to close the feedback loop.

Different approaches – unified platform

Across the platform, Dynatrace delivers end-to-end, full-stack visibility across cloud infrastructure, applications, and AI workloads, including agent behavior, decision paths, and cost, along with governance at machine scale. These capabilities serve human-led and agent-led teams differently, but from the same unified platform.

The measure of success in software delivery is shifting from human productivity metrics to a new KPI: the percentage of human intervention required. Observability is what makes that progress possible. The question for every technology leader is no longer whether to adopt AI-first, but how quickly they can close the visibility gap before competitors do to drive massive growth in innovation and productivity.

The post AI agents are redefining software development—but they’re flying blind without observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ai-agents-are-redefining-software-development-but-theyre-flying-blind-without-observability/feed/ 0
Dynatrace MCP Server for Atlassian Rovo: Investigate production problems without leaving Jira or JSM https://www.dynatrace.com/news/blog/dynatrace-mcp-server-for-atlassian-rovo-investigate-production-problems-without-leaving-jira-or-jsm/ https://www.dynatrace.com/news/blog/dynatrace-mcp-server-for-atlassian-rovo-investigate-production-problems-without-leaving-jira-or-jsm/#respond Wed, 27 May 2026 19:37:02 +0000 https://www.dynatrace.com/news/?p=74203 Atlassian and Dynatrace

A developer triaging a bug in Jira, or an on-call engineer responding to an alert in Jira Service Management (JSM), can ask their Rovo agent to investigate production issues in Dynatrace using plain language, without leaving Jira or JSM. The Rovo agent answers any relevant questions, calls the appropriate Dynatrace tools, and posts the results […]

The post Dynatrace MCP Server for Atlassian Rovo: Investigate production problems without leaving Jira or JSM appeared first on Dynatrace news.

]]>
Atlassian and Dynatrace

A developer triaging a bug in Jira, or an on-call engineer responding to an alert in Jira Service Management (JSM), can ask their Rovo agent to investigate production issues in Dynatrace using plain language, without leaving Jira or JSM. The Rovo agent answers any relevant questions, calls the appropriate Dynatrace tools, and posts the results back to the Jira workspace. Admins can now complete setup in minutes. Every call runs with the permissions of the requesting user, every action is logged, and data access and cost remain under central control.

Move from “ask the platform team” to “ask the agent”

If your organization uses Dynatrace, much of your production knowledge may reside with the Dynatrace experts on your organization’s platform or SRE team. Everyone else (developers triaging Jira bugs, on-call responders running down JSM alerts, or support engineers handling escalations) either learns enough Dynatrace to investigate production issues on their own or pings the platform team and waits for a response.

The Dynatrace Model Context Protocol Server for Rovo moves this valuable production expertise into the Rovo agent, where it’s accessible to all Rovo users. The Rovo agent holds the Atlassian context (the bug, the alert, the service it relates to) and calls Dynatrace for the production context (problems, topology, root cause). Meanwhile, the developer asking the question gets a usable answer directly in the tool where they’re already working.

Setup completes in minutes. The integration is usable on the first prompt.
Setup completes in minutes. The integration is usable on the first prompt.

How does an on-call engineer use Rovo to triage a JSM alert?

A JSM alert is triggered when a service degradation is detected. The on-call engineer opens the alert’s response panel and asks the Rovo Ops Agent, Atlassian’s built-in AI agent for JSM incident response, to investigate.

Rovo Ops calls the Dynatrace MCP Server, retrieves the open problem, the root cause, and the related signals, and returns a single answer: what went wrong, what’s related, and what to do next. The on-call engineer either acts on this problem context directly or escalates the alert, with the full analysis already attached.

How does a developer triage a Jira bug with Rovo and Dynatrace?

Let’s say Jira provides details of a bug in a service that a developer doesn’t own. The developer asks the Rovo agent a question in the issue’s chat panel: “What’s going on with this service right now?” Rovo returns the open Dynatrace problem, the elevated error rate, and the dependency that’s causing the issue.

Rovo’s answer posts as a comment on the bug in Rovo, so the next person who opens the ticket sees the investigation is already done. From there, the bug typically closes as a duplicate of the active incident or routes to the team that owns the failing deployment, in minutes rather than hours.

How to roll out the Dynatrace MCP Server in Atlassian Rovo

MCP Server setup is a configuration task, not a project. The Dynatrace MCP Server is pre-approved by Atlassian as an external MCP integration and is included with Dynatrace SaaS at no extra cost. In a few clicks, an admin connects Dynatrace using the Rovo admin UI, authenticates against the Dynatrace tenant, and selects which tools to expose. Once connected, the integration is available across Jira, Jira Service Management, and Confluence.

The Dynatrace MCP Server tool set is ready for immediate use, providing data retrieval through Grail, topology and entity context, root cause analysis, active security findings, time-series forecasting, and change-point analysis.

Security, governance, and cost stay under administrator control

Admins keep control over security, governance, audit, and cost on both the Atlassian and Dynatrace sides.

  • Per-user enforcement, end-to-end. Every Dynatrace call runs as the requesting user via OAuth 2.1, with authorization based on the user’s Dynatrace permissions. On the Atlassian side, Rovo and Rovo Ops only see the Jira and JSM data that the user is already entitled to see.
  • Per-tool selection. From a checklist in the admin panel, administrators choose which Dynatrace tools the Rovo agents in the organization can call. (New tools released later must wait for admin approval before they become available.)
  • Audit trail on both sides. Every Rovo invocation and every Dynatrace MCP call names the same person, with no stitching required between platforms.
  • Usage and costs are observable in Dynatrace. Tool call volume can be observed with Dynatrace, and costs can be clearly attributed to data owners.
  • Adding more Dynatrace users doesn’t increase your cost. Dynatrace consumption is priced on data, not per user, so onboarding more developers and on-call engineers doesn’t instantly impact your billed costs.

Get started with the Dynatrace MCP Server for Rovo now

The integration is generally available. To connect it:

  1. In Atlassian Jira or JSM, go to Atlassian AdministrationRovoRovo MCP server.
  2. Add the Dynatrace MCP Server and authenticate against your Dynatrace tenant.
  3. Select the Dynatrace tools you want to expose to your Rovo agents.

Full setup instructions are available in the Atlassian documentation and the Dynatrace MCP Server documentation.

Because the MCP Server is included with Dynatrace SaaS, you can easily set up a pilot program. Just connect the integration for one dev team, allow them access to a narrow set of tools, and then monitor the team’s usage and related costs. The patterns that emerge (which tools are called, by whom, and at what cost) can serve as a basis for a confident wider rollout.

The post Dynatrace MCP Server for Atlassian Rovo: Investigate production problems without leaving Jira or JSM appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-mcp-server-for-atlassian-rovo-investigate-production-problems-without-leaving-jira-or-jsm/feed/ 0
Dynatrace Release Radar 04.26 https://www.dynatrace.com/news/blog/dynatrace-release-radar-04-26-whats-new-and-why-it-matters/ https://www.dynatrace.com/news/blog/dynatrace-release-radar-04-26-whats-new-and-why-it-matters/#respond Wed, 27 May 2026 17:07:44 +0000 https://www.dynatrace.com/news/?p=74179 Release Radar

This series covers recent Dynatrace releases and updates, focusing on what’s new, what’s changed, and how these recent enhancements can benefit you and your organization. Each post covers newly available capabilities and points you toward where to explore them.

The post Dynatrace Release Radar 04.26 appeared first on Dynatrace news.

]]>
Release Radar

The April 2026 Dynatrace SaaS releases bring six updates aimed at a familiar problem: too much manual work between a signal and an answer. The updates focus on native cloud visibility, deeper Kubernetes insights, consistent severity handling, faster investigations, and smoother analytics work.

Explore all updates hands-on in the Release Radar launchpad.

A native Azure experience in Clouds

What does Dynatrace add for Azure practitioners?

Dynatrace now extends the enhanced Clouds experience to Microsoft Azure, putting Azure subscriptions on the same footing as AWS. Metrics, logs, metadata, and topology now sit in one managed view, so Azure teams can move from inventory to investigation without stitching the picture together by hand.

What you get out of the box

  • Opinionated insights and ready-made dashboards built from enriched Azure telemetry, so investigations start with answers instead of a blank canvas.
  • Pre-configured health alerts for Azure, created and managed directly in Clouds, with drill-down, search, and filtering directly in the Clouds app.
  • Broad metric coverage for any Azure Monitor native platform metric across Azure services.
  • A rich Azure topology inventory that periodically scans Azure environments and enriches resources with native metadata such as tags and subscription IDs, all queryable with Dynatrace Query Language (DQL).
  • Simple onboarding and lifecycle management that turns Azure subscriptions into native Dynatrace connections and manages them centrally.

Dynatrace also adds drilldowns from cloud entities into the relevant Dynatrace experiences, so teams can keep moving instead of bouncing between cloud and platform views.

The new Clouds experience for Azure lets you optimize cloud operations at scale.
The new Clouds experience for Azure lets you optimize cloud operations at scale.

Kubernetes visibility for autoscaling and custom resources

What’s new in Kubernetes observability?

Dynatrace extends Kubernetes visibility to two additional object types that SREs and platform teams rely on daily: Horizontal Pod Autoscalers (HPA) and Custom Resources (CRs).

Horizontal Pod Autoscaler as a first-class object

HPA is now a first-class object in enhanced Kubernetes visibility. You can see when scaling kicked in, what triggered it, and how desired and actual replica counts lined up next to the workloads involved.

Custom Resource insights

You can monitor up to five Custom Resources per cluster, surfaced the same way as built-in Kubernetes objects. This brings CRD-heavy ecosystems such as Argo, Istio, Cert-Manager, Kyverno, and operator-managed databases into the same investigation scope as the rest of your cluster.

For clusters connected through cloud integrations, the Kubernetes cluster details page now exposes the underlying cloud configuration (EKS, AKS, or GKE) in YAML or JSON, making cloud-side and cluster-side state accessible in one place.

HorizontalPodAutoscaler visibility in the Kubernetes app experience.
HorizontalPodAutoscaler visibility in the Kubernetes app experience.

A unified severity model for alerts and problems

What is event.severity in Dynatrace?

Dynatrace introduces a standardized event.severity field for alerts and problems, aligned with the ITIL Incident Management framework. Severity is stored in Grail as an integer from 1 (Critical) to 5 (Informational) and is shown as a human-readable label across the platform.

Severity levels at a glance

Value Label Description
1 Critical Major business disruption; service outage
2 Major Significant impact; workaround may exist
3 Minor Limited or non-critical impact
4 Warning Low impact; no business disruption
5 Informational No business impact

Severity automatically propagates from correlated alerts to the parent problem, with the highest severity always taking precedence. This gives teams one severity model to filter on, route with, and escalate from.

You can now:

  • Filter the problem feed by severity
  • Display a severity column with visual icons in problem lists
  • Use severity as a condition in Workflows for alert routing and notifications
Event severity in the Problems app experience.
Event severity in the Problems app experience.

Faster Investigations with Smartscape navigation

What changed in Smartscape?

Smartscape now offers all six ready-made views, such as vertical topology, horizontal topology, and visual resolution path, just a click away in a persistent side panel. You no longer need to return to the landing page in the middle of an investigation.

The new Recent views section shows your latest investigations, making it easy to reopen them, compare them, and keep working as you test a root-cause hypothesis.

The result is less backtracking in the middle of an incident.

The new sidebar navigation in Smartscape
The new sidebar navigation in Smartscape

Dashboards and notebooks: productivity improvements

What’s new for dashboard authors and analysts?

The latest release adds several practical upgrades for team members who build dashboards and work in notebooks every day.

  • Treemap visualization for identifying dominant categories in hierarchical data, such as requests per service by Kubernetes namespace.
    Treemap visualization example
    Treemap visualization example
  • Dashboard variables for dynamic coloring and thresholds, so visual conditions stay in sync with environment or team selectors.
    Use dashboard variables for dynamic coloring and threshold conditions
    Use dashboard variables for dynamic coloring and threshold conditions
  • Centralized tile indicator controls, allowing you to show or hide warnings, descriptions, and custom timeframes at the dashboard level.
    Select or clear tile indicators on a dashboard
    Select or clear tile indicators on a dashboard
  • Direct image upload in Markdown using a built-in image library shared across Dashboards, Notebooks, and the Launcher.
    Upload image directly in Markdown
    Upload image directly in Markdown
  • Row marker coloring for tables, making it easier to visually group related rows without sacrificing readability.
    Highlight table rows with color markers in Dashboards and Notebooks
    Highlight table rows with color markers in Dashboards and Notebooks

User experience improvements

Why does the platform feel faster?

This release smooths out the path from the first symptom to root cause analysis. Tracing and services workflows now handle high-span traces more reliably, show timing more clearly, and surface useful sample traces earlier.

Table-first workflows also benefit from richer entity-detail tables, better filtering, and clearer structure, helping teams answer more questions without switching views. Navigation patterns, overlays, and error messaging are now more consistent across the platform, reducing mental overhead when time is tight.

Explorer new table experience with entity details, alerts, and schema links
New Explorer table experience with entity details, alerts, and schema links

Why these updates matter

Taken together, these updates eliminate inefficiencies in the work that teams do every day. Cloud operations teams, Kubernetes SREs, on-call engineers, and analytics authors get richer context, faster paths to answers, and simpler ways to share what they find. This is where Dynatrace earns its keep under pressure.

Explore the updates live in the Release Radar launchpad.

The post Dynatrace Release Radar 04.26 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-release-radar-04-26-whats-new-and-why-it-matters/feed/ 0
Scaling enterprise AI with confidence: Dynatrace joins the Dell Technologies AI Ecosystem Program https://www.dynatrace.com/news/blog/scaling-enterprise-ai-with-confidence-dynatrace-joins-the-dell-technologies-ai-ecosystem-program/ https://www.dynatrace.com/news/blog/scaling-enterprise-ai-with-confidence-dynatrace-joins-the-dell-technologies-ai-ecosystem-program/#respond Tue, 19 May 2026 19:02:13 +0000 https://www.dynatrace.com/news/?p=74013 Dynatrace and Dell Technologies

Most enterprises have moved past the deployment problem. The harder question is what those workloads are doing in production: where GPU spend is going, how agent chains are behaving, and whether compliance teams can answer when regulators ask. When the answers aren’t clear, the consequences land fast and are rarely contained to one team. That’s […]

The post Scaling enterprise AI with confidence: Dynatrace joins the Dell Technologies AI Ecosystem Program appeared first on Dynatrace news.

]]>
Dynatrace and Dell Technologies

Most enterprises have moved past the deployment problem. The harder question is what those workloads are doing in production: where GPU spend is going, how agent chains are behaving, and whether compliance teams can answer when regulators ask. When the answers aren’t clear, the consequences land fast and are rarely contained to one team.

That’s why Dynatrace is joining the Dell Technologies AI Ecosystem Program, bringing full-stack AI and LLM observability natively into a broad and integrated AI infrastructure ecosystem. Dell delivers the validated, integrated infrastructure to run AI at scale. Dynatrace brings the observability, automation, and governance to operate it with confidence, with visibility from GPU infrastructure to model behavior to end-user experience. Together, they give enterprises the control to match the scale they’ve already built.

The real challenge: AI at enterprise scale

Running AI in a pilot is very different from running it at scale across the business with real users, regulated data, and demanding SLAs. As we’ve worked with enterprises across industries, these failure patterns come up repeatedly:

Cost

As enterprises scale AI, costs spiral rapidly and unpredictably across model providers, GPU clusters, and inference APIs without clear line of sight into what is driving spend or whether it’s delivering value.

Observability gaps

Traditional monitoring tools weren’t built for AI pipelines. Fragmented observability across GPU clusters, orchestration layers, and inference APIs creates blind spots while LLM latency and token throughput fluctuations under load remain difficult to diagnose and even harder to predict.

Agentic complexity

Multi-step agent workflows introduce cascading failure modes. A silent error in one tool call can corrupt downstream decisions across the entire chain.

Compliance & governance

Enterprises need continuous monitoring to detect model drift, hallucinations, and unsafe outputs before they impact end users. Regulated industries need audit trails, data governance, and behavioral monitoring that most AI monitoring bolt-ons simply weren’t built for.

These aren’t edge cases. They’re the norm. And they’re the reason so many AI initiatives stall between pilot and production.

“Agentic AI changes what observability has to do. You’re no longer watching one model respond to one prompt. In agentic AI, every transaction can be unique, and you’re tracing chains of autonomous decisions across dozens of tools and services. That’s the problem Dynatrace was built to solve and Dell AI Factory is exactly the foundation enterprises need to take AI to production at scale.”

— Steve Tack, Chief Product Officer, Dynatrace

Scale AI workloads with confidence

Dynatrace can be integrated into Dell AI Factory environments to cover end-to-end observability of agentic AI and LLM workloads. The goal is straightforward: no blind spots, no surprises, and no manual investigation when something goes wrong. Here’s what that looks like in practice:

  • Unified AI observability to monitor the AI stack. Prompts, Model calls and downstream services, in a single platform that replaces the fragmented tooling most teams rely on today.
  • Automated prevention and remediation with Dynatrace Intelligence®. When AI workloads behave unexpectedly, Dynatrace Intelligence detects anomalies in real time and triggers automated remediation to minimize or eliminate downstream consequences.
  • End-to-end agentic AI tracing. Distributed tracing across multi-step agent chains, tool calls, RAG pipelines, and external integrations gives teams visibility into how AI agent decisions are made and where they go wrong.
  • Automatic topology mapping with Smartscape®. Maps every component in your Dell AI Factory environment, showing in real time how infrastructure, services, and AI models depend on and affect each other.
  • Built-in data governance and audit trails. Track data flows, model decisions, and AI service behavior with governance capabilities designed for regulated industries not retrofitted to them after the fact.
  • Faster resolution with Dynatrace Assist. Natural language querying and AI-generated remediation recommendations help operations teams resolve issues faster, even without deep AI infrastructure expertise.

Built for the industries where AI is becoming mission critical

AI is no longer an experiment. It’s become core infrastructure for the world’s most demanding enterprises, embedded in the decisions, workflows, and customer experiences that keep businesses running. When AI is mission critical, a failure isn’t a learning opportunity; it’s a negative business impact. Tolerance for poor visibility, unexplained latency, or untraceable decisions drops to zero. That’s precisely where Dynatrace AI Observability comes in, giving teams the visibility, control, and real-time intelligence to keep AI running when it matters most.

“The enterprises winning with AI aren’t running one model in one department. They’re operationalizing AI across the business. Dynatrace joining the Dell Technologies AI Ecosystem Program gives those customers the observability foundation to expand AI workloads on Dell infrastructure with the reliability, governance, and efficiency that enterprise-scale demands.”

— Brad Maltz, Senior Director of AI Solutions, Dell Technologies

What this means for joint customers

For organizations deploying on Dell AI Factory infrastructure, the combination of Dell’s validated hardware and software stack with Dynatrace’s intelligent observability platform means:

  • Scale with confidence. Expand production AI across the business without losing visibility or control.
  • Higher AI reliability. Proactive anomaly detection surfaces issues early; moving teams from reactive firefighting to confident operations.
  • Lower risk at scale. Broad stack visibility reduces the unknowns that make executive teams cautious in moving AI to production at scale.
  • Improved ROI on AI investment. When AI workloads run efficiently and every GPU hour is visible, teams can continuously optimize performance and cost.

End-to-end observability isn’t a nice-to-have for AI. It’s a prerequisite for trust, and trust is what turns AI investments into business outcomes. We’re proud to bring that capability to the Dell AI Factory ecosystem, and we’re excited about how this deepening of our relationship with Dell can unlock incredible value for our joint customers on their AI journeys.

Learn more about Dynatrace AI observability today, or reach out to your Dynatrace account team.

The post Scaling enterprise AI with confidence: Dynatrace joins the Dell Technologies AI Ecosystem Program appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/scaling-enterprise-ai-with-confidence-dynatrace-joins-the-dell-technologies-ai-ecosystem-program/feed/ 0
2026 GigaOm Radar Report for Kubernetes Observability https://www.dynatrace.com/gigaom-radar-for-kubernetes-observability/ Thu, 07 May 2026 12:00:25 +0000 https://www.dynatrace.com/news/?post_type=analyst-reports&p=66593 The post 2026 GigaOm Radar Report for Kubernetes Observability appeared first on Dynatrace news.

]]>
The post 2026 GigaOm Radar Report for Kubernetes Observability appeared first on Dynatrace news.

]]>
Dynatrace expands AI Coding Agent monitoring for Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK https://www.dynatrace.com/news/blog/dynatrace-expands-ai-coding-agent-monitoring/ https://www.dynatrace.com/news/blog/dynatrace-expands-ai-coding-agent-monitoring/#respond Thu, 30 Apr 2026 14:39:57 +0000 https://www.dynatrace.com/news/?p=73871 Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK

AI coding agents are a core part of how modern engineering teams build, review, deploy, and troubleshoot software. But as usage grows, so do the operational questions: Which agents are being adopted? What are the associated costs? How reliable are coding agents within real developer workflows? Which tools do they invoke, and where are they slowing down, failing, or creating unnecessary risk in production?

The post Dynatrace expands AI Coding Agent monitoring for Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK appeared first on Dynatrace news.

]]>
Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK

Dynatrace helps you answer these questions by extending AI observability for a new wave of coding agents, including Claude Code, Google Gemini CLI, OpenAI Codex CLI, OpenCode, and GitHub Copilot SDK. Together, these integrations give engineering leaders, platform teams, and developers a consistent way to understand agent activity, token consumption, costs, tool behavior, and runtime impact: without forcing teams to stitch together fragmented telemetry across terminals, SDKs, dashboards, and development workflows. Dynatrace public AI agent instrumentation examples on GitHub demonstrate how to provide industry leading observability that drives performance, cost efficiency, and governance across complex, distributed AI-driven systems—all through a unified Dynatrace platform experience that developers can access directly via MCP without leaving their IDE.

From agent activity to engineering insight

As organizations adopt multiple coding agents, new adoption challenges emerge. One team might use Claude Code in the terminal, another may build internal tools with GitHub Copilot SDK, while others experiment with Gemini CLI or Codex CLI. Platform teams want visibility into usage, availability, and costs. Engineering leaders want to know whether agents improve delivery. Security and governance teams want confidence that all prompts, tool usage, and actions can be monitored appropriately.

Dynatrace provides a practical answer to these challenges: a single observability layer for agile development workflows. For agents that emit OpenTelemetry directly, such as Claude Code, Gemini CLI, and Codex CLI, Dynatrace can ingest telemetry related to sessions, tokens, costs, tool executions, errors, and performance. For GitHub Copilot workflows, Dynatrace adds production context, software delivery automation, and GitHub-based integrations that connect agent activity to real engineering workflows.

The payoff is clear. Developers gain visibility into how agents behave in real work. Platform teams can track adoption, usage trends, and cost signals. Engineering leaders can correlate agent activity with commits, pull requests, and delivery outcomes. And with an MCP-enabled production context, teams can connect coding-agent actions to what is happening in production.

“Before we instrumented Claude Code, we had no easy way to break down how our engineers actually used AI, which models, for what tasks, and at what cost. Now we can pinpoint inefficient model use and guide usage toward better cost-performance tradeoffs.”
— Markus Heimbach, Senior Director Software Development

Anthropic Claude code monitoring dashboard in Dynatrace

Multiple coding agent experiences, one observability strategy

Each coding agent has a different operating model, which is why a common observability layer matters.

Claude Code

Claude Code already supports built-in OpenTelemetry, making it easy to send metrics and logs to Dynatrace with no code changes. Teams can track sessions, tokens, costs, tool activity, API health, and engineering output such as commits and pull requests. Logs, dashboards, and alerts help teams investigate failures, spot latency spikes, and catch unusual spend or error patterns early.

Gemini CLI

Gemini CLI includes OpenTelemetry-based observability and preconfigured dashboards, making it a strong fit for Dynatrace AI observability. Teams can correlate agent activity with broader platform signals and move quickly from raw telemetry to action. This includes debugging failed runs, identifying slow or error-prone tool calls, and alerting on cost or reliability regressions.

Codex CLI

Codex CLI supports opt-in OpenTelemetry monitoring, giving teams a path to audit usage and strengthen governance across CLI, IDE, and app experiences. With Dynatrace, logs and traces help investigate request flows, delays, and failures across agent workflows. Alerts can flag degraded reliability, unexpected behavior, or rising token consumption before they become larger issues.

GitHub Copilot SDK

GitHub Copilot SDK lets teams embed agentic workflows directly into applications, while Dynatrace adds live observability and security context. This matters because embedded agents become part of real engineering and production-adjacent workflows. Dynatrace helps trace execution paths, use logs for debugging and auditability, and set alerts for failures, latency, or policy-relevant events.

OpenCode

OpenCode is a terminal-based AI coding agent that helps developers work through coding tasks directly from the command line. Because OpenCode ships with native OpenTelemetry support, teams can route telemetry to Dynatrace without code changes by setting standard OTLP environment variables. With Dynatrace, teams can track LLM call volume, session activity, tool usage, request latency, and workflow behavior across real developer sessions. Traces help teams inspect LLM requests, tool executions, session lifecycle events, message processing, file snapshots, or diff operations.

Across all operating models, the value is the same: one strategy for monitoring adoption and impact, understanding costs, logging and tracing agent activity, alerting on reliability issues, and debugging real-world workflows as coding agents scale across the enterprise.

Distributed Tracing dashboard in Dynatrace

Why this matters now

Teams are no longer asking whether coding agents are useful. They’re asking how to drive adoption, scale them safely, govern them consistently, and prove their impact. That requires visibility into usage, cost, reliability, and engineering outcomes across teams and tools. Dynatrace helps organizations make that shift with the observability and production context needed to expand coding-agent adoption with confidence.

The coding-agent market is moving fast. Claude Code, GitHub Copilot SDK, Google Gemini CLI, and OpenAI Codex CLI each represent a different path toward agentic software delivery, from terminal-based workflows to embedded SDKs and governed local execution. At the same time, Dynatrace has been expanding its developer-facing AI surface with the Dynatrace MCP Server, GitHub Copilot integrations, and AI observability capabilities built to connect agent behavior with real production systems. The timing matters because teams are no longer evaluating whether coding agents are useful. They’re deciding how to drive adoption, scale up usage safely, govern usage consistently, and measure real impact.

Prompt activity dashboard in Dynatrace

Ready to see AI coding agents through a Dynatrace lens?

With Dynatrace, teams can understand adoption, spend, reliability, tool behavior, and engineering outcomes in one place, while giving agents access to the live production context they need to make better decisions.

Whether your developers are working in Claude Code, building on GitHub Copilot SDK, experimenting with Gemini CLI, or adopting Codex CLI, Dynatrace helps bring observability, governance, and production awareness into the heart of agentic software delivery.

Public examples already demonstrate this approach for Claude Code, and the broader Dynatrace MCP and AI observability ecosystem provides the foundation to extend the same value across the next generation of coding agents.

Ready to learn more?

In our Git repository, you’ll find step-by-step examples for supported coding-agent workflows, including how to configure OpenTelemetry export, send telemetry data to Dynatrace, and use the provided dashboards to analyze the activity of your AI coding agents.

Visit our Git repo for detailed instructions and AI Coding Agent instrumentation examples

The post Dynatrace expands AI Coding Agent monitoring for Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-expands-ai-coding-agent-monitoring/feed/ 0
Observability is a team sport https://www.dynatrace.com/news/blog/observability-is-a-team-sport/ https://www.dynatrace.com/news/blog/observability-is-a-team-sport/#respond Wed, 29 Apr 2026 16:23:20 +0000 https://www.dynatrace.com/news/?p=73858 Dynatrace and OpenTelemetry

“How do I structure my observability team?” is one of the most common questions folks leading software teams ask me. My advice: Don’t create a centralized “observability team” that’s responsible for all the observability within an organization. Observability shouldn’t exist as a silo. It touches many parts of an organization, from development to production, and […]

The post Observability is a team sport appeared first on Dynatrace news.

]]>
Dynatrace and OpenTelemetry

“How do I structure my observability team?” is one of the most common questions folks leading software teams ask me. My advice: Don’t create a centralized “observability team” that’s responsible for all the observability within an organization.

Observability shouldn’t exist as a silo. It touches many parts of an organization, from development to production, and should be treated as a team sport.

As we know, our systems can only be considered observable if they emit telemetry. No data means that we can’t understand what is happening in our systems. Fortunately, the OpenTelemetry® (OTel) ecosystem from the Cloud Native Computing Foundation (CNCF) has become the de-facto standard for instrumenting, generating, collecting, and exporting telemetry data.

What does this mean for observability adoption in an organization? Let’s dig in.

Observability is everyone’s responsibility

Reliability can’t happen without observability. Observability must be looked at holistically. It is not the sole responsibility of any one team or individual. Everyone has an important part to play, and to a certain extent, the parts weave into each other.

Instrumenting code

There are two types of OpenTelemetry instrumentation:

Code-based instrumentation should be done by application developers, and not by an “observability team.” Developers know their applications best. Asking someone else to instrument your application is like asking someone else to write your code comments. Please never do that.

Zero-code instrumentation usually involves a shim or bytecode instrumentation wrapper around your code. If you’re a developer writing code in a language that supports OpenTelematry auto-instrumentation, you should understand how to implement both zero-code and code-based instrumentation. In doing so, you can use the instrumentation to troubleshoot your own code.

In some environments, zero-code instrumentation may be managed by the OTel Operator. If this is the case, the responsibility often falls to SRE or platform engineering teams. Event in those cases, developers should understand; at least at a high level, how zero-code instrumentation is configured with the OTel Operator.

Managing observability infrastructure

Observability infrastructure still needs to be managed, whether you’re using a SaaS vendor (e.g. Dynatrace) or an open source stack. If you’re using OpenTelemetry, chances are you’re managing at least one OTel Collector, and perhaps many. If you’re running your applications on Kubernetes, you’ll likely deploy and manage Collectors within the cluster as well. In most organizations, this responsibility falls under platform engineering or SRE teams, and these teams are essential to robust, reliable software delivery in large, complex environments.

That said, developers should still understand how the OpenTelemetry Collector is configured. It’s true that you don’t need to go through a Collector to send OTel data to an observability backend for non-production. However, the Collector still offers some nice things that direct-from-application doesn’t (e.g. batching data, masking data, and automatic retries), and I still highly recommend using it, even in development.

Making CI/CD pipelines observable

DevOps engineers can’t escape observability either, because guess what? We can make CI/CD pipelines observable too. While CI/CD pipelines may not be a production environment that external users interact with, they most certainly are a production environment that internal users interact with (i.e. software engineers, platform engineers, and SREs).

CI/CD pipelines are defined by code, and like it or not, that code can still fail. Making our application code observable helps us make sense of things when they fail in production. So, it stands to reason that having pipeline observability can help us understand what’s going on when CI/CD pipelines fail.

There’s been some great buzz around the observability of CI/CD pipelines, especially now that there’s an official OTel CI/CD Special Interest Group (SIG). This will give our favorite CI/CD tools a shared language for the observability of CI/CD pipelines, creating a foundation for them to support OpenTelemetry tools in this context.

We’re not there yet, which means that right now we must stitch a few tools together to achieve CI/CD observability. Fortunately, things are moving nicely in this space, and if you haven’t considered CI/CD pipeline observability in your organization before, now’s the time to start thinking about it. To learn more about what’s happening with OTel CI/CD observability, check out the #otel-cicd channel on CNCF Slack.

Troubleshooting

The beauty of observability is that once you instrument your code, you put the ability to troubleshoot in the hands of many. Consider the ripple effect when developers instrument their code:

  • Developers: Instrumentation allows developers to debug their code as they’re writing it.
  • QA testers: Instrumentation allows testers to troubleshoot failed tests, allowing them to file more detailed bug reports. If QAs can’t track down the issue, then it means that there is missing instrumentation that developers need to add to their code. This turns observability into a quality gate.
  • SREs: Instrumentation allows SREs to troubleshoot production issues, gain insight into system performance, and ensure overall system reliability.

Ensuring adherence to observability practices

Remember how I advised against creating an “observability team” responsible for all observability within an organization? I still stand by that. That said, I do believe that organizations should have an observability team responsible for enterprise-wide observability oversight and advocacy. A team that defines and disseminates observability standards and practices within that organization. This team would need to stay up to date in the latest observability practices, vendor offerings, and the OpenTelemetry  ecosystem— not just as an observer, but also as a project contributor, while also encouraging developers, platform engineers, and SREs to contribute.

This “observability practices team,” can’t, however, exist on an island. First off, it needs to be aligned with leadership to ensure that everyone is on the same page when it comes to observability. The team also needs support from individual practitioners. As a result, the team also needs to work with developers, SREs, platform engineers, QAs, and DevOps engineers to ensure that the practices and standards that it comes up with make sense.

If observability is to be a team sport, it needs coordination and guidance. There should be guardrails in place, to ensure that you have standard tooling, practices, and enforcement of said practices. Practices and standards include things like standard Collector configurations, and standard attributes emitted to your chosen observability backend(s).

Standardizing tooling is important because I’ve seen far too many “tool jungles” in organizations, where each team or department has their own tooling and practices, and it ends up being a recipe for disaster. Too much redundancy and overlap.

In addition, the observability practices team should not be responsible for instrumenting developers’ code, nor should it be managing infrastructure. It’s there to work with these other groups and to make sure that things are done right.

Final thoughts

Observability weaves its way into various aspects of an organization. It’s not just a developer concern. It’s not just an SRE concern. It’s not just a QA concern. It’s certainly not the concern of a single “observability team.” Doing so downplays its importance, takes away our collective responsibility towards observability, and dilutes the promise of observability. The only way to make this work is by ensuring that the teams participating in this team sport that we call observability don’t operate in silos.

The post Observability is a team sport appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/observability-is-a-team-sport/feed/ 0
Dynatrace for AI: Teach your AI coding agent how to use Dynatrace https://www.dynatrace.com/news/blog/dynatrace-for-ai-teach-your-ai-coding-agent-how-to-use-dynatrace/ https://www.dynatrace.com/news/blog/dynatrace-for-ai-teach-your-ai-coding-agent-how-to-use-dynatrace/#respond Thu, 23 Apr 2026 16:58:48 +0000 https://www.dynatrace.com/news/?p=73813 Agentic ecosystem

Introducing Dynatrace for AI, an open-source collection of agent skills and prompts that give any skills-compatible AI coding assistant the domain expertise it needs to work productively and accurately with Dynatrace.

The post Dynatrace for AI: Teach your AI coding agent how to use Dynatrace appeared first on Dynatrace news.

]]>
Agentic ecosystem

If you’ve already wired an AI coding assistant up to Dynatrace, through the MCP server, the Dynatrace CLI (dtctl), or a custom agent you built yourself, you’ve seen your agent have difficulty interpreting data or calling for fields that don’t exist. This makes sense, your agent may be making assumptions based upon training that isn’t relevant. It lacks the skills to understand how to get the best value from Dynatrace. That is where Dynatrace for AI fills the gap.

What are agent skills?

Agent skills are an open format for packaging domain knowledge that AI agents can load on demand. A skill is a folder containing a SKILL.md file with focused instructions, examples, and optional reference material. Compatible agents, such as Claude Code, GitHub Copilot, Cursor, Cline, or others, discover installed skills and load the full content only when it’s relevant to the task at hand.

The net effect: you can install dozens of skills without bloating an agent’s context window. Agents pull in exactly what’s relevant when it’s relevant, and ignore the rest.

Install Dynatrace agent skills via a terminal
Figure 1. Install Dynatrace agent skills via a terminal

Built for agents working with Dynatrace

Dynatrace for AI is a curated set of skills that give an agent the three things it needs to efficiently do real work on Dynatrace:

  • Access to Dynatrace data and insights: through DQL queries against Grail®, Smartscape® dependency graph, or problem records.
  • Dynatrace expertise: the syntax rules, entity-model distinctions, and query patterns that separate a working query from one that looks correct but returns nothing.
  • Task-level starting points: ready-made prompt templates for common engineering workflows, so teams don’t have to invent the approach from scratch.

Skills don’t connect to Dynatrace directly. You have to pair them with the MCP server or dtctl to perform live queries and initiate actions. Together, they turn an agent with generic observability intuition into one that easily extracts value from Dynatrace.

Complement your agent with domain expertise

The first release of Dynatrace for AI agent skills is focused on the workflows that engineering teams run every day:

  • DQL fundamentals: covering the pipeline model, core data objects, and when to use fetch, timeseries, or smartscapeNodes to prevent failures that typically come from models trained on generic query-language data.
  • Observability across the stack: services, traces, logs, frontends, and problems, each covering the entity model, key fields, and query patterns that make answers correct rather than merely plausible.
  • Infrastructure and cloud: covering Kubernetes, AWS, and hosts.
  • Platform tasks worth delegating: providing programmatic creation of dashboards and notebooks

Prompt templates for common workflows

Alongside the skills, the repo hosts a small set of prompt templates you can use as structured starting points to invoke the right skills for specific tasks. These save teams from having to design their approach from scratch and make outcomes more consistent across agents and users.

Current templates include:

  • Performance regression: walks the agent through comparing RED metrics before and after a deployment, correlating any regression with distributed traces, and summarizing the root cause.
  • Daily standup: pulls the last 24 hours of problems, deployment activity, and notable anomalies for a team’s services, so anyone can walk into a standup with the relevant production context already framed.
  • Troubleshoot a problem: takes a problem ID and guides the agent through root-cause analysis, including affected entities, correlated events, relevant logs and traces, and creates a structured summary for the incident channel.

These are a starting point, not a ceiling, designed to be forked and shaped to your team’s on-call runbooks.

What Dynatrace for AI is and what it isn’t

Skills and prompts are a knowledge and workflow layer. They don’t connect to your Dynatrace environment, define what actions your agent can take, or set guardrails. That’s the job of the tool you pair them with and your Dynatrace permission model.

The quality of what your agent can produce also depends on the entities your environment is instrumented to capture. Skills help agents ask better questions of data, but they don’t control what data is collected.

Think of this skill as onboarding a smart new hire who already knows software, but needs to learn your platform. The skills are the platform user guide; your observability data is the work itself.

Get started

It’s super simple to install the skills and prompts in one go. Just run:

npx skills add dynatrace/dynatrace-for-ai

…or activate the skills as a Claude Code plugin:

claude plugin marketplace add dynatrace/dynatrace-for-ai
claude plugin install dynatrace@dynatrace-for-ai

Make sure your agent can reach Dynatrace, then try a real agent-skill task. A few good example starting prompts:

  • “Compare the error rate of the checkout service over the last hour vs the same hour yesterday.”
  • “Are any pods in the production namespace restarting or getting OOM-killed right now?”
  • “Use the performance-regression prompt to check the deployment I just shipped.”

The difference in output quality is immediate: fewer corrections, cleaner queries, and answers that accurately reflect how Dynatrace continuously models your environment in real-time.

The Dynatrace for AI project is open source and actively developed. Issues, discussions, and pull requests are all welcome, especially from teams running agent skills against real workloads. We’d love to hear from you.

Make your agents work smarter; teach them how to use Dynatrace.

The post Dynatrace for AI: Teach your AI coding agent how to use Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-for-ai-teach-your-ai-coding-agent-how-to-use-dynatrace/feed/ 0