GenAI | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 19 Mar 2026 13:45:23 +0000 en hourly 1 Announcing agentic framework support and General Availability of the Dynatrace AI Observability app https://www.dynatrace.com/news/blog/announcing-agentic-framework-support-and-general-availability-of-the-dynatrace-ai-observability-app/ https://www.dynatrace.com/news/blog/announcing-agentic-framework-support-and-general-availability-of-the-dynatrace-ai-observability-app/#respond Wed, 28 Jan 2026 16:55:26 +0000 https://www.dynatrace.com/news/?p=72664 Agentic ecosystem

As agentic AI becomes mission-critical, systems that reason, act, and self-optimize introduce new operational challenges. Their dynamic and non-deterministic behavior makes them difficult to debug, they can drive unexpected cost spikes, and they inherently lack the auditability required for reliable, enterprise-grade use. Today, we’re excited to announce expanded support for leading agentic frameworks and protocols, […]

The post Announcing agentic framework support and General Availability of the Dynatrace AI Observability app appeared first on Dynatrace news.

]]>
Agentic ecosystem


As agentic AI becomes mission-critical, systems that reason, act, and self-optimize introduce new operational challenges. Their dynamic and non-deterministic behavior makes them difficult to debug, they can drive unexpected cost spikes, and they inherently lack the auditability required for reliable, enterprise-grade use. Today, we’re excited to announce expanded support for leading agentic frameworks and protocols, along with a new dedicated AI Observability app. With this support, you can build, run, and debug agentic AI applications with confidence across AWS, Azure, and Google Cloud.

What’s new: Broader agentic technology support

Dynatrace supports a broad and rapidly growing ecosystem of agentic AI frameworks and protocols, unifying telemetry from these frameworks via OpenTelemetry and OpenLLMetry into a single, correlated observability model, delivering end‑to‑end visibility across clouds, models, tools, and agents from one platform.

  • Amazon Bedrock AgentCore – Dynatrace offers observability for Amazon Bedrock AgentCore agents by collecting metrics such as token usage, model behavior, latency, and errors. This integration provides unified tracing, cost, performance, and guardrail monitoring, along with ready-made dashboards and intelligent anomaly detection and forecasting, helping teams quickly and effectively monitor, troubleshoot, and optimize complex autonomous agent workflows.
  • Amazon Bedrock Strands – Dynatrace supports the Amazon Bedrock Strands Agents SDK, enabling comprehensive visibility into agentic AI systems. By instrumenting Strands-based AI agents with Dynatrace, organizations can monitor agent behavior, tool usage, and dependencies end to end. This helps ensure performance, reliability, and operational insight across distributed environments, supporting the confident development and operation of agentic AI use cases such as chatbots, recommendation systems, and autonomous workflows.
  • LangChain Agents – Dynatrace provides observability for applications built with the LangChain framework, enabling the monitoring of performance, cost, and reliability of Large Language Model (LLM) applications and agents.
  • Google Agent Development Kit (ADK) – Dynatrace provides observability for applications built with the Google Agent Development Kit (ADK), enabling visibility into agent execution, dependencies, and performance. This helps teams understand runtime behavior and maintain reliability as agent-based applications
  • OpenAI Agents SDK – Dynatrace provides observability for observing applications built with the OpenAI Agents SDK, enabling monitoring of agent workflows, model interactions, latency, and errors. This supports improved operational insight, troubleshooting, and performance optimization for agentic AI applications.
  • MCP AI Agent–  Dynatrace provides deep visibility into AI agents communicating via the Model Context Protocol (MCP). By observing both AI agents and MCP servers, organizations gain end-to-end insight into execution flows through tracing, enabling data-driven decisions, performance and cost optimization, and governance for complex agent workflows.
Agentic AI Observability for popular agentic frameworks, powered by OpenTelemetry and OpenLLMetry
Figure 1. Agentic AI Observability for popular agentic frameworks, powered by OpenTelemetry and OpenLLMetry

This agentic coverage is on top of the 40+ LLM technologies that Dynatrace already supports, including OpenAI, Amazon Bedrock, Google Gemini and Vertex, Anthropic, LangChain, NVIDIA, and more.

We’re working closely across AWS, Microsoft Azure, and Google Cloud ecosystems to ensure you have consistent, enterprise‑grade observability for your multi‑AI and multi‑cloud applications.

See it in action in the new AI Observability experience

The AI Observability app is now Generally Available, delivering a purpose-built experience for observing AI workloads end-to-end from agents and LLMs to orchestration layers, emerging protocols, and tools. It gives engineering teams deep, production-ready visibility into how AI systems behave in real time, allowing them to validate changes faster, reduce risk, and confidently ship AI-powered features at scale.

Unlike generic observability views, the AI Observability app is designed specifically for agentic and LLM-driven systems, making it easy to understand complex multi-step interactions, reason about cost and performance trade-offs, and troubleshoot issues across models, tools, and dependencies.

Key capabilities

  • End‑to‑end observability for agentic AI
    • Monitor agent interactions, tool usage, dependencies, latency, and reliability
    • Track token consumption, cost trends, and caching impact
  • Tracing and debugging for complex flows
    • Follow prompts, tool calls, and model invocations from the initial request to the final response
    • Jump from high‑level health to prompt‑level traces in a couple of clicks
  • Actionable insights at scale
    • Rapid A/B testing across model and prompt variants for faster validation
    • Identify bottlenecks and optimize resource utilization with ready‑made dashboards and drill‑downs
  • Security, privacy, and governance
    • Enterprise‑grade controls, auditability, and policy‑aligned routing
    • Guardrail outcomes (for example, toxicity, PII, or denied topics) are surfaced so you can monitor behavior and trends. (Note that guardrail enforcement occurs at the model/provider; Dynatrace captures and visualizes provider‑reported outcomes.)
The Dynatrace AI Observability experience.
Video 1. The Dynatrace AI Observability experience.

Who this solution is for and why it matters

The Dynatrace AI Observability solution is for enterprise teams, including developers, DevOps, SREs, and business leaders who need deep, real-time insights into their cloud native  AI-powered applications and customer experience in a single unified view.

Who benefits the most from this solution?

  • AI Engineering and Data Science: This group includes practitioners who develop and optimize models. They use LLM observability to track metrics related to model performance, such as identifying hallucinations and biases, validating changes, and improving prompt engineering practices.
  • Software Developers: These individuals benefit from observability by gaining insights into application-level performance, which helps them debug and improve overall code quality. Observability tools allow for faster iteration in development cycles.
  • Site Reliability Engineers (SRE): These teams ensure the reliability and performance of AI applications in production environments. They use observability to identify system-level bottlenecks and failures, and to respond swiftly to operational challenges.
  • Application Security Teams: Although not traditionally the primary users, security teams can leverage AI observability to identify and mitigate emerging threats specific to AI applications, such as prompt-injection attacks and data leaks.
  • Compliance and Governance Teams: Responsible for ensuring adherence to regulatory requirements and internal policies, these teams rely on observability to audit model behavior and to identify potential biases or harmful outputs.

What’s next: Agent topology view with Smartscape

We’re committed to further enhancing these capabilities. As agentic systems evolve into distributed networks of models, tools, and decisions, observability must move beyond traces and metrics. Our next focus is the Agentic Topology View, bringing Smartscape-grade visualization to agent execution flows so teams can see how agents interact, invoke tools, propagate errors, and improve performance end to end.

This agentic topology becomes the foundation for a deeper developer experience by connecting production telemetry with prompt management and evaluation workflows. By unifying agent topology, prompt lifecycle, and LLM-as-judge scoring in a single system, we’re helping teams systematically improve the reliability, performance, and quality of agentic AI at enterprise scale.

Agent topology visualizes agent execution flows, showing how they interact with one another.
Video 2. Agent topology visualizes agent execution flows, showing how they interact with one another.

Get started today

Want to “kick the tires” with some example code? Let’s make agentic AI observable, governable, and reliably fast.

The post Announcing agentic framework support and General Availability of the Dynatrace AI Observability app appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/announcing-agentic-framework-support-and-general-availability-of-the-dynatrace-ai-observability-app/feed/ 0
Optimizing AI ROI from DevOps and IT Operations: The rising need for AI/LLM observability https://www.dynatrace.com/news/blog/optimizing-ai-roi-from-devops-and-it-operations/ https://www.dynatrace.com/news/blog/optimizing-ai-roi-from-devops-and-it-operations/#respond Wed, 03 Dec 2025 18:09:08 +0000 https://www.dynatrace.com/news/?p=72110 Blog thumbnail

Every organization is adopting GenAI across its infrastructure and application stacks. It’s important that IT operations teams seek a seat at the table because large swaths of models will be deployed across every technology. For example, the use of cloud migrations, GenAI large language models, small language models, and specialized models will drive productivity, cost […]

The post Optimizing AI ROI from DevOps and IT Operations: The rising need for AI/LLM observability appeared first on Dynatrace news.

]]>
Blog thumbnail

Every organization is adopting GenAI across its infrastructure and application stacks. It’s important that IT operations teams seek a seat at the table because large swaths of models will be deployed across every technology. For example, the use of cloud migrations, GenAI large language models, small language models, and specialized models will drive productivity, cost savings, and business returns. Every customer is considering and attempting to measure their business returns from their AI investments; transparency into the data, system and model performance and drift, security, and quality are critical areas where IT operations, DevOps, SREs, and platform engineering teams can play a critical role in optimizing business returns and reducing business risks. So, where should you start the conversation?

Executives can use observability to reduce business risks and increase AI ROI by understanding how observability capabilities play a role in delivering across the core AI value categories of productivity, customer impact, cost optimization, innovation, and quality. For example, observability improves customer satisfaction by reducing the mean time to resolution and mean time to understanding. In addition, it can improve cross-team collaboration and data access to deliver cost efficiencies.

To reduce business risks and increase ROI in GenAI use cases, technology executives should plan to manage rising complexity, and, as part of continuous evaluation, executives should consider GenAI performance across the following dimensions:

  • System performance: Monitoring the system performance of GenAI applications encompasses measuring operational performance characteristics similar to those of traditional applications, including at the software and infrastructure layers and the model. Model system performance monitoring includes the measurement of metrics such as model response latency, error rates (including failure to respond), and API failures.
  • Quality performance: It is crucial for organizations to monitor the output quality of GenAI and AI applications. Quality includes accuracy of responses and model drift, where data used to train models no longer produces accurate or relevant results.
  • Governance: Model governance of GenAI often encompasses monitoring and enforcing legal requirements and the organization’s ethics policies. Ongoing monitoring is necessary, including the adoption of guardrails to prevent the delivery of outputs that don’t comply with laws or company policies.
  • Security: In addition to the theft of private information or loss of intellectual property, organizations must protect against security risks that are specific to GenAI applications. Prompt injection and jailbreaks are two emerging attacks. Monitoring tools that detect these and other security issues are critical to risk management.
  • Cost: Monitoring the cost of delivering a GenAI application is a multitiered undertaking. Depending on the application, organizations may incur costs for each query and response to a model, in addition to costs associated with the underlying infrastructure required to deliver the application. The ability to collect the right cost information and analyze it on a per-application basis will be key to the ability of an organization to determine ROI.

Organizations must base the measurement of each performance dimension on its ability to derive outcomes that drive business value. Each GenAI application should support a targeted outcome, such as improved productivity, increased revenue, new revenue streams, or enhanced customer satisfaction. Connecting the dots between GenAI performance dimensions and business value requires defining measurements that matter to the business and collecting, correlating, and analyzing the data to understand the app’s ability to deliver that value.

For technology executives, AI observability is fast becoming essential for managing the operational complexity and business outcomes from AI initiatives. It provides the visibility needed to demonstrate ROI, ensure reliable AI applications, and make informed decisions based on critical data that supports every AI use case.

Monitor, optimize, and secure Generative AI applications, LLMs, and agentic workflows — improving performance, explainability, and compliance.

Learn more, or try Dynatrace for free!

The post Optimizing AI ROI from DevOps and IT Operations: The rising need for AI/LLM observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/optimizing-ai-roi-from-devops-and-it-operations/feed/ 0
AWS publishes Dynatrace-developed blueprint for secure Amazon Bedrock access at scale https://www.dynatrace.com/news/blog/aws-publishes-dynatrace-developed-blueprint-for-secure-amazon-bedrock-access-at-scale/ https://www.dynatrace.com/news/blog/aws-publishes-dynatrace-developed-blueprint-for-secure-amazon-bedrock-access-at-scale/#respond Wed, 19 Nov 2025 10:00:20 +0000 https://www.dynatrace.com/news/?p=71904 AWS icon and agentic AI

Enterprises are rapidly expanding their use of generative AI with Amazon Bedrock to power intelligent agents and automate workflows. As adoption grows, so does the need for governance, control, and accountability. To address these challenges, Dynatrace, an early pioneer in AI at scale, has developed a robust AI gateway architecture. In collaboration with our partners at AWS, we’re now sharing this architecture as a reusable reference pattern that allows any organization to securely and efficiently control access to Amazon Bedrock services at scale.

The post AWS publishes Dynatrace-developed blueprint for secure Amazon Bedrock access at scale appeared first on Dynatrace news.

]]>
AWS icon and agentic AI

Amazon Bedrock provides enterprises with fully managed access to leading foundation models through a single API, eliminating the complexity of managing underlying AI infrastructure. This simplicity accelerates innovation but also prompts enterprises to consider how best to govern and secure access to Amazon Bedrock as they’re using it at scale.

Without a secure AI gateway in place, organizations can quickly face challenges such as:

  • Uncontrolled access and data exposure: Without integrated authentication and authorization, anyone with credentials can invoke models or send sensitive data without oversight.
  • Compliance and audit gaps: Without consistent tracking and isolation, it’s difficult to demonstrate adherence to internal policies or regulatory requirements.
  • Operational fragility: Developers must manage credentials and request signing manually, adding complexity and security risk.

These are the same challenges Dynatrace encountered while scaling its own generative AI workloads. In response, our engineering teams developed a secure AI gateway for Amazon Bedrock, which has proven effective in serving our global user base. We’re now sharing a reusable reference architecture for the AI gateway in close collaboration with our partners at AWS.

Reference architecture of the Secure API Gateway.
Figure 1. Reference architecture of the Secure API Gateway.

Enterprise-grade governance for real-world use cases

The Secure AI Gateway extends Amazon Bedrock with enterprise-grade governance and control. Built on Amazon API Gateway, the solution integrates seamlessly into existing enterprise environments and provides:

  • Strong authentication and authorization through integration with corporate identity systems.
  • Usage quotas and throttling to manage cost and ensure fair resource distribution.
  • Multi-tenant support and tenant isolation with detailed usage tracking for security, auditability, and compliance.
  • Zero-code compatibility with Bedrock features: Once the AI Gateway is deployed, all existing Bedrock capabilities remain available without any integration code changes.

Proven within Dynatrace’s own platform, this reference pattern provides enterprises with a practical path to securely operationalize Bedrock, maintaining the speed and flexibility developers expect while introducing the control and transparency that enterprise governance demands.

Find all the details and the full technical walkthrough here: AWS: Building a Secure AI Gateway to Amazon Bedrock.

AI Observability for continuous insights after deployment

Securing access is only the first step; ensuring everything continues to work as intended is the next. With Dynatrace observability for Bedrock-based workloads, your teams gain continuous insight into performance, reliability, and cost, verifying that governance controls remain effective and that AI workloads perform as expected.

You can read more about our solution here: Deliver secure, safe, and trustworthy GenAI applications with Amazon Bedrock and Dynatrace.

The post AWS publishes Dynatrace-developed blueprint for secure Amazon Bedrock access at scale appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/aws-publishes-dynatrace-developed-blueprint-for-secure-amazon-bedrock-access-at-scale/feed/ 0
Announcing Amazon Bedrock AgentCore Agent Observability https://www.dynatrace.com/news/blog/announcing-amazon-bedrock-agentcore-agent-observability/ https://www.dynatrace.com/news/blog/announcing-amazon-bedrock-agentcore-agent-observability/#respond Tue, 18 Nov 2025 14:00:07 +0000 https://www.dynatrace.com/news/?p=71891 Dynatrace and Amazon Bedrock AgentCore

Dynatrace now provides native, end-to-end observability for Amazon Bedrock AgentCore agents, delivering unified tracing, cost and latency analytics, and guardrail monitoring out of the box. By ingesting OpenTelemetry signals enriched with generative AI semantic attributes, Dynatrace allows easy monitoring of agent workflows, faster troubleshooting, and more effective control over spending through intelligent anomaly detection and forecasting.

The post Announcing Amazon Bedrock AgentCore Agent Observability appeared first on Dynatrace news.

]]>
Dynatrace and Amazon Bedrock AgentCore

Teams can transition from setup to insights in minutes using a lightweight OTLP configuration and ready-made dashboards.

Unified view of AWS AgentCore service health and model performance
Figure 1. Unified view of AWS AgentCore service health and model performance

Agentic observability is evolving

Agentic AI systems are quickly moving from proof-of-concept to production, giving customers the ability to automate complex workflows, invoke a variety of different tools and APIs, and coordinate tasks across multiple services. However, traditional monitoring overlooks critical AI-specific signals, such as token consumption, model behavior, and guardrail outcomes. Teams struggle to trace non-linear agent flows, establish baselines for dynamic systems, and maintain predictable costs as usage scales. Without purpose-built observability, organizations risk degraded experiences, higher costs, and compliance gaps as agent complexity grows.

As agentic AI moves from pilot programs to production, organizations are automating complex, cross-system workflows with Amazon Bedrock AgentCore. However, most monitoring stacks weren’t designed for emergent, tool-driven behaviors and, therefore, leave blind spots around correctness, safety, and cost. Teams struggle to trace non-linear flows, establish baselines for dynamic systems, build agentic workflows, and keep token-driven spend under control as usage scales.

The observability gap in AI agent deployments

While AI agents offer significant benefits, including improved employee productivity, increased efficiency, and competitive advantage, among others, an observability gap remains, creating the following challenges:

  • Complex multi-step workflows
    AI agents run non-linear, multi-system sequences with inter-agent dependencies, making data flow and responsibility hard to trace. This obscures where time is spent and who is responsible for failures in the chain.
  • Limitations of traditional metrics
    Basic operational metrics often overlook AI reasoning errors and quality issues that don’t significantly affect CPU or p95 latency. Without AI-specific telemetry, subtle degradations often slip through.
  • Continuous underlying agent and LLM model version changes
    Your system might be robust today, but upstream model and version updates can alter behavior, latency, and costs, forcing continuous adaptation to prevent regressions and incidents. Proactive detection of model-induced changes is crucial to maintaining stable quality and safety over time.
  • Scalability and quality challenges
    As deployments grow, telemetry volume and coordination overhead surge while token usage and API calls remain untracked. This breaks cost predictability and quality control, leading to issues such as hallucinations and model drift. Multi-agent logic evolves constantly, so “normal” is a moving target. Baselines drift, complicating anomaly detection and root-cause analysis.

Without addressing these challenges, organizations face risks, from degraded user experiences and spiraling costs to compliance violations and reputational damage.

New enhancements for teams building with Amazon Bedrock AgentCore

The new Dynatrace AI Observability app embeds Amazon Bedrock AgentCore observability into a dedicated end-to-end experience, featuring out-of-the-box analytics, auto-instrumentation, targeted GenAI metrics, debugging flows, and ready-made dashboards to address all observability gaps in agent deployments. Support is available for over 20 technologies, including Amazon Bedrock, OpenAI, Gemini/Vertex, Anthropic, and LangChain.

These enhancements enable teams to take advantage of the following benefits:

  • End-to-end distributed tracing
    Trace every interaction from user prompt to model reasoning to tool calls, so you can pinpoint bottlenecks, errors, or costly loops in seconds. Filter by model, provider, token usage, latency, and more to accelerate root-cause analysis.
  • Enriched GenAI telemetry data, out of the box
    Each LLM and tool invocation emits spans with prompts, completions, token counts (for both prompts and completions), finish reasons, model IDs, latency, and errors, utilizing GenAI semantic attributes. Orchestration layers (for example, actions, HTTP durations, and step names) are captured for the complete workflow context.
  • Cost, performance, and safety insights
    Use intelligent forecasting to detect cost and performance anomalies in token consumption and latency. Monitor guardrails for toxicity, PII, and denied topics to build trust and meet compliance requirements.
  • Simple OTLP setup, fast time to value
    AgentCore already emits telemetry; simply register the OpenTelemetry export to Dynatrace once. Use your Dynatrace OTLP endpoint and token, and you’re streaming signals into the Dynatrace Grail® data lakehouse with no code rewrites. Ready-made dashboards for Amazon Bedrock let you verify ingestion and gain instant insights.
AgentCore end-to-end tracing for the multi-step autonomous agent workflow, available in our GitHub repository
Figure 2. AgentCore end-to-end tracing for the multi-step autonomous agent workflow, available in our GitHub repository.

What’s next

We’re investing in a deeper Amazon Bedrock model and provider insights, expanded guardrail analytics, and additional automation so you can attach remediation playbooks to cost or safety anomalies.

Additionally, we’ll introduce a new agent visualization and topology experience that visualizes your AgentCore agents, LLM services, tool backends, and dependencies, allowing you to understand real-time topology and data flows across the entire stack.

Navigate from the topology map to traces to follow agent behavior step-by-step across services, protocols, and external calls, pinpointing hotspots, ownership, and blast radius more quickly.

Expect tighter integrations with popular orchestration frameworks and more dashboards for common agent patterns, such as retrieval, multi-agent collaboration, and tool-heavy workflows.

Get started with Dynatrace AI Observability for Amazon Bedrock AgentCore agents

Ready to learn more? Have a look at our GitHub repository.

Start instrumenting your agents today. Open the Amazon Bedrock AI Observability dashboard in Dynatrace to verify telemetry and begin your analysis.

The post Announcing Amazon Bedrock AgentCore Agent Observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/announcing-amazon-bedrock-agentcore-agent-observability/feed/ 0
Unlocking productivity and trust: Dynatrace observability in NVIDIA AI Factory https://www.dynatrace.com/news/blog/unlocking-productivity-and-trust-dynatrace-observability-in-nvidia-ai-factory-environments/ https://www.dynatrace.com/news/blog/unlocking-productivity-and-trust-dynatrace-observability-in-nvidia-ai-factory-environments/#respond Tue, 28 Oct 2025 18:30:05 +0000 https://www.dynatrace.com/news/?p=71582 Davis CoPilot for NVIDIA

The NVIDIA Enterprise AI Factory addresses the rapidly evolving needs for AI infrastructure to support the rise of agentic AI. Since its launch, customers have leveraged this validated design to build agents by following structured methodology and recommended frameworks, which simplifies deployment and configuration while facilitating the implementation of AI factories in both on-premises and […]

The post Unlocking productivity and trust: Dynatrace observability in NVIDIA AI Factory appeared first on Dynatrace news.

]]>
Davis CoPilot for NVIDIA

The NVIDIA Enterprise AI Factory addresses the rapidly evolving needs for AI infrastructure to support the rise of agentic AI. Since its launch, customers have leveraged this validated design to build agents by following structured methodology and recommended frameworks, which simplifies deployment and configuration while facilitating the implementation of AI factories in both on-premises and hybrid cloud environments.

Dynatrace has been an integral part of this initiative. Dynatrace full-stack AI and LLM observability helps organizations move forward with confidence in building their AI and agentic AI initiatives.

Observable AI: Turn a black box into a glass box to build confidence

With the publication of comprehensive guidelines, it’s simpler than ever for Dynatrace customers to set up and start monitoring their full-stack NVIDIA enterprise AI infrastructure, including its key tiers and components. Covering the infrastructure layer from GPUs to Kubernetes, NVIDIA NIM microservices, NVIDIA NeMo, and other technologies up to the application layer, Dynatrace observability enables customers to confidently run and operate complex AI workflows on NVIDIA infrastructure.

NVIDIA Enterprise AI Factory for Agents including components covered by ecosystem partners (such as Observability). Picture taken from NVIDIA Enterprise AI Factory - Design Guide White Paper
Figure 1: NVIDIA Enterprise AI Factory for Agents, including components covered by ecosystem partners (such as Observability). Picture taken from NVIDIA Enterprise AI Factory – Design Guide White Paper

In parallel, Dynatrace has worked to significantly advance our AI and LLM observability offering by introducing the following:

Dynatrace AI Observability
Figure 2: Dynatrace AI Observability

These improvements address challenges such as missing observability insights, scale, sovereignty, and trust. This empowers organizations to operationalize AI by building trust and monitoring guardrails; providing analytics capabilities to detect user-facing issues; helping SREs and AI-native engineers maintain performance, reliability, and security; and reducing cost across the agentic, AI, and LLM stack.

Privacy and security lead the way to scaling AI with confidence

AI is delivering significant productivity improvements, with 66% of senior executives reporting positive trends in productivity, according to PwC’s AI Agent Survey. This momentum is driving the demand to manage AI expenditures, enhance the decision-making quality of agents, and optimize development through visibility into AI components’ behavior in production environments — from pilot projects to full-scale operations.

However, sensitive data considerations and strict compliance requirements often impede progress, preventing organizations from fully realizing the benefits of AI adoption. As enterprises prioritize data privacy, regulatory compliance, and data sovereignty, there is an increasing need for high-performance NVIDIA AI infrastructure alongside frameworks designed to preserve control, trust, and autonomy in AI development.

In a recent blog on sovereign AI, NVIDIA shares strategies for nations and enterprises to develop AI factories that uphold local governance, security, and cultural values. Combining such factories with the Dynatrace advanced observability solution enables organizations to operationalize AI at scale — building secure and scalable agents, deployed on premises or in hybrid environments.

From privacy needs to public-sector requirements: NVIDIA AI Factory for Government

At NVIDIA GTC Washington, D.C. today, NVIDIA AI Factory for Government was announced, in support of the needs for regulated environments to drive AI initiatives. The U.S. Office of Management and Budget’s decision to establish scorecards for agencies’ AI maturity and management is in line with a 2024 Gartner Research forecast that more than 60% of government organizations will be prioritizing their investments in business automation by 2026 — up from 35% in 2022. The NVIDIA AI Factory for Government is a full-stack, end-to-end reference design that brings the power of reasoning AI to federal organizations. It helps organizations unlock productivity gains just like it does for enterprises, from service delivery to threat detection and day-to-day operations.

Built on the experience of deploying internal AI factories, the reference design offers guidance for deploying agentic AI, physical AI, and high-performance computing workloads on premises and in hybrid cloud environments, while meeting the compliance needs of federal and other secure organizations. The NVIDIA AI Factory for Government reference design includes NVIDIA Blackwell accelerated computing and NVIDIA networking, NVIDIA-Certified Systems, NVIDIA AI Enterprise software, NVIDIA Nemotron open models, and third-party software from AI leaders, all validated by NVIDIA.

Dynatrace delivers trusted observability and automation for regulated environments

Dynatrace has always been committed to supporting the public sector and other industries with regulatory requirements by providing customers with capabilities to control data flow through its lifecycle and manage sensitive data from ingestion to deletion, as well as global deployment options to meet data residency requirements, configurable retention times for different data types and use cases, unique encryption keys for customer’s stored data, and more.

Our dedication is reflected in customers’ success stories from regulated industries, as well as a growing list of global and local certifications, such as ISO 27001, SOC 2 Type II, CSA STAR 2, ENS, Tisax, and others. Find out more about our certifications and supported compliance frameworks in our Trust Center. For organizations also navigating evolving sovereignty requirements, our approach to digital sovereignty demonstrates how Dynatrace combines technical innovation with policy alignment to deliver trusted solutions globally.

Benefit from full-stack observability for end-to-end validated design

Dynatrace observability with the NVIDIA AI Factory for Government reference design enables organizations to accelerate the deployments of their AI agents and applications for federal and enterprise environments, and benefit from real-time, AI-powered insights.

These benefits range from improved scalability and performance to reduced complexity and total cost of ownership by simplifying processes, mitigating deployment risks to improved data security and compliance.

Visit the Dynatrace Playground to experience the possibilities of AI and LLM observability, and discover how Dynatrace is accelerating enterprise AI at scale.

Dynatrace and the Dynatrace logo are trademarks of the Dynatrace, Inc. group of companies. All other trademarks are the property of their respective owners.

The post Unlocking productivity and trust: Dynatrace observability in NVIDIA AI Factory appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/unlocking-productivity-and-trust-dynatrace-observability-in-nvidia-ai-factory-environments/feed/ 0
The rise of agentic AI part 7: introducing data governance and audit trails for AI services https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-7-introducing-data-governance-and-audit-trails-for-ai-services/ https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-7-introducing-data-governance-and-audit-trails-for-ai-services/#respond Tue, 14 Oct 2025 15:59:34 +0000 https://www.dynatrace.com/news/?p=71363 Dynatrace Agentic AI

Your AI investments can’t reach their potential without effective AI governance. AI governance is a challenge that demands unprecedented agility, proactive measures, and comprehensive oversight to manage complexity. With Dynatrace, you’re prepared for whatever comes next. Stay compliant and build trust in your AI systems AI regulation is tightening, and non-compliance is becoming a huge […]

The post The rise of agentic AI part 7: introducing data governance and audit trails for AI services appeared first on Dynatrace news.

]]>
Dynatrace Agentic AI
  • Your AI investments can’t reach their potential without effective AI governance.
  • AI governance is a challenge that demands unprecedented agility, proactive measures, and comprehensive oversight to manage complexity.
  • With Dynatrace, you’re prepared for whatever comes next.

Stay compliant and build trust in your AI systems

AI regulation is tightening, and non-compliance is becoming a huge risk to broader, production-scale AI adoption. Penalties are only part of the impact: reputational damage, customer mistrust, and stalled innovation can cripple even forward-looking organizations. That’s why we’re introducing data governance and audit trails for AI observability: a scalable way to manage, monitor, and secure the AI data lifecycle with end-to-end lineage, retention controls, and evidentiary records of model and user interactions.

Our platform helps you turn governance into a competitive advantage. Built-in audit support helps with emerging regulations like the EU AI Act, and alignment to industry standard frameworks such as NIST AI and ISO/IEC 42001:2023.

The hidden challenges of AI data governance

The complexity of compliance

AI regulations are becoming stricter, and new regulations are on the horizon. Organizations must maintain detailed records of AI activities for years, ensure transparency of data and processes, and align retention policies with legal requirements. These measures are imperative for trust and safety, but they introduce significant challenges. For instance, AI-related events are often scattered across multiple systems, applications, and teams, complicating efforts to create a unified audit trail. Default retention periods can fall short of regulatory needs, and manual governance processes are error-prone and infeasible at scale.

The risk of non-compliance

Failing to meet regulatory standards risks hefty fines and penalties, but the market consequences, reputational damage, and loss of customer trust are even worse. Without the right tools, organizations will struggle to manage the growing complexity of AI data governance and reap the full benefits of AI investments.

Introducing Dynatrace data governance and audit trails

Dynatrace has a long history of empowering organizations to tackle complex challenges with AI-driven solutions. Building on this expertise, we’re introducing a new set of capabilities designed to simplify compliance, enhance transparency, and streamline data management. With Dynatrace, you can:

  • Automatically retain AI-related events for up to 10 years in Grail®, our secure data lakehouse.
  • Monitor and capture events from platforms like Amazon Bedrock, tracking everything from model deployments to fine-tuning activities.
  • Leverage OpenTelemetry to collect real-time traces and metrics of AI workloads, along with every AI user interaction, giving you a complete picture of your AI ecosystem.

Data governance audit in Dynatrace screenshot

Close the compliance gap with embedded oversight

What sets Dynatrace apart is seamless integration with your existing workflows. With OpenPipeline® on Grail, you can route AI-related events to custom storage buckets with extended retention, automatically, and without forcing teams to change tools or processes. This allows long-term auditability and helps meet sector-specific compliance requirements that might require special retention and auditability measures.
Once configured, Dynatrace can automatically route and store events, creating a reliable and transparent audit trail. This helps to reduce fragmentation, tool sprawl, and manual effort traditionally associated with data governance.

Imagine being able to trace every user interaction, model training session, or deployment event with just a few clicks. Dynatrace makes this possible by consolidating fragmented data into a single, coherent view. Whether you’re responding to a regulatory inquiry or optimizing your AI models, you’ll have the insights you need, when you need them.

Simplified and instant data filtering with Dynatrace segments

Not all audit data carries the same compliance weight. For global enterprises with complex IT environments, the ability to instantly filter data by precise criteria is essential for accelerating compliance across diverse regulations, from strict local regulatory transparency obligations to lighter regimes elsewhere.

Dynatrace segments make it simple to break down and filter data to match your analysis needs and regulatory requirements:

  • Targeted compliance views: Instantly filter audit data by region, environment, platform, model, or custom criteria to align with diverse regulatory requirements.
  • Dynamic adaptability: Segments automatically update, for example, when new LLM models or environments are introduced, minimizing manual maintenance and keeping governance current.
  • Reusable assets: Leverage a single dashboard or notebook across multiple use cases by simply applying different segments, reducing duplication of effort.
  • Noise reduction: Exclude irrelevant data such as development or test logs to keep compliance and observability focused on what truly matters.
  • Custom team-context: Provide different teams (for example, compliance, data science, operations) with clear, filtered views of their audit data, ensuring ownership and audit-readiness across departments.

From observability to trusted automation

The future of AI governance lies in proactive, automated solutions that not only meet today’s regulations but also anticipate tomorrow’s challenges. With Dynatrace, you’re not just complying—you’re building a foundation of trust and reliability that scales with your business. By capturing and integrating AI events into a unified platform, Dynatrace transforms compliance from a burden into a strategic advantage.

Get started today

Ready to simplify your AI data governance?

Try it out yourself on the Dynatrace playground. Or, learn how to configure AI governance in our documentation.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two of the Rise of Agentic AI blog series explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.

The post The rise of agentic AI part 7: introducing data governance and audit trails for AI services appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-7-introducing-data-governance-and-audit-trails-for-ai-services/feed/ 0
Dynatrace achieves AWS Generative AI Competency: A new milestone in observability and AI https://www.dynatrace.com/news/blog/dynatrace-achieves-aws-generative-ai-competency/ https://www.dynatrace.com/news/blog/dynatrace-achieves-aws-generative-ai-competency/#respond Tue, 30 Sep 2025 12:11:34 +0000 https://www.dynatrace.com/news/?p=71159 Dynatrace | AWS

Enterprise adoption of generative AI is showing no signs of slowing down, and it’s easy to understand why; organizations in every vertical aim to reap its benefits, including increased efficiency, routine task automation, and content generation, ultimately creating a competitive advantage. To better help organizations maximize the benefit and full potential of generative AI, Dynatrace […]

The post Dynatrace achieves AWS Generative AI Competency: A new milestone in observability and AI appeared first on Dynatrace news.

]]>
Dynatrace | AWS

Enterprise adoption of generative AI is showing no signs of slowing down, and it’s easy to understand why; organizations in every vertical aim to reap its benefits, including increased efficiency, routine task automation, and content generation, ultimately creating a competitive advantage. To better help organizations maximize the benefit and full potential of generative AI, Dynatrace has achieved the Amazon Web Services Generative AI (GenAI) Competency.

With this milestone, Dynatrace reinforces its position as a leading observability partner, backed by a proven track record of innovation and customer success on AWS. Building on its achievement of earning the AWS Machine Learning Competency, Dynatrace continues to drive advancements in generative AI.

Weighing the importance of this milestone for Dynatrace customers

This competency is more than just a badge; it’s a validation of how Dynatrace can help organizations safely, efficiently, and cost-effectively adopt generative AI in their business. AWS awards these competencies after rigorous technical validation and proven customer success. This means organizations can trust that Dynatrace solutions are designed to deliver measurable outcomes on AWS.

For existing customers, this competency reaffirms the Dynatrace commitment to continued innovation alongside AWS. This ensures the Dynatrace AI-powered observability platform evolves with the latest advancements in AI, future-proofing organizations’ existing investments as generative AI capabilities become core to modern cloud workloads.

For new customers, Dynatrace provides a trusted, proven foundation for observability and AI adoption on AWS. Whether an organization is exploring GenAI for customer engagement, automation, or new digital experiences, Dynatrace ensures these systems are reliable, secure, and optimized at every step.

Graph showing a layered approach to AI observability for agentic AI reliability
The Dynatrace layered approach to AI observability

Looking ahead with AI-powered observability on AWS

As organizations increasingly adopt generative AI, observability becomes a critical enabler. By leveraging Dynatrace causal AI, predictive insights, and seamless AWS integrations, organizations can maintain control over costs, risks, and performance while driving innovation, enhancing competitive advantage, and delivering exceptional customer experiences.

Whether you’re building, scaling, or fine-tuning GenAI application, Dynatrace and AWS Bedrock empower you to transform your observability. With end-to-end visibility into AI workloads, their interactions in full context of your business, and cloud-native applications, you can optimize performance, troubleshoot effectively, and maximize the value of your GenAI investments with greater confidence and precision.

Learn more

The post Dynatrace achieves AWS Generative AI Competency: A new milestone in observability and AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-achieves-aws-generative-ai-competency/feed/ 0
The rise of agentic AI part 6: Introducing AI Model Versioning and A/B testing for smarter LLM services https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-6-introducing-ai-model-versioning-and-a-b-testing-for-smarter-llm-services/ https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-6-introducing-ai-model-versioning-and-a-b-testing-for-smarter-llm-services/#respond Thu, 25 Sep 2025 16:39:51 +0000 https://www.dynatrace.com/news/?p=71137 Agentic AI - model versioning

Debug, optimize, and secure your AI models with confidence As agentic AI applications and systems gain traction, delivering reliable, high‑performing LLMs and agents becomes challenging due to heterogeneous stacks, non‑deterministic behavior, and cost sensitivity across multi‑cloud runtimes. Reliable delivery and deployment to production requires end-to-end telemetry across the full chain: UI/services → orchestration/agents (LangChain, LlamaIndex, […]

The post The rise of agentic AI part 6: Introducing AI Model Versioning and A/B testing for smarter LLM services appeared first on Dynatrace news.

]]>
Agentic AI - model versioning

Debug, optimize, and secure your AI models with confidence

As agentic AI applications and systems gain traction, delivering reliable, high‑performing LLMs and agents becomes challenging due to heterogeneous stacks, non‑deterministic behavior, and cost sensitivity across multi‑cloud runtimes. Reliable delivery and deployment to production requires end-to-end telemetry across the full chain:
UI/services → orchestration/agents (LangChain, LlamaIndex, MCP/A2A) → RAG pipeline (embedding + vector DB) → model gateway (OpenAI, Azure/OpenAI, Bedrock, Gemini, Mistral, DeepSeek) → GPU/infra. To support deterministic rollouts and continuous model improvement, teams need standardized tracing/metrics, guardrail signal capture, and automated cost and performance governance.

The hidden challenges of AI model management

The invisible bottlenecks

AI models, especially LLMs, are prone to issues like hallucinations, degraded performance, and incorrect outputs. Debugging these problems is often like finding a needle in a haystack. Existing tools fall short in providing a unified view to compare prompts, datasets, or model versions, making it hard to identify regressions or improvements.

The impact of deprecation and automatic upgrades on cost, performance, and quality

The rapid pace of innovation in the AI space means that providers like OpenAI and Anthropic frequently release new versions of their models, such as ChatGPT 5 or Anthropic Opus 4.1.

While these updates often promise better performance and new capabilities, they can also introduce significant risks for your AI services:

  • Deprecation of older versions: Providers may discontinue support for older models, forcing you to adopt newer versions without sufficient time to test their impact.
  • Automatic upgrades: Many AI providers automatically update their underlying models, which can lead to unexpected changes in behavior, degraded performance, or even broken workflows.
  • Compatibility issues: Changes in model behavior, such as output format or token usage, can disrupt your application’s functionality, requiring adjustments to prompts, configurations, or integrations.

Tracking token usage and managing costs is another uphill battle. Add to this the risk of prompt injection attacks and data leaks, and it’s clear that traditional methods are no longer sufficient

The new AI Model Versioning and A/B testing

Ship better models with confidence. In a single view, compare models and versions to validate improvements and spot bottlenecks across latency, reliability, token usage, cost, and output quality, then drill into prompt-level differences to confirm why a variant wins. When something breaks, follow the request end to end with distributed tracing: from input through orchestration steps and model calls to completion, so you can pinpoint exactly where an error or slowdown originated.

Compare models and versions: Detect bottlenecks and validate improvements in a single view.

Trace prompt failures: Debug errors from input to output with our Distributed Tracing solution.

Monitor costs and token usage: Gain real-time insights into token consumption and cost implications.

Detect security and guardrail risks: Identify and alert on vulnerabilities like prompt injection attacks, toxic responses, or captured PII.

Attach your own attributes like user session, feedback, or dataset ID for additional debugging information.

AI Observability model versioning and A/B testing

How it works

With AI Model Versioning, you can track metadata such as model version, dataset ID, and hyperparameters.

A/B testing lets you expose different user segments to model variations, providing data-driven insights into performance metrics like accuracy and cost.

Instrument in minutes: Use the supported OpenTelemetry-based SDK to instrument your service to capture prompts, completions, token usage, errors, and guardrail signals.
You can also enrich spans with attributes like model.version, dataset.id, user/session, and feedback for deeper analysis. (You can read more about this here.)

Start analyzing out of the box: Once data is flowing, the AI Observability app provides ready-made dashboards and distributed tracing so you can compare models/versions, monitor costs and tokens, and debug prompt failures end to end. No extra setup is required; you can try it out on the Dynatrace Playground right now.

 AI Model Versioning, you can track metadata such as model version, data video thumbnail

By combining observability, AI-driven insights, and organizational knowledge, we’re enabling systems that don’t just react but learn and adapt. Each critical issue or incident you resolve fuels a living knowledge base, paving the way for proactive incident prevention through alerting.

What’s next?

We’re committed to enhancing these capabilities further. Upcoming updates will include a dedicated app experience for multi-model and multi-cloud setups, advanced visualization tools, enhanced security features, intelligent forecasting, and alerting for cost/performance and guardrail optimization.

Get started today

Ready to revolutionize your AI services? Here’s how:

  1. Sign up for a free trial.
  2. Install the AI Observability app.
  3. Explore the AI Model Versioning ready-made dashboard, or check it out on our playground

Together, let’s build smarter, more reliable AI systems.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two of the Rise of Agentic AI blog series explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part seven introduces data governance and audit trails for AI services.

The post The rise of agentic AI part 6: Introducing AI Model Versioning and A/B testing for smarter LLM services appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-6-introducing-ai-model-versioning-and-a-b-testing-for-smarter-llm-services/feed/ 0
Shaping the Future: Autonomous Intelligence by Dynatrace https://www.dynatrace.com/news/blog/shaping-the-future-autonomous-intelligence-by-dynatrace/ https://www.dynatrace.com/news/blog/shaping-the-future-autonomous-intelligence-by-dynatrace/#respond Tue, 05 Aug 2025 15:31:08 +0000 https://www.dynatrace.com/news/?p=70228 Dynatrace for Executives: Leveraging Agentic AI

In my frequent interactions with customers implementing agentic AI, the expectations of two key audiences—executives and developers—quickly become apparent. Executives are actively exploring how to implement agentic AI, with a strong focus on unlocking significant productivity gains. They expect AI automation to free up engineering time, fix software automatically, prevent outages, and take over the […]

The post Shaping the Future: Autonomous Intelligence by Dynatrace appeared first on Dynatrace news.

]]>
Dynatrace for Executives: Leveraging Agentic AI

In my frequent interactions with customers implementing agentic AI, the expectations of two key audiences—executives and developers—quickly become apparent.

Executives are actively exploring how to implement agentic AI, with a strong focus on unlocking significant productivity gains. They expect AI automation to free up engineering time, fix software automatically, prevent outages, and take over the majority of 80% of non-feature tasks.

Developers are rapidly adopting AI for convenience and efficiency in their day-to-day work; it’s becoming as essential to them as internet access. For example, GitHub Copilot usage among developers rose from 17% in 2023 to 45% in 2024. They want AI to bring context and suggest precise error repairs, generate tests automatically, auto-collect information to fix vulnerabilities, and recommend optimizations based on real production insights.

The market is embracing agentic AI with growing excitement. KPMG’s AI Pulse Survey, 68% of business leaders plan to invest between $50 million and $250 million in generative and agentic AI technologies this year alone, up from 45% in 2024. Enterprises see it as a strategic priority and as an enabler for smarter automation. While the potential is real, the requirements to make the use of agentic AI robust and secure need a solid foundation.

Key insights

  • Agentic AI is powerful, but only as good as its foundation. While rapidly adopting agentic AI for its promise of autonomous action, the market also realizes that it requires more than a prompt-based agent. To be both reliable and precise, agentic AI must combine the creative problem-solving capabilities of probabilistic models like large language models with the rigor and accuracy of deterministic algorithms.
  • Agentic AI amplifies the value of Dynatrace AI. Thousands of organizations already benefit from Dynatrace AI capabilities: preventive operations, real-time insights, and improved productivity and reliability. Agentic AI will extend this foundation by enabling more autonomy, accelerating intelligent action and decision-making across cloud-native ecosystems.
  • Autonomous intelligence shifts human responsibilities from step-by-step instructions to goal setting and supervision. As Dynatrace is evolving into autonomous intelligence, we enable auto-remediation, auto-protection and auto-optimization, based on business-relevant goals. Rather than scripting every action, humans define high-level objectives and Dynatrace determines and executes the most effective path, while explaining every step and allowing human supervision. This shift requires structured, context-rich knowledge, causal reasoning, and AI agents that operate with trust, clarity, and precision.
  • Real-time, contextual data is a non-negotiable prerequisite. Agentic AI must not operate blindly only on its general-purpose model; it needs a fast memory, business-specific context, and the ability to synthesize signals across systems. Dynatrace Grail®, offers the only foundation that provides access to real-time insights from petabytes of structured and unstructured information without predefined schemas or indexing. Grail makes it possible for the user to ask any question, any time, and receive instant answers with organizations’ digital environment context in mind, revealing relationships and dependencies across the digital ecosystem as a directed graph connecting the right dots across tech and business.
  • AI-driven autonomy and insights work most effectively when brought across all organization. Dynatrace enables teams (from developers and site reliability engineers to operations and business or administration) to make smarter, faster decisions at every level.
  • 2026 update: The fusion of deterministic AI and agentic AI within Dynatrace Intelligence enables organizations to build and employ agentic frameworks that are not only capable but also reliable.

Context as foundation for reliable agentic AI

Imagine your car won’t start, and you ask an online car assistant for help. Most would start by asking you vague questions or suggesting generic fixes (“Try a new battery”) because they don’t understand or know the context of the problem. The next one might tell you: “Your engine is entirely broken. You need a new one.” Now, imagine instead you bring the car to an automotive expert who not only sees the reason for not starting but also instantly analyzes the entire build of your car down to the exact configuration of parts, how they interact, and even what parts were installed in what order. They don’t just know that the motor and screw exist: they know the screw holds the ignition coil to the engine block, and not the other way around.

This is how agentic AI works with Dynatrace. Agentic AI works like a team of experts who know your car inside out: every screw and why and how the vehicle was built. It’s not guessing but rather operating with architectural clarity, automatically pinpointing the root cause because it understands how everything is connected. With Davis® AI Root Cause Analysis, Dynatrace analyzes more than three million problems accurately and at scale every 24 hours, every day.

Instead of fumbling through 100,000 parts, it navigates a precise causal (say, the 50 services that actually influence the outcome) thanks to Dynatrace Smartscape. It doesn’t reach for every tool in the shed, but instead picks the right one for your specific digital system, every time.

So, similarly in IT: instead of general comments (“Your system seems slow, maybe scale your servers”), engineering teams get granular insights: “User slowdown originates from a failed API call in payment service, due to a misconfigured feature flag introduced in deployment of branch ‘calculation update in payment service’.” That’s how Dynatrace delivers context in action.

Today’s AI-powered automation in Dynatrace already shows agentic behavior

Dynatrace has long been operating at the intersection of data, intelligence, and automation. In fact, many capabilities typically associated with agentic AI, such as autonomous root cause detection, preventive operations, causal inference (which today has become causal AI), and self-healing production environments, have already been running across our platform for a decade.

Take this example: Dynatrace automatically detects a capacity issue, anticipates seasonal fluctuations, rates it by customer and business impact, and recalibrates the production environment across a customer’s hyperscaler setup, all end-to-end. It carries out full analytical and planning steps, creates reconfiguration plans, and only then notifies a human for final governance. This isn’t hypothetical: thousands of customers around the world are already leveraging our trusted causal and predictive AI in production workloads that run their businesses. And hundreds are taking the next step, adopting preventive operations by carefully adding generative AI to automatically draft remediation workflows, simulate outcomes, and enhance decision-making, shaping the future of intelligent automation.

Shaping the Future with Agentic AI: Autonomous Intelligence by Dynatrace
Example of the Dynatrace Problems app, where the service owner gets automatically tasked with a problem.

Dynatrace AI capabilities flag and remediate problems, surface insights, and feed them into IDEs. This process triggers ticket creation to the responsible teams and aligns them around automatically planned actions, including learning from past incidents while incorporating real-time facts in context.

To further evolve from automation to autonomy, Dynatrace magnifies its capabilities with agentic AI and delivers three reliable agentic AI requirements through an architecture built for intelligent action.

Leveraging agentic AI for redefined observability with Dynatrace

The future of observability is being redefined by a powerful triad: Knowledge, Reasoning, and Actioning.

  1. Knowledge. Dynatrace transforms contextual full-stack observability data into fact-based, real-time knowledge optimized for AI access. The Grail massive parallel processing data lakehouse is schema- and index-free, boosting AI agents with limitless query permutations. Grail works in tandem with Dynatrace Smartscape dynamic topology, an auto-discovered, continuously updated knowledge graph. This allows AI to deliver precise insights efficiently and at petabyte scale, eliminating the need for redundant queries (hence, also the increased cost) while maintaining full context and performance integrity.
  2. Reasoning. Dynatrace unifies causal, predictive, and generative AI to power expert AI agents that optimize the blend of deterministic logic with probabilistic and stochastic models, to provide precision and fact-based trustworthy decision-making, while minimizing risks of hallucinations. This enables context-aware decisions with built-in enterprise-grade safety, compliance, and observability of AI itself, ensuring transparency and trust to not only achieve a capable AI, but also a reliable one.
  3. Actioning. Dynatrace turns high-level objectives into intelligent, automated actions, where humans define the goals and AI determines the best way to achieve them, both reactively and proactively. With AutomationEngine, AppEngine, and OpenFeature, it remediates, optimizes, and even triggers systemic fixes, transforming observability into a strategic business enabler.
Leveraging agentic AI for redefined observability with Dynatrace.
Leveraging agentic AI for redefined observability with Dynatrace.

Last, but not least: AI is already powering production workloads across global enterprises, but not all AI is created equal. To deliver real value, it must be reliable, context-aware, and purpose-built for an organization’s digital environment. Dynatrace is engineered to meet those demands, magnified with agentic AI that answers organizations’ specific needs and business outcomes.

The post Shaping the Future: Autonomous Intelligence by Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/shaping-the-future-autonomous-intelligence-by-dynatrace/feed/ 0
The rise of agentic AI part 5: Developing and monitoring multi-agent applications with OpenAI Agents SDK on Azure AI Foundry https://www.dynatrace.com/news/blog/building-agentic-ai-applications-with-openai-agents-sdk/ https://www.dynatrace.com/news/blog/building-agentic-ai-applications-with-openai-agents-sdk/#respond Mon, 04 Aug 2025 15:36:15 +0000 https://www.dynatrace.com/news/?p=70239 Building agentic AI applications with OpenAI Agents SDK

As agentic AI applications gain ground, the trick becomes how to build multi-agent systems quickly with all the connective tissue built in. In this fifth installment of our series, The Rise of Agentic AI, we explain how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.

The post The rise of agentic AI part 5: Developing and monitoring multi-agent applications with OpenAI Agents SDK on Azure AI Foundry appeared first on Dynatrace news.

]]>
Building agentic AI applications with OpenAI Agents SDK

Recently, OpenAI released a customer service agents demo built using the OpenAI Agents Python SDK that showcases an example multi-agent system at work. With the OpenAI Agents SDK, you can build agentic AI applications with the help of agents, handoffs, guardrails, tools (built-in and custom), and built-in tracing. These capabilities support the core pattern of knowledge, reasoning, and actioning as the foundation for scalable and trustworthy automation, first introduced by Dynatrace CTO Bernd Greifeneder.

In this blog post, we explain how to build an OpenAI agents SDK-based agentic application and instrument the agents and app with AI-powered observability from Dynatrace. Dynatrace can help you see agent executions, tool usages, and prompt flows from initial request to final response for quick root cause analysis and troubleshooting.

To illustrate the capabilities of the OpenAI Agents SDK and agent framework with Azure OpenAI on Azure AI Foundry, we have built our multi-agent solution using the OpenAI customer service agents demo mentioned above as a reference and modified it for our use cases.

About our sample agentic AI application

Our multi-agent system enables users to research, summarize and translate across a range of topics and content. The system consists of four agents:

  • Welcome Agent: Engages the user, reasons with Azure OpenAI to analyze the prompt, identifies the intent, and passes it to the right agent to start processing.
  • Researcher Agent: Searches the web and analyzes the results using OpenAI.
  • Summarizer Agent: Summarizes content, including search results, text, PDF, CSV, and more, using Anthropic Claude.
  • Translator agent: Translates queries and inputs into any user-requested language using OpenAI.
OpenAI Agent SDK sample app architecture
Figure 1: Azure OpenAI Agent SDK setup for demo application in Github

Next, we want the multi-agent system to perform two distinct scenarios:

  1. Context history: In a specified chat session, the entire chat history and context is available for the duration of the session, while the individual prompts might be handed off to different agents for processing.
  2. Composite queries: The app orchestrates multiple different agents for different purposes, such as Research, Translate, Summary, and Welcome, so users can engage to process a composite prompt with multiple sub-queries.

Understanding multi-agent frameworks and handoff workflows

There are some key differences between the agent frameworks. Unlike the A2A protocol, the OpenAI framework does not explicitly have a central registry for agents. Instead, OpenAI agents use the concept of “handoffs” orchestrated by the OpenAI Agent Framework.

OpenAI framework agent handoffs

While orchestrator-led coordination offers a more deterministic and structured workflow, agent-to-agent handoffs provide significant advantages in adaptability and modularity. These handoffs enable agents to collaborate dynamically, making it possible to handle complex, multi-step queries with greater flexibility. This approach focuses on a more decentralized and scalable system, allowing agents to specialize and respond to changing requirements in real-time.

Here are two example scenarios to illustrate the agent-to-agent collaboration in chat sessions, with context, as well as delivering multi-agent query processing.

Show the user prompts for a composite query and multi-agent workflow

For example: “Research Michael Jordan, then summarize in 40 words or less, and then translate to French.”

Welcome Agent user prompt and composite query for the sample agentic AI application
Figure 2: User prompt -> Welcome Agent -> Identifies as multi-step workflow -> Handoff -> Researcher
Researcher Agent, Handoff, and Summarizer activities of the multi-agent workflow
Figure 3: Researcher Agent processes -> Handoff -> Summarizer
Handoff to Translator agent in multi-agent workflow
Figure 4: Summarizer -> Summary -> Handoff to Translator -> Summary results in French
Additional user input triggering translator, researcher, and response in the sample agentic AI application
Figure 5: User chat continues with Context and History -> Translator handoff -> Researcher -> Response
Researcher agent handing off to the translator for translation to Hindi
Figure 6: Researcher -> Handoff -> Translator to translate results to Hindi, keeping context and history

Multi-agent processing for CSV files uploaded

This example includes sample customer data to showcase multi-agent workflow processes with context and history in the chat session.

customer-uploaded CSV file and multi-agent triggers in the sample agentic AI application
Figure 7: Customer Data CSV -> Summarize file -> Welcome Agent -> Handoff -> Summarizer
Countries listed in the CSV file of the sample agentic AI application
Figure 8: “What Countries are listed in the file” -> Summarizer Handoff -> Researcher -> results
Research on the first country in summary
Figure 9: “Research on the 1st country in summary” -> uses context, history -> Researcher -> Results

Overall, the agent-to-agent handoffs worked well (and with context) during all the session runs. Tracing and debugging can be achieved by instrumenting the SDK with OpenTelemetry and sending the data to Dynatrace’s built-in AI Observability solution for Azure OpenAI. You can easily capture the multi-agent workflow for a given prompt on the Azure AI Foundry platform dashboard. Find the code examples in our GitHub repository.

Set up tracing using Python

Using Python, you can set up the tracing by changing a few simple lines of code in your agent framework and core component:

from traceloop.sdk import Traceloop Traceloop.init( app_name="openai-cs-agents", api_endpoint="https://wkf10640.live.dynatrace.com/api/v2/otlp", disable_batch=True, headers=headers, should_enrich_metrics=True, ) 
with tracer.start_as_current_span(name="update_seat", kind=trace.SpanKind.INTERNAL) as span: 
    context.context.confirmation_number = confirmation_number 
    context.context.seat_number = new_seat 
    assert context.context.flight_number is not None, "Flight number is required" 
    return f"Updated seat to {new_seat} for confirmation number {confirmation_number}"

You can see the results right away in distributed tracing:

Results of the OpenAI chat
Figure 10: Multi-agent workflow trace view in Distributed Tracing
Reviewing all OpenAI consumption statistics with Dynatrace AI Observability
Figure 11: How to review all your OpenAI consumption on Dynatrace with AI Observability

OpenAI orchestration

Within the OpenAI framework, there are two approaches to orchestrating agents:

  1. Allow the LLM to make decisions: Use the intelligence of an LLM to plan, reason, and decide what steps to take.
  2. Orchestrate with code: Use code to determine the flow of agents.

Overall, the OpenAI Agents SDK is comprehensive and easy to get running with some minor code changes, this time with OpenAI’s Codex assistant.

OpenAI Agents SDK Codex assistant code example
Figure 12: Codex example

Multiple frameworks and toolkits are quickly ramping up to make multi-agent systems a reality. We foresee this space evolving and innovating rapidly.

The evolution of multi-agent systems

As agentic AI continues to advance, multi-agent applications are poised to play a transformative role in reshaping how applications operate. These systems enable dynamic, context-aware collaboration between specialized agents, empowering businesses to tackle increasingly complex workflows. From helping with automation, orchestrating large-scale data analysis, multi-agent systems will unlock new levels of efficiency, scalability, and innovation.

Tools like the OpenAI Agents SDK on Azure AI Foundry and Azure AI Studio are at the forefront of this evolution. By providing built-in capabilities such as agent handoffs, guardrails, and tracing, the SDK simplifies the development and monitoring of multi-agent workflows. These features make it easier for organizations to deploy responsible, secure, and robust AI systems and also ensure transparency and trustworthiness in their operations. These are key factors for widespread adoption.

Looking ahead, we can expect rapid innovation in this space. Emerging standards like MCP, A2A protocols, and frameworks such as OpenAI Agents are creating a vibrant ecosystem for multi-agent interoperability. The focus will likely shift toward even more intelligent and reliable orchestration, where agents autonomously plan, reason, and adapt to dynamic environments.

AI Observability for agentic AI applications

To keep pace with these advancements, we believe that observability must evolve in lockstep to ensure transparency across heterogeneous agent ecosystems. Advancements in observability tools, such as the Dynatrace AI Observability solution, are essential to help create more reliable and scalable AI frameworks at the enterprise level.

The future of multi-agent systems holds immense potential, with the OpenAI SDK marking the starting point. We’re just at the beginning of what’s possible. As this technology evolves, it will gradually become more stable and reliable, ultimately transforming the way we approach automation, collaboration, and AI-powered problem-solving across industries.

Check out our GitHub repo for detailed code examples for OpenAI Agents, AWS Strands, Google ADK, and start building your own AI Observability solutions today.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two explores how monitoring A2A and MCP communications results in better, more effective agentic AI. This blog post covers AI agent observability and monitoring, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.

Together, these capabilities make it possible to achieve robust, scalable observability in agentic AI environments so teams can build reliable and trustworthy applications and services.

The post The rise of agentic AI part 5: Developing and monitoring multi-agent applications with OpenAI Agents SDK on Azure AI Foundry appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/building-agentic-ai-applications-with-openai-agents-sdk/feed/ 0
Dynatrace 3rd-generation platform: Built for the world of Autonomous Intelligence https://www.dynatrace.com/news/blog/dynatrace-3rd-gen-platform/ https://www.dynatrace.com/news/blog/dynatrace-3rd-gen-platform/#respond Tue, 22 Jul 2025 06:45:50 +0000 https://www.dynatrace.com/news/?p=70120 Dynatrace paving the way to autonomous intelligence

The world has become software-defined, distributed, and complex, creating a widening gap between digital complexity and our ability to manage and govern the systems that run businesses and organizations. To close this gap, we reimagine how observability works. It is no longer enough to collect and analyze telemetry after the fact. Organizations need trusted, intelligent […]

The post Dynatrace 3rd-generation platform: Built for the world of Autonomous Intelligence appeared first on Dynatrace news.

]]>
Dynatrace paving the way to autonomous intelligence

The world has become software-defined, distributed, and complex, creating a widening gap between digital complexity and our ability to manage and govern the systems that run businesses and organizations. To close this gap, we reimagine how observability works. It is no longer enough to collect and analyze telemetry after the fact. Organizations need trusted, intelligent systems that turn real-time data into reliable knowledge, apply advanced AI to reason through that knowledge, and take action to optimize outcomes at every level of the business.

This is the foundation of the Dynatrace 3rd-generation platform. We’ve spent the past two decades shaping the observability market. Today, we are transforming it from a rear-view mirror into a real-time control system for the modern enterprise. Thousands of organizations are already using Dynatrace 3rd generation to turn data into decisions and decisions into action. The result is faster innovation and stronger business results across every layer of the business.

A new model built on knowledge, reasoning, and actioning

Dynatrace 3rd generation introduces a new standard for observability and automation based on three foundational capabilities:

  • Knowledge: The Dynatrace platform turns petabytes of real-time data into a continuously updated, queryable knowledge graph. Powered by Grail and Smartscape, it provides trustworthy, fact-based insights with real-time context across dynamic environments.
  • Reasoning: Causal, predictive, and generative AI models work together to derive intelligent decisions. These models are context-aware, transparent, and built for enterprise-grade safety and compliance.
  • Actioning: Dynatrace enables users to define goals and let intelligent automation determine the best path forward, through innovations like AutomationEngine, AppEngine, and OpenFeature. This shifts operations from reactive remediation to preventive operations and continuous improvement.

Together, these capabilities form the foundation for autonomous intelligence. Dynatrace doesn’t just provide visibility; it enables systems to understand and act. By continuously converting real-time data into trustworthy insights, applying AI to reason through business and technical context, and triggering intelligent, goal-based actions, Dynatrace transforms observability into a real-time engine for automation and impact.

This is not about removing humans from the loop. It’s about empowering teams to define outcomes and rely on the system to carry out the best path forward. As with any leadership decision, autonomy depends on the quality of information and confidence in its context. The same principle applies to AI systems. Dynatrace 3rd generation gives organizations confidence, allowing them to scale decision-making with speed and trust.

Trusted knowledge, not just data

Legacy observability platforms focus on collecting telemetry data. But to support real-time decisions, teams need a trusted knowledge foundation. Dynatrace eliminates silos between metrics, traces, logs, events, user sessions, and security signals by unifying them in Grail, our schema-on-read, massively parallel data lakehouse.

For the first time, users can run any query at any time, with Grail supporting dramatically higher concurrency than traditional observability platforms. There’s no cold storage, no indexing, and no need for rehydration. This unlocks a goldmine of observability data and turns it into reliable, real-time answers.

That data is then contextualized in real time by Smartscape, our dynamic topology engine, and made instantly accessible to AI agents. The result is not just visibility, but deep, evolving system knowledge, providing machine-speed decisions no other platform can match.

AI that reasons with real-time context

Dynatrace has long set the standard for causal AI in observability. With the 3rd generation platform, we expand that foundation by combining causal AI with predictive and generative models. These AI types work together to support decisions at machine speed, with full context. Whether it’s automatically identifying the root cause of a service degradation, forecasting capacity needs, or evaluating how to improve online customer experiences, Dynatrace AI operates with the reliability and transparency required in enterprise environments.

Now, organizations can pursue modernization, transformation, and agentic AI initiatives with greater confidence. Dynatrace helps make AI accessible and actionable by reducing friction, delivering answers precisely when and where they are needed. Davis CoPilot enables natural language queries, workflow generation, and seamless integration into IDEs. It provides intelligent assistance at every step and supports a broad range of use cases across observability, security, and business operations with explainability, precision, and trust.

Automation that adapts to your goals

Traditional automation is limited by what is explicitly scripted. Dynatrace has taken a different approach. With the 3rd generation platform, you define high-level goals, and the platform determines the best way to achieve them. By grounding automation in real-time, high-quality data and precise causal analytics, Dynatrace ensures that actions are driven by accurate understanding, not assumptions.

This goal-based automation can resolve incidents, optimize performance, reduce cost, and even generate pull requests that improve code quality.

For example, preventive cloud operations allow site reliability engineers to move from firefighting to strategic orchestration. Instead of chasing alerts, SREs can focus on managing service-level objectives and improving business outcomes. Dynatrace handles the rest.

Built for the future of cloud and AI

The complexity of cloud-native architectures, Kubernetes deployments, and emerging agentic AI models are already testing the limits of traditional observability. Dynatrace 3rd generation is designed for the future.

By unifying telemetry, security data, and business context into a single real-time graph powered by Grail, Dynatrace provides the AI-powered intelligence required to operate modern systems with confidence. And by embedding automation throughout the platform, teams can scale faster than headcount, without sacrificing control or trust.

Turning observability into intelligent action

Organizations like TELUS and Air France-KLM are already seeing results: faster resolution, improved resiliency, and reduced downtime.

“By combining our Agentic AI initiatives with Dynatrace’s AI Observability capabilities, we’ve successfully optimized our development and operations workflows. We’re driving innovation and delivering measurable business impact while reducing downtime.”

– TELUS

“The AI and predictive capabilities from Dynatrace were a differentiator. We’re confident that any problem that arises can be dealt with quickly, dramatically reducing operational and revenue impact.”

– Air France-KLM

The combined impact of agentic AI initiatives and AI-powered observability extends beyond IT. These outcomes drive business performance, from greater availability and productivity to better customer experiences.

What this means for your organization

The Dynatrace 3rd-generation platform defines the path forward to a future where software can understand, reason, and act. It helps your organization move from reactive to proactive, from fragmented tools to unified intelligence, and from scripted automation to AI-driven operations.

Whether you are focused on cloud modernization, application security, cost optimization, or AI governance, Dynatrace provides a foundation to build and scale with confidence.

This is the evolution of observability. One built on context, driven by reasoning, and capable of taking action.

Learn more and see what’s possible with Dynatrace 3rd generation.

The post Dynatrace 3rd-generation platform: Built for the world of Autonomous Intelligence appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-3rd-gen-platform/feed/ 0
From data to insights with Dynatrace Dashboards https://www.dynatrace.com/news/blog/from-data-to-insights-with-dynatrace-dashboards/ https://www.dynatrace.com/news/blog/from-data-to-insights-with-dynatrace-dashboards/#respond Fri, 11 Jul 2025 13:42:19 +0000 https://www.dynatrace.com/news/?p=69842 Dynatrace dashboards

We had one main goal in mind when designing Dynatrace® Dashboards: reimagine how our customers consume and interact with their observability data. Built for speed, clarity, and collaboration, the Dashboards app helps teams easily explore, visualize, and act on telemetry data. From natural language queries to advanced visualizations, Dashboards streamlines your workflow and reveals critical insights at any scale. Whether you're monitoring infrastructure, applications, or AI workloads, Dashboards adapts to your needs, turning raw data into real-time insights.

The post From data to insights with Dynatrace Dashboards appeared first on Dynatrace news.

]]>
Dynatrace dashboards

We’ll walk you through a real-world example of monitoring OpenAI APIs in production to show you what this looks like in action.

In practice: Create a dashboard monitoring OpenAI LLM APIs

Imagine you’re on a platform team at a SaaS company that recently integrated OpenAI to power features like smart search, summarization, or chatbots. With these capabilities now live, your next challenge is ensuring they perform reliably, scale efficiently, and stay within budget. This is where Dynatrace shines—helping you transform telemetry into insights that drive action.

Let’s walk through all the steps to create just such a dashboard, and dig deeper to:

  • Find and add (OpenAI telemetry) data with ease.
  • Tailor visualizations to understand token usage, latency, and error metrics easily.
  • See what matters: filter and segment data by LLM model, service, or environment.
  • Predict and prevent issues: avoid model response slowdowns and cost spikes.

Find and add (OpenAI telemetry) data with ease

Creating a new dashboard begins with identifying and understanding the relevant data for your use case. Monitoring LLM APIs requires the visualization of key metrics like request volume, latency, or error rates per model. With Dashboards, exploring your data is intuitive, providing multiple ways to search for and analyze data.

  • Start with a ready-made dashboard that provides instant insights
  • Explore data using a simple-to-use point-and-click interface—ideal for getting started by quickly adding tiles
  • Utilize the full power of Grail by writing your own DQL query or utilizing Davis CoPilot® to transform your natural language prompts into DQL queries.

As an experienced Dynatrace user, you’re familiar with exploring data in context with our purpose-built apps like Kubernetes, Logs, or Distributed Traces, and how to add visualizations from those apps to your dashboards.

Let’s look at some of these approaches in the following sections.

Start the journey with ready-made dashboards

You don’t have to start from scratch. Dynatrace offers many ready-made dashboards as part of Dynatrace® Apps and purpose-built extensions to serve dedicated use cases. As the leading observability solution for monitoring AI workloads, we offer dashboards for all major AI and LLM stacks, including agentic frameworks such as OpenAI, Anthropic, Amazon Bedrock, or NVIDIA. These dashboards provide instant value, whether you’re monitoring performance or debugging expensive prompts. By delivering real-time insights into request volume, latency, cost, and service health, they not only save you time but also create a solid foundation for tailoring their experience to your needs.

Let’s start our journey by opening the ready-made dashboard for OpenAI and creating a copy of it. To follow along, locate the Dashboards app on the Dynatrace Playground.

Duplicate the ready-made dashboard to customize it.
Figure 1. Duplicate the ready-made dashboard to customize it.

Add further tiles to analyze token usage

Next, let’s add another tile to visualize the overall prompt token usage by type: input vs. output for OpenAI services. From discussions with our platform observability team, we know that all relevant metrics sent to Dynatrace using OpenTelemetry are available as custom metrics prefixed with gen_ai. We add a metrics tile and type gen_ai into the search field. This instantly surfaces all related telemetry. A few clicks later, applying data splits and aggregations, we have two more tiles, demonstrating how simple it is to turn raw telemetry into actionable insights:

  • pie chart that shows the balance between input and output tokens
  • line chart that tracks how the usage evolves over time

Visualizing overall prompt token usage video thumbnail
Figure 2. Visualizing overall prompt token usage.

For further insights into the exploration and transformation possibilities in Dashboards, check out our blog post on transforming data into insights.

Leverage the power of Dynatrace Grail

Not sure where to start, which metric to use, or how to quickly advance with the power of Dynatrace Query Language (DQL) and Grail® data lakehouse? That’s where Davis CoPilot® comes in. Built directly into Dashboards and Notebooks, Davis CoPilot allows you to interact with your data using plain language—no need to write queries or know exact metric names. Just type something like Visualize token usage by input and output types, and the AI will help you instantly generate the appropriate query, taking you from question to insight in seconds.
CoPilot Token Usage video thumbnail
Figure 3. Use Davis CoPilot to create and visualize queries instantly.

Tailor visualizations to easily understand token usage, latency, and errors

As someone responsible for monitoring systems or ensuring service reliability, you know how important it is to get the right insights at a glance. Dynatrace helps you build intuitive dashboards that focus on what matters most: understanding your data and taking action on it.

Once the data is set and a tile added, Dynatrace automatically suggests the most suitable visualization. For example, when tracking API token usage by type over time, a line chart is recommended to highlight trends and fluctuations.

A suitable line chart visualization is automatically suggested.
Figure 4. A suitable line chart visualization is automatically suggested.

Dynatrace also applies other smart defaults based on the context of the visualized data. For example, when you add a metric that tracks the usage of example prompts and split it by the prompt name, sparklines are automatically included to show trends over time—no extra configuration needed. And if you’re already a power user, the newly added search speeds up your dashboard creation journey by offering a way to instantly jump to any configuration without the need to scroll around. But there’s a lot more that helps improve the user journey. We harmonized the settings of individual visualization types, ensuring that already defined configurations, such as color palettes or units, persist, even if you change the type.

The settings of individual visualization types are enhanced and harmonized.
Figure 5. The settings of individual visualization types are enhanced and harmonized.

We’ve also made many updates to the chart plotting features of our pre-existing visualizations. For example, the single value tile, which used to be a basic number display, is now a highly expressive component. You can now enrich the single-value tile with icons, apply color thresholds to flag anomalies, add sparklines to show trends, and add value and trend labels that provide additional context for the charted value and give it meaning.

The single value tile now also includes sparklines and other options.
Figure 6. The single value tile now also includes sparklines and other options.

Plotting the values on a map benefits many signals. Consider displaying token usage or prompts issued per destination. The map component has a rich set of customization options—such as color rules, pin shapes, and unit formatting—explicitly designed to support the visualization of geographic data.

Use the map visualization to display data geographically and to uncover location-based patterns.
Figure 7. Use the map visualization to display data geographically and to uncover location-based patterns.

See what matters: filter and segment data by LLM model, service, or environment

To make a dashboard truly actionable, the next step is to add filters and segmentations. This allows you to tailor one view dynamically for different audiences, environments, or services, all within a single dashboard. For example, you might filter an OpenAI dashboard by environment (production, staging, test) or model type (GPT-4.1, o3, o3-mini) to focus on what matters most in each context.

Dynatrace offers powerful ways to filter data:

  • Using reusable segments, multidimensional global filters can be applied to all tiles and data. This is ideal for applying a specific (user) context, such as environment, team, or cluster. Segments are persisted across navigation between apps, allowing for simple drill-down journeys.
  • With variables, we introduce Dashboard-specific filters, offering fine-grained control for each tile, perfect for filtering information such as LLM model type or feature toggles.

If you want a more in-depth tutorial, check out our latest blog post on filtering.

Predict and prevent issues: avoid model response slowdowns and cost spikes

Dashboards aren’t meant to be stared at all day. In most organizations, they’re often left untouched until something goes wrong. That’s when dashboards become invaluable: surfacing the correct data at the right time to help teams quickly understand, diagnose, and resolve issues.

From passive observation to proactive action, Dynatrace bridges the gap with interactive, AI-powered dashboards that don’t just visualize data; they empower you to act on it. You can add alerts and forecasts directly from charts with just a few clicks.

For example, suppose your dashboard tracks OpenAI model response times and associated costs. In this case, you can set an alert to notify your team if the average response time exceeds a certain threshold for a defined period directly from within the chart. This ensures you’re reacting to issues and anticipating them before users are impacted.

You can interact with your data directly on your charts, for example, zoom in/out and set up instant alerts.
Figure 8. You can interact with your data directly on your charts, for example, zoom in/out and set up instant alerts.

Another popular example of proactive monitoring is cost forecasting. Our dashboard already tracks cost trends over time—such as “prompt costs” and “complete costs,” for example—with a line chart highlighting weekly fluctuations.

By enabling forecasting, Dynatrace projects future spending based on historical usage patterns. This helps you anticipate budget overruns, adjust resource allocation, and make informed decisions before costs spiral. The predicted budget spend is shown alongside a table highlighting the “Top 10 expensive prompts.” This allows teams to identify which workloads or user actions contribute most to spending, ideal for optimization efforts or chargeback models.

Utilize AI-powered forecasting to predict future costs.
Figure 9. Utilize AI-powered forecasting to predict future costs.

Share with teams: secure, flexible collaboration

The next step is to share our dashboard with the right people, ensuring teams are aligned across job roles and departmental boundaries. The new Dynatrace Dashboards supports flexible sharing options for collaboration within your organization.

Fine-grained collaboration settings allow you to:

  • Share a document with specific users or groups applying either view or edit permissions.
  • Roll out a dashboard to users in the environment.
  • Generate a link that works for any authenticated user in your environment—ideal for broad internal visibility without managing individual access.

Ready to try it out yourself?

Dynatrace Dashboards redefine how teams interact with observability data. Whether you’re monitoring LLM APIs, optimizing cloud costs, or ensuring service reliability, Dashboards empowers you to:

  • Explore data intuitively.
  • Visualize insights using smart defaults and rich customization options.
  • Segment and filter your data dynamically, offering tailored views for use cases.
  • Act proactively on data anomalies using forecasting and creating alerts in context.

Experience the power of Dashboards: Head over to the Dynatrace Playground and browse the ready-made dashboards or create your own, following the steps described in this blog post.

The post From data to insights with Dynatrace Dashboards appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/from-data-to-insights-with-dynatrace-dashboards/feed/ 0
The rise of agentic AI part 4: Dynatrace delivers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM https://www.dynatrace.com/news/blog/full-stack-observability-for-nvidia-blackwell-and-nim-based-ai/ https://www.dynatrace.com/news/blog/full-stack-observability-for-nvidia-blackwell-and-nim-based-ai/#respond Fri, 20 Jun 2025 06:00:05 +0000 https://www.dynatrace.com/news/?p=69115 Davis CoPilot for NVIDIA

The Dynatrace® unified, AI-powered observability platform delivers full-stack AI and LLM observability, including of NVIDIA Blackwell and NVIDIA NIM systems, and AI-driven insights to meet the scale and complexity of enterprise AI deployments. In this fourth installment of our series, The Rise of Agentic AI, we explore how the Dynatrace integration with NVIDIA systems provides enterprises with all the insights needed to detect customer-facing issues, helping IT teams maintain performance, reliability, and security across their AI workloads.

The post The rise of agentic AI part 4: Dynatrace delivers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM appeared first on Dynatrace news.

]]>
Davis CoPilot for NVIDIA

NVIDIA Blackwell systems provide high-performance infrastructure for enterprise AI, and now, thanks to the Dynatrace integration with the NVIDIA Enterprise AI Factory reference design, enterprises can add Dynatrace Full-Stack Observability to NVIDIA Blackwell infrastructure. This magnifies the value of the NVIDIA Blackwell platform by providing real-time performance insights, anomaly detection, and dependency mapping.

Keep high performance and security top of mind with unified observability and security

Figure 1. The Dynatrace AI Observability platform
Figure 1. The Dynatrace AI Observability platform

Dynatrace aligns with high data security and privacy standards typical of on-premises NVIDIA Blackwell deployments, particularly in regulated industries such as finance and healthcare. Its unified data model, Smartscape® topology mapping, and Davis® AI engine provide deep visibility into the full stack—from GPU metrics and containerized workloads to distributed applications and user experiences, enabling tailored observability for workloads running on  NVIDIA Blackwell. Integrating NVIDIA Data Center GPU Manager or other telemetry sources is straightforward, allowing teams to monitor GPU health, utilization, thermal thresholds, and memory bandwidth alongside traditional infrastructure metrics.

Dynatrace technology allows for automated discovery and instrumentation of services running on NVIDIA Blackwell-accelerated systems. Whether monitoring high-throughput GPU compute tasks, Kubernetes clusters, or microservices, Dynatrace ensures low-overhead performance monitoring with minimal manual configuration.

AI-powered, real-time insights improve performance and explainability

With Dynatrace Full-Stack AI Observability, you can monitor real-time performance, trace prompts end-to-end, and ensure compliance, optimizing cost and throughput for your AI and LLM workflows and agents, offering various use cases such as

  • Monitor service health and performance, tracking real-time metrics and offering clear visibility into service incidents.
  • Validate service quality by measuring response speed or identifying performance hotspots.
  • End-to-end tracing and debugging pinpoint the root cause of errors and failures in the LLM chain, troubleshoot issues in complex pipelines, and trace dependencies across the entire system spanning multiple LLMs, RAG pipelines, and agentic frameworks.
Figure 2. Sample dashboards provided for tracking service health and performance
Figure 2. Sample dashboards are provided for tracking service health and performance

Unified AI-powered observability

Dynatrace delivers full stack observability for your LLMs and Generative AI applications running on NVIDIA Blackwell systems. Its ability to provide visibility into complex, high-performance environments allows enterprises to fully leverage Blackwell’s capabilities while maintaining operational excellence and system reliability, improving the performance, explainability, and compliance of your AI workloads and agents.

Figure 3. Dig deeper into the possibilities of AI and LLM observability on the Dynatrace Playground
Figure 3. Dig deeper into the possibilities of AI and LLM observability on the Dynatrace Playground

Visit the Dynatrace Playground to learn more and gain hands-on experience with prepopulated data, so you can experience the possibilities of AI and LLM observability with Dynatrace. If you’re interested in using Dynatrace for your own AI workloads, visit our documentation and start benefiting from full stack observability for AI and LLM.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.

The post The rise of agentic AI part 4: Dynatrace delivers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/full-stack-observability-for-nvidia-blackwell-and-nim-based-ai/feed/ 0
Auth0 monitoring with Dynatrace for more secure authentications https://www.dynatrace.com/news/blog/auth0-monitoring-with-dynatrace-for-more-secure-authentications/ https://www.dynatrace.com/news/blog/auth0-monitoring-with-dynatrace-for-more-secure-authentications/#respond Thu, 12 Jun 2025 12:00:06 +0000 https://www.dynatrace.com/news/?p=69442 Dynatrace AI-powered observability is now on Google Cloud

Understanding user authentication patterns and security events is critical for maintaining robust application security. Auth0, a leader in identity management, is a secure and customizable identity platform that simplifies authentication and authorization for applications of any scale. Auth0 generates rich logs detailing every authentication event. Meanwhile, Dynatrace provides powerful observability across your entire technology stack. […]

The post Auth0 monitoring with Dynatrace for more secure authentications appeared first on Dynatrace news.

]]>
Dynatrace AI-powered observability is now on Google Cloud

Understanding user authentication patterns and security events is critical for maintaining robust application security. Auth0, a leader in identity management, is a secure and customizable identity platform that simplifies authentication and authorization for applications of any scale. Auth0 generates rich logs detailing every authentication event. Meanwhile, Dynatrace provides powerful observability across your entire technology stack. Auth0 monitoring with the Dynatrace observability and security platform enables organizations to gain unprecedented insights into authentication behaviors, security anomalies, and the relationship between identity events and application performance.

This blog explores how to implement the power of Auth0 monitoring, the benefits it provides, and real-world examples of how it enhances your security posture and user experience monitoring.

Why monitor Auth0 with Dynatrace?

Auth0 logs capture detailed information about authentication events, including:

  • Login successes and failures
  • Password changes and reset requests
  • Multi-factor authentication events
  • Permission and access changes
  • Anomalous security events
  • API token usage
  • User management operations

While the Auth0 dashboard provides basic log viewing capabilities, integrating these logs with the advanced analytics of the Dynatrace platform offers several significant advantages:

  • End-to-end observability. Connect authentication events with application performance metrics, infrastructure health, and user experience data.
  • Advanced visualization. Create comprehensive dashboards showing authentication patterns alongside other system metrics.
  • Proactive alerting. Set up intelligent alerts based on complex combinations of authentication and application metrics.
  • Security monitoring. Quickly identify suspicious login patterns or authentication failures.
  • Compliance requirements. Generate comprehensive audit trails to meet regulatory compliance requirements.

Ingesting Auth0 logs into Dynatrace

Before you begin, ensure you have the following:

  • An active Auth0 account with admin access
  • A Dynatrace environment with appropriate permissions
  • API access to both platforms

First, we’ll set up the data extraction, then transform and enrich the data with Dynatrace OpenPipeline™, and finally, the analytics part, set up dashboards and engage Davis AI.

Workflow diagram showing the order of operations in the Auth0 data ingestion process as part of Auth0 monitoring
Figure 1. Process for ingesting Auth0 logs into Dynatrace for Auth0 monitoring.

Data extraction

Auth0 provides multiple ways to export logs. For Dynatrace integration, use the Auth0 Event Stream capability. For detailed step-by-step instructions, see Dynatrace Log Streaming on the Auth0 Marketplace. Authentication for ingesting logs and connecting to the correct tenant are key to integrating the data. If the extracted data contains personally identifiable information (PII), you can perform the data extraction in a pre-production environment.

Data transformation

Once ingested, Dynatrace OpenPipeline, a unified, high-scale stream-processing technology, automatically contextualizes incoming data and enriches signals by adding metadata and links to other relevant data signals. In addition, OpenPipeline ingests and processes data securely and compliantly.

For this Auth0 example, we set up a new log pipeline and added the following three OpenPipeline processing steps:

  • Set log level and status for ingested logs
  • Mask and replace pattern
  • Route incoming log data to a separate bucket

We provide sample snippets below that you can add to your own pipeline.

Set log level and status for ingested Auth0 logs

First, let’s add the log level and status for the ingested Auth0 log signals using Dynatrace OpenPipeline processing instructions (processors) expressed in the Dynatrace Query Language (DQL).

Screenshot showing Dynatrace OpenPipeline with processing instructions expressed in Data Query Language (DQL) as part of Auth0 monitoring
Figure 2. Configure the log level and status for Auth0 logs in Dynatrace OpenPipeline using DQL.
fieldsAdd status = if(status == "NONE" AND (startsWith(data.type, "f") OR contains(data.type, "failed")), "ERROR", else:status)
| fieldsAdd status = if(status == "NONE" AND (startsWith(data.type, "s") OR contains(data.type, "succeed")), "INFO", else:status)
| fieldsAdd status = if(status == "NONE" AND (contains(data.type, "exceed") OR contains(data.type, "limit")), "WARN", else:status)
| fieldsAdd status = if(status == "NONE", "INFO", else:status)
| fieldsAdd loglevel = status
| fieldsAdd data.date = data.date
| fieldsAdd timestamp = data.date

Masking and replacing pattern

Dynatrace applies this masking pattern to data from specific geographic regions, which can vary based on area codes. When using the example code below, add a DQL processor, paste the example code, and alter it based on your individual content.

parse content, "LD ([a-zA-Z0-9.!#$%&*+-/=?^_{|}~]+ '@' LD '.' ALNUM'.'? ALNUM?):email"
| parse data.user_name, "LD:email"
| fieldsAdd content = replacePattern(content, "([a-zA-Z0-9.!#$%&*+-/=?^_{|}~]+ '@' LD '.' ALNUM'.'? ALNUM?)", hashMd5(email))
| fieldsAdd data.user_name = replacePattern(data.user_name, "LD:email", hashMd5(email))
|fieldsRemove email

Routing incoming log data to a separate bucket

Log data in Dynatrace can be stored in different buckets for addressing security, compliance, or performance objectives. The sensitivity of Auth0 data calls for storing its logs in a dedicated bucket with long term retention and access rights limited to tightly restricted admin users.

Process diagram showing the stages for routing logs through OpenPipeline to Grail.
Figure 3. Process for routing logs through OpenPipeline into the Dynatrace Grail data lakehouse.

Log analytics and insights

Once the data is stored in Dynatrace, we can start analyzing data.

First, let’s create a dashboard for simple visual data exploration purposes. It’s very straightforward to create your own dashboard—you can either use the built-in explore data interface, translate natural language prompts into DQL statements using the Dynatrace natural language AI assistant, Davis CoPilot™, or create the DQL statement on your own.

In the second step, we remove and mask sensitive data to address compliance and privacy requirements. You can learn more about applying sensitive data masking on capture using either Dynatrace OneAgent® or OpenTelemetry. As an alternative, you can also use field-level permissions to apply data masking on read.

Once this is done, we can concentrate on visualizing data. Dynatrace dashboards provide the flexibility to view data with any required dimensions.

By using the detailed information captured by Auth0, you can easily identify system health and spot user behavior trends. For example, you can track trends for the following use cases:

  • Login successes and failures
  • Password changes and reset requests
  • Multi-factor authentication events
  • Permission and access changes
  • Anomalous security events
  • API token usage
  • User management operations
Screenshot showing a dashboard with Auth0 authentication data, such as login successes and failures and password chagnes
Figure 4: Sample dashboard showing statistics from Auth0, including successful and failed logins, password changes, and multi-factor authentication events.

Using Davis AI for advanced anomaly detection and forecasting

Davis AI can identify anomalies in data and predict future trends. While anomaly detection ensures that administrators will be notified once a metric is no longer operating within its boundaries, forecasting is used to predict the future based on historic values.

Davis AI is fully integrated in Dashboards. To learn how to add your own anomaly detectors or set-up forecasting, see the blog Better dashboarding with Dynatrace Davis AI.

Alerting incorporated in dashboard

The following example chart shows the number of times the failed SMS count has breached the auto adaptive threshold. These conditions can generate an alert and notifications via Slack or email as required.

Screenshot showing a line graph of failed SMS count exceeds its threshold and triggers an alert.
Figure 5. Dynatrace sends alerts for SMS counts that exceed their thresholds.

Predicting the number of sign-ups

Davis AI forecast analysis predicts future numeric values of any time series. It can even process external datasets or the results of any data query if it can be displayed as a numeric time series, such as occurrences over time.

A line graph showing trends from actual data that can then be forecast by Davis AI.
Figure 6. Davis AI can forecast trends from data.

Advanced use cases for Auth0 monitoring

Once you have Auth0 data in the Dynatrace platform, you can do some advanced analytics.

Correlating authentication events with application performance

One of the most powerful aspects of this Auth0 monitoring integration is the ability to see how authentication processes impact overall application performance. For example:

  • Create a custom dashboard showing login response times alongside application load metrics.
  • Set up alerts when authentication response times exceed thresholds.
  • Analyze how authentication traffic spikes affect backend services.

Security anomaly detection

Configure Davis AI to detect unusual authentication patterns:

  • Sudden increases in failed login attempts
  • Authentication attempts from unusual locations
  • Password reset patterns that deviate from normal behaviour
  • Unusual activity on dormant accounts

User behavior analytics

Combine Auth0 logs with Dynatrace Real User Monitoring to create comprehensive user behavior profiles:

  • Typical login times and locations for specific user segments
  • Authentication method preferences (password vs. social logins vs. SSO)
  • Device and browser usage patterns during authentication
  • User journey patterns following successful authentication

Auth0 monitoring best practices

Based on our experiences, we recommend applying the following best practices:

  1. Filter judiciously. Auth0 generates extensive logs; only send relevant events to Dynatrace.
  2. Respect PII. Be careful with personally identifiable information in logs; consider masking.
  3. Set appropriate retention. Configure log retention based on security requirements and compliance needs.
  4. Monitor integration health. Create monitoring for the integration itself to ensure log delivery.
  5. Start small. Begin with key authentication events before expanding to full log streaming.

Auth0 monitoring places authentication events in context

Integrating Auth0 logs with Dynatrace creates a powerful security and performance monitoring solution that bridges the gap between identity management and application observability. This integration enables organizations to:

  • Detect security threats earlier through correlated analysis
  • Optimize authentication flows based on performance data
  • Understand the relationship between authentication and user experience
  • Create comprehensive security analytics
  • Streamline troubleshooting for authentication-related issues

In an era where digital identity is the cornerstone of security, having comprehensive observability of authentication events within your broader application monitoring strategy is invaluable. The Auth0-Dynatrace logs integration provides this critical capability, empowering organizations to enhance both security and user experience simultaneously.

The post Auth0 monitoring with Dynatrace for more secure authentications appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/auth0-monitoring-with-dynatrace-for-more-secure-authentications/feed/ 0
The rise of agentic AI part 2: Scaling MCP best practices for seamless developers’ experience in the IDE with Cline https://www.dynatrace.com/news/blog/mcp-best-practices-cline-live-debugger-developer-experience/ https://www.dynatrace.com/news/blog/mcp-best-practices-cline-live-debugger-developer-experience/#respond Wed, 11 Jun 2025 16:51:18 +0000 https://www.dynatrace.com/news/?p=69450 abstract image showing connected dots and waves representing MCP best practices for agentic AI

Agentic AI is dramatically altering how engineering teams respond to incidents and debug use cases. Monitoring agent communications using model communications protocol (MCP) and live debugging are key to gaining insight into how AI models are performing. In this second installment of our series, The Rise of Agentic AI, we explore some MCP best practices that Dynatrace customer TELUS uses to massively accelerate agentic AI issue resolution.

The post The rise of agentic AI part 2: Scaling MCP best practices for seamless developers’ experience in the IDE with Cline appeared first on Dynatrace news.

]]>
abstract image showing connected dots and waves representing MCP best practices for agentic AI

Agentic AI is revolutionizing—and dramatically accelerating—how engineering teams respond to incidents and debug use cases. By monitoring how AI agents communicate using standards such as Model Context Protocol (MCP), performance engineers can gain deep insights into how systems and AI models themselves are performing. And crucially, automate incident responses using coding assistants.

The next step in this scenario is to employ some MCP best practices, such as helping full-stack engineers to debug code at runtime without disrupting operations. Fortunately, Dynatrace has the answer with Live Debugger. First debuted earlier in 2025, Live Debugger gives developers instant access to code-level troubleshooting data in any environment, including production, and is now generally available.

With all these capabilities coalesced in an integrated development environment (IDE), developers can interact with AI using natural language to massively accelerate issue resolution.

Key takeaways:
  • Unify AI debugging resources in one place. Debugging agentic AI workflows starts with combining Dynatrace features with custom MCPs in an IDE.
  • Leverage an AI coding assistant to add context. An AI coding assistant, such as Cline, helps add vital context and improve prompt engineering using natural language queries.
  • Use best practices to configure MCPs and Cline. Combining Cline with MCPs and on-the-fly debugging reduces guesswork and manual overhead.

MCP best practices TELUS uses to accelerate incident response

As an early adopter of Dynatrace Live Debugger, TELUS harnessed the power of Davis® AI, MCP, and real-time debug data to gain a deep understanding of their code, enabling every engineer to ask questions in natural language about their code behavior at runtime.

At a recent Dynatrace Guild meeting, TELUS SREs Dana Harrison and Cheng Li demonstrated how they’re combining Dynatrace features like Live Debugger with custom MCPs for the ultimate purpose-built troubleshooting setup in their Visual Studio Code IDE. By integrating multiple resources in one place, the TELUS team can blend next-generation debugging with agentic AI workflows. This integration reduces developer context-switching and enables developers to use straightforward, natural-language prompts.

With a fully instrumented environment that monitors every service end-to-end, the TELUS team leverages Dynatrace Live Debugger for Node.js and Java microservices to monitor running code in real time to reduce mean time to resolution. They initially tested this approach in non-production but are expanding into production with strict auditing and data masking.

The team also uses MCP in tandem with Cline AI, a coding assistant capable of gathering logs, metrics, and trace data through natural language commands. By unifying multiple data sources under MCP, TELUS avoids manually juggling different dashboards and tools. Developers can simply interact with the AI tools, which fetch relevant context on demand, even correlating Live Debugger snapshots with real-time logs. As a result, the entire debugging and incident investigation process remains inside the IDE, which significantly streamlines performance engineering and accelerates issue resolution.

How MCPs empower AI agents for debugging use cases

As an open standard, MCP connects AI agents to relevant data sources, such as repositories, tools, or external APIs. Instead of bespoke integrations for each data silo, MCP provides a universal interface to connect multiple relevant sources to feed the right context to the models and agents. This interface simplifies how agents access relevant context, leading to better task outcomes, execution, and more consistent performance across complex environments. For reference, you can try out the Dynatrace MCP server on GitHub.

architecture diagram showing Dynatrace MCP monitoring reference architecture
Figure 1. Dynatrace MCP server reference architecture.

MCP best practices: How Cline AI helps to bring in context

Cline AI is a coding assistant in Visual Studio Code that leverages multiple MCP servers to streamline data retrieval from various back-end systems. With plain-language queries (“Search for a specific service,” “Fetch logs from GKE,” or “Investigate errors”), developers can prompt Cline AI to automatically contact the relevant MCP server.

By becoming better at prompt engineering and handling the underlying API calls and correlating responses from Dynatrace MCP, native Kubernetes logs, or even third-party services, Cline AI provides an all-in-one investigation workflow right in the editor. This means developers don’t have to switch among multiple tools or memorize specialized APIs; they simply ask Cline AI what they need in natural language, and the MCP layer manages the rest.

Cline supports MCP-client features, such as dynamic tool discovery, prompt reusability, and adaptive resource access. It also enables powerful features like custom instruction, cline-rule, and memory banks.

TELUS’ best practices for configuring MCPs and Cline

The TELUS team followed these MCP best practices while configuring Cline.

Install Cline AI as VS Code extension

First, the TELUS team installed Cline AI as a VS Code extension, pointing it to their AI proxy, Fuel iX. Fuel iX lets them choose among various large language models based on the user’s requirements and desires.

Use careful prompt engineering and Cline rules to direct specific MCP tools

Cline AI can then interpret a developer’s natural-language prompt—such as “Investigate errors in our Node.js service”, “What is this service about”—and select the appropriate MCP endpoint(s) to gather data. Directing to specific MCP tools for more accurate results can be fine-tuned through careful prompt engineering, creating detailed .clinerules file definitions, and directives built into the MCPs themselves – including predicted requests and response formatting.

Configure each MCP server for a specific data domain

Each MCP server at TELUS is responsible for interacting with a specific data domain: for instance, they have a Dynatrace MCP server to fetch metrics and trace data, a GCP Logs MCP server to pull container and other forwarded logs, and additional servers for services like JIRA issues or PagerDuty.

Rely on centralized agentic AI commands, not vendor-specific APIs

By standardizing how these data sources are exposed, TELUS ensures that Cline AI never has to manage direct integrations or vendor-specific APIs. Instead, the assistant simply issues commands to the MCP servers, which internally handle authentication, query templates, and data normalization. This architecture not only allows TELUS to maintain clear boundaries between data retrieval logic and AI-driven workflows, but also accelerates developer onboarding, making it simple for any team member to debug or investigate issues from within VS Code by asking Cline AI, rather than switching between specialized tools or dashboards.

​​Use MCPs as a standardized endpoint

​In this context, MCP servers act as standardized endpoints, each responsible for a particular data source or application. By relying on Cline for coding assistance, the TELUS team can type a natural-language command (“Investigate issues in our Node.js service”), and Cline will automatically query Dynatrace Grail or GKE logs using the relevant MCP server.

This approach centralizes the complexities of data retrieval in one place and lets the AI assistant produce a consolidated summary or recommended fix, all within the IDE.

Step-by-step guide to set up Cline with Dynatrace

  1. Install Cline through the VS Code Marketplace.
    screenshot showing Cline configuration as part of MCP best practices
  2. Configure Cline with the required credentials, usually an API endpoint and key.
    screenshot showing Cline configuration
  3. Install your MCP. This example uses an internal package, but there are currently many MCPs published for installation.
    screenshot showing MCP installation
  4. Validate that the Dynatrace MCP server is up and running.
    screenshot showing MCP validation
  5. Ask Cline your questions. The agent formats them according to the needs of the MCP(s) you installed.
    Screenshot showing asking a question of Cline
  6. Example response of a production authorization service problem pulled and summarized from Dynatrace.
    screenshot showing results of the question asked of Cline as part of MCP best practices

The objective of this combined setup (Live Debugger + MCP + Cline) is fast and intelligent troubleshooting. Instead of juggling multiple dashboards and CLI tools, an engineer can see code snapshots, error traces, logs, and even recently created JIRA tickets within one session. As mentioned, MCP standardizes queries across diverse systems, so the AI agents can correlate all relevant information, such as referencing a single trace ID to pull logs from GKE, configuration details from the Kubernetes cluster, or known issues in PagerDuty. This AI agent-powered workflow saves time and reduces context switching.

The TELUS team found that combining these insights with on-the-fly debugging significantly reduces the guesswork and manual overhead that typically slows resolution times.

The road ahead: Developer-first observability and IDE-centric debugging at TELUS with Dynatrace

Although TELUS currently uses Live Debugger in mostly non-production environments, they plan to enable it in production using strict governance. Their upcoming strategy involves role-based permissions for debugging sessions, mandatory security training on data privacy, and well-defined data retention policies to keep snapshots transient.

When combined with MCP’s universal integration points, like Dynatrace MCP, developers can quickly drill down into real production issues, see the exact variables at fault, and either push a fix or revert a misconfiguration with minimal downtime.

Looking ahead, the TELUS team plans to integrate Davis Copilot APIs to further simplify natural language interactions with Dynatrace.

In practice, this would allow an AI assistant to automatically translate human-readable queries into Dynatrace Query Language (DQL), making it even easier to retrieve the right metrics, traces, or logs. By layering Davis CoPilot™ on top of their existing MCP approach, TELUS aims to reduce manual DQL writing while offering developers a powerful yet streamlined way to perform more advanced data analysis through simple, intuitive prompts.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.

More MCP best practices and resources for Dynatrace Live Debugger

The post The rise of agentic AI part 2: Scaling MCP best practices for seamless developers’ experience in the IDE with Cline appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/mcp-best-practices-cline-live-debugger-developer-experience/feed/ 0
The rise of agentic AI part 1: Understanding MCP, A2A, and the future of automation https://www.dynatrace.com/news/blog/agentic-ai-how-mcp-and-ai-agents-drive-the-latest-automation-revolution/ https://www.dynatrace.com/news/blog/agentic-ai-how-mcp-and-ai-agents-drive-the-latest-automation-revolution/#respond Tue, 13 May 2025 07:40:45 +0000 https://www.dynatrace.com/news/?p=69029 multiple robot icons linked like a network on a dark background asking the question, what is agentic AI? And what is Model Context Protocol? also represents AI agent observability and Amazon Bedrock agents monitoring

Agentic AI systems—independent AI agents that perform tasks by reasoning, learning, and adapting—are radically changing how enterprises automate tasks and orchestrate complex workflows. In this first installment of our series, The Rise of Agentic AI, we explore agentic AI and how the agents communicate using Agent2Agent and model context protocol (MCP).

The post The rise of agentic AI part 1: Understanding MCP, A2A, and the future of automation appeared first on Dynatrace news.

]]>
multiple robot icons linked like a network on a dark background asking the question, what is agentic AI? And what is Model Context Protocol? also represents AI agent observability and Amazon Bedrock agents monitoring

By now, everyone is aware of generative AI fueled by large language models (LLMs) and generative pre-trained transformers (GPTs). The next level of innovation is agentic AI and the autonomous AI agents that drive it. Using Model Context Protocol (MCP) to facilitate agent-to-agent communication, these systems are revolutionizing how enterprises automate tasks and orchestrate complex workflows.

Powered by LLMs, vector databases, retrieval augmented generation (RAG) pipelines and additional tools, these AI agents are expanding extensively, giving rise to multi-agent systems, cross-agent protocols, and context-sharing standards. But these autonomous agents also introduce new challenges in monitoring, debugging, and security.

We’ll examine in detail the fundamentals of AI agents, models, and the emerging standards that help them communicate, like Agent2Agent (A2A) and Model Context Protocol (MCP).

Key takeaways:
  • Autonomous AI agents are the backbone of agentic AI. These services combine to deliver adaptable automated tasks.
  • AI agents depend on LLMs and orchestration logic. These technologies maintain the agent’s state, session memory, context, and reasoning strategies.
  • Agents depend on protocols, such as A2A and MCP, to effectively communicate. Models and agents need these protocols to manage multi-agent communication.

What is agentic AI?

Agentic AI is an artificial intelligence system made up of independent agents that can take initiative and perform sequences of actions to complete tasks by reasoning, learning, and adapting to changing circumstances.

Dynatrace Chief Technologist Alois Reitbauer described agentic AI this way:

Alois Reitbauer

“It’s really delegating a task to software the way you would delegate it to a human. Say if you wanted to do travel booking, give it some complexity and freedom and some decision points it can make. Like, I have to go to Vegas, I need a hotel, I need a couple of good restaurants to go to, we’re going to be 50 people, fix it with my schedule.”
– Alois Reitbauer in The New Stack

Agentic AI systems rely on AI agents to perform the tasks that lead to the desired outcome.

What are AI agents?

An AI agent is a self-directed autonomous application that harnesses large language model (LLM) reasoning, tool usage, and context-awareness from numerous data sources to carry out tasks.

Agents can think and act independently without outside intervention. Agents can think through chain-of-thought, plan, execute (Reason+Act=ReAct), and refine their actions as needed. Businesses are looking into adopting these autonomous agents for applications such as customer service automation, supply-chain optimization, and content generation.

How do AI agents operate?

AI agents operate similarly to a Michelin-starred chef in a busy kitchen: They continuously gather information, plan, execute, and adjust to reach their desired end goal.

In the chef analogy, the cook surveys orders and available ingredients, decides on a suitable recipe, and then refines the approach based on feedback or resource constraints.

Agents do the same thing in a computational context. Specifically, they observe the world (for example, a user request or a set of data), perform internal reasoning about the best course of action, then carry out the steps needed to fulfill the request. This cycle allows them to respond adaptively to changing conditions, much as a chef would substitute ingredients or modify a dish mid-preparation.

Underpinning this iterative loop is the orchestration layer, which maintains the agent’s state, session memory, and reasoning strategies (such as ReAct, Chain-of-Thought, or Tree-of-Thoughts). Large language models (such as OpenAI’s GPT, Anthropic Claude, Google Gemini, Amazon Nova) provide the core reasoning capability for the agent. The model “thinks” about the user’s query. But the agent gains its power by incorporating additional frameworks or tools that can fetch external information or execute actions in the real world. One way to fetch and provide tools and information is through a unified protocol called Model Context Protocol (MCP).

Additionally, the orchestration layer ensures that multiple rounds of reasoning, tool usage, and tool outputs are all tracked and synthesized before the agent returns a final response to the user. Agents follow these steps in a structured way, so they can produce more accurate, context-rich answers and easily manage complex tasks.

architecture diagram that shows multiple agents interacting with an agentic application
Figure 1. Autonomous agent workflows and task execution.

What is the difference between models and agents?

A model (like a large language model) simply generates outputs based on its training data and the given prompt, typically without any built-in mechanism for session memory, external actions, or complex decision loops and validations.

An agent, on the other hand, includes the model but goes further. It maintains a stateful process (managing multi-turn conversations and thought processes), uses external tools to gather fresh data or perform actions, and follows a defined orchestration logic (such as ReAct and chain-of-thought). Thus, while a model is a core reasoning component, an agent adds the surrounding structure and capabilities needed for autonomous, goal-directed behavior.

What is Agent2Agent (A2A)? How multiple agents communicate with each other

As enterprises slowly adopt multiple specialized agents, interoperability of these services becomes crucial to create reliable experiences. To achieve this, A2A from Google helps to create an open protocol that enables agents—regardless of vendor or framework—to securely exchange information, coordinate actions, and integrate capabilities. By specifying tasks, capabilities, and artifacts in a standardized JSON-based lifecycle model, A2A fosters multi-agent collaboration across otherwise siloed systems.

A2A protocol enables agents to share updates and delegate tasks without overhead. However, direct communication between agents only solves half the problem: These agents also need relevant, up-to-date data and context to drive decisions and be equipped with the right toolset to execute actions.

Without a unified method for accessing diverse data sources, even the most capable multi-agent ecosystem remains limited in scope. The open-source project Model Context Protocol (MCP) fills this gap.

architecture diagram showing two agents using different protocols communicating using A2A protocol as part of an AI agent monitoring and MCP monitoring scheme.
Figure 2. Agent-to-agent communication.

What is Model Context Protocol? How MCPs empower agents

As an open standard, the Model Context Protocol (MCP) connects AI agents to relevant data sources, such as repositories, tools, or external APIs. Instead of the above mentioned integrations for each data silo, MCP provides a universal interface like USB-C to connect multiple relevant sources to feed the right context to the models and agents. This universality simplifies how agents access relevant context, leading to better task outcomes, execution and more consistent performance across complex environments. For managing complex tasks like the ones highlighted above, the Dynatrace MCP server on GitHub helps to get real-time end-to-end observability and MCP data into your daily workflow.

architecture diagram showing Dynatrace MCP monitoring reference architecture
Figure 3. Dynatrace MCP server reference architecture.

What’s next: Monitoring A2A and MCP for better agentic AI

As these technologies evolve, we can expect deeper integrations between agent orchestration protocols (A2A and MCP) and open observability frameworks, delivering end-to-end visibility from data ingestion to cross-agent collaboration. Likewise, as standards converge, organizations will rapidly compose advanced AI solutions while retaining full transparency and control, paving the way for even greater scalability, resilience, and confidence in autonomous agents.

Read more

  • Part two of the Rise of Agentic AI blog series explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.
Check out Dynatrace MCP and Dynatrace AI Observability for AI agent monitoring and MCP monitoring at scale.

The post The rise of agentic AI part 1: Understanding MCP, A2A, and the future of automation appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/agentic-ai-how-mcp-and-ai-agents-drive-the-latest-automation-revolution/feed/ 0
Power dashboarding part 2: Dynatrace Dashboards tutorial to gain better, faster answers using AI and formatting https://www.dynatrace.com/news/blog/dynatrace-dashboard-tutorial-ai-and-formatting/ https://www.dynatrace.com/news/blog/dynatrace-dashboard-tutorial-ai-and-formatting/#respond Mon, 31 Mar 2025 13:00:42 +0000 https://www.dynatrace.com/news/?p=68452 Abstract image showing charts and graphs as part of the Dynatrace dashboard tutorial to filter data and demonstrate dashboard filtering

Part 2 of our power dashboarding series delves into how to leverage AI to make additional insights available at your fingertips. In this Dynatrace Dashboards tutorial, we'll also explore formatting options to enhance the aesthetics of your dashboards and guide the user's eye to the most relevant information. By the end, you'll have a visually appealing, AI-enhanced dashboard that delivers better, faster answers. Let's dive in and make your data shine!

The post Power dashboarding part 2: Dynatrace Dashboards tutorial to gain better, faster answers using AI and formatting appeared first on Dynatrace news.

]]>
Abstract image showing charts and graphs as part of the Dynatrace dashboard tutorial to filter data and demonstrate dashboard filtering


Welcome back to our power dashboarding blog series, data enthusiasts! Today, our Dynatrace Dashboards tutorial will dive into some exciting features to level up our dashboards with little effort:

  • Create charts effortlessly with Davis CoPilot™ using a natural language interface.
  • Analyze your charts with AI to gain instant insights into trends and anomalies.
  • Guide users with highlight formatting based on conditions, optimized visuals, and icons/emojis.

You can either continue with the custom infrastructure metrics dashboard you created in Part I or use the dashboard we prepared here (Dynatrace login required). By the end of this session, your enhanced dashboard will look like this:

Dashboard that shows the final product of the Dynatrace dashboard tutorial
Figure 1. Our enhanced host monitoring dashboard that highlights disk usage includes AI forecasting for CPU usage.

Looking for something? Query your data with natural language

Davis CoPilot is an excellent virtual assistant that helps you create queries using natural language. While the Explore interface is useful for quickly visualizing known metrics, Davis CoPilot is great for exploring your data when you know your desired outcome but are unfamiliar with the available data. exploring your data when you know your desired outcome but are unfamiliar with the available data.

In our Dynatrace Dashboards tutorial, we want to add a chart that shows the bytes in and out per host over time to enhance visibility into network traffic. By tracking these metrics, we can identify any unusual spikes or drops in network activity, which might indicate performance issues or bottlenecks. To simplify the data and make it more readable, we’ll show only the entity names instead of the host IDs. This approach helps you quickly pinpoint potential problems and ensures efficient monitoring of your infrastructure.

Since we’re unsure which metric to use, we turn to CoPilot, using natural language to express what we’re looking for. For this tile, we’ll use the following prompt:

“Show me the bytes in and out by host over time, add the entity names for the hosts, and remove the host id.”

Detail screen showing how to add a Davis CoPilot tile to your dashboard.
Figure 2. Add a Davis CoPilot tile to your dashboard.

Create a chart with Davis CoPilot, visualizing bytes in/bytes out

  1. Select + to create a Davis CoPilot tile.
  2. In the options pane, select the Data tab and enter the natural language prompt above in the text field. Select Run.
  3. Go to the Visual tab and update the visualization to Categorical. The resulting tile displays the transmitted and received bytes, split by host. By default, dashboards show data from the last two hours. You can adjust the timeframe at any time using the selector located in the upper right corner of your dashboard.
    Result of adding the chart generated by the Davis CoPilot query.
    Figure 3. Add Transmit/Receive (Tx/Rx) bytes using Davis CoPilot.

That’s how easy querying your data and creating charts with Davis CoPilot can be. For more information on optimizing your prompts and best practices, check out the topic Tips for writing better prompts.

Now, let’s take our Dynatrace Dashboards tutorial a step further by using AI to analyze our charts and gain additional insights with the predictive capabilities of the Davis® AI Analyzer.

Upgrade your chart with an AI-powered trend forecast

In our dashboard, we already have a tile displaying CPU usage. Let’s enhance this chart with a trend outlook to gain insights into future CPU usage patterns. By monitoring and predicting CPU usage, we can ensure optimal performance, prevent potential bottlenecks, and proactively manage resources to maintain smooth system operations.

Add Davis® AI Analyzer to the CPU usage % chart

  1. Go to the Data tab and expand the Davis AI section at the bottom.
  2. Select the Activate Davis AI Analyzer toggle and select Forecast from the drop-down menu.
  3. Leave the default parameters and select Run.
  4. Go to the Visual tab and select the Davis AI Analysis category.
  5. Select the Chart option from the Davis AI Analysis category.
    Video of the Dynatrace dashboard tutorial showing how to add the Davis AI forecasting line chart
    Figure 4. Add Forecasting powered by Davis® AI to a line chart.

You may have noticed the option to activate anomaly detection when adding forecasting. This feature provides immediate insights into unusual spikes or drops in your telemetry data. For more information, have a look at our recent blog post on utilizing Davis AI on dashboards.

Formatting Dynatrace Dashboards for enhanced usability and faster insights

By adding formatting to dashboards, you can improve usability, highlight important information, and make data easy to digest. Dashboards offer various ways to enhance data appearance, from simple color updates to conditional formatting based on rules.

To make it easy for users to quickly identify what’s important, format your charts to clearly display that information. We’ll now review and improve each dashboard tile to ensure it is user-friendly and easy to understand.

Access the customization options by going to the Visual tab of your selected tile.

Here, you can reassign data mapping, modify data representation, adjust the X/Y axis, change data colors, set thresholds, and assign units and formats to provide more context to your data. Depending on the selected visualization, you have various options for modifying the tile’s visuals.

Improve the DNS errors single-value tile

Let’s start by adjusting the visuals of the DNS errors tile. We want to make the information instantly clear to the user and modify the sparkline with ticks to better illustrate how the metric changes over time.

  1. Select the tile and go to the Visual tab.
  2. Expand the Single value section.
  3. Using the Show label option, update the label to DNS errors.
  4. Enable the Show icon option and select Network icon.
  5. Expand the Trend section. Update the arrow colors to red, blue, and green. Red indicates that the number of errors is increasing.
  6. Expand the Sparkline section. Set the Variant option to Area and enable Show ticks.
    Video showing how to fix visuals for the single value tile.
    Figure 5. Improve visuals for the DNS errors single-value tile.

Improving the Inodes pie chart

For the pie chart displaying the total number of Inodes, we’ll adjust the colors to ensure consistency across the dashboard and minimize distractions. Blue will be used as the primary color to maintain a uniform palette. In other instances, it often also makes sense to use colors to indicate functional differences or common categories across different charts.

  1. Select the tile and go to the Visual tab.
  2. Select the Color section and expand it.
  3. Expand the Series colors option and select the Blue-steel category.
    Total of Inodes tile showing blue-steel color optimizations in the Dynatrace dashboard tutorial
    Figure 6. We improve the color palette of the pie chart.

Improving the Memory used % chart

To further reduce noise in our Dynatrace Dashboards tutorial, we’ll hide all displayed legends related to entity hostnames, retain only the essential fields, and update the color palette.

  1. Select the tile and go to the Visual tab.
  2. Select the Y-axis section, expand it and delete from the label to only leave entity.name.
  3. Go to the Legend and tooltip section and expand it.
  4. On the Displayed fields option, expand the list and uncheck dt.entity.host field.
  5. On the Show legend option, disable the toggle.
  6. Go to the Color section and expand it.
  7. In the Series colors option, select the list to expand it.
  8. Select the Blue-steel category.
    Memory usage percentage tile showing optimized colors as a result of the Dynatrace dashboards tutorial.
    Figure 7. Optimize the labels and colors of the memory used % chart.

Use conditional formatting to highlight information in tables

Using conditional formatting in a table helps highlight important data points and trends, making it easier to spot anomalies and patterns at a glance. It also enhances readability by visually distinguishing between values based on predefined criteria.

Let’s enhance the visuals for the table displaying disk usage. We’ll hide some of the fields currently displayed and add a threshold for average disk usage by host using conditional formatting.

  1. Select the Columns section and expand it.
  2. On the Displayed fields option, select the list and uncheck the timeframe and interval fields.
  3. Go to the Cells section, expand it and change the Apply threshold color to option from Value to Background.
  4. Go to the Threshold section and select it.
  5. Select + Threshold. This will add a new threshold with three predetermined rules with empty values.
  6. On the Field option, select value.A.
  7. Change the three operators for the rules from to .
  8. Modify each of the Value fields as displayed in the image below:
    • Green: No action required.
    • Yellow: Attention, potentially action required.
    • Red: Warning, action required.
  9. Go back to the Cells section and enable the Show threshold in row option.
    Detail screen showing color coding for different thresholds
    Figure 8. Add conditional formatting to the table to indicate levels of concern.
    Video that shows how to apply conditional formatting to the table.
    Figure 9. Apply conditional formatting to a table.

Add icons and emojis in markup tiles

Including icons on your dashboard can enhance visual appeal and make it easier to quickly identify and interpret key information. Icons can also improve user experience by providing intuitive visual cues that guide users through the data. Dashboards support emojis and icons for that purpose, making it easier for users to identify the information they’re looking for quickly.

Add emojis/icons

  1. Select Windows key + . (period) on Windows or Fn + e on Mac in any markdown or text field.
  2. We’ll add a 🖥️ computer icon to the main title of our dashboard and incorporate more emojis in the subtitles. This will help users to faster navigate to the information they’re looking for.
    Dynatrace dashboards tutorial video showing how to add icons to markdown tiles.
    Figure 10. Add icons to markdown tiles.

Dynatrace Dashboards tutorial: Where to go from here

With this Dynatrace Dashboard tutorial, you now have a solid understanding of dashboarding basics and are familiar with several customization options.

If you want to continue your learning journey, here are some recommended sources for further reading.

The post Power dashboarding part 2: Dynatrace Dashboards tutorial to gain better, faster answers using AI and formatting appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-dashboard-tutorial-ai-and-formatting/feed/ 0
Deliver secure, safe, and trustworthy GenAI applications with Amazon Bedrock and Dynatrace https://www.dynatrace.com/news/blog/deliver-secure-safe-and-trustworthy-genai-applications-with-amazon-bedrock-and-dynatrace/ https://www.dynatrace.com/news/blog/deliver-secure-safe-and-trustworthy-genai-applications-with-amazon-bedrock-and-dynatrace/#respond Wed, 12 Mar 2025 18:56:47 +0000 https://www.dynatrace.com/news/?p=68271 Gen AI graphic

Every software development team grappling with Generative AI (GenAI) and LLM-based applications knows the challenge: how to observe, monitor, and secure production-level workloads at scale. Traditional debugging approaches, logs, and occasional remote breakpoint instrumentation can’t easily keep pace with cloud-native AI deployments, where performance, compliance, and costs are all on the line. How can you […]

The post Deliver secure, safe, and trustworthy GenAI applications with Amazon Bedrock and Dynatrace appeared first on Dynatrace news.

]]>
Gen AI graphic

Every software development team grappling with Generative AI (GenAI) and LLM-based applications knows the challenge: how to observe, monitor, and secure production-level workloads at scale. Traditional debugging approaches, logs, and occasional remote breakpoint instrumentation can’t easily keep pace with cloud-native AI deployments, where performance, compliance, and costs are all on the line. How can you gain insights that drive innovation and reliability in AI initiatives without breaking the bank?
Dynatrace helps enhance your AI strategy with practical, actionable knowledge to maximize benefits while managing costs effectively.

Amazon Bedrock, equipped with Dynatrace Davis® AI and LLM observability, gives you end-to-end insight into the Generative AI stack, from code-level visibility and performance metrics to GenAI-specific guardrails.

Developers deserve a frictionless troubleshooting experience and fast access to real-time data—no more guesswork or costly redeployments. Here’s how Dynatrace, combined with Amazon Bedrock, arms teams with instant intelligence from dev to production, helping to accelerate innovation while keeping performance, costs, and compliance in check.

Introducing Amazon Bedrock and Dynatrace Observability

Amazon Bedrock is a serverless service for building and scaling Generative AI applications easily with foundation models (FM). It provides an easy way to select, integrate, and customize foundation models with enterprise data using techniques like retrieval-augmented generation (RAG), fine-tuning, or continued pre-training.

Dynatrace is an all-in-one observability platform that automatically collects production insights, traces, logs, metrics, and real-time application data at scale.  With powerful Davis AI engine Dynatrace notifies teams about production-level issues before they disrupt users, helps predict resource usage,costs, and performance issues, and delivers guardrails that protect data and maintain compliance.

Together, Amazon Bedrock and Dynatrace provide an end-to-end observability solution for AI applications:

  • Predictive operations: Proactive usage and cost forecasting to reduce unexpected operational expenses and token usage.
  • Production performance monitoring: Service uptime, service health, CPU, GPU, memory, token usage, and real-time cost and performance metrics.
  • Guardrail analysis: Detect hallucinations, track prompt injections, mitigate PII leakage, and ensure brand-safe outputs.
  • Full-stack tracing: Track each user request across multiple FMs, vector databases, orchestrators (LangChain), and custom business logic.
  • Compliance: Document all inputs and outputs, maintaining full data lineage from prompt to response to build a clear audit trail and ensure compliance with regulatory standards.

Video overview of Amazon Bedrock dashboard with Dynatrace AI and LLM Observability solution
Figure 1. Video overview of Amazon Bedrock dashboard with Dynatrace AI and LLM Observability solution.

How it works

Dynatrace seamlessly instruments your LLM-based workloads using Traceloop OpenLLMetry, which augments standard OpenTelemetry data with AI-specific KPIs (for example, token usage, prompt length, and model version).

Combined with Amazon Bedrock, you can:

  • Spin up your AI model on Amazon Bedrock—choose from providers like AI21, Anthropic, Cohere, Stability AI, Mistral AI, Meta, or Amazon’s own Nova/Titan foundation models.
  • Automatically instrument your application with OpenTelemetry.
  • Configure OpenLLMetry to capture specialized LLM details as spans and metrics, like model name, completion time, token count, token cost, and prompt text.
  • Send unified data to Dynatrace for analysis alongside your logs, metrics, and traces.

Behind the scenes, Dynatrace merges the standard telemetry with these advanced AI attributes, surfaces them in real-time dashboards, and applies AI-driven analytics to discover anomalies, forecast usage costs, and diagnose root causes.

Distributed Tracing overview of an Amazon Bedrock request with LangChain
Figure 2. Distributed Tracing overview of an Amazon Bedrock request with LangChain.

How to set up and instrument your data with OpenLLMetry

Traceloop OpenLLMetry is an open source extension that standardizes LLM and Generative AI data collection. By layering on top of OpenTelemetry standards, OpenLLMetry captures the critical metrics you can’t get by default—like the number of tokens, model temperature, or guardrail triggers.

Here’s how to set it up for Amazon Bedrock:

  1. Install OpenLLMetry in your Python or Node.js environment:
 pip install traceloop-sdk
from traceloop.sdk import Traceloop

headers = {

'Authorization': f"Api-Token {environ.get('DYNATRACE_TEAM_KEY')}"

}

Traceloop.init(

app_name=environ.get('DYNATRACE_APP_NAME'),

api_endpoint=environ.get('DYNATRACE_URL'),

headers=headers

)
  1. Configure environment variables to send data to Dynatrace via your ingest token:
 DYNATRACE_URL =https://123abcde.live.dynatrace.com/api/v2/otlp 

DYNATRACE_TEAM_KEY=dt0.....
  1. You can optionally add OpenLLMetry decorators or instrumentation to your LLM calls (for example, with LangChain or direct Bedrock SDK calls).

When your application queries Amazon Bedrock, OpenLLMetry automatically captures:

  • Prompt tokens vs. completion tokens
  • Finish reason (did the LLM stop due to a user request, or was the max token limit reached?)
  • Model type (which Amazon foundation model or third-party model is used?)
  • Performance: Response time, throughput, and error rate
    • Guardrail activations: Toxicity, PII, denied topics, and hallucinations
    • System, prompt, and completion messages and roles

This data is instantly correlated in Dynatrace so you can visualize or alert on critical thresholds (for example, if your average token usage spikes or your overall cost forecast grows beyond budget).

Overview of observability data flowing into Dynatrace from a travel agent application running in a Kubernetes cluster powered with Amazon Bedrock, where OpenLLMetry instruments the data
Figure 3. Overview of observability data flowing into Dynatrace from a travel agent application running in a Kubernetes cluster powered with Amazon Bedrock, where OpenLLMetry instruments the data.

How to debug incorrect responses in production

Let’s walk through a real-world scenario:

Your production travel agent application—powered by Amazon Bedrock and Dynatrace—gives users incorrect travel recommendations. Perhaps it suggests flights or hotels that don’t exist or mixes up time zones. This isn’t just a minor inconvenience; it jeopardizes user experience and can directly impact revenue and trust.

Here’s how Dynatrace helps you trace and resolve the issue quickly:

Proactive alerting with Davis AI

You receive an alert from Dynatrace Davis AI anomaly detection indicating incorrect system behavior. There might be a spike in “incorrect itinerary” complaints or conversation outcomes flagged as “nonsensical.” Davis AI correlates the unusual LLM responses with application telemetry and usage patterns, so you immediately know something is off in the recommendation flow.

Full-stack end-to-end tracing

In Dynatrace Distributed Tracing, you see the entire transaction trace for the affected user session. This includes front-end requests, back-end aggregator logic, calls to Amazon Bedrock, and any vector database lookups performed for retrieval-augmented generation (RAG). Rather than sifting through multiple logs, you have a single timeline that reveals exactly where the LLM call returned unexpected data.

Inspecting the GenAI model details

By drilling down into the span data enriched by OpenLLMetry, you can see:

  • Prompt and completion text and tokens used.
  • The specific foundation model version (for example, anthropic.claude-v1 or amazon.nova).
  • Temperature setting and max token limits.
  • Any error codes or guardrail triggers.

This clarity helps you pinpoint if the model produces off-base recommendations because of a misaligned temperature, an out-of-date context, or a mismatch in user inputs.

Root cause analysis

With Dynatrace, you quickly correlate the LLM anomaly to a specific function in your microservice code. You discover that an external data source used for itinerary validation had missing or stale updates, causing the LLM prompt to reference invalid flights. You’ve found the “why” without manually spelunking logs in disparate systems.

Resolving and validating

A fix might involve updating your data pipeline or refining the prompt logic. You can deploy the change and watch in near real-time as Dynatrace collects new traces and logs. Davis AI recognizes that the anomaly is cleared, confirming that your fix resolved the incorrect responses—no guesswork required.

You can find the code example for our travel agent application here for review, and the dashboard on our Dynatrace Playground instance.

Overview of Amazon Bedrock service health, performance, quality, and guardrails
Figure 4. Overview of Amazon Bedrock service health, performance, quality, and guardrails.

Summary

By integrating Amazon Bedrock with Dynatrace end-to-end observability, you not only catch issues early but also trace them across your entire AI stack to the root cause. Building or scaling Generative AI applications with Amazon Bedrock requires robust insights into your environment—from model usage and performance metrics to cost forecasts and guardrail efficacy.

Dynatrace helps you scale with:

  • Complete end-to-end tracing across your services, external data pipelines, and LLM calls.
  • Predictive analytics that forecast AI resource usage and cost trends, letting you proactively manage budgets.
  • Unified dashboards that bring performance, cost, code-level data, logs, metrics, and audit events together.
  • Compliance and governance that integrate security checks, data masking, and guardrail analysis.

Whether you’re a developer racing to put your latest AI-powered application or a new feature into production or an SRE ensuring your system meets enterprise-grade SLAs, Dynatrace and Amazon Bedrock help you to create frictionless AI applications, focusing on performance and observability—at any scale, in production, with no downtime.

Useful resources

The post Deliver secure, safe, and trustworthy GenAI applications with Amazon Bedrock and Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/deliver-secure-safe-and-trustworthy-genai-applications-with-amazon-bedrock-and-dynatrace/feed/ 0
Better dashboarding with Dynatrace Davis AI: Instant meaningful insights https://www.dynatrace.com/news/blog/better-dashboarding-with-dynatrace-davis-ai/ https://www.dynatrace.com/news/blog/better-dashboarding-with-dynatrace-davis-ai/#respond Tue, 21 Jan 2025 21:38:07 +0000 https://www.dynatrace.com/news/?p=67370 abstract image showing connected dots and waves representing MCP best practices for agentic AI

Discover the value of Davis® AI when working with dashboards for observability, security, or business use cases. Quickly spot anomalies by activating Davis AI on any numeric time series chart data. Stay ahead with visual, AI-powered forecasting, or get new insights into your data with just a few clicks by leveraging Davis CoPilot™.

The post Better dashboarding with Dynatrace Davis AI: Instant meaningful insights appeared first on Dynatrace news.

]]>
abstract image showing connected dots and waves representing MCP best practices for agentic AI

Ensuring smooth operations is no small feat, whether you’re in charge of application performance, IT infrastructure, or business processes. Chances are, you’re a seasoned expert who visualizes meticulously identified key metrics across several sophisticated charts. Your trained eye can interpret them at a glance, a skill that sets you apart.

However, your responsibilities might change or expand, and you need to work with unfamiliar data sets. The market is saturated with tools for building eye-catching dashboards, but ultimately, it comes down to interpreting the presented information. This is where Davis AI for exploratory analytics can make all the difference.

Activate Davis AI to analyze charts within seconds
Figure 1. Activate Davis AI to analyze charts within seconds

Davis AI can help you expand your dashboards and dive deeper into your available data to extract additional information. Our customers value the nearly unlimited possibilities for querying and joining data on the Dynatrace platform, with the option of instant, real-time visualization of query results. Whether you’re an expert or an occasional user, our recently launched Davis CoPilot will enable you to get instant results without the need to write complex queries yourself. Have a look at our recent Davis CoPilot blog post for more information and practical use cases.

If you’ve already created your dashboards, now is the time to use Davis AI to identify anomalies or predict future trends without restricting use cases.

Leverage Davis AI for anomaly detection and instant insights

“My chart shows a peak at 8:00 AM. Do I need to investigate this further?” You might be regularly confronted with this or similar questions. Davis AI machine learning capabilities will help you identify actual anomalies within seconds, enabling you to focus resources on issues that matter.

Based on your requirements, you can select one of three approaches for Davis AI anomaly detection directly from any time series chart:

  • Auto-Adaptive Threshold: This dynamic, machine-learning-driven approach automatically adjusts reference thresholds based on a rolling seven-day analysis, continuously adapting to changes in metric behavior over time. For example, if you’re monitoring network traffic and the average over the past 7 days is 500 Mbps, the threshold will adapt to this baseline. An anomaly will be identified if traffic suddenly drops below 200 Mbps or above 800 Mbps, helping you identify unusual spikes or drops.
  • Seasonal Baseline: Ideal for metrics with predictable seasonal patterns, this option leverages Davis AI to create a confidence band based on historical data, accounting for expected variations. For instance, in a web shop, sales might vary by day of the week. Using a seasonal baseline, you can monitor sales performance based on the past fourteen days. An anomaly is identified if sales on a Friday are significantly lower than on previous Fridays, indicating a potential issue.
  • Static Threshold: This approach defines a fixed threshold suitable for well-known processes or when specific threshold values are critical. For example, if you have an SLA guaranteeing 95% uptime, you can set a static threshold to alert you whenever uptime drops below this value, ensuring you meet your service commitments.

Davis AI is particularly powerful because it can be applied to any numeric time series chart independently of data source or use case.

The following example will monitor an end-to-end order flow utilizing business events displayed on a Dynatrace dashboard. By leveraging Davis AI anomaly detection, we can identify potentially fraudulent behavior by activating anomaly detection on the Average order size chart. As shown in the chart below on the lower left, most values fall within the band of acceptable response time (highlighted in green), with only one spike occurring at 5:00 AM. Since this spike was outside the expected range, an anomaly was identified.

Apply Davis AI anomaly detection to detect fraudulent behavior in a business process
Figure 2. Apply Davis AI anomaly detection to detect fraudulent behavior in a business process
  • Application Observability: Identify unexpected error rate increases in application performance, helping pinpoint and resolve issues quickly.
  • Digital Experience Management: Monitor user interaction patterns to spot anomalies in website or app performance that could affect user experience, such as slow page load times.
  • FinOps: Track irregularities in cloud spending or resource usage, enabling cost optimization and preventing budget overruns.

Davis AI forecast analysis predicts future numeric values of any time series. It can even process external datasets or the results of any data query if it can be displayed as a numeric time series, such as occurrences over time.

The forecast is created instantly, even for large data sets, and updates dynamically whenever filter settings are changed.

In application performance management, acting with foresight is paramount. Maintaining reliability and scalability requires a good grasp of resource management; predicting future demands helps prevent resource shortages, avoid over-provisioning, and maintain cost efficiency.

On this SRE dashboard, we utilize Davis AI to forecast and visualize future resource utilization:

SRE dashboard monitoring the four golden signals and forecasting resource utilization
Figure 3. SRE dashboard monitoring the four golden signals and forecasting resource utilization

Other potential applications for forecasting include:

  • Kubernetes: Forecasting helps dynamically scale Kubernetes clusters by predicting future resource needs. This ensures optimal resource utilization and cost efficiency. Forecasting can identify potential anomalies in node performance, helping to prevent issues before they impact the system.
  • Business: Using information on past order volumes, businesses can predict future sales trends, helping to manage inventory levels and effectively plan marketing strategies.

AIOps: Utilize Davis AI to predict and prevent

Utilizing the Dynatrace AutomationEngine, Davis AI forecasting capabilities can even trigger automated actions. One of our customers’ SRE teams needed to increase disk space to avoid ongoing over- and under-provisioning, which was time-consuming and annoying. Now, with Davis AI forecasting capabilities, the target disk size is predicted automatically, and an automated task for disk resizing is triggered when necessary.

If you want to further explore the possibilities for prediction and prevention management with Dashboards, have a look at our example dashboard in the Dynatrace Playground.

Prevent incidents through predictive maintenance and capacity management
Figure 4. Prevent incidents through predictive maintenance and capacity management

Experience Davis AI in action

To experience the possibilities of Davis AI, look at this short introduction video by Andreas Grabner:
How to chart and forecast any data point

To explore the depth of functionality of Dynatrace Dashboards yourself and get first-hand experience, try out the app in the Dynatrace Playground.

The post Better dashboarding with Dynatrace Davis AI: Instant meaningful insights appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/better-dashboarding-with-dynatrace-davis-ai/feed/ 0
How to implement an AIOps strategy at scale https://www.dynatrace.com/news/blog/how-to-implement-an-aiops-strategy-at-scale/ https://www.dynatrace.com/news/blog/how-to-implement-an-aiops-strategy-at-scale/#respond Fri, 20 Dec 2024 15:42:15 +0000 https://www.dynatrace.com/news/?p=67149 AIOps strategy

Imagine a day when your IT team resolves critical incidents before users even notice them. In today’s multicloud world, complexity is the norm. But what if an AIOps strategy could transform that complexity into your organization’s greatest advantage? IT operations teams must monitor, maintain, and optimize a broader mix of applications, infrastructure, and technologies, with […]

The post How to implement an AIOps strategy at scale appeared first on Dynatrace news.

]]>
AIOps strategy

Imagine a day when your IT team resolves critical incidents before users even notice them. In today’s multicloud world, complexity is the norm. But what if an AIOps strategy could transform that complexity into your organization’s greatest advantage?

IT operations teams must monitor, maintain, and optimize a broader mix of applications, infrastructure, and technologies, with newer digital business solutions only adding to the strain. To manage it, organizations are turning to AI to help automate tasks such as anomaly detection, root-cause analysis, and incident response. But some IT operations teams are taking this approach a step further, applying multiple forms of AI to accelerate and enhance all business operations.

As AI is evolving, it’s important for organizations to understand the different types of AI and the keys to implementing them to achieve an AIOps strategy that moves from reactive to predictive problem solving.

The three kinds of AI that are key to a successful AIOps strategy

AI has evolved beyond the traditional correlation and probability-based approaches. Now, organizations turn to multiple forms of AI, such as causal, generative, and predictive AI, to manage cloud environments, secure data, and improve business decision-making. Therefore, implementing a successful AIOps strategy requires a deeper understanding of these types of AI.

Causal AI: Think of a global retail chain instantly pinpointing the root cause of a checkout slowdown across thousands of stores. Causal AI uses real-time, contextual data and causal dependencies for precise root-cause analysis and issue prevention. This establishes a business environment safeguarded by automated health monitoring and risk remediation.

Predictive AI: Imagine anticipating server outages hours before they occur, allowing seamless customer experiences. Predictive AI analyzes data patterns and trends, using statistical algorithms and other advanced machine learning techniques to anticipate future system behavior. This means technical users and business leaders can be much more proactive and innovative in the face of constant change.

Generative AI: Consider using generative AI to automate repetitive tasks, freeing up teams for innovation. Generative AI trains on large and diverse data sources, boosting business productivity and efficiency.  But when used in combination with causal and predictive AI and trained on real-time, high-fidelity observability data, generative AI can help organizations accelerate productivity and automate workflows.

The key strategy is using GenAI in conjunction with causal and predictive AI to understand your data and environment using natural conversation, and not as a way to correlate disparate events or anticipate the future. GenAI alone has limitations in these areas.

Implementing AIOps at scale

The multicloud complexity challenge demands a new approach—one that moves beyond toolchains that are stitched together. That’s where Davis AI™ comes in, combining the strengths of three AI capabilities in one: predictive, causal, and generative AI. Davis provides advanced analytics and proactive problem-solving to deliver out-of-the-box insights. Additionally, Davis doesn’t require extensive configuration or integrations. As a result, IT teams and business decision-makers can take immediate advantage of AI and scale it across the organization.

This power-of-three AI and observability approach democratizes the value of complex data analytics for both technical and nontechnical users. By unifying operations, security, development, and business teams with a single, complete, real-time view of activity across every cloud, app, and system, teams have an intuitive, self-service observability environment for problem-solving security and business events. Unified observability and AI put users in control of finding the answers they need from data, without needing to be experts in query languages.

Take your AIOps strategy from reactive to proactive

In addition to unifying teams, Dynatrace enables organizations to cut down on redundant observability tools and vendors, simplifying the overall view and improving insight coherency. And with new standards and capabilities in AI, teams can scale these efforts across the business for a more streamlined, inclusive, and collaborative business approach to resolving operational issues and driving your business forward.

Tool sprawl is an obvious problem for organizations. But choosing the wrong end-to-end observability platform can make things worse—and more expensive. Contact us today to request a demo of the AI-powered unified observability and security platform from Dynatrace and take your first step toward eliminating tool sprawl.

eBook: Developing an AIOps strategy for cloud observability

Download our free eBook to learn the best practices for developing an AIOps strategy that drives efficiency, innovation, and better business outcomes

The post How to implement an AIOps strategy at scale appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-to-implement-an-aiops-strategy-at-scale/feed/ 0
Transform data into insights with Dynatrace Dashboards and Notebooks https://www.dynatrace.com/news/blog/transform-data-into-insights-with-dynatrace-dashboards-and-notebooks/ https://www.dynatrace.com/news/blog/transform-data-into-insights-with-dynatrace-dashboards-and-notebooks/#respond Wed, 16 Oct 2024 18:45:55 +0000 https://www.dynatrace.com/news/?p=66227 Explore Kubernetes metrics graphic

When we launched the new Dynatrace experience, we introduced major updates to the platform, including Grail™, our innovative data lakehouse unifying observability, security, and business data, and Dynatrace Query Language (DQL) for accessing and exploring unified data. While Grail and DQL opened up nearly limitless possibilities for data exploration, mastering DQL was necessary to fully […]

The post Transform data into insights with Dynatrace Dashboards and Notebooks appeared first on Dynatrace news.

]]>
Explore Kubernetes metrics graphic


When we launched the new Dynatrace experience, we introduced major updates to the platform, including Grail™, our innovative data lakehouse unifying observability, security, and business data, and Dynatrace Query Language (DQL) for accessing and exploring unified data. While Grail and DQL opened up nearly limitless possibilities for data exploration, mastering DQL was necessary to fully leverage the power of Grail. Our latest enhancements to the Dynatrace Dashboards and Notebooks apps make learning DQL optional in your day-to-day work, speeding up your troubleshooting and optimization tasks.

In this blog post, we look at these enhancements, exploring methods for monitoring your Kubernetes environment and showcasing how modern dashboards can transform your data. Furthermore, we illustrate how these methods work seamlessly with Dashboards and Notebooks to enhance their effectiveness.

Get real-time insights by transforming complex data into dynamic, interactive dashboards

The many paths to building a dashboard or notebook

Getting started with the new Dashboards is now easier than ever, offering unprecedented ease and capabilities for exploring your data. We’ve not only improved how you interact with data in dashboards and notebooks, we also enhanced the way that underlying data can be shared across apps. These updates expand your options for exploration and creation, helping you to build your dashboards and notebooks quicker and more intuitively.

You can now:

Let’s look at each of these paths through an end-to-end use case focused on Kubernetes monitoring.

Kickstart your creation journey using ready-made dashboards and notebooks

Creating dashboards and notebooks from scratch can take time, particularly when figuring out available data and how to best use it. Ready-made dashboards and notebooks address this concern by offering pre-configured data visualizations and filters designed for common scenarios like troubleshooting and optimization.

These ready-made dashboards offer your platform engineers, who oversee Kubernetes environments, immediate and comprehensive data visibility. This allows platform engineers to focus on high-value tasks like resolving issues and optimizing performance rather than spending time on data discovery and exploration.

Kickstarting the dashboard creation process is, however, just one advantage of ready-made dashboards. Let’s assume you’re already using the new Kubernetes app, which offers a comprehensive overview of your Kubernetes environments and their telemetry. There are cases where more flexible data presentation is needed. Our new ready-made dashboards for Kubernetes not only provide instant insights into your clusters, nodes, workloads, or pods but also enable you to extend and customize the data shown in the Kubernetes app, leveraging the context-rich data from Dynatrace Grail. So, for example, if you need to seamlessly integrate metrics with logs for your workloads, you can create a customized view based on the pre-configured dashboard that consolidates all critical signals in one place, which is particularly essential for troubleshooting.

Finding ready-made dashboards is straightforward. Navigate to the list of dashboards and set the filter at the top left to Ready-made. Select the title of any dashboard that interests you, or use the search bar to narrow down the results.

Visualization: Leverage ready-made dashboards to create yours video thumbnail

Accelerate data exploration with seamless integration between apps

In developing the new Dynatrace experience, our goal was to integrate apps seamlessly by sharing the context when navigating between them (known as “intent”), much like sharing a photo from your smartphone to social media. This approach acknowledges that in any organization, software doesn’t work in isolation; boundaries and responsibilities are often blurred. This is even more true for critical scenarios like troubleshooting, which often requires more than the capabilities of a single person or app.

Let’s make this more tangible by using the Kubernetes cluster dashboard and demonstrating how this concept helps you to:

  • Seamlessly navigate between apps while maintaining context.
  • Effortlessly explore data in Dynatrace and create dashboards from it.

When working with the Kubernetes cluster dashboard, you have two options for digging deeper into further analysis, both using the Kubernetes app. You can use dynamic markdown links, which include the values of the actual dashboard variables, or you can utilize the open-with feature (the “intent” concept), which uses the actual context of the dashboard tile you’re viewing. With this latter approach, you even have the choice of passing a single value (Open field with) or all underlying data (Open record with) for the respective element (row, series, cells, etc.) when navigating to another app.

Visualization: Accelerate data exploration with seamless integration between apps video thumbnail

Next, let’s use the Kubernetes app to investigate more metrics. The intent concept and the open with feature can also be applied in reverse to include data or specific visualizations from an app on a particular dashboard. An example of this is shown in the video above, where we incorporated network-related metrics into the Kubernetes cluster dashboard.

Start from scratch with the new Explore interface for metrics in Dashboards and Notebooks

Once you’ve learned how to monitor your Kubernetes cluster using a ready-made dashboard and extending it with context from other apps, the next step is understanding how to create and extend such dashboards using the Dashboards or Notebooks app.

Exploring and adding metrics from scratch

Let’s revisit our example from the last chapter and add the same Kubernetes network metrics, this time by using the new Explore metric interface that allows you to:

  • Browse and add multiple metrics to a single tile
  • Apply basic commands such as aggregation, filter, and split
  • Use expressions to do calculations based on previously added metrics

Visualization: Exploring and adding metrics from scratch video thumbnail

The revised Explore interface, as shown in the clip above, now includes logs, metrics, events and business events, offering an improved filtering experience that enables you to:

  • Type ahead to add, edit, or remove available filters
  • Control how filters are applied via a rich set of operators (=, !=, in, not in, >=, <=, >, <)
  • Place wildcards before and after your filter values to automatically generate the best matching DQL when using startsWith, endsWith, or contains.
  • Control how filters are combined with logical operators, such as AND or OR
  • Easily filter entities by ID, name, or tags in the web UI
  • Get suggestions for metric values and entities (IDs, names, tags) for all data types

Build your dashboard effortlessly with only a few clicks

Blend metrics with data from Explore logs for a more comprehensive view to start log analysis

With the enhanced Explore Logs interface, retrieving and viewing logs from your Kubernetes workloads is straightforward. By incorporating a new tile, you can integrate these logs into your dashboard along with key metrics, such as the new Kubernetes network metrics we added earlier.

Leverage dashboards to monitor your environment in real time through log data. Once you identify an anomaly that requires your attention, you can start troubleshooting by delving further into the issue using the Open with option and the intent mechanism in the new Logs app. This app provides advanced analytics, such as highlighting related surrounding traces and pinpointing the root cause, as illustrated in the example below.

Visualization: Enhanced Explore Logs interface video thumbnail

Leveraging the capabilities of Grail and Smartscape® topology, Dynatrace seamlessly integrates logs, metrics, and traces to offer enhanced context for troubleshooting and in-depth analysis. This integration facilitates a comprehensive understanding of individual transactions through the Distributed Traces app and aids in pinpointing the root cause of issues when using the Problems app.

Intuitive data access with Davis CoPilot AI assistant

There’s also a brand new and completely different option for analyzing data using natural language; using the power of generative AI, Davis CoPilot™ converts your conversational prompts into accurate DQL commands, allowing both non-technical users as well as experienced data analysts to make data-driven decisions faster than ever before.

Davis CoPilot in Dynatrace screenshot

To learn more about how Davis CoPilot empowers you and your teams, see our blog post, Announcing General Availability of Davis CoPilot: Your new AI assistant.

Search metrics from anywhere

A speedy way to begin your data exploration journey, particularly if you already know which data you need, is to utilize our global search feature to effortlessly find and explore metrics from anywhere on the Dynatrace platform. This feature lets you explore any available metric and add it to Notebooks or Dashboards.

Imagine a colleague mentioning a newly released metric for Kubernetes during a coffee break. Rather than manually exploring the Kubernetes app you can simply open the Dynatrace global search and enter “Kubernetes network.” The relevant metrics are then immediately displayed alongside further details.

This efficient method allows you to easily browse and identify the appropriate metrics; adding them to your notebooks and dashboards requires just a single click.

Browse and identify the appropriate metrics in Dynatrace screenshot

Get started discovering and exploring your data

It has never been easier to analyze data within Dynatrace. Kickstart your data exploration journey and familiarize yourself with ready-made dashboards and the new Explore data interface. By the way, we also added new data visualization capabilities by adding new chart types and chart interactions.

Curious about our latest releases and upcoming features for Dashboards and Notebooks? Check out our community roadmap thread to stay updated!

Are you ready to try out the new Explore Data features?

The post Transform data into insights with Dynatrace Dashboards and Notebooks appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/transform-data-into-insights-with-dynatrace-dashboards-and-notebooks/feed/ 0