LLM observability | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 05 Mar 2026 09:01:57 +0000 en hourly 1 Optimizing AI ROI from DevOps and IT Operations: The rising need for AI/LLM observability https://www.dynatrace.com/news/blog/optimizing-ai-roi-from-devops-and-it-operations/ https://www.dynatrace.com/news/blog/optimizing-ai-roi-from-devops-and-it-operations/#respond Wed, 03 Dec 2025 18:09:08 +0000 https://www.dynatrace.com/news/?p=72110 Blog thumbnail

Every organization is adopting GenAI across its infrastructure and application stacks. It’s important that IT operations teams seek a seat at the table because large swaths of models will be deployed across every technology. For example, the use of cloud migrations, GenAI large language models, small language models, and specialized models will drive productivity, cost […]

The post Optimizing AI ROI from DevOps and IT Operations: The rising need for AI/LLM observability appeared first on Dynatrace news.

]]>
Blog thumbnail

Every organization is adopting GenAI across its infrastructure and application stacks. It’s important that IT operations teams seek a seat at the table because large swaths of models will be deployed across every technology. For example, the use of cloud migrations, GenAI large language models, small language models, and specialized models will drive productivity, cost savings, and business returns. Every customer is considering and attempting to measure their business returns from their AI investments; transparency into the data, system and model performance and drift, security, and quality are critical areas where IT operations, DevOps, SREs, and platform engineering teams can play a critical role in optimizing business returns and reducing business risks. So, where should you start the conversation?

Executives can use observability to reduce business risks and increase AI ROI by understanding how observability capabilities play a role in delivering across the core AI value categories of productivity, customer impact, cost optimization, innovation, and quality. For example, observability improves customer satisfaction by reducing the mean time to resolution and mean time to understanding. In addition, it can improve cross-team collaboration and data access to deliver cost efficiencies.

To reduce business risks and increase ROI in GenAI use cases, technology executives should plan to manage rising complexity, and, as part of continuous evaluation, executives should consider GenAI performance across the following dimensions:

  • System performance: Monitoring the system performance of GenAI applications encompasses measuring operational performance characteristics similar to those of traditional applications, including at the software and infrastructure layers and the model. Model system performance monitoring includes the measurement of metrics such as model response latency, error rates (including failure to respond), and API failures.
  • Quality performance: It is crucial for organizations to monitor the output quality of GenAI and AI applications. Quality includes accuracy of responses and model drift, where data used to train models no longer produces accurate or relevant results.
  • Governance: Model governance of GenAI often encompasses monitoring and enforcing legal requirements and the organization’s ethics policies. Ongoing monitoring is necessary, including the adoption of guardrails to prevent the delivery of outputs that don’t comply with laws or company policies.
  • Security: In addition to the theft of private information or loss of intellectual property, organizations must protect against security risks that are specific to GenAI applications. Prompt injection and jailbreaks are two emerging attacks. Monitoring tools that detect these and other security issues are critical to risk management.
  • Cost: Monitoring the cost of delivering a GenAI application is a multitiered undertaking. Depending on the application, organizations may incur costs for each query and response to a model, in addition to costs associated with the underlying infrastructure required to deliver the application. The ability to collect the right cost information and analyze it on a per-application basis will be key to the ability of an organization to determine ROI.

Organizations must base the measurement of each performance dimension on its ability to derive outcomes that drive business value. Each GenAI application should support a targeted outcome, such as improved productivity, increased revenue, new revenue streams, or enhanced customer satisfaction. Connecting the dots between GenAI performance dimensions and business value requires defining measurements that matter to the business and collecting, correlating, and analyzing the data to understand the app’s ability to deliver that value.

For technology executives, AI observability is fast becoming essential for managing the operational complexity and business outcomes from AI initiatives. It provides the visibility needed to demonstrate ROI, ensure reliable AI applications, and make informed decisions based on critical data that supports every AI use case.

Monitor, optimize, and secure Generative AI applications, LLMs, and agentic workflows — improving performance, explainability, and compliance.

Learn more, or try Dynatrace for free!

The post Optimizing AI ROI from DevOps and IT Operations: The rising need for AI/LLM observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/optimizing-ai-roi-from-devops-and-it-operations/feed/ 0
Unlocking productivity and trust: Dynatrace observability in NVIDIA AI Factory https://www.dynatrace.com/news/blog/unlocking-productivity-and-trust-dynatrace-observability-in-nvidia-ai-factory-environments/ https://www.dynatrace.com/news/blog/unlocking-productivity-and-trust-dynatrace-observability-in-nvidia-ai-factory-environments/#respond Tue, 28 Oct 2025 18:30:05 +0000 https://www.dynatrace.com/news/?p=71582 Davis CoPilot for NVIDIA

The NVIDIA Enterprise AI Factory addresses the rapidly evolving needs for AI infrastructure to support the rise of agentic AI. Since its launch, customers have leveraged this validated design to build agents by following structured methodology and recommended frameworks, which simplifies deployment and configuration while facilitating the implementation of AI factories in both on-premises and […]

The post Unlocking productivity and trust: Dynatrace observability in NVIDIA AI Factory appeared first on Dynatrace news.

]]>
Davis CoPilot for NVIDIA

The NVIDIA Enterprise AI Factory addresses the rapidly evolving needs for AI infrastructure to support the rise of agentic AI. Since its launch, customers have leveraged this validated design to build agents by following structured methodology and recommended frameworks, which simplifies deployment and configuration while facilitating the implementation of AI factories in both on-premises and hybrid cloud environments.

Dynatrace has been an integral part of this initiative. Dynatrace full-stack AI and LLM observability helps organizations move forward with confidence in building their AI and agentic AI initiatives.

Observable AI: Turn a black box into a glass box to build confidence

With the publication of comprehensive guidelines, it’s simpler than ever for Dynatrace customers to set up and start monitoring their full-stack NVIDIA enterprise AI infrastructure, including its key tiers and components. Covering the infrastructure layer from GPUs to Kubernetes, NVIDIA NIM microservices, NVIDIA NeMo, and other technologies up to the application layer, Dynatrace observability enables customers to confidently run and operate complex AI workflows on NVIDIA infrastructure.

NVIDIA Enterprise AI Factory for Agents including components covered by ecosystem partners (such as Observability). Picture taken from NVIDIA Enterprise AI Factory - Design Guide White Paper
Figure 1: NVIDIA Enterprise AI Factory for Agents, including components covered by ecosystem partners (such as Observability). Picture taken from NVIDIA Enterprise AI Factory – Design Guide White Paper

In parallel, Dynatrace has worked to significantly advance our AI and LLM observability offering by introducing the following:

Dynatrace AI Observability
Figure 2: Dynatrace AI Observability

These improvements address challenges such as missing observability insights, scale, sovereignty, and trust. This empowers organizations to operationalize AI by building trust and monitoring guardrails; providing analytics capabilities to detect user-facing issues; helping SREs and AI-native engineers maintain performance, reliability, and security; and reducing cost across the agentic, AI, and LLM stack.

Privacy and security lead the way to scaling AI with confidence

AI is delivering significant productivity improvements, with 66% of senior executives reporting positive trends in productivity, according to PwC’s AI Agent Survey. This momentum is driving the demand to manage AI expenditures, enhance the decision-making quality of agents, and optimize development through visibility into AI components’ behavior in production environments — from pilot projects to full-scale operations.

However, sensitive data considerations and strict compliance requirements often impede progress, preventing organizations from fully realizing the benefits of AI adoption. As enterprises prioritize data privacy, regulatory compliance, and data sovereignty, there is an increasing need for high-performance NVIDIA AI infrastructure alongside frameworks designed to preserve control, trust, and autonomy in AI development.

In a recent blog on sovereign AI, NVIDIA shares strategies for nations and enterprises to develop AI factories that uphold local governance, security, and cultural values. Combining such factories with the Dynatrace advanced observability solution enables organizations to operationalize AI at scale — building secure and scalable agents, deployed on premises or in hybrid environments.

From privacy needs to public-sector requirements: NVIDIA AI Factory for Government

At NVIDIA GTC Washington, D.C. today, NVIDIA AI Factory for Government was announced, in support of the needs for regulated environments to drive AI initiatives. The U.S. Office of Management and Budget’s decision to establish scorecards for agencies’ AI maturity and management is in line with a 2024 Gartner Research forecast that more than 60% of government organizations will be prioritizing their investments in business automation by 2026 — up from 35% in 2022. The NVIDIA AI Factory for Government is a full-stack, end-to-end reference design that brings the power of reasoning AI to federal organizations. It helps organizations unlock productivity gains just like it does for enterprises, from service delivery to threat detection and day-to-day operations.

Built on the experience of deploying internal AI factories, the reference design offers guidance for deploying agentic AI, physical AI, and high-performance computing workloads on premises and in hybrid cloud environments, while meeting the compliance needs of federal and other secure organizations. The NVIDIA AI Factory for Government reference design includes NVIDIA Blackwell accelerated computing and NVIDIA networking, NVIDIA-Certified Systems, NVIDIA AI Enterprise software, NVIDIA Nemotron open models, and third-party software from AI leaders, all validated by NVIDIA.

Dynatrace delivers trusted observability and automation for regulated environments

Dynatrace has always been committed to supporting the public sector and other industries with regulatory requirements by providing customers with capabilities to control data flow through its lifecycle and manage sensitive data from ingestion to deletion, as well as global deployment options to meet data residency requirements, configurable retention times for different data types and use cases, unique encryption keys for customer’s stored data, and more.

Our dedication is reflected in customers’ success stories from regulated industries, as well as a growing list of global and local certifications, such as ISO 27001, SOC 2 Type II, CSA STAR 2, ENS, Tisax, and others. Find out more about our certifications and supported compliance frameworks in our Trust Center. For organizations also navigating evolving sovereignty requirements, our approach to digital sovereignty demonstrates how Dynatrace combines technical innovation with policy alignment to deliver trusted solutions globally.

Benefit from full-stack observability for end-to-end validated design

Dynatrace observability with the NVIDIA AI Factory for Government reference design enables organizations to accelerate the deployments of their AI agents and applications for federal and enterprise environments, and benefit from real-time, AI-powered insights.

These benefits range from improved scalability and performance to reduced complexity and total cost of ownership by simplifying processes, mitigating deployment risks to improved data security and compliance.

Visit the Dynatrace Playground to experience the possibilities of AI and LLM observability, and discover how Dynatrace is accelerating enterprise AI at scale.

Dynatrace and the Dynatrace logo are trademarks of the Dynatrace, Inc. group of companies. All other trademarks are the property of their respective owners.

The post Unlocking productivity and trust: Dynatrace observability in NVIDIA AI Factory appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/unlocking-productivity-and-trust-dynatrace-observability-in-nvidia-ai-factory-environments/feed/ 0
From data to insights with Dynatrace Dashboards https://www.dynatrace.com/news/blog/from-data-to-insights-with-dynatrace-dashboards/ https://www.dynatrace.com/news/blog/from-data-to-insights-with-dynatrace-dashboards/#respond Fri, 11 Jul 2025 13:42:19 +0000 https://www.dynatrace.com/news/?p=69842 Dynatrace dashboards

We had one main goal in mind when designing Dynatrace® Dashboards: reimagine how our customers consume and interact with their observability data. Built for speed, clarity, and collaboration, the Dashboards app helps teams easily explore, visualize, and act on telemetry data. From natural language queries to advanced visualizations, Dashboards streamlines your workflow and reveals critical insights at any scale. Whether you're monitoring infrastructure, applications, or AI workloads, Dashboards adapts to your needs, turning raw data into real-time insights.

The post From data to insights with Dynatrace Dashboards appeared first on Dynatrace news.

]]>
Dynatrace dashboards

We’ll walk you through a real-world example of monitoring OpenAI APIs in production to show you what this looks like in action.

In practice: Create a dashboard monitoring OpenAI LLM APIs

Imagine you’re on a platform team at a SaaS company that recently integrated OpenAI to power features like smart search, summarization, or chatbots. With these capabilities now live, your next challenge is ensuring they perform reliably, scale efficiently, and stay within budget. This is where Dynatrace shines—helping you transform telemetry into insights that drive action.

Let’s walk through all the steps to create just such a dashboard, and dig deeper to:

  • Find and add (OpenAI telemetry) data with ease.
  • Tailor visualizations to understand token usage, latency, and error metrics easily.
  • See what matters: filter and segment data by LLM model, service, or environment.
  • Predict and prevent issues: avoid model response slowdowns and cost spikes.

Find and add (OpenAI telemetry) data with ease

Creating a new dashboard begins with identifying and understanding the relevant data for your use case. Monitoring LLM APIs requires the visualization of key metrics like request volume, latency, or error rates per model. With Dashboards, exploring your data is intuitive, providing multiple ways to search for and analyze data.

  • Start with a ready-made dashboard that provides instant insights
  • Explore data using a simple-to-use point-and-click interface—ideal for getting started by quickly adding tiles
  • Utilize the full power of Grail by writing your own DQL query or utilizing Davis CoPilot® to transform your natural language prompts into DQL queries.

As an experienced Dynatrace user, you’re familiar with exploring data in context with our purpose-built apps like Kubernetes, Logs, or Distributed Traces, and how to add visualizations from those apps to your dashboards.

Let’s look at some of these approaches in the following sections.

Start the journey with ready-made dashboards

You don’t have to start from scratch. Dynatrace offers many ready-made dashboards as part of Dynatrace® Apps and purpose-built extensions to serve dedicated use cases. As the leading observability solution for monitoring AI workloads, we offer dashboards for all major AI and LLM stacks, including agentic frameworks such as OpenAI, Anthropic, Amazon Bedrock, or NVIDIA. These dashboards provide instant value, whether you’re monitoring performance or debugging expensive prompts. By delivering real-time insights into request volume, latency, cost, and service health, they not only save you time but also create a solid foundation for tailoring their experience to your needs.

Let’s start our journey by opening the ready-made dashboard for OpenAI and creating a copy of it. To follow along, locate the Dashboards app on the Dynatrace Playground.

Duplicate the ready-made dashboard to customize it.
Figure 1. Duplicate the ready-made dashboard to customize it.

Add further tiles to analyze token usage

Next, let’s add another tile to visualize the overall prompt token usage by type: input vs. output for OpenAI services. From discussions with our platform observability team, we know that all relevant metrics sent to Dynatrace using OpenTelemetry are available as custom metrics prefixed with gen_ai. We add a metrics tile and type gen_ai into the search field. This instantly surfaces all related telemetry. A few clicks later, applying data splits and aggregations, we have two more tiles, demonstrating how simple it is to turn raw telemetry into actionable insights:

  • pie chart that shows the balance between input and output tokens
  • line chart that tracks how the usage evolves over time

Visualizing overall prompt token usage video thumbnail
Figure 2. Visualizing overall prompt token usage.

For further insights into the exploration and transformation possibilities in Dashboards, check out our blog post on transforming data into insights.

Leverage the power of Dynatrace Grail

Not sure where to start, which metric to use, or how to quickly advance with the power of Dynatrace Query Language (DQL) and Grail® data lakehouse? That’s where Davis CoPilot® comes in. Built directly into Dashboards and Notebooks, Davis CoPilot allows you to interact with your data using plain language—no need to write queries or know exact metric names. Just type something like Visualize token usage by input and output types, and the AI will help you instantly generate the appropriate query, taking you from question to insight in seconds.
CoPilot Token Usage video thumbnail
Figure 3. Use Davis CoPilot to create and visualize queries instantly.

Tailor visualizations to easily understand token usage, latency, and errors

As someone responsible for monitoring systems or ensuring service reliability, you know how important it is to get the right insights at a glance. Dynatrace helps you build intuitive dashboards that focus on what matters most: understanding your data and taking action on it.

Once the data is set and a tile added, Dynatrace automatically suggests the most suitable visualization. For example, when tracking API token usage by type over time, a line chart is recommended to highlight trends and fluctuations.

A suitable line chart visualization is automatically suggested.
Figure 4. A suitable line chart visualization is automatically suggested.

Dynatrace also applies other smart defaults based on the context of the visualized data. For example, when you add a metric that tracks the usage of example prompts and split it by the prompt name, sparklines are automatically included to show trends over time—no extra configuration needed. And if you’re already a power user, the newly added search speeds up your dashboard creation journey by offering a way to instantly jump to any configuration without the need to scroll around. But there’s a lot more that helps improve the user journey. We harmonized the settings of individual visualization types, ensuring that already defined configurations, such as color palettes or units, persist, even if you change the type.

The settings of individual visualization types are enhanced and harmonized.
Figure 5. The settings of individual visualization types are enhanced and harmonized.

We’ve also made many updates to the chart plotting features of our pre-existing visualizations. For example, the single value tile, which used to be a basic number display, is now a highly expressive component. You can now enrich the single-value tile with icons, apply color thresholds to flag anomalies, add sparklines to show trends, and add value and trend labels that provide additional context for the charted value and give it meaning.

The single value tile now also includes sparklines and other options.
Figure 6. The single value tile now also includes sparklines and other options.

Plotting the values on a map benefits many signals. Consider displaying token usage or prompts issued per destination. The map component has a rich set of customization options—such as color rules, pin shapes, and unit formatting—explicitly designed to support the visualization of geographic data.

Use the map visualization to display data geographically and to uncover location-based patterns.
Figure 7. Use the map visualization to display data geographically and to uncover location-based patterns.

See what matters: filter and segment data by LLM model, service, or environment

To make a dashboard truly actionable, the next step is to add filters and segmentations. This allows you to tailor one view dynamically for different audiences, environments, or services, all within a single dashboard. For example, you might filter an OpenAI dashboard by environment (production, staging, test) or model type (GPT-4.1, o3, o3-mini) to focus on what matters most in each context.

Dynatrace offers powerful ways to filter data:

  • Using reusable segments, multidimensional global filters can be applied to all tiles and data. This is ideal for applying a specific (user) context, such as environment, team, or cluster. Segments are persisted across navigation between apps, allowing for simple drill-down journeys.
  • With variables, we introduce Dashboard-specific filters, offering fine-grained control for each tile, perfect for filtering information such as LLM model type or feature toggles.

If you want a more in-depth tutorial, check out our latest blog post on filtering.

Predict and prevent issues: avoid model response slowdowns and cost spikes

Dashboards aren’t meant to be stared at all day. In most organizations, they’re often left untouched until something goes wrong. That’s when dashboards become invaluable: surfacing the correct data at the right time to help teams quickly understand, diagnose, and resolve issues.

From passive observation to proactive action, Dynatrace bridges the gap with interactive, AI-powered dashboards that don’t just visualize data; they empower you to act on it. You can add alerts and forecasts directly from charts with just a few clicks.

For example, suppose your dashboard tracks OpenAI model response times and associated costs. In this case, you can set an alert to notify your team if the average response time exceeds a certain threshold for a defined period directly from within the chart. This ensures you’re reacting to issues and anticipating them before users are impacted.

You can interact with your data directly on your charts, for example, zoom in/out and set up instant alerts.
Figure 8. You can interact with your data directly on your charts, for example, zoom in/out and set up instant alerts.

Another popular example of proactive monitoring is cost forecasting. Our dashboard already tracks cost trends over time—such as “prompt costs” and “complete costs,” for example—with a line chart highlighting weekly fluctuations.

By enabling forecasting, Dynatrace projects future spending based on historical usage patterns. This helps you anticipate budget overruns, adjust resource allocation, and make informed decisions before costs spiral. The predicted budget spend is shown alongside a table highlighting the “Top 10 expensive prompts.” This allows teams to identify which workloads or user actions contribute most to spending, ideal for optimization efforts or chargeback models.

Utilize AI-powered forecasting to predict future costs.
Figure 9. Utilize AI-powered forecasting to predict future costs.

Share with teams: secure, flexible collaboration

The next step is to share our dashboard with the right people, ensuring teams are aligned across job roles and departmental boundaries. The new Dynatrace Dashboards supports flexible sharing options for collaboration within your organization.

Fine-grained collaboration settings allow you to:

  • Share a document with specific users or groups applying either view or edit permissions.
  • Roll out a dashboard to users in the environment.
  • Generate a link that works for any authenticated user in your environment—ideal for broad internal visibility without managing individual access.

Ready to try it out yourself?

Dynatrace Dashboards redefine how teams interact with observability data. Whether you’re monitoring LLM APIs, optimizing cloud costs, or ensuring service reliability, Dashboards empowers you to:

  • Explore data intuitively.
  • Visualize insights using smart defaults and rich customization options.
  • Segment and filter your data dynamically, offering tailored views for use cases.
  • Act proactively on data anomalies using forecasting and creating alerts in context.

Experience the power of Dashboards: Head over to the Dynatrace Playground and browse the ready-made dashboards or create your own, following the steps described in this blog post.

The post From data to insights with Dynatrace Dashboards appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/from-data-to-insights-with-dynatrace-dashboards/feed/ 0
The rise of agentic AI part 4: Dynatrace delivers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM https://www.dynatrace.com/news/blog/full-stack-observability-for-nvidia-blackwell-and-nim-based-ai/ https://www.dynatrace.com/news/blog/full-stack-observability-for-nvidia-blackwell-and-nim-based-ai/#respond Fri, 20 Jun 2025 06:00:05 +0000 https://www.dynatrace.com/news/?p=69115 Davis CoPilot for NVIDIA

The Dynatrace® unified, AI-powered observability platform delivers full-stack AI and LLM observability, including of NVIDIA Blackwell and NVIDIA NIM systems, and AI-driven insights to meet the scale and complexity of enterprise AI deployments. In this fourth installment of our series, The Rise of Agentic AI, we explore how the Dynatrace integration with NVIDIA systems provides enterprises with all the insights needed to detect customer-facing issues, helping IT teams maintain performance, reliability, and security across their AI workloads.

The post The rise of agentic AI part 4: Dynatrace delivers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM appeared first on Dynatrace news.

]]>
Davis CoPilot for NVIDIA

NVIDIA Blackwell systems provide high-performance infrastructure for enterprise AI, and now, thanks to the Dynatrace integration with the NVIDIA Enterprise AI Factory reference design, enterprises can add Dynatrace Full-Stack Observability to NVIDIA Blackwell infrastructure. This magnifies the value of the NVIDIA Blackwell platform by providing real-time performance insights, anomaly detection, and dependency mapping.

Keep high performance and security top of mind with unified observability and security

Figure 1. The Dynatrace AI Observability platform
Figure 1. The Dynatrace AI Observability platform

Dynatrace aligns with high data security and privacy standards typical of on-premises NVIDIA Blackwell deployments, particularly in regulated industries such as finance and healthcare. Its unified data model, Smartscape® topology mapping, and Davis® AI engine provide deep visibility into the full stack—from GPU metrics and containerized workloads to distributed applications and user experiences, enabling tailored observability for workloads running on  NVIDIA Blackwell. Integrating NVIDIA Data Center GPU Manager or other telemetry sources is straightforward, allowing teams to monitor GPU health, utilization, thermal thresholds, and memory bandwidth alongside traditional infrastructure metrics.

Dynatrace technology allows for automated discovery and instrumentation of services running on NVIDIA Blackwell-accelerated systems. Whether monitoring high-throughput GPU compute tasks, Kubernetes clusters, or microservices, Dynatrace ensures low-overhead performance monitoring with minimal manual configuration.

AI-powered, real-time insights improve performance and explainability

With Dynatrace Full-Stack AI Observability, you can monitor real-time performance, trace prompts end-to-end, and ensure compliance, optimizing cost and throughput for your AI and LLM workflows and agents, offering various use cases such as

  • Monitor service health and performance, tracking real-time metrics and offering clear visibility into service incidents.
  • Validate service quality by measuring response speed or identifying performance hotspots.
  • End-to-end tracing and debugging pinpoint the root cause of errors and failures in the LLM chain, troubleshoot issues in complex pipelines, and trace dependencies across the entire system spanning multiple LLMs, RAG pipelines, and agentic frameworks.
Figure 2. Sample dashboards provided for tracking service health and performance
Figure 2. Sample dashboards are provided for tracking service health and performance

Unified AI-powered observability

Dynatrace delivers full stack observability for your LLMs and Generative AI applications running on NVIDIA Blackwell systems. Its ability to provide visibility into complex, high-performance environments allows enterprises to fully leverage Blackwell’s capabilities while maintaining operational excellence and system reliability, improving the performance, explainability, and compliance of your AI workloads and agents.

Figure 3. Dig deeper into the possibilities of AI and LLM observability on the Dynatrace Playground
Figure 3. Dig deeper into the possibilities of AI and LLM observability on the Dynatrace Playground

Visit the Dynatrace Playground to learn more and gain hands-on experience with prepopulated data, so you can experience the possibilities of AI and LLM observability with Dynatrace. If you’re interested in using Dynatrace for your own AI workloads, visit our documentation and start benefiting from full stack observability for AI and LLM.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.

The post The rise of agentic AI part 4: Dynatrace delivers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/full-stack-observability-for-nvidia-blackwell-and-nim-based-ai/feed/ 0
Deliver secure, safe, and trustworthy GenAI applications with Amazon Bedrock and Dynatrace https://www.dynatrace.com/news/blog/deliver-secure-safe-and-trustworthy-genai-applications-with-amazon-bedrock-and-dynatrace/ https://www.dynatrace.com/news/blog/deliver-secure-safe-and-trustworthy-genai-applications-with-amazon-bedrock-and-dynatrace/#respond Wed, 12 Mar 2025 18:56:47 +0000 https://www.dynatrace.com/news/?p=68271 Gen AI graphic

Every software development team grappling with Generative AI (GenAI) and LLM-based applications knows the challenge: how to observe, monitor, and secure production-level workloads at scale. Traditional debugging approaches, logs, and occasional remote breakpoint instrumentation can’t easily keep pace with cloud-native AI deployments, where performance, compliance, and costs are all on the line. How can you […]

The post Deliver secure, safe, and trustworthy GenAI applications with Amazon Bedrock and Dynatrace appeared first on Dynatrace news.

]]>
Gen AI graphic

Every software development team grappling with Generative AI (GenAI) and LLM-based applications knows the challenge: how to observe, monitor, and secure production-level workloads at scale. Traditional debugging approaches, logs, and occasional remote breakpoint instrumentation can’t easily keep pace with cloud-native AI deployments, where performance, compliance, and costs are all on the line. How can you gain insights that drive innovation and reliability in AI initiatives without breaking the bank?
Dynatrace helps enhance your AI strategy with practical, actionable knowledge to maximize benefits while managing costs effectively.

Amazon Bedrock, equipped with Dynatrace Davis® AI and LLM observability, gives you end-to-end insight into the Generative AI stack, from code-level visibility and performance metrics to GenAI-specific guardrails.

Developers deserve a frictionless troubleshooting experience and fast access to real-time data—no more guesswork or costly redeployments. Here’s how Dynatrace, combined with Amazon Bedrock, arms teams with instant intelligence from dev to production, helping to accelerate innovation while keeping performance, costs, and compliance in check.

Introducing Amazon Bedrock and Dynatrace Observability

Amazon Bedrock is a serverless service for building and scaling Generative AI applications easily with foundation models (FM). It provides an easy way to select, integrate, and customize foundation models with enterprise data using techniques like retrieval-augmented generation (RAG), fine-tuning, or continued pre-training.

Dynatrace is an all-in-one observability platform that automatically collects production insights, traces, logs, metrics, and real-time application data at scale.  With powerful Davis AI engine Dynatrace notifies teams about production-level issues before they disrupt users, helps predict resource usage,costs, and performance issues, and delivers guardrails that protect data and maintain compliance.

Together, Amazon Bedrock and Dynatrace provide an end-to-end observability solution for AI applications:

  • Predictive operations: Proactive usage and cost forecasting to reduce unexpected operational expenses and token usage.
  • Production performance monitoring: Service uptime, service health, CPU, GPU, memory, token usage, and real-time cost and performance metrics.
  • Guardrail analysis: Detect hallucinations, track prompt injections, mitigate PII leakage, and ensure brand-safe outputs.
  • Full-stack tracing: Track each user request across multiple FMs, vector databases, orchestrators (LangChain), and custom business logic.
  • Compliance: Document all inputs and outputs, maintaining full data lineage from prompt to response to build a clear audit trail and ensure compliance with regulatory standards.

Video overview of Amazon Bedrock dashboard with Dynatrace AI and LLM Observability solution
Figure 1. Video overview of Amazon Bedrock dashboard with Dynatrace AI and LLM Observability solution.

How it works

Dynatrace seamlessly instruments your LLM-based workloads using Traceloop OpenLLMetry, which augments standard OpenTelemetry data with AI-specific KPIs (for example, token usage, prompt length, and model version).

Combined with Amazon Bedrock, you can:

  • Spin up your AI model on Amazon Bedrock—choose from providers like AI21, Anthropic, Cohere, Stability AI, Mistral AI, Meta, or Amazon’s own Nova/Titan foundation models.
  • Automatically instrument your application with OpenTelemetry.
  • Configure OpenLLMetry to capture specialized LLM details as spans and metrics, like model name, completion time, token count, token cost, and prompt text.
  • Send unified data to Dynatrace for analysis alongside your logs, metrics, and traces.

Behind the scenes, Dynatrace merges the standard telemetry with these advanced AI attributes, surfaces them in real-time dashboards, and applies AI-driven analytics to discover anomalies, forecast usage costs, and diagnose root causes.

Distributed Tracing overview of an Amazon Bedrock request with LangChain
Figure 2. Distributed Tracing overview of an Amazon Bedrock request with LangChain.

How to set up and instrument your data with OpenLLMetry

Traceloop OpenLLMetry is an open source extension that standardizes LLM and Generative AI data collection. By layering on top of OpenTelemetry standards, OpenLLMetry captures the critical metrics you can’t get by default—like the number of tokens, model temperature, or guardrail triggers.

Here’s how to set it up for Amazon Bedrock:

  1. Install OpenLLMetry in your Python or Node.js environment:
 pip install traceloop-sdk
from traceloop.sdk import Traceloop

headers = {

'Authorization': f"Api-Token {environ.get('DYNATRACE_TEAM_KEY')}"

}

Traceloop.init(

app_name=environ.get('DYNATRACE_APP_NAME'),

api_endpoint=environ.get('DYNATRACE_URL'),

headers=headers

)
  1. Configure environment variables to send data to Dynatrace via your ingest token:
 DYNATRACE_URL =https://123abcde.live.dynatrace.com/api/v2/otlp 

DYNATRACE_TEAM_KEY=dt0.....
  1. You can optionally add OpenLLMetry decorators or instrumentation to your LLM calls (for example, with LangChain or direct Bedrock SDK calls).

When your application queries Amazon Bedrock, OpenLLMetry automatically captures:

  • Prompt tokens vs. completion tokens
  • Finish reason (did the LLM stop due to a user request, or was the max token limit reached?)
  • Model type (which Amazon foundation model or third-party model is used?)
  • Performance: Response time, throughput, and error rate
    • Guardrail activations: Toxicity, PII, denied topics, and hallucinations
    • System, prompt, and completion messages and roles

This data is instantly correlated in Dynatrace so you can visualize or alert on critical thresholds (for example, if your average token usage spikes or your overall cost forecast grows beyond budget).

Overview of observability data flowing into Dynatrace from a travel agent application running in a Kubernetes cluster powered with Amazon Bedrock, where OpenLLMetry instruments the data
Figure 3. Overview of observability data flowing into Dynatrace from a travel agent application running in a Kubernetes cluster powered with Amazon Bedrock, where OpenLLMetry instruments the data.

How to debug incorrect responses in production

Let’s walk through a real-world scenario:

Your production travel agent application—powered by Amazon Bedrock and Dynatrace—gives users incorrect travel recommendations. Perhaps it suggests flights or hotels that don’t exist or mixes up time zones. This isn’t just a minor inconvenience; it jeopardizes user experience and can directly impact revenue and trust.

Here’s how Dynatrace helps you trace and resolve the issue quickly:

Proactive alerting with Davis AI

You receive an alert from Dynatrace Davis AI anomaly detection indicating incorrect system behavior. There might be a spike in “incorrect itinerary” complaints or conversation outcomes flagged as “nonsensical.” Davis AI correlates the unusual LLM responses with application telemetry and usage patterns, so you immediately know something is off in the recommendation flow.

Full-stack end-to-end tracing

In Dynatrace Distributed Tracing, you see the entire transaction trace for the affected user session. This includes front-end requests, back-end aggregator logic, calls to Amazon Bedrock, and any vector database lookups performed for retrieval-augmented generation (RAG). Rather than sifting through multiple logs, you have a single timeline that reveals exactly where the LLM call returned unexpected data.

Inspecting the GenAI model details

By drilling down into the span data enriched by OpenLLMetry, you can see:

  • Prompt and completion text and tokens used.
  • The specific foundation model version (for example, anthropic.claude-v1 or amazon.nova).
  • Temperature setting and max token limits.
  • Any error codes or guardrail triggers.

This clarity helps you pinpoint if the model produces off-base recommendations because of a misaligned temperature, an out-of-date context, or a mismatch in user inputs.

Root cause analysis

With Dynatrace, you quickly correlate the LLM anomaly to a specific function in your microservice code. You discover that an external data source used for itinerary validation had missing or stale updates, causing the LLM prompt to reference invalid flights. You’ve found the “why” without manually spelunking logs in disparate systems.

Resolving and validating

A fix might involve updating your data pipeline or refining the prompt logic. You can deploy the change and watch in near real-time as Dynatrace collects new traces and logs. Davis AI recognizes that the anomaly is cleared, confirming that your fix resolved the incorrect responses—no guesswork required.

You can find the code example for our travel agent application here for review, and the dashboard on our Dynatrace Playground instance.

Overview of Amazon Bedrock service health, performance, quality, and guardrails
Figure 4. Overview of Amazon Bedrock service health, performance, quality, and guardrails.

Summary

By integrating Amazon Bedrock with Dynatrace end-to-end observability, you not only catch issues early but also trace them across your entire AI stack to the root cause. Building or scaling Generative AI applications with Amazon Bedrock requires robust insights into your environment—from model usage and performance metrics to cost forecasts and guardrail efficacy.

Dynatrace helps you scale with:

  • Complete end-to-end tracing across your services, external data pipelines, and LLM calls.
  • Predictive analytics that forecast AI resource usage and cost trends, letting you proactively manage budgets.
  • Unified dashboards that bring performance, cost, code-level data, logs, metrics, and audit events together.
  • Compliance and governance that integrate security checks, data masking, and guardrail analysis.

Whether you’re a developer racing to put your latest AI-powered application or a new feature into production or an SRE ensuring your system meets enterprise-grade SLAs, Dynatrace and Amazon Bedrock help you to create frictionless AI applications, focusing on performance and observability—at any scale, in production, with no downtime.

Useful resources

The post Deliver secure, safe, and trustworthy GenAI applications with Amazon Bedrock and Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/deliver-secure-safe-and-trustworthy-genai-applications-with-amazon-bedrock-and-dynatrace/feed/ 0
Dynatrace accelerates business transformation with new AI observability solution https://www.dynatrace.com/news/blog/dynatrace-accelerates-business-transformation-with-new-ai-observability-solution/ https://www.dynatrace.com/news/blog/dynatrace-accelerates-business-transformation-with-new-ai-observability-solution/#respond Wed, 31 Jan 2024 17:00:34 +0000 https://www.dynatrace.com/news/?p=61661 Davis CoPilot

Adoption of artificial intelligence (AI) is increasingly imperative for any organization that hopes to remain competitive in the future. However, the benefits of AI are not as straightforward as they might first appear.

The post Dynatrace accelerates business transformation with new AI observability solution appeared first on Dynatrace news.

]]>
Davis CoPilot

While off-the-shelf models assist many organizations in initiating their journeys with generative AI (GenAI), scaling AI for enterprise use presents formidable challenges. It requires specialized talent, a new technology stack for managing and deploying models, an ample budget for rising compute costs, and end-to-end security. Many organizations haven’t even considered which use cases will bring them the biggest return on AI investment.

This blog post explores how AI observability enables organizations to predict and control costs, performance, and data reliability. It also shows how data observability relates to business outcomes as organizations embrace generative AI.

Challenges of deploying AI applications in the enterprise

There are several reasons for organizations to consider AI observability as they adopt AI.

Unpredictable costs. Many organizations face significant challenges in pursuing their cloud migration initiatives, which often accompany or precede AI initiatives. Insufficient consideration of total lifecycle costs during the strategic planning phase is a critical issue that forces some organizations that initially pioneer AI to retreat from the cloud due to the pressures of unforeseen costs. Worse, the costs associated with GenAI aren’t straightforward, are often multi-layered, and can be five times higher than traditional cloud services.

Service reliability. GenAI represents a radical shift from command-based interaction models and graphic user interfaces, which have dominated computing for the past 50 years. While GenAI tools have catalyzed more conversational and natural interactions between humans and machines, service reliability is an issue. GenAI is prone to erratic behavior due to unforeseen data scenarios or underlying system issues. Failure to provide timely and accurate answers erodes user trust, hinders adoption, and harms retention. Research by VC firm Sequoia indicates that the use of large language model (LLM) applications lags behind traditional consumer applications, with only 14% of users active daily.

Service quality. The quality and accuracy of information are crucial. However, correct answers might not be immediately apparent. LLMs are prone to “hallucinations” that exacerbate human bias and struggle to produce highly personalized outputs. AI hallucination is a phenomenon where an LLM perceives patterns that are nonexistent or imperceptible to human observers, creating outputs that are nonsensical or altogether inaccurate.

Consequently, AI model drift and hallucinations emerge as primary concerns. For example, a Stanford University and UC Berkeley team noted in a research study that ChatGPT behavior deteriorates over time. The team found that the behavior of the same LLM service devolved in a relatively short time, highlighting the need for continuous observability of LLM quality.

Retrieval-augmented generation emerges as the standard architecture for LLM-based applications

Given that LLMs can generate factually incorrect or nonsensical responses, retrieval-augmented generation (RAG) has emerged as an industry standard for building GenAI applications. RAG augments user prompts with relevant data retrieved from outside the LLM. Augmenting LLM input in this way reduces apparent knowledge gaps in the training data and limits AI hallucinations.

The RAG process begins by summarizing and converting user prompts into queries that are sent to a search platform that uses semantic similarities to find relevant data in vector databases, semantic caches, or other online data sources. Retrieved data is then submitted to the LLM along with the prompt to provide complete context for the LLM to create its response.

Using the example of a chatbot, once the user submits a natural language prompt, RAG summarizes that prompt using semantic data. The converted data is transmitted to a search platform that searches for relevant data related to the query. Related data is sorted based on a relevance score—for example, semantic distances (a quantitative measure of the relatedness between two or more data points). The most relevant data is submitted to the LLM along with the prompt. The LLM then synthesizes the retrieved data with the augmented prompt and its internal training data to create a response that can be sent back to the user.

Sample Retrieval-augmented generation (RAG) architecture
Figure 1: Sample RAG architecture

While this approach significantly improves the response quality of GenAI applications, it also introduces new challenges. Data dependencies and framework intricacies require observing the lifecycle of an AI-powered application end to end, from infrastructure and model performance to semantic caches and workflow orchestration.

Dynatrace provides end-to-end observability of AI applications

As AI systems grow in complexity, a holistic approach to the observability of AI-powered applications becomes even more crucial. Bringing together metrics, logs, traces, problem analytics, and root-cause information in dashboards and notebooks, Dynatrace offers an end-to-end unified operational view of cloud applications.

Dynatrace AI observability capabilities
Figure 2: Dynatrace AI observability capabilities

Dynatrace AI observability allows Dynatrace to observe the complete AI stack of modern applications, from foundational models and vector database metrics to orchestration frameworks covering modern RAG architectures, providing you with visibility into the entire lifecycle of modern applications across various layers:

  • Infrastructure: Utilization, saturation, and errors
  • Models: Accuracy, precision/recall, and explainability
  • Semantic caches and vector databases: Volume and distribution
  • Orchestration: Performance, versions, and degradation
  • Application health: Availability, latency, and reliability

Observe AI infrastructure and the environmental impact of machine learning

Though the cost of training LLMs and operating GenAI applications is significant, it’s not the sole factor companies should consider in their AI strategies. Development and demand for AI tools come with a growing concern about their environmental cost. Building LLMs consumes vast amounts of electricity and generates substantial heat. Researchers estimate it took 1,287 megawatt hours to create ChatGPT, releasing 552 tons of CO2. This is equivalent to driving 123 gas-powered cars for a whole year. But energy consumption isn’t limited to training models—their usage contributes significantly more. For example, generating an image requires as much power as fully charging your smartphone.

Estimates show that NVIDIA, a semiconductor manufacturer, could release 1.5 million AI server units annually by 2027, consuming 75.4+ terawatt hours yearly—more than the annual consumption of some countries.

Monitoring NVIDIA GPUs with Dynatrace

To help companies build more sustainable products, Dynatrace seamlessly integrates with AI infrastructure such as Amazon Elastic Inference, Google Tensor Processing Unit, and NVIDIA GPU, enabling monitoring of infrastructure data, including temperature, memory utilization, and process usage to ultimately support carbon-reduction carbon-reduction initiatives.

Observing AI models

Running AI models at scale can be resource-intensive. Model observability provides visibility into resource consumption and operation costs, aiding in optimization and ensuring the most efficient use of available resources.

Integrations with cloud services and custom models such as OpenAI, Amazon Translate, Amazon Textract, Azure Computer Vision, and Azure Custom Vision provide a robust framework for model monitoring. For production models, this provides observability of service-level agreement (SLA) performance metrics, such as token consumption, latency, availability, response time, and error count.

Model observability with Dynatrace

Beyond SLAs, the emergence of machine learning technical debt poses an additional challenge for model observability. Managing regressions and model drift is crucial when deploying and monitoring machine learning models in operation, especially as new data comes in.

To observe model drift and accuracy, companies can use holdout evaluation sets for comparison to model data. For model explainability, they can implement custom regression tests, providing indicators of model reputation and behavior over time.

AI model regression tests with Dynatrace
Figure 5: AI model regression tests with Dynatrace

Observing semantic caches and vector databases

The RAG framework has proven to be a cost-effective and easy-to-implement approach to enhancing the performance of LLM-powered apps by feeding LLMs with contextually relevant information, eliminating the need to constantly retrain and update models while mitigating the risk of hallucination.

However, RAG is not perfect and raises various challenges, particularly concerning the use of vector databases and semantic caches. To address the challenge of these retrieval and generation aspects, Dynatrace provides monitoring capabilities to semantic caches and vector databases such as Milvus, Weaviate, and Chroma.

Vector database observability in Dynatrace
Figure 6: Vector database observability in Dynatrace

This empowers customers to capture the effectiveness of retrieval-augmented generation systems, giving them the tools to optimize prompt engineering, search and retrieval, and overall resource utilization.

Observing orchestration frameworks

The knowledge utilized by LLMs and other models is limited to the data on which these models are trained. Building AI applications that can integrate private data or data introduced after a model’s training cutoff date requires augmenting the model with the specific information it needs using prompt engineering and retrieval-augmented generation.

Orchestration frameworks such as LangChain provide application developers with several components designed to help build RAG applications—first by providing a pipeline for ingesting data from external data sources and indexing it.

RAG indexing
Figure 7: RAG indexing

Second, they support the actual RAG chain, taking user queries at runtime and retrieving relevant data from the index, then passing that to the model.

RAG search and retrieval
Figure 8: RAG search and retrieval

Integrating with frameworks such as LangChain makes the tracing of distributed requests seamless. This capability enables early detection of emerging system issues, thereby preventing performance degradation before outages occur. Organizations benefit from detailed workflow analysis, resource allocation insights, and execution insights, end to end from prompt to response.

Dynatrace provides insights into costs, prompt and completion sampling, error tracking, and performance metrics using the logs, metrics, and traces of each specific LangChain task, as shown below.

LangChain workflow analysis in Dynatrace
Figure 9: LangChain workflow analysis in Dynatrace

AI observability helps you get the most out of AI for business success

The required initial investment in GenAI is high, and organizations often only see a return on investment (ROI) through increased efficiency and reduced costs. However, organizations must consider which use cases will bring them the biggest ROI. Finding a balance between complexity and impact must be a priority for organizations that adopt AI strategies.

Organizations need to stay on top of AI developments, and AI adoption is not a one-time event for which they can plan. AI adoption requires an ongoing mindset shift where organizations examine and observe all processes of building and enhancing services beyond the technical stack of their AI applications. Dynatrace helps to optimize customer experiences end to end, tying AI cost to business cases and sustainability, enabling organizations to deliver reliable, new AI-backed services with the help of predictive orchestrations. Dynatrace Real User Monitoring (RUM) capabilities identify performance bottlenecks and root causes automatically, fulfilling the demand for a holistic approach that includes an understanding of intricate system designs and nuanced hidden costs.

Our commitment to customer success is exemplified by how we utilize AI observability for our own benefit, delivering successful AI-based applications at Dynatrace. For instance, based on token counts, we derived a development strategy that not only enhanced reliability but also paved the way for investigating prompt engineering possibilities, as well as deriving measures to better design RAG pipelines to reduce response times. Observing cache hit rates also allowed us to detect model drift in the embedding computations of the AzureOpenAI endpoints, helping Microsoft identify and resolve a bug.

The future of generative AI observability

From facilitating growth to increasing efficiency and reducing costs, GenAI adoption represents a paradigm shift in the industry—a trend consistent with the historical pattern where core technological innovations disrupt prevailing business paradigms. Throughout business history, the advent of pivotal technologies has consistently led to disruptive shifts. Enterprises that fail to adapt to these innovations face extinction. Despite 93% of companies acknowledging the risks of integrating GenAI, the risk mitigation gap hinders companies from progressing at their desired pace.

With GenAI set to become a $1.3 trillion market by 2032, Dynatrace fills this critical risk mitigation gap with AI observability today. Our technologies are evolving quickly from the feedback we receive from clients and partners, helping us to assist them in building successful GenAI applications at scale. Dynatrace offers an expansive suite of nearly 700 integrations to provide unparalleled insights into every facet of your AI stack. From monitoring infrastructure and models to dissecting service chains, Dynatrace provides a comprehensive observability and security solution.

To leverage these integrations and embark on a journey toward optimized AI performance, explore the AI/ML Observability documentation for seamless onboarding. Join us in redefining the standards of AI service quality and reliability.

The post Dynatrace accelerates business transformation with new AI observability solution appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-accelerates-business-transformation-with-new-ai-observability-solution/feed/ 0