Mrudula Bangera | Dynatrace news https://www.dynatrace.com/news/blog/author/mrudula-bangera/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 09 Jul 2026 11:17:49 +0000 en hourly 1 Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/ https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/#respond Thu, 02 Jul 2026 23:42:51 +0000 https://www.dynatrace.com/news/?p=74651 NVIDIA and Dynatrace

Enterprise AI is rapidly evolving from standalone models to agentic AI systems, where multiple AI agents collaborate to gather information, reason across data sources, and generate complex outputs. These systems unlock powerful new capabilities, but they also introduce significant operational challenges. Organizations must be able to observe, govern, and optimize AI agents, models, and infrastructure in real […]

The post Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q appeared first on Dynatrace news.

]]>
NVIDIA and Dynatrace

Enterprise AI is rapidly evolving from standalone models to agentic AI systems, where multiple AI agents collaborate to gather information, reason across data sources, and generate complex outputs. These systems unlock powerful new capabilities, but they also introduce significant operational challenges. Organizations must be able to observe, govern, and optimize AI agents, models, and infrastructure in real time.

Dynatrace helps support this need by providing broad visibility across key layers of the AI stack—from agent orchestration and model inference to GPU infrastructure and enterprise applications. With Dynatrace, teams can monitor AI workflows, understand model behavior, optimize costs, and help improve reliability as agentic systems scale.

Why agentic AI needs full-stack observability

As organizations build GPU-accelerated platforms for AI training and inference, understanding system behavior becomes increasingly complex, with bottlenecks potentially occurring anywhere – from GPU utilization, model latency, token consumption, and downstream service dependencies.

Dynatrace connects these layers through full-stack AI observability, designed to help teams monitor model performance, trace multi-agent workflows, track GPU and infrastructure utilization, detect bottlenecks across AI pipelines, and potentially accelerate troubleshooting with AI-powered root cause analysis.

This unified visibility helps organizations run AI workloads with the same reliability, efficiency, and operational confidence expected from modern enterprise systems.

This unified visibility helps organizations operate AI workloads with improved visibility and operational confidence. By integrating with NVIDIA AI–Q Blueprint and the NVIDIA Agent Toolkit, Dynatrace enriches agent reasoning with high-quality operational telemetry while at the same time helping teams govern and identify opportunities to optimize costs.

How Dynatrace addresses Agentic AI

Dynatrace is designed to assist your team with monitoring infrastructure usage and model behavior and detecting pipeline bottlenecks and token consumption while improving reliability by accelerating troubleshooting and root cause analysis. It also provides a unified view of AI workflows from agent to model down to the infrastructure, allowing organizations to support responsible AI operations, manage cost, improve performance and support agentic workflows at scale.

Every agentic deployment is customized with different agents, tools, models, and data pipelines; therefore, observability is an important capability for understanding how these systems behave in production. The complexity arises as agents interact with multiple enterprise data sources, including:

  • internal datasets
  • external web and knowledge repositories
  • proprietary research systems
  • models served through NVIDIA NIM and Nemotron

Dynatrace can serve as operational data source for AI agents that may help improve the quality of generated insights and enable more informed decision-making. With flexible integration across customized AI-Q implementations, this architecture also lays out the groundwork for automated analysis, research, and decision making.

How Dynatrace integrates NVIDIA AI-Q

By combining NVIDIA’s AI-Q Blueprint with Dynatrace AI observability, organizations gain the transparency and operational intelligence needed to govern, optimize, and scale complex AI systems.

Dynatrace integrates into AI-Q environments in two ways.

1. Observability and cost intelligence for Agentic AI workflows

The NVIDIA Agent Toolkit generates lightweight OpenTelemetry traces that Dynatrace ingests to visualize agent workflows and model interactions.

Dynatrace automatically maps the underlying infrastructure supporting AIQ deployments including NVIDIA NIM and Nemotron microservices and enriches telemetry with AI-specific signals such as:

  • token usage
  • inference latency
  • model metadata
  • GPU utilization

This provides comprehensive visibility across key components including:

  • AI models and inference workloads
  • agent orchestration pipelines
  • GPU and infrastructure resources
  • enterprise data interactions

With these insights, teams can quickly detect performance bottlenecks across agent pipelines, monitor GPU utilization and overall infrastructure health, and identify inefficient model usage. This visibility can help organizations identify cost optimization opportunities associated with AI workloads. Together, these capabilities position observability as important components for building reliable and scalable AI systems.

2. Dynatrace as a high-quality data source for AI agents

Dynatrace can also serve as an operational intelligence source for AI agents.

Through Model Context Protocol (MCP) integrations, Dynatrace exposes telemetry that agents can use in their reasoning workflows, including:

  • infrastructure performance metrics
  • operational incidents and problems
  • deployment and reliability trends
  • system behavior and resource consumption

This allows AI agents to incorporate real-time operational insights into their decision-making. Instead of relying solely on external data, agents gain contextual awareness of enterprise systems, which may support more informed outputs Dynatrace ingests NVIDIA Agent Toolkit OpenTelemetry traces, model telemetry, and infra metrics exposing operational context via MCP.

Together, these technologies create a powerful foundation for deploying deep research in the enterprise as reflected in the picture below.

Dynatrace AI Observability - NVIDIA
Figure 1: Dynatrace providing AI Observability for NVIDIA AI-Q

AI-Q use cases

The following are illustrative examples of what becomes possible when AI-Q-based research agents incorporate Dynatrace operational data and insights into their reasoning workflows. While NVIDIA AI-Q is a reference framework rather than a formal certified Dynatrace integration, these scenarios show how agentic research systems could use Dynatrace AI observability to generate richer analysis, identify patterns, and support more informed decisions.

Infrastructure migration analysis

AI agents combine Dynatrace operational telemetry such as performance trends, incidents, and deployment velocity with infrastructure and cloud cost data to evaluate platform migration scenarios (for example, OpenShift to AKS). The system produces data-driven recommendations with quantified tradeoffs to support strategic decisions.

Large-scale incident analysis

By analyzing thousands of historical problems, AI agents can identify recurring patterns, understand infrastructure behavior, and correlate technical issues with business KPIs. This enables deep operational insights and long-form analysis that would be difficult and time-consuming for humans to produce.

AI cost governance and optimization

Enterprises can use observability data from Dynatrace to analyze token consumption, model usage, and inefficient data interactions across AI workloads. Agents can identify patterns and suggest potential optimizations such as more efficient models or improved workflows.

Software delivery and reliability insights

DevOps and SRE teams can use agentic analysis to correlate deployments with incidents, assess build quality trends, forecast reliability risks, and identify engineering priorities—using Dynatrace as the trusted operational data source.

Get started today

The post Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/feed/ 0
Beyond LLM-as-a-judge: Establishing LLM evaluations as a foundation for trustworthy agentic AI systems https://www.dynatrace.com/news/blog/llm-evaluations-as-a-foundation-for-trustworthy-agentic-ai-systems/ https://www.dynatrace.com/news/blog/llm-evaluations-as-a-foundation-for-trustworthy-agentic-ai-systems/#respond Fri, 26 Jun 2026 15:54:20 +0000 https://www.dynatrace.com/news/?p=74677

Large language models and agents are rapidly transforming how organizations build software, automate workflows, and interact with data. From copilots to autonomous agents, AI-powered systems are increasingly responsible for answering questions, generating code, and supporting operational decisions. But as organizations move from experimentation to production, measuring performance reliably is no longer optional; this is where LLM evaluations become essential.

The post Beyond LLM-as-a-judge: Establishing LLM evaluations as a foundation for trustworthy agentic AI systems appeared first on Dynatrace news.

]]>

This is the second post in our series on LLM evaluations. In the companion post, Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals, we showed you how to run online evaluations against real GenAI prompt traces and bring quality scores into Dynatrace AI Observability alongside latency, cost, and errors. This post steps back to the fundamentals: what evaluations are, how they work, and the methods teams use to measure AI quality.

Just as traditional software relies on testing frameworks to ensure reliability, AI systems require robust evaluation frameworks to measure the quality, accuracy, and safety of model outputs. Evals are the primary mechanism by which teams build trust in, iterate on, and responsibly deploy AI systems. Without them, organizations may risk deploying systems that produce unreliable answers, hallucinate facts, or quietly degrade in performance over time.

Key takeaways

  • Evaluations are how teams move from “the LLM feels right” to “we can prove the LLM works.”
  • There is no single best evaluation method. The right approach depends on what you’re measuring and why.
  • LLM-as-a-Judge is one powerful tool within the broader evaluation ecosystem, not synonymous with evals as a whole.
  • Online and offline evaluations serve complementary roles: offline for development, online for production monitoring.
  • A mature evaluation strategy combines code-based, model-based, and human-based methods.
  • Evals should be treated as living artifacts — maintained, versioned, and improved over time like any other engineering asset.

Why LLM evaluation is fundamentally different from traditional testing

Traditional software produces deterministic outputs — the same input consistently returns the same result, making pass/fail testing straightforward. LLMs are probabilistic systems: the same prompt can produce different responses depending on context, temperature, and model behavior. This variability makes conventional testing methods insufficient.

Instead of verifying a single correct output, teams must evaluate across multiple dimensions simultaneously:

  • Correctness— does the response answer the question accurately?
  • Relevance — is the output aligned with the user’s intent?
  • Faithfulness — is the response grounded in source data, not invented?
  • Safety and bias — does the output comply with organizational policies?

This transforms evaluation from simple pass/fail checks into continuous measurement of AI quality.

Prompt stream with evaluation results shown in AI Observability app
Figure 1. Prompt stream with evaluation results shown in AI Observability app

The hallucination problem

The most well-known consequence of probabilistic generation is hallucination — when a model produces plausible-sounding but factually incorrect information. This happens because LLMs predict likely word sequences rather than verify facts, which enables powerful reasoning but introduces serious risk in enterprise environments where accuracy is critical.

Addressing this requires evaluation frameworks that track signals like factual accuracy, semantic similarity, groundedness in source data, and consistency across responses. These metrics transform subjective quality judgments into measurable, improvable signals.

What is an LLM evaluation?

An LLM evaluation is a systematic process of testing a model or AI-powered system to determine whether it meets a defined standard of quality. That standard could be factual accuracy, helpfulness, safety, tone, latency, cost-efficiency, or any other measurable dimension that matters to the application.

Evaluations translate vague product goals (“the assistant should be helpful and safe”) into concrete, repeatable measurements. They allow teams to:

  • Catch regressions when a model is updated, or a prompt is changed.
  • Compare candidates — different models, prompt versions, or retrieval strategies — objectively.
  • Build accountability by producing evidence that a system behaves as intended.
  • Accelerate iteration by giving developers fast, structured feedback loops.

Evals exist on a spectrum of formality, from a small hand-curated test set run locally, to a large, automated pipeline running thousands of test cases in CI/CD on every deployment.

AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability
Figure 2. AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability

How do LLM evaluations operate?

At their core, evaluations follow a consistent pattern regardless of their complexity:

  1. Define the task and success criteria. What should the LLM model do, and how will you know when it does it correctly? This is the hardest and most important step.
  2. Assemble a dataset. A set of inputs (prompts, user messages, documents) paired with expected outputs or grading rubrics. Datasets can be human-curated, synthetically generated, or sampled from production traffic.
  3. Run inference. Pass the inputs through the system under test and collect outputs.
  4. Score the outputs. Apply a scoring method — a function, a model, or a human — to assess how well each output meets the success criteria.
  5. Aggregate and analyze. Roll up scores into metrics (accuracy, pass rate, average score), visualize distributions, and compare against baselines or previous runs.
  6. Act on results. Use the findings to accept or reject a change, file a bug, update a prompt, or trigger retraining.

This loop can run manually during development, automatically in CI/CD pipelines, or continuously against live production traffic.

What’s the difference between LLM evaluations and LLM-as-a-Judge?

This is one of the most common points of confusion in the space.

LLM evaluations are the broader discipline — the full process described above. They encompass everything from how you define success to how you collect test data to how you score outputs to how you act on results.

LLM-as-a-Judge is one specific scoring method that can be used within an evaluation pipeline. It involves using a language model (often a strong general-purpose model like GPT-5 or Claude Sonnet 4.6) to automatically assess the quality of another model’s outputs.

Think of it this way: evaluations are the framework, and LLM-as-a-Judge is one type of grader you can plug into that framework — alongside code-based graders, human graders, or embedding-based similarity checks.

  • LLM-as-a-judge handles open-ended, subjective dimensions (tone, creativity, helpfulness) that are hard to capture in code.
  • It scales to large datasets without human effort.
  • It can be surprisingly well-calibrated when prompts and rubrics are carefully designed.

Limitations

  • Inherent biases of the LLM model used to judge (verbosity bias, position bias, self-preference).
  • Requires prompt engineering and validation to ensure the judge is grading what you intend.
  • Adds cost and latency to the evaluation pipeline.
  • Not appropriate for tasks with clear ground-truth answers where code-based checks suffice.

Code-based evaluations

Code-based evaluations use deterministic functions — written in Python or any language — to score model outputs. No secondary LLM model is involved.

How it works

You write a function that takes the model output as input and returns a score. The function might check for exact string matches, run regex patterns, execute generated code and test it, parse JSON and validate its structure, call an external API to verify a fact, or compare numerical results.

Common patterns

  • Exact match — does the output equal the expected answer?
  • Contains / regex match — does the output include a required phrase or follow a required format?
  • Execution-based — for code generation tasks, run the output and check whether tests pass.
  • Structured output validation — parse JSON/XML outputs and verify schema and values.
  • Tool call verification — for agentic tasks, did the model call the right tool with the right parameters?

Strengths

  • Fully deterministic and reproducible.
  • Fast and cheap to run at scale.
  • Easy to understand, debug, and audit.
  • No dependence on a secondary model’s judgment.

Limitations

  • Cannot handle open-ended or subjective quality dimensions.
  • Requires knowing the exact expected output or a verifiable property of the output.
  • Brittle for tasks where there are many valid correct outputs (for example, summarization, creative writing).

Code-based LLM evals are the first tool to reach for whenever a task has a clear, verifiable answer. They form the backbone of any reliable eval suite.

Online vs. offline evaluations

These two modes are not competing approaches — they’re complementary phases of a complete evaluation strategy.

Offline evaluations

Offline evals run against a static, pre-collected dataset before a system reaches production. They’re the evaluation equivalent of unit and integration tests in software development.

  • When: During development, before deploying a new model, prompt, or retrieval change.
  • Dataset: Curated, labeled, or synthetically generated. Often maintained in version control.
  • Latency: Can run in batch; speed is less critical.
  • Use cases: Regression testing, model comparison, prompt optimization, safety red-teaming, fine-tune evaluation.

Key advantage: Full control over the test distribution and ground-truth labels.

Key limitation: The dataset may not reflect real user behavior or the long tail of production inputs.

Online Evaluations

Online evals run against live production traffic in real time or near real time. They observe what is actually happening when real users interact with the system.

  • When: Continuously, in production.
  • Dataset: Real user inputs — unlabeled, unpredictable, and representative.
  • Latency: Must be fast or asynchronous to avoid slowing down user-facing requests.
  • Use cases: Production monitoring, anomaly detection, drift detection, A/B testing, continuous quality assurance.

Key advantage: Captures real-world usage patterns, prompts, and failure modes from production traffic, giving teams the most representative signal for monitoring AI quality over time.

Key limitation: No pre-defined labels; scoring must rely on heuristics, implicit signals (thumbs up/down, re-prompts), or async LLM-as-a-Judge pipelines.

Get started with LLM evaluations today

The field of LLM evals is evolving rapidly. As enterprises deploy increasingly autonomous AI systems, evaluation can play an important role in improving AI accuracy, reliability, and safety.

Here are the trends worth watching and investing in:

  1. Evaluation-driven development. Treat evals as a first-class engineering artifact. Write eval cases before building features, maintain them in version control, and integrate them into CI/CD pipelines — mirroring test-driven development practices from software engineering.
  2. Agentic and multi-step evaluation. As AI systems move from single-turn Q&A to multi-step agents that use tools and maintain state, evaluations must evolve to assess full trajectories rather than just individual outputs. This includes evaluating tool use, planning quality, error recovery, and task completion over long horizons.
  3. Adversarial and safety evals. Red-teaming — probing a system for failures, biases, and unsafe behaviors — is becoming a standard part of the eval lifecycle, especially as regulatory requirements around AI safety mature.
  4. Human-in-the-loop calibration. Even automated eval pipelines benefit from periodic human review to catch drift in what the judge model or scoring function is measuring. Building lightweight human-annotation workflows alongside automated evaluations yields a more reliable signal over time.
  5. Standardization and benchmarking. The industry is moving toward shared benchmarks and eval frameworks (for example, HELM, MMLU, LMSYS Chatbot Arena, OpenAI Evals) that allow apples-to-apples comparisons across models. Building internal evals that complement these public benchmarks will be an increasingly important capability for any team deploying LLMs.
  6. Cost-aware evaluation. As evals scale, cost becomes a real constraint. Emerging approaches include training lightweight specialized judge models, using embedding-based similarity as a cheap first filter, and intelligently sampling which examples need expensive LLM-as-a-Judge scoring.

Organizations that invest early in robust evaluation frameworks and combine them with AI observability will be positioned to scale AI safely across their operations.

Ready to put this into practice?

See our companion blog post, Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals, to learn how dt-evals lets you run LLM-as-a-judge evaluations on real GenAI traces and turn AI quality into a queryable, trendable, and alertable signal inside Dynatrace AI Observability.

Because in the end, AI systems are only as trustworthy as the processes used to evaluate them.

The post Beyond LLM-as-a-judge: Establishing LLM evaluations as a foundation for trustworthy agentic AI systems appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/llm-evaluations-as-a-foundation-for-trustworthy-agentic-ai-systems/feed/ 0
Scaling enterprise AI with confidence: Dynatrace joins the Dell Technologies AI Ecosystem Program https://www.dynatrace.com/news/blog/scaling-enterprise-ai-with-confidence-dynatrace-joins-the-dell-technologies-ai-ecosystem-program/ https://www.dynatrace.com/news/blog/scaling-enterprise-ai-with-confidence-dynatrace-joins-the-dell-technologies-ai-ecosystem-program/#respond Tue, 19 May 2026 19:02:13 +0000 https://www.dynatrace.com/news/?p=74013 Dynatrace and Dell Technologies

Most enterprises have moved past the deployment problem. The harder question is what those workloads are doing in production: where GPU spend is going, how agent chains are behaving, and whether compliance teams can answer when regulators ask. When the answers aren’t clear, the consequences land fast and are rarely contained to one team. That’s […]

The post Scaling enterprise AI with confidence: Dynatrace joins the Dell Technologies AI Ecosystem Program appeared first on Dynatrace news.

]]>
Dynatrace and Dell Technologies

Most enterprises have moved past the deployment problem. The harder question is what those workloads are doing in production: where GPU spend is going, how agent chains are behaving, and whether compliance teams can answer when regulators ask. When the answers aren’t clear, the consequences land fast and are rarely contained to one team.

That’s why Dynatrace is joining the Dell Technologies AI Ecosystem Program, bringing full-stack AI and LLM observability natively into a broad and integrated AI infrastructure ecosystem. Dell delivers the validated, integrated infrastructure to run AI at scale. Dynatrace brings the observability, automation, and governance to operate it with confidence, with visibility from GPU infrastructure to model behavior to end-user experience. Together, they give enterprises the control to match the scale they’ve already built.

The real challenge: AI at enterprise scale

Running AI in a pilot is very different from running it at scale across the business with real users, regulated data, and demanding SLAs. As we’ve worked with enterprises across industries, these failure patterns come up repeatedly:

Cost

As enterprises scale AI, costs spiral rapidly and unpredictably across model providers, GPU clusters, and inference APIs without clear line of sight into what is driving spend or whether it’s delivering value.

Observability gaps

Traditional monitoring tools weren’t built for AI pipelines. Fragmented observability across GPU clusters, orchestration layers, and inference APIs creates blind spots while LLM latency and token throughput fluctuations under load remain difficult to diagnose and even harder to predict.

Agentic complexity

Multi-step agent workflows introduce cascading failure modes. A silent error in one tool call can corrupt downstream decisions across the entire chain.

Compliance & governance

Enterprises need continuous monitoring to detect model drift, hallucinations, and unsafe outputs before they impact end users. Regulated industries need audit trails, data governance, and behavioral monitoring that most AI monitoring bolt-ons simply weren’t built for.

These aren’t edge cases. They’re the norm. And they’re the reason so many AI initiatives stall between pilot and production.

“Agentic AI changes what observability has to do. You’re no longer watching one model respond to one prompt. In agentic AI, every transaction can be unique, and you’re tracing chains of autonomous decisions across dozens of tools and services. That’s the problem Dynatrace was built to solve and Dell AI Factory is exactly the foundation enterprises need to take AI to production at scale.”

— Steve Tack, Chief Product Officer, Dynatrace

Scale AI workloads with confidence

Dynatrace can be integrated into Dell AI Factory environments to cover end-to-end observability of agentic AI and LLM workloads. The goal is straightforward: no blind spots, no surprises, and no manual investigation when something goes wrong. Here’s what that looks like in practice:

  • Unified AI observability to monitor the AI stack. Prompts, Model calls and downstream services, in a single platform that replaces the fragmented tooling most teams rely on today.
  • Automated prevention and remediation with Dynatrace Intelligence®. When AI workloads behave unexpectedly, Dynatrace Intelligence detects anomalies in real time and triggers automated remediation to minimize or eliminate downstream consequences.
  • End-to-end agentic AI tracing. Distributed tracing across multi-step agent chains, tool calls, RAG pipelines, and external integrations gives teams visibility into how AI agent decisions are made and where they go wrong.
  • Automatic topology mapping with Smartscape®. Maps every component in your Dell AI Factory environment, showing in real time how infrastructure, services, and AI models depend on and affect each other.
  • Built-in data governance and audit trails. Track data flows, model decisions, and AI service behavior with governance capabilities designed for regulated industries not retrofitted to them after the fact.
  • Faster resolution with Dynatrace Assist. Natural language querying and AI-generated remediation recommendations help operations teams resolve issues faster, even without deep AI infrastructure expertise.

Built for the industries where AI is becoming mission critical

AI is no longer an experiment. It’s become core infrastructure for the world’s most demanding enterprises, embedded in the decisions, workflows, and customer experiences that keep businesses running. When AI is mission critical, a failure isn’t a learning opportunity; it’s a negative business impact. Tolerance for poor visibility, unexplained latency, or untraceable decisions drops to zero. That’s precisely where Dynatrace AI Observability comes in, giving teams the visibility, control, and real-time intelligence to keep AI running when it matters most.

“The enterprises winning with AI aren’t running one model in one department. They’re operationalizing AI across the business. Dynatrace joining the Dell Technologies AI Ecosystem Program gives those customers the observability foundation to expand AI workloads on Dell infrastructure with the reliability, governance, and efficiency that enterprise-scale demands.”

— Brad Maltz, Senior Director of AI Solutions, Dell Technologies

What this means for joint customers

For organizations deploying on Dell AI Factory infrastructure, the combination of Dell’s validated hardware and software stack with Dynatrace’s intelligent observability platform means:

  • Scale with confidence. Expand production AI across the business without losing visibility or control.
  • Higher AI reliability. Proactive anomaly detection surfaces issues early; moving teams from reactive firefighting to confident operations.
  • Lower risk at scale. Broad stack visibility reduces the unknowns that make executive teams cautious in moving AI to production at scale.
  • Improved ROI on AI investment. When AI workloads run efficiently and every GPU hour is visible, teams can continuously optimize performance and cost.

End-to-end observability isn’t a nice-to-have for AI. It’s a prerequisite for trust, and trust is what turns AI investments into business outcomes. We’re proud to bring that capability to the Dell AI Factory ecosystem, and we’re excited about how this deepening of our relationship with Dell can unlock incredible value for our joint customers on their AI journeys.

Learn more about Dynatrace AI observability today, or reach out to your Dynatrace account team.

The post Scaling enterprise AI with confidence: Dynatrace joins the Dell Technologies AI Ecosystem Program appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/scaling-enterprise-ai-with-confidence-dynatrace-joins-the-dell-technologies-ai-ecosystem-program/feed/ 0
Announcing agentic framework support and General Availability of the Dynatrace AI Observability app https://www.dynatrace.com/news/blog/announcing-agentic-framework-support-and-general-availability-of-the-dynatrace-ai-observability-app/ https://www.dynatrace.com/news/blog/announcing-agentic-framework-support-and-general-availability-of-the-dynatrace-ai-observability-app/#respond Wed, 28 Jan 2026 16:55:26 +0000 https://www.dynatrace.com/news/?p=72664 Agentic ecosystem

As agentic AI becomes mission-critical, systems that reason, act, and self-optimize introduce new operational challenges. Their dynamic and non-deterministic behavior makes them difficult to debug, they can drive unexpected cost spikes, and they inherently lack the auditability required for reliable, enterprise-grade use. Today, we’re excited to announce expanded support for leading agentic frameworks and protocols, […]

The post Announcing agentic framework support and General Availability of the Dynatrace AI Observability app appeared first on Dynatrace news.

]]>
Agentic ecosystem


As agentic AI becomes mission-critical, systems that reason, act, and self-optimize introduce new operational challenges. Their dynamic and non-deterministic behavior makes them difficult to debug, they can drive unexpected cost spikes, and they inherently lack the auditability required for reliable, enterprise-grade use. Today, we’re excited to announce expanded support for leading agentic frameworks and protocols, along with a new dedicated AI Observability app. With this support, you can build, run, and debug agentic AI applications with confidence across AWS, Azure, and Google Cloud.

What’s new: Broader agentic technology support

Dynatrace supports a broad and rapidly growing ecosystem of agentic AI frameworks and protocols, unifying telemetry from these frameworks via OpenTelemetry and OpenLLMetry into a single, correlated observability model, delivering end‑to‑end visibility across clouds, models, tools, and agents from one platform.

  • Amazon Bedrock AgentCore – Dynatrace offers observability for Amazon Bedrock AgentCore agents by collecting metrics such as token usage, model behavior, latency, and errors. This integration provides unified tracing, cost, performance, and guardrail monitoring, along with ready-made dashboards and intelligent anomaly detection and forecasting, helping teams quickly and effectively monitor, troubleshoot, and optimize complex autonomous agent workflows.
  • Amazon Bedrock Strands – Dynatrace supports the Amazon Bedrock Strands Agents SDK, enabling comprehensive visibility into agentic AI systems. By instrumenting Strands-based AI agents with Dynatrace, organizations can monitor agent behavior, tool usage, and dependencies end to end. This helps ensure performance, reliability, and operational insight across distributed environments, supporting the confident development and operation of agentic AI use cases such as chatbots, recommendation systems, and autonomous workflows.
  • LangChain Agents – Dynatrace provides observability for applications built with the LangChain framework, enabling the monitoring of performance, cost, and reliability of Large Language Model (LLM) applications and agents.
  • Google Agent Development Kit (ADK) – Dynatrace provides observability for applications built with the Google Agent Development Kit (ADK), enabling visibility into agent execution, dependencies, and performance. This helps teams understand runtime behavior and maintain reliability as agent-based applications
  • OpenAI Agents SDK – Dynatrace provides observability for observing applications built with the OpenAI Agents SDK, enabling monitoring of agent workflows, model interactions, latency, and errors. This supports improved operational insight, troubleshooting, and performance optimization for agentic AI applications.
  • MCP AI Agent–  Dynatrace provides deep visibility into AI agents communicating via the Model Context Protocol (MCP). By observing both AI agents and MCP servers, organizations gain end-to-end insight into execution flows through tracing, enabling data-driven decisions, performance and cost optimization, and governance for complex agent workflows.
Agentic AI Observability for popular agentic frameworks, powered by OpenTelemetry and OpenLLMetry
Figure 1. Agentic AI Observability for popular agentic frameworks, powered by OpenTelemetry and OpenLLMetry

This agentic coverage is on top of the 40+ LLM technologies that Dynatrace already supports, including OpenAI, Amazon Bedrock, Google Gemini and Vertex, Anthropic, LangChain, NVIDIA, and more.

We’re working closely across AWS, Microsoft Azure, and Google Cloud ecosystems to ensure you have consistent, enterprise‑grade observability for your multi‑AI and multi‑cloud applications.

See it in action in the new AI Observability experience

The AI Observability app is now Generally Available, delivering a purpose-built experience for observing AI workloads end-to-end from agents and LLMs to orchestration layers, emerging protocols, and tools. It gives engineering teams deep, production-ready visibility into how AI systems behave in real time, allowing them to validate changes faster, reduce risk, and confidently ship AI-powered features at scale.

Unlike generic observability views, the AI Observability app is designed specifically for agentic and LLM-driven systems, making it easy to understand complex multi-step interactions, reason about cost and performance trade-offs, and troubleshoot issues across models, tools, and dependencies.

Key capabilities

  • End‑to‑end observability for agentic AI
    • Monitor agent interactions, tool usage, dependencies, latency, and reliability
    • Track token consumption, cost trends, and caching impact
  • Tracing and debugging for complex flows
    • Follow prompts, tool calls, and model invocations from the initial request to the final response
    • Jump from high‑level health to prompt‑level traces in a couple of clicks
  • Actionable insights at scale
    • Rapid A/B testing across model and prompt variants for faster validation
    • Identify bottlenecks and optimize resource utilization with ready‑made dashboards and drill‑downs
  • Security, privacy, and governance
    • Enterprise‑grade controls, auditability, and policy‑aligned routing
    • Guardrail outcomes (for example, toxicity, PII, or denied topics) are surfaced so you can monitor behavior and trends. (Note that guardrail enforcement occurs at the model/provider; Dynatrace captures and visualizes provider‑reported outcomes.)
The Dynatrace AI Observability experience.
Video 1. The Dynatrace AI Observability experience.

Who this solution is for and why it matters

The Dynatrace AI Observability solution is for enterprise teams, including developers, DevOps, SREs, and business leaders who need deep, real-time insights into their cloud native  AI-powered applications and customer experience in a single unified view.

Who benefits the most from this solution?

  • AI Engineering and Data Science: This group includes practitioners who develop and optimize models. They use LLM observability to track metrics related to model performance, such as identifying hallucinations and biases, validating changes, and improving prompt engineering practices.
  • Software Developers: These individuals benefit from observability by gaining insights into application-level performance, which helps them debug and improve overall code quality. Observability tools allow for faster iteration in development cycles.
  • Site Reliability Engineers (SRE): These teams ensure the reliability and performance of AI applications in production environments. They use observability to identify system-level bottlenecks and failures, and to respond swiftly to operational challenges.
  • Application Security Teams: Although not traditionally the primary users, security teams can leverage AI observability to identify and mitigate emerging threats specific to AI applications, such as prompt-injection attacks and data leaks.
  • Compliance and Governance Teams: Responsible for ensuring adherence to regulatory requirements and internal policies, these teams rely on observability to audit model behavior and to identify potential biases or harmful outputs.

What’s next: Agent topology view with Smartscape

We’re committed to further enhancing these capabilities. As agentic systems evolve into distributed networks of models, tools, and decisions, observability must move beyond traces and metrics. Our next focus is the Agentic Topology View, bringing Smartscape-grade visualization to agent execution flows so teams can see how agents interact, invoke tools, propagate errors, and improve performance end to end.

This agentic topology becomes the foundation for a deeper developer experience by connecting production telemetry with prompt management and evaluation workflows. By unifying agent topology, prompt lifecycle, and LLM-as-judge scoring in a single system, we’re helping teams systematically improve the reliability, performance, and quality of agentic AI at enterprise scale.

Agent topology visualizes agent execution flows, showing how they interact with one another.
Video 2. Agent topology visualizes agent execution flows, showing how they interact with one another.

Get started today

Want to “kick the tires” with some example code? Let’s make agentic AI observable, governable, and reliably fast.

The post Announcing agentic framework support and General Availability of the Dynatrace AI Observability app appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/announcing-agentic-framework-support-and-general-availability-of-the-dynatrace-ai-observability-app/feed/ 0
The rise of agentic AI part 6: Introducing AI Model Versioning and A/B testing for smarter LLM services https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-6-introducing-ai-model-versioning-and-a-b-testing-for-smarter-llm-services/ https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-6-introducing-ai-model-versioning-and-a-b-testing-for-smarter-llm-services/#respond Thu, 25 Sep 2025 16:39:51 +0000 https://www.dynatrace.com/news/?p=71137 Agentic AI - model versioning

Debug, optimize, and secure your AI models with confidence As agentic AI applications and systems gain traction, delivering reliable, high‑performing LLMs and agents becomes challenging due to heterogeneous stacks, non‑deterministic behavior, and cost sensitivity across multi‑cloud runtimes. Reliable delivery and deployment to production requires end-to-end telemetry across the full chain: UI/services → orchestration/agents (LangChain, LlamaIndex, […]

The post The rise of agentic AI part 6: Introducing AI Model Versioning and A/B testing for smarter LLM services appeared first on Dynatrace news.

]]>
Agentic AI - model versioning

Debug, optimize, and secure your AI models with confidence

As agentic AI applications and systems gain traction, delivering reliable, high‑performing LLMs and agents becomes challenging due to heterogeneous stacks, non‑deterministic behavior, and cost sensitivity across multi‑cloud runtimes. Reliable delivery and deployment to production requires end-to-end telemetry across the full chain:
UI/services → orchestration/agents (LangChain, LlamaIndex, MCP/A2A) → RAG pipeline (embedding + vector DB) → model gateway (OpenAI, Azure/OpenAI, Bedrock, Gemini, Mistral, DeepSeek) → GPU/infra. To support deterministic rollouts and continuous model improvement, teams need standardized tracing/metrics, guardrail signal capture, and automated cost and performance governance.

The hidden challenges of AI model management

The invisible bottlenecks

AI models, especially LLMs, are prone to issues like hallucinations, degraded performance, and incorrect outputs. Debugging these problems is often like finding a needle in a haystack. Existing tools fall short in providing a unified view to compare prompts, datasets, or model versions, making it hard to identify regressions or improvements.

The impact of deprecation and automatic upgrades on cost, performance, and quality

The rapid pace of innovation in the AI space means that providers like OpenAI and Anthropic frequently release new versions of their models, such as ChatGPT 5 or Anthropic Opus 4.1.

While these updates often promise better performance and new capabilities, they can also introduce significant risks for your AI services:

  • Deprecation of older versions: Providers may discontinue support for older models, forcing you to adopt newer versions without sufficient time to test their impact.
  • Automatic upgrades: Many AI providers automatically update their underlying models, which can lead to unexpected changes in behavior, degraded performance, or even broken workflows.
  • Compatibility issues: Changes in model behavior, such as output format or token usage, can disrupt your application’s functionality, requiring adjustments to prompts, configurations, or integrations.

Tracking token usage and managing costs is another uphill battle. Add to this the risk of prompt injection attacks and data leaks, and it’s clear that traditional methods are no longer sufficient

The new AI Model Versioning and A/B testing

Ship better models with confidence. In a single view, compare models and versions to validate improvements and spot bottlenecks across latency, reliability, token usage, cost, and output quality, then drill into prompt-level differences to confirm why a variant wins. When something breaks, follow the request end to end with distributed tracing: from input through orchestration steps and model calls to completion, so you can pinpoint exactly where an error or slowdown originated.

Compare models and versions: Detect bottlenecks and validate improvements in a single view.

Trace prompt failures: Debug errors from input to output with our Distributed Tracing solution.

Monitor costs and token usage: Gain real-time insights into token consumption and cost implications.

Detect security and guardrail risks: Identify and alert on vulnerabilities like prompt injection attacks, toxic responses, or captured PII.

Attach your own attributes like user session, feedback, or dataset ID for additional debugging information.

AI Observability model versioning and A/B testing

How it works

With AI Model Versioning, you can track metadata such as model version, dataset ID, and hyperparameters.

A/B testing lets you expose different user segments to model variations, providing data-driven insights into performance metrics like accuracy and cost.

Instrument in minutes: Use the supported OpenTelemetry-based SDK to instrument your service to capture prompts, completions, token usage, errors, and guardrail signals.
You can also enrich spans with attributes like model.version, dataset.id, user/session, and feedback for deeper analysis. (You can read more about this here.)

Start analyzing out of the box: Once data is flowing, the AI Observability app provides ready-made dashboards and distributed tracing so you can compare models/versions, monitor costs and tokens, and debug prompt failures end to end. No extra setup is required; you can try it out on the Dynatrace Playground right now.

 AI Model Versioning, you can track metadata such as model version, data video thumbnail

By combining observability, AI-driven insights, and organizational knowledge, we’re enabling systems that don’t just react but learn and adapt. Each critical issue or incident you resolve fuels a living knowledge base, paving the way for proactive incident prevention through alerting.

What’s next?

We’re committed to enhancing these capabilities further. Upcoming updates will include a dedicated app experience for multi-model and multi-cloud setups, advanced visualization tools, enhanced security features, intelligent forecasting, and alerting for cost/performance and guardrail optimization.

Get started today

Ready to revolutionize your AI services? Here’s how:

  1. Sign up for a free trial.
  2. Install the AI Observability app.
  3. Explore the AI Model Versioning ready-made dashboard, or check it out on our playground

Together, let’s build smarter, more reliable AI systems.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two of the Rise of Agentic AI blog series explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part seven introduces data governance and audit trails for AI services.

The post The rise of agentic AI part 6: Introducing AI Model Versioning and A/B testing for smarter LLM services appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-6-introducing-ai-model-versioning-and-a-b-testing-for-smarter-llm-services/feed/ 0