The rise of agentic AI | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 19 Mar 2026 13:45:23 +0000 en hourly 1 The rise of agentic AI part 7: introducing data governance and audit trails for AI services https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-7-introducing-data-governance-and-audit-trails-for-ai-services/ https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-7-introducing-data-governance-and-audit-trails-for-ai-services/#respond Tue, 14 Oct 2025 15:59:34 +0000 https://www.dynatrace.com/news/?p=71363 Dynatrace Agentic AI

Your AI investments can’t reach their potential without effective AI governance. AI governance is a challenge that demands unprecedented agility, proactive measures, and comprehensive oversight to manage complexity. With Dynatrace, you’re prepared for whatever comes next. Stay compliant and build trust in your AI systems AI regulation is tightening, and non-compliance is becoming a huge […]

The post The rise of agentic AI part 7: introducing data governance and audit trails for AI services appeared first on Dynatrace news.

]]>
Dynatrace Agentic AI
  • Your AI investments can’t reach their potential without effective AI governance.
  • AI governance is a challenge that demands unprecedented agility, proactive measures, and comprehensive oversight to manage complexity.
  • With Dynatrace, you’re prepared for whatever comes next.

Stay compliant and build trust in your AI systems

AI regulation is tightening, and non-compliance is becoming a huge risk to broader, production-scale AI adoption. Penalties are only part of the impact: reputational damage, customer mistrust, and stalled innovation can cripple even forward-looking organizations. That’s why we’re introducing data governance and audit trails for AI observability: a scalable way to manage, monitor, and secure the AI data lifecycle with end-to-end lineage, retention controls, and evidentiary records of model and user interactions.

Our platform helps you turn governance into a competitive advantage. Built-in audit support helps with emerging regulations like the EU AI Act, and alignment to industry standard frameworks such as NIST AI and ISO/IEC 42001:2023.

The hidden challenges of AI data governance

The complexity of compliance

AI regulations are becoming stricter, and new regulations are on the horizon. Organizations must maintain detailed records of AI activities for years, ensure transparency of data and processes, and align retention policies with legal requirements. These measures are imperative for trust and safety, but they introduce significant challenges. For instance, AI-related events are often scattered across multiple systems, applications, and teams, complicating efforts to create a unified audit trail. Default retention periods can fall short of regulatory needs, and manual governance processes are error-prone and infeasible at scale.

The risk of non-compliance

Failing to meet regulatory standards risks hefty fines and penalties, but the market consequences, reputational damage, and loss of customer trust are even worse. Without the right tools, organizations will struggle to manage the growing complexity of AI data governance and reap the full benefits of AI investments.

Introducing Dynatrace data governance and audit trails

Dynatrace has a long history of empowering organizations to tackle complex challenges with AI-driven solutions. Building on this expertise, we’re introducing a new set of capabilities designed to simplify compliance, enhance transparency, and streamline data management. With Dynatrace, you can:

  • Automatically retain AI-related events for up to 10 years in Grail®, our secure data lakehouse.
  • Monitor and capture events from platforms like Amazon Bedrock, tracking everything from model deployments to fine-tuning activities.
  • Leverage OpenTelemetry to collect real-time traces and metrics of AI workloads, along with every AI user interaction, giving you a complete picture of your AI ecosystem.

Data governance audit in Dynatrace screenshot

Close the compliance gap with embedded oversight

What sets Dynatrace apart is seamless integration with your existing workflows. With OpenPipeline® on Grail, you can route AI-related events to custom storage buckets with extended retention, automatically, and without forcing teams to change tools or processes. This allows long-term auditability and helps meet sector-specific compliance requirements that might require special retention and auditability measures.
Once configured, Dynatrace can automatically route and store events, creating a reliable and transparent audit trail. This helps to reduce fragmentation, tool sprawl, and manual effort traditionally associated with data governance.

Imagine being able to trace every user interaction, model training session, or deployment event with just a few clicks. Dynatrace makes this possible by consolidating fragmented data into a single, coherent view. Whether you’re responding to a regulatory inquiry or optimizing your AI models, you’ll have the insights you need, when you need them.

Simplified and instant data filtering with Dynatrace segments

Not all audit data carries the same compliance weight. For global enterprises with complex IT environments, the ability to instantly filter data by precise criteria is essential for accelerating compliance across diverse regulations, from strict local regulatory transparency obligations to lighter regimes elsewhere.

Dynatrace segments make it simple to break down and filter data to match your analysis needs and regulatory requirements:

  • Targeted compliance views: Instantly filter audit data by region, environment, platform, model, or custom criteria to align with diverse regulatory requirements.
  • Dynamic adaptability: Segments automatically update, for example, when new LLM models or environments are introduced, minimizing manual maintenance and keeping governance current.
  • Reusable assets: Leverage a single dashboard or notebook across multiple use cases by simply applying different segments, reducing duplication of effort.
  • Noise reduction: Exclude irrelevant data such as development or test logs to keep compliance and observability focused on what truly matters.
  • Custom team-context: Provide different teams (for example, compliance, data science, operations) with clear, filtered views of their audit data, ensuring ownership and audit-readiness across departments.

From observability to trusted automation

The future of AI governance lies in proactive, automated solutions that not only meet today’s regulations but also anticipate tomorrow’s challenges. With Dynatrace, you’re not just complying—you’re building a foundation of trust and reliability that scales with your business. By capturing and integrating AI events into a unified platform, Dynatrace transforms compliance from a burden into a strategic advantage.

Get started today

Ready to simplify your AI data governance?

Try it out yourself on the Dynatrace playground. Or, learn how to configure AI governance in our documentation.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two of the Rise of Agentic AI blog series explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.

The post The rise of agentic AI part 7: introducing data governance and audit trails for AI services appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-7-introducing-data-governance-and-audit-trails-for-ai-services/feed/ 0
The rise of agentic AI part 6: Introducing AI Model Versioning and A/B testing for smarter LLM services https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-6-introducing-ai-model-versioning-and-a-b-testing-for-smarter-llm-services/ https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-6-introducing-ai-model-versioning-and-a-b-testing-for-smarter-llm-services/#respond Thu, 25 Sep 2025 16:39:51 +0000 https://www.dynatrace.com/news/?p=71137 Agentic AI - model versioning

Debug, optimize, and secure your AI models with confidence As agentic AI applications and systems gain traction, delivering reliable, high‑performing LLMs and agents becomes challenging due to heterogeneous stacks, non‑deterministic behavior, and cost sensitivity across multi‑cloud runtimes. Reliable delivery and deployment to production requires end-to-end telemetry across the full chain: UI/services → orchestration/agents (LangChain, LlamaIndex, […]

The post The rise of agentic AI part 6: Introducing AI Model Versioning and A/B testing for smarter LLM services appeared first on Dynatrace news.

]]>
Agentic AI - model versioning

Debug, optimize, and secure your AI models with confidence

As agentic AI applications and systems gain traction, delivering reliable, high‑performing LLMs and agents becomes challenging due to heterogeneous stacks, non‑deterministic behavior, and cost sensitivity across multi‑cloud runtimes. Reliable delivery and deployment to production requires end-to-end telemetry across the full chain:
UI/services → orchestration/agents (LangChain, LlamaIndex, MCP/A2A) → RAG pipeline (embedding + vector DB) → model gateway (OpenAI, Azure/OpenAI, Bedrock, Gemini, Mistral, DeepSeek) → GPU/infra. To support deterministic rollouts and continuous model improvement, teams need standardized tracing/metrics, guardrail signal capture, and automated cost and performance governance.

The hidden challenges of AI model management

The invisible bottlenecks

AI models, especially LLMs, are prone to issues like hallucinations, degraded performance, and incorrect outputs. Debugging these problems is often like finding a needle in a haystack. Existing tools fall short in providing a unified view to compare prompts, datasets, or model versions, making it hard to identify regressions or improvements.

The impact of deprecation and automatic upgrades on cost, performance, and quality

The rapid pace of innovation in the AI space means that providers like OpenAI and Anthropic frequently release new versions of their models, such as ChatGPT 5 or Anthropic Opus 4.1.

While these updates often promise better performance and new capabilities, they can also introduce significant risks for your AI services:

  • Deprecation of older versions: Providers may discontinue support for older models, forcing you to adopt newer versions without sufficient time to test their impact.
  • Automatic upgrades: Many AI providers automatically update their underlying models, which can lead to unexpected changes in behavior, degraded performance, or even broken workflows.
  • Compatibility issues: Changes in model behavior, such as output format or token usage, can disrupt your application’s functionality, requiring adjustments to prompts, configurations, or integrations.

Tracking token usage and managing costs is another uphill battle. Add to this the risk of prompt injection attacks and data leaks, and it’s clear that traditional methods are no longer sufficient

The new AI Model Versioning and A/B testing

Ship better models with confidence. In a single view, compare models and versions to validate improvements and spot bottlenecks across latency, reliability, token usage, cost, and output quality, then drill into prompt-level differences to confirm why a variant wins. When something breaks, follow the request end to end with distributed tracing: from input through orchestration steps and model calls to completion, so you can pinpoint exactly where an error or slowdown originated.

Compare models and versions: Detect bottlenecks and validate improvements in a single view.

Trace prompt failures: Debug errors from input to output with our Distributed Tracing solution.

Monitor costs and token usage: Gain real-time insights into token consumption and cost implications.

Detect security and guardrail risks: Identify and alert on vulnerabilities like prompt injection attacks, toxic responses, or captured PII.

Attach your own attributes like user session, feedback, or dataset ID for additional debugging information.

AI Observability model versioning and A/B testing

How it works

With AI Model Versioning, you can track metadata such as model version, dataset ID, and hyperparameters.

A/B testing lets you expose different user segments to model variations, providing data-driven insights into performance metrics like accuracy and cost.

Instrument in minutes: Use the supported OpenTelemetry-based SDK to instrument your service to capture prompts, completions, token usage, errors, and guardrail signals.
You can also enrich spans with attributes like model.version, dataset.id, user/session, and feedback for deeper analysis. (You can read more about this here.)

Start analyzing out of the box: Once data is flowing, the AI Observability app provides ready-made dashboards and distributed tracing so you can compare models/versions, monitor costs and tokens, and debug prompt failures end to end. No extra setup is required; you can try it out on the Dynatrace Playground right now.

 AI Model Versioning, you can track metadata such as model version, data video thumbnail

By combining observability, AI-driven insights, and organizational knowledge, we’re enabling systems that don’t just react but learn and adapt. Each critical issue or incident you resolve fuels a living knowledge base, paving the way for proactive incident prevention through alerting.

What’s next?

We’re committed to enhancing these capabilities further. Upcoming updates will include a dedicated app experience for multi-model and multi-cloud setups, advanced visualization tools, enhanced security features, intelligent forecasting, and alerting for cost/performance and guardrail optimization.

Get started today

Ready to revolutionize your AI services? Here’s how:

  1. Sign up for a free trial.
  2. Install the AI Observability app.
  3. Explore the AI Model Versioning ready-made dashboard, or check it out on our playground

Together, let’s build smarter, more reliable AI systems.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two of the Rise of Agentic AI blog series explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part seven introduces data governance and audit trails for AI services.

The post The rise of agentic AI part 6: Introducing AI Model Versioning and A/B testing for smarter LLM services appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-rise-of-agentic-ai-part-6-introducing-ai-model-versioning-and-a-b-testing-for-smarter-llm-services/feed/ 0
The rise of agentic AI part 5: Developing and monitoring multi-agent applications with OpenAI Agents SDK on Azure AI Foundry https://www.dynatrace.com/news/blog/building-agentic-ai-applications-with-openai-agents-sdk/ https://www.dynatrace.com/news/blog/building-agentic-ai-applications-with-openai-agents-sdk/#respond Mon, 04 Aug 2025 15:36:15 +0000 https://www.dynatrace.com/news/?p=70239 Building agentic AI applications with OpenAI Agents SDK

As agentic AI applications gain ground, the trick becomes how to build multi-agent systems quickly with all the connective tissue built in. In this fifth installment of our series, The Rise of Agentic AI, we explain how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.

The post The rise of agentic AI part 5: Developing and monitoring multi-agent applications with OpenAI Agents SDK on Azure AI Foundry appeared first on Dynatrace news.

]]>
Building agentic AI applications with OpenAI Agents SDK

Recently, OpenAI released a customer service agents demo built using the OpenAI Agents Python SDK that showcases an example multi-agent system at work. With the OpenAI Agents SDK, you can build agentic AI applications with the help of agents, handoffs, guardrails, tools (built-in and custom), and built-in tracing. These capabilities support the core pattern of knowledge, reasoning, and actioning as the foundation for scalable and trustworthy automation, first introduced by Dynatrace CTO Bernd Greifeneder.

In this blog post, we explain how to build an OpenAI agents SDK-based agentic application and instrument the agents and app with AI-powered observability from Dynatrace. Dynatrace can help you see agent executions, tool usages, and prompt flows from initial request to final response for quick root cause analysis and troubleshooting.

To illustrate the capabilities of the OpenAI Agents SDK and agent framework with Azure OpenAI on Azure AI Foundry, we have built our multi-agent solution using the OpenAI customer service agents demo mentioned above as a reference and modified it for our use cases.

About our sample agentic AI application

Our multi-agent system enables users to research, summarize and translate across a range of topics and content. The system consists of four agents:

  • Welcome Agent: Engages the user, reasons with Azure OpenAI to analyze the prompt, identifies the intent, and passes it to the right agent to start processing.
  • Researcher Agent: Searches the web and analyzes the results using OpenAI.
  • Summarizer Agent: Summarizes content, including search results, text, PDF, CSV, and more, using Anthropic Claude.
  • Translator agent: Translates queries and inputs into any user-requested language using OpenAI.
OpenAI Agent SDK sample app architecture
Figure 1: Azure OpenAI Agent SDK setup for demo application in Github

Next, we want the multi-agent system to perform two distinct scenarios:

  1. Context history: In a specified chat session, the entire chat history and context is available for the duration of the session, while the individual prompts might be handed off to different agents for processing.
  2. Composite queries: The app orchestrates multiple different agents for different purposes, such as Research, Translate, Summary, and Welcome, so users can engage to process a composite prompt with multiple sub-queries.

Understanding multi-agent frameworks and handoff workflows

There are some key differences between the agent frameworks. Unlike the A2A protocol, the OpenAI framework does not explicitly have a central registry for agents. Instead, OpenAI agents use the concept of “handoffs” orchestrated by the OpenAI Agent Framework.

OpenAI framework agent handoffs

While orchestrator-led coordination offers a more deterministic and structured workflow, agent-to-agent handoffs provide significant advantages in adaptability and modularity. These handoffs enable agents to collaborate dynamically, making it possible to handle complex, multi-step queries with greater flexibility. This approach focuses on a more decentralized and scalable system, allowing agents to specialize and respond to changing requirements in real-time.

Here are two example scenarios to illustrate the agent-to-agent collaboration in chat sessions, with context, as well as delivering multi-agent query processing.

Show the user prompts for a composite query and multi-agent workflow

For example: “Research Michael Jordan, then summarize in 40 words or less, and then translate to French.”

Welcome Agent user prompt and composite query for the sample agentic AI application
Figure 2: User prompt -> Welcome Agent -> Identifies as multi-step workflow -> Handoff -> Researcher
Researcher Agent, Handoff, and Summarizer activities of the multi-agent workflow
Figure 3: Researcher Agent processes -> Handoff -> Summarizer
Handoff to Translator agent in multi-agent workflow
Figure 4: Summarizer -> Summary -> Handoff to Translator -> Summary results in French
Additional user input triggering translator, researcher, and response in the sample agentic AI application
Figure 5: User chat continues with Context and History -> Translator handoff -> Researcher -> Response
Researcher agent handing off to the translator for translation to Hindi
Figure 6: Researcher -> Handoff -> Translator to translate results to Hindi, keeping context and history

Multi-agent processing for CSV files uploaded

This example includes sample customer data to showcase multi-agent workflow processes with context and history in the chat session.

customer-uploaded CSV file and multi-agent triggers in the sample agentic AI application
Figure 7: Customer Data CSV -> Summarize file -> Welcome Agent -> Handoff -> Summarizer
Countries listed in the CSV file of the sample agentic AI application
Figure 8: “What Countries are listed in the file” -> Summarizer Handoff -> Researcher -> results
Research on the first country in summary
Figure 9: “Research on the 1st country in summary” -> uses context, history -> Researcher -> Results

Overall, the agent-to-agent handoffs worked well (and with context) during all the session runs. Tracing and debugging can be achieved by instrumenting the SDK with OpenTelemetry and sending the data to Dynatrace’s built-in AI Observability solution for Azure OpenAI. You can easily capture the multi-agent workflow for a given prompt on the Azure AI Foundry platform dashboard. Find the code examples in our GitHub repository.

Set up tracing using Python

Using Python, you can set up the tracing by changing a few simple lines of code in your agent framework and core component:

from traceloop.sdk import Traceloop Traceloop.init( app_name="openai-cs-agents", api_endpoint="https://wkf10640.live.dynatrace.com/api/v2/otlp", disable_batch=True, headers=headers, should_enrich_metrics=True, ) 
with tracer.start_as_current_span(name="update_seat", kind=trace.SpanKind.INTERNAL) as span: 
    context.context.confirmation_number = confirmation_number 
    context.context.seat_number = new_seat 
    assert context.context.flight_number is not None, "Flight number is required" 
    return f"Updated seat to {new_seat} for confirmation number {confirmation_number}"

You can see the results right away in distributed tracing:

Results of the OpenAI chat
Figure 10: Multi-agent workflow trace view in Distributed Tracing
Reviewing all OpenAI consumption statistics with Dynatrace AI Observability
Figure 11: How to review all your OpenAI consumption on Dynatrace with AI Observability

OpenAI orchestration

Within the OpenAI framework, there are two approaches to orchestrating agents:

  1. Allow the LLM to make decisions: Use the intelligence of an LLM to plan, reason, and decide what steps to take.
  2. Orchestrate with code: Use code to determine the flow of agents.

Overall, the OpenAI Agents SDK is comprehensive and easy to get running with some minor code changes, this time with OpenAI’s Codex assistant.

OpenAI Agents SDK Codex assistant code example
Figure 12: Codex example

Multiple frameworks and toolkits are quickly ramping up to make multi-agent systems a reality. We foresee this space evolving and innovating rapidly.

The evolution of multi-agent systems

As agentic AI continues to advance, multi-agent applications are poised to play a transformative role in reshaping how applications operate. These systems enable dynamic, context-aware collaboration between specialized agents, empowering businesses to tackle increasingly complex workflows. From helping with automation, orchestrating large-scale data analysis, multi-agent systems will unlock new levels of efficiency, scalability, and innovation.

Tools like the OpenAI Agents SDK on Azure AI Foundry and Azure AI Studio are at the forefront of this evolution. By providing built-in capabilities such as agent handoffs, guardrails, and tracing, the SDK simplifies the development and monitoring of multi-agent workflows. These features make it easier for organizations to deploy responsible, secure, and robust AI systems and also ensure transparency and trustworthiness in their operations. These are key factors for widespread adoption.

Looking ahead, we can expect rapid innovation in this space. Emerging standards like MCP, A2A protocols, and frameworks such as OpenAI Agents are creating a vibrant ecosystem for multi-agent interoperability. The focus will likely shift toward even more intelligent and reliable orchestration, where agents autonomously plan, reason, and adapt to dynamic environments.

AI Observability for agentic AI applications

To keep pace with these advancements, we believe that observability must evolve in lockstep to ensure transparency across heterogeneous agent ecosystems. Advancements in observability tools, such as the Dynatrace AI Observability solution, are essential to help create more reliable and scalable AI frameworks at the enterprise level.

The future of multi-agent systems holds immense potential, with the OpenAI SDK marking the starting point. We’re just at the beginning of what’s possible. As this technology evolves, it will gradually become more stable and reliable, ultimately transforming the way we approach automation, collaboration, and AI-powered problem-solving across industries.

Check out our GitHub repo for detailed code examples for OpenAI Agents, AWS Strands, Google ADK, and start building your own AI Observability solutions today.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two explores how monitoring A2A and MCP communications results in better, more effective agentic AI. This blog post covers AI agent observability and monitoring, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.

Together, these capabilities make it possible to achieve robust, scalable observability in agentic AI environments so teams can build reliable and trustworthy applications and services.

The post The rise of agentic AI part 5: Developing and monitoring multi-agent applications with OpenAI Agents SDK on Azure AI Foundry appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/building-agentic-ai-applications-with-openai-agents-sdk/feed/ 0
The rise of agentic AI part 4: Dynatrace delivers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM https://www.dynatrace.com/news/blog/full-stack-observability-for-nvidia-blackwell-and-nim-based-ai/ https://www.dynatrace.com/news/blog/full-stack-observability-for-nvidia-blackwell-and-nim-based-ai/#respond Fri, 20 Jun 2025 06:00:05 +0000 https://www.dynatrace.com/news/?p=69115 Davis CoPilot for NVIDIA

The Dynatrace® unified, AI-powered observability platform delivers full-stack AI and LLM observability, including of NVIDIA Blackwell and NVIDIA NIM systems, and AI-driven insights to meet the scale and complexity of enterprise AI deployments. In this fourth installment of our series, The Rise of Agentic AI, we explore how the Dynatrace integration with NVIDIA systems provides enterprises with all the insights needed to detect customer-facing issues, helping IT teams maintain performance, reliability, and security across their AI workloads.

The post The rise of agentic AI part 4: Dynatrace delivers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM appeared first on Dynatrace news.

]]>
Davis CoPilot for NVIDIA

NVIDIA Blackwell systems provide high-performance infrastructure for enterprise AI, and now, thanks to the Dynatrace integration with the NVIDIA Enterprise AI Factory reference design, enterprises can add Dynatrace Full-Stack Observability to NVIDIA Blackwell infrastructure. This magnifies the value of the NVIDIA Blackwell platform by providing real-time performance insights, anomaly detection, and dependency mapping.

Keep high performance and security top of mind with unified observability and security

Figure 1. The Dynatrace AI Observability platform
Figure 1. The Dynatrace AI Observability platform

Dynatrace aligns with high data security and privacy standards typical of on-premises NVIDIA Blackwell deployments, particularly in regulated industries such as finance and healthcare. Its unified data model, Smartscape® topology mapping, and Davis® AI engine provide deep visibility into the full stack—from GPU metrics and containerized workloads to distributed applications and user experiences, enabling tailored observability for workloads running on  NVIDIA Blackwell. Integrating NVIDIA Data Center GPU Manager or other telemetry sources is straightforward, allowing teams to monitor GPU health, utilization, thermal thresholds, and memory bandwidth alongside traditional infrastructure metrics.

Dynatrace technology allows for automated discovery and instrumentation of services running on NVIDIA Blackwell-accelerated systems. Whether monitoring high-throughput GPU compute tasks, Kubernetes clusters, or microservices, Dynatrace ensures low-overhead performance monitoring with minimal manual configuration.

AI-powered, real-time insights improve performance and explainability

With Dynatrace Full-Stack AI Observability, you can monitor real-time performance, trace prompts end-to-end, and ensure compliance, optimizing cost and throughput for your AI and LLM workflows and agents, offering various use cases such as

  • Monitor service health and performance, tracking real-time metrics and offering clear visibility into service incidents.
  • Validate service quality by measuring response speed or identifying performance hotspots.
  • End-to-end tracing and debugging pinpoint the root cause of errors and failures in the LLM chain, troubleshoot issues in complex pipelines, and trace dependencies across the entire system spanning multiple LLMs, RAG pipelines, and agentic frameworks.
Figure 2. Sample dashboards provided for tracking service health and performance
Figure 2. Sample dashboards are provided for tracking service health and performance

Unified AI-powered observability

Dynatrace delivers full stack observability for your LLMs and Generative AI applications running on NVIDIA Blackwell systems. Its ability to provide visibility into complex, high-performance environments allows enterprises to fully leverage Blackwell’s capabilities while maintaining operational excellence and system reliability, improving the performance, explainability, and compliance of your AI workloads and agents.

Figure 3. Dig deeper into the possibilities of AI and LLM observability on the Dynatrace Playground
Figure 3. Dig deeper into the possibilities of AI and LLM observability on the Dynatrace Playground

Visit the Dynatrace Playground to learn more and gain hands-on experience with prepopulated data, so you can experience the possibilities of AI and LLM observability with Dynatrace. If you’re interested in using Dynatrace for your own AI workloads, visit our documentation and start benefiting from full stack observability for AI and LLM.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.

The post The rise of agentic AI part 4: Dynatrace delivers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/full-stack-observability-for-nvidia-blackwell-and-nim-based-ai/feed/ 0
The rise of agentic AI part 3: Amazon Bedrock Agents monitoring and how observability optimizes AI agents at scale https://www.dynatrace.com/news/blog/ai-agent-observability-amazon-bedrock-agents-monitoring/ https://www.dynatrace.com/news/blog/ai-agent-observability-amazon-bedrock-agents-monitoring/#respond Tue, 17 Jun 2025 19:00:54 +0000 https://www.dynatrace.com/news/?p=69101 AWS icon and agentic AI

Model-building platforms like Amazon Bedrock provide the foundation for successful agentic AI applications. But effective cross-agent communication requires standardized telemetry. In this third installment of our series, The Rise of Agentic AI, we explain how standardizing and instrumenting tracing and logging, and monitoring Amazon Bedrock Agents helps to debug and deliver better performing agentic AI applications.

The post The rise of agentic AI part 3: Amazon Bedrock Agents monitoring and how observability optimizes AI agents at scale appeared first on Dynatrace news.

]]>
AWS icon and agentic AI

The next big wave in artificial intelligence is agentic AI, which harnesses autonomous agents to perform tasks by reasoning, learning, and adapting to changing circumstances. The success and efficiency of agentic AI systems depend on how well these AI agents communicate. Facilitating this communication requires monitoring AI agents and their underlying communication protocols, such as Model Context Protocol (MCP).

In this blog post, we explain how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.

Key takeaways
  • Effective cross-agent communication requires standardized telemetry. For foundational model-building platforms like Amazon Bedrock, OpenTelemetry-based solutions provide standardization and instrumentation for tracing and logging to debug at scale.
  • End-to-end observability is a key best practice for monitoring agentic AI. AI agent observability best practices include using GenAI semantic conventions with traditional logs, traces, and instrumentation.
  • Observability helps deliver effective agentic AI results in the context of the whole stack. AI agent observability and Amazon Bedrock Agents monitoring help deliver better performance, ensure compliance, and provide detailed debugging tools.

Cross-agent communication requires standardized telemetry

Given the non-deterministic nature of large language models (LLMs) and dynamic cross-agent communication, organizations need standardized telemetry. OpenTelemetry-based GenAI semantic convention libraries are emerging to unify logging, metrics, and tracing in multi-agent ecosystems. Likewise, these standardized instrumentation libraries let you collect and analyze data from each step in an agent’s decision or communication chain on Dynatrace. Observability of each step lets you monitor the communications among your agents and evaluate their health and performance, regulatory compliance, and debugging.

Architecture of travel agent application using Amazon Bedrock Agents and monitoring it with Dynatrace through OpenTelemetry
Figure 1. Architecture of travel agent application using Amazon Bedrock Agents and monitoring it with Dynatrace through OpenTelemetry.

Scale and monitor Amazon Bedrock Agents with Dynatrace

Amazon Bedrock Agents provide an easy way to build and scale generative AI applications with foundation models.

Amazon Bedrock is a fully managed service that offers a choice of high-performing foundation models from leading AI companies, such as AI21 Labs, Anthropic, Cohere, Luma, Meta, Mistral AI, poolside, and Stability AI—or from Amazon’s own model, Amazon Nova—all through a single API. In addition, Amazon Bedrock Agents also provide the broad set of capabilities teams need to build generative AI applications with security, privacy, and responsible AI best practices.

Dynatrace provides an AI-powered, unified observability and security solution for tracking and revealing the full context of used technologies and service interaction topology. Using Dynatrace for AI agent monitoring and MCP monitoring, teams can analyze security vulnerabilities and observe metrics, traces, logs, and business events in real time—automatically and securely.

“With the rise of agents, the need for deep visibility and real-time insights is more essential than ever. Through this partnership, AWS and Dynatrace are uniquely positioned to deliver performance, cost, and quality insights alongside robust compliance monitoring—empowering customers to innovate with confidence.”
– Atul Deo, Director of Amazon Bedrock

Best practices for Agent-to-Agent (A2A) and MCP monitoring

architecture diagram that shows multiple agents interacting with an agentic application
Figure 2. Autonomous agent workflows and task execution.

As with hybrid and cloud-based environments, context-based observability of AI agents and models is essential for efficient and healthy outcomes. Here are some best practices:

  • Adopt common semantic conventions. Standardize metrics and trace attributes—for example, gen_ai.agent.operation.name and gen_ai.agent.name—across different frameworks.
  • Use logs and traces for Amazon Bedrock Agents. Log critical task lifecycle events—capability discovery, artifact creation, agent collaboration steps, API calls—so teams can replay and debug complex interactions and detect hallucinations.
  • Instrument thoroughly. Bake observability into agent frameworks using external OpenTelemetry libraries or by manually instrumenting calls. Ensure each agent’s start, stop, and reasoning steps, like tools, knowledge base, and guardrails, are captured consistently.
  • Secure communication. Enforce enterprise-grade authentication and authorization within agent-to-agent traffic. Use well-defined protocols like A2A to avoid unauthorized data exposure.
  • Continuous feedback. Feed observability insights into iterative retraining or fine-tuning for improved agent reliability.
Screenshot of an example trace showing debugging an Amazon Bedrock agent workflow with Dynatrace AI observability.
Figure 3. Debugging an Amazon Bedrock Agents workflow with Dynatrace AI Observability.

With Amazon Bedrock and the Dynatrace AI Observability solution, you can cover the following use cases for agent observability:

Monitor AI agent service health and performance

  • Detect bottlenecks by tracking real-time metrics, including request counts, durations, and error rates.
  • Manage service costs with automated cost calculations for each request.
  • Stay on track with service-level objectives (SLOs).

Monitor guardrails to ensure compliance

  • Monitor your safeguards customized to application requirements and responsible AI policies.
  • Validate toxicity, filtered content, and denied topics to ensure compliance.
  • Prevent leaks of personally identifiable information (PII).
  • Prevent quality degradation by validating models and usage patterns in real time.

End-to-end tracing and debugging

  • Achieve complete visibility of prompt flows, from initial request to final response, for faster root cause analysis.
  • Capture detailed debug data to troubleshoot issues in complex pipelines.
  • Streamline workflows with granular tracing of LLM prompts, including response latency and model-level metrics.
  • Resolve issues more quickly by pinpointing exact problem areas in prompts, tokens, or system integrations.
Dashboard showing Amazon Bedrock agents monitoring details, such as service health, guardrails, and performance debugging
Figure 4: Dynatrace AI Observability for Amazon Bedrock Agents dashboard covering service health, guardrails, performance, and debugging.

Future of AI agent observability and Amazon Bedrock Agents monitoring

We expect to see deeper integrations between agent orchestration protocols (A2A, MCP) and open observability frameworks, delivering end-to-end visibility from data ingestion to cross-agent collaboration. As standards converge, organizations will rapidly compose advanced AI solutions while retaining full transparency and control, paving the way for even greater scalability, resilience, and confidence in autonomous agents.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.
For AI agent observability and MCP monitoring at scale, check out Dynatrace AI Observability solution and the observability agent samples from Dynatrace on the AWS Labs GitHub site.

The post The rise of agentic AI part 3: Amazon Bedrock Agents monitoring and how observability optimizes AI agents at scale appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ai-agent-observability-amazon-bedrock-agents-monitoring/feed/ 0
The rise of agentic AI part 2: Scaling MCP best practices for seamless developers’ experience in the IDE with Cline https://www.dynatrace.com/news/blog/mcp-best-practices-cline-live-debugger-developer-experience/ https://www.dynatrace.com/news/blog/mcp-best-practices-cline-live-debugger-developer-experience/#respond Wed, 11 Jun 2025 16:51:18 +0000 https://www.dynatrace.com/news/?p=69450 abstract image showing connected dots and waves representing MCP best practices for agentic AI

Agentic AI is dramatically altering how engineering teams respond to incidents and debug use cases. Monitoring agent communications using model communications protocol (MCP) and live debugging are key to gaining insight into how AI models are performing. In this second installment of our series, The Rise of Agentic AI, we explore some MCP best practices that Dynatrace customer TELUS uses to massively accelerate agentic AI issue resolution.

The post The rise of agentic AI part 2: Scaling MCP best practices for seamless developers’ experience in the IDE with Cline appeared first on Dynatrace news.

]]>
abstract image showing connected dots and waves representing MCP best practices for agentic AI

Agentic AI is revolutionizing—and dramatically accelerating—how engineering teams respond to incidents and debug use cases. By monitoring how AI agents communicate using standards such as Model Context Protocol (MCP), performance engineers can gain deep insights into how systems and AI models themselves are performing. And crucially, automate incident responses using coding assistants.

The next step in this scenario is to employ some MCP best practices, such as helping full-stack engineers to debug code at runtime without disrupting operations. Fortunately, Dynatrace has the answer with Live Debugger. First debuted earlier in 2025, Live Debugger gives developers instant access to code-level troubleshooting data in any environment, including production, and is now generally available.

With all these capabilities coalesced in an integrated development environment (IDE), developers can interact with AI using natural language to massively accelerate issue resolution.

Key takeaways:
  • Unify AI debugging resources in one place. Debugging agentic AI workflows starts with combining Dynatrace features with custom MCPs in an IDE.
  • Leverage an AI coding assistant to add context. An AI coding assistant, such as Cline, helps add vital context and improve prompt engineering using natural language queries.
  • Use best practices to configure MCPs and Cline. Combining Cline with MCPs and on-the-fly debugging reduces guesswork and manual overhead.

MCP best practices TELUS uses to accelerate incident response

As an early adopter of Dynatrace Live Debugger, TELUS harnessed the power of Davis® AI, MCP, and real-time debug data to gain a deep understanding of their code, enabling every engineer to ask questions in natural language about their code behavior at runtime.

At a recent Dynatrace Guild meeting, TELUS SREs Dana Harrison and Cheng Li demonstrated how they’re combining Dynatrace features like Live Debugger with custom MCPs for the ultimate purpose-built troubleshooting setup in their Visual Studio Code IDE. By integrating multiple resources in one place, the TELUS team can blend next-generation debugging with agentic AI workflows. This integration reduces developer context-switching and enables developers to use straightforward, natural-language prompts.

With a fully instrumented environment that monitors every service end-to-end, the TELUS team leverages Dynatrace Live Debugger for Node.js and Java microservices to monitor running code in real time to reduce mean time to resolution. They initially tested this approach in non-production but are expanding into production with strict auditing and data masking.

The team also uses MCP in tandem with Cline AI, a coding assistant capable of gathering logs, metrics, and trace data through natural language commands. By unifying multiple data sources under MCP, TELUS avoids manually juggling different dashboards and tools. Developers can simply interact with the AI tools, which fetch relevant context on demand, even correlating Live Debugger snapshots with real-time logs. As a result, the entire debugging and incident investigation process remains inside the IDE, which significantly streamlines performance engineering and accelerates issue resolution.

How MCPs empower AI agents for debugging use cases

As an open standard, MCP connects AI agents to relevant data sources, such as repositories, tools, or external APIs. Instead of bespoke integrations for each data silo, MCP provides a universal interface to connect multiple relevant sources to feed the right context to the models and agents. This interface simplifies how agents access relevant context, leading to better task outcomes, execution, and more consistent performance across complex environments. For reference, you can try out the Dynatrace MCP server on GitHub.

architecture diagram showing Dynatrace MCP monitoring reference architecture
Figure 1. Dynatrace MCP server reference architecture.

MCP best practices: How Cline AI helps to bring in context

Cline AI is a coding assistant in Visual Studio Code that leverages multiple MCP servers to streamline data retrieval from various back-end systems. With plain-language queries (“Search for a specific service,” “Fetch logs from GKE,” or “Investigate errors”), developers can prompt Cline AI to automatically contact the relevant MCP server.

By becoming better at prompt engineering and handling the underlying API calls and correlating responses from Dynatrace MCP, native Kubernetes logs, or even third-party services, Cline AI provides an all-in-one investigation workflow right in the editor. This means developers don’t have to switch among multiple tools or memorize specialized APIs; they simply ask Cline AI what they need in natural language, and the MCP layer manages the rest.

Cline supports MCP-client features, such as dynamic tool discovery, prompt reusability, and adaptive resource access. It also enables powerful features like custom instruction, cline-rule, and memory banks.

TELUS’ best practices for configuring MCPs and Cline

The TELUS team followed these MCP best practices while configuring Cline.

Install Cline AI as VS Code extension

First, the TELUS team installed Cline AI as a VS Code extension, pointing it to their AI proxy, Fuel iX. Fuel iX lets them choose among various large language models based on the user’s requirements and desires.

Use careful prompt engineering and Cline rules to direct specific MCP tools

Cline AI can then interpret a developer’s natural-language prompt—such as “Investigate errors in our Node.js service”, “What is this service about”—and select the appropriate MCP endpoint(s) to gather data. Directing to specific MCP tools for more accurate results can be fine-tuned through careful prompt engineering, creating detailed .clinerules file definitions, and directives built into the MCPs themselves – including predicted requests and response formatting.

Configure each MCP server for a specific data domain

Each MCP server at TELUS is responsible for interacting with a specific data domain: for instance, they have a Dynatrace MCP server to fetch metrics and trace data, a GCP Logs MCP server to pull container and other forwarded logs, and additional servers for services like JIRA issues or PagerDuty.

Rely on centralized agentic AI commands, not vendor-specific APIs

By standardizing how these data sources are exposed, TELUS ensures that Cline AI never has to manage direct integrations or vendor-specific APIs. Instead, the assistant simply issues commands to the MCP servers, which internally handle authentication, query templates, and data normalization. This architecture not only allows TELUS to maintain clear boundaries between data retrieval logic and AI-driven workflows, but also accelerates developer onboarding, making it simple for any team member to debug or investigate issues from within VS Code by asking Cline AI, rather than switching between specialized tools or dashboards.

​​Use MCPs as a standardized endpoint

​In this context, MCP servers act as standardized endpoints, each responsible for a particular data source or application. By relying on Cline for coding assistance, the TELUS team can type a natural-language command (“Investigate issues in our Node.js service”), and Cline will automatically query Dynatrace Grail or GKE logs using the relevant MCP server.

This approach centralizes the complexities of data retrieval in one place and lets the AI assistant produce a consolidated summary or recommended fix, all within the IDE.

Step-by-step guide to set up Cline with Dynatrace

  1. Install Cline through the VS Code Marketplace.
    screenshot showing Cline configuration as part of MCP best practices
  2. Configure Cline with the required credentials, usually an API endpoint and key.
    screenshot showing Cline configuration
  3. Install your MCP. This example uses an internal package, but there are currently many MCPs published for installation.
    screenshot showing MCP installation
  4. Validate that the Dynatrace MCP server is up and running.
    screenshot showing MCP validation
  5. Ask Cline your questions. The agent formats them according to the needs of the MCP(s) you installed.
    Screenshot showing asking a question of Cline
  6. Example response of a production authorization service problem pulled and summarized from Dynatrace.
    screenshot showing results of the question asked of Cline as part of MCP best practices

The objective of this combined setup (Live Debugger + MCP + Cline) is fast and intelligent troubleshooting. Instead of juggling multiple dashboards and CLI tools, an engineer can see code snapshots, error traces, logs, and even recently created JIRA tickets within one session. As mentioned, MCP standardizes queries across diverse systems, so the AI agents can correlate all relevant information, such as referencing a single trace ID to pull logs from GKE, configuration details from the Kubernetes cluster, or known issues in PagerDuty. This AI agent-powered workflow saves time and reduces context switching.

The TELUS team found that combining these insights with on-the-fly debugging significantly reduces the guesswork and manual overhead that typically slows resolution times.

The road ahead: Developer-first observability and IDE-centric debugging at TELUS with Dynatrace

Although TELUS currently uses Live Debugger in mostly non-production environments, they plan to enable it in production using strict governance. Their upcoming strategy involves role-based permissions for debugging sessions, mandatory security training on data privacy, and well-defined data retention policies to keep snapshots transient.

When combined with MCP’s universal integration points, like Dynatrace MCP, developers can quickly drill down into real production issues, see the exact variables at fault, and either push a fix or revert a misconfiguration with minimal downtime.

Looking ahead, the TELUS team plans to integrate Davis Copilot APIs to further simplify natural language interactions with Dynatrace.

In practice, this would allow an AI assistant to automatically translate human-readable queries into Dynatrace Query Language (DQL), making it even easier to retrieve the right metrics, traces, or logs. By layering Davis CoPilot™ on top of their existing MCP approach, TELUS aims to reduce manual DQL writing while offering developers a powerful yet streamlined way to perform more advanced data analysis through simple, intuitive prompts.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.

More MCP best practices and resources for Dynatrace Live Debugger

The post The rise of agentic AI part 2: Scaling MCP best practices for seamless developers’ experience in the IDE with Cline appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/mcp-best-practices-cline-live-debugger-developer-experience/feed/ 0
The rise of agentic AI part 1: Understanding MCP, A2A, and the future of automation https://www.dynatrace.com/news/blog/agentic-ai-how-mcp-and-ai-agents-drive-the-latest-automation-revolution/ https://www.dynatrace.com/news/blog/agentic-ai-how-mcp-and-ai-agents-drive-the-latest-automation-revolution/#respond Tue, 13 May 2025 07:40:45 +0000 https://www.dynatrace.com/news/?p=69029 multiple robot icons linked like a network on a dark background asking the question, what is agentic AI? And what is Model Context Protocol? also represents AI agent observability and Amazon Bedrock agents monitoring

Agentic AI systems—independent AI agents that perform tasks by reasoning, learning, and adapting—are radically changing how enterprises automate tasks and orchestrate complex workflows. In this first installment of our series, The Rise of Agentic AI, we explore agentic AI and how the agents communicate using Agent2Agent and model context protocol (MCP).

The post The rise of agentic AI part 1: Understanding MCP, A2A, and the future of automation appeared first on Dynatrace news.

]]>
multiple robot icons linked like a network on a dark background asking the question, what is agentic AI? And what is Model Context Protocol? also represents AI agent observability and Amazon Bedrock agents monitoring

By now, everyone is aware of generative AI fueled by large language models (LLMs) and generative pre-trained transformers (GPTs). The next level of innovation is agentic AI and the autonomous AI agents that drive it. Using Model Context Protocol (MCP) to facilitate agent-to-agent communication, these systems are revolutionizing how enterprises automate tasks and orchestrate complex workflows.

Powered by LLMs, vector databases, retrieval augmented generation (RAG) pipelines and additional tools, these AI agents are expanding extensively, giving rise to multi-agent systems, cross-agent protocols, and context-sharing standards. But these autonomous agents also introduce new challenges in monitoring, debugging, and security.

We’ll examine in detail the fundamentals of AI agents, models, and the emerging standards that help them communicate, like Agent2Agent (A2A) and Model Context Protocol (MCP).

Key takeaways:
  • Autonomous AI agents are the backbone of agentic AI. These services combine to deliver adaptable automated tasks.
  • AI agents depend on LLMs and orchestration logic. These technologies maintain the agent’s state, session memory, context, and reasoning strategies.
  • Agents depend on protocols, such as A2A and MCP, to effectively communicate. Models and agents need these protocols to manage multi-agent communication.

What is agentic AI?

Agentic AI is an artificial intelligence system made up of independent agents that can take initiative and perform sequences of actions to complete tasks by reasoning, learning, and adapting to changing circumstances.

Dynatrace Chief Technologist Alois Reitbauer described agentic AI this way:

Alois Reitbauer

“It’s really delegating a task to software the way you would delegate it to a human. Say if you wanted to do travel booking, give it some complexity and freedom and some decision points it can make. Like, I have to go to Vegas, I need a hotel, I need a couple of good restaurants to go to, we’re going to be 50 people, fix it with my schedule.”
– Alois Reitbauer in The New Stack

Agentic AI systems rely on AI agents to perform the tasks that lead to the desired outcome.

What are AI agents?

An AI agent is a self-directed autonomous application that harnesses large language model (LLM) reasoning, tool usage, and context-awareness from numerous data sources to carry out tasks.

Agents can think and act independently without outside intervention. Agents can think through chain-of-thought, plan, execute (Reason+Act=ReAct), and refine their actions as needed. Businesses are looking into adopting these autonomous agents for applications such as customer service automation, supply-chain optimization, and content generation.

How do AI agents operate?

AI agents operate similarly to a Michelin-starred chef in a busy kitchen: They continuously gather information, plan, execute, and adjust to reach their desired end goal.

In the chef analogy, the cook surveys orders and available ingredients, decides on a suitable recipe, and then refines the approach based on feedback or resource constraints.

Agents do the same thing in a computational context. Specifically, they observe the world (for example, a user request or a set of data), perform internal reasoning about the best course of action, then carry out the steps needed to fulfill the request. This cycle allows them to respond adaptively to changing conditions, much as a chef would substitute ingredients or modify a dish mid-preparation.

Underpinning this iterative loop is the orchestration layer, which maintains the agent’s state, session memory, and reasoning strategies (such as ReAct, Chain-of-Thought, or Tree-of-Thoughts). Large language models (such as OpenAI’s GPT, Anthropic Claude, Google Gemini, Amazon Nova) provide the core reasoning capability for the agent. The model “thinks” about the user’s query. But the agent gains its power by incorporating additional frameworks or tools that can fetch external information or execute actions in the real world. One way to fetch and provide tools and information is through a unified protocol called Model Context Protocol (MCP).

Additionally, the orchestration layer ensures that multiple rounds of reasoning, tool usage, and tool outputs are all tracked and synthesized before the agent returns a final response to the user. Agents follow these steps in a structured way, so they can produce more accurate, context-rich answers and easily manage complex tasks.

architecture diagram that shows multiple agents interacting with an agentic application
Figure 1. Autonomous agent workflows and task execution.

What is the difference between models and agents?

A model (like a large language model) simply generates outputs based on its training data and the given prompt, typically without any built-in mechanism for session memory, external actions, or complex decision loops and validations.

An agent, on the other hand, includes the model but goes further. It maintains a stateful process (managing multi-turn conversations and thought processes), uses external tools to gather fresh data or perform actions, and follows a defined orchestration logic (such as ReAct and chain-of-thought). Thus, while a model is a core reasoning component, an agent adds the surrounding structure and capabilities needed for autonomous, goal-directed behavior.

What is Agent2Agent (A2A)? How multiple agents communicate with each other

As enterprises slowly adopt multiple specialized agents, interoperability of these services becomes crucial to create reliable experiences. To achieve this, A2A from Google helps to create an open protocol that enables agents—regardless of vendor or framework—to securely exchange information, coordinate actions, and integrate capabilities. By specifying tasks, capabilities, and artifacts in a standardized JSON-based lifecycle model, A2A fosters multi-agent collaboration across otherwise siloed systems.

A2A protocol enables agents to share updates and delegate tasks without overhead. However, direct communication between agents only solves half the problem: These agents also need relevant, up-to-date data and context to drive decisions and be equipped with the right toolset to execute actions.

Without a unified method for accessing diverse data sources, even the most capable multi-agent ecosystem remains limited in scope. The open-source project Model Context Protocol (MCP) fills this gap.

architecture diagram showing two agents using different protocols communicating using A2A protocol as part of an AI agent monitoring and MCP monitoring scheme.
Figure 2. Agent-to-agent communication.

What is Model Context Protocol? How MCPs empower agents

As an open standard, the Model Context Protocol (MCP) connects AI agents to relevant data sources, such as repositories, tools, or external APIs. Instead of the above mentioned integrations for each data silo, MCP provides a universal interface like USB-C to connect multiple relevant sources to feed the right context to the models and agents. This universality simplifies how agents access relevant context, leading to better task outcomes, execution and more consistent performance across complex environments. For managing complex tasks like the ones highlighted above, the Dynatrace MCP server on GitHub helps to get real-time end-to-end observability and MCP data into your daily workflow.

architecture diagram showing Dynatrace MCP monitoring reference architecture
Figure 3. Dynatrace MCP server reference architecture.

What’s next: Monitoring A2A and MCP for better agentic AI

As these technologies evolve, we can expect deeper integrations between agent orchestration protocols (A2A and MCP) and open observability frameworks, delivering end-to-end visibility from data ingestion to cross-agent collaboration. Likewise, as standards converge, organizations will rapidly compose advanced AI solutions while retaining full transparency and control, paving the way for even greater scalability, resilience, and confidence in autonomous agents.

Read more

  • Part two of the Rise of Agentic AI blog series explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part three explains how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.
Check out Dynatrace MCP and Dynatrace AI Observability for AI agent monitoring and MCP monitoring at scale.

The post The rise of agentic AI part 1: Understanding MCP, A2A, and the future of automation appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/agentic-ai-how-mcp-and-ai-agents-drive-the-latest-automation-revolution/feed/ 0