AI Archives | Dynatrace news https://www.dynatrace.com/news/category/artificial-intelligence/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 09 Jul 2026 11:17:49 +0000 en hourly 1 Quit trying to keep up with every new AI tool and keep building https://www.dynatrace.com/news/blog/quit-trying-to-keep-up-with-every-new-ai-tool-and-keep-building/ https://www.dynatrace.com/news/blog/quit-trying-to-keep-up-with-every-new-ai-tool-and-keep-building/#respond Mon, 06 Jul 2026 20:25:58 +0000 https://www.dynatrace.com/news/?p=74727 Dynatrace Intelligence

I think a lot of developers are quietly afraid right now. Not because AI is going to replace them any time soon. Most experienced developers don’t actually believe that. The fear is more subtle than that. They’re afraid of getting left behind, learning the wrong tools or building the wrong way. Afraid that younger developers […]

The post Quit trying to keep up with every new AI tool and keep building appeared first on Dynatrace news.

]]>
Dynatrace Intelligence

I think a lot of developers are quietly afraid right now. Not because AI is going to replace them any time soon. Most experienced developers don’t actually believe that.

The fear is more subtle than that. They’re afraid of getting left behind, learning the wrong tools or building the wrong way. Afraid that younger developers are adapting faster or that others understand all of this better.

I feel that pressure too.

I use Claude® Code CLI constantly. Every day. At this point I can’t imagine building software without it. I’m dramatically faster than I was a year ago, and the amount of experimentation I can do now is honestly kind of incredible. Yet when I open Reddit, LinkedIn, BlueSky, Twitter, YouTube, even TikTok, suddenly I’m hearing about:

  • New coding agents
  • Orchestration systems
  • Autonomous workflows
  • MCP setups
  • Self-healing applications
  • AI IDEs
  • AI browsers
  • AI terminals
  • AI everything

And every one of them is apparently “the future.” How do you know what to ignore and what to adopt?

Look for people who are working in public and demonstrating real value

There’s a lot of confusion around productivity in AI conversations. People equate productivity with:

  • More generated code
  • Faster feature completion
  • One-shot demos
  • Or “look what I built in six minutes”

But none of that necessarily creates value. I can generate a giant pull request full of AI code this afternoon. That doesn’t mean it survives:

  • Integration tests
  • Security review
  • Observability requirements
  • Real users doing unpredictable human things

Real productivity is delivering value to users more quickly. AI helps with that, but developers waste enormous amounts of time chasing hype right now. Good marketing does not automatically equal good value. I’ve seen tools receive massive amounts of attention that, once I dug deeper, simply didn’t improve my workflow in any meaningful way.

The tools that matter are the ones that genuinely make you more productive while still allowing you to maintain understanding and control over what you’re building. That balance matters to me. I don’t want to become disconnected from the software itself.

Yes, developers should pay attention to what other people are doing. Part of being a software engineer has always been committing yourself to continuous learning. You should read Reddit. You should follow developers on social media. You should pay attention to emerging workflows and ideas.

But you also need to learn how to distinguish hype from value. Look for the people who are actually building things in public. Look for the people sharing process instead of certainty.

Ask yourself: “Was this content created to share knowledge, or to drive engagement?”

Those are very different things.

The people I trust most right now are usually the ones showing unfinished work, explaining tradeoffs, discussing failures, and sharing lessons learned from actual implementation. Not the people pretending they’ve solved software development permanently because they chained six agents together in a demo video.

That’s exactly how I ended up changing workflows myself. A friend of mine seemed to be moving leaps and bounds faster than I was. He had piles of documentation, a giant application, and everything looked polished and organized. So, I asked him to show me what he was using.

That was the first time I saw Claude Code CLI in action. Everything clicked. It had access to the filesystem. It could read the actual code I wanted modified. It could maintain context across the project. It could document what it was doing and why.

For the first time, AI coding stopped feeling like a clever autocomplete tool and started feeling like a genuine development partner. I instantly saw productivity, and it was real, not hype.

Keeping learning by doing

The developers who are adapting well right now are not the ones consuming the most AI content. They’re the ones building with AI consistently.

You need reps. The same way you learn almost anything else. You do not become a better golfer by watching YouTube videos about golf swings. At some point you need to stand over the ball and embarrass yourself for a while.

AI coding is the same thing. You need to build things. Not startup ideas. Not the next billion-dollar SaaS. Not some massive architecture astronaut project. Solve a problem in your own house. Maybe your family needs a better way to plan meals. Maybe you want to track something like the number of times you turn off light switches. Maybe you want a better view into your finances.

Those are the best projects right now because you already understand what success looks like. You can judge whether a tool helped you succeed. And you develop real skill that will transfer to whatever the next big thing is.

Learn more about Jeff’s approach to software engineering with AI at 31 Days of Vibe Coding.

The post Quit trying to keep up with every new AI tool and keep building appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/quit-trying-to-keep-up-with-every-new-ai-tool-and-keep-building/feed/ 0
Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/ https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/#respond Thu, 02 Jul 2026 23:42:51 +0000 https://www.dynatrace.com/news/?p=74651 NVIDIA and Dynatrace

Enterprise AI is rapidly evolving from standalone models to agentic AI systems, where multiple AI agents collaborate to gather information, reason across data sources, and generate complex outputs. These systems unlock powerful new capabilities, but they also introduce significant operational challenges. Organizations must be able to observe, govern, and optimize AI agents, models, and infrastructure in real […]

The post Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q appeared first on Dynatrace news.

]]>
NVIDIA and Dynatrace

Enterprise AI is rapidly evolving from standalone models to agentic AI systems, where multiple AI agents collaborate to gather information, reason across data sources, and generate complex outputs. These systems unlock powerful new capabilities, but they also introduce significant operational challenges. Organizations must be able to observe, govern, and optimize AI agents, models, and infrastructure in real time.

Dynatrace helps support this need by providing broad visibility across key layers of the AI stack—from agent orchestration and model inference to GPU infrastructure and enterprise applications. With Dynatrace, teams can monitor AI workflows, understand model behavior, optimize costs, and help improve reliability as agentic systems scale.

Why agentic AI needs full-stack observability

As organizations build GPU-accelerated platforms for AI training and inference, understanding system behavior becomes increasingly complex, with bottlenecks potentially occurring anywhere – from GPU utilization, model latency, token consumption, and downstream service dependencies.

Dynatrace connects these layers through full-stack AI observability, designed to help teams monitor model performance, trace multi-agent workflows, track GPU and infrastructure utilization, detect bottlenecks across AI pipelines, and potentially accelerate troubleshooting with AI-powered root cause analysis.

This unified visibility helps organizations run AI workloads with the same reliability, efficiency, and operational confidence expected from modern enterprise systems.

This unified visibility helps organizations operate AI workloads with improved visibility and operational confidence. By integrating with NVIDIA AI–Q Blueprint and the NVIDIA Agent Toolkit, Dynatrace enriches agent reasoning with high-quality operational telemetry while at the same time helping teams govern and identify opportunities to optimize costs.

How Dynatrace addresses Agentic AI

Dynatrace is designed to assist your team with monitoring infrastructure usage and model behavior and detecting pipeline bottlenecks and token consumption while improving reliability by accelerating troubleshooting and root cause analysis. It also provides a unified view of AI workflows from agent to model down to the infrastructure, allowing organizations to support responsible AI operations, manage cost, improve performance and support agentic workflows at scale.

Every agentic deployment is customized with different agents, tools, models, and data pipelines; therefore, observability is an important capability for understanding how these systems behave in production. The complexity arises as agents interact with multiple enterprise data sources, including:

  • internal datasets
  • external web and knowledge repositories
  • proprietary research systems
  • models served through NVIDIA NIM and Nemotron

Dynatrace can serve as operational data source for AI agents that may help improve the quality of generated insights and enable more informed decision-making. With flexible integration across customized AI-Q implementations, this architecture also lays out the groundwork for automated analysis, research, and decision making.

How Dynatrace integrates NVIDIA AI-Q

By combining NVIDIA’s AI-Q Blueprint with Dynatrace AI observability, organizations gain the transparency and operational intelligence needed to govern, optimize, and scale complex AI systems.

Dynatrace integrates into AI-Q environments in two ways.

1. Observability and cost intelligence for Agentic AI workflows

The NVIDIA Agent Toolkit generates lightweight OpenTelemetry traces that Dynatrace ingests to visualize agent workflows and model interactions.

Dynatrace automatically maps the underlying infrastructure supporting AIQ deployments including NVIDIA NIM and Nemotron microservices and enriches telemetry with AI-specific signals such as:

  • token usage
  • inference latency
  • model metadata
  • GPU utilization

This provides comprehensive visibility across key components including:

  • AI models and inference workloads
  • agent orchestration pipelines
  • GPU and infrastructure resources
  • enterprise data interactions

With these insights, teams can quickly detect performance bottlenecks across agent pipelines, monitor GPU utilization and overall infrastructure health, and identify inefficient model usage. This visibility can help organizations identify cost optimization opportunities associated with AI workloads. Together, these capabilities position observability as important components for building reliable and scalable AI systems.

2. Dynatrace as a high-quality data source for AI agents

Dynatrace can also serve as an operational intelligence source for AI agents.

Through Model Context Protocol (MCP) integrations, Dynatrace exposes telemetry that agents can use in their reasoning workflows, including:

  • infrastructure performance metrics
  • operational incidents and problems
  • deployment and reliability trends
  • system behavior and resource consumption

This allows AI agents to incorporate real-time operational insights into their decision-making. Instead of relying solely on external data, agents gain contextual awareness of enterprise systems, which may support more informed outputs Dynatrace ingests NVIDIA Agent Toolkit OpenTelemetry traces, model telemetry, and infra metrics exposing operational context via MCP.

Together, these technologies create a powerful foundation for deploying deep research in the enterprise as reflected in the picture below.

Dynatrace AI Observability - NVIDIA
Figure 1: Dynatrace providing AI Observability for NVIDIA AI-Q

AI-Q use cases

The following are illustrative examples of what becomes possible when AI-Q-based research agents incorporate Dynatrace operational data and insights into their reasoning workflows. While NVIDIA AI-Q is a reference framework rather than a formal certified Dynatrace integration, these scenarios show how agentic research systems could use Dynatrace AI observability to generate richer analysis, identify patterns, and support more informed decisions.

Infrastructure migration analysis

AI agents combine Dynatrace operational telemetry such as performance trends, incidents, and deployment velocity with infrastructure and cloud cost data to evaluate platform migration scenarios (for example, OpenShift to AKS). The system produces data-driven recommendations with quantified tradeoffs to support strategic decisions.

Large-scale incident analysis

By analyzing thousands of historical problems, AI agents can identify recurring patterns, understand infrastructure behavior, and correlate technical issues with business KPIs. This enables deep operational insights and long-form analysis that would be difficult and time-consuming for humans to produce.

AI cost governance and optimization

Enterprises can use observability data from Dynatrace to analyze token consumption, model usage, and inefficient data interactions across AI workloads. Agents can identify patterns and suggest potential optimizations such as more efficient models or improved workflows.

Software delivery and reliability insights

DevOps and SRE teams can use agentic analysis to correlate deployments with incidents, assess build quality trends, forecast reliability risks, and identify engineering priorities—using Dynatrace as the trusted operational data source.

Get started today

The post Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/feed/ 0
Beyond LLM-as-a-judge: Establishing LLM evaluations as a foundation for trustworthy agentic AI systems https://www.dynatrace.com/news/blog/llm-evaluations-as-a-foundation-for-trustworthy-agentic-ai-systems/ https://www.dynatrace.com/news/blog/llm-evaluations-as-a-foundation-for-trustworthy-agentic-ai-systems/#respond Fri, 26 Jun 2026 15:54:20 +0000 https://www.dynatrace.com/news/?p=74677

Large language models and agents are rapidly transforming how organizations build software, automate workflows, and interact with data. From copilots to autonomous agents, AI-powered systems are increasingly responsible for answering questions, generating code, and supporting operational decisions. But as organizations move from experimentation to production, measuring performance reliably is no longer optional; this is where LLM evaluations become essential.

The post Beyond LLM-as-a-judge: Establishing LLM evaluations as a foundation for trustworthy agentic AI systems appeared first on Dynatrace news.

]]>

This is the second post in our series on LLM evaluations. In the companion post, Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals, we showed you how to run online evaluations against real GenAI prompt traces and bring quality scores into Dynatrace AI Observability alongside latency, cost, and errors. This post steps back to the fundamentals: what evaluations are, how they work, and the methods teams use to measure AI quality.

Just as traditional software relies on testing frameworks to ensure reliability, AI systems require robust evaluation frameworks to measure the quality, accuracy, and safety of model outputs. Evals are the primary mechanism by which teams build trust in, iterate on, and responsibly deploy AI systems. Without them, organizations may risk deploying systems that produce unreliable answers, hallucinate facts, or quietly degrade in performance over time.

Key takeaways

  • Evaluations are how teams move from “the LLM feels right” to “we can prove the LLM works.”
  • There is no single best evaluation method. The right approach depends on what you’re measuring and why.
  • LLM-as-a-Judge is one powerful tool within the broader evaluation ecosystem, not synonymous with evals as a whole.
  • Online and offline evaluations serve complementary roles: offline for development, online for production monitoring.
  • A mature evaluation strategy combines code-based, model-based, and human-based methods.
  • Evals should be treated as living artifacts — maintained, versioned, and improved over time like any other engineering asset.

Why LLM evaluation is fundamentally different from traditional testing

Traditional software produces deterministic outputs — the same input consistently returns the same result, making pass/fail testing straightforward. LLMs are probabilistic systems: the same prompt can produce different responses depending on context, temperature, and model behavior. This variability makes conventional testing methods insufficient.

Instead of verifying a single correct output, teams must evaluate across multiple dimensions simultaneously:

  • Correctness— does the response answer the question accurately?
  • Relevance — is the output aligned with the user’s intent?
  • Faithfulness — is the response grounded in source data, not invented?
  • Safety and bias — does the output comply with organizational policies?

This transforms evaluation from simple pass/fail checks into continuous measurement of AI quality.

Prompt stream with evaluation results shown in AI Observability app
Figure 1. Prompt stream with evaluation results shown in AI Observability app

The hallucination problem

The most well-known consequence of probabilistic generation is hallucination — when a model produces plausible-sounding but factually incorrect information. This happens because LLMs predict likely word sequences rather than verify facts, which enables powerful reasoning but introduces serious risk in enterprise environments where accuracy is critical.

Addressing this requires evaluation frameworks that track signals like factual accuracy, semantic similarity, groundedness in source data, and consistency across responses. These metrics transform subjective quality judgments into measurable, improvable signals.

What is an LLM evaluation?

An LLM evaluation is a systematic process of testing a model or AI-powered system to determine whether it meets a defined standard of quality. That standard could be factual accuracy, helpfulness, safety, tone, latency, cost-efficiency, or any other measurable dimension that matters to the application.

Evaluations translate vague product goals (“the assistant should be helpful and safe”) into concrete, repeatable measurements. They allow teams to:

  • Catch regressions when a model is updated, or a prompt is changed.
  • Compare candidates — different models, prompt versions, or retrieval strategies — objectively.
  • Build accountability by producing evidence that a system behaves as intended.
  • Accelerate iteration by giving developers fast, structured feedback loops.

Evals exist on a spectrum of formality, from a small hand-curated test set run locally, to a large, automated pipeline running thousands of test cases in CI/CD on every deployment.

AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability
Figure 2. AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability

How do LLM evaluations operate?

At their core, evaluations follow a consistent pattern regardless of their complexity:

  1. Define the task and success criteria. What should the LLM model do, and how will you know when it does it correctly? This is the hardest and most important step.
  2. Assemble a dataset. A set of inputs (prompts, user messages, documents) paired with expected outputs or grading rubrics. Datasets can be human-curated, synthetically generated, or sampled from production traffic.
  3. Run inference. Pass the inputs through the system under test and collect outputs.
  4. Score the outputs. Apply a scoring method — a function, a model, or a human — to assess how well each output meets the success criteria.
  5. Aggregate and analyze. Roll up scores into metrics (accuracy, pass rate, average score), visualize distributions, and compare against baselines or previous runs.
  6. Act on results. Use the findings to accept or reject a change, file a bug, update a prompt, or trigger retraining.

This loop can run manually during development, automatically in CI/CD pipelines, or continuously against live production traffic.

What’s the difference between LLM evaluations and LLM-as-a-Judge?

This is one of the most common points of confusion in the space.

LLM evaluations are the broader discipline — the full process described above. They encompass everything from how you define success to how you collect test data to how you score outputs to how you act on results.

LLM-as-a-Judge is one specific scoring method that can be used within an evaluation pipeline. It involves using a language model (often a strong general-purpose model like GPT-5 or Claude Sonnet 4.6) to automatically assess the quality of another model’s outputs.

Think of it this way: evaluations are the framework, and LLM-as-a-Judge is one type of grader you can plug into that framework — alongside code-based graders, human graders, or embedding-based similarity checks.

  • LLM-as-a-judge handles open-ended, subjective dimensions (tone, creativity, helpfulness) that are hard to capture in code.
  • It scales to large datasets without human effort.
  • It can be surprisingly well-calibrated when prompts and rubrics are carefully designed.

Limitations

  • Inherent biases of the LLM model used to judge (verbosity bias, position bias, self-preference).
  • Requires prompt engineering and validation to ensure the judge is grading what you intend.
  • Adds cost and latency to the evaluation pipeline.
  • Not appropriate for tasks with clear ground-truth answers where code-based checks suffice.

Code-based evaluations

Code-based evaluations use deterministic functions — written in Python or any language — to score model outputs. No secondary LLM model is involved.

How it works

You write a function that takes the model output as input and returns a score. The function might check for exact string matches, run regex patterns, execute generated code and test it, parse JSON and validate its structure, call an external API to verify a fact, or compare numerical results.

Common patterns

  • Exact match — does the output equal the expected answer?
  • Contains / regex match — does the output include a required phrase or follow a required format?
  • Execution-based — for code generation tasks, run the output and check whether tests pass.
  • Structured output validation — parse JSON/XML outputs and verify schema and values.
  • Tool call verification — for agentic tasks, did the model call the right tool with the right parameters?

Strengths

  • Fully deterministic and reproducible.
  • Fast and cheap to run at scale.
  • Easy to understand, debug, and audit.
  • No dependence on a secondary model’s judgment.

Limitations

  • Cannot handle open-ended or subjective quality dimensions.
  • Requires knowing the exact expected output or a verifiable property of the output.
  • Brittle for tasks where there are many valid correct outputs (for example, summarization, creative writing).

Code-based LLM evals are the first tool to reach for whenever a task has a clear, verifiable answer. They form the backbone of any reliable eval suite.

Online vs. offline evaluations

These two modes are not competing approaches — they’re complementary phases of a complete evaluation strategy.

Offline evaluations

Offline evals run against a static, pre-collected dataset before a system reaches production. They’re the evaluation equivalent of unit and integration tests in software development.

  • When: During development, before deploying a new model, prompt, or retrieval change.
  • Dataset: Curated, labeled, or synthetically generated. Often maintained in version control.
  • Latency: Can run in batch; speed is less critical.
  • Use cases: Regression testing, model comparison, prompt optimization, safety red-teaming, fine-tune evaluation.

Key advantage: Full control over the test distribution and ground-truth labels.

Key limitation: The dataset may not reflect real user behavior or the long tail of production inputs.

Online Evaluations

Online evals run against live production traffic in real time or near real time. They observe what is actually happening when real users interact with the system.

  • When: Continuously, in production.
  • Dataset: Real user inputs — unlabeled, unpredictable, and representative.
  • Latency: Must be fast or asynchronous to avoid slowing down user-facing requests.
  • Use cases: Production monitoring, anomaly detection, drift detection, A/B testing, continuous quality assurance.

Key advantage: Captures real-world usage patterns, prompts, and failure modes from production traffic, giving teams the most representative signal for monitoring AI quality over time.

Key limitation: No pre-defined labels; scoring must rely on heuristics, implicit signals (thumbs up/down, re-prompts), or async LLM-as-a-Judge pipelines.

Get started with LLM evaluations today

The field of LLM evals is evolving rapidly. As enterprises deploy increasingly autonomous AI systems, evaluation can play an important role in improving AI accuracy, reliability, and safety.

Here are the trends worth watching and investing in:

  1. Evaluation-driven development. Treat evals as a first-class engineering artifact. Write eval cases before building features, maintain them in version control, and integrate them into CI/CD pipelines — mirroring test-driven development practices from software engineering.
  2. Agentic and multi-step evaluation. As AI systems move from single-turn Q&A to multi-step agents that use tools and maintain state, evaluations must evolve to assess full trajectories rather than just individual outputs. This includes evaluating tool use, planning quality, error recovery, and task completion over long horizons.
  3. Adversarial and safety evals. Red-teaming — probing a system for failures, biases, and unsafe behaviors — is becoming a standard part of the eval lifecycle, especially as regulatory requirements around AI safety mature.
  4. Human-in-the-loop calibration. Even automated eval pipelines benefit from periodic human review to catch drift in what the judge model or scoring function is measuring. Building lightweight human-annotation workflows alongside automated evaluations yields a more reliable signal over time.
  5. Standardization and benchmarking. The industry is moving toward shared benchmarks and eval frameworks (for example, HELM, MMLU, LMSYS Chatbot Arena, OpenAI Evals) that allow apples-to-apples comparisons across models. Building internal evals that complement these public benchmarks will be an increasingly important capability for any team deploying LLMs.
  6. Cost-aware evaluation. As evals scale, cost becomes a real constraint. Emerging approaches include training lightweight specialized judge models, using embedding-based similarity as a cheap first filter, and intelligently sampling which examples need expensive LLM-as-a-Judge scoring.

Organizations that invest early in robust evaluation frameworks and combine them with AI observability will be positioned to scale AI safely across their operations.

Ready to put this into practice?

See our companion blog post, Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals, to learn how dt-evals lets you run LLM-as-a-judge evaluations on real GenAI traces and turn AI quality into a queryable, trendable, and alertable signal inside Dynatrace AI Observability.

Because in the end, AI systems are only as trustworthy as the processes used to evaluate them.

The post Beyond LLM-as-a-judge: Establishing LLM evaluations as a foundation for trustworthy agentic AI systems appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/llm-evaluations-as-a-foundation-for-trustworthy-agentic-ai-systems/feed/ 0
Log management for AI workloads: How to bring your logs and telemetry plan into the AI-first century https://www.dynatrace.com/news/blog/2026-log-management-for-ai-workloads-action-plan/ https://www.dynatrace.com/news/blog/2026-log-management-for-ai-workloads-action-plan/#respond Wed, 24 Jun 2026 12:19:44 +0000 https://www.dynatrace.com/news/?p=74644 2026 Logs Report action plan blog

AI is stretching the boundaries of traditional log management. More data without context slows insight, increases risk, and stalls AI progress. Teams need to rethink how they capture, process, and use telemetry from ingest to analysis. This action plan outlines how to unify telemetry, optimize pipelines, and turn data into real‑time, trusted intelligence teams can […]

The post Log management for AI workloads: How to bring your logs and telemetry plan into the AI-first century appeared first on Dynatrace news.

]]>
2026 Logs Report action plan blog

AI is stretching the boundaries of traditional log management. More data without context slows insight, increases risk, and stalls AI progress. Teams need to rethink how they capture, process, and use telemetry from ingest to analysis. This action plan outlines how to unify telemetry, optimize pipelines, and turn data into real‑time, trusted intelligence teams can use to scale AI operations with confidence.

AI workloads aren’t just increasing telemetry—they’re exposing the limits of how teams capture, store, and use it. Traditional log management was built for predictable systems and finite telemetry. AI systems break those assumptions. The result: more telemetry, less context, and increasing cost pressures to stay operational. The shift is how to manage logs better—and changing how teams capture, process, and make logs telemetry available across the entire lifecycle.

Here’s how that shift looks in practice:

Traditional log management AI-ready log management
Indexing and re-indexing, schema-first, archiving and rehydrating  Schema-on-read, always hydrated, always queryable 
Tool-specific context  Unified telemetry context 
Reactive troubleshooting  Preventive operations 

The State of Log Management 2026 research report—based on a global survey of 450 senior IT leaders—examines how AI is reshaping log economics, instrumentation, and observability strategies, and what shifts technical leaders must make to telemetry capture, storage, and management to support and scale agentic AI projects.

5 actions to kickstart your new log management plan

  • Create a single source of truth for AI systems by centralizing all telemetry in a unified, continuously queryable context layer and platform.
  • Establish causation across AI systems by automatically unifying logs with metrics, traces, and lifecycle context—not relying on logs alone.
  • Control costs without losing visibility by optimizing telemetry before ingest and eliminating indexing, archiving, and rehydration dependencies.
  • Standardize and govern telemetry at ingest to ensure data quality, compliance, and real-time usability at AI scale.
  • Enable preventive AI operations by turning contextual telemetry into real-time insight and automated remediation.

Why should teams unify telemetry on a single observability platform for AI workloads?

Unifying telemetry reduces manual correlation, preserves context, and keeps logs, metrics, and traces continuously queryable as telemetry scale increases.

AI workloads are exacerbating an existing problem by fragmenting even more telemetry across tools just as systems require more context. Teams now use an average of seven log tools, forcing manual correlation that doesn’t scale.

  • Unify all telemetry—logs, metrics, traces, security signals, user behavior, business events—into a single, continuously queryable context layer where telemetry is correlated automatically.
  • Enrich telemetry at ingest starting at the edge with shared technical and business context to explain system behavior as dependencies multiply.
  • Democratize access using intuitive querying so more teams can validate AI behavior and act faster with confidence.

How do logs and traces work together for reliable and explainable autonomous operations?

Logs don’t explain AI behavior independently. Understanding comes from unifying logs with traces and other telemetry signals optimized throughout the telemetry lifecycle.

Autonomous systems demand deterministic signals that explain what happened, why it happened, and how to respond—something logs alone can’t fully provide.

  • Instrument logs to capture AI‑specific details at every inference layer to preserve the exact sequence of events.
  • Correlate logs with traces automatically to establish causation and pinpoint root causes.
  • Automate remediation using continuously enriched telemetry to enable reliable, explainable autonomous operations at scale.

How can teams optimize log management costs without sacrificing insight?

Teams can manage costs by retaining high‑value telemetry without rigid schemas, indexing overhead, or rehydration delays that limit analysis.

Managing log costs involves data strategy, not just storage. Logs consume nearly half of observability budgets, yet even after reducing volume by filtering, masking, and aggregating, 50% of organizations don’t collect or discard an average of 86% of logs specifically to manage costs, and 74% say indexing and rehydration costs are barriers to value.

  • Ingest and retain telemetry without rigid schemas or indexes, eliminating the need to predict questions in advance.
  • Store exabytes of data in one queryable layer, avoiding cold archives and rehydration costs and delays.
  • Analyze telemetry in full context to reduce waste and maximize business value from AI‑generated data.

What changes to instrumentation and ingest should teams make to support AI workloads?

Teams must optimize telemetry before ingest—standardizing instrumentation and automating parsing and configurations—so data remains high-quality, contextual, and continuously queryable at AI scale.

Fragmented instrumentation and brittle ingest pipelines slow insight and delay AI projects from reaching production. 85% of organizations struggle to ingest logs at AI scale, and 80% say turning telemetry into insight delays AI initiatives.

  • Standardize instrumentation across logs, traces, metrics, and other telemetry signals to maintain context and reduce downstream correlation.
  • Streamline ingestion in real time by automating parsing, configurations, and enrichment to retain only high‑value, compliant data.
  • Sustain telemetry at scale with an always‑queryable data layer that supports real‑time analytics and automation.

Why are preventive operations critical to AI-native environments?

As AI workloads increase telemetry volume and autonomous operations, teams need detailed intelligence about what’s happening in AI output to predictively detect early signals of unexpected results.

Because reactive troubleshooting can’t keep up with autonomous systems, teams need real‑time, contextual telemetry to detect drift and prevent failures early. 84% say customer trust in AI depends on their ability to use log analytics to predict and prevent problems.

  • Correlate logs with end‑to‑end traces automatically to create a reliable understanding of AI behavior before failures escalate.
  • Analyze AI telemetry in real time and full context to detect early signs of drift or degradation.
  • Automate response to reduce risk and scale AI‑driven operations safely.

Upleveling log management to advance trustworthy agentic AI

Expanding AI workloads demand more from log management—an approach built on unified observability, open, optimized ingest at massive scale, and real‑time analytics without rigid schemas, indexing overhead, or rehydration delays. Logs remain the accountability anchor, but trust emerges only when all telemetry signals come together in context.

Get the State of Log Management 2026 report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI operations.

FAQ: Log management action plan for AI workloads

Why do AI workloads require a different log management approach?

AI workloads are variable and generate significantly more telemetry, which demands explainability, reliability, and cost control capabilities that traditional log architectures weren’t designed to support.

What is the first step teams should take to modernize log management for AI?

Unify logs, metrics, and traces on a single observability platform so telemetry is always available in context and doesn’t require manual correlation.

How can organizations reduce log management costs without losing insight?

By starting observability at the edge and retaining high‑value telemetry without rigid schemas, indexes, or rehydration delays, teams avoid discarding data while controlling cost.

Why aren’t logs alone enough to support autonomous operations?

Logs show what happened and why, but traces show how and what’s affected; together they explain AI behavior and enable reliable, automated remediation.

What enables preventive operations in AI‑native environments?

Real‑time analysis of contextual telemetry that detects early signs of drift or degradation and triggers automated guardrails before failures escalate.

The post Log management for AI workloads: How to bring your logs and telemetry plan into the AI-first century appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/2026-log-management-for-ai-workloads-action-plan/feed/ 0
How AI workloads are changing what logs must deliver, forcing a new strategy https://www.dynatrace.com/news/blog/log-management-research-key-findings/ https://www.dynatrace.com/news/blog/log-management-research-key-findings/#respond Wed, 17 Jun 2026 12:34:03 +0000 https://www.dynatrace.com/news/?p=73713 Future of log management: 2026 research shows AI demands of logs

AI workloads are redefining what logs must deliver—and exposing where traditional approaches fall short. New research reveals how rising scale, cost pressures, and missing context are reshaping log management. The path forward is clear: unify logs with traces and other telemetry in context to turn fragmented signals into reliable insight and build the foundation for trusted, scalable AI operations.

The post How AI workloads are changing what logs must deliver, forcing a new strategy appeared first on Dynatrace news.

]]>
Future of log management: 2026 research shows AI demands of logs

New Dynatrace research reveals how logs are becoming the accountability anchor for AI systems and why cost-driven trade-offs (sampling, cold storage, and discarding data) increase operational risk. In fact, with legacy log management tools now consuming 45% of observability budgets, 67% say the costs of these tools outweigh their value, indicating a need to modernize quickly (source: Dynatrace State of Log Management 2026 research).

3 takeaways for engineering and business leaders

  • Scale is accelerating. Log and telemetry volume increased 93% on average with 1 in 5 organizations seeing growth above 150%.
  • Costs are breaking traditional models. Existing log management tools consume 45% of observability budgets with respondents estimating an average annual spend near $2.5M.
  • Context with traces is the path to AI trust. 73% say logs reveal only part of what’s happening with AI workloads, and 70% rank traces as a top source for evaluating AI performance and behavior. But 65% rely heavily on logs because they don’t have easy visibility into other telemetry signals.

Logs capture the precise details of events from cloud, AI, and infrastructure. As AI workloads and agentic systems (AI systems that act autonomously) make up an increasing proportion of technology stacks, logs serve as a crucial shared language for humans and agents to reason and troubleshoot system state and autonomously build and optimize infrastructure and applications.

Accountability depends on placing high-fidelity log telemetry in context: connecting it with traces, metrics, security events, user behavior, and business signals to understand AI behavior and guide remediation at scale.

How is AI breaking the economics of traditional log management?

AI workloads break the financials of traditional logging approaches by driving massive telemetry growth that forces many teams to not collect or discard data to avoid runaway costs.

Over the past year, AI workloads triggered a 93% average increase in log and telemetry volume, with one in five experiencing expansions above 150%. At the same time, teams rely on an average of seven different log and telemetry tools, forcing manual correlation that doesn’t scale.

The financial impact is just as stark. According to the report, existing log management tools now consume 45% of observability budgets, with average annual spend among survey respondents nearing an estimated $2.5M per organization. To contain costs, many teams limit ingestion, sample telemetry, or push logs into cold storage, losing crucial context and increasing security, compliance, and operational risk. In fact, 67% say the cost of existing log management tools now outweighs their value. And even after filtering, masking, and aggregating to reduce volume, the limitations of traditional logging tools force teams into imprecise compromises, resulting in half of organizations not collecting or discarding 86% of logs specifically to manage costs.

Traditional vs. AI-native log management

What changes

Traditional log management
(cost-first)

AI-native log management
(unified observability)

How teams find answers  Manual stitching slows analysis; teams spend 58% of analysis time correlating telemetry  Automated ingestion, enrichment, and correlation reduces manual work and accelerates time to answers 
AI trust and validation  Logs alone are incomplete; 73% say logs reveal only part of what’s happening in AI workloads  Context builds trust: logs + traces (top-ranked by 70%) + metrics + events show behavior and causality 
Readiness for what’s next  79% worry current ingest/storage won’t meet future needs; instrumentation lags AI requirements  Updated instrumentation starting at the edge and open, automated processing at scale (supported by 81%) for accelerated AI innovation 
Data strategy  Many teams control spend by reducing ingest using sampling, cold storage, or discarding data  Retain high-fidelity telemetry at scale without rehydration/indexing friction 
Tooling approach  Fragmented tooling; teams use an average of seven tools, forcing manual correlation  One real-time observability context layer that unifies logs with metrics, traces, security events, and business signals 
Operational impact  Blind spots and risks grow as half of teams don’t collect or discard 86% of logs  Answers, not guesses; more context preserved for faster diagnosis and safer automation 

What must change in instrumentation and ingestion for autonomous systems?

To build trust in AI workloads and advance the business value of autonomous decision-making, optimizing and streamlining telemetry must start before ingest and be open and automated at a massive scale.

Traditional logging architectures weren’t built for AI‑driven scale or autonomy.

  • As AI workloads proliferate, 79% of technical leaders worry their current ingest and storage approaches won’t meet future needs.
  • 80% say they must update instrumentation to support new AI‑specific metrics.

The operational toll is significant. Teams spend 58% of their analysis time stitching together logs, metrics, and traces before extracting insight, which slows decisions and delays AI projects from moving into production.

To break this bottleneck, 81% of organizations say log ingestion and processing must be open and automated at massive scale, enabling real‑time analysis without rigid schemas, indexing overhead, or rehydration delays.

Why do logs need an observability ecosystem to build AI trust?

While logs are a crucial component of AI observability, they need context to tell the whole story.

  • 73% of organizations say logs reveal only part of what’s happening in AI workloads
  • Teams spend 58% of analysis time stitching telemetry together
  • 72% say standalone log management tools are obsolete—AI workloads demand a platform approach that combines all types of telemetry in one place to accelerate time to answers

As a result, most organizations rely on logs alongside other telemetry signals—especially traces, which 70% rank as the top source for evaluating AI performance and behavior.

Trust in AI develops when analysis is based on high-fidelity telemetry in context:

  • Logs provide the fact basis
  • Traces expose flow and causality
  • Metrics quantify performance
  • Security events surface risk
  • Business signals connect system behavior to outcomes

Unified observability turns this telemetry into a coherent narrative—enabling teams to validate AI behavior, assess impact radius, guide remediation, and move from reactive troubleshooting to preventive operations.

What does AI-native log management look like with unified observability?

Unified observability transforms log management for the AI era by optimizing telemetry instrumentation and ingest, unifying telemetry with context, and eliminating crippling cost-cutting measures.

With exponential increase in telemetry volumes due to AI workloads, the path forward can’t just be shrinking log ingest to manage costs. AI innovation depends on leaning into telemetry volume with the right strategy and capabilities. The report points to key actions leaders can take to adopt AI-native log management practices.

  • Centralize all telemetry (logs, traces, metrics, events) in a unified, continuously queryable context layer to eliminate silos and scale AI visibility.
  • Automatically correlate logs with traces and lifecycle context to establish causation, understand AI behavior end to end, and enable reliable autonomous operations.
  • Control log costs before ingest while retaining full-fidelity telemetry in exabyte-capacity storage with no rigid schemas, indexes, cold archives, or rehydration.
  • Standardize instrumentation and optimize ingestion to capture high-value, governed telemetry that tracks agent actions, reasoning traces, and lifecycle events.
  • Enable preventive operations by detecting early signals and automating remediation to reduce risk, strengthen reliability, and safely scale AI projects.

Fragmented vs Unified log management

The goal is reliable operations and AI accountability at scale. Together, logs, traces, metrics, and events in context can enable preventive operations, reliable autonomy, and confident decision‑making as AI systems move from pilots into production. Organizations that build this unified foundation will be best positioned to scale AI without sacrificing trust.

Download the State of Log Management 2026 report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

FAQ: State of Log management 2026

Why is log management getting harder in the AI era?

Because AI workloads are driving rapid telemetry growth—organizations saw a 93% average increase in log and telemetry volume in the past year.

Why do log management costs feel out of control?

Existing log management tools consume 45% of observability budgets, and average annual spend is nearing $2.5M per organization.

Do teams still see value in log management at today’s price?

Not consistently. 67% of respondents say costs of existing log management tools now outweigh their value.

Why do teams discard so many logs?

Cost pressure. Using existing tools, 50% of organizations don’t collect or discard 86% of their logs on average, often using sampling or limiting ingestion specifically to reduce spend.

What’s the biggest operational bottleneck with today’s tooling?

Manual correlation. Teams spend 58% of analysis time stitching together logs, metrics, and traces before they can extract insight.

Why aren’t standalone log tools enough for AI workloads?

Because logs alone rarely tell the whole story. 72% say standalone log management tools are obsolete, and 73% say logs reveal only part of what’s happening in AI workloads.

What telemetry signal do teams rely on most to evaluate AI behavior?

Traces—70% rank them as the top source for evaluating AI performance and behavior beyond logs alone.

The post How AI workloads are changing what logs must deliver, forcing a new strategy appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/log-management-research-key-findings/feed/ 0
Orchestrate multicloud AI agents for autonomous incident resolution https://www.dynatrace.com/news/blog/orchestrate-multicloud-ai-agents-for-autonomous-incident-resolution/ https://www.dynatrace.com/news/blog/orchestrate-multicloud-ai-agents-for-autonomous-incident-resolution/#respond Mon, 15 Jun 2026 20:11:14 +0000 https://www.dynatrace.com/news/?p=74557 Observability data

Cloud SRE Agents is a Dynatrace app that orchestrates AWS®, Azure®, and Google® AI agents for automated investigation and resolution assistance for incidents across multicloud environments. Cloud SRE Agents routes identified issues based on configurable rules, centralizes its findings, and provides a single audit trail for autonomous operations.

The post Orchestrate multicloud AI agents for autonomous incident resolution appeared first on Dynatrace news.

]]>
Observability data

Organizations are evolving from human-driven operations to supervised autonomous operations, where AI investigates, recommends, and remediates, and humans stay in control of what matters most. A big part of delivering on that vision is working with the agents that customers already run in their cloud environments.

Harness the power of hyperscale agents

Each hyperscaler has AI agents that automatically investigate and help resolve production incidents using native cloud telemetry and tools. They act like embedded site reliability engineers, analyzing issues and recommending or executing remediation steps without waiting for a human to start the process.

AWS DevOps Agent provides investigation and remediation in AWS using native tooling. An Azure SRE Agent specializes in investigating and remediating Azure issues. And Google Gemini Cloud Assist is for incident analysis across Google Cloud Platform (GCP).

Over the past year, we’ve published how Dynatrace supercharges each of these cloud agents individually. When an issue occurs, Dynatrace Intelligence combines causal, predictive, and agentic AI using the Smartscape dependency graph to automatically link related symptoms and root causes across the environment into one unified problem card.

When Dynatrace integrates with the AWS DevOps Agent, dependency-aware root cause analysis combines with AWS frontier-agent capabilities, and joint customers report up to 70% reductions in mean time to resolution. When Azure SRE Agent connects with Dynatrace, deterministic, causation-based AI flows directly into Azure-native remediation workflows, cutting the back-and-forth between teams. And with Google Gemini Cloud Assist, Dynatrace delivers the same production context layer to GCP-hosted incidents: precise root cause, full topology, real business impact.

Problem detected by Dynatrace Intelligence, investigated and remediated by AWS DevOps Agent (see documentation in the right-hand panel)
Figure 1. Problem detected by Dynatrace Intelligence, investigated and remediated by AWS DevOps Agent (see documentation in the right-hand panel)

From integrations to intelligent orchestration

Many enterprises run workloads across AWS, Azure, and Google Cloud simultaneously, and managing three separate integrations with separate routing logic and separate cost controls is its own operational tax. Cloud SRE Agents provides a single orchestration layer that routes problems to specific hyperscaler agents based on configurable profiles to see everything happening across all three cloud agents.

The Cloud SRE Agents app writes findings back to Dynatrace, and provides your team with measurable visibility into autonomous actions.

The Overview tab's interactive graph shows a live view of problems and their activity status, grouped by related SRE agent.
Figure 2. The Overview tab’s interactive graph shows a live view of problems and their activity status, grouped by related SRE agent.

How Cloud SRE Agents works

When Dynatrace Intelligence detects a problem and identifies the root cause, Cloud SRE Agents calls dedicated cloud-native agents from AWS, Azure, and Google Cloud to retrieve deeper insights from the sources that only they can reach: CloudTrail history, Azure subscription policy, GCP project IAM, recent deployments, and native runbooks. These agents run in parallel, gathering evidence as soon as the problem is detected. Their findings, and, where applicable, the recommended remediation path, are displayed in the same Dynatrace problem view that the on-call SRE is already using in their day-to-day workflow.

One view. No tab-switching. The work starts without you.

Three workflows do the orchestration in the background:

  • Investigate evaluates your Interaction Profiles and dispatches matching problems to the right agents in parallel.
  • Periodic Tasks polls each cloud provider for completion, detects stalled or timed-out investigations, and writes findings back as problem annotations.
  • Event Handlers normalize the cloud-provider event stream so every action correlates back to its originating problem, end to end.

Cloud SRE Agents has the insights and intelligence to decide which agent gets which problem, tracks each run to completion, and brings the answers back together in a single view. The Overview tab provides a real-time, interactive network graph of problems, agents, and activities. The replay view allows the user to step back in time and get an overview of what has happened when, as well as the status of each investigation.

Replay functionality in the Cloud SRE Agents Overview
Figure 3. Replay functionality in the Cloud SRE Agents Overview

Intelligent routing with Interaction Profiles

In agentic operations, routing rules make the difference between turning autonomous systems loose on every alert and pointing them precisely where they earn their keep. Interaction Profiles are how you express routing judgment in Cloud SRE Agents. Each profile pairs a set of conditions with the agent or agents that should handle the problems flagged by the profile, and evaluates the conditions whenever Dynatrace Intelligence detects a problem.

The conditions you can write are deliberately broad. You can route by the cloud account, subscription, or project an incident touches; by problem category (availability, error, slowdown, resource contention); by affected entity type (a Kubernetes cluster, a database, a Lambda function); by tag, label, or any custom attribute carried in the problem record. Conditions combine with AND/OR logic and nest as deeply as you need, keeping real production routing policy inside the app rather than spilling into custom workflows or scripts.

Three ways teams put it to work

Route problems to the right cloud, automatically

A spike in Lambda error rates belongs to AWS DevOps Agent. An Azure App Service degradation calls for Azure SRE Agent. A Pub/Sub latency issue lands with Gemini Cloud Assist. In a multicloud estate, none of those decisions should fall to a human at 2:00 AM. A profile filtered by AWS Account ID, Azure Subscription ID, or GCP Project ID, then narrowed by resource type or tag, settles the routing question once. Every matching problem is automatically routed to the right specialist with the right cloud-native context.

Optimize spend with budget-aware routing

Cloud AI agents do work, and that work has a cost. Cloud SRE Agents lets you set a Monthly Duration Budget per agent and gate dispatch on it via a Has Available Budget filter: once the budget is exhausted, new investigations either stop (in strict enforcement mode) or proceed with a logged warning. The duration figure itself is a proxy, derived from Dynatrace event timestamps rather than the cloud provider’s clock, which makes it useful as a circuit breaker and directional signal, not a substitute for AWS, Azure, or GCP usage reports. The governance value is what matters: you decide how much autonomous investigation you’re willing to underwrite each month, and the system holds the line.

Tier autonomous investigation by problem type and entity

Not every Dynatrace problem warrants an autonomous investigation. Problem Category filters let you dispatch agents only to the problem categories that warrant it, for example, availability or error problems that require immediate action, rather than slowdowns or custom alerts where human triage might still be the right call. Layer on Entity Type filters, and you can further focus on specific infrastructure tiers (hosts, services, process groups, Kubernetes clusters). The result is a tiered model: high-severity issues receive immediate autonomous investigation, lower-severity signals queue for human review, and your team controls the threshold.

Governance that makes autonomous work measurable

Agentic operations earn trust when teams can see what the agents did, why, and whether it worked. Cloud SRE Agents treats that as a first-class concern, with two views built for the two audiences who care about it.

The Activity tab is the audit trail. Every investigation and mitigation appears as a card on a unified timeline; expand any card to see the agent’s full findings, the evidence it pulled, and the action it took or recommended. Each response can be rated Good, OK, or Bad, building a quality signal grounded in what your team actually saw rather than what the system predicted. When a single problem triggers work across multiple agents, those activities roll up to a single status (in progress, done, or stalled), so you always know where things stand without having to reconstruct the run from individual records.

Activity tab showing an expanded investigation card with agent findings and rating control.
Figure 4. Activity tab showing an expanded investigation card with agent findings and rating control.

The Statistics tab is where autonomous operations become a number you can show to a leadership team: problems handled, mitigations executed, average investigation time, MTTR and MTTI trends, success rates, and satisfaction scores broken down by agent. The same view doubles as a directional cost lens, since agent working time is the dominant driver on the cloud side of the bill. Treat the number as a trend signal and a circuit-breaker input, not a billing record (reconcile against AWS, Azure, and GCP usage reports for exact spend), and it makes the case for expanding agentic coverage with evidence rather than anecdote.

The Statistics tab shows key metrics and per-agent insights across a selected time range.
Figure 5. The Statistics tab shows key metrics and per-agent insights across a selected time range.

Why production context multiplies the value

What changes Cloud SRE Agents from a smart dispatcher into something more is what Dynatrace Intelligence contributes before an agent ever begins its analysis. Dynatrace delivers deterministic, causation-based root cause analysis grounded in Dynatrace’s Smartscape real-time dependency mapping, alongside business impact assessment and correlated telemetry. That context shapes the entire direction of the investigation. A cloud agent arriving with that foundation starts from “this specific service on this specific host is the root cause, and here’s the customer impact” rather than “something is wrong somewhere in this account.”

The numbers reflect it. According to AWS, organizations using the AWS DevOps Agent with Dynatrace see up to a 75% reduction in mean time to resolution.

Western Governors University, which runs a fully online learning environment for 200,000 students, uses AWS DevOps Agent with Dynatrace to automate cross-system correlation that previously required manual effort across multiple tools. At a larger scale, United Airlines transports more than 500,000 passengers daily across a hybrid environment that includes more than 500 AWS accounts, 20,000 Lambda functions, and 38,000 OneAgent deployments.

The team’s description of the before and after status is direct: previously, multiple tools with overlapping functions created gaps and black boxes during troubleshooting. With AWS DevOps Agent and Dynatrace, Dynatrace identifies the responsible layer, the agent investigates and provides resolution steps, and everything surfaces in a single Dynatrace view. No 3:00 AM tool-switching required.

Get started

For a closer look at the individual integrations, read the posts on AWS DevOps Agent and Dynatrace and Azure SRE Agent and Dynatrace, or see how Dynatrace Intelligence powers autonomous operations. To put your cloud agents to work today, install Cloud SRE Agents from the Dynatrace Hub. Cloud SRE Agents is currently available as a community-supported app.

Harness the power of your hyperscaler agents

The post Orchestrate multicloud AI agents for autonomous incident resolution appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/orchestrate-multicloud-ai-agents-for-autonomous-incident-resolution/feed/ 0
Dynatrace observability is now a Kiro power https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/ https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/#respond Fri, 12 Jun 2026 21:12:32 +0000 https://www.dynatrace.com/news/?p=74536

In this blog, we'll introduce the Kiro power for Dynatrace, show what it unlocks for developers, and walk you through how to get it up and running.

The post Dynatrace observability is now a Kiro power appeared first on Dynatrace news.

]]>

What is the Kiro power for Dynatrace?

The Kiro power for Dynatrace delivers live observability data, root cause analysis, and remediation suggestions directly into the Kiro IDE, with no JSON editing or manual MCP setup.

Kiro is an AI-powered IDE that helps developers move from idea to working code through spec-driven development and an agentic assistant. To make the assistant genuinely useful in unfamiliar domains, Kiro recently introduced powers: curated, partner-validated bundles of MCP servers, steering files, and best practices that install with a single click and load on demand when a relevant task comes up. Install a power, and Kiro’s agent gains specialized expertise the moment you need it.

For Dynatrace customers already working in Kiro, it’s the shortest path yet from code to production insight. For developers new to Dynatrace, it’s a one-click way to ground Kiro’s reasoning in real facts from your environment, not guesses.

Why this matters for developers

Developers have historically been one step removed from production. When something breaks after deployment, the path to figuring out what went wrong usually runs through a Site Reliability Engineering (SRE) or operations team, and AI coding assistants can’t automatically and reliably remediate issues in software they’re unfamiliar with. Agents that can write code are guessing about how their code behaves in production unless they have access to real telemetry data.

The Dynatrace Kiro power for Dynatrace closes this gap through Dynatrace Intelligence, the agentic operations system at the core of the Dynatrace platform. Kiro’s answers are grounded in deterministic, causal AI and real-time production data, not probabilistic guesses.

When a developer starts a task by writing a prompt, Kiro evaluates the conversation, identifies the relevant power using keywords, and dynamically activates power. Kiro then loads Dynatrace MCP tools and power instructions, providing skills to investigate problems, query live observability data, surface root causes, and even execute and verify remediations.
Figure 1. When a developer starts a task by writing a prompt, Kiro evaluates the conversation, identifies the relevant power using keywords, and dynamically activates the power. Kiro then loads Dynatrace MCP tools and the power instructions, providing the skills needed to investigate problems, query live observability data, surface root causes, and even execute and verify remediations.

With the tools provided by the Kiro power, developers can:

  • Investigate live incidents and get root cause analysis directly in Kiro chat
  • Query metrics, logs, and traces from production using natural language
  • Surface security vulnerabilities affecting the code they’re working on
  • Get remediation suggestions grounded in what’s actually happening in their environment

“Using Kiro powers for Dynatrace has been a total game-changer in the observability space. Deep-dive root cause analysis of complex system issues that once required lengthy manual intervention now happens in seconds, giving us unprecedented speed and confidence.”

Mike Kobush, Sr. Software Performance Engineer, NAIC

How to install the Kiro power for Dynatrace

Getting started takes only a few steps. Once installed, the Kiro power activates automatically when Kiro detects a relevant task. Mention an incident, a slow service, or anything that needs production context, and the Dynatrace tools and guidance will load in Kiro chat.

Prerequisites

  • A Dynatrace account. If you don’t already have one, you can start a free 15-day trial.
  • Kiro installed on your system.

Prepare the Dynatrace connection

First, create a Dynatrace Platform Token, which Kiro will use to authenticate. Then add the required permissions for the Dynatrace MCP server.

Install the Kiro power

The power can be installed from either the Kiro IDE or the Kiro powers website. For this walkthrough, we’ll use the IDE.

  1. Launch the Kiro IDE.
  2. Select the Ghosty icon with the lightning bolt to open the powers panel.
  3. Select Dynatrace Observability from the Recommended
  4. Select Install. The power is registered with placeholder values for the Dynatrace URL and token. Therefore, Kiro will show an error message that the MCP server can’t be reached.
  5. To complete the configuration, select Open Settings and replace the placeholders with your environment details.

Configure your tenant and token

In the settings file, replace the two placeholders:

Placeholder Replace with
YOUR_DT_URL https://TENANT_ID.apps.dynatrace.com/platform-reserved/mcp-gateway/v0.1/servers/dynatrace-mcp/mcp. Replace TENANT_ID with your Dynatrace environment ID (visible in your environment URL, for example https://<ENVIRONMENT_ID>.apps.dynatrace.com/ui).
YOUR_BEARER_TOKEN The Dynatrace platform token you created earlier (for example, dt0s16.XXXXX).

Start asking questions

Open a new chat in Kiro and start interacting with your Dynatrace environment using natural language. Query active problems or security vulnerabilities, request a root cause analysis to identify critical issues in production, or pull related logs and traces, all without leaving the IDE.

See it in action

The short demo below walks through installing the Kiro power for Dynatrace, verifying the connection, and running a first query against your environment to list the top 10 vulnerabilities detected by Dynatrace.

Installing and activating the Kiro power for Dynatrace (video)
Figure 2. Installing and activating the Kiro power for Dynatrace (video)

Get started with the Kiro power for Dynatrace

Kiro powers transform what used to be a stitching exercise (MCP servers here, steering files there, custom instructions somewhere else) into one single, ready-to-use bundle. The Kiro power for Dynatrace applies the same idea to observability: live production insight, causal root cause analysis, and remediation grounded in real telemetry, all available the moment a developer needs them.

The result is a tighter loop between writing code and understanding how it behaves in production. Less waiting for diagnostic data from someone else. Less guesswork from an AI assistant operating without context. And, more time spent on the work that actually matters.

Ready to try it? The Kiro Power for Dynatrace is publicly available: install it from kiro.dev or the Kiro IDE and start asking your environment questions.

Using Kiro and the Kiro power for Dynatrace root cause analysis (video)
Figure 3. Using Kiro and the Kiro power for Dynatrace root cause analysis (video)
Experience the Kiro power for Dynatrace for yourself.

The post Dynatrace observability is now a Kiro power appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-observability-is-now-a-kiro-power/feed/ 0
From reactive to proactive: How NAIC embedded AI‑powered observability directly into the IDE https://www.dynatrace.com/news/blog/how-naic-embedded-ai-powered-observability-directly-into-the-ide/ https://www.dynatrace.com/news/blog/how-naic-embedded-ai-powered-observability-directly-into-the-ide/#respond Fri, 12 Jun 2026 17:54:29 +0000 https://www.dynatrace.com/news/?p=74532 Achieving enhanced observability for Alibaba Cloud in multi-cloud environments with Dynatrace

Every developer knows the feeling: You’re in your IDE when something breaks. Error rates spike, alerts fire, and suddenly you’re out of the flow. Michael Kobush, Performance Engineer III at the National Association of Insurance Commissioners (NAIC®), wanted to eliminate the gap between development and runtime. Instead of switching tools or waiting on SRE support, […]

The post From reactive to proactive: How NAIC embedded AI‑powered observability directly into the IDE appeared first on Dynatrace news.

]]>
Achieving enhanced observability for Alibaba Cloud in multi-cloud environments with Dynatrace

Every developer knows the feeling: You’re in your IDE when something breaks. Error rates spike, alerts fire, and suddenly you’re out of the flow. Michael Kobush, Performance Engineer III at the National Association of Insurance Commissioners (NAIC®), wanted to eliminate the gap between development and runtime. Instead of switching tools or waiting on SRE support, NAIC set out to bring production insight directly into the developer workflow.

Let’s take a look at how NAIC embedded real-time observability directly into their development workflow and reduced investigation time to a few minutes.

The problem: Context switching kills developer productivity

Developers lose time the moment they leave their IDE, jumping between views of logs, metrics, and traces simply to understand what has changed.

For NAIC, this friction was slowing down their teams. Developers didn’t have access to production context, creating a dependency on SRE teams whenever investigations were needed. An analysis that should have taken minutes routinely took 45 minutes to an hour. Root-cause identification required manual correlation across multiple systems, a process that was neither scalable nor sustainable.

At Dynatrace Perform 2026, Kobush demonstrated how his team uses Kiro and Dynatrace at NAIC: Real-time observability in your IDE: How NAIC uses Kiro powers to drive developer productivity

The solution: Kiro powers and intelligent observability

Kiro is AWS’s agentic AI-powered IDE that takes a spec-driven approach to software development by turning natural language prompts into structured requirements, architecture designs, and implementation tasks to carry code from prototype to production.

NAIC installed the Dynatrace power for Kiro, one of Kiro’s installable powers that
dynamically connect domain-specific tools and context to the agent. Once connected, Kiro gives developers and AI agents access to Dynatrace data and insights, helping them pinpoint root causes and receive remediation recommendations directly in their workflow.

No switching between tools. No waiting on another team.

The aha moment: Root-cause analysis in minutes, not hours

The first prompt NAIC ran after connecting Kiro to Dynatrace set the tone for everything that followed. Kobush typed a single line into Kiro: “Tell me about problem P-18576.”

Within 30 seconds, Kiro returned a full problem summary with details and recommendations, pulling everything from Dynatrace automatically. Then, he pushed further: “Give me a really deep dive root-cause analysis of what happened.”

In under two minutes, Kiro returned a full root-cause analysis correlating telemetry, infrastructure signals, historical incidents, and the current problem from Dynatrace into a structured response that included:

  • An executive summary
  • Detailed problem context
  • Infrastructure analysis
  • Technical root-cause analysis
  • Remediation strategies
  • Conclusions and next steps

A preliminary assessment that would have previously taken 45 minutes to an hour was now done in minutes. More importantly, it wasn’t just faster; it gave the team a clear, connected view of how services, infrastructure, and dependencies contributed to the issue.

Beyond root cause: Automation across the entire workflow

What makes this more than just a faster diagnostic tool is how NAIC extended Kiro’s capabilities to automate the full incident response workflow.

Using Kiro’s steering files feature, NAIC configured Kiro to automatically generate a structured Markdown file whenever a root-cause analysis was completed. That file includes:

  • Relevant DQL queries used during the investigation
  • Direct links to the Dynatrace dashboards and data sources that surfaced the issue
  • A clear summary of findings

With a Targetprocess MCP also connected, Kiro can take that analysis and populate a ticket directly, automatically loading all relevant context and sending it to the development team. For NAIC, this means the handoff from investigation to remediation is essentially hands-off. This level of automation doesn’t just save time; it creates consistent, repeatable workflows with built-in guardrails. Every incident gets the same structured, data-rich documentation, regardless of who’s investigating it or when.

This isn’t just about faster incident response. It changes how teams build and release software—giving developers immediate feedback on how their changes behave in real environments.

Proactive alerting: Catching problems before they crash

Root-cause analysis after the fact is valuable. With observability embedded directly into the workflow, teams can detect issues earlier in development and respond faster in production, closing the gap between building and operating software.

After noticing that a specific process had crashed, Kobush asked Kiro to set up an alerting profile that would trigger both before the crash, based on stress signals visible in the logs, and at the point of the crash. Kiro analyzed historical log data, identified pre-crash indicators, and built the alert profile automatically.

The result: NAIC’s team now receives early warning signals before a process fails, giving engineers time to intervene rather than react.

This shift from reactive to proactive operations is central to what the Dynatrace and AWS partnership enables. When observability data is embedded in the developer workflow rather than siloed in a separate platform, the entire engineering organization is better equipped to prevent incidents, not just resolve them.

Debugging a sneaky production bug

Perhaps the most telling story from NAIC’s experience with Kiro occurred during a routine error-rate investigation.

An application error rate had increased unexpectedly. Kobush asked Kiro to investigate. Two minutes later, Kiro identified the culprit: A developer had left debug code in the development environment, and it had made its way into production. Every time a user triggered that code path, it threw errors.

When Kobush sent the Markdown report to the developer, the response was immediate: “How did you find that? I’ve been looking for that.”

Kiro leveraged correlated logs, traces, systems context, and historical behavior from Dynatrace to pinpoint exactly where the issue originated.

Start embedding observability into your development workflow

NAIC’s experience highlights a broader shift: When developers, AI assistants, and systems all operate from the same runtime context, debugging becomes faster, releases become safer, and teams spend less time chasing issues and more time building.

The broader message from Kobush is simple: “I’m not a developer. I have a degree in biology and a minor in chemistry… But this, to me, is a game changer in the observability space. I can do things in seconds that would take me hours.”

The productivity gap between observability data and developer action is a solvable problem.

For DevOps engineers, SREs, and platform teams looking to accelerate incident resolution, reduce context switching, and move from reactive troubleshooting to proactive operations, the Dynatrace and Kiro integration offers a practical, immediately actionable path forward.

For developers, this means fewer interruptions, faster answers, and the ability to stay in flow, even when issues arise.

For more information on how Dynatrace and AWS work together, and to access integration best practices, read our guide, Master AI Observability, or come and see us at an AWS Summit near you.

The post From reactive to proactive: How NAIC embedded AI‑powered observability directly into the IDE appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-naic-embedded-ai-powered-observability-directly-into-the-ide/feed/ 0
Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals https://www.dynatrace.com/news/blog/evaluate-llm-and-agent-quality-in-dynatrace-ai-observability/ https://www.dynatrace.com/news/blog/evaluate-llm-and-agent-quality-in-dynatrace-ai-observability/#respond Thu, 11 Jun 2026 19:44:47 +0000 https://www.dynatrace.com/news/?p=74476

AI applications fail in ways that differ from traditional software. They can return responses quickly, with no errors, and still deliver answers that are inaccurate, ungrounded, unsafe, or unusable. That's why AI quality can't be treated as a side project.

The post Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals appeared first on Dynatrace news.

]]>


For AI systems, reliability is defined by response quality, factual grounding, data security, and usability — and those signals need to live alongside the same observability data teams already trust to monitor performance and availability.

When evaluation scores are isolated in notebooks, spreadsheets, standalone tools, or CI logs, they’re hard to operationalize. By bringing AI quality metrics into Dynatrace AI Observability—next to latency, cost, errors, traces, and user behavior—teams can connect poor responses and hallucinations directly to the prompts, models, retrieval contexts, tool calls, services, and traces that produced them.

What is dt-evals?

dt-evals is an open source CLI for evaluating LLM and agent quality from real GenAI traces, agentic interactions. Teams can run online evaluations against live or recent interactions, score outputs with an LLM judge, and send structured results back to Dynatrace AI Observability so quality becomes visible, queryable, trendable, and actionable.

A minor prompt edit, model change, or retrieval update to an AI application can improve one behavior while quietly breaking another. The challenge to tracking down where and why these systems break is that evaluation results are often maintained outside the operational workflow, making it difficult to connect a low score to the exact trace, prompt, model version, retrieval context, tool call, or service that produced the unwanted behavior.

Dynatrace AI Observability closes this loop. With dt-evals and the AI Observability Evaluation Preview teams can pull recent gen_ai.*  spans, score real interactions with an LLM judge, and write structured evaluation results back as business events. These scores can be viewed with the originating trace, queried for custom analysis, trended in dashboards, and used to trigger alerts or workflow-driven remediation.

A failing faithfulness score is no longer just a number in a report. It’s now an operational signal.

What are LLM evaluations?

An LLM evaluation system scores an AI response against a range of quality and safety dimensions. Common examples include whether the answer is relevant to the question, faithful to the provided context, free of hallucinations, safe for users, complete enough to be useful, and resistant to prompt-injection attempts.

LLM evaluations are typically applied in two modes:

Offline evaluations run before release against a fixed test set or curated trace dataset. These are used to compare a proposed prompt, model, retriever, or agent-tool change against a known baseline before shipping. For example, replay 500 representative support questions in CI and block the release if faithfulness drops below the configured threshold.

Online evaluations run after deployment against sampled production or user traffic. Use online evaluations to detect regressions caused by live inputs, changing retrieval results, tool behavior, traffic mix, or model drift. For example, evaluate 10% of support-agent traces from the last hour and alert the team if hallucination failures exceed the configured window.

With dt-evals, you can run evaluations from the command line, use them in CI/CD, or schedule them to detect quality regressions autonomously after deployment as a post-processing quality gate for your AI agents and LLM output.

AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability
Figure 1. AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability

Run evaluations from the command line

dt-evals is an open source evaluation toolkit for teams that want to bring their own data, judge provider, and evaluation logic while keeping traces, scores, dashboards, and alerts connected.

Install the CLI:

npm install -g @dynatrace-oss/dt-evals

Or run it directly with npx:

npx @dynatrace-oss/dt-evals <command>

A typical first run has three steps:

  1. Configure your environment and judge provider (Bring Your Own AI API key):

dt-evals configure

  1. Verify your local setup and connection:

dt-evals doctor

  1. Run evaluations on recent GenAI traces:

dt-evals run --since 1h --sample 10

This command evaluates traces from the last hour and samples 10% of them. In other words, dt-evals evaluates roughly one out of every ten matching traces, including the prompt and completion messages associated with each selected trace.

During configuration, you provide the connection to your Dynatrace environment and the credentials for the LLM judge provider you want to use. dt-evals does not require teams to send evaluations through a fixed provider. You bring your own judge credentials and control where evaluation execution happens.

For CI/CD use cases, run in CI mode:

dt-evals run --since 6h –ci

In CI mode, dt-evals emits machine-readable output and can fail the pipeline when a configured threshold is breached. This makes quality checks part of the same delivery process used for prompt changes, model upgrades, retrieval updates, and agent releases.

dt-evals in action
Video 1. dt-evals in action

Bring your own LLM judge provider

The “LLM-as-judge” evaluation approach involves using an AI model to score another model or agent response. The LLM judge needs to come from an AI provider your team trusts and has approved for the type of data being evaluated.

dt-evals supports common LLM judge AI models and inference providers, including OpenAI, Anthropic, Google/Vertex/Gemini, AWS Bedrock, and Azure OpenAI. Depending on the package and configuration path you use. Teams provide their own credentials, choose the judge model, and can tune execution settings such as thresholds and concurrency.

This matters for both governance and cost control. Teams can decide which LLM provider is assigned to evaluate which traffic, how many judge calls run in parallel, and where evaluation results are stored.

Which quality and safety dimensions are evaluated by dt-evals?

dt-evals supports built-in LLM-as-Judge evaluators for a range of quality and safety dimensions, including:

  • Relevance: Does the response answer the user’s question?
  • Faithfulness: Is the response supported by the provided context?
  • Hallucination: Does the response invent facts that are not present in the available context?
  • Answer completeness: Does the response fully address the user’s request?
  • Context relevance: Is the retrieved or supplied context useful for answering the question?
  • Factual accuracy: Does the response match an expected or known-correct answer?
  • Summarization quality: Does the summary preserve the important information?
  • Conciseness: Is the response direct with no unnecessary detail?
  • Fluency: Is the response clear and readable?
  • Toxicity: Does the response contain harmful or abusive content?
  • Bias: Does the response show unfair or inappropriate bias?
  • PII leakage: Does the response expose sensitive personal information?
  • Prompt injection: Did the input or response show signs of instruction manipulation?
  • User frustration: Does the interaction suggest the user is blocked or dissatisfied?
  • Drift: Are scores changing meaningfully compared with prior behavior?

A Retrieval Augmented Generation (RAG) application might focus on faithfulness, hallucination, context relevance, and answer completeness. A customer-facing support agent might focus on relevance, fluency, bias, toxicity, and prompt-injection risk. An internal assistant might add custom checks for tone, policy compliance, or whether the answer includes required next steps.

Add custom evaluations

Built-in metrics are useful, but most production AI systems also need checks that are specific to the business, domain, or workflow.

Custom evaluations let teams define their own judge prompts, scoring rules, labels, and thresholds. For example, a support team can create a custom evaluator that checks whether an answer includes a required troubleshooting step before recommending escalation. A financial services team can check whether responses include the required disclaimers. A platform team can check whether an agent uses the correct tool before answering.

The critical point is that custom evaluators run through the same pipeline as built-in evaluators. They can produce the same structured results, appear alongside other scores, and be used in dashboards, alerts, and release checks.

A typical configuration defines the target service, judge provider, sampling strategy, enabled metrics, and thresholds:

schemaVersion: 1
name: support-agent-prod

dynatrace:
  environmentUrl: https://your-env.apps.dynatrace.com
  platformToken: dt0s16.xxxxx

judge:
  provider: openai
  model: gpt-5.5

scope:
  service: support-agent
  since: 1h
  sampling:
    strategy: random
    percent: 10

metrics:
  enabled:
    - faithfulness
    - hallucination
    - relevance
    - drift

alerts:
  thresholds:
    faithfulness: 0.7
    relevance: 0.7

Evaluation results in the AI Observability app

Evaluation results appear directly in the AI Observability app, so teams don’t have to jump between a trace view, an eval report, and a separate dashboard to understand what happened.

In the Prompts view, teams can filter for prompts with evaluation scores and inspect row-level verdicts. Score badges such as relevance, fluency, bias, faithfulness, or toxicity make response quality easy to scan without opening every trace.

This is useful when triaging a regression. Instead of starting with a generic failure count, teams can quickly see which prompts failed, which evaluator failed them, and whether the issue is isolated or widespread.

AI Observability App Prompts stream with evaluation results
Figure 2: AI Observability App Prompts stream with evaluation results

From an individual prompt or trace, the Evaluations tab shows run-level details, including the evaluation name, score, provider, judge model, method, and supporting metadata.

That detail matters because a failed score is only useful if teams can explain it. Engineers and evaluation owners can move from a low score to the exact prompt, response, trace, model, evaluator, and rationale that produced it.

Prompt detail view with evaluation results and trace context in Dynatrace AI Observability
Figure 3:  Prompt detail view with evaluation results and trace context in Dynatrace AI Observability

Query, trend, and alert on evaluation scores

Because dt-evals writes results back as structured events, evaluation scores can be analyzed with the rest of your telemetry.

Teams can ask questions such as:

  • Which evaluator has the lowest average score?
  • Which services are producing the most failed evaluations?
  • Did quality drop after a model or prompt change?
  • Are hallucinations increasing over time?
  • Is quality improving at the cost of latency or token usage?

For example, to get average score by evaluator you could write this query:

fetch bizevents
| filter event.type == "gen_ai.evaluation.result"
| summarize avg_score = avg(gen_ai.evaluation.score.value),
    by: { gen_ai.evaluation.name }
| sort avg_score asc 
Querying failed evaluations by service and evaluator in Dynatrace AI Observability
Figure 4: Querying failed evaluations by service and evaluator in Dynatrace AI Observability

Failed evaluations by service and metric can be determined with this query:

fetch bizevents
| filter event.type == "gen_ai.evaluation.result"
| filter gen_ai.evaluation.score.label == "fail"
| summarize failures = count(),
    by: { dt.service.name, gen_ai.evaluation.name }
| sort failures desc 
Average evaluation scores by evaluator in Dynatrace AI Observability
Figure 5: Average evaluation scores by evaluator in Dynatrace AI Observability

Trending is where evaluation data becomes more useful than a point-in-time report. A single failed score can show an issue. A trend can show whether quality is drifting slowly, whether a release caused a sudden drop, or whether a fix actually improved behavior over time.

On the AI Evaluation & LLM App Performance dashboard, teams can track trends in quality score, pass rate, failed evaluations, drift detections, evaluator health, run cadence, and pass/fail volume over time.

AI Evaluation &amp; LLM App Performance dashboard
Video 2: AI Evaluation & LLM App Performance dashboard

How to turn quality regressions into alerts

Evaluation results can also drive alerts. For example, a support agent team may want to notify the AI team when hallucinations appear in production, or when faithfulness drops for more than a few minutes.

name: support-agent-prod

alerts:
  notifications:
    - name: hallucination-detected
      metric: hallucination
      condition: count > 0
      window: 5m
      channel:
        type: slack
        connection: ai-observability-slack
        channel: "#ai-alerts"

    - name: faithfulness-regression
      metric: faithfulness
      condition: fail_rate > 10%
      window: 15m
      channel:
        type: email
        connection: ai-team-email
        to: [ai-team@example.com] 

Deploy the alerts with:

dt-evals alerts list ./support-agent-prod.yaml
dt-evals alerts apply ./support-agent-prod.yaml

Once applied, Dynatrace runs these checks continuously as Workflows. If hallucinations appear in the last five minutes, the team gets a Slack alert. If more than 10% of faithfulness checks fail over 15 minutes, the AI team receives an email. This turns LLM quality from something teams inspect manually into something Dynatrace can monitor and route automatically.

For continuous alerting, evaluation runs need to happen continuously or on a schedule. You can run dt-evals in CI for release checks (see our example here), run it manually during investigation, or deploy a scheduled runner for ongoing production evaluation. An alert is only as fresh as the evaluation results it carries.

Once configured, quality signals can be routed to the teams that need to act. If hallucinations appear in the last five minutes, the team can receive a Slack alert. If more than 10% of faithfulness checks fail over 15 minutes, the AI team can receive an email. This turns LLM quality from something teams inspect manually into something they can monitor and route automatically.

Close the loop in the AI software delivery lifecycle

Evaluation gates are most useful when they meet developers where they already work. Because dt-evals writes evaluation results back into the observability data layer, those results are not limited to dashboards or post-release reviews. They can be queried, inspected, and acted on from development workflows, CI/CD pipelines, and agentic coding environments.

For example, a team can run dt-evals after any change to a prompt, model, retriever, or agent tool, and then use dtctl (Dynatrace CLI tool for AI Agents) to query the resulting evaluation data, inspect related traces, review dashboards, or validate whether a release threshold was met. In an AI-assisted workflow, tools such as Claude Code, Cursor, GitHub Copilot, or an internal agent harness can leverage MCP or CLI access to bring that same observability context into the developer’s daily workflow.

That closes the loop of the AI software delivery lifecycle: teams can evaluate behavior, control rollout decisions, remediate regressions, and feed production learning back into the next development cycle. Quality signals are no longer in a separate report; they’ve become a part of how AI software is built, shipped, and operated.

Bring evaluations into the release process

Evaluation support is not just for inspection after something breaks. It can also help prevent regressions before they reach users.

Overview of a typical release workflow for an AI app with dt-evals
Figure 6: Overview of a typical release workflow for an AI app with dt-evals

This makes AI quality part of the release process. Teams can gate changes based on relevance, faithfulness, hallucination risk, prompt-injection risk, toxicity, or custom metrics, rather than relying solely on latency and error rate.

What makes this meaningful is that it’s the same pipeline teams already run. AI quality gates sit alongside the latency, error rate, and SLO gates teams have been using for years. There’s no second CI system, no second platform to learn, no second dashboard to monitor. Quality becomes one more dimension of the release decision, gated the same way performance is gated, by the same platform, in the same pipeline.

Coming next

The current experience makes evaluation results visible and actionable inside Dynatrace AI Observability. Next, the focus is on making evaluation workflows easier to run at scale and easier to compare across changes.

Planned improvements include targeted and bulk trace evaluations, custom evaluation libraries, evaluator versioning and lineage, baseline comparisons, experiment views, native quality gates, and deeper visibility into online evaluations.

These capabilities will help teams compare prompt and model variants, understand quality versus cost and latency tradeoffs, and detect sustained quality regressions before they affect more users.

Start today

To get started, check out the Git repository. You’ll need:

  • Node.js 20 or later
  • A Dynatrace environment with GenAI spans and the AI Observability app installed
  • Credentials for the judge provider you want to use
  • A service, trace sample, or CI workflow you want to evaluate

Then, install the CLI:

npm install -g @dynatrace-oss/dt-evals

Configure your service and judge provider:

dt-evals configure

Run your first evaluation:

dt-evals run --since 1h --sample 10

With Dynatrace AI Observability and dt-evals, teams can bring LLM and agent evaluations into the operational loop, where they can trace behavior, score outputs, trend results, alert on regressions, and gate releases before silent failures reach production.

The post Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/evaluate-llm-and-agent-quality-in-dynatrace-ai-observability/feed/ 0
Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/ https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/#respond Mon, 08 Jun 2026 15:03:09 +0000 https://www.dynatrace.com/news/?p=74422 Blog OTP Observability for Agentic AI

Agentic AI is breaking the mold of what organizations need from observability. Fragmented, correlation-dependent observability platforms are no longer “good enough.” Enterprises with dynamic, hybrid environments require observability that provides real-time, precise answers, so AI agents can prevent problems, automate workflows, and deliver better, more secure software.

The post Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI appeared first on Dynatrace news.

]]>
Blog OTP Observability for Agentic AI

As more agentic AI projects come online, the observability market is abuzz with familiar promises: tool consolidation, AI-powered insights, and faster remediation through smarter tools. On the surface, this sounds like progress. But beneath the excitement, many discussions are framed around the wrong question.

The real issue isn’t about how to adopt autonomous operations; it’s about ensuring AI agents are operating reliably and resolving problems without introducing new ones. When evaluating new observability solutions, the question should be:

Can this observability solution accurately analyze complex, dynamic telemetry in context so AI agents can act autonomously with trust, precision, and reliability?

As systems become increasingly agent driven, observability is crossing a structural boundary. Approaches designed for environments where only humans decide and act must adapt to a world where agents increasingly operate autonomously with human oversight, while keeping organizations informed.

Rethinking observability for the agentic age

Observability platforms were initially intended to support engineers in delivering reliable applications, services, and infrastructure to users, and alert them in the event of a problem. Dashboards, alerts, and correlation helped teams investigate incidents, piece together what happened, diagnose issues, decide on next steps, and resolve the problem. This model worked when changes were pushed manually.

The assumption was that more data, better correlation, and cleaner interfaces will lead to increased visibility and improved operational decision making.

Agentic AI systems break that assumption.

With faster release cycles and AI-generated code, manual investigations can no longer keep pace. Moreover, observability platforms must now provide actionable insights to both humans and AI agents.

As agents begin operating as autonomous participants in software environments by triggering mitigations, scaling infrastructure, and optimizing behavior in real time, observability can no longer function solely as a human interface. It must also provide AI systems with a reliable, contextual fact basis that agents can act on programmatically. Machines can’t rely on dashboards and alerts. They require a deterministic foundation of unified, real-time data that delivers accurate, context-rich answers at exabyte scale.

Agentic systems break the mold of “good enough”

Many observability platforms layer probabilistic AI on top of siloed data. They use LLMs to correlate signals and rank likely causes—but they can’t always determine correctness.

“Probabilistic” means that the same input will generate a different output based on a probability distribution of predefined outputs, delivering a different answer when the same problem occurs. This approach is also prone to hallucinations, requiring additional human validation, which can increase operational overhead and token costs, delay resolution of business-critical issues, and divert resources from strategic initiatives.

Enterprise-grade observability must now answer: Is this insight reliable enough for autonomous action?

AI built on siloed data is inherently unreliable. Autonomous systems depend on deterministic, contextual, and trustworthy data to act reliably.

“Deterministic” means that the same input always results in the same output by using factual data to trace the exact causal changes that created the issue. When agentic AI systems act on business-critical applications, the cost of being “mostly right” becomes operationally unacceptable.

This is where a subtle but critical divide appears in the market. Aggregating signals and correlating anomalies can surface patterns. Patterns alone are not a solid basis for decisions, and without deterministic understanding, AI systems inherit that uncertainty and can propagate it downstream.

To drive reliable enterprise autonomous operations, AI agents require a unified, AI-powered observability platform that can analyze exabytes of data in real time and across models to pinpoint root cause, delivering actionable answers in context of what’s affected and its business impact.

From correlated guesses to deterministic answers

This shift in the demands of observability hinges on a clear distinction:

  • Probabilistic AI correlates signals that happened around the same time and therefore appear related, pulling information from fragmented data stores to propose a likely root cause.
  • Deterministic AI uses causal analysis to pinpoint what happened and why, recommend remediation actions, and identify business impact.

Probabilistic AI is intended to narrow the search space and direct engineers toward potential resolution, but it still requires interpretation.

Deterministic AI establishes sequence, dependency, and impact, enabling systems to decide safely without waiting for humans to connect the dots.

Auto‑remediation, auto-prevention, and auto-optimization all depend on this leap. A platform that unifies telemetry only at the UI layer may deliver data and potential root cause, but it can’t compensate for fragmented understanding and missing context underneath. When context is pieced together after the fact, confidence is never guaranteed.

You can’t automate what you don’t precisely understand.

Context driven observability as the control plane for AI

In an autonomous enterprise, observability doesn’t sit beside execution; it’s embedded within it. This integration requires that teams adopt a new mindset toward observability architecture.

Because more AI workloads are happening at the source, telemetry must be optimized and streamlined before ingest, not after the fact, from the edge to the back end. Data access must be unified, context-aware, and always-hydrated on a massive scale. Answers must be explicit, not implicit, and they must be informed by automatic, real-time dependency mapping.

Likewise, intelligence must combine deterministic and agentic AI—not as add‑ons, but as a single reasoning system from ingest to execution.

In this model:

  • AI agents can become the primary consumers of observability data.
  • Humans can shift toward strategy, architecture, oversight, and exception handling.
  • Observability evolves from a reactive lens into a control plane for autonomous operations.

Observability purpose-built for autonomous operations ensures successful agentic AI initiatives

This moment represents an architectural transition, not just an incremental upgrade cycle. Correlation-dependent observability that uses probabilistic AI can be extended, augmented, and rebranded, but it will always carry the limitations of approximation and human validation.

The next era belongs to an observability platform that’s built for machine understanding from the start: a unified, context driven architecture that delivers deterministic answers at machine speed, precision, and scale.

Do you want more data or better decisions? Learn why enterprises are switching to Dynatrace.

The post Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/feed/ 0
Port and Dynatrace: One-prompt incident triage with the Dynatrace MCP Server https://www.dynatrace.com/news/blog/port-and-dynatrace-one-prompt-incident-triage/ https://www.dynatrace.com/news/blog/port-and-dynatrace-one-prompt-incident-triage/#respond Fri, 05 Jun 2026 12:28:09 +0000 https://www.dynatrace.com/news/?p=74412 Agent graphic

The Dynatrace MCP Server is now available in Port via Port MCP Connectors. A single OAuth flow connects it in minutes. Set up Port AI to communicate with Dynatrace, GitHub, Slack, and your service catalog in a single conversation, correlating production signals, code, and ownership in a single agent run rather than three browser tabs […]

The post Port and Dynatrace: One-prompt incident triage with the Dynatrace MCP Server appeared first on Dynatrace news.

]]>
Agent graphic

The Dynatrace MCP Server is now available in Port via Port MCP Connectors. A single OAuth flow connects it in minutes. Set up Port AI to communicate with Dynatrace, GitHub, Slack, and your service catalog in a single conversation, correlating production signals, code, and ownership in a single agent run rather than three browser tabs and a manual handoff. For incident triage, this turns a multi-tool investigation into a single prompt: just ask Port AI what’s wrong; you’ll get details about the failing service, the error signature, the file and function, and the suspect commit.

Ground every Port AI conversation in live production data from Dynatrace

Port is an agentic engineering platform that platform teams use to organize their software development lifecycle. It gives software and DevOps teams a central place where engineers can find system information, take action on it, and route work across their connected tools, without waiting on IT or operations.

Port AI is the assistant that queries the catalog using natural language. Through Port MCP Connectors, the same chat also reaches external systems, such as Dynatrace, GitHub, and Slack.

Dynatrace complements Port by providing context-rich observability and security insights, right where you need them:

What Dynatrace brings What Port brings
Live observability signal across logs, traces, and metrics Service ownership and team responsibility
Dependency topology between affected services On-call rotation and escalation paths
Open problems with root cause already identified Recent deploys and commit history per service
Security vulnerabilities and exposures detected in running services, with severity and affected entities Remediation ownership and the team that’s accountable for the fix

The result: Team members ask a question in the Port AI chat, where they’re already working, and get back complete answers that no single tool could produce on its own.

Complete triage run with Port AI [VIDEO]
Figure 1. Complete triage run with Port AI [VIDEO]

Incident triage from a single Port AI prompt

The Dynatrace MCP Server provides Port AI with a set of tools it can call during any conversation. For incident triage, the most relevant needs are:

  • Query production data. Logs, traces, metrics, and events from across the Dynatrace tenant returned in structured form.
  • List open problems. An overview of all active problems on the tenant.
  • Get problem details. Root cause, causal chain, and affected entities for a specific problem.
  • Get troubleshooting guidance. Relevant troubleshooting guides matched to a problem description.

The full toolset also covers security findings, entity and topology lookup, query generation, forecasting, and more. See the Dynatrace Hub for the complete list. In combination with Port AI, these capabilities turn the Port AI chat into a single place to ask production questions.

The example below walks through the triage of an incident affecting broker_service, a fictional service. The same investigation, done manually without this integration, starts in Dynatrace (where the failing service shows up in seconds), then jumps to GitHub to scan recent commits, to Slack to confirm ownership, and back to a doc to write up the summary. With this Port AI integration, those steps run in a single agent run, with the Dynatrace signal at the center of the chain.

The scenario begins when broker_service degrades and lands as a new incident, INC-1003. In Port AI chat, an SRE asks, “Help me understand the root cause of INC-1003.”

Port correlates signals from Dynatrace, GitHub, and Slack in a single agent run.
Figure 2. Port correlates signals from Dynatrace, GitHub, and Slack in a single agent run.

Port AI loads the ai-incident-triage skill and runs the following steps:

  1. Resolve the incident in Port’s catalog. Port AI looks up the incident entity and pulls the affected service identifier.
  2. Query Dynatrace. The Dynatrace MCP Server queries logs and traces for broker_service. The response carries the first-seen failure timestamp and the top error signature.
  3. Find the suspect commit in GitHub. Port AI passes the failure window to the GitHub MCP Server, locates the failing function, and lists the commits to that file. One commit aligns with the first-seen timestamp.
  4. Return a structured triage summary. Port AI returns the failing service, the error signature, the file and function, and the suspect commit.
  5. Post to Slack and close the loop. The Slack MCP Server posts the same summary to #incident-updates (Figure 2). The incident entity records a triaged_at timestamp through a Port self-service action.
The triage summary is posted to the Slack #incident-updates channel.
Figure 3. The triage summary is posted to the Slack #incident-updates channel.

Within a single agent run, the engineer receives a triage summary in Slack that already includes the live Dynatrace signal.

The same Dynatrace integration allows many more use cases, such as deployment correlation, on-call summaries, and postmortem drafts, each built as a Port AI skill.

Security triage works just as easily: ask Port AI about a vulnerability; the Dynatrace MCP Server returns the affected running services, severity, and exposed entities, while Port resolves ownership and routes the fix to the accountable team.

Roll it out across teams, govern centrally

The integration is designed for organization-wide rollout. Admins maintain central control over which tools are exposed and who can access them, while each query remains scoped to the user’s existing permissions.

  • Per-tool selection. Admins choose which Dynatrace tools Port AI can call across the organization. Sensitive tools can be scoped to selected groups.
  • Per-user authentication. Each user authenticates to Dynatrace through OAuth. Queries return only the data that their existing Dynatrace permissions already allow.
  • Audit trail on both sides. Every Port AI call and every Dynatrace MCP call names the same person, with no stitching required between platforms.
  • Scales without new workflows. The same per-user model that works for a pilot team works for hundreds of developers. No separate access-request workflow needed.

Get started: connect Port with Dynatrace

The Dynatrace MCP Server connects to Port through a single OAuth flow. Setup takes a few minutes and is done once by an admin.

  • Add Dynatrace as a data source. In Port, go to Data Sources > + Data source > MCP Servers. Select Dynatrace, fill in the connector details, and select Connect to authenticate with your Dynatrace tenant.
  • Expose the tools you want Port AI to use. Under Allowed Tools, add the Dynatrace capabilities you want available to your organization (querying production data, listing problems, getting problem details, finding troubleshooting guidance). Select Publish.
  • Register the incident triage skill. Add the ai-incident-triage skill from Port’s skill library and point it at the incident entities in your catalog.

Once published, the integration is available to every authenticated Port user in your organization. Each user authenticates to Dynatrace individually through OAuth on first use.

For full setup walkthroughs, see Port documentation

Port MCP Connectors documentation: connector setup and admin configuration

Triage incidents with AI: full skill walkthrough

The post Port and Dynatrace: One-prompt incident triage with the Dynatrace MCP Server appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/port-and-dynatrace-one-prompt-incident-triage/feed/ 0
AI agents are redefining software development—but they’re flying blind without observability https://www.dynatrace.com/news/blog/ai-agents-are-redefining-software-development-but-theyre-flying-blind-without-observability/ https://www.dynatrace.com/news/blog/ai-agents-are-redefining-software-development-but-theyre-flying-blind-without-observability/#respond Thu, 28 May 2026 17:09:42 +0000 https://www.dynatrace.com/news/?p=74210 AI agents are redefining software development

Imagine a team of AI agents building, deploying, and running software at machine speed—yet unable to see what’s happening in production. This is the new reality for enterprise technology leaders. As one Fortune 500 CTO told us, “Speed is now the primary driver of innovation, forcing organizations to rethink processes, compliance, and roles; it’s a […]

The post AI agents are redefining software development—but they’re flying blind without observability appeared first on Dynatrace news.

]]>
AI agents are redefining software development

Imagine a team of AI agents building, deploying, and running software at machine speed—yet unable to see what’s happening in production. This is the new reality for enterprise technology leaders. As one Fortune 500 CTO told us, “Speed is now the primary driver of innovation, forcing organizations to rethink processes, compliance, and roles; it’s a necessity for innovation teams.”

Observability—real-time visibility into how software behaves in production—has become the critical enabler for both human-led and agent-led teams. Without it, AI agents are powerful but blind.


Key executive insights

  1. Software production is being redefined by AI agents. This transformation is a structural shift, not a trend.
  2. The world is bimodal again. Human-led and agent-led environments coexist.
  3. AI agents are powerful but blind. Without rich context from production, they cannot deliver reliably.
  4. Observability is a crucial enabler to gradually transform from human-led to agent-led operations. Observability is what allows organizations to industrialize software delivery with confidence.
  5. The new KPI for agent-led teams is the percentage of human intervention required. The lower the number, the better the AI is working.

The market reality: A bimodal world

Organizations are accelerating AI adoption not because it is trendy, but because it is existential. Companies that fail to transform risk being outpaced by competitors that can deliver software faster, cheaper, and at higher quality. CTOs and CIOs are making statements like “speed over compliance” not out of recklessness, but because they recognize that without radical acceleration, their businesses face disruption.

At the frontier of this shift is a fundamentally new way of building software: AI-first development. In these environments, 100% of coding, testing, deployment, operations, bug fixing, and optimization are performed by AI agents. The human role shifts to specification, goal setting, supervision, and correction. Intellectual property moves from the code to the specification—code becomes a generated artifact, not the source of truth. With a complete, well-architected spec, agents can fully rebuild the software from it again.

This creates a bimodal operating environment:

  • Human-led teams—the majority today—are existing operations, SREs, and developers augmenting their workflows with AI. They follow the traditional SDLC, increasingly supported by AI agents that auto-prevent, auto-remediate, and auto-optimize, which reduces manual effort and achieves more with the same resources.
  • Agent-led teams—growing fast—are innovation groups operating in full AI development life cycle (AIDLC) mode. Swarms of AI agents build, deploy, and run software end-to-end. Humans write specifications and intent, not code. For these teams, the KPI is no longer “how many story points were solved?” but “what percentage of human intervention is required?”

Observability enables a reliable transition to autonomous operations

In the early 2010s, a similar bimodal pattern emerged with cloud: one team running thousands of servers on-premises, another in stealth mode on AWS. The pattern is repeating now with AI.

Why not switch everything to agent-led right away? Because existing systems follow processes, compliance, and technology stacks that can’t be immediately automated in an AI-first way. Moreover, it’s too risky to move all business-critical systems simultaneously. The safer path: start with an innovation team, build less critical applications first, and only when those are successful and trusted, begin migrating more of the business-critical services.

New foundation models that arrived in early 2026 have accelerated the path to fully autonomous operations, making agent-led teams realistic at small scale today, with large scale within sight. These systems focus on AI-first software generation first, with a clear goal to eventually master operational challenges (resilience, performance, scale, security) entirely with agents as well.

Observability plays a critical role not only in making both modes work reliably, but also in enabling the transformation from the first mode to the second. The context observability provides—understanding existing system behavior, dependencies, and requirements—is exactly what agents need to create the reliable and scalable software. Observability is what makes both modes work, and it is the critical bridge between them.

The core problem: AI agents are blind

AI agents can code, deploy, refactor, and operate software faster than humans ever could. But there is one thing AI cannot do without help: AI has no awareness of what happens in production. It’s blind to the real world: without real-time feedback from running software—in development and production —agents make decisions without context and without understanding their consequences. They operate at speed, but without sight.

77% of IT teams still lack full visibility across hybrid environments (IBM Institute for Business Value, 2025). If you can’t see it, you can’t scale it. Observability is not optional for AI-first operations, it’s a prerequisite.

The Dynatrace response: Real-time observability for both worlds

Dynatrace addresses both sides of this bimodal reality: a complementary response to the two speeds at which enterprises now operate.

For human-led teams: Autonomous operations at scale

This year, Dynatrace launched Dynatrace Intelligence: a full agentic operations system that orchestrates dozens of agents that auto-prevent, auto-remediate, and auto-optimize across site reliability, development, and application security. These AI agents deliver the following value in production:

  • SRE Agent: Kubernetes troubleshooting, infrastructure optimization, and automated incident resolution – reducing mean time to resolution at scale.
  • Developer Agent: Surfaces production context during deployment, validates changes, and prevents issues before they reach customers.
  • Security Agent: Identifies vulnerabilities, triages threats, and accelerates security response, all in real time.

The deterministic foundation underneath: what separates Dynatrace agents from others is its deterministic foundation: real-time, full-stack, and cross-model root-cause analysis, anomaly detection, and forecasting, all grounded by data in a unified, purpose-built data lakehouse that delivers accurate, contextual answers from exabytes of information. This is not AI that guesses; it’s AI that reasons from facts. Benchmarks from internal testing and observed customer use cases: 12× higher success rate in SRE use cases, 3× faster problem resolution, 2.5× lower token cost.

Ecosystem integrations that extend intelligence beyond the platform: Dynatrace Intelligence extends into third-party tools to drive autonomous actions across development, SRE, and ITOps workflows.

For agent-led teams: develop and run software reliably

Dynatrace enables AI-first teams to let swarms of agents to build and run software reliably, providing real-world awareness from observability, run-time context across development, security, and operations, and self-optimization toward SLAs, cost, and resilience. Key capabilities include:

  • Agentic observability: Closed-loop autonomous operations where observability agents coordinate with coding and deployment agents to self-heal.
  • AI and cloud observability: Full-stack visibility across cloud infrastructure and AI workloads, covering resilience, performance, security, user experience, and LLM evaluations to assess the quality and reliability of agent outputs, helping identify potential inaccuracies, hallucinations, or risks.
  • AI data lakehouse (Grail): Real-time context engine that provides long-term memory for agent decisions—sub-second, API-native, at an exabyte scale.

The goal: a closed loop where agents detect issues, resolve them, and ship the fix—autonomously, 24/7.

Dynatrace is on the same bimodal journey – our entire business runs on Dynatrace Intelligence in human-led mode, with agents taking over more tasks continuously, while our AI-first offering and new services are built and operated entirely by agent swarms, using our own observability to close the feedback loop.

Different approaches – unified platform

Across the platform, Dynatrace delivers end-to-end, full-stack visibility across cloud infrastructure, applications, and AI workloads, including agent behavior, decision paths, and cost, along with governance at machine scale. These capabilities serve human-led and agent-led teams differently, but from the same unified platform.

The measure of success in software delivery is shifting from human productivity metrics to a new KPI: the percentage of human intervention required. Observability is what makes that progress possible. The question for every technology leader is no longer whether to adopt AI-first, but how quickly they can close the visibility gap before competitors do to drive massive growth in innovation and productivity.

The post AI agents are redefining software development—but they’re flying blind without observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ai-agents-are-redefining-software-development-but-theyre-flying-blind-without-observability/feed/ 0
Dynatrace MCP Server for Atlassian Rovo: Investigate production problems without leaving Jira or JSM https://www.dynatrace.com/news/blog/dynatrace-mcp-server-for-atlassian-rovo-investigate-production-problems-without-leaving-jira-or-jsm/ https://www.dynatrace.com/news/blog/dynatrace-mcp-server-for-atlassian-rovo-investigate-production-problems-without-leaving-jira-or-jsm/#respond Wed, 27 May 2026 19:37:02 +0000 https://www.dynatrace.com/news/?p=74203 Atlassian and Dynatrace

A developer triaging a bug in Jira, or an on-call engineer responding to an alert in Jira Service Management (JSM), can ask their Rovo agent to investigate production issues in Dynatrace using plain language, without leaving Jira or JSM. The Rovo agent answers any relevant questions, calls the appropriate Dynatrace tools, and posts the results […]

The post Dynatrace MCP Server for Atlassian Rovo: Investigate production problems without leaving Jira or JSM appeared first on Dynatrace news.

]]>
Atlassian and Dynatrace

A developer triaging a bug in Jira, or an on-call engineer responding to an alert in Jira Service Management (JSM), can ask their Rovo agent to investigate production issues in Dynatrace using plain language, without leaving Jira or JSM. The Rovo agent answers any relevant questions, calls the appropriate Dynatrace tools, and posts the results back to the Jira workspace. Admins can now complete setup in minutes. Every call runs with the permissions of the requesting user, every action is logged, and data access and cost remain under central control.

Move from “ask the platform team” to “ask the agent”

If your organization uses Dynatrace, much of your production knowledge may reside with the Dynatrace experts on your organization’s platform or SRE team. Everyone else (developers triaging Jira bugs, on-call responders running down JSM alerts, or support engineers handling escalations) either learns enough Dynatrace to investigate production issues on their own or pings the platform team and waits for a response.

The Dynatrace Model Context Protocol Server for Rovo moves this valuable production expertise into the Rovo agent, where it’s accessible to all Rovo users. The Rovo agent holds the Atlassian context (the bug, the alert, the service it relates to) and calls Dynatrace for the production context (problems, topology, root cause). Meanwhile, the developer asking the question gets a usable answer directly in the tool where they’re already working.

Setup completes in minutes. The integration is usable on the first prompt.
Setup completes in minutes. The integration is usable on the first prompt.

How does an on-call engineer use Rovo to triage a JSM alert?

A JSM alert is triggered when a service degradation is detected. The on-call engineer opens the alert’s response panel and asks the Rovo Ops Agent, Atlassian’s built-in AI agent for JSM incident response, to investigate.

Rovo Ops calls the Dynatrace MCP Server, retrieves the open problem, the root cause, and the related signals, and returns a single answer: what went wrong, what’s related, and what to do next. The on-call engineer either acts on this problem context directly or escalates the alert, with the full analysis already attached.

How does a developer triage a Jira bug with Rovo and Dynatrace?

Let’s say Jira provides details of a bug in a service that a developer doesn’t own. The developer asks the Rovo agent a question in the issue’s chat panel: “What’s going on with this service right now?” Rovo returns the open Dynatrace problem, the elevated error rate, and the dependency that’s causing the issue.

Rovo’s answer posts as a comment on the bug in Rovo, so the next person who opens the ticket sees the investigation is already done. From there, the bug typically closes as a duplicate of the active incident or routes to the team that owns the failing deployment, in minutes rather than hours.

How to roll out the Dynatrace MCP Server in Atlassian Rovo

MCP Server setup is a configuration task, not a project. The Dynatrace MCP Server is pre-approved by Atlassian as an external MCP integration and is included with Dynatrace SaaS at no extra cost. In a few clicks, an admin connects Dynatrace using the Rovo admin UI, authenticates against the Dynatrace tenant, and selects which tools to expose. Once connected, the integration is available across Jira, Jira Service Management, and Confluence.

The Dynatrace MCP Server tool set is ready for immediate use, providing data retrieval through Grail, topology and entity context, root cause analysis, active security findings, time-series forecasting, and change-point analysis.

Security, governance, and cost stay under administrator control

Admins keep control over security, governance, audit, and cost on both the Atlassian and Dynatrace sides.

  • Per-user enforcement, end-to-end. Every Dynatrace call runs as the requesting user via OAuth 2.1, with authorization based on the user’s Dynatrace permissions. On the Atlassian side, Rovo and Rovo Ops only see the Jira and JSM data that the user is already entitled to see.
  • Per-tool selection. From a checklist in the admin panel, administrators choose which Dynatrace tools the Rovo agents in the organization can call. (New tools released later must wait for admin approval before they become available.)
  • Audit trail on both sides. Every Rovo invocation and every Dynatrace MCP call names the same person, with no stitching required between platforms.
  • Usage and costs are observable in Dynatrace. Tool call volume can be observed with Dynatrace, and costs can be clearly attributed to data owners.
  • Adding more Dynatrace users doesn’t increase your cost. Dynatrace consumption is priced on data, not per user, so onboarding more developers and on-call engineers doesn’t instantly impact your billed costs.

Get started with the Dynatrace MCP Server for Rovo now

The integration is generally available. To connect it:

  1. In Atlassian Jira or JSM, go to Atlassian AdministrationRovoRovo MCP server.
  2. Add the Dynatrace MCP Server and authenticate against your Dynatrace tenant.
  3. Select the Dynatrace tools you want to expose to your Rovo agents.

Full setup instructions are available in the Atlassian documentation and the Dynatrace MCP Server documentation.

Because the MCP Server is included with Dynatrace SaaS, you can easily set up a pilot program. Just connect the integration for one dev team, allow them access to a narrow set of tools, and then monitor the team’s usage and related costs. The patterns that emerge (which tools are called, by whom, and at what cost) can serve as a basis for a confident wider rollout.

The post Dynatrace MCP Server for Atlassian Rovo: Investigate production problems without leaving Jira or JSM appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-mcp-server-for-atlassian-rovo-investigate-production-problems-without-leaving-jira-or-jsm/feed/ 0
Scaling enterprise AI with confidence: Dynatrace joins the Dell Technologies AI Ecosystem Program https://www.dynatrace.com/news/blog/scaling-enterprise-ai-with-confidence-dynatrace-joins-the-dell-technologies-ai-ecosystem-program/ https://www.dynatrace.com/news/blog/scaling-enterprise-ai-with-confidence-dynatrace-joins-the-dell-technologies-ai-ecosystem-program/#respond Tue, 19 May 2026 19:02:13 +0000 https://www.dynatrace.com/news/?p=74013 Dynatrace and Dell Technologies

Most enterprises have moved past the deployment problem. The harder question is what those workloads are doing in production: where GPU spend is going, how agent chains are behaving, and whether compliance teams can answer when regulators ask. When the answers aren’t clear, the consequences land fast and are rarely contained to one team. That’s […]

The post Scaling enterprise AI with confidence: Dynatrace joins the Dell Technologies AI Ecosystem Program appeared first on Dynatrace news.

]]>
Dynatrace and Dell Technologies

Most enterprises have moved past the deployment problem. The harder question is what those workloads are doing in production: where GPU spend is going, how agent chains are behaving, and whether compliance teams can answer when regulators ask. When the answers aren’t clear, the consequences land fast and are rarely contained to one team.

That’s why Dynatrace is joining the Dell Technologies AI Ecosystem Program, bringing full-stack AI and LLM observability natively into a broad and integrated AI infrastructure ecosystem. Dell delivers the validated, integrated infrastructure to run AI at scale. Dynatrace brings the observability, automation, and governance to operate it with confidence, with visibility from GPU infrastructure to model behavior to end-user experience. Together, they give enterprises the control to match the scale they’ve already built.

The real challenge: AI at enterprise scale

Running AI in a pilot is very different from running it at scale across the business with real users, regulated data, and demanding SLAs. As we’ve worked with enterprises across industries, these failure patterns come up repeatedly:

Cost

As enterprises scale AI, costs spiral rapidly and unpredictably across model providers, GPU clusters, and inference APIs without clear line of sight into what is driving spend or whether it’s delivering value.

Observability gaps

Traditional monitoring tools weren’t built for AI pipelines. Fragmented observability across GPU clusters, orchestration layers, and inference APIs creates blind spots while LLM latency and token throughput fluctuations under load remain difficult to diagnose and even harder to predict.

Agentic complexity

Multi-step agent workflows introduce cascading failure modes. A silent error in one tool call can corrupt downstream decisions across the entire chain.

Compliance & governance

Enterprises need continuous monitoring to detect model drift, hallucinations, and unsafe outputs before they impact end users. Regulated industries need audit trails, data governance, and behavioral monitoring that most AI monitoring bolt-ons simply weren’t built for.

These aren’t edge cases. They’re the norm. And they’re the reason so many AI initiatives stall between pilot and production.

“Agentic AI changes what observability has to do. You’re no longer watching one model respond to one prompt. In agentic AI, every transaction can be unique, and you’re tracing chains of autonomous decisions across dozens of tools and services. That’s the problem Dynatrace was built to solve and Dell AI Factory is exactly the foundation enterprises need to take AI to production at scale.”

— Steve Tack, Chief Product Officer, Dynatrace

Scale AI workloads with confidence

Dynatrace can be integrated into Dell AI Factory environments to cover end-to-end observability of agentic AI and LLM workloads. The goal is straightforward: no blind spots, no surprises, and no manual investigation when something goes wrong. Here’s what that looks like in practice:

  • Unified AI observability to monitor the AI stack. Prompts, Model calls and downstream services, in a single platform that replaces the fragmented tooling most teams rely on today.
  • Automated prevention and remediation with Dynatrace Intelligence®. When AI workloads behave unexpectedly, Dynatrace Intelligence detects anomalies in real time and triggers automated remediation to minimize or eliminate downstream consequences.
  • End-to-end agentic AI tracing. Distributed tracing across multi-step agent chains, tool calls, RAG pipelines, and external integrations gives teams visibility into how AI agent decisions are made and where they go wrong.
  • Automatic topology mapping with Smartscape®. Maps every component in your Dell AI Factory environment, showing in real time how infrastructure, services, and AI models depend on and affect each other.
  • Built-in data governance and audit trails. Track data flows, model decisions, and AI service behavior with governance capabilities designed for regulated industries not retrofitted to them after the fact.
  • Faster resolution with Dynatrace Assist. Natural language querying and AI-generated remediation recommendations help operations teams resolve issues faster, even without deep AI infrastructure expertise.

Built for the industries where AI is becoming mission critical

AI is no longer an experiment. It’s become core infrastructure for the world’s most demanding enterprises, embedded in the decisions, workflows, and customer experiences that keep businesses running. When AI is mission critical, a failure isn’t a learning opportunity; it’s a negative business impact. Tolerance for poor visibility, unexplained latency, or untraceable decisions drops to zero. That’s precisely where Dynatrace AI Observability comes in, giving teams the visibility, control, and real-time intelligence to keep AI running when it matters most.

“The enterprises winning with AI aren’t running one model in one department. They’re operationalizing AI across the business. Dynatrace joining the Dell Technologies AI Ecosystem Program gives those customers the observability foundation to expand AI workloads on Dell infrastructure with the reliability, governance, and efficiency that enterprise-scale demands.”

— Brad Maltz, Senior Director of AI Solutions, Dell Technologies

What this means for joint customers

For organizations deploying on Dell AI Factory infrastructure, the combination of Dell’s validated hardware and software stack with Dynatrace’s intelligent observability platform means:

  • Scale with confidence. Expand production AI across the business without losing visibility or control.
  • Higher AI reliability. Proactive anomaly detection surfaces issues early; moving teams from reactive firefighting to confident operations.
  • Lower risk at scale. Broad stack visibility reduces the unknowns that make executive teams cautious in moving AI to production at scale.
  • Improved ROI on AI investment. When AI workloads run efficiently and every GPU hour is visible, teams can continuously optimize performance and cost.

End-to-end observability isn’t a nice-to-have for AI. It’s a prerequisite for trust, and trust is what turns AI investments into business outcomes. We’re proud to bring that capability to the Dell AI Factory ecosystem, and we’re excited about how this deepening of our relationship with Dell can unlock incredible value for our joint customers on their AI journeys.

Learn more about Dynatrace AI observability today, or reach out to your Dynatrace account team.

The post Scaling enterprise AI with confidence: Dynatrace joins the Dell Technologies AI Ecosystem Program appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/scaling-enterprise-ai-with-confidence-dynatrace-joins-the-dell-technologies-ai-ecosystem-program/feed/ 0
Dynatrace expands AI Coding Agent monitoring for Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK https://www.dynatrace.com/news/blog/dynatrace-expands-ai-coding-agent-monitoring/ https://www.dynatrace.com/news/blog/dynatrace-expands-ai-coding-agent-monitoring/#respond Thu, 30 Apr 2026 14:39:57 +0000 https://www.dynatrace.com/news/?p=73871 Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK

AI coding agents are a core part of how modern engineering teams build, review, deploy, and troubleshoot software. But as usage grows, so do the operational questions: Which agents are being adopted? What are the associated costs? How reliable are coding agents within real developer workflows? Which tools do they invoke, and where are they slowing down, failing, or creating unnecessary risk in production?

The post Dynatrace expands AI Coding Agent monitoring for Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK appeared first on Dynatrace news.

]]>
Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK

Dynatrace helps you answer these questions by extending AI observability for a new wave of coding agents, including Claude Code, Google Gemini CLI, OpenAI Codex CLI, OpenCode, and GitHub Copilot SDK. Together, these integrations give engineering leaders, platform teams, and developers a consistent way to understand agent activity, token consumption, costs, tool behavior, and runtime impact: without forcing teams to stitch together fragmented telemetry across terminals, SDKs, dashboards, and development workflows. Dynatrace public AI agent instrumentation examples on GitHub demonstrate how to provide industry leading observability that drives performance, cost efficiency, and governance across complex, distributed AI-driven systems—all through a unified Dynatrace platform experience that developers can access directly via MCP without leaving their IDE.

From agent activity to engineering insight

As organizations adopt multiple coding agents, new adoption challenges emerge. One team might use Claude Code in the terminal, another may build internal tools with GitHub Copilot SDK, while others experiment with Gemini CLI or Codex CLI. Platform teams want visibility into usage, availability, and costs. Engineering leaders want to know whether agents improve delivery. Security and governance teams want confidence that all prompts, tool usage, and actions can be monitored appropriately.

Dynatrace provides a practical answer to these challenges: a single observability layer for agile development workflows. For agents that emit OpenTelemetry directly, such as Claude Code, Gemini CLI, and Codex CLI, Dynatrace can ingest telemetry related to sessions, tokens, costs, tool executions, errors, and performance. For GitHub Copilot workflows, Dynatrace adds production context, software delivery automation, and GitHub-based integrations that connect agent activity to real engineering workflows.

The payoff is clear. Developers gain visibility into how agents behave in real work. Platform teams can track adoption, usage trends, and cost signals. Engineering leaders can correlate agent activity with commits, pull requests, and delivery outcomes. And with an MCP-enabled production context, teams can connect coding-agent actions to what is happening in production.

“Before we instrumented Claude Code, we had no easy way to break down how our engineers actually used AI, which models, for what tasks, and at what cost. Now we can pinpoint inefficient model use and guide usage toward better cost-performance tradeoffs.”
— Markus Heimbach, Senior Director Software Development

Anthropic Claude code monitoring dashboard in Dynatrace

Multiple coding agent experiences, one observability strategy

Each coding agent has a different operating model, which is why a common observability layer matters.

Claude Code

Claude Code already supports built-in OpenTelemetry, making it easy to send metrics and logs to Dynatrace with no code changes. Teams can track sessions, tokens, costs, tool activity, API health, and engineering output such as commits and pull requests. Logs, dashboards, and alerts help teams investigate failures, spot latency spikes, and catch unusual spend or error patterns early.

Gemini CLI

Gemini CLI includes OpenTelemetry-based observability and preconfigured dashboards, making it a strong fit for Dynatrace AI observability. Teams can correlate agent activity with broader platform signals and move quickly from raw telemetry to action. This includes debugging failed runs, identifying slow or error-prone tool calls, and alerting on cost or reliability regressions.

Codex CLI

Codex CLI supports opt-in OpenTelemetry monitoring, giving teams a path to audit usage and strengthen governance across CLI, IDE, and app experiences. With Dynatrace, logs and traces help investigate request flows, delays, and failures across agent workflows. Alerts can flag degraded reliability, unexpected behavior, or rising token consumption before they become larger issues.

GitHub Copilot SDK

GitHub Copilot SDK lets teams embed agentic workflows directly into applications, while Dynatrace adds live observability and security context. This matters because embedded agents become part of real engineering and production-adjacent workflows. Dynatrace helps trace execution paths, use logs for debugging and auditability, and set alerts for failures, latency, or policy-relevant events.

OpenCode

OpenCode is a terminal-based AI coding agent that helps developers work through coding tasks directly from the command line. Because OpenCode ships with native OpenTelemetry support, teams can route telemetry to Dynatrace without code changes by setting standard OTLP environment variables. With Dynatrace, teams can track LLM call volume, session activity, tool usage, request latency, and workflow behavior across real developer sessions. Traces help teams inspect LLM requests, tool executions, session lifecycle events, message processing, file snapshots, or diff operations.

Across all operating models, the value is the same: one strategy for monitoring adoption and impact, understanding costs, logging and tracing agent activity, alerting on reliability issues, and debugging real-world workflows as coding agents scale across the enterprise.

Distributed Tracing dashboard in Dynatrace

Why this matters now

Teams are no longer asking whether coding agents are useful. They’re asking how to drive adoption, scale them safely, govern them consistently, and prove their impact. That requires visibility into usage, cost, reliability, and engineering outcomes across teams and tools. Dynatrace helps organizations make that shift with the observability and production context needed to expand coding-agent adoption with confidence.

The coding-agent market is moving fast. Claude Code, GitHub Copilot SDK, Google Gemini CLI, and OpenAI Codex CLI each represent a different path toward agentic software delivery, from terminal-based workflows to embedded SDKs and governed local execution. At the same time, Dynatrace has been expanding its developer-facing AI surface with the Dynatrace MCP Server, GitHub Copilot integrations, and AI observability capabilities built to connect agent behavior with real production systems. The timing matters because teams are no longer evaluating whether coding agents are useful. They’re deciding how to drive adoption, scale up usage safely, govern usage consistently, and measure real impact.

Prompt activity dashboard in Dynatrace

Ready to see AI coding agents through a Dynatrace lens?

With Dynatrace, teams can understand adoption, spend, reliability, tool behavior, and engineering outcomes in one place, while giving agents access to the live production context they need to make better decisions.

Whether your developers are working in Claude Code, building on GitHub Copilot SDK, experimenting with Gemini CLI, or adopting Codex CLI, Dynatrace helps bring observability, governance, and production awareness into the heart of agentic software delivery.

Public examples already demonstrate this approach for Claude Code, and the broader Dynatrace MCP and AI observability ecosystem provides the foundation to extend the same value across the next generation of coding agents.

Ready to learn more?

In our Git repository, you’ll find step-by-step examples for supported coding-agent workflows, including how to configure OpenTelemetry export, send telemetry data to Dynatrace, and use the provided dashboards to analyze the activity of your AI coding agents.

Visit our Git repo for detailed instructions and AI Coding Agent instrumentation examples

The post Dynatrace expands AI Coding Agent monitoring for Claude Code, Google Gemini CLI, Codex CLI, OpenCode, and GitHub Copilot SDK appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-expands-ai-coding-agent-monitoring/feed/ 0
Dynatrace for AI: Teach your AI coding agent how to use Dynatrace https://www.dynatrace.com/news/blog/dynatrace-for-ai-teach-your-ai-coding-agent-how-to-use-dynatrace/ https://www.dynatrace.com/news/blog/dynatrace-for-ai-teach-your-ai-coding-agent-how-to-use-dynatrace/#respond Thu, 23 Apr 2026 16:58:48 +0000 https://www.dynatrace.com/news/?p=73813 Agentic ecosystem

Introducing Dynatrace for AI, an open-source collection of agent skills and prompts that give any skills-compatible AI coding assistant the domain expertise it needs to work productively and accurately with Dynatrace.

The post Dynatrace for AI: Teach your AI coding agent how to use Dynatrace appeared first on Dynatrace news.

]]>
Agentic ecosystem

If you’ve already wired an AI coding assistant up to Dynatrace, through the MCP server, the Dynatrace CLI (dtctl), or a custom agent you built yourself, you’ve seen your agent have difficulty interpreting data or calling for fields that don’t exist. This makes sense, your agent may be making assumptions based upon training that isn’t relevant. It lacks the skills to understand how to get the best value from Dynatrace. That is where Dynatrace for AI fills the gap.

What are agent skills?

Agent skills are an open format for packaging domain knowledge that AI agents can load on demand. A skill is a folder containing a SKILL.md file with focused instructions, examples, and optional reference material. Compatible agents, such as Claude Code, GitHub Copilot, Cursor, Cline, or others, discover installed skills and load the full content only when it’s relevant to the task at hand.

The net effect: you can install dozens of skills without bloating an agent’s context window. Agents pull in exactly what’s relevant when it’s relevant, and ignore the rest.

Install Dynatrace agent skills via a terminal
Figure 1. Install Dynatrace agent skills via a terminal

Built for agents working with Dynatrace

Dynatrace for AI is a curated set of skills that give an agent the three things it needs to efficiently do real work on Dynatrace:

  • Access to Dynatrace data and insights: through DQL queries against Grail®, Smartscape® dependency graph, or problem records.
  • Dynatrace expertise: the syntax rules, entity-model distinctions, and query patterns that separate a working query from one that looks correct but returns nothing.
  • Task-level starting points: ready-made prompt templates for common engineering workflows, so teams don’t have to invent the approach from scratch.

Skills don’t connect to Dynatrace directly. You have to pair them with the MCP server or dtctl to perform live queries and initiate actions. Together, they turn an agent with generic observability intuition into one that easily extracts value from Dynatrace.

Complement your agent with domain expertise

The first release of Dynatrace for AI agent skills is focused on the workflows that engineering teams run every day:

  • DQL fundamentals: covering the pipeline model, core data objects, and when to use fetch, timeseries, or smartscapeNodes to prevent failures that typically come from models trained on generic query-language data.
  • Observability across the stack: services, traces, logs, frontends, and problems, each covering the entity model, key fields, and query patterns that make answers correct rather than merely plausible.
  • Infrastructure and cloud: covering Kubernetes, AWS, and hosts.
  • Platform tasks worth delegating: providing programmatic creation of dashboards and notebooks

Prompt templates for common workflows

Alongside the skills, the repo hosts a small set of prompt templates you can use as structured starting points to invoke the right skills for specific tasks. These save teams from having to design their approach from scratch and make outcomes more consistent across agents and users.

Current templates include:

  • Performance regression: walks the agent through comparing RED metrics before and after a deployment, correlating any regression with distributed traces, and summarizing the root cause.
  • Daily standup: pulls the last 24 hours of problems, deployment activity, and notable anomalies for a team’s services, so anyone can walk into a standup with the relevant production context already framed.
  • Troubleshoot a problem: takes a problem ID and guides the agent through root-cause analysis, including affected entities, correlated events, relevant logs and traces, and creates a structured summary for the incident channel.

These are a starting point, not a ceiling, designed to be forked and shaped to your team’s on-call runbooks.

What Dynatrace for AI is and what it isn’t

Skills and prompts are a knowledge and workflow layer. They don’t connect to your Dynatrace environment, define what actions your agent can take, or set guardrails. That’s the job of the tool you pair them with and your Dynatrace permission model.

The quality of what your agent can produce also depends on the entities your environment is instrumented to capture. Skills help agents ask better questions of data, but they don’t control what data is collected.

Think of this skill as onboarding a smart new hire who already knows software, but needs to learn your platform. The skills are the platform user guide; your observability data is the work itself.

Get started

It’s super simple to install the skills and prompts in one go. Just run:

npx skills add dynatrace/dynatrace-for-ai

…or activate the skills as a Claude Code plugin:

claude plugin marketplace add dynatrace/dynatrace-for-ai
claude plugin install dynatrace@dynatrace-for-ai

Make sure your agent can reach Dynatrace, then try a real agent-skill task. A few good example starting prompts:

  • “Compare the error rate of the checkout service over the last hour vs the same hour yesterday.”
  • “Are any pods in the production namespace restarting or getting OOM-killed right now?”
  • “Use the performance-regression prompt to check the deployment I just shipped.”

The difference in output quality is immediate: fewer corrections, cleaner queries, and answers that accurately reflect how Dynatrace continuously models your environment in real-time.

The Dynatrace for AI project is open source and actively developed. Issues, discussions, and pull requests are all welcome, especially from teams running agent skills against real workloads. We’d love to hear from you.

Make your agents work smarter; teach them how to use Dynatrace.

The post Dynatrace for AI: Teach your AI coding agent how to use Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-for-ai-teach-your-ai-coding-agent-how-to-use-dynatrace/feed/ 0
The rise of the AI workforce: Enterprises need a new operating model https://www.dynatrace.com/news/blog/the-rise-of-the-ai-workforce-enterprises-need-a-new-operating-model/ https://www.dynatrace.com/news/blog/the-rise-of-the-ai-workforce-enterprises-need-a-new-operating-model/#respond Tue, 21 Apr 2026 15:49:54 +0000 https://www.dynatrace.com/news/?p=73763 Dynatrace and Google Cloud

A profound shift is underway in enterprise software delivery. Developers can now generate, modify, and deploy systems faster than ever using AI. But understanding what those systems are doing in production is getting harder, not easier. What began as simple prompt and response interactions with LLMs has evolved into something far more powerful: a distributed […]

The post The rise of the AI workforce: Enterprises need a new operating model appeared first on Dynatrace news.

]]>
Dynatrace and Google Cloud

A profound shift is underway in enterprise software delivery.

Developers can now generate, modify, and deploy systems faster than ever using AI. But understanding what those systems are doing in production is getting harder, not easier.

What began as simple prompt and response interactions with LLMs has evolved into something far more powerful: a distributed system of humans and AI agents working together to build, run, and adapt software. This isn’t a theoretical future. It’s happening now, and it’s reshaping how organizations must architect, operate, and govern AI at scale.

The industry has moved beyond monolithic LLMs. According to Merlin Yamssi, AI/ML CoE Lead for Partner Engineering at Google, who spoke at Dynatrace Perform this year, the early era of “LLM + prompt” model broke down quickly in real systems: no context, inconsistent behavior, and no way to act safely. That drove a rapid evolution toward retrieval, tool use, and ultimately agents that can plan, act, and collaborate across systems.

Today, AI systems behave less like single models and more like teams: one agent retrieves context, another writes code, another validates changes, and another evaluated impact in production. This shift is redefining enterprise expectations. AI is no longer here just to respond. It’s here to work.

Three forces reshaping enterprise AI

Three major trends are accelerating this transformation.

  1. Inputs are no longer just text, systems must interpret complex signals. As Yamssi put it, “You can show an image to a model, and then it will understand the image … and think on the image.”
  2. Execution is no longer linear, systems coordinate across agents.
  3. Data is no longer static, systems operate on constantly evolving context. “Essentially, [this turns] all the vast enterprise data into an active conversation,” Yamssi said.

Together, these forces are creating not just better models, but an entirely new AI operating model.

Why traditional AI architectures can’t keep up

Legacy architectures weren’t designed for distributed, autonomous AI systems. Single model approaches are rigid, difficult to debug, and prone to hallucinations. As organizations adopt multiagent systems, complexity skyrockets.

These are not only model problems; they are also distributed systems problems.

Emerging risks include the following:

  • Agents lose context as they hand tasks to one another.
  • Infinite loops occur where agents trigger each other endlessly, consuming tokens and budget.
  • Opaque decision chains can make it difficult to understand why an agent acted.
  • Token usage explodes without visibility or guardrails.

“You cannot really go into production with something that looks like a black box,” Yamssi added.

Enterprises should treat AI like a distributed system, not a chatbot.

What modern AI applications require

To support an AI workforce, organizations need a vertically integrated stack that spans five foundational layers: infrastructure, data, models, platform, and applications. Google Cloud is one of the few providers offering all five layers in a unified architecture, with hooks for observability at each layer.

On top of this foundation, a new class of agent specific tooling is emerging:

  • ADKs for building reasoning capable agents
  • MCP for standardized access to tools and enterprise systems
  • A2A protocols enabling seamless agent-to-agent collaboration across environments
  • Agent engines capable of running thousands of agents at scale

This is the new AI application stack, and it calls for a new operational model.

The missing layer: Observability for the AI workforce

Observability is increasingly about decisions, not just systems. As multiagent systems scale, observability can serve as a control plane.

Without deep visibility, enterprises face black box behavior, unpredictable costs, and operational risk. Dynatrace and Google Cloud are working to address this gap, including integrating observability capabilities with Gemini Enterprise, A2A, MCP, and other parts of the AI stack.

Modern AI observability should help reveal:

  • How decisions are made
  • How agents coordinate
  • How costs and behavior evolve in real time
  • Where failures originate across reasoning

This is a shift from monitoring applications to monitoring reasoning, decisions, and collaboration. Enterprises often need visibility from the infrastructure all the way to the data, the LLM, the agent, and the application.

For developers, this changes the job entirely. You’re no longer debugging a service. You’re debugging a system of agents, decisions, and interactions across code, cloud, and runtime behavior.

From insight to action: The path to autonomous operations

According to recent Dynatrace research, 50% of respondents have agentic AI projects in production for limited use cases, and 44% have projects in broad adoption. Further, 72% of respondents have 2-10 agentic AI projects. With agentic AI already in production and growing, observability becomes essential. Once organizations can observe their AI workforce, they can begin to automate operations.

Observability must surface not just what happened, but why. Dynatrace and Google Cloud are already enabling this through integrations with Gemini Cloud Assist, which can recommend infrastructure or application changes based on observed issues.

This unlocks a new operational loop:

  • Observe systems behavior across code, runtime, and agents
  • Diagnose causal relationships, not just symptoms
  • Recommend changes grounded in runtime context
  • Remediate automatically or in collaboration with developers and agents

The result can be faster recovery, lower cost, and safer AI deployment at scale.

Accelerate your AI workforce strategy with Dynatrace on Google Cloud

At Perform, this shift was clear: The challenge is no longer generating code or deploying models. It’s understanding and controlling how these systems behave once they’re running.

The organizations that solve this will define the next generation of software delivery. If you want to go deeper, watch the full Perform session: The AI workforce: Advancing agentic collaboration through observability.

Dynatrace and the Dynatrace logo are trademarks of the Dynatrace, Inc. group of companies. All other trademarks are the property of their respective owners. © 2026 Dynatrace LLC

The post The rise of the AI workforce: Enterprises need a new operating model appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-rise-of-the-ai-workforce-enterprises-need-a-new-operating-model/feed/ 0
Dynatrace AI agents begin working for you on day one, and are built to grow with you https://www.dynatrace.com/news/blog/dynatrace-ai-agents-begin-working-for-you-on-day-one-and-are-built-to-grow-with-you/ https://www.dynatrace.com/news/blog/dynatrace-ai-agents-begin-working-for-you-on-day-one-and-are-built-to-grow-with-you/#respond Fri, 03 Apr 2026 15:44:42 +0000 https://www.dynatrace.com/news/?p=73625 Agents graphic

AI agents are everywhere in tech conversations right now, but what agents can you actually use today to make your job easier? In Dynatrace, ready-made agents help developers, SREs, and IT operations teams investigate issues, understand system behavior, and reduce manual work using the data they trust every day. Dynatrace ready-made agents are not concepts or previews; they're available now, integrated into existing Dynatrace workflows, and designed to solve real operational problems. For teams ready to go further, Dynatrace agents lay the groundwork for autonomous operations.

This blog shows what Dynatrace ready-made agents are, how to get value from them quickly, and how to decide which agents are relevant for you, using concrete examples rather than promises.

The post Dynatrace AI agents begin working for you on day one, and are built to grow with you appeared first on Dynatrace news.

]]>
Agents graphic

From generic AI to task‑focused operational agents

Dynatrace ready‑made agents are purpose‑built capabilities that apply Dynatrace intelligence to specific, recurring operational tasks. Each agent focuses on a clearly defined problem, such as explaining why a service is slow, summarizing unusual behavior in an environment, or helping you understand what changed and why it matters. These agents are designed to take a question or a signal based on the exact data that is in your environment and organization and turn it into a useful answer you can act on.

Because Dynatrace agents are ready‑made, there is no need to define prompts, train models, or design behavior from scratch. Each agent already knows:

  • What type of input to expect,
  • Which Dynatrace signals and context it should use,
  • And what output types are most useful for each addressed problem type.

All available ready-made Dynatrace agents can be found in Dynatrace Hub.

Trigger agent actions with Dynatrace Workflows and the Dynatrace MCP Server

Ready‑made agents can be triggered automatically as part of Dynatrace Workflows or available wherever you already work via the Dynatrace MCP Server.

Using agents in Dynatrace Workflows

Dynatrace Workflows lets you run agents in response to events or on a schedule. Instead of manually asking questions about potential problems and remediation steps, the workflow autonomously responds to changes in your environment.

For example, the Kubernetes Troubleshooting Agent runs nine parallel queries for data enrichment, and Dynatrace Intelligence turns all the information into a structured diagnosis. Customize the agents to your needs, including instructions for human approval steps and automated remediation.

Dynatrace Kubernetes Troubleshooting Agent in action.
Figure 1. Dynatrace Kubernetes Troubleshooting Agent in action.

The fastest way to get started is with Dynatrace ready-made agentic workflow templates, currently available in a preview release. Instead of building from scratch, you get proven automations that summarize issues, suggest remediation, and deliver insights directly to the tools your teams already use.

Figure 2. Agentic workflow templates available in preview
Figure 2. Agentic workflow templates available in preview

Power users can go further by building their own agentic workflows that combine Dynatrace Intelligence actions with any trigger, data source, or integration in Workflows. Use cases range from auto-scaling Kubernetes clusters based on Dynatrace Intelligence forecasts to generating query-cost-optimization recommendations for stakeholders, to virtually any other automation your environment requires.

Using agents through the Dynatrace MCP Server

The Dynatrace MCP Server makes the agents available outside the Dynatrace web UI, without requiring you to deploy or operate any additional infrastructure. You can connect Dynatrace to any MCP‑compatible client in minutes, with no server to install, host, or maintain.

Through the tools exposed by the MCP Server, you can use natural language to query data in Grail®, check system health, and get problem analyses and remediation recommendations. This brings Dynatrace directly into the tools you already use, such as your IDE, Claude Code and Cowork, Microsoft Copilot, Slack, or automation platforms like n8n. The Dynatrace MCP Server also powers integrations with systems like Azure SRE, AWS DevOps, GitHub Copilot, Atlassian Rovo Ops, Amazon Q, and others.

Dynatrace MCP server in Visual Studio Code with GitHub Copilot
Figure 3. Dynatrace MCP server in Visual Studio Code with GitHub Copilot

This means agents are no longer tied to a single interface. You can ask Dynatrace questions and get grounded, production‑ready answers wherever you work, using the same agents and intelligence that power Assist and workflows.

Dynatrace Assist: a simple way to test ready-made agents

The quickest way to use a ready‑made agent and see how it works before you start creating a workflow is with . Dynatrace Assist lets you ask questions about your environment using natural language, without switching tools or setting anything up.

A simple way to start is with a real problem you already have. For example, when a service becomes slow, open Assist and ask a question such as “Summarize the open problems and highlight those that need immediate attention.” Assist interprets the question, evaluates the environment you’re working in, and pulls together relevant data and context using Dynatrace Intelligence. Instead of manually navigating metrics, traces, logs, and dependencies, you get an explanation grounded in what is actually happening in your system.

Continuing your conversation with Assist, you can refine the question or follow suggested drill‑downs. Assist supports this as a single flow, helping you move from an initial question to deeper analysis and, where applicable, to next steps. You’re not configuring an agent or defining behavior. You’re simply asking a question and letting Dynatrace coordinate the right intelligence and ready‑made agents behind the scenes.

Dynatrace Assist
Figure 5. Dynatrace Assist

This makes Assist your lowest‑friction entry point for using Dynatrace agents. You get a concrete result quickly, using the same data and context you already rely on in your daily work.

What’s next?

If you haven’t already, open Dynatrace Playground, or your Dynatrace tenant, and ask Dynatrace Assist a question to see the ready-made agents in action.

The post Dynatrace AI agents begin working for you on day one, and are built to grow with you appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-ai-agents-begin-working-for-you-on-day-one-and-are-built-to-grow-with-you/feed/ 0
Bring real-time production insights into Claude Code with the Dynatrace MCP Server https://www.dynatrace.com/news/blog/bring-real-time-production-insights-into-claude-code-with-the-dynatrace-mcp-server/ https://www.dynatrace.com/news/blog/bring-real-time-production-insights-into-claude-code-with-the-dynatrace-mcp-server/#respond Mon, 30 Mar 2026 18:36:42 +0000 https://www.dynatrace.com/news/?p=73599 Dynatrace and Claude logos

Get immediate production visibility inside Claude Code, the next-gen AI coding assistant from Anthropic. The Dynatrace MCP server can now be used as a ready-to-use connector for Claude Code, Cowork, and Chat. Connect in minutes to query logs, traces, and problems; conduct live debugging from your terminal, and validate AI-first workflows with real data. Model […]

The post Bring real-time production insights into Claude Code with the Dynatrace MCP Server appeared first on Dynatrace news.

]]>
Dynatrace and Claude logos

Get immediate production visibility inside Claude Code, the next-gen AI coding assistant from Anthropic. The Dynatrace MCP server can now be used as a ready-to-use connector for Claude Code, Cowork, and Chat. Connect in minutes to query logs, traces, and problems; conduct live debugging from your terminal, and validate AI-first workflows with real data.

Model Context Protocol (MCP) is the standard for connecting AI assistants to live tools and data. The Dynatrace MCP server brings your full observability and security context into every Claude session. This means fewer context switches, less time hunting through dashboards, and better answers that are grounded in what’s actually happening in your environment.

How to connect Claude Code with Dynatrace

Getting started is straightforward. Open Claude, go to Connectors, and search for “Dynatrace.” Install the Dynatrace MCP Server connector, follow the setup steps, and you’re connected.

Figure 1. Dynatrace MCP Server connector setup in Claude
Figure 1. Dynatrace MCP Server connector setup in Claude

How Dynatrace MCP Server and Claude give you visibility into your production data

Whether you’re investigating a production issue, reviewing a deployment, or working through a security vulnerability, you can ask Claude in plain language and get answers backed by your actual Dynatrace production data. This is not just documentation summaries or generic guidance, but a real window into production.

Troubleshoot without leaving Claude Code

Let’s say you receive a Jira ticket with details of an error. Instead of switching to dashboards and digging through logs to find out what went wrong, you can now ask Claude. Dynatrace pulls all root-cause information, related logs, metrics, traces, and CPU and memory profiling data from your production environment directly into the session. You can query these details by impact, filter by service, and request remediation suggestions, all without knowing in advance where to look or how to write the DQL query.

Figure 2. Claude gets the problem description and production data via the Dynatrace MCP Server.
Figure 2. Claude gets the problem description and production data via the Dynatrace MCP Server.

This same workflow applies when you’re checking for vulnerabilities in your running workloads, verifying whether a recent deployment introduced regressions, or proactively catching performance issues before they hit production.

You can now also use dtctl, the open source CLI for the Dynatrace platform, in Claude Code alongside the Dynatrace MCP server to manage dashboards, run workflows, and execute DQL from your terminal.

Ready to get started? Install Dynatrace MCP Server in Claude Code today.

The post Bring real-time production insights into Claude Code with the Dynatrace MCP Server appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/bring-real-time-production-insights-into-claude-code-with-the-dynatrace-mcp-server/feed/ 0
dtctl: The Dynatrace observability CLI that’s built for AI agents and humans https://www.dynatrace.com/news/blog/dtctl-the-dynatrace-observability-cli-thats-built-for-ai-agents-and-humans/ https://www.dynatrace.com/news/blog/dtctl-the-dynatrace-observability-cli-thats-built-for-ai-agents-and-humans/#respond Mon, 23 Mar 2026 19:38:51 +0000 https://www.dynatrace.com/news/?p=73510 AI agents graphic

As AI agents take on more operational tasks, the tools they use to interact with platforms matter. MCP (Model Context Protocol) is emerging as the standard for structured agent-tool interaction, but it adds an abstraction layer that not every workflow needs. Sometimes you just need to run a command, get a result, and act.

The post dtctl: The Dynatrace observability CLI that’s built for AI agents and humans appeared first on Dynatrace news.

]]>
AI agents graphic

What is dtctl?

dtctl (short for “Dynatrace control”) is the open-source CLI for the Dynatrace platform; it’s kubectl-inspired, terminal-native, and designed for both AI agents and humans. Platform engineers, SREs, and developers use it to manage workflows, dashboards, queries, and settings from the command line.

With dtctl, AI agents use the same command line interface and commands that your platform engineers, SREs, and developers use to autonomously manage workflows, dashboards, queries, and settings.

Three ways to connect agents to Dynatrace

Dynatrace offers three access patterns for AI agents and automation, all governed by the same IAM permission system. Regardless of which path you choose, an agent can access only the data and operations permitted by its authentication scopes. The right choice depends on your use case; in practice, most teams use more than one approach.

  • MCP (Model Context Protocol): A standardized, schema-driven interface in which every tool call is declared upfront and validated automatically. Agents get structured, predictable interactions while platform teams get strict control over which tools are exposed and a full audit trail of every call.
  • API: Access to the full platform surface with maximum flexibility. The API is ideal when you need endpoints that higher-level tools can’t cover. The tradeoff: you build and maintain everything yourself, from endpoint selection to token lifecycle, pagination, and rate-limit retries.
  • dtctl (CLI): A terminal-native and composable command line interface that’s built for execution speed, with minimal setup: Just run a command, inspect the result, adjust, and repeat. This offers a natural fit for tight iteration loops, scripting, and agents that need to get things done without overhead.

Why is CLI gaining traction for agent workflows?

CLIs have been the primary interface for automation since the early days of Unix, and for good reason. Output is structured, syntax is predictable, and there’s no protocol overhead. MCP mirrors how humans interact with tools and includes discovery, negotiation, and structured handshakes. This is valuable when you need it; however, every schema negotiation round-trip adds tokens and latency to the agent’s context window. CLI skips that entirely and cuts straight to execution.

Examples of dtctl capabilities
Figure 1. (video) Examples of dtctl capabilities

Use the same CLI tool and commands for human engineers, scripts, and AI agents

Whether you’re a platform engineer or an AI agent, dtctl gives you the same powerful Dynatrace interface. Like kubectl or git, dtctl follows a simple verb-noun syntax: Just state what it is you want to do, then state what you want to act on. What makes dtctl stand out is that it’s designed to support human engineers and AI agents equally.

  • A single interface for everything. Workflows, dashboards, notebooks, queries, SLOs, Dynatrace Intelligence, and more, all accessible through the same consistent set of commands. There’s no need to stitch together multiple API endpoints or learn different tools for different resources.
  • Built for AI agents. dtctl lets agents discover all available commands at runtime, no documentation needed, no upfront configuration. When running inside an AI agent, dtctl automatically switches to structured output that agents can parse and act on, including follow-up suggestions and error context.
  • Built for humans, too. Use tab-autocomplete resource shortcuts like db for dashboards and wf for workflows, –mine to filter your own resources, and an edit command that opens YAML in your $EDITOR and uploads on save. Because dtctl follows familiar command-line patterns, experienced users move fast from day one.
  • Managing multiple environments is simple. Switch contexts between dev, staging, and production with a single command. Authenticate via SSO or API token, run dtctl doctor to verify the setup, and you’re ready to go.

For all technical details and the full command reference, visit the dtctl repository on GitHub.

An AI agent modifies a Dynatrace workflow end-to-end

What makes AI agents truly useful is their ability to close the loop: They can discover data, make changes, verify results, and fix what’s broken all without handing control back to a human operator.

Here’s what such a scenario looks like using dtctl. In this example, a Dynatrace workflow queries the number of Kubernetes pods and sends an email report. The goal is to enhance the query so that it lists every pod with its respective node tolerations, making it a cross-entity query that explores the data model, identifies the correct relationships, and iterates until the output is correct.

Using GitHub Copilot in VS Code, the agent works through the full cycle autonomously:

  • Discover: Explores the Dynatrace data model to find the right entities and relationships: how pods connect to nodes and where tolerations are stored.
  • Iterate: Refines the query step by step until the output matches the goal, then updates the workflow and rebuilds the email report.
  • Apply and run: Pushes the updated workflow to Dynatrace and executes it.
  • Verify: Checks whether the workflow ran successfully and produced the expected results.
  • Fix: If something fails, it reads the error, adjusts the query, and tries again.

The human operator defines only the intent, and the agent handles the rest. Watch this full video walkthrough to see it in action.

Create and modify dashboards without leaving your code editor

One concrete example of what dtctl enables is dashboard creation and management directly from the terminal. Because Dynatrace dashboards are structured data, they can be version-controlled, templated, and automated just like any other code artifact.

To illustrate this, the OpenClaw Gateway Monitoring dashboard shown below was created end‑to‑end in just a few minutes. A developer used GitHub Copilot in VS Code to create an AI observability dashboard similar to others, tailored specifically for monitoring OpenClaw. The agent then pulled existing Dynatrace AI observability dashboards as templates, adapted the layouts and queries to OpenClaw’s monitoring needs, and deployed the results using dtctl; all this was managed by the developer without leaving their IDE.

A custom AI Observability dashboard, created using dtctl.
Figure 2. A custom AI Observability dashboard, created using dtctl.

For day-to-day management, the workflow remains the same, whether executed by an agent or a human: just pull a dashboard, adjust queries or filters, preview the changes, and save the dashbaord. Coding agents can be instructed to update queries across multiple dashboards with a single prompt. And it takes just a single command to promote dashboards from dev to production, or to roll them back instantly if something breaks.

You can take dtctl even further: wire dashboard definitions into your CI/CD pipeline so that, as a service evolves, its dashboards evolve automatically.

Try out dtctl today

dtctl is fully open source and available at dynatrace-oss/dtctl on GitHub, with documentation, skills, and examples to get you started.

Our dtctl Quick Start Guide walks you through the complete setup in under five minutes:

  1. Connect dtctl to your environment.
  2. Run your first query.
  3. Pull a dashboard.
  4. Modify the dashboard queries.
  5. Save your dashboard.

From there, install the agent skill to teach GitHub Copilot, Claude Code, or Cursor how to operate your Dynatrace environment.

The project is in active development. If you hit a bug or have a use case to share, open a GitHub issue or start a discussion, and help us shape the roadmap.

The post dtctl: The Dynatrace observability CLI that’s built for AI agents and humans appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dtctl-the-dynatrace-observability-cli-thats-built-for-ai-agents-and-humans/feed/ 0
10 things I learned writing 49,000 words about vibe coding https://www.dynatrace.com/news/blog/10-things-i-learned-writing-49000-words-about-vibe-coding/ https://www.dynatrace.com/news/blog/10-things-i-learned-writing-49000-words-about-vibe-coding/#respond Fri, 13 Mar 2026 18:03:17 +0000 https://www.dynatrace.com/news/?p=73437 31 Days of Vibecoding

In January 2026, I ran an experiment: I published one blog post every single day for 31 days about building production software with AI. Not toy demos. Not “hello world” chatbots. Real, shipped, supported product. The whole series was built around one project: collectyourcards.com, a sports card collection app I built entirely with Claude Code […]

The post 10 things I learned writing 49,000 words about vibe coding appeared first on Dynatrace news.

]]>
31 Days of Vibecoding

In January 2026, I ran an experiment: I published one blog post every single day for 31 days about building production software with AI. Not toy demos. Not “hello world” chatbots. Real, shipped, supported product. The whole series was built around one project: collectyourcards.com, a sports card collection app I built entirely with Claude Code over three months of nights and weekends. 1,000,000+ cards, an achievement system, universal sub-200ms search, Excel exports. All of it.

I called it 31 Days of Vibe Coding. “Vibe coding” is the increasingly common practice of describing features to an AI assistant and iterating on generated code instead of writing every line manually.

I documented my journey on a website and newsletter that covered my learnings in bite sized chunks. Check out the full series. If any of these lessons resonated, the detailed posts have the code, the prompts, and all the messy parts I couldn’t fit here. I also appeared on the PurePerformance podcast to discuss my learnings. Writing about and discussing my work helped me better understand what I learned. Here are a few important insights I gleaned from my work.

1. You’re the architect. AI is the junior developer.

This is the framing that makes everything else work. You decide what to build, how it should be structured, and what trade-offs are acceptable. AI helps you build it faster. When you flip that relationship and start expecting AI to make design decisions for you, things go sideways fast.

I learned this the hard way when I described a notification system to Claude and asked it to “figure out the best approach.” It picked polling every 30 seconds. That works fine for 10 users. At 10,000 users, that’s 333 requests per second hitting your server for no reason. When I instead described my proposed approach and asked Claude to poke holes in it, it identified five real problems I hadn’t considered, including the scaling issue. Same tool, completely different outcome, because I stayed in the architect role.

AI doesn’t replace your judgment. It accelerates your execution.

2. Write specs, not prompts.

Early on, I was typing things like “build me a card search feature” into Claude and getting back something that technically worked but was missing pagination, fuzzy matching, error handling, and observability. The output was exactly as vague as my input.

The fix was treating every feature like a spec. I started writing GitHub Issues as full feature specifications before ever opening a conversation with AI. Context, intent, constraints, examples, and how to verify it works. When I built the achievement system for collectyourcards.com (1,200+ achievements across 14 categories), the entire thing was driven by a single well-written GitHub Issue. AI even helped me write the spec before a single line of code existed, and then created follow-up issues during implementation for database indexes, notification hooks, and edge cases I’d missed.

Vague input gets vague output. Specific input gets production-ready output. Every time.

3. Break every feature into phases.

Never ask AI to build a complex feature in one shot. I tried this exactly once with the achievement system. Forty-five minutes of generated code spread across dozens of files, and none of it worked because everything was interdependent and untestable. I couldn’t even figure out where to start debugging.

I threw it all away and started over with phases.

  • Phase 1: core achievement engine with 150 basic achievements. That shipped in one session and worked immediately.
  • Phase 2 added categories.
  • Phase 3 added the notification system. Each phase produced working, deployable, independently verifiable software. The total time was less than my failed single-shot attempt.

This applies to everything. My universal search feature went from basic text matching to fuzzy search to multi-entity results to filters to autocomplete, each phase shipping on its own. If Phase 1 alone isn’t useful, your phases are too small. If Phase 1 takes more than one session, your phases are too big.

4. Commit before every AI operation.

Git is your undo button. This sounds obvious until you skip it once and lose an hour of your life.

I asked Claude to change a five-word error message in a validation function. Simple, right? When I looked at the diff afterward, it had touched 12 files. It renamed the function from `validateUser` to `checkUserStatus`, refactored the validation logic, and updated three other files that referenced the original function name. The app crashed. Without a prior commit, I spent 50 minutes manually figuring out what had changed and reversing the damage.

Now I commit before every significant AI operation. `WIP: before achievement refactor`. `WIP: before search update`. It takes five seconds. When AI changes something you didn’t ask for (and it will, inevitably), you run `git diff`, see exactly what happened, and revert cleanly. Think of it like saving before a boss fight in a video game.

5. Configure your AI once, not every conversation.

AI has no memory between conversations. Every new session starts from zero. For weeks, I was repeating the same instructions: “Use TypeScript, use Prisma, use the service layer pattern, add structured logging, don’t use console.log.” Every. Single. Time.

The fix was a CLAUDE.md file in my project root that defines my tech stack, coding patterns, and explicit “Always” and “Never” lists. I added annotated pattern files that show, not describe, how services should be structured. I added a common mistakes file documenting errors AI kept making so it would stop making them.

Before configuration, asking Claude to “create an endpoint to update user email” produced JavaScript with raw SQL and console.log. After configuration, the exact same prompt produced TypeScript with strict types, Prisma, a proper service layer, and structured telemetry logging. Same AI, same prompt, dramatically different output, because the context was already there.

6. AI builds for the happy path. You have to demand the rest.

AI-generated code is optimistic. Really optimistic. It assumes your database is always up, your network is always fast, and your users always send valid JSON. It writes code that works perfectly when everything goes right and fails silently when anything goes wrong.

I built a card-fetch endpoint that looked clean and correct. Eight lines of straightforward code. When I ran a security audit (by asking Claude to switch into adversarial “penetration tester” mode), it found four vulnerabilities: no ownership check (any user could fetch any other user’s cards), sensitive data exposure (returning all owner fields including email), no rate limiting, and no access logging. The fix was about 15 lines, but it never would have written them unless I explicitly asked.

Security, operability, edge cases, and test coverage don’t happen unless you ask for them separately. I started using what I call the “3am test” for every feature: if this broke at 3am, would I know it happened? Would I know why? Would I know how to fix it? If the answer to any of those is no, the feature isn’t done.

7. Observability replaces manual code review.

When you’re shipping AI-generated code, you’re not reading every line. You can’t. The volume is too high and the iteration speed is too fast. So how do you know it works? You watch it run.

I add structured logging, metrics, and tracing to every new feature on Day one. When a user gets locked out of their account on collectyourcards.com, I can diagnose why in 30 seconds by looking at the authentication telemetry: user ID, IP, user agent, error message, all structured and searchable. Before I had this, the same diagnosis took hours of guessing.

This isn’t optional when you’re building with AI. It’s the substitute for the line-by-line code review you’re no longer doing. OpenTelemetry with traces, metrics, and logs became the foundation of every feature I built. If I couldn’t observe it in production, I didn’t ship it.

8. Use AI to review AI.

One of the most useful patterns I found was using AI to check AI’s own work. But there’s a trick: run multiple focused review passes, not one general review.

A single “review this code” pass might catch three or four issues. Four separate focused passes (security, performance, edge cases, maintainability) routinely caught 12 or more. I’d ask Claude to review a user registration endpoint once for security vulnerabilities, once for performance problems, once for missing edge cases, and once from the perspective of someone maintaining this code at 3am. Each pass found things the others missed.

I also used AI to write tests for AI-generated code, to refactor AI-generated code, and to find edge cases I’d never think of. A `formatUserName()` function seems simple until AI points out: What about right-to-left text? Emoji in names? HTML injection? A user who entered no name at all (producing “undefined undefined” on the profile page, which actually happened)?

9. Measure whether it’s actually helping.

Here’s the honest part. I spent 31 days writing about how AI makes you faster and more productive. And when I actually measured my output, my features-shipped-per-week, bugs-in-production, and time-from-issue-to-deploy were… about the same as before.

I felt faster. The moment-to-moment experience of coding with AI feels incredible. Code appears instantly. Features take shape in minutes instead of hours. But the total cycle time, including prompt crafting, output review, fixing hallucinations, and debugging unexpected changes, often washed out the speed gains.

Where AI genuinely helped: boilerplate generation, code review, exploring unfamiliar APIs, enumerating edge cases, and writing documentation. Where it consistently hurt: novel problems with no clear pattern, subtle bugs that required deep context, and any situation where it over-engineered a simple solution. If you’re not tracking which category your work falls into, you’re flying blind.

10. Document what you learn, because AI won’t remember (yet).

AI’s context degrades within a single session and vanishes completely between sessions. Around message 30 or 40 in a long conversation, I’d start seeing Claude reference models that didn’t exist, import from paths that were never created, and contradict decisions it made 20 messages earlier.

The fix is aggressive documentation. Progress docs that summarize what was built, what decisions were made, and what’s next. A mistakes file that records patterns AI keeps getting wrong (with “wrong” and “right” examples) so you can reference it in future sessions. Proactive compaction at natural breakpoints, not after quality has already collapsed.

Your configuration files, your mistakes file, and your progress docs are the substitute for the institutional knowledge a human teammate would accumulate over months. AI doesn’t learn from working with you–at least not yet. But your documentation can simulate that learning for every future session.

How I changed over 31 days

After 31 days and nearly 49,000 words, the biggest shift was in how I think about what I’m building. I plan more carefully because the cost of planning is low and the cost of a bad AI-generated implementation is high. I commit more often. I monitor everything. I write better specs than I ever did when I was the one writing every line.

AI didn’t make me a faster coder. It made me a better architect.

The post 10 things I learned writing 49,000 words about vibe coding appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/10-things-i-learned-writing-49000-words-about-vibe-coding/feed/ 0
Observe API responses to runtime behavior: Connect Postman’s Agent Mode with Dynatrace https://www.dynatrace.com/news/blog/connect-postman-agent-mode-with-dynatrace/ https://www.dynatrace.com/news/blog/connect-postman-agent-mode-with-dynatrace/#respond Thu, 12 Mar 2026 12:58:47 +0000 https://www.dynatrace.com/news/?p=73368 Dynatrace and Postman

Dynatrace and Postman announced an expansion of their technology alliance to bring real-time observability directly into AI-assisted API workflows with the launch of Postman’s Agent Mode. With the Dynatrace MCP Server now available in Postman’s MCP Catalog, developers can securely connect Agent Mode to trusted Dynatrace telemetry and production context without leaving the Postman environment. By connecting Postman Agent Mode with Dynatrace, developers can move beyond validating API responses to understanding how those APIs behave in real systems, under real load, and across real dependencies. The result is faster insight, fewer blind spots, and greater confidence that APIs perform as expected when it matters most.

The post Observe API responses to runtime behavior: Connect Postman’s Agent Mode with Dynatrace appeared first on Dynatrace news.

]]>
Dynatrace and Postman

What Agent Mode can do with Dynatrace observability data

Once connected with Dynatrace MCP Server, Postman’s Agent Mode works with real runtime data rather than just API requests and responses. When an API test fails or behaves unexpectedly, the agent can correlate the behavior with Dynatrace telemetry such as service metrics, traces, logs, and detected anomalies.

This allows developers to quickly understand whether an issue is caused by the API itself or by what’s happening behind it, such as a slow dependency, a failing backend service, or a recent change impacting runtime behavior. Because this context comes directly from Dynatrace, developers can ask follow‑up questions in natural language and get answers grounded in production data. All of this happens inside the Postman workflow, without switching tools or manually querying observability dashboards.

The result is a faster feedback loop: APIs can be tested, validated, and debugged with awareness of how they actually behave in live environments, not just how they respond in isolation.

Explore Dynatrace use cases with Postman

This integration is most valuable in situations where API behavior needs to be understood in the context of how systems actually run, not just how endpoints respond in isolation.

Explore Dynatrace use cases with Postman

Common scenarios

Instantly check Dynatrace environment health in Postman

Quickly connect Dynatrace as an MCP server in Postman, set up environment variables, and use natural language prompts to get a real-time overview of open problems in your monitored environment. This streamlines troubleshooting by surfacing actionable issues directly in Postman, saving you time and reducing context switching.

Summary of open problems identified by Dynatrace.
Figure 1. Summary of open problems identified by Dynatrace.

Diagnose and resolve API failures with Dynatrace insights

By running API tests and intentionally triggering failures, you can leverage the MCP server to pinpoint the root causes of errors, such as backend issues or logic bugs, using the real-time insights from Dynatrace. This empowers developers and testers to quickly distinguish between code and infrastructure problems, accelerating debugging and improving application reliability.

Dynatrace analyzes frontend performance metrics
Figure 2. Dynatrace analyzes frontend performance metrics

Get a holistic application health assessment in Postman

You can ask the MCP server in Postman broad, natural-language questions about the overall health of your application and the services monitored by Dynatrace. The server provides a summary of performance and health across multiple services, highlights slow endpoints, and offers actionable recommendations. This allows you to quickly identify bottlenecks and focus on critical issues, supporting proactive performance management and efficient troubleshooting.

Holistic overview of an application monitored by Dynatrace in Postman
Figure 3. Holistic overview of an application monitored by Dynatrace in Postman

Automate continuous health checks with Postman collections

You can have Postman automatically generate collections that call Dynatrace endpoints on a schedule, enabling ongoing performance and health monitoring between releases. This ensures teams are proactively alerted to issues, supporting continuous delivery and higher service quality.

Configuring scheduled monitoring to automatically receive insights from Dynatrace to Postman
Figure 4. Configuring scheduled monitoring to automatically receive insights from Dynatrace to Postman

Reduce guesswork in AI-assisted workflows

AI‑assisted development is only as effective as the context an agent can access. Without reliable runtime data, agents are limited to reasoning from API definitions, test results, and assumptions, making it difficult to distinguish real system issues from surface‑level symptoms.

By connecting Postman’s Agent Mode to Dynatrace through MCP, agents can ground their reasoning in trusted observability data from live systems. This gives the agent access to the same production signals developers rely on today, such as service health, performance trends, errors, and dependencies, rather than relying on inferred behavior alone.

The result is more actionable AI assistance. Instead of guessing why an API behaves a certain way, the agent can explain what is happening in the system and why, reducing false conclusions and improving the quality of recommendations in AI‑driven workflows.

Try Postman’s Agent Mode with Dynatrace

The Dynatrace MCP Server is available in the Postman MCP Catalog and can be connected to Postman’s Agent Mode to bring runtime observability data directly into API workflows.

For teams already using Dynatrace and Postman, this is a straightforward way to add production context to AI‑assisted API development. For teams exploring Agent Mode, it provides a practical foundation for grounding agent workflows in real system behavior.

The post Observe API responses to runtime behavior: Connect Postman’s Agent Mode with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/connect-postman-agent-mode-with-dynatrace/feed/ 0
The runtime reckoning: How the agentic evolution is reshaping security https://www.dynatrace.com/news/blog/the-runtime-reckoning-how-the-agentic-evolution-is-reshaping-security/ https://www.dynatrace.com/news/blog/the-runtime-reckoning-how-the-agentic-evolution-is-reshaping-security/#respond Thu, 05 Mar 2026 18:50:47 +0000 https://www.dynatrace.com/news/?p=73311 Observability graphic

AI is fundamentally reshaping the speed and scale of cyberattacks. Financially and geopolitically motivated threat actors are executing sophisticated attacks against Fortune 500 enterprises that frequently bypass perimeter defenses.

The post The runtime reckoning: How the agentic evolution is reshaping security appeared first on Dynatrace news.

]]>
Observability graphic

AI is compressing attack timelines

The organizations pulling ahead are those with runtime visibility into the applications that directly generate and impact revenue.

Stats about AI usage

What AI is changing

In a report published in November 2025, Anthropic documented the first large-scale AI-orchestrated cyberattack campaign, in which AI autonomously performed 80–90% of attack operations with minimal human intervention. This signaled a turning point.

The threat actors causing the most damage today aren’t traditional nation-state actors or advanced persistent threats (APTs)—they’re financially motivated collectives like Scattered Spider. AI hasn’t invented fundamentally new attack techniques yet, but it has significantly increased attack volume and velocity—CrowdStrike observed an 89% increase in attacks by AI-powered adversaries in 2025 alone.1 Techniques once reserved for state-sponsored groups are now accessible to cybercriminals through AI-assisted tooling.

The same automation accelerating attacks also allows for faster defense. But with an average time of only 29 minutes from initial access to lateral movement, defenders relying on periodic assessments are structurally disadvantaged. The capability that matters now is continuous exposure assessment—analyzing real-time signals across all cloud assets and acting before adversaries complete their objectives.

Two battlefronts: One adversary

Traditional attack vectors haven’t disappeared—supply chains, endpoints, physical security, and human-targeted attacks are increasing just as rapidly. But business-critical applications have become a focal point of attack operations because lower entry barriers make them accessible to more threat actors. The applications that process transactions, manage customer data, and orchestrate supply chains are precisely where runtime compromise has the greatest business impact.

Organizations must drive parallel initiatives: one focused on people, identity governance, and communications; another on runtime protection, continuous exposure management, and application-layer visibility. These require different tools and skills—but critically, connected processes and connected insights. Siloed security functions are precisely what sophisticated attackers exploit.

Why runtime is the new frontline

Initiatives like Anthropic’s Claude Code Security and OpenAI’s reasoning-based vulnerability detection are collapsing scanning stages and shortening developer feedback loops. This is real progress. But production remains the definitive validation point. AI-generated code and accelerated release cycles introduce risks that surface only at runtime—pipeline scanners offer helpful signals but lack the reliability and context of a unified runtime view.

In an agentic ecosystem, detection and response can’t be focused on the perimeter—they must be intrinsic to the application and infrastructure. Environment-aware malware and prompt injection against enterprise AI agents aren’t visible at the perimeter. They’re visible at runtime. The XZ Utils backdoor (CVE-2024-3094) demonstrated this perfectly: malicious code passed all static analysis, activating only when loaded by sshd on targeted Linux distributions in production.7

Autonomous workflows that lack visibility into security risk operate with an incomplete picture. Integrating security context into agentic decision-making—rather than treating it as a separate operational domain—will be essential as enterprises scale these capabilities.

– IDC Link, Dynatrace Perform 2026: From Observability to Supervised Autonomous Operations (Doc #lcUS54307526, February 2026)

Security that’s embedded, not bolted on

The convergence of observability and security isn’t theoretical—it’s operational. IDC analysis notes that organizations expect the next generation of security to be embedded in solutions rather than bolted on. Dynatrace application security capabilities—runtime vulnerability analytics and runtime application protection—operate within Grail alongside observability and business data, sharing the same contextual mapping and causal dependency graph.

Resilience as a competitive advantage

Fortune 500 organizations that embed security into procurement, prioritize supplier maturity assessments, and integrate threat intelligence into operations are positioned to protect both infrastructure and the applications that drive revenue. Forward-thinking organizations recognize the agentic evolution as an opportunity to align security investments with business outcomes and build operational resilience that allows confident growth.

In a world where adversaries move from initial access to lateral movement in 29 minutes or less and autonomous agents make decisions at machine speed, the organizations that thrive will be those with unified runtime visibility.

Clarifying the competitive narrative

There is a prevailing belief that frontier AI labs—Anthropic with Claude Code Security, OpenAI with Codex5—are making traditional security tools obsolete. This belief is partially correct, but it fundamentally misunderstands which tools are being displaced.

What frontier labs are changing

These capabilities collapse scanning stages in the IDE and CI/CD pipeline. They find vulnerabilities through reasoning rather than pattern matching, uncovering business-logic flaws that static analysis misses. This is genuine progress—and it is expected to reduce certain vulnerability classes over time. SQL injection, for example, may become less prevalent as LLM-generated code matures and shift-left tools are integrated throughout the agent development lifecycle.

What frontier labs don’t address

Frontier lab security tools operate before deployment. They don’t see how code behaves in production—outside of API integrations that push context to them. They can’t detect configuration drift, environment-aware malware that activates only at runtime, or prompt injection against live AI agents. Anthropic describes Claude Code Security as an evolution of static analysis: rather than matching known patterns, it “reads and reasons about your code the way a human security researcher would.”4 Static analysis operates on code before it runs. These tools sit at completely different points in the security lifecycle from runtime protection.

What Dynatrace surfaces today

These capabilities deliver findings that frontier lab tools can’t—evidence of what is actually happening in production, not predictions based on code analysis.

Runtime Vulnerability Analytics

Identifies vulnerabilities in the context of actual execution—not theoretical exposure, but real risk based on how code runs in production.

Runtime Application Protection

Detects and blocks exploit attempts as they happen—the defensive layer that shift-left tools structurally can’t provide.

Security Posture Management

Surfaces misconfigurations and compliance gaps in live cloud infrastructure, including Kubernetes environments.

The evolving threat landscape

The Open Worldwide Application Security Project (OWASP) Top 10 for Agentic Applications (2026) report, provides consensus-based guidance and introduces new categories—Agent Behavior Hijacking, Tool Misuse and Exploitation, and Identity and Privilege Abuse6—directly relevant for organizations deploying autonomous agents. For any organization running business-critical applications, runtime visibility remains the validation layer that confirms whether security controls actually work.

What this means in practice

Organizations using Dynatrace already have the core capabilities required for agentic security.

  • Security findings flow directly to development and SRE teams through workflows that include clear remediation guidance.
  • Natural-language queries on security findings are already available through Dynatrace Intelligence.
  • Dynatrace Intelligence allows agentic remediation, allowing automation to act within goals and constraints defined by humans and for humans to retain the flexibility to stay in the loop (see Agentic workflows in Dynatrace documentation).

What sets the Dynatrace approach apart is the combination of deterministic runtime findings with agentic workflows—enabling remediation that is reliable, repeatable, and grounded in real production evidence.


Article citations

  1. CrowdStrike 2026 Global Threat Report. Average eCrime breakout time (initial access to lateral movement) fell to 29 minutes in 2025, a 65% increase in speed from 2024. Fastest observed: 27 seconds. CrowdStrike also observed a 89% increase in attacks by AI-enabled adversaries compared with 2024. crowdstrike.com/en-us/blog/crowdstrike-2026-global-threat-report-findings
  2. Fortune, “Feds are hunting teenage hacking groups like ‘Scattered Spider’ who have targeted $1 trillion worth of the Fortune 500 since 2022,” January 2026.
  3. Anthropic, “Disrupting the first reported AI-orchestrated cyber espionage campaign,” published November 2025. Attack detected mid-September 2025. AI executed 80–90% of tactical operations independently.
  4. Anthropic, “Claude Code Security,” anthropic.com/news/claude-code-security, 2026.
  5. OpenAI, “Codex,” platform.openai.com/docs/codex.
  6. OWASP GenAI Security Project, “Top 10 for Agentic Applications 2026,” genai.owasp.org, December 2025.
  7. CVE-2024-3094 (XZ Utils). Backdoor activated only when loaded by sshd on targeted Linux distributions — bypassing all static analysis. Wired, April 2024.
  8. Dynatrace Intelligence: “supports in-context natural language for investigation and guided next steps.” docs.dynatrace.com/docs/dynatrace-intelligence

The post The runtime reckoning: How the agentic evolution is reshaping security appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-runtime-reckoning-how-the-agentic-evolution-is-reshaping-security/feed/ 0
Fueling visual insights with MCP applications for complex data analysis https://www.dynatrace.com/news/blog/fueling-visual-insights-with-mcp-applications-for-complex-data-analysis/ https://www.dynatrace.com/news/blog/fueling-visual-insights-with-mcp-applications-for-complex-data-analysis/#respond Thu, 26 Feb 2026 18:17:48 +0000 https://www.dynatrace.com/news/?p=73193 MCP Server AI-Assistants

The Dynatrace platform provides expanded data access via the Dynatrace MCP server and several APIs that tie into agentic integrations from our ecosystem, our IDE plugins, CLIs, and Dynatrace® Apps. The Dynatrace MCP server has been a cornerstone of making this innovation available to our developer audience right in the IDE. However, MCP is heavily […]

The post Fueling visual insights with MCP applications for complex data analysis appeared first on Dynatrace news.

]]>
MCP Server AI-Assistants

The Dynatrace platform provides expanded data access via the Dynatrace MCP server and several APIs that tie into agentic integrations from our ecosystem, our IDE plugins, CLIs, and Dynatrace® Apps. The Dynatrace MCP server has been a cornerstone of making this innovation available to our developer audience right in the IDE. However, MCP is heavily relying on text-based output. This blog introduces new Dynatrace MCP Server application support, which we’re happy to contribute to the community, upgrading UI capabilities by enabling data visualization. It empowers developers and organizations to build their own agentic platforms and visualize experience data, without relying solely on LLM reasoning.

Why are MCP applications so valuable to Developers?

A picture is worth a thousand words

Charts, color cues, and visual tables are here for a reason. We just need a quick look to determine whether something is right or wrong, and with MCP Apps, those visuals can appear in the same conversation with Dynatrace Assist or when interacting with Dynatrace Intelligence from within Kiro, GitHub Copilot, or any other MCP integration, where we ask questions while avoiding context switching and speeding up analysis and result checking. Once we get the insights we need, we can dive even further into the platform applications with comprehensive context.

For example, during a problem investigation, we observe a spike in requests across all endpoints, especially for the v1/trade/long/process endpoint.

DQL chart in Dynatrace

Dynamic visual results

Interactive UIs let users modify results per their needs. With the new visual results, we can filter, pivot, and drill into results without re-issuing prompts; the model and UI remain in sync and continue the conversation with a richer context.

Actionability built in

Visuals can include action buttons, for example, Open in Notebooks, which calls back to the server to drill into the data stored in Dynatrace Grail® unified data lakehouse. You can create dashboards and notebooks directly from any page without changing context.

Human verification

Human involvement in data analysis is one of the factors that differentiate these processes from purely LLM-based solutions. Keeping humans in the loop to validate and interpret results visually is one of the most important factors for organizations that want to maintain control over their data analysis process.

MCP App architecture

MCP App Data Flow architecture

MCP extension protocol integration

The Dynatrace MCP Server uses the @modelcontextprotocol/ext-apps library to register interactive UI applications alongside traditional text-based tool responses. This allows the MCP client (such as GitHub Copilot) to extend the standard MCP results with visual capabilities.

Tool-to-UI binding

Tools are registered via the _meta.ui.resourceUri property that links the tool’s output to its visual counterpart. When execute_dql returns results, the host knows to render the associated UI app as well.

Data flow

Structured output parsing

The tool returns results as structured markdown containing metadata and a JSON block. The UI app parses this response, extracting record counts, warnings, and the actual data records.

Real-time client-side rendering

A lightweight TypeScript-based web-app (execute-dql.ts) receives the tool result via the app.ontoolresult callback and dynamically builds an interactive HTML table—complete with sortable columns, hover states, and expandable JSON cells.

Single-file bundling

When using Vite with vite-plugin-singlefile, the HTML, CSS, and TypeScript are bundled into a self-contained HTML resource that the MCP server serves on demand.

See it in action

Take a look at this example video showing an investigation into slowdowns.

Dynatrace CoPilot demo

In the following screenshots, you can see examples of how data that is structured in a visual way helps us to ask the right questions or continue investigations promptly.

In reviewing error logs, we see frequent errors appearing in easytrade-broker-service.

DQL chart in Dynatrace

While analyzing service degradation due to the slowdown, we found that two services were affected for approximately 30 minutes.

DQL chart in Dynatrace

What about security?

Let’s call out the elephant in the room: security is a major concern with MCP servers, and adding HTML rendering on top is likely to raise eyebrows.

There are a couple of security measures and best practices in place to protect our users:

  1. Dynatrace MCP Apps are always rendered in a sandboxed iframe controlled by the host.
  2. Our MCP apps communicate only with the MCP host via JSON RPC. No direct communication with the internet occurs with Dynatrace MCP apps.
  3. We make use of battle-tested React components, inheriting all security-relevant features of React.

Get started

Dynatrace MCP Apps is already included in the latest version of our local MCP Server, and will soon be added to our remote server as well. If you’re already using the local MCP server, you only need the updated version, which should be automatically available in your Dynatrace environment. If the update doesn’t happen automatically, or if you aren’t already using the local Dynatrace MCP server, install it via the GitHub MCP registry and start experimenting using the provided quickstart.

To better understand how you can uplevel your AI assistants with live production insights from Dynatrace, read our most recent blog post or explore the Dynatrace Agentic AI ecosystem in Dynatrace Hub.

To learn more about MCP Apps and the MCP protocol, visit the Model Context Protocol website.

The post Fueling visual insights with MCP applications for complex data analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/fueling-visual-insights-with-mcp-applications-for-complex-data-analysis/feed/ 0