observability | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 16 Jul 2026 09:08:15 +0000 en hourly 1 Dynatrace named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms for the 16th consecutive time https://www.dynatrace.com/news/blog/2026-gartner-magic-quadrant-observability-platforms/ https://www.dynatrace.com/news/blog/2026-gartner-magic-quadrant-observability-platforms/#respond Wed, 15 Jul 2026 16:45:57 +0000 https://www.dynatrace.com/news/?p=74637 GartnerMQ-2026

​Observability began in a world where software was more predictable. You could usually see what went wrong and where. AI changes that. Now a system can be technically healthy and still produce a bad answer, break a policy, or quietly burn money at scale.​ That shift is forcing observability to evolve quickly. It must account […]

The post Dynatrace named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms for the 16th consecutive time appeared first on Dynatrace news.

]]>
GartnerMQ-2026

​Observability began in a world where software was more predictable. You could usually see what went wrong and where. AI changes that. Now a system can be technically healthy and still produce a bad answer, break a policy, or quietly burn money at scale.​

That shift is forcing observability to evolve quickly. It must account for behavior, judgment, cost, and risk, not just uptime. That’s why we’re proud to share that Dynatrace has been named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms, marking the 16th time Dynatrace has been recognized as a Leader in this report. We think Dynatrace continues to deliver consistent business value for our customers at scale as technology evolves.​

We believe this recognition reflects a simple reality: Teams need a unified view across applications, infrastructure, cloud environments, and AI systems in one place, with the context to get answers—not guesses—about what is happening, why it matters, and what to do next.

The AI era demands end-to-end visibility​

Observability was already under pressure. Cloud-native architectures, distributed systems, and constantly changing environments made it harder to follow cause and effect across the stack. AI raises the stakes.​

AI systems do not fail like traditional software. They can look healthy at the system level while producing low-quality outputs, violating policy, exposing sensitive data, or quietly consuming far more resources than expected. Those problems often sit outside the view of conventional monitoring tools. And when observability is fragmented across separate products, teams lose the context they need to understand how AI behavior connects to infrastructure, applications, and business outcomes.​

Organizations need a complete, connected view from GPU to business outcome. That is what makes it possible to adopt AI safely, operate it confidently, and avoid trading speed for risk.

The Dynatrace difference: A unified platform, powered by AI and built for AI​

Dynatrace approaches observability with a unified platform, not a collection of fragmented tools. Built on Grail®, Smartscape®, and Dynatrace Intelligence — with integrations into the tools, clouds, and AI agents your teams already rely on — the platform brings together a unified data foundation, real-time contextual understanding, and AI-powered intelligence to help teams understand complex systems and act with confidence.​

That matters for two reasons:​

  • Dynatrace is powered by AI. Dynatrace applies deterministic AI to deliver precise, trustworthy answers across applications, infrastructure, and cloud environments. Dynatrace agentic AI can then act on those answers, helping teams move faster, resolve issues sooner, and execute at scale with confidence. ​
  • Dynatrace is built for AI. AI is now becoming part of the software stack itself, and it introduces new failure modes that traditional observability cannot fully explain. Dynatrace gives teams visibility into what their AI is actually doing — not just how the system is performing — with insight into performance, cost, quality, and compliance, all on the same platform and with no additional tooling required.​

Together, these capabilities help organizations move beyond isolated dashboards and alerts. They make it possible to observe, analyze, and automate across modern environments with the context required to keep AI systems reliable, governed, and aligned to business goals.​

We believe this recognition reflects where the market is going

The observability market is changing quickly as organizations invest in AI-powered applications, modernize technology stacks, and look for ways to reduce operational complexity. In this environment, platform depth, unified data, context, and AI matter more than ever.​

​We feel Dynatrace’s continued recognition as a Leader reflects the strength of this approach: A platform that helps customers unify and contextualize data across complex environments, transform it into actionable answers, and support intelligent automation at scale. We think sixteen times as a Leader also speaks to consistent delivery through wave after wave of technology change.​

Read the full Gartner® report​

We’re proud of this recognition, and grateful to the customers, partners, and teams that continue to push observability forward with us.​

Read the 2026 Gartner® Magic Quadrant™ for Observability Platforms to learn more about why Dynatrace was recognized as a Leader and how we think the category continues to evolve in the AI era.

Access the 2026 Gartner® Magic Quadrant™ for Observability Platforms report.

FAQ

What does it mean that Dynatrace was named a Leader in the Gartner® Magic Quadrant™ for Observability Platforms?

The Gartner Magic Quadrant evaluates vendors based on Completeness of Vision and Ability to Execute. We believe Dynatrace’s position as a Leader reflects the strength of our unified observability platform and our ability to help customers manage modern complexity at scale in the AI-era.

Why does observability need to change in the AI era?

AI introduces new kinds of operational risk. A system can appear healthy while still producing poor outputs, violating guardrails, or increasing cost. Teams need observability that can connect AI behavior to the rest of the environment and provide context across performance, cost, quality, and compliance. 

What makes the Dynatrace approach different?

Dynatrace combines a unified data foundation, real-time topology and context, and AI-powered intelligence in one platform. That lets teams move from fragmented signals to precise answers and intelligent action, without relying on separate tools to understand AI systems. Dynatrace offers a combination of:

– A unified platform with Grail® lakehouse, exabyte-scale data foundation
– Business-aware insights with Smartscape® real-time topology and contextual understanding
– Answers, not guesses and governed agentic automation with Dynatrace Intelligence

Gartner Disclaimer

Gartner, Magic Quadrant for Observability Platforms, Padraig Byrne, Martin Caren, D.B. Cummings, Neil Young, 13 July 2026

Gartner does not endorse any vendor, product or service depicted in its research publications and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner research publications consist of the opinions of Gartner’s Research & Advisory organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this research, including any warranties of merchantability or fitness for a particular purpose.

Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates

Dynatrace was recognized as Compuware from 2010-2014.

The post Dynatrace named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms for the 16th consecutive time appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/2026-gartner-magic-quadrant-observability-platforms/feed/ 0
Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/ https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/#respond Thu, 02 Jul 2026 23:42:51 +0000 https://www.dynatrace.com/news/?p=74651 NVIDIA and Dynatrace

Enterprise AI is rapidly evolving from standalone models to agentic AI systems, where multiple AI agents collaborate to gather information, reason across data sources, and generate complex outputs. These systems unlock powerful new capabilities, but they also introduce significant operational challenges. Organizations must be able to observe, govern, and optimize AI agents, models, and infrastructure in real […]

The post Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q appeared first on Dynatrace news.

]]>
NVIDIA and Dynatrace

Enterprise AI is rapidly evolving from standalone models to agentic AI systems, where multiple AI agents collaborate to gather information, reason across data sources, and generate complex outputs. These systems unlock powerful new capabilities, but they also introduce significant operational challenges. Organizations must be able to observe, govern, and optimize AI agents, models, and infrastructure in real time.

Dynatrace helps support this need by providing broad visibility across key layers of the AI stack—from agent orchestration and model inference to GPU infrastructure and enterprise applications. With Dynatrace, teams can monitor AI workflows, understand model behavior, optimize costs, and help improve reliability as agentic systems scale.

Why agentic AI needs full-stack observability

As organizations build GPU-accelerated platforms for AI training and inference, understanding system behavior becomes increasingly complex, with bottlenecks potentially occurring anywhere – from GPU utilization, model latency, token consumption, and downstream service dependencies.

Dynatrace connects these layers through full-stack AI observability, designed to help teams monitor model performance, trace multi-agent workflows, track GPU and infrastructure utilization, detect bottlenecks across AI pipelines, and potentially accelerate troubleshooting with AI-powered root cause analysis.

This unified visibility helps organizations run AI workloads with the same reliability, efficiency, and operational confidence expected from modern enterprise systems.

This unified visibility helps organizations operate AI workloads with improved visibility and operational confidence. By integrating with NVIDIA AI–Q Blueprint and the NVIDIA Agent Toolkit, Dynatrace enriches agent reasoning with high-quality operational telemetry while at the same time helping teams govern and identify opportunities to optimize costs.

How Dynatrace addresses Agentic AI

Dynatrace is designed to assist your team with monitoring infrastructure usage and model behavior and detecting pipeline bottlenecks and token consumption while improving reliability by accelerating troubleshooting and root cause analysis. It also provides a unified view of AI workflows from agent to model down to the infrastructure, allowing organizations to support responsible AI operations, manage cost, improve performance and support agentic workflows at scale.

Every agentic deployment is customized with different agents, tools, models, and data pipelines; therefore, observability is an important capability for understanding how these systems behave in production. The complexity arises as agents interact with multiple enterprise data sources, including:

  • internal datasets
  • external web and knowledge repositories
  • proprietary research systems
  • models served through NVIDIA NIM and Nemotron

Dynatrace can serve as operational data source for AI agents that may help improve the quality of generated insights and enable more informed decision-making. With flexible integration across customized AI-Q implementations, this architecture also lays out the groundwork for automated analysis, research, and decision making.

How Dynatrace integrates NVIDIA AI-Q

By combining NVIDIA’s AI-Q Blueprint with Dynatrace AI observability, organizations gain the transparency and operational intelligence needed to govern, optimize, and scale complex AI systems.

Dynatrace integrates into AI-Q environments in two ways.

1. Observability and cost intelligence for Agentic AI workflows

The NVIDIA Agent Toolkit generates lightweight OpenTelemetry traces that Dynatrace ingests to visualize agent workflows and model interactions.

Dynatrace automatically maps the underlying infrastructure supporting AIQ deployments including NVIDIA NIM and Nemotron microservices and enriches telemetry with AI-specific signals such as:

  • token usage
  • inference latency
  • model metadata
  • GPU utilization

This provides comprehensive visibility across key components including:

  • AI models and inference workloads
  • agent orchestration pipelines
  • GPU and infrastructure resources
  • enterprise data interactions

With these insights, teams can quickly detect performance bottlenecks across agent pipelines, monitor GPU utilization and overall infrastructure health, and identify inefficient model usage. This visibility can help organizations identify cost optimization opportunities associated with AI workloads. Together, these capabilities position observability as important components for building reliable and scalable AI systems.

2. Dynatrace as a high-quality data source for AI agents

Dynatrace can also serve as an operational intelligence source for AI agents.

Through Model Context Protocol (MCP) integrations, Dynatrace exposes telemetry that agents can use in their reasoning workflows, including:

  • infrastructure performance metrics
  • operational incidents and problems
  • deployment and reliability trends
  • system behavior and resource consumption

This allows AI agents to incorporate real-time operational insights into their decision-making. Instead of relying solely on external data, agents gain contextual awareness of enterprise systems, which may support more informed outputs Dynatrace ingests NVIDIA Agent Toolkit OpenTelemetry traces, model telemetry, and infra metrics exposing operational context via MCP.

Together, these technologies create a powerful foundation for deploying deep research in the enterprise as reflected in the picture below.

Dynatrace AI Observability - NVIDIA
Figure 1: Dynatrace providing AI Observability for NVIDIA AI-Q

AI-Q use cases

The following are illustrative examples of what becomes possible when AI-Q-based research agents incorporate Dynatrace operational data and insights into their reasoning workflows. While NVIDIA AI-Q is a reference framework rather than a formal certified Dynatrace integration, these scenarios show how agentic research systems could use Dynatrace AI observability to generate richer analysis, identify patterns, and support more informed decisions.

Infrastructure migration analysis

AI agents combine Dynatrace operational telemetry such as performance trends, incidents, and deployment velocity with infrastructure and cloud cost data to evaluate platform migration scenarios (for example, OpenShift to AKS). The system produces data-driven recommendations with quantified tradeoffs to support strategic decisions.

Large-scale incident analysis

By analyzing thousands of historical problems, AI agents can identify recurring patterns, understand infrastructure behavior, and correlate technical issues with business KPIs. This enables deep operational insights and long-form analysis that would be difficult and time-consuming for humans to produce.

AI cost governance and optimization

Enterprises can use observability data from Dynatrace to analyze token consumption, model usage, and inefficient data interactions across AI workloads. Agents can identify patterns and suggest potential optimizations such as more efficient models or improved workflows.

Software delivery and reliability insights

DevOps and SRE teams can use agentic analysis to correlate deployments with incidents, assess build quality trends, forecast reliability risks, and identify engineering priorities—using Dynatrace as the trusted operational data source.

Get started today

The post Smarter, safer Agentic AI: Dynatrace observability meets NVIDIA AI-Q appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-observability-meets-nvidia-ai-q/feed/ 0
Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/ https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/#respond Mon, 08 Jun 2026 15:03:09 +0000 https://www.dynatrace.com/news/?p=74422 Blog OTP Observability for Agentic AI

Agentic AI is breaking the mold of what organizations need from observability. Fragmented, correlation-dependent observability platforms are no longer “good enough.” Enterprises with dynamic, hybrid environments require observability that provides real-time, precise answers, so AI agents can prevent problems, automate workflows, and deliver better, more secure software.

The post Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI appeared first on Dynatrace news.

]]>
Blog OTP Observability for Agentic AI

As more agentic AI projects come online, the observability market is abuzz with familiar promises: tool consolidation, AI-powered insights, and faster remediation through smarter tools. On the surface, this sounds like progress. But beneath the excitement, many discussions are framed around the wrong question.

The real issue isn’t about how to adopt autonomous operations; it’s about ensuring AI agents are operating reliably and resolving problems without introducing new ones. When evaluating new observability solutions, the question should be:

Can this observability solution accurately analyze complex, dynamic telemetry in context so AI agents can act autonomously with trust, precision, and reliability?

As systems become increasingly agent driven, observability is crossing a structural boundary. Approaches designed for environments where only humans decide and act must adapt to a world where agents increasingly operate autonomously with human oversight, while keeping organizations informed.

Rethinking observability for the agentic age

Observability platforms were initially intended to support engineers in delivering reliable applications, services, and infrastructure to users, and alert them in the event of a problem. Dashboards, alerts, and correlation helped teams investigate incidents, piece together what happened, diagnose issues, decide on next steps, and resolve the problem. This model worked when changes were pushed manually.

The assumption was that more data, better correlation, and cleaner interfaces will lead to increased visibility and improved operational decision making.

Agentic AI systems break that assumption.

With faster release cycles and AI-generated code, manual investigations can no longer keep pace. Moreover, observability platforms must now provide actionable insights to both humans and AI agents.

As agents begin operating as autonomous participants in software environments by triggering mitigations, scaling infrastructure, and optimizing behavior in real time, observability can no longer function solely as a human interface. It must also provide AI systems with a reliable, contextual fact basis that agents can act on programmatically. Machines can’t rely on dashboards and alerts. They require a deterministic foundation of unified, real-time data that delivers accurate, context-rich answers at exabyte scale.

Agentic systems break the mold of “good enough”

Many observability platforms layer probabilistic AI on top of siloed data. They use LLMs to correlate signals and rank likely causes—but they can’t always determine correctness.

“Probabilistic” means that the same input will generate a different output based on a probability distribution of predefined outputs, delivering a different answer when the same problem occurs. This approach is also prone to hallucinations, requiring additional human validation, which can increase operational overhead and token costs, delay resolution of business-critical issues, and divert resources from strategic initiatives.

Enterprise-grade observability must now answer: Is this insight reliable enough for autonomous action?

AI built on siloed data is inherently unreliable. Autonomous systems depend on deterministic, contextual, and trustworthy data to act reliably.

“Deterministic” means that the same input always results in the same output by using factual data to trace the exact causal changes that created the issue. When agentic AI systems act on business-critical applications, the cost of being “mostly right” becomes operationally unacceptable.

This is where a subtle but critical divide appears in the market. Aggregating signals and correlating anomalies can surface patterns. Patterns alone are not a solid basis for decisions, and without deterministic understanding, AI systems inherit that uncertainty and can propagate it downstream.

To drive reliable enterprise autonomous operations, AI agents require a unified, AI-powered observability platform that can analyze exabytes of data in real time and across models to pinpoint root cause, delivering actionable answers in context of what’s affected and its business impact.

From correlated guesses to deterministic answers

This shift in the demands of observability hinges on a clear distinction:

  • Probabilistic AI correlates signals that happened around the same time and therefore appear related, pulling information from fragmented data stores to propose a likely root cause.
  • Deterministic AI uses causal analysis to pinpoint what happened and why, recommend remediation actions, and identify business impact.

Probabilistic AI is intended to narrow the search space and direct engineers toward potential resolution, but it still requires interpretation.

Deterministic AI establishes sequence, dependency, and impact, enabling systems to decide safely without waiting for humans to connect the dots.

Auto‑remediation, auto-prevention, and auto-optimization all depend on this leap. A platform that unifies telemetry only at the UI layer may deliver data and potential root cause, but it can’t compensate for fragmented understanding and missing context underneath. When context is pieced together after the fact, confidence is never guaranteed.

You can’t automate what you don’t precisely understand.

Context driven observability as the control plane for AI

In an autonomous enterprise, observability doesn’t sit beside execution; it’s embedded within it. This integration requires that teams adopt a new mindset toward observability architecture.

Because more AI workloads are happening at the source, telemetry must be optimized and streamlined before ingest, not after the fact, from the edge to the back end. Data access must be unified, context-aware, and always-hydrated on a massive scale. Answers must be explicit, not implicit, and they must be informed by automatic, real-time dependency mapping.

Likewise, intelligence must combine deterministic and agentic AI—not as add‑ons, but as a single reasoning system from ingest to execution.

In this model:

  • AI agents can become the primary consumers of observability data.
  • Humans can shift toward strategy, architecture, oversight, and exception handling.
  • Observability evolves from a reactive lens into a control plane for autonomous operations.

Observability purpose-built for autonomous operations ensures successful agentic AI initiatives

This moment represents an architectural transition, not just an incremental upgrade cycle. Correlation-dependent observability that uses probabilistic AI can be extended, augmented, and rebranded, but it will always carry the limitations of approximation and human validation.

The next era belongs to an observability platform that’s built for machine understanding from the start: a unified, context driven architecture that delivers deterministic answers at machine speed, precision, and scale.

Do you want more data or better decisions? Learn why enterprises are switching to Dynatrace.

The post Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/feed/ 0
The rise of business observability https://www.dynatrace.com/news/blog/the-rise-of-business-observability/ https://www.dynatrace.com/news/blog/the-rise-of-business-observability/#respond Sun, 07 Jun 2026 14:10:28 +0000 https://www.dynatrace.com/news/?p=73757 Business Process analytics

Every organization runs on data. But for most leaders, the challenge is clarity. Traditional dashboards and reports often arrive too late, and system-level metrics don’t explain the business consequences of technical events. When a payment API slows by 200 milliseconds, what does that mean for daily revenue? When user sessions drop, is it a performance […]

The post The rise of business observability appeared first on Dynatrace news.

]]>
Business Process analytics

Every organization runs on data. But for most leaders, the challenge is clarity. Traditional dashboards and reports often arrive too late, and system-level metrics don’t explain the business consequences of technical events.

When a payment API slows by 200 milliseconds, what does that mean for daily revenue? When user sessions drop, is it a performance issue or a recent pricing change? These are the questions business observability is designed to answer.

Business observability connects technical telemetry to business outcomes, creating a shared understanding of how digital performance drives, or hinders, organizational results. It’s less about adding more data, and more about connecting the right data in real time to the decisions that matter.

Importantly, business observability is broader than observing individual business processes. While end‑to‑end workflows like “order to fulfillment” or “claim to payout” are one expression of business observability, the discipline also encompasses digital experience, customer behavior, revenue impact, operational efficiency, and risk. Business observability focuses on understanding how technology performance influences business outcomes across the organization and not just how a single process executes.

Moving beyond traditional monitoring

Traditional monitoring focuses on uptime, latency, and error rates. These are essential for engineers but often disconnected from business context. Business observability closes that gap by linking system health directly to business performance.

That connection – the ability to translate technical telemetry into business insight – is what distinguishes business observability from conventional monitoring.

What makes business observability different

Business observability builds on traditional monitoring by expanding its scope from system health to business outcomes. It introduces three essential capabilities that operate across customer experience, revenue‑generating transactions, and end‑to‑end processes:

  • Business context integration: IT metrics are linked with business KPIs, translating system behavior into measurable impact.
  • End-to-end business process visibility: Processes like “loan approval to disbursement,” “claim to payout,” or “application to onboarding” are monitored across applications, infrastructure, and external systems.
  • Real-time decision support: Business observability surfaces business implications in real time, allowing teams to respond to issues before they escalate.

Together, these principles give organizations a live understanding of how every digital interaction affects outcomes, from conversions and revenue to efficiency and satisfaction.

Where business observability gets its data

Business observability relies on unifying multiple data streams into a single, contextual view of performance:

  • Business events: Structured, contextualized data that represents key business outcomes, such as completed checkouts, submitted claims, bookings, or failed transactions.
  • Logs: Time-stamped records of system activity that show what happened and when, often containing critical technical and business data.
  • Real user monitoring (RUM): Continuous insight into how users actually experience applications and digital services across devices, channels, and geographies.
  • External business tools: Data from tools such as ERP, CRM, billing, and payment platforms that provide commercial, operational, and customer context.
  • Instrumentation and agents: Technologies that collect telemetry and business-relevant data from applications and services, ranging from manual instrumentation to automated approaches that observe activity as it flows through systems with minimal configuration effort.
  • OpenTelemetry: A standardized instrumentation framework that enables consistent collection of metrics, logs, and traces across hybrid and cloud-native environments.

When correlated, these data sources bridge the gap between technical operations and business results, providing a comprehensive view of how technology supports, or disrupts, performance.

How organizations put business observability to work

Modern organizations apply business observability in three key ways. Organizations may pursue these use cases through a variety of approaches, ranging from manually instrumented metrics and custom reporting to more automated, integrated platforms that reduce effort and time to insight.

  1. Drive real-time decisions with IT context: When issues arise, teams and business leaders can see the business impact immediately. A sudden dip in conversions, for example, can be traced to a misconfigured API or third-party service outage, enabling immediate, targeted response.
  2. Track and optimize business processes: Complex workflows—like “order to fulfillment” or “quote to claim”—are mapped from end to end. Teams can pinpoint where time, cost, or customer satisfaction are being lost and address bottlenecks before they affect outcomes.
  3. Accelerate sustainability and reduce costs: By correlating business events with resource usage, organizations can identify inefficiencies in automation, cloud consumption, or scheduling that inflate cost and carbon impact.

Business observability, in action

Each of these examples shows how business observability does more than surface anomalies. It provides the context to act on them. By connecting business events, telemetry, and user experience data, organizations move from simply detecting issues to understanding their impact, cause, and resolution path in real time.

In practice, achieving this level of insight often depends on how business and technical data are captured and correlated. While some organizations rely on custom instrumentation, manual analysis, or post‑incident reporting, others use more automated approaches that make it possible to detect impact, trace root cause, and act in real time.

Financial services

A large financial institution monitored loan application volume as a business event rather than relying solely on system health metrics. When completed applications began declining, traditional monitoring showed no infrastructure failures. Business observability revealed that timeouts from a third-party credit scoring service were affecting only new applicants following a recent integration change. By correlating business events with external service performance, the bank isolated the issue quickly and rolled back the configuration—preventing lost loan volume and downstream compliance exposure.

Payments

A global payment services provider processes billions of transactions annually and needs to understand not just whether systems are available, but how performance impacts transaction success and revenue. By connecting transaction latency and failure rates directly to payment outcomes, the organization can quantify the financial impact of technical issues in real time. This shared visibility allows both internal teams and customers to see how payment flows are performing and proactively optimize transaction speed and reliability.

Travel and hospitality

A large travel platform aggregates booking data across partners, regions, and channels. Business observability enables teams to track quotes, bookings, and completed reservations as business events, segmented by partner and geography. When booking volume dips, teams can immediately determine whether the cause is a partner integration issue, a regional performance problem, or a downstream service slowdown—allowing rapid response to protect revenue across the ecosystem.

Retail and consumer services

A national restaurant chain observed high abandonment rates during online reservation flows. Rather than treating this as a generic user experience issue, business observability correlated real user behavior with backend availability and booking outcomes. This insight enabled automated recovery workflows that re-engaged customers who abandoned reservations due to technical or availability issues—recovering lost bookings without manual intervention.

Aviation and transportation

An international airline struggled with fragmented visibility across booking, pricing, and fulfillment systems. By modeling bookings as end-to-end business processes, business observability provided real-time insight into how technical issues affected customer bookings and operational teams. This shared context improved collaboration between IT, call centers, and operations, reducing response times and improving passenger experience during disruptions.

Insurance and regulated industries

An insurer tracked claims submissions as business events and noticed rising exception rates on mobile channels. Business observability revealed that document uploads from newer devices exceeded a backend file-size limit introduced during a recent update. Because business events, logs, and user session data were correlated in real time, teams deployed a same-day fix and notified affected customers—preventing claim backlogs and demonstrating operational transparency to regulators.

Manufacturing and public sector

Organizations running complex, multi-step production or licensing processes use business observability to monitor each step as it moves across applications, infrastructure, and external systems. In manufacturing and government services alike, correlating process steps with technical events allows teams to identify bottlenecks that delay outcomes—such as throttled payment validation or downstream capacity limits—and resolve them without disrupting customer- or citizen-facing services.

Enabling data-driven executive insight

For executives, business observability shifts technology from a cost center to a source of executive intelligence, providing leaders with real‑time visibility into revenue, customer experience, operational performance, and risk. It enables:

  • Real-time business health monitoring: Live dashboards show key performance indicators such as order volume, claim processing times, or fulfillment success rates.
  • Anomaly detection: Advanced analytics identify deviations in both technical and business metrics before they escalate.
  • Impact analysis: Teams can immediately quantify how issues affect revenue, engagement, or satisfaction, and prioritize based on real business value.
  • Trend analysis: Historical and real-time data combine to forecast outcomes and guide strategic decisions.

Improving collaboration between business and IT

One of the most transformative aspect of business observability is the way it unifies language and priorities across teams.

  • IT teams can prioritize work by business impact instead of technical urgency.
  • Business stakeholders gain visibility into technical dependencies that influence performance.
  • Cross-functional teams align around shared outcomes rather than isolated metrics.

This shared context strengthens trust and speeds response, particularly during digital transformation, where both technology performance and customer experience are constantly evolving.

Laying the groundwork for business observability

While business observability may be expressed through dashboards, process views, or executive metrics, its success depends less on how data is visualized and more on how consistently business outcomes are connected to technical signals across teams. Adopting business observability effectively requires thoughtful preparation:

  • Data integration: Pull from diverse sources – applications, infrastructure, and business systems – to form a cohesive view.
  • Shared metrics: Define how business KPIs map to technical signals.
  • Organizational alignment: Ensure teams are trained and incentivized to act on shared insights.
  • Platform scalability: Choose an observability solution that supports hybrid, cloud, and partner ecosystems as data volume grows.

The road ahead

Business observability represents a shift from reactive monitoring to outcome-driven intelligence. Rather than focusing solely on system health or individual process performance, it enables organizations to understand in real time how technology influences revenue, customer experience, operational efficiency, and strategic decision‑making. For executives, it means real-time visibility into how technology influences outcomes. For IT teams, it means prioritizing based on business value. For organizations, it means a unified, data-driven way to make decisions with confidence.

In practice, many organizations attempt to reach this level of insight through manually instrumented metrics, custom dashboards, and offline analysis – approaches that require ongoing effort and often delay understanding when it matters most. Dynatrace Business Observability takes a different approach, capturing business events from multiple sources and automatically correlating them in real time with full‑stack telemetry. This delivers the context leaders need to act decisively, without the manual overhead of traditional approaches.

Take Dynatrace for a spin

FAQs: Business observability

What is business observability in simple terms?

Business observability is the ability to understand how technical performance affects business outcomes in real time. It connects IT telemetry, such as logs, traces, and metrics, with business data such as transactions, claims, or orders, so teams can see both the cause and the consequence of an issue in one view.

How is business observability different from traditional monitoring?

Traditional monitoring reports on system health (uptime, latency, or resource use) without showing the business impact. Business observability goes further by linking these technical metrics with key performance indicators (KPIs), providing the context to understand how technical changes influence revenue, customer experience, and operational efficiency.

What types of data does business observability rely on?

Business observability combines multiple data sources, including:

  • Business data from transactions or processes
  • Logs and traces from applications and infrastructure
  • Real user monitoring (RUM) for end-user experience
  • Data from external business tools like CRM, ERP, or payment systems
  • Instrumentation frameworks such as OneAgent and OpenTelemetry for consistent telemetry across environments

Who benefits most from business observability?

Executives gain real-time visibility into how technology affects performance and outcomes. IT and engineering teams gain business context to prioritize fixes based on impact. Together, these perspectives drive faster decision-making and closer alignment between technology operations and business goals.

What are common use cases for business observability?

Typical use cases include:

  • Detecting and resolving process slowdowns in finance, healthcare, or logistics workflows
  • Tracking and optimizing customer journeys and digital transactions
  • Measuring the business impact of new releases or integrations
  • Identifying inefficiencies that increase costs or carbon footprint

How does business observability support sustainability and cost efficiency?

By correlating business events with resource consumption, business observability helps identify where automation, compute, or storage are overused. This enables teams to optimize cloud spend, reduce energy consumption, and track sustainability metrics alongside operational performance.

What challenges do organizations face when implementing business observability?

Key challenges include integrating data from multiple systems, defining the right shared KPIs between business and IT, and ensuring teams are trained to act on insights collaboratively. Success depends on cross-functional alignment as much as on technology.

Is business observability only relevant for large enterprises?

No. Any organization that relies on digital processes – from mid-sized financial firms to healthcare networks or government agencies – can benefit. The ability to link technical performance to business results is valuable at any scale

The post The rise of business observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-rise-of-business-observability/feed/ 0
OpenTelemetry graduates: A milestone for the observability Open Source community https://www.dynatrace.com/news/blog/opentelemetry-graduates-a-milestone-for-the-observability-open-source-community/ https://www.dynatrace.com/news/blog/opentelemetry-graduates-a-milestone-for-the-observability-open-source-community/#respond Fri, 29 May 2026 17:13:57 +0000 https://www.dynatrace.com/news/?p=74222 OpenTelemetry logo icon

In May 2026, OpenTelemetry (OTel) officially graduated from Cloud Native Computing Foundation (CNCF). This milestone marks the cloud native ecosystem’s achievement of production readiness and maturity, thanks to the efforts of hundreds of companies and thousands of developers who believed in an open standard for observability and built it together. OpenTelemetry was already the standard. […]

The post OpenTelemetry graduates: A milestone for the observability Open Source community appeared first on Dynatrace news.

]]>
OpenTelemetry logo icon

In May 2026, OpenTelemetry (OTel) officially graduated from Cloud Native Computing Foundation (CNCF). This milestone marks the cloud native ecosystem’s achievement of production readiness and maturity, thanks to the efforts of hundreds of companies and thousands of developers who believed in an open standard for observability and built it together.

OpenTelemetry was already the standard. Now it’s official.

If you’ve been shipping production code over the last few years, OpenTelemetry has almost certainly touched your tech stack, whether through its SDKs and collectors, or traces, metrics, and logs. For many teams, OTel has quietly become part of the default toolbox for building and operating modern applications. It solves a real problem: the industry needed a common language for telemetry data, and OpenTelemetry became that language.

CNCF graduation reflects the strength of the ecosystem: a diverse contributor base, widespread vendor support with proven production readiness, comprehensive security audits and a governance model built by the community, for the community and for future sustainability.

Why a shared standard changes everything

Graduation formalizes OpenTelemetry as the common protocol and shared language for observability:

  1. A standard protocol allows different tools, open source and commercial, to work together.
  2. Semantic conventions define how telemetry is named and structured. When all systems speak the same language, correlation becomes possible at scale.
  3. OpenTelemetry decouples instrumentation from backend analytics, enabling teams to export telemetry data to any backend system and switch analytics platforms without rewriting code.
  4. Standardized, high-quality telemetry data allows automation, anomaly detection, and AI-driven insights. A consistent protocol becomes critical as systems grow more complex and autonomous.

Dynatrace loves OpenTelemetry and open source

Dynatrace has been involved in shaping OpenTelemetry from its early days, contributing to the specification, semantic conventions, Collector, and many other areas, ensuring the standard works at enterprise scale. With over 46,000 contributions and 54,000 commits, Dynatrace is one of the top contributors to the project.

Our focus has always been clear: make OpenTelemetry production-ready without compromising its open, vendor-neutral model.

Beyond OpenTelemetry, Dynatrace actively contributes to over 30 open source projects, including W3C Trace Context, and integrations with Kubernetes, JMeter, and more.

Frequently asked questions

Is OpenTelemetry stable after CNCF graduation?

Yes. Graduation confirms that the core specifications, APIs, and data model are stable and suitable for long-term production use. Teams can adopt OpenTelemetry with confidence across environments and use cases.

Does OpenTelemetry lock teams into a specific vendor or backend?

No. OpenTelemetry decouples instrumentation from backend analytics. Teams can export telemetry data to any backend and switch platforms without rewriting instrumentation code.

How does Dynatrace support OpenTelemetry?

Dynatrace supports you wherever you are. You can ingest pure OpenTelemetry data natively; no proprietary agents are required. From there, you can query billions of spans in seconds, correlate metrics, logs, and traces automatically across petabytes of data, and get the full observability context.

Want to see OpenTelemetry in action?

Learn more about OpenTelemetry at Dynatrace Hub.

The post OpenTelemetry graduates: A milestone for the observability Open Source community appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/opentelemetry-graduates-a-milestone-for-the-observability-open-source-community/feed/ 0
Observability is a team sport https://www.dynatrace.com/news/blog/observability-is-a-team-sport/ https://www.dynatrace.com/news/blog/observability-is-a-team-sport/#respond Wed, 29 Apr 2026 16:23:20 +0000 https://www.dynatrace.com/news/?p=73858 Dynatrace and OpenTelemetry

“How do I structure my observability team?” is one of the most common questions folks leading software teams ask me. My advice: Don’t create a centralized “observability team” that’s responsible for all the observability within an organization. Observability shouldn’t exist as a silo. It touches many parts of an organization, from development to production, and […]

The post Observability is a team sport appeared first on Dynatrace news.

]]>
Dynatrace and OpenTelemetry

“How do I structure my observability team?” is one of the most common questions folks leading software teams ask me. My advice: Don’t create a centralized “observability team” that’s responsible for all the observability within an organization.

Observability shouldn’t exist as a silo. It touches many parts of an organization, from development to production, and should be treated as a team sport.

As we know, our systems can only be considered observable if they emit telemetry. No data means that we can’t understand what is happening in our systems. Fortunately, the OpenTelemetry® (OTel) ecosystem from the Cloud Native Computing Foundation (CNCF) has become the de-facto standard for instrumenting, generating, collecting, and exporting telemetry data.

What does this mean for observability adoption in an organization? Let’s dig in.

Observability is everyone’s responsibility

Reliability can’t happen without observability. Observability must be looked at holistically. It is not the sole responsibility of any one team or individual. Everyone has an important part to play, and to a certain extent, the parts weave into each other.

Instrumenting code

There are two types of OpenTelemetry instrumentation:

Code-based instrumentation should be done by application developers, and not by an “observability team.” Developers know their applications best. Asking someone else to instrument your application is like asking someone else to write your code comments. Please never do that.

Zero-code instrumentation usually involves a shim or bytecode instrumentation wrapper around your code. If you’re a developer writing code in a language that supports OpenTelematry auto-instrumentation, you should understand how to implement both zero-code and code-based instrumentation. In doing so, you can use the instrumentation to troubleshoot your own code.

In some environments, zero-code instrumentation may be managed by the OTel Operator. If this is the case, the responsibility often falls to SRE or platform engineering teams. Event in those cases, developers should understand; at least at a high level, how zero-code instrumentation is configured with the OTel Operator.

Managing observability infrastructure

Observability infrastructure still needs to be managed, whether you’re using a SaaS vendor (e.g. Dynatrace) or an open source stack. If you’re using OpenTelemetry, chances are you’re managing at least one OTel Collector, and perhaps many. If you’re running your applications on Kubernetes, you’ll likely deploy and manage Collectors within the cluster as well. In most organizations, this responsibility falls under platform engineering or SRE teams, and these teams are essential to robust, reliable software delivery in large, complex environments.

That said, developers should still understand how the OpenTelemetry Collector is configured. It’s true that you don’t need to go through a Collector to send OTel data to an observability backend for non-production. However, the Collector still offers some nice things that direct-from-application doesn’t (e.g. batching data, masking data, and automatic retries), and I still highly recommend using it, even in development.

Making CI/CD pipelines observable

DevOps engineers can’t escape observability either, because guess what? We can make CI/CD pipelines observable too. While CI/CD pipelines may not be a production environment that external users interact with, they most certainly are a production environment that internal users interact with (i.e. software engineers, platform engineers, and SREs).

CI/CD pipelines are defined by code, and like it or not, that code can still fail. Making our application code observable helps us make sense of things when they fail in production. So, it stands to reason that having pipeline observability can help us understand what’s going on when CI/CD pipelines fail.

There’s been some great buzz around the observability of CI/CD pipelines, especially now that there’s an official OTel CI/CD Special Interest Group (SIG). This will give our favorite CI/CD tools a shared language for the observability of CI/CD pipelines, creating a foundation for them to support OpenTelemetry tools in this context.

We’re not there yet, which means that right now we must stitch a few tools together to achieve CI/CD observability. Fortunately, things are moving nicely in this space, and if you haven’t considered CI/CD pipeline observability in your organization before, now’s the time to start thinking about it. To learn more about what’s happening with OTel CI/CD observability, check out the #otel-cicd channel on CNCF Slack.

Troubleshooting

The beauty of observability is that once you instrument your code, you put the ability to troubleshoot in the hands of many. Consider the ripple effect when developers instrument their code:

  • Developers: Instrumentation allows developers to debug their code as they’re writing it.
  • QA testers: Instrumentation allows testers to troubleshoot failed tests, allowing them to file more detailed bug reports. If QAs can’t track down the issue, then it means that there is missing instrumentation that developers need to add to their code. This turns observability into a quality gate.
  • SREs: Instrumentation allows SREs to troubleshoot production issues, gain insight into system performance, and ensure overall system reliability.

Ensuring adherence to observability practices

Remember how I advised against creating an “observability team” responsible for all observability within an organization? I still stand by that. That said, I do believe that organizations should have an observability team responsible for enterprise-wide observability oversight and advocacy. A team that defines and disseminates observability standards and practices within that organization. This team would need to stay up to date in the latest observability practices, vendor offerings, and the OpenTelemetry  ecosystem— not just as an observer, but also as a project contributor, while also encouraging developers, platform engineers, and SREs to contribute.

This “observability practices team,” can’t, however, exist on an island. First off, it needs to be aligned with leadership to ensure that everyone is on the same page when it comes to observability. The team also needs support from individual practitioners. As a result, the team also needs to work with developers, SREs, platform engineers, QAs, and DevOps engineers to ensure that the practices and standards that it comes up with make sense.

If observability is to be a team sport, it needs coordination and guidance. There should be guardrails in place, to ensure that you have standard tooling, practices, and enforcement of said practices. Practices and standards include things like standard Collector configurations, and standard attributes emitted to your chosen observability backend(s).

Standardizing tooling is important because I’ve seen far too many “tool jungles” in organizations, where each team or department has their own tooling and practices, and it ends up being a recipe for disaster. Too much redundancy and overlap.

In addition, the observability practices team should not be responsible for instrumenting developers’ code, nor should it be managing infrastructure. It’s there to work with these other groups and to make sure that things are done right.

Final thoughts

Observability weaves its way into various aspects of an organization. It’s not just a developer concern. It’s not just an SRE concern. It’s not just a QA concern. It’s certainly not the concern of a single “observability team.” Doing so downplays its importance, takes away our collective responsibility towards observability, and dilutes the promise of observability. The only way to make this work is by ensuring that the teams participating in this team sport that we call observability don’t operate in silos.

The post Observability is a team sport appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/observability-is-a-team-sport/feed/ 0
Pipeline Groups in Dynatrace OpenPipeline: Enterprise-grade governance explained https://www.dynatrace.com/news/blog/pipeline-groups-in-dynatrace-openpipeline-enterprise-grade-governance-explained/ https://www.dynatrace.com/news/blog/pipeline-groups-in-dynatrace-openpipeline-enterprise-grade-governance-explained/#respond Wed, 08 Apr 2026 23:11:32 +0000 https://www.dynatrace.com/news/?p=73675 OpenPipeline logo

As organizations scale their observability practice, a familiar tension emerges: platform engineering teams need to enforce consistent configuration for pipeline ingestion, such as security context, cost allocation, and compliance-driven routing, while application teams need the freedom to tailor data processing for their specific workloads.

The post Pipeline Groups in Dynatrace OpenPipeline: Enterprise-grade governance explained appeared first on Dynatrace news.

]]>
OpenPipeline logo

In practice, most organizations attempt to solve this with one of two compromises: lock everything down centrally and become a bottleneck for every pipeline change, or copy rules across pipelines and hope nothing drifts. Neither approach scales. And as data requirements shift faster than ever to address new services, new regulatory mandates, and new object types to observe, the cost of such compromise keeps growing.

Pipeline Groups change that dynamic. Generally available now in Dynatrace SaaS version 1.332, Pipeline Groups let platform engineering teams, the teams with admin permissions over the observability platform’s configuration, mandate and standardize ingest behavior across many pipelines, while safely delegating day-to-day configuration of individual pipelines to the teams best suited to own them.

The problem: Pipeline governance at scale

In large enterprises, dozens of teams may each operate their own custom pipelines. Common requirements, security context enrichment, cost allocation tagging, and storage bucket assignment must be applied consistently across the board. However, without a structured mechanism to enforce such standards, platform teams face a choice that gets harder with every new team, every new data source, and every new regulation:

  • Lock everything down and become the gatekeeper for every pipeline change. Every adjustment, no matter how small, becomes a ticket, a review, a delay. Innovation stalls. Teams wait days for changes that should take minutes.
  • Copy rules everywhere and accept the risk of configuration drift. What starts as a manageable set of shared rules gradually fragments—one team forgets to apply the latest cost allocation tag, another skips the security enrichment step, a third routes data to the wrong bucket. The inconsistencies compound silently until they surface as compliance gaps or billing surprises.

Customers consistently tell us that this approach doesn’t scale. What they need is a way to separate what must always happen from what teams should be free to decide, and to encode that separation directly into the platform, not just into process documents that team members forget to follow.

What are Pipeline Groups?

A Pipeline Group is a first-class configuration object that separates a global pipeline list into distinct sets; the pipelines in these sets are then members of the group. A Pipeline Group orchestrates a special type of reusable pipeline to define what happens before or after the stages of the member pipelines.

Pipeline Groups allow you to:

  • Organize pipelines into groups with clearly defined membership.
  • Define execution order so processing happens in a predictable, layered sequence.
  • Control stage execution for the member pipelines so that teams can turn on or turn off specific pipeline stages.
  • Enforce mandatory global processing that no team can bypass or override.
  • Allow team-level customization within boundaries defined by the platform team.

Pipeline Groups determine how data flows through your pipelines. This ensures that governance isn’t an afterthought bolted onto a pipeline configuration, but rather it’s built into the execution model itself.

Figure 1: Pipeline Groups determine how data flows through pipelines
Figure 1: Pipeline Groups determine how data flows through pipelines

How Pipeline Group ownership works

The design behind Pipeline Groups draws a deliberate line between group-level configuration and member-pipeline configuration. How organizations map this division to team ownership is up to them. Pipeline Groups provide the mechanism, not a prescriptive policy. The pattern we see most often is straightforward: platform or SRE teams own the pipeline groups, and application teams own their member pipelines within those groups.

To make this more concrete, consider a platform team that’s responsible for observability across multiple business units. They create a Pipeline Group that adds a business segment field to every record for organizational attribution, applies cost allocation tags for accurate chargeback, sets security context so sensitive data is handled consistently, and assigns data to the correct storage bucket based on retention and compliance needs. These are the rules that must always apply, and because they operate at the group level, no member pipeline can bypass or override them.

Member pipeline ownership can be more granular than a simple admin-vs-team split. While a group is always admin-owned, the individual pipelines that make up the Pipeline Group can each have different owners. For example, a pipeline that enriches every record with organizational metadata or classifies data sensitivity might be globally relevant and owned by the platform team, or by a specialist team like the security team. A pipeline that handles domain-specific compliance logic, such as PCI field masking for a financial services division, might be owned by that compliance team because they have the expertise that the platform team lacks. This arrangement reflects how responsibility within that organization is distributed.

Within that group, application teams each get their own member pipeline. One team configures parsing for a specific log format. Another sets up filtering to reduce noise. A third extracts metrics tailored to their services. Each team works independently within its own pipeline without touching the platform-level rules above and without needing to coordinate with other teams or the platform team for routine changes.

The value of this separation becomes clear when things change, as they do frequently in large enterprises. Say that one of those business units onboards a new microservice that generates a completely new log format. Under a centralized-only model, the team would file a request, wait for the platform team to update the pipeline, validate, and iterate. With Pipeline Groups, the team simply adds their parsing and filtering rules to their own member pipeline. The mandatory enrichment, cost tagging, and routing are already guaranteed by the group. The new service is observable in hours, not weeks, and the platform team doesn’t need to be involved at all.

Or picture the reverse direction: a new regulatory requirement lands, say, a mandate that all log data from EU-based services must be assigned to region-specific storage. The platform team updates the group configuration once. Every member pipeline inherits the change immediately.

Why the flexibility of Pipeline Groups matters

The pace at which enterprise data requirements evolve has fundamentally changed. New services are spun up in days, not months. Regulatory landscapes shift across jurisdictions. The types of software entities that organizations need to observe, from traditional infrastructure to AI model outputs, edge devices, and third-party SaaS telemetry, keep expanding. In this environment, a rigid, centralized-only pipeline configuration becomes a constraint rather than a benefit.

Pipeline Groups are built for this reality. They give enterprises a mechanism that:

  1. Creates room for innovation. Application teams can iterate on their pipeline configurations independently, experimenting with new parsing rules, extracting new metrics, and adapting to new data shapes without waiting for central approval on every change.
  2. Eliminates bottlenecks. Platform teams define the rules once and let the system enforce them. Teams are freed from serving as gatekeepers for routine changes and can focus on architecture, standards, and strategy.
  3. Ensures that ever-changing guardrails are applied. As compliance requirements, security policies, or cost structures evolve, platform teams can update them all in one place. The changes propagate automatically to every pipeline in the Pipeline Group.

Pipeline Groups give platform teams the governance controls and flexibility they need

Pipeline Groups give enterprise platform teams the governance controls they’ve been asking for, mandatory processing, centralized ownership of sensitive stages, and structured delegation, all without taking away the flexibility that makes Dynatrace OpenPipeline® valuable to individual teams in the first place.

Ready to get started with OpenPipeline and Pipeline Groups?

If your organization manages observability across multiple teams and struggles with consistency, Pipeline Groups are now generally available.

Learn how to define global guardrails once and empower your application teams to build their own pipelines confidently within standardized boundaries.

The post Pipeline Groups in Dynatrace OpenPipeline: Enterprise-grade governance explained appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/pipeline-groups-in-dynatrace-openpipeline-enterprise-grade-governance-explained/feed/ 0
Platform engineering success requires comprehensive observability https://www.dynatrace.com/news/blog/platform-engineering-success-requires-comprehensive-observability/ https://www.dynatrace.com/news/blog/platform-engineering-success-requires-comprehensive-observability/#respond Wed, 04 Mar 2026 13:18:17 +0000 https://www.dynatrace.com/news/?p=73747 3D Futuristic circuit background

Platform engineering has emerged as a critical discipline for organizations managing cloud-native complexity. However, building internal development platforms (IDPs) is only part of the challenge. Without comprehensive observability, even well-architected platforms can become operational blind spots that hinder developer productivity. The platform engineering imperative Modern organizations face unprecedented infrastructure complexity. The DevOps principle of “you […]

The post Platform engineering success requires comprehensive observability appeared first on Dynatrace news.

]]>
3D Futuristic circuit background

Platform engineering has emerged as a critical discipline for organizations managing cloud-native complexity. However, building internal development platforms (IDPs) is only part of the challenge. Without comprehensive observability, even well-architected platforms can become operational blind spots that hinder developer productivity.

The platform engineering imperative

Modern organizations face unprecedented infrastructure complexity. The DevOps principle of “you build it, you run it” struggles to scale when development teams must navigate dozens of interconnected tools, distributed systems, and the expansive Kubernetes ecosystem. Platform engineering addresses this challenge by creating abstraction layers that provide self-service capabilities while shielding developers from underlying complexity.

Within this context, Kubernetes is best understood as a foundation, rather than a finished platform – a strong starting point for building the internal experiences teams actually rely on. That perspective captures platform engineering’s core mission: enabling organizations to scale development operations through thoughtful abstraction and intelligent automation.

Why platform observability matters

Platform observability goes beyond traditional infrastructure monitoring. When an IDP becomes the foundation for all development activities, its health directly shapes the speed, reliability, and quality of everything built on top of it. Any degradation creates immediate ripple effects, including blocked deployments, delayed releases, unresolved incidents, and a noticeable drop in developer productivity.

But the impact goes deeper. Clear, consistent insight into platform health is essential for developer experience (DevEx). When developers can quickly understand what’s happening inside the platform, they resolve issues faster, experience fewer sources of friction, and build greater trust in the platform itself. Strong platform observability reduces guesswork, accelerates debugging, and helps developers ship code with confidence, rather than wrestling with hidden blockers.

Effective platform observability must address multiple dimensions:

  • Platform availability and performance. Help all core services to operate reliably and within expected thresholds.
  • Usage patterns and adoption. Understand how teams consume platform services and where friction emerges.
  • Security and compliance. Continuously monitor for vulnerabilities, misconfigurations, and policy violations.
  • Resource utilization. Optimize capacity and infrastructure costs with real-time and historical data.
  • Success metrics. Measure whether platform investments translate into improved DevEx, faster delivery, and tangible business value.

Observable platform architecture

A typical IDP consists of multiple interconnected layers, each requiring tailored observability approaches.

  • Infrastructure layer (Kubernetes-based): Traditional metrics, including CPU, memory, network performance, and cluster health indicators.
  • Platform services layer: Service mesh performance, policy engine effectiveness, security vault availability, and inter-service communication patterns.
  • Delivery services layer: CI/CD pipeline performance, deployment frequency, success rates, and GitOps synchronization status.
  • Self-service interface layer: Developer portal performance, template usage patterns, and user adoption metrics.

Key performance indicators for platform success

Organizations should track platform effectiveness across five critical areas:

Adoption metrics

  • Active user count and growth trends
  • Services deployed through the platform
  • Team adoption rates across business units
  • Feature utilization patterns

Developer experience

  • Net Promoter Score (NPS) from internal development teams
  • Time to first successful deployment for new team members
  • Self-service request fulfillment times
  • Support ticket volume and resolution times

Delivery performance

  • DORA metrics, including deployment frequency, lead time for changes, mean time to recovery, and change failure rate
  • Pipeline success rates and failure analysis
  • Time to production for new features and services

Platform reliability

  • Availability metrics for critical platform components
  • Error rates and performance of self-service APIs
  • Mean time to resolution for platform incidents
  • Service-level objective achievement

Financial efficiency

  • Infrastructure cost per service or team
  • Resource utilization rates across platform components
  • Cost allocation accuracy and chargeback effectiveness
  • Total cost of ownership optimization

Real-world implementation examples and lessons learned

Container registry optimization

A large financial services organization implemented comprehensive monitoring for their container registry, tracking API availability, authentication performance, storage utilization, and security scan results. Analysis revealed that deployment slowdowns occurred consistently during morning hours due to multiple teams scheduling builds simultaneously. By implementing intelligent build scheduling, they reduced average deployment times by 35%.

CI/CD pipeline performance analysis

An enterprise technology company discovered that their deployment times had increased 40% over three months through systematic pipeline monitoring. Deep analysis revealed that security scanning steps were experiencing performance degradation as codebases grew. They optimized scanner configurations and implemented parallel processing, restoring deployment times to acceptable levels.

Developer portal engagement

A global manufacturing company found that while their Backstage-based developer portal had strong initial adoption, sustained engagement declined over time. Usage analytics revealed that complex service templates were creating friction for development teams. Simplifying templates and improving documentation increased sustained engagement by 65%.

Predictive capacity management

Organizations can leverage observability data for proactive capacity planning. By monitoring resource utilization trends and applying predictive algorithms, platform teams can forecast capacity needs and trigger scaling operations before users experience performance effects. This approach is particularly valuable for resources requiring extended provisioning times.

Implementation strategy

Successful platform observability requires a systematic approach:

  • Define clear objectives. Establish specific, measurable goals for platform performance and user experience.
  • Implement comprehensive instrumentation. Ensure all platform components emit relevant telemetry data.
  • Establish performance baselines. Document normal operational patterns before implementing alerting thresholds.
  • Automate validation workflows. Create automated processes that validate new deployments and configuration changes.
  • Connect monitoring to action. Link observability data to notification systems and automated remediation workflows.
  • Create stakeholder-specific dashboards. Develop tailored views for different audiences, from developers to executive leadership.
  • Measure business impact. Connect platform metrics to broader organizational outcomes and key performance indicators.

Common implementation challenges

Organizations should watch out for these pitfalls:

  • Data overload without actionable insights. Focus on metrics that drive specific actions, rather than comprehensive data collection.
  • Siloed monitoring approaches. Implement the observability strategy to span all platform layers and component interactions.
  • Neglecting user experience metrics. Balance technical performance metrics with developer satisfaction and productivity indicators.
  • Over-reliance on manual processes. Automate deployment validation and incident response where possible.
  • Unclear ownership models. Establish clear responsibility for platform components to enable efficient incident response.

Strategic value of platform observability

Platform engineering represents a significant investment for most organizations. Comprehensive observability helps to ensure that this investment delivers measurable returns through improved developer productivity, reduced operational overhead, and accelerated time to market.

The most successful platform teams treat their platforms as products, using observability data to drive roadmap decisions, optimize resource allocation, and demonstrate business value to stakeholders. They move beyond reactive troubleshooting to proactive optimization and strategic planning.

Organizations that implement thoughtful platform observability strategies position themselves to scale development operations effectively while maintaining operational excellence. In today’s competitive landscape, this capability can provide significant competitive advantages through faster innovation cycles and more reliable service delivery.

Turn your platform into a productivity powerhouse with data-driven observability. Start your free trial

Platform engineering and observability FAQs

Why is observability essential for platform engineering?

Because an internal development platform (IDP) touches every application and service, blind spots can stall deployments, reduce developer productivity, and delay releases. Observability ensures the platform actually accelerates development, rather than slowing it down.

How is platform observability different from traditional infrastructure monitoring?

Traditional monitoring focuses on system health, including CPU, memory, and uptime. Platform observability adds visibility into developer usage patterns, CI/CD pipeline health, self-service adoption, compliance posture, and business impact.

What are the key dimensions of platform observability?

  • Availability and performance: Are components running smoothly?
  • Usage and adoption: Are developers using the platform effectively?
  • Security and compliance: Are vulnerabilities and policy violations detected early?
  • Resource utilization: Are resources optimized to control costs?
  • Business outcomes: Is the platform delivering measurable value?

How can organizations measure platform engineering success?

Organizations should track developer productivity metrics like deployment frequency and lead time. Platform adoption rates indicate whether self-service capabilities meet developer needs. Security and compliance measurements demonstrate risk reduction across development workflows. These metrics should align with business outcomes like faster time-to-market and customer satisfaction.

The post Platform engineering success requires comprehensive observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/platform-engineering-success-requires-comprehensive-observability/feed/ 0
Building trust in agentic AI: An observability‑led 90‑day action plan https://www.dynatrace.com/news/blog/agentic-ai-report-new-observability-strategy/ https://www.dynatrace.com/news/blog/agentic-ai-report-new-observability-strategy/#respond Thu, 05 Feb 2026 17:06:38 +0000 https://www.dynatrace.com/news/?p=73000 Pulse of Agentic AI Report - Action plan

Agentic AI is gaining traction quickly in pursuit of autonomous operations. But establishing the trust, reliability, and governance required to derive real business value is proving more challenging. New Dynatrace research suggests ways leaders can pair human oversight with observability as a real‑time control plane for scaling agentic AI safely from pilot to production.

The post Building trust in agentic AI: An observability‑led 90‑day action plan appeared first on Dynatrace news.

]]>
Pulse of Agentic AI Report - Action plan

New research from Dynatrace reveals how organizations are adopting agentic AI to drive greater business value through automating operations. But as teams push toward AI‑driven automation at scale, they’re also confronting a core challenge: the variable, context-dependent nature of AI systems makes it difficult to establish the reliability, safety, and governance needed to fully realize ROI.

Context is key for AI systems to avoid losing track of instructions, hallucinating missing dependencies, or misinterpreting evolving system states—especially during extended multi‑step tasks. Some of the technical challenges AI agents present include:

  • Context fragmentation: As tasks grow more complex, agents cannot reliably hold, retrieve, or apply the full operational context they need for accurate decisions.
  • Unpredictable autonomy: Small gaps or inconsistencies in context can cause cascading errors that affect downstream systems, workflows, and data integrity.
  • Lack of verifiable control signals: Without real‑time, fact‑based grounding, agents cannot validate their own assumptions or detect deviations, making it extremely difficult for leaders to operationalize autonomy safely.

These issues explain why agentic AI is accelerating but still challenged to become “production‑ready” without a new foundational layer of observability, governance, and human oversight.

The emerging reality: What the 2026 Pulse of Agentic AI reveals

The 2026 Pulse of Agentic AI is a global survey of 919 senior leaders and decision makers directly involved in or responsible for agentic AI development and implementation. Results show that agentic AI is advancing rapidly but encountering structural barriers on the path to scalable autonomy.

  • Agentic AI is moving quickly from experimentation into real operations. Most organizations (72%) now run 2-10 agentic AI initiatives, and 50% have at least some production deployments. Adoption is strongest where reliability and risk sensitivity are highest: IT operations (70%), data processing (51%), and cybersecurity (49%), where automation can deliver fast, measurable gains.
  • Maturity is uneven. While investment is rising and expectations for ROI are high—44% have projects in broad adoption in select departments—only 23% have projects in mature, enterprise-wide adoption. The primary blocker is not ambition, but trust. Leaders cite security and data privacy (52%), and technical challenges (51%)—especially limited visibility into agent behavior and difficulty defining when agents can act autonomously versus when humans must intervene.
  • AI operations forge a new role for human oversight. Most agentic decisions are reviewed or validated by people (69%), and 44% rely on manual methods to monitor agent interactions—slowing scale and increasing operational risk. These findings make one conclusion clear: agentic AI cannot reach its potential through experimentation alone. Scaling autonomy requires stronger governance, clearer decision boundaries, and real‑time observability that connects AI behavior to system reliability and business outcomes.

From insight to execution: Why AI projects are stalling and how observability enables results

The research makes clear why many agentic AI initiatives stall before delivering full business value.

Leading organizations are already using observability as more than a monitoring tool

Observability is becoming the foundation for scaling agentic AI safely. Nearly seven in ten respondents apply observability during implementation to integrate agents with existing systems, monitor data quality, and detect anomalies. As agentic systems move into production, observability is increasingly used to track agent performance in real time, validate outputs, and correlate AI behavior with reliability, efficiency, and risk.

Observability data alone is not enough

At the same time, the research exposes a clear gap: many teams still rely on manual reviews to understand agent interactions, slowing scale and limiting trust. Respondents consistently point to limited real‑time visibility and weak connections between technical signals and business outcomes as barriers to autonomy.

Observability must become a fact-based control plane for agentic AI

This is the inflection point. Organizations that treat observability as a real‑time control plane—governing decisions, enforcing guardrails, and grounding AI actions in facts—are better positioned to expand autonomy with confidence. The following 90‑day action plan translates these proven practices into practical steps leaders can take now.

A 90‑day action plan for execs and IT leads

Operationalizing agentic AI requires moving deliberately—from experimentation to governed, observable autonomy. The first 90 days should focus on building AI trust, resilience, and measurable business impact.

days 1-30

Establish foundations and governance.

First, define clear decision boundaries for when agents can act autonomously versus when human approval is required.

Inventory active agentic AI initiatives, assess their business criticality, and identify where visibility gaps exist.

Stand up a baseline observability layer that instruments AI agents, workflows, and data paths, capturing logs, metrics, traces, and contextual signals among agents and infrastructure needed for validation and auditability.

days 31-60

Build trust and controlled autonomy.

Define clear roles for human‑in‑the‑loop operations, placing human judgment in the drivers’ seat for intent and accountability while agents perform tasks and perfect execution.

Set up observability‑driven data‑quality checks, drift detection, and alerts.

Promote observability from passive monitoring to active control by enforcing rules, detecting anomalous behavior in real time, and correlating agent actions with reliability, cost, and performance outcomes.

Secure two quick wins: Implement these trust factors for two high-criticality cases to harden these guardrails to create a template for other use cases.

days 61-90

Scale with confidence.

Graduate proven use cases from supervised to higher levels of autonomy, beginning with repeatable, high‑ROI workflows.

Embed AI observability into operational reviews and executive KPIs.

Establish a continuous improvement cycle to safely expand autonomous operations across the business.

The bottom line: Autonomy only scales with trust

Agentic AI is here—and it’s accelerating. The organizations that win the next phase of AI transformation will be those that implement autonomy with control to minimize risk:

  • Build incrementally, moving from supervised to autonomous operations
  • Ground all agent decisions in deterministic observability data
  • Redesign human roles to guide, not replace, human judgment
  • Treat reliability, safety, and transparency as business‑critical capabilities

With a well‑structured 90‑day plan, enterprises can convert experimentation into operational advantage—unlocking the resilience, scalability, and efficiency that agentic AI promises, while keeping humans firmly in control of outcomes.

Download the full report for a deeper look into agentic AI adoption trends, maturity criteria, KPI breakdowns, and stage-specific observability priorities.

The post Building trust in agentic AI: An observability‑led 90‑day action plan appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/agentic-ai-report-new-observability-strategy/feed/ 0
Database monitoring made easy with Dynatrace’s AI-powered solution https://www.dynatrace.com/news/blog/database-monitoring-made-easy-with-dynatraces-ai-powered-solution/ https://www.dynatrace.com/news/blog/database-monitoring-made-easy-with-dynatraces-ai-powered-solution/#respond Wed, 21 Jan 2026 16:21:06 +0000 https://www.dynatrace.com/news/?p=72573 Business process observability graphic

Modern organizations face growing challenges in managing thousands of distributed databases across clouds, regions, and teams, often leading to blind spots and reactive firefighting. Dynatrace’s new Database Monitoring app changes the game by providing unified visibility and AI-powered insights across your entire database landscape. With proactive health scoring, query-level analytics, and seamless integration with application […]

The post Database monitoring made easy with Dynatrace’s AI-powered solution appeared first on Dynatrace news.

]]>
Business process observability graphic

Modern organizations face growing challenges in managing thousands of distributed databases across clouds, regions, and teams, often leading to blind spots and reactive firefighting. Dynatrace’s new Database Monitoring app changes the game by providing unified visibility and AI-powered insights across your entire database landscape. With proactive health scoring, query-level analytics, and seamless integration with application monitoring, teams can quickly identify and resolve issues before they impact users. This solution empowers developers, SREs, and platform engineers to collaborate effectively, reduce complexity, and take ownership of database performance with confidence.

Organizations today rely on thousands of databases spread across clouds, regions, and teams; traditional monitoring tools simply can’t keep up.

In the AI era, databases are increasingly ephemeral; workloads are massively data-intensive, and debugging efforts span across complex, distributed systems. Simply put, traditional tools were never designed for this pace or scale. Fragmented tooling and data silos add complexity, while usage-based costs limit observability, creating blind spots and slowing issue detection. Teams often end up firefighting problems, such as slow queries and deadlocks, leading to operational fatigue and delayed resolutions that impact performance and innovation.

At the same time, as database reliability shifts from DBAs to engineering teams, visibility gaps and expertise challenges grow. Developers often lack feedback on how their code impacts performance, which slows down releases, increases risk, and makes collaboration with SRE and platform teams more challenging. With many developers lacking in-depth database knowledge, it’s crucial to deliver tools that simplify management and provide actionable insights, allowing teams to innovate faster without compromising reliability.

It’s time to transform database observability with an AI-native approach designed specifically for developers, SREs, and platform engineers. Dynatrace’s new Database Monitoring solution is revolutionizing database observability with unified visibility across your entire database estate, from on-premises to hybrid and multicloud environments.

Unified observability across your entire database estate

The new Dynatrace Databases app, powered by the Dynatrace® unified observability platform, takes database monitoring to the next level with deep, actionable insights.

You can now monitor entire database clusters, gain deep AI-powered insights with remediation plans and alerts, and access granular visibility to reduce outages and downtime. Essentially, this means you can now go beyond basic health checks and gain visibility into:

  • Query-level analytics and transaction tracing across services to pinpoint performance bottlenecks.
  • Schema and index behavior analysis to understand structural inefficiencies.
  • Preventive alerts and insights that help you act before issues impact users.
  • Historical and real-time performance trends for smarter optimization and capacity planning.

With all critical data in one place, practitioners get a holistic view of their entire ecosystem. What used to take significant time, resources, and tools is now available in one central place.

Database monitoring Explorer dashboard in Dynatrace

The Databases application was designed to monitor a diverse range of database technologies across on-premises, public cloud, and hybrid cloud environments. With native support for cloud-specific services, Dynatrace delivers comprehensive coverage regardless of your infrastructure strategy.

From reactive to proactive database monitoring

Proactive database monitoring starts here. Instead of waiting for problems to surface and reacting to outages and slowdowns, teams can now prioritize the necessary adjustments to keep their databases healthy and performing optimally, and to ensure their applications run smoothly. This proactive approach drives performance improvements beyond query optimization, allowing you to address issues before they impact users.

With granular database views and new metrics, including replication status, cache hit ratios, lock contention, and server pulse overviews, practitioners gain the clarity needed to make informed decisions.

With real-time intelligence, Databases eliminates silos to help you optimize performance, reduce time to resolution, and deliver a seamless user experience. Combined with APM integration, Dynatrace delivers end-to-end visibility from services down to individual database queries, so you can understand how the layers interact.

How it works

After a seamless and quick integration, Databases immediately provides you with an overview of everything happening across all your databases.

Database monitoring Explorer dashboard in Dynatrace

You can dive deeper into the instances you’re responsible for. Go to the Overview tab and view the Health score, which indicates the instance’s health.

The health score is an innovative feature that allows you to be proactive with your database fleets and make sure they always meet your application needs. In this example, you can see how a low health score indicates a database health issue related to the instance’s Cache Hit Ratio.

Database monitoring health score dashboard in Dynatrace

Once you have a clear indication of what’s preventing a database from maximizing its potential, you can dive deeper into that specific database and the database calls reported as the root cause of the problem, all from the health score.

Database monitoring Activity metrics dashboard in Dynatrace

In the above graph, we can see that the problem is within the airbases. Now that we know which database is causing the problem, we can access it and debug the queries that are contributing to the low database Cache Hit Ratio.

Seamless debugging across databases and services

Experience the full power of Dynatrace’s unified platform by taking your database monitoring a step further. Monitoring your databases and services together in Dynatrace gives you a holistic view of your application ecosystem. This integration allows you to quickly get to the root cause, identify the issues, and transition effortlessly from the Database app to the Services app for deeper analysis, quick and efficient debugging, and resolution.

Let’s take a look at an example. Let’s say one of your databases experiences a Cache Hit Ratio problem, as depicted in the screenshot below.

Database monitoring Activity metrics dashboard in Dynatrace

From the database host overview page, go to the Calling Services tab. Here, you can pinpoint which service is causing, or is impacted by, the issue based on certain metrics:

Database monitoring Calling services dashboard in Dynatrace

Once you identify the problematic service in the Databases app, go to the Services app. There, you can debug the root cause by analyzing the problematic queries in detail.

Database monitoring Database queries dashboard in Dynatrace

Get started

By breaking down silos and delivering intuitive, actionable insights, the new Dynatrace Database Monitoring solution helps teams collaborate seamlessly, reduce complexity, and take ownership of database performance with confidence. From unified observability and AI-powered analytics to proactive health scoring and end-to-end debugging across databases and services, Dynatrace transforms how you manage and optimize your database landscape.

Ready to experience Dynatrace database monitoring for yourself? Explore the full capabilities today and see how effortless it can be.

The post Database monitoring made easy with Dynatrace’s AI-powered solution appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/database-monitoring-made-easy-with-dynatraces-ai-powered-solution/feed/ 0
Data in context: How Dynatrace solves the OpenTelemetry analytics challenge https://www.dynatrace.com/news/blog/data-in-context-how-dynatrace-solves-the-opentelemetry-analytics-challenge/ https://www.dynatrace.com/news/blog/data-in-context-how-dynatrace-solves-the-opentelemetry-analytics-challenge/#respond Thu, 15 Jan 2026 18:45:02 +0000 https://www.dynatrace.com/news/?p=72469 OpenTelemetry logo

Discover a new era of enterprise-grade observability with OpenTelemetry and Dynatrace. Our latest enhancements unlock powerful possibilities for modern cloud native teams with mass data analysis (MDA) at scale.

The post Data in context: How Dynatrace solves the OpenTelemetry analytics challenge appeared first on Dynatrace news.

]]>
OpenTelemetry logo

The real OpenTelemetry challenge: When OTel meets scale

Standardizing on OpenTelemetry gives teams flexibility and control in modern cloud native architectures. The instrumentation works. The data flows. The collection is solved. However, as organizations scale out OTel, they encounter the analytics gap- and the cost continues to climb without clear returns. Can your observability platform turn millions of spans and log lines into clear, contextual answers without manual correlation or vendor lock‑in?

While having OTel data is exciting, many teams find themselves asking, ‘What’s the actual payoff?’ Engineers may occasionally explore telemetry data, but without intelligent analytics to connect the dots, the data often remains underutilized. Meanwhile, managers struggle to quantify the value of their investment, especially as costs climb with scale. This is the natural challenge of DIY setups, where teams focus on data collection but rarely think critically about turning that data into actionable insights.

Anyone can analyze a single trace; that’s easy. Deriving answers from millions? That’s the hard part. The promise remains unfulfilled. Until now.

Why OpenTelemetry + Dynatrace changes everything for teams

Dynatrace closes this gap. With Dynatrace, all telemetry signals are combined to give you the insights you need:

  • Comprehensive failure analysis across large trace sets
  • Response-time insights revealing performance patterns at scale, with logs and exceptions in context
  • Deep visibility into database and queuing systems
  • AI-powered intelligence delivering contextualized insights
  • Enterprise operational controls providing cost allocation, secure data handling, and scalable telemetry management

Your OpenTelemetry data transforms into actionable contextualized answers that help you move faster, ship confidently, and get back to building.

Complete enterprise coverage for OpenTelemetry

Dynatrace brings OpenTelemetry for Enterprise to life through specialized analysis designed for practitioners troubleshooting issues in Kubernetes environments, available in our extended Services app. AI-powered intelligence, including anomaly detection, ensures you get answers faster, without losing context.

Failure analysis: from single traces to mass insights

When your booking service fails, you can see the complete story: failed traces, related log entries, specific database statements, exceptions- all automatically correlated in one view. The Failure Analysis also provides a visual investigation of your OTel data.

Failure analysis comparing timeframes with detailed log insights
Figure 1. Failure analysis comparing timeframes with detailed log insights

Time-based comparisons allow you to overlay current failures against previous windows, instantly identifying regressions- what was stable yesterday and failing today becomes obvious. And there’s more: advanced visualization distinguishes between different failure types and severities, automatically categorizing them so you can prioritize based on actual user impact.

It’s a new, intuitive way to explore data visually with full context- analyzing failure patterns across your entire architecture and understanding how problems flow through distributed systems. Derive answers from millions of spans that individual trace inspection would never reveal.

Response time analysis with full telemetry context

Mass data analysis extends to performance insights. Dynatrace delivers response time analysis and comparisons built for practitioners. The platform allows you to easily compare failures between two time windows to spot exactly when things degraded.

See how response times correlate with database performance, downstream dependencies, like calls or queue interactions, and infrastructure resource utilization.
Dive in visually, explore the correlated context, and understand what’s happening- all in one place.

Response time analysis comparing timeframes with full telemetry context
Figure 2. Response time analysis comparing timeframes with full telemetry context

Database queries: understand service-to-database interactions

Modern services thrive or fail based on their interactions with their databases. Dynatrace provides comprehensive database analysis for OpenTelemetry-instrumented services, showing exactly what your services are doing against your databases.

Database query analysis revealing service-to-database interactions, query performance, error rates, and high-impact queries
Figure 3. Database query analysis revealing service-to-database interactions, query performance, error rates, and high-impact queries

Get immediate visibility into your most expensive queries across your entire environment- whether it’s Cassandra, SQL, or other databases. Queries are automatically ranked by cumulative duration (query count times query duration), surfacing what’s actually costing you performance.

When troubleshooting an individual service, you immediately see which database calls are problematic. You can also examine patterns across many services: identifying query problems, detecting spikes, or discovering when services start overwhelming databases with inefficient calls. The view aligns with how your architecture actually functions.

Cloud native queuing systems support

Modern applications stream data through Kafka, RabbitMQ, MQTT, and SQS, sending thousands of messages per second through distributed architectures. Dynatrace delivers comprehensive visibility into these message processing interactions with dedicated metrics, dashboarding, and alerting designed specifically for how modern streaming systems actually operate.

See which services are publishing or receiving messages from which queues, with full performance metrics. Advanced filtering lets you explore your entire environment or drill into a specific service’s queue interactions.

Full visibility into message processing to identify bottlenecks and service issues
Figure 4. Full visibility into message processing to identify bottlenecks and service issues

Exception analysis: uncover patterns and failures

We’ve only scratched the surface of how service analysis capabilities can make an impact. From the Services app, you can seamlessly navigate to related traces in the Distributed Tracing app, which now includes extended Exception Analysis. This enhancement surfaces exceptions across traces with readable stack traces, aggregated insights, and visual markers to highlight problematic spans.

By analyzing exceptions in context, teams can quickly identify patterns, prioritize fixes, and reduce MTTR. Whether leveraging OneAgent or OpenTelemetry, no critical issue goes unnoticed, providing complete visibility and reliability across modern environments.

Get a complete view of exceptions across traces, with trends, failure rates, and detailed stack traces
Figure 5. Get a complete view of exceptions across traces, with trends, failure rates, and detailed stack traces

AI-powered intelligence: pinpoint the needle in the haystack

Dynatrace AI delivers actionable insights through baselining, anomaly detection, and precise alerting, continuously learning your environment’s behavior. From day one, these capabilities surface meaningful deviations with full context, enabling teams to act quickly and confidently.

By analyzing OpenTelemetry data, you can detect trends, predict potential issues, and get intelligent, context-rich alerts. This ensures teams can focus on what matters most- resolving problems faster and optimizing performance- without manual effort or guesswork.

Enterprise operational controls that scale

Beyond analytics, enterprise teams need operational capabilities that work with OpenTelemetry data:

  • Primary fields and tags: Use your existing Kubernetes labels and cloud tags (AWS, Azure) to filter and organize telemetry data. Filter by namespace, cluster, deployment, or custom business dimensions to focus on what matters most.
  • Cost allocation: Track and understand costs by subscription, project, or resource group to optimize spending and ensure efficient resource usage.
  • Pipeline routing and processing: Route telemetry data to specific pipelines based on cloud provider, region, or cluster. Control how data flows through your observability stack to improve efficiency and ensure compliance.
  • Bucket assignment: Assign data storage by environment, account, or custom dimensions. Optimize retention and costs while adapting to operational requirements.
  • Security context: Tag data with permissions and access controls, so teams see only the namespaces and services they’re authorized to access.

These aren’t add-ons; they’re core platform capabilities that work identically whether you use OpenTelemetry or OneAgent instrumentation.

The bottom line

Success comes from choosing the analytics platform designed for practitioners in modern environments- one that delivers insights across millions of signals with AI-powered intelligence. Dynatrace meets you where you are with the “Data in Context” advantage: every signal works together with the enterprise capabilities that cloud native environments demand.

The enhanced Services app and the Distributed Tracing app are now available for Dynatrace Platform Subscription (DPS) customers.

Check out the Dynatrace Playground to experience OpenTelemetry for Enterprise firsthand.

Join us at Perform in Las Vegas, January 26-29!

The post Data in context: How Dynatrace solves the OpenTelemetry analytics challenge appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/data-in-context-how-dynatrace-solves-the-opentelemetry-analytics-challenge/feed/ 0
Accelerate SNMP network device observability with Dynatrace Discovery & Coverage https://www.dynatrace.com/news/blog/accelerate-snmp-network-device-observability-with-dynatrace-discovery-coverage/ https://www.dynatrace.com/news/blog/accelerate-snmp-network-device-observability-with-dynatrace-discovery-coverage/#respond Tue, 06 Jan 2026 20:23:04 +0000 https://www.dynatrace.com/news/?p=72365 SNMP Autodiscovery

When onboarding network devices for observability, challenges often arise related to inconsistent or partial monitoring coverage or inefficient processes. Such challenges make it difficult to ensure that all devices are properly and uniformly monitored and can provide actionable insights. Managing network devices at scale exacerbates the problem, as organizations contend with thousands of devices from […]

The post Accelerate SNMP network device observability with Dynatrace Discovery & Coverage appeared first on Dynatrace news.

]]>
SNMP Autodiscovery

When onboarding network devices for observability, challenges often arise related to inconsistent or partial monitoring coverage or inefficient processes.

Such challenges make it difficult to ensure that all devices are properly and uniformly monitored and can provide actionable insights. Managing network devices at scale exacerbates the problem, as organizations contend with thousands of devices from a diverse range of vendors. This can have an adverse impact on your ability to maintain and troubleshoot your networks.

Network monitoring tools often lack integration with the rest of the infrastructure, making it even more challenging to analyze network monitoring data in context.

To achieve consistent, end-to-end monitoring, you need a tool that allows deterministic network observability and automatically contextualizes all ingested data. Let’s take a look at how Dynatrace does this.

Simplified and accelerated network monitoring

When onboarding network devices to an observability platform, it’s not uncommon to spend a disproportionate amount of time configuring network observability. Despite spending valuable time in the process, the uncertainty that some devices might be overlooked remains. Furthermore, in today’s dynamic environments, devices can be added or removed at a moment’s notice and at a rapid pace. This creates a burden for the operations team to constantly prove that the level of observability is at the right level.

To gain a comprehensive overview of the state of observability across all your environments, Dynatrace has expanded the Discovery & Coverage application to include network device monitoring capabilities. This ensures that there are no blind spots in this domain and that no network device is left unmonitored.

The Discovery & Coverage app achieves that by scanning predefined IP ranges or subnets. Based on the outcome of the discovery process, the journey continues by offering the possibility to automatically onboard the right device extension and create Network Availability monitors. Statistics are provided for both extension and availability coverage, providing tangible data to measure the completion of the observability enablement process.

Manage network devices at scale across distributed environments

SNMP (Simple Network Management Protocol) provides a standardized framework for monitoring and managing devices on IP networks. Its simplicity, scalability, and compatibility with a wide range of hardware make it an ideal choice for network management across diverse environments.

However, managing and monitoring SNMP across devices from multiple vendors in large, distributed, or even siloed networks can become cumbersome. Additional complexity is introduced when various teams own and manage their load-balancing devices or application firewalls and core switches individually.

Historically, IP Address Management (IPAM) tools have been effective at mapping entire IP networks, but they struggle to leverage the observability potential of discovered endpoints. Health information on SNMP devices is often isolated, and discovered devices are not placed in the correct context or topology, thereby failing to fulfill the goal of the discovery process: associating SNMP device data with the IT environments where the devices reside. This is where Dynatrace excels.

Leverage the power of the Dynatrace platform for your SNMP devices

Driving tool consolidation and integrating auto discovery and monitoring into your observability solution reduces costs and boosts operational excellence. Additionally, using the Discovery & Coverage app for auto discovery of networking devices means you’ll spend less time manually tracking devices and reduce the chance of lacking the right extension or using the wrong extension for a device.

The Dynatrace end-to-end observability approach makes your discovered network devices available in Dynatrace Grail® and the Dynatrace platform, allowing you to build innovative functionality while simultaneously reducing tool sprawl.

Your discovered devices appear in the Infrastructure and Operations app as network devices, with essential properties automatically populated. Dynatrace and vendor-provided extensions offered in Dynatrace Hub can enhance the observability level and insights of your individual devices. Further, you’re automatically notified when predefined thresholds are exceeded or when anomalies are detected.

Your network devices will function as part of an integrated network, not as standalone entities. For Dynatrace SaaS customers, network devices are readily available in Grail and can be queried using Dynatrace Query Language (DQL) to create value-added insights in the Davis Anomaly Detection app or to create automations and workflows, such as auto-generated tickets for external systems or notifications sent to teams via Slack or PagerDuty.

Get started with network device autodiscovery in Dynatrace

The SNMP autodiscovery capability is provided by the Discovery & Coverage app. Open the app, navigate to the Network coverage tab, and select Configure scanning.

SNMP network device observability configuration in Dynatrace

Edit SNMP network device observability configuration in Dynatrace

Once a configuration is activated, the discovery process begins.

The SNMP Autodiscovery extension regularly scans your IPv4 and IPv6 address lists, ranges, or subnets for devices with SNMP agents. If devices with matching SNMP parameters are present, they will be added to the environment.

Coverage reports for each device configuration provide an overview of the level of observability for each discovered device, answering questions such as: Is each device polled with the correct extension? Or, is the Network Availability Monitor installed on the management interface?

SNMP network device observability in Dynatrace

The Discovery & Coverage app allows you to add a device type-matching extension in bulk. The Poll icon opens the following windows, enabling the specific extensions with one select.

SNMP network device observability poll device in Dynatrace

Additionally, you can centrally configure Network Availability Monitors (NAM) to continuously probe the health and availability of your devices. This ranges from simple ICMP ping tests to advanced probes, depending on your requirements and setup.

SNMP network device observability create a ping monitor in Dynatrace

Take the first step to simplifying device onboarding

SNMP network device autodiscovery is a powerful tool for network administrators. It simplifies network management, improves efficiency, and allows for easy scalability. By leveraging this feature, organizations today ensure their networks are always accurately represented and observed.

Remember, a well-managed network with thorough observability and health monitors is the backbone of any successful organization. So, embrace the power of SNMP autodiscovery and take your network management to the next level!

If you have already deployed network observability using the extension apps, you need to download and open the Discovery & Coverage app. By configuring autodiscovery, you’ll ensure that no part of your network remains unattended and that you’ve configured the right level of observability in your environment. From now on, the Dynatrace network device onboarding process is greatly simplified and accelerated.

If you’re not yet a Dynatrace customer, consider starting a free trial, opening the Discovery & Coverage app, and navigating to the Network coverage tab. You can then immediately discover SNMP devices in one of your environment’s IP networks and add the right level of observability to the devices.

The post Accelerate SNMP network device observability with Dynatrace Discovery & Coverage appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/accelerate-snmp-network-device-observability-with-dynatrace-discovery-coverage/feed/ 0
Is it vibe coding if you know what you’re doing? https://www.dynatrace.com/news/blog/is-it-vibe-coding-if-you-know-what-youre-doing/ https://www.dynatrace.com/news/blog/is-it-vibe-coding-if-you-know-what-youre-doing/#respond Tue, 23 Dec 2025 20:15:07 +0000 https://www.dynatrace.com/news/?p=72309 Agentic AI icon

When people first started saying they were using AI to build full apps, I was wildly skeptical. You’ve heard these stories too, right? “I built a complete SaaS platform in one weekend using ChatGPT!” Sure, you did. But I kept hearing it. Over and over. Whole apps. Entire startups. Just prompts. So, I decided to […]

The post Is it vibe coding if you know what you’re doing? appeared first on Dynatrace news.

]]>
Agentic AI icon

When people first started saying they were using AI to build full apps, I was wildly skeptical. You’ve heard these stories too, right? “I built a complete SaaS platform in one weekend using ChatGPT!” Sure, you did.

But I kept hearing it. Over and over. Whole apps. Entire startups. Just prompts. So, I decided to test the hype. I’ve had the idea to build an app to help collectors track and manage their sports card collections for years. I’ve tried Excel, Google Sheets, Airtable, and even Retool. They are all excellent tools but weren’t the right ones for the job I was trying to accomplish. But after a friend showed me how he was using Claude Code in his CLI, the light bulb went on.

Two weeks later, I had a working site. Authentication. A database. A clean interface. All built with AI tools, a little stubbornness, and an embarrassing number of “let’s try that again” moments.

That’s when it hit me: It’s not vibe coding if you know what you’re doing.

The CLAUDE.md rulebook

The first lesson I learned: AI needs boundaries. So, I made a file called CLAUDE.md. Think of it like an employee handbook for a very confident robot. It spells out how I want things done — naming conventions, frameworks, the tone of my responses; even how to handle uncertainty. Without it, Claude will take creative liberties. With it, it’s like working with someone who mostly listens and only occasionally decides your database schema needs “personality.” You don’t need to go all in on “spec-driven development” when you’re first starting out, you just need some guide rails.

Second, design consistency is not AI’s strong suit. One minute it’s using Material UI, the next it’s gone rogue with pastel buttons and drop shadows from 2007. I started a design guide just to keep it on the rails. It made things infinitely easier to describe and control, all in one place.

An excerpt from the rulebook

# Claude Development Notes

## 🚨 CRITICAL: HOW TO WORK WITH THIS USER (READ FIRST - EVERY SESSION)

### Core Working Principles - NEVER FORGET THESE

1. **BE A SYSTEMATIC CODE ANALYST, NOT A GUESSER**

- ALWAYS search for existing code patterns, dependencies, and conflicts BEFORE attempting fixes

- Use grep, find, and comprehensive code analysis to understand root causes systematically

- Don't apply "band-aid" fixes - identify and solve the underlying architectural issue

- Leverage your file search and cross-reference capabilities instead of making the user debug manually

- This applies to CSS, JavaScript, database queries, API endpoints, configuration - EVERYTHING

2. **THINK LIKE AN EXPERT WITH DEEP SYSTEM KNOWLEDGE**

- Consider how changes affect the entire system: dependencies, imports, inheritance, scoping

- Look for naming conflicts, architectural patterns, and existing conventions

- Analyze the broader codebase structure and established patterns before making changes

- Understand the user is building maintainable, scalable, isolated components

- Apply this expertise to ALL aspects: styling, logic, data flow, security, performance

You can view the full rulebook here.

Don’t give AI the keys to the kingdom

Those initial guard rails will get you started, but they aren’t enough. You’re going to need reliable source control. Badly. AI can (and will) make sweeping changes in seconds. Whole sections of your project — just gone. It’s not malicious, it’s just eager. Commit early, commit often, and keep your fingers hovering over the ESC key to interrupt bad ideas.

Next: Write tests. I know. You won’t want to. Neither did I. But AI code that “works” can still be deeply wrong. I can’t tell you how many times I’ve seen it confidently try to pass data through functions that don’t even exist.

Perhaps most importantly: Never — and I mean never — let AI near your production database or environment. That’s not a metaphor. Twice, my AI coding partner has seen differences in our ORM and our database schema, and twice I have seen it decide to drop all of the database tables to create something that matches the ORM. Thankfully, it was only the development database, but if it had access to production, it would have dropped those tables too.

I have a rule: Staging is for the robots; production is for adults.

“AI forgets. Teammates inherit.”

This is painfully true. AI forgets everything. Context, logic, variable names. All of it. And when you come back to that project later (or someone else does), there’s no memory of why anything was done the way it was. That’s why documentation still matters. Probably more than ever. Because you’re not just documenting for other people anymore — you’re documenting for yourself when your future self wonders, “Why did past me let the robot do this?”

Observability, or “Please tell me what just happened”

This is where tools like Dynatrace come in. I know, I work there, but I’d say this either way. If you’re going to let AI write your code, you need a way to see what it’s doing. Not just logs — observability. Dynatrace tells me what’s happening across the system: slow API calls, inefficient queries, the occasional “why is this endpoint melting the CPU?” moment. It’s not about catching mistakes — it’s about catching surprises. And AI delivers plenty of those.

“Don’t asume corectness”

The typo’s intentional. Mostly. But the point stands: AI code looks right more often than it is right. I’ve had it produce entire components that look perfect and fail silently. It’s confident. It’s articulate. It’s wrong. So, check everything. Run it. Test it. Review it. Don’t just asume corectness.

The big takeaway

Here’s what surprised me the most: Building with AI doesn’t feel like cheating. It feels like collaborating. You’re still making all the decisions — the AI just types faster than you do. It’s like pairing with someone who’s read every Stack Overflow thread but still can’t quite tell when it’s about to create a security hole.

You’re not outsourcing the work. You’re accelerating it. AI doesn’t replace developers. It amplifies them. The people who’ll thrive in this new era aren’t the fastest typists — they’re the clearest thinkers. They’re the best communicators.

I don’t think “vibe coding” is bad. It’s just unstructured. And once you’ve got a plan — a CLAUDE.md file, a solid process, some observability, and a healthy dose of skepticism — it stops being “vibe coding” and starts being modern engineering.

AI can write code. But you still have to lead.

It’s not vibe coding if you know what you’re doing. It’s just engineering with better tools.

Ready to learn more about building real software with AI? Follow Jeff’s 31 Days of Vibe Coding series.

The post Is it vibe coding if you know what you’re doing? appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/is-it-vibe-coding-if-you-know-what-youre-doing/feed/ 0
From black box to glass box https://www.dynatrace.com/news/blog/confidence-where-it-counts/ https://www.dynatrace.com/news/blog/confidence-where-it-counts/#respond Wed, 15 Oct 2025 11:38:03 +0000 https://www.dynatrace.com/news/?p=71427 Confidence where it counts

The Dynatrace 3rd-generation platform continues to evolve, helping teams see, explain, and trust the AI shaping their business.

The post From black box to glass box appeared first on Dynatrace news.

]]>
Confidence where it counts

Key insights

  • Confidence where it counts: Real time, contextual data turns AI from a black box into decisions you can see, explain, and trust.
  • Outcomes, not features: Teams reduce risk, control cost, and improve answer quality by grounding AI in better data and context.
  • Role-based paths: Leaders, platform and SRE, AI engineering and ML ops, application developers, and security each have clear ways to act.
  • Dynatrace 3rd generation foundation: Grail unifies telemetry, Smartscape maps live dependencies, and Davis AI turns insight into explainable actions.

The shift is underway

Your AI helped ship a feature customers love. Then a bad answer slips through, support tickets rise, and no one can explain why. Was it a model change, a prompt tweak, or missing context from an upstream service? When AI behaves like a black box, you cannot manage risk, cost, or trust.

You’re not alone. Budgets and expectations reflect the push to make AI observable and accountable. In our 2025 State of Observability data, 70% of organizations increased observability spend in the last year, and 75% expect to increase it again. Leaders often see the biggest returns from optimizing model configurations, detecting anomalies in model outputs, and automating remediation. Most believe that AI decisions still require a human check because proof matters.

At Dynatrace, we believe the difference is the data. Grail keeps observability, security, and business telemetry together in real-time, connected context. Smartscape maintains a real-time map of services, dependencies, and releases. Davis AI reasons over this context to explain cause and effect and suggest what to do next. Together they make AI explainable and governable, so teams can move faster without losing trust.

Here’s how your teams can build confidence in AI

Choose models with real context

Run evaluations on real workloads, not synthetic tests. Compare latency, cost per answer, and relevancy side by side, then use prompt traces to see why outputs differ. Your team picks the right AI model for the job with evidence, not guesswork. Use canary rollouts to validate the winning configuration on a small slice of traffic, then scale with confidence while a policy records who approved the change.

  • This showcases: AI model evaluation and versioning across providers, prompt tracing and debugging, cost and performance insights tied to real services.

Explain every decision on demand

Capture inputs, prompts, context, and model versions so you can show how a decision was made. Export an auditable bundle, attach it to your review, and keep evidence with the workflow. Approvals move faster because proof is built in. Evidence travels with the workflow so leaders can audit decisions without meetings.

  • This showcases: Exportable audit evidence, lineage across prompts and context, review-ready artifacts for governance.

Spot and prevent drift

Watch for quality changes after releases or upstream shifts in agentic AI systems that span multiple models. Because data lives together in Grail and dependencies are mapped in Smartscape, Davis AI detects drift, explains the likely root cause, and recommends next steps. Your team rolls back or tunes with confidence and documents the outcome.

  • This showcases: End-to-end drift detection and explanation across services and models, causal analysis tied to releases, and actionable root cause context.

Move fast with guardrails

Set policy rules that pause risky paths for humans and promote safe improvements automatically. When a policy is triggered, the flow pauses for approval. When targets are met, changes move forward on their own. Speed and accountability rise together.

  • This showcases: Policy-driven automation, human-in-the-loop controls, measurable targets for quality and cost.

One place to work, together

Platform and SRE, AI engineering and ML ops, developers, and security see the same facts and act in the same space. A built-in experience brings multi-cloud and multi-model views together with alerts, traces, and reviews. Teams focus on outcomes because data and context are already aligned, and agentic tracing makes complex multi-LLM paths explainable.

  • This showcases: Unified, role-aware experience in the Dynatrace 3rd generation platform, powered by Grail for data, Smartscape for live topology and dependencies, and Davis AI for reasoning.

How it works

Dynatrace brings key, differentiated capabilities together, so AI becomes observable, governable, and improvable. Grail stores and relates telemetry with context as a single source of truth. Smartscape maps real-time topology and dependencies. Davis AI analyzes that knowledge to explain issues, correlate cause and effect, and drive or recommend actions. These explainable insights show up in built-in experiences so teams can act quickly in one place, with the same context everyone trusts. Together, they give you a glass-box view of AI across your environment.

Use cases you can try today

Platform Engineering and SRE

Use progressive delivery to your advantage and run a canary release with two model versions on a real service. Compare key metrics and data points, such as latency and error rates, then set an automated rollback if quality drifts beyond the threshold.

AI engineering and ML ops

A/B test two providers for a summarization workload. Measure token usage, cost per answer, and relevancy. Use prompt debugging to tune instructions, then promote the winning configuration.

Application developers

Trace a problematic response from UI to model call. Inspect the prompt, context, and dependencies, then commit a configuration change to fix a latency regression.

Security and compliance

Export an audit bundle for a high-risk workflow. Attach it to your review ticket and record a human approval step. Automate export on schedule.

The market signal

Leaders are investing to make AI observable and governable. From our 2025 State of Observability report, 70% increased observability budgets last year, with 75% of those surveyed planning to increase again next year. Teams expect the biggest return from optimizing model configurations, detecting anomalies in model outputs, and automated remediation. Ninety-eight percent already use AI to support security compliance.

What this means for your organization

For leaders

  • Improve decision quality with model evaluations grounded in a real workload context.
  • Reduce risk with audit trails that satisfy compliance.
  • Control spend by measuring cost next to performance.

For Platform Engineering and SRE teams

  • Standardize model rollouts with versioning and repeatable A/B tests.
  • Tie prompt traces to services, releases, and alerts to cut MTTR.
  • Automate safe rollbacks or configuration changes with policy guardrails.

For AI engineering and ML ops

  • Compare models and versions on latency, cost per answer, and relevancy.
  • Use prompt debugging to tune instructions and ground responses in context.
  • Track drift and trigger-governed actions when quality drops.

For application developers

  • See how AI components behave in production next to service health.
  • Debug prompts without leaving the release context.
  • Ship changes with data on performance and cost impact.

For security and compliance

  • Export signed audit trails that show inputs, outputs, models, and context.
  • Prove who did what, when, and why for policy and regulatory reviews.
  • Route risky outputs for human approval and record the decision.
Gain confidence where it counts

The post From black box to glass box appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/confidence-where-it-counts/feed/ 0
The State of Observability 2025: Business impact, key trends, and a 90-day plan for decision-makers https://www.dynatrace.com/news/blog/ai-observability-business-impact-2025/ https://www.dynatrace.com/news/blog/ai-observability-business-impact-2025/#respond Tue, 07 Oct 2025 11:46:45 +0000 https://www.dynatrace.com/news/?p=71285 Sate of Observability 2025 - action plan

Although organizations are universally adopting AI, moving from pilot to production and sustainable scaling present new challenges. Results from the State of Observability 2025 report suggest some ways organizations can use observability data in a 90-day action plan to drive measurable business results.

The post The State of Observability 2025: Business impact, key trends, and a 90-day plan for decision-makers appeared first on Dynatrace news.

]]>
Sate of Observability 2025 - action plan

Organizations are integrating artificial intelligence into their operations at a rapid pace. This transformation is changing how businesses work, innovate, and compete. The State of Observability 2025 report confirms that while 100% of responding organizations are now using AI, how they’re using it is often fragmented.

Senior IT and business leaders should pursue a unified strategy to link AI initiatives with clear business results. A practical solution that’s gaining traction is AI-powered observability, which is evolving from a technical monitoring platform or tool suite into a strategic control plane for AI transformation.

The emergence of AI technologies within observability presents a novel opportunity for leaders to drive tangible business value from data across the full stack. Insights from our research highlight several key trends that are reshaping priorities so you can create a new action plans for sustainable growth, efficiency, and resilience.

Key takeaways from The State of Observability 2025 report

  • Observability is a fast-growing AI use case. With 75% of organizations increasing their observability budgets, it’s clear that leaders see it as a critical investment for managing AI. In fact, AI capabilities are now the #1 criterion for selecting an observability solution.
  • The AI trust gap is real. Humans are still very much in the loop. A significant 69% of AI-powered decisions are verified by humans, and one in four leaders believes improving trust in AI should be a top priority.
  • AI-powered observability encompasses application security, DevOps, and sustainability. Nearly all security leaders (98%) use AI for security compliance, and 69% have increased budgets for AI-powered threat detection. At the same time, more than 70% of organizations use observability to manage sustainability initiatives.
  • Business observability is on the rise: While only 28% of organizations currently use AI to align observability data with business KPIs, the opportunity is clear. Leaders are moving toward real-time solutions that connect technical performance directly to customer experience and business agility.

These findings illustrate that observability is no longer just about keeping systems running. It’s about optimizing performance, reducing risk, and aligning every aspect of your technology stack with strategic business goals.

How AI-driven insights translate into business results

Being able to understand what’s happening in all dimensions of your operating environments presents some clear business benefits. Here are just a few.

Lower risk and faster response
With AI-assisted detection and guided remediation, teams can reduce the impact of incidents and significantly cut response times.

Lower unit cost and carbon impact
By correlating observability telemetry with cloud spend, energy usage (kWh), and CO₂ emissions, leaders can uncover operational waste and identify clear opportunities for savings.

Stronger security posture
Integrating security and observability enhances compliance, extends threat visibility, and improves the overall quality of incident response.

Greater AI trust and accountability
Human-verified guardrails and comprehensive audit trails improve the transparency and trustworthiness of AI-driven actions.

Clear KPI alignment
It’s now possible to tightly link services and customer journeys to business-critical metrics like MTTR, SLO attainment, cost per request, revenue at risk, and customer experience, enabling informed, real-time decisions.

While these insights are a good start, turning them into an action plan is the critical next step.

A 90-day action plan to drive measurable results and understand your business

For executives looking to deliver measurable ROI from AI projects by harnessing the power of AI-driven observability, here’s an actionable 90-day plan.

days
1-30

Instrument what matters

Begin by mapping your top five revenue or mission-critical customer journeys. Identify and close telemetry gaps across logs, traces, metrics, and real-user experience data to create a complete picture of performance.

days
30-60

Connect to business KPIs

First, establish a scorecard that links technical metrics to business outcomes. Include MTTR, Mean Time to Detection (MTTD), SLOs, cost per request, revenue at risk, customer experience, and a security incident score. Ingest cloud billing data and tag costs to specific services to gain financial visibility.

Next, secure two quick wins

Security. Pilot AI-assisted threat detection and guided response on one high-value service. Measure and report the improvement in time-to-contain threats.

Cost. Link service utilization to cloud spend and carbon emissions (kWh and CO₂e). Identify one clear source of waste, remove it, and report the financial and environmental savings.

days
60-90

Automate with guardrails

Select your two most frequent operational responses and add generative AI to automatically draft remediation workflows, simulate outcomes, and enhance decision-making. Implement a human-in-the-loop approval process for policy checks and rollbacks to maintain control and build trust. Track the outcomes with a live dashboard to demonstrate success.

Why Dynatrace for reliable agentic AI projects

Dynatrace provides the context and controls leaders need to run AI like a business program:

  • Contextual analytics of unified observability, security, and business data.
  • Advanced predictive, causal, and generative AI to provide deterministic answers and validate generative AI results.
  • Preventive operations through ecosystem workflow automation capabilities.

If you’re seeking to turn AI-driven observability into a source of competitive advantage, explore what’s possible with Dynatrace and take the next step toward resilient, agentic AI projects.

Download the full 2025 State of Observability report.

The post The State of Observability 2025: Business impact, key trends, and a 90-day plan for decision-makers appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ai-observability-business-impact-2025/feed/ 0
State of Observability 2025: AI use cases are growing as business leaders seek to build AI trust and ROI https://www.dynatrace.com/news/blog/state-of-observability-2025-ai-trust-roi/ https://www.dynatrace.com/news/blog/state-of-observability-2025-ai-trust-roi/#respond Tue, 07 Oct 2025 11:45:36 +0000 https://www.dynatrace.com/news/?p=71274 Sate of Observability 2025 - findings

AI adoption is universal, but its business impact is not. The Dynatrace annual research report on the state of observability reveals the effects of wider trends in AI adoption. This year’s report shows how observability, once a reactive IT tool, has evolved into the central control plane for AI transformation.

The post State of Observability 2025: AI use cases are growing as business leaders seek to build AI trust and ROI appeared first on Dynatrace news.

]]>
Sate of Observability 2025 - findings

Executives and technology leaders are prioritizing AI observability to reduce risk, lower unit cost, and accelerate delivery, aligning to business objectives.

The State of Observability 2025 report reveals how organizations are moving from experimenting with AI to integrating it into core operations. Not surprisingly, 100% of responding organizations now use AI in some capacity. But this universal adoption isn’t uniform. Data management, AI governance, and security are the most common AI use cases, with observability growing significantly.

As organizations seek to realize ROI on their overall AI investments, observability is clearly emerging as the key to unlocking AI value while mitigating its inherent risks. In other words, AI observability is becoming a prerequisite for the success of AI initiatives.

Why observability is now a C-suite imperative

Executives now recognize that a comprehensive observability strategy is essential for reducing AI risk, lowering unit costs, and accelerating service delivery. Observability is emerging as a vital intelligence layer for managing complex AI initiatives and aligning them with strategic business goals. Further, AI capabilities within observability platforms are becoming a determining factor for selecting an observability vendor.

Findings from the State of Observability 2025 report

The report’s findings underscore this shift:

  • Observability budgets are increasing: 70% of organizations increased their observability budgets this year, and 75% plan to increase them again next year. These increases signal the importance and value leaders are placing on this capability for the success of their business goals.
  • AI capabilities are now the #1 criterion for choosing an observability platform: For the first time, AI capabilities (29%) have surpassed cloud compatibility as the primary criterion for selecting an observability platform. This highlights the market’s demand for intelligent, automated solutions.
  • The AI trust gap is real: Despite widespread AI adoption, a significant trust gap remains. Humans verify 69% of all AI-driven decisions, and 70% of organizations increased budgets for trust and transparency initiatives this year. This indicates that while leaders are eager to use AI, they require guardrails designed to enhance its reliability.

AI is expanding the value of observability across security, sustainability, DevOps, and more

Using AI for security compliance, sustainability, and real-time DevOps automation initiatives is on the rise, fueling the evolution of agentic AI—autonomous systems that plan and execute tasks.

AI-powered threat detection is influencing budget priorities

Security is a prime example of how AI and observability are converging. A staggering 98% of security leaders report using AI to manage security compliance, and 69% are increasing budgets for AI-powered threat detection. Enhancing threat visibility is the top expected growth area for AI over the next five years. By converging security data with observability telemetry, organizations gain faster time-to-contain and fewer customer-impacting incidents.

AI pays dividends for sustainability and managing costs

The scope of observability is also expanding to include environmental sustainability. Our research shows that 70% of organizations use observability to monitor and manage their sustainability initiatives, which in most cases also drives cost reductions. A full 64% report growing budgets for observability-aligned sustainability efforts. Correlating telemetry with resource consumption reduces cost per request and CO₂ emissions by linking telemetry to spend and energy.

Real-time DevSecOps automation is giving rise to agentic AI

The ongoing expansion of AI into combined DevOps and security (DevSecOps) automation represents another powerful shift. Up to 50% of DevSecOps leaders currently use real-time automation, with adoption expected to grow to 70% in five years, driven by use cases like security risk mitigation and anomaly detection. The focus on agentic AI promises high ROI (41%) and is reshaping incident response, infrastructure management, and debugging. Real-time observability with natural language interaction results in AI systems with a shorter time to value through safe, policy-gated actions.

From data to business impact: Closing the KPI gap

While the potential is clear, many organizations are still working to connect observability data to tangible business outcomes. Currently, only 28% use AI to align observability data with key performance indicators (KPIs). This “KPI gap” represents a significant opportunity.

Leaders who successfully bridge this gap can transform their operations. About 22% of leaders report that converging real-time data and AI-driven automation with observability positively impacts business agility, so they can respond more quickly to market changes and customer demands. By tying technical performance metrics like mean time to resolution (MTTR) and service level objectives (SLOs) directly to business metrics like cost per request, revenue at risk, and customer experience scores, leaders can gain real-time insight into how technology performance affects business agility and financial efficiency.

The mandate for observability in the AI era

As organizations increasingly rely on AI, they are also turning to observability to make these complex systems more explainable, reliable, and auditable. Observability is no longer just about monitoring systems. It’s about providing the intelligence and control needed to steer the enterprise through AI transformation. AI-driven observability provides the foundation for lowering risk, strengthening security, and aligning every technological decision with strategic business value.

To explore these findings in greater detail and build a comprehensive strategy, get the State of Observability Report 2025 below.

Download the full State of Observability 2025 report.

The post State of Observability 2025: AI use cases are growing as business leaders seek to build AI trust and ROI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/state-of-observability-2025-ai-trust-roi/feed/ 0
Self-service observability: Empower engineers to get the most value out of your observability data https://www.dynatrace.com/news/blog/self-service-observability-empower-engineers-to-get-the-most-value-out-of-your-observability-data/ https://www.dynatrace.com/news/blog/self-service-observability-empower-engineers-to-get-the-most-value-out-of-your-observability-data/#respond Wed, 01 Oct 2025 16:47:24 +0000 https://www.dynatrace.com/news/?p=71195 Observability data

What good is your observability data if the engineers who need it have to spend a lot of time looking for it? Why is it important to deliver data right into the engineering workflows? Events, logs, traces, and metrics are only valuable if the right people can use them effectively.

The post Self-service observability: Empower engineers to get the most value out of your observability data appeared first on Dynatrace news.

]]>
Observability data

The Dynatrace unified platform experience delivers actionable insights to engineers without the hassle. With customizable launchpads, seamless Backstage integration, the RedHat DevHub plugin, and automated deployment validation with Site Reliability Guardian, engineers get the most important information front and center, allowing them to prioritize work and make informed decisions.

In this blog post, you’ll learn how platform teams can bring observability to developers, right where they need it.

Focus on what matters with launchpads

Dynatrace launchpads let you create customizable home pages that offer an easily digestible view of your environment, tailored to the needs of engineering teams. Launchpads can address various use cases, from onboarding and learning, serving as an entry point for occasional users, or providing an opinionated view for daily operations across different users, teams, and departments.

You can pin quick access to relevant Dynatrace® Apps, link to external tools or specific entries in Dashboards, Notebooks, Workflows, Problems, and more.

For example, the launchpad below was built for a developer team. It covers their daily routines, relevant content, and documents. Using Launchpads this way ensures that even occasional users know where to find key apps, how to resolve common issues, and how to debug problems—using Dynatrace and external resources.

Read our Launchpads blog and explore different Launchpads on our Dynatrace Playground tenant!

This home page is the entry point for a team of developers who work with Dynatrace on a daily basis. It includes links to daily routines, relevant content, and documents.
Figure 1. This home page is the entry point for a team of developers who work with Dynatrace on a daily basis. It includes links to daily routines, relevant content, and documents.

Bring observability to developers with developer portal integrations

Developer portals like Backstage have become increasingly popular in recent years, based on their ability to centralize access to tools, services, and documentation. By offering a consistent interface, they help developers navigate complex ecosystems more effectively and reduce time spent on context switching.

The seamless integration between Dynatrace and Backstage allows developers to pull observability and security data from Dynatrace and display it in software components you manage through the Backstage Software Catalog. This plugin allows you to provide smart links to Dynatrace apps, context-rich overview tables, and Site Reliability Guardian results and logs, all directly in Backstage!

Read our documentation on Backstage integration, and to get started, check out the Backstage Dynatrace plugins monitoring & observability page on the Dynatrace Hub.

This is a team’s Backstage developer portal, directly accessed from Dynatrace through deep links. Data from Dynatrace is displayed directly in the portal.
Figure 2. This is a team’s Backstage developer portal, directly accessed from Dynatrace through deep links. Data from Dynatrace is displayed directly in the portal.

Teams using the RedHat Developer Hub can also benefit from the Dynatrace integration. This plugin seamlessly integrates observability and security data from Dynatrace into the portal, enabling teams to monitor and operate their software components more effectively.

You can display real-time insights alongside managed components and add links to Dynatrace Apps for deeper analysis and root cause investigation.

This page comes from the Red Hat Developer Hub, showing data from Dynatrace displayed directly in the portal.
Figure 3. This page comes from the Red Hat Developer Hub, showing data from Dynatrace displayed directly in the portal.

Validate deployments automatically with Site Reliability Guardian

Dynatrace’s Site Reliability Guardian (SRG) embeds observability directly into the engineering workflow, allowing for fast, automated validation of service health during every change.

Automated change impact analysis

SRG automatically analyses the impact of deployments on performance, availability, and capacity. Powered by Davis® AI, it detects regressions and anomalies early—before they hit production.

SLO validation

Engineers can define service-level objectives for their software components to ensure they meet expected quality standards. See how your software components rate against company benchmarks and get insights on where and how to improve to deliver the best quality. SRG validates these objectives regularly or on demand, ensuring services stay performant and resilient.

Unified overview

SRG validation results are stored in Dynatrace Grail® data lakehouse and visualized in a single-page summary, showing the pass/fail status of the last validations and a detailed evaluation of individual objectives. This allows you to easily spot trends and degressions over time. Visualization can be enriched with custom metadata, such as release versions, build IDs, and environment identifiers.

Automated release validation, CI/CD integrated

Trigger SRG validations directly from your CI/CD pipeline, to get instant feedback and observability insights into the current health and performance state of your component.

For more information, check out our documentation and then go to the Hub and set up your first Site Reliability Guardian! If you’re looking for inspiration, check out our Site Reliability Guardian config-as-code samples.

Conclusion

Observability as a self-service is a reality with Dynatrace. Our unified platform experience allows platform teams to connect observability with developer portals, workflows, and automated deployment validation, giving engineers direct access to the data they need—without bottlenecks.

By empowering engineers to curate and automate their data according to their needs, teams can work more effectively and bring reliability checks earlier into the development lifecycle.

Next steps

Ready to try out these self-service observability enhancements yourself? Explore Dynatrace integrations and start customizing your observability and developer experience!

The post Self-service observability: Empower engineers to get the most value out of your observability data appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/self-service-observability-empower-engineers-to-get-the-most-value-out-of-your-observability-data/feed/ 0
Dynatrace achieves AWS Generative AI Competency: A new milestone in observability and AI https://www.dynatrace.com/news/blog/dynatrace-achieves-aws-generative-ai-competency/ https://www.dynatrace.com/news/blog/dynatrace-achieves-aws-generative-ai-competency/#respond Tue, 30 Sep 2025 12:11:34 +0000 https://www.dynatrace.com/news/?p=71159 Dynatrace | AWS

Enterprise adoption of generative AI is showing no signs of slowing down, and it’s easy to understand why; organizations in every vertical aim to reap its benefits, including increased efficiency, routine task automation, and content generation, ultimately creating a competitive advantage. To better help organizations maximize the benefit and full potential of generative AI, Dynatrace […]

The post Dynatrace achieves AWS Generative AI Competency: A new milestone in observability and AI appeared first on Dynatrace news.

]]>
Dynatrace | AWS

Enterprise adoption of generative AI is showing no signs of slowing down, and it’s easy to understand why; organizations in every vertical aim to reap its benefits, including increased efficiency, routine task automation, and content generation, ultimately creating a competitive advantage. To better help organizations maximize the benefit and full potential of generative AI, Dynatrace has achieved the Amazon Web Services Generative AI (GenAI) Competency.

With this milestone, Dynatrace reinforces its position as a leading observability partner, backed by a proven track record of innovation and customer success on AWS. Building on its achievement of earning the AWS Machine Learning Competency, Dynatrace continues to drive advancements in generative AI.

Weighing the importance of this milestone for Dynatrace customers

This competency is more than just a badge; it’s a validation of how Dynatrace can help organizations safely, efficiently, and cost-effectively adopt generative AI in their business. AWS awards these competencies after rigorous technical validation and proven customer success. This means organizations can trust that Dynatrace solutions are designed to deliver measurable outcomes on AWS.

For existing customers, this competency reaffirms the Dynatrace commitment to continued innovation alongside AWS. This ensures the Dynatrace AI-powered observability platform evolves with the latest advancements in AI, future-proofing organizations’ existing investments as generative AI capabilities become core to modern cloud workloads.

For new customers, Dynatrace provides a trusted, proven foundation for observability and AI adoption on AWS. Whether an organization is exploring GenAI for customer engagement, automation, or new digital experiences, Dynatrace ensures these systems are reliable, secure, and optimized at every step.

Graph showing a layered approach to AI observability for agentic AI reliability
The Dynatrace layered approach to AI observability

Looking ahead with AI-powered observability on AWS

As organizations increasingly adopt generative AI, observability becomes a critical enabler. By leveraging Dynatrace causal AI, predictive insights, and seamless AWS integrations, organizations can maintain control over costs, risks, and performance while driving innovation, enhancing competitive advantage, and delivering exceptional customer experiences.

Whether you’re building, scaling, or fine-tuning GenAI application, Dynatrace and AWS Bedrock empower you to transform your observability. With end-to-end visibility into AI workloads, their interactions in full context of your business, and cloud-native applications, you can optimize performance, troubleshoot effectively, and maximize the value of your GenAI investments with greater confidence and precision.

Learn more

The post Dynatrace achieves AWS Generative AI Competency: A new milestone in observability and AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-achieves-aws-generative-ai-competency/feed/ 0
Delivering agentic AI reliability: Why AI Observability is imperative https://www.dynatrace.com/news/blog/agentic-ai-reliability-depends-on-ai-observability/ https://www.dynatrace.com/news/blog/agentic-ai-reliability-depends-on-ai-observability/#respond Wed, 10 Sep 2025 18:12:33 +0000 https://www.dynatrace.com/news/?p=70885 Dynatrace for Executives: AI Observability

As AI investment accelerates, a gap is emerging between ambition and execution. IDC projects1 that by 2028, AI spending will make up 16.4% of total IT expenditures. However, Gartner, Inc.2 predicts over 40% of agentic AI projects will be canceled by end of 2027. Likewise, a CIO survey found that 88% of AI pilots fail […]

The post Delivering agentic AI reliability: Why AI Observability is imperative appeared first on Dynatrace news.

]]>
Dynatrace for Executives: AI Observability

As AI investment accelerates, a gap is emerging between ambition and execution. IDC projects1 that by 2028, AI spending will make up 16.4% of total IT expenditures. However, Gartner, Inc.2 predicts over 40% of agentic AI projects will be canceled by end of 2027. Likewise, a CIO survey found that 88% of AI pilots fail to reach production due to unclear objectives, insufficient data readiness, and a lack of in-house expertise. These findings place the expected return on research and innovation firmly at risk, as organizations invest in bespoke models and agentic AI that lack a clear, scalable outcome.

Nonetheless, another Gartner, Inc. article3 predicts that by 2028, 33% of enterprise software applications will include agentic AI, up from less than 1% in 2024, enabling 15% of day-to-day work decisions to be made autonomously by agentic AI systems. Consequently, a Forrester blog4 predicts that 40% of highly regulated enterprises will combine data and AI governance in a move toward a more integrated, transparent, accountable, and ethically responsible approach to AI.

These trends are not contradictory—they show how the market is searching for the right formula to adopt AI, and specifically agentic AI. Successful agentic AI outcomes are predicated on trust in AI’s reliability, security, and alignment with business goals and strategies to not fall behind competitors. Achieving that trust requires AI-native observability that’s deeply integrated with both data and strategic objectives.

Key insights for executives

  • Every modern cloud-native enterprise project will also be an AI-native project – either because of first party AI or through invoking agentic AI services. Preparing for AI adoption is among the top drivers for cloud strategy and investment. Likewise, 63% of top-performing companies increase their cloud budgets to be able to leverage AI.
  • Visibility into reliability and governance of AI interactions has emerged as a new responsibility for executives to realize the value of AI investments while managing risks. From analyst firms in the US to regulators in the EU – increased oversight, link to business goals and regulation of AI systems has become mandatory.
  • Unifying observability signals with AI-powered analytics provides a strategic advantage for AI transformation. By converging observability and AI, teams can accelerate moving projects from pilot production and advance trust and transparency in AI.
  • Dynatrace sets the standard for cloud- and AI-native software, including tracing and logging of AI behavior, predicting and optimizing AI resource utilization, and protecting from unintended AI behavior through runtime security.
  • Dynatrace delivers unified, full-stack visibility across cloud infrastructure, AI workloads—from chat interfaces and prompts to models, tools, and GPUs running on Kubernetes—plus customer experiences and the business layer, all in a single pane of glass, to confidently deliver advanced, AI-powered cloud-native services via a rapidly growing number of 40+ technologies and integrations with hyperscalers and major agentic frameworks providers.

The rise of AI comes with a rise in complexity—and executive responsibility

Organizations generally find themselves maturing their AI implementations along five phases with growing complexity and risks:

Graph showing the evolution of AI usage
Figure 1. Evolution of AI usage
  1. Prompt engineering (generative Al hype). Single step human language prompts a large language model (LLM) for automated text processing and assistance.
  2. Retrieval augmented generation (embedding Al in digital services). Multi-step prompt engineering and LLM access for customer support, automation, and decision-making.
  3. Fine-tuned models. Additional model(s) put on top of existing ones for increased accuracy and domain-aware responses.
  4. Multimodal GenAI. Combination of various modalities beyond text—such as video, audio, imaging and others—that further increase heterogeneity and processing power of services and their interdependences.
  5. Agentic Al. Multiple AI agents and cloud native digital services intensively interacting with each other to autonomously fulfill a specific goal. Agentic AI can double the number of deployed digital service instances and massively increase IT complexity.

The necessity of AI observability for agentic AI reliability

As the complexity of AI implementations increases, observability becomes an essential feedback channel to properly orchestrate and moderate reliable agentic AI outcomes.

Even the early phase implementations show the need to observe AI, tune experience, manage cost, provide guardrails and govern AI responsibly. As the complexity grows, the risks also increase, making deep, context-rich observability of AI strictly mandatory.

7 important reasons for continuously observing AI

  1. Business value. Validate AI investments against business goals and verify end-user value of AI services. Gain business insights from observability data.
  2. Cost and performance control. Monitor and control expenses and sustainability associated with AI operations and investments.
  3. Security. Increase awareness of interactions among AI services, reducing the risk of hacking and malicious influence. Leverage converged observability and security offerings to minimize risk and cost.
  4. Compliance. Monitor that AI output is ethical, unbiased, and adheres to guardrails for meeting regulatory compliance requirements and providing traceability for audits. Expect high volumes of logs and traces to observe AI behaviors and keep audit trails.
  5. Accuracy. Verify that AI agents function properly and precisely, generating quality output. Use observability to deeply check run-time behaviors and
  6. Reliability. Provide traceability and root-cause analysis to verify AI agent health, scalability, performance, and availability.
  7. Collaboration. Govern communications among agent-to-agent and agent-to-human, and provide the means to keep humans in control to override and take responsibility. Automate events from observability platforms that integrate with enterprise ecosystems.

With Dynatrace, executives can solve one of the biggest challenges of managing return on AI investment: Balancing innovation speed with risk, cost, and value.

Increase AI success with AI Observability from Dynatrace

Graph showing a layered approach to AI observability for agentic AI reliability
Figure 2. The Dynatrace layered approach to AI observability

AI is not a single component. Agentic AI in particular is composed of multiple layers and technologies, each observed within a holistic context. Dynatrace provides complete coverage of all layers that allows teams to observe the complete AI stack of modern cloud- and AI-native applications. The layers consist of the following:

  • Business – track outcome: does it create productivity gains, does it deflect support tickets, does it act autonomously and is the investment worth it
  • Infrastructure – utilization, saturation, errors
  • Models – accuracy, precision/recall, explainability
  • Semantic caches and vector databases – volume, distribution
  • Orchestration – performance, versions, degradation
  • Agentic layer – autonomous agents, MCPs
  • Application health – availability, latency, reliability

Dynatrace automatically observes and analyzes complex multicloud and agentic AI systems. By securely unifying and storing all data in context, the Grail® data lakehouse with massively parallel processing unifies all data signals with full context and is continuously updated by Dynatrace Smartscape® real-time dependency mapping technology.

Davis® AI combines predictive, causal, and generative AI to provide deterministic answers and insights, which drive AutomationEngine actions and inform teams with recommendations to optimize productivity, performance, and cost. With these advantages, teams can embrace AI with confidence, make better decisions faster, and innovate at speed—without compromising trust, performance, reliability, or control.

Figure 3. Dynatrace large observability and security coverage of AI technologies keeps growing fast
Figure 3. Dynatrace large observability and security coverage of AI technologies keeps growing fast

Why Dynatrace for reliable agentic AI projects

Top Fortune 500™ organizations use Dynatrace to not only maximize return on investment (ROI) in AI technologies, but across their cloud- and enterprise stacks. Dynatrace leverages partnerships with hyperscalers and major AI framework providers to provide customers with observability for the latest technologies in this fast-moving space.

The recent announcement of our collaboration with NVIDIA is an example of our commitment to providing differentiated AI observability. Dynatrace AI observability delivers real-time, end-to-end observability into AI and LLM workloads—from infrastructure and applications to model performance and end-user experiences. This empowers enterprises to accelerate innovation, ensure compliance, and confidently scale mission-critical AI, all while maintaining reliability and efficiency across their cloud environments.

_____________________________________________

1 IDC Market Forecast, “Worldwide Artificial Intelligence IT Spending Forecast, 2024–2028,” October 2024, https://my.idc.com/getdoc.jsp?containerId=US52635424&pageType=PRINTFRIENDLY.

2 Gartner Press Release, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027,” June 25, 2025, https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027

GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.

3 Gartner Article, “Intelligent Agents in AI Really Can Work Alone. Here’s How.,” by Tom Coshow, October 01, 2024, https://www.gartner.com/en/articles/intelligent-agent-in-ai.

4 “Predictions 2025: An AI Reality Check Paves The Path For Long-Term Success,” Forrester Research, Inc., by Jayesh Chaurasia and Sudha Maheshwari, October 22, 2024, https://www.forrester.com/blogs/predictions-2025-artificial-intelligence/.

The post Delivering agentic AI reliability: Why AI Observability is imperative appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/agentic-ai-reliability-depends-on-ai-observability/feed/ 0
OpenTelemetry and Dynatrace: Complete unified observability analytics for modern applications https://www.dynatrace.com/news/blog/opentelemetry-and-dynatrace-the-complete-analytics-platform-for-modern-observability/ https://www.dynatrace.com/news/blog/opentelemetry-and-dynatrace-the-complete-analytics-platform-for-modern-observability/#respond Thu, 21 Aug 2025 15:58:26 +0000 https://www.dynatrace.com/news/?p=70856 Dynatrace and OpenTelemetry

The freedom to choose your observability stack matters. Whether you're standardizing on OpenTelemetry (OTel) for maximum flexibility and team autonomy, future-proofing your architecture, or simply gaining control over your telemetry pipeline, the choice is yours to make. But here's the reality check every engineering team faces: collecting telemetry data is just the beginning. The real question is, what happens next?

The post OpenTelemetry and Dynatrace: Complete unified observability analytics for modern applications appeared first on Dynatrace news.

]]>
Dynatrace and OpenTelemetry

OpenTelemetry excels at capturing data from any environment and service: traces flowing from microservices, metrics streaming from containers and infrastructure hosts, and logs capturing the application and service lifecycles. OTel does exactly what it was designed to do: standardize telemetry collection. But here’s what OpenTelemetry doesn’t do by design: unified observability that analyzes that data, correlates it across services, or turns it into actionable, intelligent insights.

This is where most organizations hit a wall and require lots of expert knowledge. Organizations today have more observability data than ever before, but somehow less visibility into what’s actually happening in their systems. Raw telemetry data becomes a burden rather than an asset. Engineers spend more time hunting through dashboards than solving actual problems.

This is the “analytics gap” that Dynatrace was built to solve, transforming your telemetry data from scattered signals into unified, AI-powered intelligence.

Why OpenTelemetry + Dynatrace changes everything

Here’s what makes this combination powerful: OpenTelemetry gives you standardized data collection. Dynatrace gives you intelligent analysis that goes far beyond the static dashboards and manual correlation work that other platforms require.

While many observability solutions leave it to you to build custom dashboards and manually connect the dots between your telemetry signals, Dynatrace transforms your OpenTelemetry data into insights that actually drive decisions. Your traces, metrics, and logs aren’t just stored; they’re automatically correlated, analyzed, and contextualized as they flow through Dynatrace OpenPipeline®.

When a trace shows latency spikes, you immediately see related log entries and metric anomalies automatically contextualized and correlated. Lightning-fast queries via Dynatrace Grail® data lakehouse process millions of spans at the speed of thought, making observability accessible to your entire team, not just the experts who know how to build complex queries and visualizations.

The result? Your OpenTelemetry investment becomes a competitive advantage, not just another data collection project that requires a team of dashboard architects to maintain.

Complete OpenTelemetry coverage

Here’s how Dynatrace helps you to get the most out of your telemetry data, without requiring additional agents or complex configurations:

Distributed tracing excellence

Native OpenTelemetry tracing delivers superior span and trace processing with dynamic visualization tools that transform complex distributed architectures into complete end-to-end visibility. But it doesn’t stop there; all your telemetry signals (logs, traces, and metrics) correlate seamlessly, giving you clear, actionable insights within the full context of your traces and services. Get simple answers to advanced questions by expanding your investigations with DQL for powerful analytics, including correlation of logs and traces.

Interactive trace waterfall view showing end-to-end request flow
Figure 1. Interactive trace waterfall view showing end-to-end request flow

Service monitoring that understands your Architecture

Comprehensive service health monitoring built on OpenTelemetry standards. Dynatrace provides intelligent service analysis, anomaly detection, and visualization that work seamlessly with your OpenTelemetry-instrumented applications. Our service monitoring goes beyond simple health checks.

When issues arise, you see exactly which services are affected and how problems cascade through your architecture, all without manual tagging, configuration, or service discovery setup with YAML files.

See exactly which services are affected and how problems flow through your architecture.
Figure 2. See exactly which services are affected and how problems flow through your architecture.

Intelligent metrics with full context

You get flexible metric ingestion for custom business metrics and standard application performance indicators. Your metrics connect directly to the services and traces that generated them. But here’s where Dynatrace takes it further: we allow you to unify all your OpenTelemetry signals into comprehensive service intelligence. Instead of analyzing metrics in isolation, you see how they connect to actual service behavior, request flows, and application logs. Every metric becomes part of a complete service story.

Full context in one service view
Figure 3. Full context in one service view

Complete log processing

Your OpenTelemetry logs are transformed from noise to narrative. Instead of searching through endless log streams and manually created dashboards, Dynatrace supports a comprehensive log ingestion and analysis pipeline, allowing you to go big with Dynatrace.

Every log event becomes part of a larger story about user journeys, interactions, services, and app behavior, all focused on your desired business outcomes and incident investigations.

Here’s where it gets powerful: you automatically get additional contextual enrichment when you direct all your telemetry signals to Dynatrace. By creating bi-directional relationships between logs and traces, where logs provide context to traces and traces illuminate relevant logs, Dynatrace evolves troubleshooting from detective power-user work into AI-driven, streamlined, and intuitive investigations.

Traces to logs video thumbnail
Video: See the full story behind every trace with correlated logs.

Kubernetes native support

For teams running OpenTelemetry in Kubernetes, Dynatrace delivers enterprise-grade support that scales with your cloud native operations. Native Kubernetes handling of spans, metrics, and logs from your Kubernetes OpenTelemetry deployments automatically collects Kubernetes context for automated enrichment: namespace, cluster, and workload relationships, all without any additional instrumentation. Your existing Kubernetes labels, AWS tags, and Azure tags become first-class filtering dimensions for all OpenTelemetry data, enabling automatic cost attribution and comprehensive data permissions using your existing RBAC patterns.

The result is that your OpenTelemetry observability inherits the same operational patterns, security boundaries, and cost structures as your Kubernetes infrastructure.

OTel spans and logs are automatically enriched with Kubernetes context.
Figure 4. OTel spans and logs are automatically enriched with Kubernetes context.

Why this matters for your team

Every organization adopting OpenTelemetry faces the same challenge: turning data collection into intelligent insights. The engineering teams that succeed are those that choose analytics platforms built specifically for OpenTelemetry data.

Dynatrace transforms your OpenTelemetry investment from a data collection project into a competitive advantage. We meet you where you are. We respect your choice to standardize OpenTelemetry by simplifying its operational complexity and enhancing it with analytics that actually deliver value.

Ready to transform your OpenTelemetry data?

Open standards have clear benefits. Industry standardization makes it easier to make sense of data coming from multiple different sources, whether it’s traces, metrics, logs, or telemetry from third-party tools. Your analytics platform can deliver intelligent insights across your entire technology stack. Your OpenTelemetry investment deserves analytics that reveal its full potential.

  • Want to explore specific OpenTelemetry capabilities with Dynatrace? Try them out on the Dynatrace Playground
  • Boost your productivity with these quick video guides for service owners working with OpenTelemetry:

Video: Easy access to your OTel Logs and Traces
Video: Analyze Service Failure from OTel Data

Video: Analyze Service Failure from OTel Data
Video: Analyze Service Failure from OTel Data

Video: Easy access to your OTel & Prometheus Service Metrics
Video: Easy access to your OTel & Prometheus Service Metrics

Join us at OpenSource Summit. We’ll be in Amsterdam August 25-27. Stop by our booth to see the magic in action!

The post OpenTelemetry and Dynatrace: Complete unified observability analytics for modern applications appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/opentelemetry-and-dynatrace-the-complete-analytics-platform-for-modern-observability/feed/ 0
Dynatrace 3rd-generation platform: Built for the world of Autonomous Intelligence https://www.dynatrace.com/news/blog/dynatrace-3rd-gen-platform/ https://www.dynatrace.com/news/blog/dynatrace-3rd-gen-platform/#respond Tue, 22 Jul 2025 06:45:50 +0000 https://www.dynatrace.com/news/?p=70120 Dynatrace paving the way to autonomous intelligence

The world has become software-defined, distributed, and complex, creating a widening gap between digital complexity and our ability to manage and govern the systems that run businesses and organizations. To close this gap, we reimagine how observability works. It is no longer enough to collect and analyze telemetry after the fact. Organizations need trusted, intelligent […]

The post Dynatrace 3rd-generation platform: Built for the world of Autonomous Intelligence appeared first on Dynatrace news.

]]>
Dynatrace paving the way to autonomous intelligence

The world has become software-defined, distributed, and complex, creating a widening gap between digital complexity and our ability to manage and govern the systems that run businesses and organizations. To close this gap, we reimagine how observability works. It is no longer enough to collect and analyze telemetry after the fact. Organizations need trusted, intelligent systems that turn real-time data into reliable knowledge, apply advanced AI to reason through that knowledge, and take action to optimize outcomes at every level of the business.

This is the foundation of the Dynatrace 3rd-generation platform. We’ve spent the past two decades shaping the observability market. Today, we are transforming it from a rear-view mirror into a real-time control system for the modern enterprise. Thousands of organizations are already using Dynatrace 3rd generation to turn data into decisions and decisions into action. The result is faster innovation and stronger business results across every layer of the business.

A new model built on knowledge, reasoning, and actioning

Dynatrace 3rd generation introduces a new standard for observability and automation based on three foundational capabilities:

  • Knowledge: The Dynatrace platform turns petabytes of real-time data into a continuously updated, queryable knowledge graph. Powered by Grail and Smartscape, it provides trustworthy, fact-based insights with real-time context across dynamic environments.
  • Reasoning: Causal, predictive, and generative AI models work together to derive intelligent decisions. These models are context-aware, transparent, and built for enterprise-grade safety and compliance.
  • Actioning: Dynatrace enables users to define goals and let intelligent automation determine the best path forward, through innovations like AutomationEngine, AppEngine, and OpenFeature. This shifts operations from reactive remediation to preventive operations and continuous improvement.

Together, these capabilities form the foundation for autonomous intelligence. Dynatrace doesn’t just provide visibility; it enables systems to understand and act. By continuously converting real-time data into trustworthy insights, applying AI to reason through business and technical context, and triggering intelligent, goal-based actions, Dynatrace transforms observability into a real-time engine for automation and impact.

This is not about removing humans from the loop. It’s about empowering teams to define outcomes and rely on the system to carry out the best path forward. As with any leadership decision, autonomy depends on the quality of information and confidence in its context. The same principle applies to AI systems. Dynatrace 3rd generation gives organizations confidence, allowing them to scale decision-making with speed and trust.

Trusted knowledge, not just data

Legacy observability platforms focus on collecting telemetry data. But to support real-time decisions, teams need a trusted knowledge foundation. Dynatrace eliminates silos between metrics, traces, logs, events, user sessions, and security signals by unifying them in Grail, our schema-on-read, massively parallel data lakehouse.

For the first time, users can run any query at any time, with Grail supporting dramatically higher concurrency than traditional observability platforms. There’s no cold storage, no indexing, and no need for rehydration. This unlocks a goldmine of observability data and turns it into reliable, real-time answers.

That data is then contextualized in real time by Smartscape, our dynamic topology engine, and made instantly accessible to AI agents. The result is not just visibility, but deep, evolving system knowledge, providing machine-speed decisions no other platform can match.

AI that reasons with real-time context

Dynatrace has long set the standard for causal AI in observability. With the 3rd generation platform, we expand that foundation by combining causal AI with predictive and generative models. These AI types work together to support decisions at machine speed, with full context. Whether it’s automatically identifying the root cause of a service degradation, forecasting capacity needs, or evaluating how to improve online customer experiences, Dynatrace AI operates with the reliability and transparency required in enterprise environments.

Now, organizations can pursue modernization, transformation, and agentic AI initiatives with greater confidence. Dynatrace helps make AI accessible and actionable by reducing friction, delivering answers precisely when and where they are needed. Davis CoPilot enables natural language queries, workflow generation, and seamless integration into IDEs. It provides intelligent assistance at every step and supports a broad range of use cases across observability, security, and business operations with explainability, precision, and trust.

Automation that adapts to your goals

Traditional automation is limited by what is explicitly scripted. Dynatrace has taken a different approach. With the 3rd generation platform, you define high-level goals, and the platform determines the best way to achieve them. By grounding automation in real-time, high-quality data and precise causal analytics, Dynatrace ensures that actions are driven by accurate understanding, not assumptions.

This goal-based automation can resolve incidents, optimize performance, reduce cost, and even generate pull requests that improve code quality.

For example, preventive cloud operations allow site reliability engineers to move from firefighting to strategic orchestration. Instead of chasing alerts, SREs can focus on managing service-level objectives and improving business outcomes. Dynatrace handles the rest.

Built for the future of cloud and AI

The complexity of cloud-native architectures, Kubernetes deployments, and emerging agentic AI models are already testing the limits of traditional observability. Dynatrace 3rd generation is designed for the future.

By unifying telemetry, security data, and business context into a single real-time graph powered by Grail, Dynatrace provides the AI-powered intelligence required to operate modern systems with confidence. And by embedding automation throughout the platform, teams can scale faster than headcount, without sacrificing control or trust.

Turning observability into intelligent action

Organizations like TELUS and Air France-KLM are already seeing results: faster resolution, improved resiliency, and reduced downtime.

“By combining our Agentic AI initiatives with Dynatrace’s AI Observability capabilities, we’ve successfully optimized our development and operations workflows. We’re driving innovation and delivering measurable business impact while reducing downtime.”

– TELUS

“The AI and predictive capabilities from Dynatrace were a differentiator. We’re confident that any problem that arises can be dealt with quickly, dramatically reducing operational and revenue impact.”

– Air France-KLM

The combined impact of agentic AI initiatives and AI-powered observability extends beyond IT. These outcomes drive business performance, from greater availability and productivity to better customer experiences.

What this means for your organization

The Dynatrace 3rd-generation platform defines the path forward to a future where software can understand, reason, and act. It helps your organization move from reactive to proactive, from fragmented tools to unified intelligence, and from scripted automation to AI-driven operations.

Whether you are focused on cloud modernization, application security, cost optimization, or AI governance, Dynatrace provides a foundation to build and scale with confidence.

This is the evolution of observability. One built on context, driven by reasoning, and capable of taking action.

Learn more and see what’s possible with Dynatrace 3rd generation.

The post Dynatrace 3rd-generation platform: Built for the world of Autonomous Intelligence appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-3rd-gen-platform/feed/ 0
Dynatrace Cloud Security and CADR: Revolutionizing cloud security with observability context https://www.dynatrace.com/news/blog/revolutionizing-cloud-security-observability-cadr/ https://www.dynatrace.com/news/blog/revolutionizing-cloud-security-observability-cadr/#respond Mon, 21 Jul 2025 13:00:18 +0000 https://www.dynatrace.com/news/?p=70068 Dynatrace Log Management & Analytics graphic

Traditional security tools can’t keep up with today’s cloud- and AI-native environments. Built for static corporate IT, they struggle with the highly dynamic, short-lived workloads behind modern digital services. Securing critical applications within each organization’s digital environments requires a new approach: one that protects production in real-time and brings end-to-end observability context into AI-powered security […]

The post Dynatrace Cloud Security and CADR: Revolutionizing cloud security with observability context appeared first on Dynatrace news.

]]>
Dynatrace Log Management & Analytics graphic

Traditional security tools can’t keep up with today’s cloud- and AI-native environments. Built for static corporate IT, they struggle with the highly dynamic, short-lived workloads behind modern digital services. Securing critical applications within each organization’s digital environments requires a new approach: one that protects production in real-time and brings end-to-end observability context into AI-powered security analytics. AI-powered Dynatrace Cloud Security meets this need and corresponds to what the market recognizes as Cloud Application Detection and Response (CADR).

Why traditional security falls short for modern workloads

Incomplete security controls: Our analytics shows that many organizations still have blind spots, risking critical vulnerabilities, and are exposed to significant attack paths despite the multiple existing security controls and tools in place. For example, 50% of Fortune 500 companies are still vulnerable to Spring4Shell vulnerability. Applications—often the main revenue drivers—remain one of the top sources of risk, frequently exploited for initial access.

Outdated compliance practices: Traditional quarterly or annual audits are becoming obsolete; auditors now validate continuous compliance, even between scheduled audits. Compliance is tightly connected to correctly configured systems, demanding real-time configuration analytics, specifically for rapidly changing cloud environments.

Rules and regulations on breach hygiene: Modern regulations (e.g., GDPR requirements in Europe or SEC in the US) demand near-instant reporting of breaches or even suspected breaches within 48 to 72 hours. Organizations need continuous, real-time insights and monitoring, not delayed, point-in-time reports derived from static log archives.

Ineffective threat detection: XDRs and SIEMs (predominantly optimized for corporate assets such as laptops, phones, email, and Office365) often lack the deep, real-time runtime visibility needed for dynamic environments like containers, microservices, and serverless functions. Gartner reviews highlight a gap in real-time, context-aware visibility for dynamic environments. These workloads require tailored, context-aware detections and the ability to investigate across ephemeral components that disappear within minutes, closing the coverage blind spot.

CADR: Security for cloud- and AI-natives

As organizations shift to the cloud, the focus has increasingly centered on securing containerized applications and microservices: the core of modern digital services. Security teams need real-time runtime protection that empowers them to take immediate, autonomous action. Unlike traditional tools that generate isolated alerts, CADR provides rich application context, enabling security operations to understand the full story behind an incident, from exploitability to understanding the impact of the attack. While the CADR market is still emerging, it represents the need to focus on applications and their underlying infrastructure at runtime. This represents a natural evolution and strategic refinement of the Cloud Native Application Protection Platform (CNAPP) category.

Application and SRE teams need to be able to make cloud security alerts operational by simplifying incident response and enabling the SOC to act with clarity and speed. Numerous tools address the various security threats and their types. However, the solution isn’t to pile up even more tools to cover each threat. Rather, it’s smarter analytics of unified data brought into context. This is where the convergence of observability and security brings the foundational difference, through end-to-end coverage including real-time analytics for logs, traces, user behavior, security events, topology, and more.

Security isn’t solely about acquiring a SIEM; it’s about effectively addressing specific challenges. Often, SIEM is a broadly used term, obscuring the true requirements of modern cloud environments.

CADR transforms this discussion: shifting from static log collection to achieving dynamic, interconnected security outcomes spanning vulnerability management, workload protection, compliance, and automated response.

The Dynatrace approach isn’t merely about replacing a SIEM; it’s about empowering organizations to ask more pertinent questions about their business-specific use cases and gain actionable insights.

So is Dynatrace a SIEM? The answer is yes—and then we delve deeper into your specific needs and desired outcomes.

The biggest barrier to effective threat response in the cloud is the lack of unified context. Application and security teams often operate without full visibility into how threats impact the broader digital service environment, business objectives, or operational ownership. Making cloud security truly actionable requires converging it with observability. Only by combining deep, real-time insights into application behavior, context and topology information, infrastructure performance, and user interactions can organizations prioritize threats accurately, identify the right teams to respond, and automate remediation with minimal human intervention. This convergence ensures the speed and precision needed to reduce risk before it escalates. It also provides entirely new indicators of compromise, impossible without observability context.

The true value of leveraging a unified platform is coverage across the attack steps from initial access, over lateral movement, to exfiltration, with the additional benefit of coverage for MITRE ATT&CK as well as MITRE ATLAS. Dynatrace implements a layered security approach by leveraging full-stack observability, real user monitoring, and automatic log collection to evolve how organizations identify indicators of compromise and achieve comprehensive coverage. These efforts are further empowered by analytics using Dynatrace Query Language (DQL) on Grail and AutomationEngine. By that, Dynatrace not only provides security findings across the full stack but adds response automation, threat detection, and investigation on top of a combined security, observability, and threat intel data set. The Dynatrace MCP server makes runtime findings accessible for both agentic AI remediation automation and information distribution. This brings, for example, runtime vulnerability remediation into the developer’s IDE.

Secure your cloud with Dynatrace. Start your 15-day free trial today.

Convergence of observability and security enables improved CADR

Leveraging the abilities of AI-powered unified observability and security, Dynatrace CADR integrates the following three critical security capabilities to secure modern applications and their infrastructure:

Threat Detection & Investigation (TDI)

  • Leverages our powerful Grail data lakehouse, Logs app, and Security Investigator capabilities to analyze security and observability data in full context. Dynatrace log management enables seamless ingestion, indexing, and querying of massive volumes of log data with lightning-fast performance and low overhead. Combined with ingested threat intelligence data, ingested third-party findings, and Dynatrace’s own security findings, this empowers real-time, high-fidelity threat detection and investigation, even across ephemeral and dynamic workloads. Together, these capabilities facilitate proactive threat hunting, deep forensics, and accelerated root cause analysis.

Runtime Vulnerability & Exposures Analytics (RVA) + Runtime Application Protection (RAP)

  • Pinpoints and prioritizes vulnerabilities and exposures in real time across applications, infrastructure, and operating systems.
  • Blocks malicious traffic from within the application at runtime, using rich observability context to trace threats from entry to impact and prevent exploitation (RAP).

Security Posture Management (SPM)

  • Continuously detects misconfigurations, standards compliance violations, and security policy issues that attackers exploit for persistence and privilege escalation.
  • Supports modern compliance needs by providing real-time status and reporting capabilities, crucial for adhering to strict reporting deadlines required by regulations.

+ Agentic AI response automation & integration

  • Automates remediation workflows through seamless integration with CI/CD pipelines and ITSM tools, reducing the window of exposure and operational friction.
  • Enables operationalization in organizations, allowing developers to fix vulnerabilities (Dev, e.g. connecting into the IDE via the Dynatrace MCP server), SREs to fix config issues, and SecOps teams to create detections and act on findings, all within familiar contexts and workflows.

+ AI-powered contextual risk prioritization

  • Dynatrace’s causal AI, when leveraged for security solutions, understands the risks thanks to the vector graph, Smartscape, that prioritizes issues based on real-time topology knowledge and access to the various attack paths. It also leverages intelligent automation for tasks such as automatically disqualifying false positives and automatically dispatching positives of vulnerabilities to responsible development teams for remediation.

Dynatrace cloud security for cloud application detection and response (CADR)

A realistic attack path, and how Dynatrace stops it

Attackers typically follow a path when targeting modern applications.

  1. Initial access: Attackers often gain access by exploiting application vulnerabilities. Dynatrace already monitors these applications for availability and performance; adding security capabilities is a simple extension. Runtime Vulnerability Analytics (RVA)/Runtime Application Protection (RAP) identifies and prevents exploitation at runtime.
  2. Persistence & privilege escalation: Attackers use misconfigurations or compliance gaps to maintain access and escalate privileges within the environment. Dynatrace Security Posture Management (SPM) identifies these risks in real time.
  3. Discovery & exfiltration: Attackers discover sensitive data and attempt to exfiltrate it. Dynatrace Threat Detection and Investigation (TDI) detects and investigates suspicious behaviour using deep observability context, providing full traceability.

By integrating RVA/RAP, SPM, and TDI, Dynatrace CADR offers a comprehensive, unified approach that maps directly to these attack stages, allowing you to detect, respond to, and manage security risks effectively across your modern workloads.

Final thoughts: Why CADR is the logical next step

  • Organizations can’t secure what they can’t see. Dynatrace not only delivers complete visibility into digital environments but also maps the dependencies between assets, providing critical topology insights. By leveraging layered security insights and converging observability and security, Dynatrace makes securing modern applications actionable, automated, and accessible to the teams who already use Dynatrace for observability, having security deeply integrated.
  • Whether consolidating basic log management or needing advanced runtime threat detection, Dynatrace offers the platform to address these needs.
  • CADR is the natural next step for any Dynatrace customer running modern workloads and the fastest path to proactive, real-time cloud application security for any organization facing the unique challenges of the modern attack surface.

The post Dynatrace Cloud Security and CADR: Revolutionizing cloud security with observability context appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/revolutionizing-cloud-security-observability-cadr/feed/ 0
Logs and traces: Why context is everything for seamless investigations https://www.dynatrace.com/news/blog/correlating-logs-and-traces-with-observability/ https://www.dynatrace.com/news/blog/correlating-logs-and-traces-with-observability/#respond Fri, 04 Jul 2025 10:05:27 +0000 https://www.dynatrace.com/news/?p=69744 logs and traces

It’s 3:00 AM. Alerts are firing. Something’s broken, latency is spiking, there’s too much noise, and you’re under pressure to find the root cause fast. You need to be able to understand how your system interacts to solve the problem as soon as possible. But systems just keep getting more complex as your organization adds […]

The post Logs and traces: Why context is everything for seamless investigations appeared first on Dynatrace news.

]]>
logs and traces

It’s 3:00 AM. Alerts are firing. Something’s broken, latency is spiking, there’s too much noise, and you’re under pressure to find the root cause fast. You need to be able to understand how your system interacts to solve the problem as soon as possible. But systems just keep getting more complex as your organization adds new technologies, AI models, and container-based microservices. That’s why, as complexity scales, so does the need for connected insights. In the world of observability, logs and traces serve distinct but complementary purposes. When used together, they unlock a powerful view into system health, performance, and behavior.

The secret lives of logs and traces

In theory, correlating logs and traces should be straightforward. However, in practice, teams often find themselves context-switching to follow the path of a trace and all the logs involved. To understand why, let’s take a closer look at the roles and responsibilities of logs and traces.

Traces: The big picture view

Traces follow the journey of a request as it moves through various services in a distributed system. They provide end-to-end observability of how different components interact, making them ideal for understanding latency, bottlenecks, and service dependencies. Distributed traces connect events into a cohesive timeline, helping engineers see how one service’s performance affects others.

Traces shine when you’re trying to answer questions like, “Where did this request slow down?” or “Which service caused the failure?”

Logs: The detailed detective work

Logs are detailed, timestamped records of events generated by applications and infrastructure. They’re rich in context, often containing error messages, debug information, and custom outputs that developers write into the code. While traces show the flow, logs show the details. Logs can exist independently of traces and are often the first place developers look when something goes wrong.

Logs shine when you’re trying to answer: “What exactly happened here?”

Don’t forget metrics and other telemetry signals

Although we’re focusing here on logs and traces, metrics and other telemetry data are also essential for observability and deeper context. For more about why it’s important to unify the full spectrum of observability signals, see What is observability and Unified observability: Why storing OpenTelemetry signals in one place matters.

The power of correlating logs and traces from a single, full-context platform

Isolated telemetry signals can lead to blind spots and wasted time searching for answers. Some of the main ways to use logs and traces are to simplify troubleshooting, enhance performance, improve security posture, and meet compliance standards.

When you can correlate logs and traces from a single source of observability data, you eliminate the constant context switching that slows down investigations. Instead of toggling between tracing tools and log viewers, you get a unified view that connects the dots fast.

Correlating logs and traces from a single platform transforms troubleshooting from a fragmented hunt into streamlined analysis, where you spend time solving problems instead of searching for information. Core technologies like Grail®, OneAgent®, and Davis® AI provide the scalable foundation while embracing open-source frameworks like OpenTelemetry for flexibility.

Cracking the case of the failed checkout: Investigating logs and traces

Not every investigation starts the same way. Sometimes a trace gives you the high-level view you need to spot an issue and dive deeper. Other times, a log entry is the first clue that something is off. In the next section, we’ll walk through two examples, one that starts with traces and the other with logs, to show how you can get the answers you need.

Scenario 1: Investigating from traces to logs

Investigating from traces to logs in Dynatrace video

While doing some routine monitoring in the Distributed Tracing app, we notice a series of failed requests in our Kubernetes prod namespace. So we filter for unsuccessful transactions to examine them more closely.

One request stands out: “/cart/checkout”. It’s a critical transaction path, and we’re seeing failures.

We dive into the trace waterfall. Just below it, we find the logs tied to each span, giving us deeper insight. That’s where we find the message:

error: failure to complete the order

Distributed Tracing requests in Dynatrace screenshot

Digging further, another log reveals the root cause: only Visa and Mastercard are accepted, which is in line with our policy, but potentially limiting our business. This raises a new question: how often is this happening?

With a single click, we pivot to the logs app, where we can search for this specific message and quantify how many transactions may have been impacted.

This approach turns scattered signals into a cohesive story, helping us move from surface-level symptoms to actionable insights with speed and precision.

Scenario 2: Investigating from logs to traces

Logs to Traces video thumbnail

No matter how you start your day, whether you are coming from PagerDuty, Slack or start directly in Dynatrace through one of the many apps like Kubernetes or the Clouds app, you can always see logs in context of your investigation.

In this scenario, we’re investigating this case from another angle, starting with the logs app using the prefiltered segment for the Kubernetes prod namespace. The view is tailored to the services we own. A quick scan reveals something suspicious: numerous errors in some of the log files.

screenshot of logs affected by errors in logs and traces investigation
Figure 1. A quick scan reveals numerous errors in some log files.

To dig deeper, we navigate in the logs app and use the content filter for “payment” and “error”, and we find several logs with the following message:

Could not charge card for user id = xxxxxxxxxxxxx

But what is causing the failure? We click Show surrounding logs, which reveals all logs associated with the trace ID. Now we can view log messages sequentially as they happened.

Investigating some of the surrounding logs, we see that the user is using a credit card other than Visa or Mastercard, which our organization doesn’t support. Now that we understand why things are failing, let’s investigate further to see if we can optimize this experience.

To understand the full impact, we pivot seamlessly to the trace view. Here, we see the full waterfall breakdown of the request: service calls, timing, and span-level metadata.

One detail stands out: it took 5 seconds for the user to receive the failure message. That’s a long time to wait just to be told their card isn’t supported.

With this insight, we can now make targeted improvements so that the user does not have to wait a long time to understand that their payment method is not supported and deliver a better user experience.

While these examples highlight how seamless navigation between logs and traces accelerates troubleshooting, they’re just one part of the story. With Dynatrace Grail and Notebooks, you can take things a step further by running advanced queries, automating repetitive tasks, and building collaborative, data-rich workflows. These tools empower teams to go beyond reactive troubleshooting and into proactive, scalable observability.

Why seamless navigation between logs and traces matters

Seamless navigation between logs and traces isn’t just a convenience; it’s a game-changer. Whether you start with a trace or a log, the ability to pivot instantly between signals means you spend less time hunting for answers and more time solving problems. It accelerates root cause analysis, improves team collaboration, and gives you the full context needed to act with confidence. This is just one example of how Dynatrace helps you move from fragmented troubleshooting to unified intelligent observability.

Ready to start investigating?

Explore Distributed Tracing and Log Management and Analytics, complete with prepopulated data in the Dynatrace Playground.

Want to get started with your own data instead?

The post Logs and traces: Why context is everything for seamless investigations appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/correlating-logs-and-traces-with-observability/feed/ 0
The rise of agentic AI part 3: Amazon Bedrock Agents monitoring and how observability optimizes AI agents at scale https://www.dynatrace.com/news/blog/ai-agent-observability-amazon-bedrock-agents-monitoring/ https://www.dynatrace.com/news/blog/ai-agent-observability-amazon-bedrock-agents-monitoring/#respond Tue, 17 Jun 2025 19:00:54 +0000 https://www.dynatrace.com/news/?p=69101 AWS icon and agentic AI

Model-building platforms like Amazon Bedrock provide the foundation for successful agentic AI applications. But effective cross-agent communication requires standardized telemetry. In this third installment of our series, The Rise of Agentic AI, we explain how standardizing and instrumenting tracing and logging, and monitoring Amazon Bedrock Agents helps to debug and deliver better performing agentic AI applications.

The post The rise of agentic AI part 3: Amazon Bedrock Agents monitoring and how observability optimizes AI agents at scale appeared first on Dynatrace news.

]]>
AWS icon and agentic AI

The next big wave in artificial intelligence is agentic AI, which harnesses autonomous agents to perform tasks by reasoning, learning, and adapting to changing circumstances. The success and efficiency of agentic AI systems depend on how well these AI agents communicate. Facilitating this communication requires monitoring AI agents and their underlying communication protocols, such as Model Context Protocol (MCP).

In this blog post, we explain how to monitor Amazon Bedrock Agents and how observability optimizes AI agents at scale.

Key takeaways
  • Effective cross-agent communication requires standardized telemetry. For foundational model-building platforms like Amazon Bedrock, OpenTelemetry-based solutions provide standardization and instrumentation for tracing and logging to debug at scale.
  • End-to-end observability is a key best practice for monitoring agentic AI. AI agent observability best practices include using GenAI semantic conventions with traditional logs, traces, and instrumentation.
  • Observability helps deliver effective agentic AI results in the context of the whole stack. AI agent observability and Amazon Bedrock Agents monitoring help deliver better performance, ensure compliance, and provide detailed debugging tools.

Cross-agent communication requires standardized telemetry

Given the non-deterministic nature of large language models (LLMs) and dynamic cross-agent communication, organizations need standardized telemetry. OpenTelemetry-based GenAI semantic convention libraries are emerging to unify logging, metrics, and tracing in multi-agent ecosystems. Likewise, these standardized instrumentation libraries let you collect and analyze data from each step in an agent’s decision or communication chain on Dynatrace. Observability of each step lets you monitor the communications among your agents and evaluate their health and performance, regulatory compliance, and debugging.

Architecture of travel agent application using Amazon Bedrock Agents and monitoring it with Dynatrace through OpenTelemetry
Figure 1. Architecture of travel agent application using Amazon Bedrock Agents and monitoring it with Dynatrace through OpenTelemetry.

Scale and monitor Amazon Bedrock Agents with Dynatrace

Amazon Bedrock Agents provide an easy way to build and scale generative AI applications with foundation models.

Amazon Bedrock is a fully managed service that offers a choice of high-performing foundation models from leading AI companies, such as AI21 Labs, Anthropic, Cohere, Luma, Meta, Mistral AI, poolside, and Stability AI—or from Amazon’s own model, Amazon Nova—all through a single API. In addition, Amazon Bedrock Agents also provide the broad set of capabilities teams need to build generative AI applications with security, privacy, and responsible AI best practices.

Dynatrace provides an AI-powered, unified observability and security solution for tracking and revealing the full context of used technologies and service interaction topology. Using Dynatrace for AI agent monitoring and MCP monitoring, teams can analyze security vulnerabilities and observe metrics, traces, logs, and business events in real time—automatically and securely.

“With the rise of agents, the need for deep visibility and real-time insights is more essential than ever. Through this partnership, AWS and Dynatrace are uniquely positioned to deliver performance, cost, and quality insights alongside robust compliance monitoring—empowering customers to innovate with confidence.”
– Atul Deo, Director of Amazon Bedrock

Best practices for Agent-to-Agent (A2A) and MCP monitoring

architecture diagram that shows multiple agents interacting with an agentic application
Figure 2. Autonomous agent workflows and task execution.

As with hybrid and cloud-based environments, context-based observability of AI agents and models is essential for efficient and healthy outcomes. Here are some best practices:

  • Adopt common semantic conventions. Standardize metrics and trace attributes—for example, gen_ai.agent.operation.name and gen_ai.agent.name—across different frameworks.
  • Use logs and traces for Amazon Bedrock Agents. Log critical task lifecycle events—capability discovery, artifact creation, agent collaboration steps, API calls—so teams can replay and debug complex interactions and detect hallucinations.
  • Instrument thoroughly. Bake observability into agent frameworks using external OpenTelemetry libraries or by manually instrumenting calls. Ensure each agent’s start, stop, and reasoning steps, like tools, knowledge base, and guardrails, are captured consistently.
  • Secure communication. Enforce enterprise-grade authentication and authorization within agent-to-agent traffic. Use well-defined protocols like A2A to avoid unauthorized data exposure.
  • Continuous feedback. Feed observability insights into iterative retraining or fine-tuning for improved agent reliability.
Screenshot of an example trace showing debugging an Amazon Bedrock agent workflow with Dynatrace AI observability.
Figure 3. Debugging an Amazon Bedrock Agents workflow with Dynatrace AI Observability.

With Amazon Bedrock and the Dynatrace AI Observability solution, you can cover the following use cases for agent observability:

Monitor AI agent service health and performance

  • Detect bottlenecks by tracking real-time metrics, including request counts, durations, and error rates.
  • Manage service costs with automated cost calculations for each request.
  • Stay on track with service-level objectives (SLOs).

Monitor guardrails to ensure compliance

  • Monitor your safeguards customized to application requirements and responsible AI policies.
  • Validate toxicity, filtered content, and denied topics to ensure compliance.
  • Prevent leaks of personally identifiable information (PII).
  • Prevent quality degradation by validating models and usage patterns in real time.

End-to-end tracing and debugging

  • Achieve complete visibility of prompt flows, from initial request to final response, for faster root cause analysis.
  • Capture detailed debug data to troubleshoot issues in complex pipelines.
  • Streamline workflows with granular tracing of LLM prompts, including response latency and model-level metrics.
  • Resolve issues more quickly by pinpointing exact problem areas in prompts, tokens, or system integrations.
Dashboard showing Amazon Bedrock agents monitoring details, such as service health, guardrails, and performance debugging
Figure 4: Dynatrace AI Observability for Amazon Bedrock Agents dashboard covering service health, guardrails, performance, and debugging.

Future of AI agent observability and Amazon Bedrock Agents monitoring

We expect to see deeper integrations between agent orchestration protocols (A2A, MCP) and open observability frameworks, delivering end-to-end visibility from data ingestion to cross-agent collaboration. As standards converge, organizations will rapidly compose advanced AI solutions while retaining full transparency and control, paving the way for even greater scalability, resilience, and confidence in autonomous agents.

Read more

  • Part one of the Rise of Agentic AI blog series covers the fundamentals of AI agents, models, and emerging communication standards such as Agent2Agent (A2A) and MCP.
  • Part two explores AI agent observability and monitoring, A2A and MCP communications, and how to scale and monitor Amazon Bedrock Agents.
  • Part four covers full-stack observability for AI with NVIDIA Blackwell and NVIDIA NIM.
  • Part five demonstrates how to build a simple agentic application using the OpenAI Agents SDK and instrument the data with Dynatrace.
  • Part six explores AI Model Versioning and A/B testing for smarter LLM services.
  • Part seven introduces data governance and audit trails for AI services.
For AI agent observability and MCP monitoring at scale, check out Dynatrace AI Observability solution and the observability agent samples from Dynatrace on the AWS Labs GitHub site.

The post The rise of agentic AI part 3: Amazon Bedrock Agents monitoring and how observability optimizes AI agents at scale appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ai-agent-observability-amazon-bedrock-agents-monitoring/feed/ 0