AIOps Archives | Dynatrace news https://www.dynatrace.com/news/category/aiops/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 09 Jul 2026 14:33:36 +0000 en hourly 1 Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/ https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/#respond Mon, 08 Jun 2026 15:03:09 +0000 https://www.dynatrace.com/news/?p=74422 Blog OTP Observability for Agentic AI

Agentic AI is breaking the mold of what organizations need from observability. Fragmented, correlation-dependent observability platforms are no longer “good enough.” Enterprises with dynamic, hybrid environments require observability that provides real-time, precise answers, so AI agents can prevent problems, automate workflows, and deliver better, more secure software.

The post Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI appeared first on Dynatrace news.

]]>
Blog OTP Observability for Agentic AI

As more agentic AI projects come online, the observability market is abuzz with familiar promises: tool consolidation, AI-powered insights, and faster remediation through smarter tools. On the surface, this sounds like progress. But beneath the excitement, many discussions are framed around the wrong question.

The real issue isn’t about how to adopt autonomous operations; it’s about ensuring AI agents are operating reliably and resolving problems without introducing new ones. When evaluating new observability solutions, the question should be:

Can this observability solution accurately analyze complex, dynamic telemetry in context so AI agents can act autonomously with trust, precision, and reliability?

As systems become increasingly agent driven, observability is crossing a structural boundary. Approaches designed for environments where only humans decide and act must adapt to a world where agents increasingly operate autonomously with human oversight, while keeping organizations informed.

Rethinking observability for the agentic age

Observability platforms were initially intended to support engineers in delivering reliable applications, services, and infrastructure to users, and alert them in the event of a problem. Dashboards, alerts, and correlation helped teams investigate incidents, piece together what happened, diagnose issues, decide on next steps, and resolve the problem. This model worked when changes were pushed manually.

The assumption was that more data, better correlation, and cleaner interfaces will lead to increased visibility and improved operational decision making.

Agentic AI systems break that assumption.

With faster release cycles and AI-generated code, manual investigations can no longer keep pace. Moreover, observability platforms must now provide actionable insights to both humans and AI agents.

As agents begin operating as autonomous participants in software environments by triggering mitigations, scaling infrastructure, and optimizing behavior in real time, observability can no longer function solely as a human interface. It must also provide AI systems with a reliable, contextual fact basis that agents can act on programmatically. Machines can’t rely on dashboards and alerts. They require a deterministic foundation of unified, real-time data that delivers accurate, context-rich answers at exabyte scale.

Agentic systems break the mold of “good enough”

Many observability platforms layer probabilistic AI on top of siloed data. They use LLMs to correlate signals and rank likely causes—but they can’t always determine correctness.

“Probabilistic” means that the same input will generate a different output based on a probability distribution of predefined outputs, delivering a different answer when the same problem occurs. This approach is also prone to hallucinations, requiring additional human validation, which can increase operational overhead and token costs, delay resolution of business-critical issues, and divert resources from strategic initiatives.

Enterprise-grade observability must now answer: Is this insight reliable enough for autonomous action?

AI built on siloed data is inherently unreliable. Autonomous systems depend on deterministic, contextual, and trustworthy data to act reliably.

“Deterministic” means that the same input always results in the same output by using factual data to trace the exact causal changes that created the issue. When agentic AI systems act on business-critical applications, the cost of being “mostly right” becomes operationally unacceptable.

This is where a subtle but critical divide appears in the market. Aggregating signals and correlating anomalies can surface patterns. Patterns alone are not a solid basis for decisions, and without deterministic understanding, AI systems inherit that uncertainty and can propagate it downstream.

To drive reliable enterprise autonomous operations, AI agents require a unified, AI-powered observability platform that can analyze exabytes of data in real time and across models to pinpoint root cause, delivering actionable answers in context of what’s affected and its business impact.

From correlated guesses to deterministic answers

This shift in the demands of observability hinges on a clear distinction:

  • Probabilistic AI correlates signals that happened around the same time and therefore appear related, pulling information from fragmented data stores to propose a likely root cause.
  • Deterministic AI uses causal analysis to pinpoint what happened and why, recommend remediation actions, and identify business impact.

Probabilistic AI is intended to narrow the search space and direct engineers toward potential resolution, but it still requires interpretation.

Deterministic AI establishes sequence, dependency, and impact, enabling systems to decide safely without waiting for humans to connect the dots.

Auto‑remediation, auto-prevention, and auto-optimization all depend on this leap. A platform that unifies telemetry only at the UI layer may deliver data and potential root cause, but it can’t compensate for fragmented understanding and missing context underneath. When context is pieced together after the fact, confidence is never guaranteed.

You can’t automate what you don’t precisely understand.

Context driven observability as the control plane for AI

In an autonomous enterprise, observability doesn’t sit beside execution; it’s embedded within it. This integration requires that teams adopt a new mindset toward observability architecture.

Because more AI workloads are happening at the source, telemetry must be optimized and streamlined before ingest, not after the fact, from the edge to the back end. Data access must be unified, context-aware, and always-hydrated on a massive scale. Answers must be explicit, not implicit, and they must be informed by automatic, real-time dependency mapping.

Likewise, intelligence must combine deterministic and agentic AI—not as add‑ons, but as a single reasoning system from ingest to execution.

In this model:

  • AI agents can become the primary consumers of observability data.
  • Humans can shift toward strategy, architecture, oversight, and exception handling.
  • Observability evolves from a reactive lens into a control plane for autonomous operations.

Observability purpose-built for autonomous operations ensures successful agentic AI initiatives

This moment represents an architectural transition, not just an incremental upgrade cycle. Correlation-dependent observability that uses probabilistic AI can be extended, augmented, and rebranded, but it will always carry the limitations of approximation and human validation.

The next era belongs to an observability platform that’s built for machine understanding from the start: a unified, context driven architecture that delivers deterministic answers at machine speed, precision, and scale.

Do you want more data or better decisions? Learn why enterprises are switching to Dynatrace.

The post Beyond correlation to autonomous action: Why “good enough” observability fails in the age of agentic AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/beyond-correlation-to-autonomous-action/feed/ 0
Dynatrace AI agents begin working for you on day one, and are built to grow with you https://www.dynatrace.com/news/blog/dynatrace-ai-agents-begin-working-for-you-on-day-one-and-are-built-to-grow-with-you/ https://www.dynatrace.com/news/blog/dynatrace-ai-agents-begin-working-for-you-on-day-one-and-are-built-to-grow-with-you/#respond Fri, 03 Apr 2026 15:44:42 +0000 https://www.dynatrace.com/news/?p=73625 Agents graphic

AI agents are everywhere in tech conversations right now, but what agents can you actually use today to make your job easier? In Dynatrace, ready-made agents help developers, SREs, and IT operations teams investigate issues, understand system behavior, and reduce manual work using the data they trust every day. Dynatrace ready-made agents are not concepts or previews; they're available now, integrated into existing Dynatrace workflows, and designed to solve real operational problems. For teams ready to go further, Dynatrace agents lay the groundwork for autonomous operations.

This blog shows what Dynatrace ready-made agents are, how to get value from them quickly, and how to decide which agents are relevant for you, using concrete examples rather than promises.

The post Dynatrace AI agents begin working for you on day one, and are built to grow with you appeared first on Dynatrace news.

]]>
Agents graphic

From generic AI to task‑focused operational agents

Dynatrace ready‑made agents are purpose‑built capabilities that apply Dynatrace intelligence to specific, recurring operational tasks. Each agent focuses on a clearly defined problem, such as explaining why a service is slow, summarizing unusual behavior in an environment, or helping you understand what changed and why it matters. These agents are designed to take a question or a signal based on the exact data that is in your environment and organization and turn it into a useful answer you can act on.

Because Dynatrace agents are ready‑made, there is no need to define prompts, train models, or design behavior from scratch. Each agent already knows:

  • What type of input to expect,
  • Which Dynatrace signals and context it should use,
  • And what output types are most useful for each addressed problem type.

All available ready-made Dynatrace agents can be found in Dynatrace Hub.

Trigger agent actions with Dynatrace Workflows and the Dynatrace MCP Server

Ready‑made agents can be triggered automatically as part of Dynatrace Workflows or available wherever you already work via the Dynatrace MCP Server.

Using agents in Dynatrace Workflows

Dynatrace Workflows lets you run agents in response to events or on a schedule. Instead of manually asking questions about potential problems and remediation steps, the workflow autonomously responds to changes in your environment.

For example, the Kubernetes Troubleshooting Agent runs nine parallel queries for data enrichment, and Dynatrace Intelligence turns all the information into a structured diagnosis. Customize the agents to your needs, including instructions for human approval steps and automated remediation.

Dynatrace Kubernetes Troubleshooting Agent in action.
Figure 1. Dynatrace Kubernetes Troubleshooting Agent in action.

The fastest way to get started is with Dynatrace ready-made agentic workflow templates, currently available in a preview release. Instead of building from scratch, you get proven automations that summarize issues, suggest remediation, and deliver insights directly to the tools your teams already use.

Figure 2. Agentic workflow templates available in preview
Figure 2. Agentic workflow templates available in preview

Power users can go further by building their own agentic workflows that combine Dynatrace Intelligence actions with any trigger, data source, or integration in Workflows. Use cases range from auto-scaling Kubernetes clusters based on Dynatrace Intelligence forecasts to generating query-cost-optimization recommendations for stakeholders, to virtually any other automation your environment requires.

Using agents through the Dynatrace MCP Server

The Dynatrace MCP Server makes the agents available outside the Dynatrace web UI, without requiring you to deploy or operate any additional infrastructure. You can connect Dynatrace to any MCP‑compatible client in minutes, with no server to install, host, or maintain.

Through the tools exposed by the MCP Server, you can use natural language to query data in Grail®, check system health, and get problem analyses and remediation recommendations. This brings Dynatrace directly into the tools you already use, such as your IDE, Claude Code and Cowork, Microsoft Copilot, Slack, or automation platforms like n8n. The Dynatrace MCP Server also powers integrations with systems like Azure SRE, AWS DevOps, GitHub Copilot, Atlassian Rovo Ops, Amazon Q, and others.

Dynatrace MCP server in Visual Studio Code with GitHub Copilot
Figure 3. Dynatrace MCP server in Visual Studio Code with GitHub Copilot

This means agents are no longer tied to a single interface. You can ask Dynatrace questions and get grounded, production‑ready answers wherever you work, using the same agents and intelligence that power Assist and workflows.

Dynatrace Assist: a simple way to test ready-made agents

The quickest way to use a ready‑made agent and see how it works before you start creating a workflow is with . Dynatrace Assist lets you ask questions about your environment using natural language, without switching tools or setting anything up.

A simple way to start is with a real problem you already have. For example, when a service becomes slow, open Assist and ask a question such as “Summarize the open problems and highlight those that need immediate attention.” Assist interprets the question, evaluates the environment you’re working in, and pulls together relevant data and context using Dynatrace Intelligence. Instead of manually navigating metrics, traces, logs, and dependencies, you get an explanation grounded in what is actually happening in your system.

Continuing your conversation with Assist, you can refine the question or follow suggested drill‑downs. Assist supports this as a single flow, helping you move from an initial question to deeper analysis and, where applicable, to next steps. You’re not configuring an agent or defining behavior. You’re simply asking a question and letting Dynatrace coordinate the right intelligence and ready‑made agents behind the scenes.

Dynatrace Assist
Figure 5. Dynatrace Assist

This makes Assist your lowest‑friction entry point for using Dynatrace agents. You get a concrete result quickly, using the same data and context you already rely on in your daily work.

What’s next?

If you haven’t already, open Dynatrace Playground, or your Dynatrace tenant, and ask Dynatrace Assist a question to see the ready-made agents in action.

The post Dynatrace AI agents begin working for you on day one, and are built to grow with you appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-ai-agents-begin-working-for-you-on-day-one-and-are-built-to-grow-with-you/feed/ 0
Dynatrace Release Radar 01.26 https://www.dynatrace.com/news/blog/dynatrace-release-radar-2026/ https://www.dynatrace.com/news/blog/dynatrace-release-radar-2026/#respond Mon, 02 Mar 2026 17:29:50 +0000 https://www.dynatrace.com/news/?p=73224 Release Radar

This series covers recent Dynatrace releases and updates, focusing on what’s new, what changed, and how it applies to you and your organization. Each post outlines newly available capabilities and points to places where you can explore them directly, helping you understand what’s relevant and what to look at next.

The post Dynatrace Release Radar 01.26 appeared first on Dynatrace news.

]]>
Release Radar

We kicked off the new year with our annual customer event, Dynatrace Perform, and many new announcements. If you weren’t able to join us in person, you can watch all the mainstage keynotes, innovation sessions, and breakouts on demand on the Dynatrace Perform 2026 webpage.

In this blog, we’ll focus on brand-new product enhancements that accelerate service troubleshooting, provide richer cloud context for AWS, and deliver meaningful improvements that reduce friction in daily workflows.

If you want to jump straight to our curated sandbox environment for the capabilities mentioned below, head over to our dedicated playground launchpad.

Dynatrace Intelligence

Our biggest news is that Dynatrace Intelligence is now available. It’s the industry’s first agentic operations system that effectively fuses deterministic insights with agentic action to deliver reliable outcomes with autonomous prevention, remediation, and optimization at scale.

Dynatrace Intelligence Marketecture

Here are the new features and capabilities now available in Dynatrace Intelligence:

  • Dynatrace MCP Server: In addition to the local MCP server that was launched in May 2025, our remote MCP server is now generally available.
  • Dynatrace Assist: The evolution of Davis CoPilot puts Dynatrace Intelligence at your fingertips. Dynatrace Assist pulls context from Grail, maps relationships using Smartscape – our real-time dependency graph – and collaborates autonomously with Dynatrace agents using the tools provided by the Dynatrace MCP server.
  • Agentic ecosystem: Whether you aim to level up collaborative operations with SRE agents, enjoy closed-loop autonomous operations with ITSM agents, or want AI-powered code repair with developer and coding agents, we’ve got you covered. Maximize the value of your tool landscape by leveraging our agentic integrations.
  • Agentic workflows: Turn your workflows into agentic automations leveraging Dynatrace Intelligence. This program is currently available in a Preview program.

Smartscape: Real-time dependency graph

The new Smartscape experience delivers a real-time dependency graph that helps practitioners move from “watching signals” to understanding true entity health and cause-and-effect across fast-changing cloud, Kubernetes, on-premises, and hybrid environments. It adds major new capabilities:

  • An all-new Smartscape app with powerful visual analytics and domain-specific views,
  • Fully native cloud entities with complete metadata (including raw cloud/Kubernetes object JSON),
  • and agentless cloud data ingest for automatic enrichment of dependencies and policy context.

This allows teams to diagnose faster, reduce MTTR, and make better architectural and operational decisions with real production context.

Smartscape dashboard

Smartscape enables exploration across millions of relationships, strengthens incident collaboration via in-context workflows like Visual Resolution Path and “view topology” actions, and improves security and governance by visualizing exposure and attack paths with real blast radius and enriched cloud native semantics like tags, ownership, cost centers, and compliance attributes.

The new Smartscape extends far beyond a standard topology; it offers domain-specific views tailored to the unique requirements of your use cases—whether application performance, cloud infrastructure, or essential business services. The following pre-configured views are now available.

  • Smartscape on Grail: discover all entities and relationships in your environment
  • Infrastructure overview: gain insights into which components are running and how they’re connected
  • Service dependency graph: see how your services are connected
  • Problem graph: understand problem impact and blast radius
  • Kubernetes overview: map your Kubernetes environment, from clusters to components
  • AWS EC2 ecosystem overview: understand your entire EC2 ecosystem and resource relationships

Have a look at our recent Smartscape blog post to learn how these enhanced views help solve real-world challenges.

Cloud Operations for AWS

Dynatrace enhanced Cloud Platform Operations expands AI-powered observability into an operations-first experience for practitioners (cloud ops, SRE, and platform teams) by unifying cloud metrics, logs, and events across AWS, Azure (see Preview program), and Google Cloud (see Preview program) in a single platform, enriched with topology-aware context for faster troubleshooting and safer automation. It introduces:

  • fully managed cloud connections with a guided wizard (no extra infrastructure),
  • expanded ingest that captures more cloud service metrics plus richer cloud events (including hyperscaler-native security alerts),
  • and automatic reuse of existing cloud tags to drive access control, ownership, cost allocation, alert routing, and preventive workflows—so teams can move from fragmented signals to clear, actionable answers at enterprise scale.

Dynatrace Dashboards

This allows users to shift from reactive monitoring to proactive cloud operations built around three outcomes: prevention (predict anomalies and trigger workflows before user impact), remediation (AI-driven RCA plus self-healing automation to cut resolution time), and optimization (continuous cost and performance efficiency via real-time insights and recommendations).

For platform teams, the big win is operational simplicity: the onboarding flow is GitOps-ready and removes the need to maintain ActiveGates for CloudWatch ingest on this path. For practitioners, the win is troubleshooting speed: reimagined exploration, resource-rich metadata, and opinionated insights reduce the time from “something’s wrong” to “here’s why.”

Real User Monitoring experience

The new Real User Monitoring (RUM) experience adds modern frontend signals that match how today’s web and mobile apps behave—for example, soft navigation for Single Page Apps (SPA), user interactions (clicks/taps/scrolls), and background requests—alongside Core Web Vitals and key mobile performance signals (including troubleshooting enhancements like application not responding and symbolication). Out of the box, teams get task-focused workflows and dashboards that connect frontend symptoms to backend reality, so you can pinpoint what’s slow or broken and shorten the path from user complaint to verified cause and fix.

Dynatrace Real User Monitoring (RUM) experience

Achieve faster validation of real user impact and clearer prioritization: Users & Sessions grounds investigations in actual sessions, Error Inspector groups and prioritizes errors with the right context, and Experience Vitals helps identify which requests/assets drive slowdowns using redesigned analysis views—so teams can reduce friction, resolve complaints with confidence, and connect experience trends to business outcomes via custom dashboards, notebooks, and DQL exploration, with built-in privacy/permission controls, and optional extended retention for deeper historical analysis (currently available in a Preview program).

AI observability

Dynatrace has expanded agentic AI observability with a broader framework and protocol support, so teams can build, run, and debug autonomous agent systems with confidence across AWS, Azure, and Google Cloud. Support now includes popular agentic ecosystems such as Amazon Bedrock AgentCore, Amazon Bedrock Strands, LangChain Agents, Google Agent Development Kit (ADK), OpenAI Agents SDK, and Model Context Protocol (MCP)—with signals unified via OpenTelemetry and OpenLLMetry into a single correlated observability model for end-to-end visibility across agents, tools, models, and dependencies.

Agent topology visualizes agent execution flows, showing how they interact with one another.
Video: Agent topology visualizes agent execution flows, showing how they interact with one another.

Alongside this expanded support, the new AI Observability app delivers a purpose-built experience to observe AI workloads end-to-end—from agents and LLMs to orchestration layers and tools—so practitioners can validate changes faster, reduce risk, and ship AI features at scale. Key capabilities include end-to-end monitoring of agent interactions and tool usage, prompt/tool/model tracing and debugging across multi-step flows, cost visibility (token consumption, cost trends, caching impact), actionable dashboards and drill-downs (including faster validation via A/B testing across model/prompt variants), and enterprise-grade security, privacy, and governance views such as surfaced guardrail outcomes for auditability and trend monitoring.

Investigations: Transform how practitioners derive actionable insights

The Investigations app provides a central starting point for exploring analytical insights across Grail data. It gives practitioners immediate access to essential investigation capabilities—such as analyzing large DQL results, pivoting queries based on metadata, reviewing investigation history, and connecting logs, metrics, events, and traces—helping practitioners quickly uncover root causes and accelerate complex investigations.

Dynatrace investigations

Improved Dashboards experience

We’ve enhanced several ready-made dashboards that improve your dashboard experience and make insights clearer, faster, and more consistent. You can duplicate and adapt them to kick-start your own dashboards.

  • The Getting started dashboard demonstrates the major types of visualizations you can use and provides example tiles and layouts.
    Dynatrace Dashboards
  • The Page performance & errors dashboard serves as a starting point for investigating page performance and web front-end navigation. It surfaces the most important web performance and reliability KPIs at a glance, highlighting key metrics such as page load time, error count, navigations, LCP, INP, and CLS.
    Dynatrace Dashboards
  • The XHR & fetch performance dashboard includes core KPIs such as request duration, time to first byte (TTFB), and fetch failure rate. These help you quickly spot slow or failing back-end calls that affect the user experience.
    Dynatrace Dashboards

Where to start this week

We encourage you to take advantage of all the efficiencies and insights these new Dynatrace capabilities provide. Depending on your role, here are the recommended next steps for SREs, Cloud Owners, and Development teams seeking faster service troubleshooting loops, richer AWS cloud context, and other meaningful improvements that reduce friction in their daily workflows.

Get started: SREs

  1. Start with Dynatrace Intelligence for faster incident loops
    1. Open Dynatrace Assist during an active issue to pull context from Grail and map relationships via Smartscape, then let it collaborate with Dynatrace agents/tools (via MCP) to accelerate triage and next steps.
    2. If you use chat/agent tooling internally, connect via the Dynatrace MCP Server (remote if you want centralized access) to make Dynatrace context available in your agentic workflows.
  1. Make Smartscape your default “blast-radius + causality” view
    1. Use the new Smartscape app and Visual Resolution Path/view topology actions to validate true upstream/downstream impact and shorten MTTR.
    2. Leverage native cloud/Kubernetes metadata (including raw object JSON) to quickly confirm “what changed” vs. “what broke.”
  1. Automate closure with agentic workflows (Preview program)
    1. Convert recurring remediation steps into agentic automations by combining Dynatrace Intelligence with Workflows for closed-loop operations (start with a high-confidence, low-risk runbook).

Get started: Cloud owners

  1. Onboard AWS with enhanced Cloud Operations first
    1. Use the fully managed cloud connection and guided wizard to bring in unified metrics, logs, and events with richer AWS context, without maintaining ActiveGates for CloudWatch ingest on this path.
    2. Ensure your cloud tags are clean and meaningful, because they’ll automatically drive ownership, access control, cost allocation, and alert routing.
  1. Operationalize outcomes: prevention, remediation, and optimization
    1. Set up alerting and dashboards around the three outcomes:
    2. Prevention: anomaly prediction + proactive workflows
    3. Remediation: AI-driven RCA + self-healing actions
    4. Optimization: continuous cost/performance efficiency using context-rich insights
    5. Use Smartscape to validate dependencies and impacts across accounts, regions, clusters, and services.
  1. Plan for multicloud setup
    1. If you’re also on Azure or GCP, use what you learn on AWS to establish a standard operating model, then extend to preview programs when ready.

Get started: Development teams

  1. Start from user impact with the new RUM experience
    1. Use Users and Sessions to reproduce issues from real sessions, then jump to Error Inspector and Experience Vitals to identify which requests/assets/interactions drive pain.
    2. For SPAs and modern apps, validate soft navigation, user interactions, and background requests alongside Core Web Vitals to quickly pinpoint frontend bottlenecks.
  1. Connect frontend symptoms to backend issues
    1. From a slow, erroring session, follow the workflow to backend services and dependencies (Smartscape helps confirm causality), shortening the path from complaints to verified root causes.
  1. If you ship AI features, instrument them with AI Observability
    1. Adopt the AI Observability app for end-to-end tracing across agents, tools, and models; use cost visibility and A/B validation to safely iterate on prompts/models.
    2. Standardize telemetry via OpenTelemetry and OpenLLMetry, and if you use agent frameworks (LangChain Agents, OpenAI Agents SDK, Google ADK, Bedrock, or MCP), start by observing one representative production flow before scaling coverage.

The post Dynatrace Release Radar 01.26 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-release-radar-2026/feed/ 0
Dynatrace Intelligence at the core of autonomous operations https://www.dynatrace.com/news/blog/dynatrace-intelligence-at-the-core-of-autonomous-operations/ https://www.dynatrace.com/news/blog/dynatrace-intelligence-at-the-core-of-autonomous-operations/#respond Wed, 28 Jan 2026 16:47:00 +0000 https://www.dynatrace.com/news/?p=72710 Dynatrace Intelligence

Executives are looking for successful ways to run their digital ecosystems with AI as cloud and AI adoption reach unprecedented complexity. Organizations are increasingly recognizing that agentic AI on its own can’t deliver the consistent, trustworthy outcomes they expect. With 65% of enterprises investing in AI‑driven monitoring and automation, leaders now need trustworthy AI‑powered observability […]

The post Dynatrace Intelligence at the core of autonomous operations appeared first on Dynatrace news.

]]>
Dynatrace Intelligence

Executives are looking for successful ways to run their digital ecosystems with AI as cloud and AI adoption reach unprecedented complexity. Organizations are increasingly recognizing that agentic AI on its own can’t deliver the consistent, trustworthy outcomes they expect. With 65% of enterprises investing in AI‑driven monitoring and automation, leaders now need trustworthy AI‑powered observability to shift from human‑driven operations to human‑supervised, autonomous digital ecosystems.

Key executive insights

  • Alongside the rapid adoption of agentic AI, Dynatrace is uniquely architected for powering real‑time autonomous operations across organizations’ digital systems while also integrating seamlessly into broader agentic ecosystems.
  • Dynatrace – pioneer of large-scale AI-powered root cause analysis – established predictive operations and now further redefines observability, taking the next step from automation to autonomous action by auto-remediating, auto-preventing and auto-optimizing.
  • Dynatrace takes the guesswork out of AI by optimizing the balance between deterministic AI, contextual analytics, and stochastic AI to drive precise agentic answers and reliable actions.
  • Dynatrace Intelligence is an agentic operations system in the Dynatrace platform, driving autonomous actions through orchestrating ready-made Dynatrace agents as well as external ecosystem agents.
  • AI engineering, AI operations, agentic SRE get enabled by Dynatrace with real-time production feedback-loops from trusted AI.

Minimizing hallucinations and avoiding large language model data processing limits

One of the biggest fears of executives who build agentic frameworks is that generative AI can hallucinate and push their agents off course. CTOs also prioritize ensuring agentic systems have instant access to high‑quality information, so multi‑step agent workflows can execute quickly and reliably.

Hallucinations aren’t minor errors – they can trigger wrong actions leading to outages, security risks, and financial exposure. To make it worse, in agentic processing inaccuracies can accumulate and amplify.

The Dynatrace AI approach reduces the risk of hallucinations by maximizing the use of deterministic AI in its agents, allowing Dynatrace to deliver responses based on real-time data rather than probabilistic guesses. Also, Dynatrace continuously observes, maintains, and enriches this context through real-time dependency graphs, real-world context, and high-performance data lakehouse analytics—fueling precise analytics and real-time answers across the entire digital services and business landscape.

This level of contextual precision is critical because large language models cannot directly process petabytes of heterogeneous observability data. Their context windows are limited, and performance often degrades as input approaches the maximum length.

Dumping large amounts of data into AI requests can reduce quality. Models have finite context windows and can underweight important details in very long prompts. Curating and structuring only the most relevant information generally yield better results than providing everything at once.

To overcome these constraints, it’s essential to rapidly distill vast amounts of data into short, high-quality context—this is where contextual analytics, dependency graphs, and the AI-optimized data lakehouse Grail become critical differentiators.

Real-time analytics with instant visualization of dependencies and interactions

Agentic AI depends not only on the model it runs on but on the quality, context, and immediacy of the data it receives—real‑time, fact‑based inputs are essential to keep pace with agentic decision‑making.

Grail, the Dynatrace data lakehouse, provides this foundation by unifying observability, security, and business data at an exabyte scale. Its schema‑on‑read approach removes indexing overhead and enables any‑question, any‑time analytics of Grail data. Grail processes metrics, logs, traces, user behavior, security events, and business signals alongside directed dependency graphs, delivering deep insights instantly through zero‑latency, always-hydrated storage.

Complementing this, Dynatrace Smartscape continuously refines a real-time dependency graph of real-world dependencies across business, teams, digital services, processes, infrastructure, risks, and more. By dynamically uncovering both vertical and horizontal dependencies, Smartscape enables teams to understand how systems, business processes, and organizational structures interact. The latest generation of Smartscape real-time dependency graph can now also be augmented with custom entities, such as business data types, ownership information, and other meta-data and ownership details, so teams immediately know who to notify for fast, fact-based remediation.

Together, Grail and Smartscape provide the technical foundation for Dynatrace Intelligence to let agentic AI act on facts, not guesses, ensuring that AI-powered decisions are accurate, scalable, and actionable, a pre-requisite for reliable autonomous operations.

Dynatrace Intelligence

Dynatrace Intelligence is an agentic operations system at the core of the Dynatrace platform that fuses deterministic AI with agentic AI to drive a new level of reliability across observability and autonomous operations. It provides a unified intelligence layer where humans define the goals, and AI executes them with precision—guided by policies, context, and guardrails.

At the foundation are deterministic agents that create a highly reliable operational core. The Root Cause Agent, powered by deterministic causal AI, delivers answers with far greater speed and precision than LLM‑only approaches. The Analytics Agent distills petabytes of Grail data lakehouse data into concise, contextual intelligence, while the Forecasting Agent scales predictive capabilities across the environment. An Operator Agent oversees, orchestrates, and coordinates agentic team efforts to ensure optimal outcomes. These foundational agents power every other agent operating on the platform.

Building on this foundation are ready‑made, domain‑specific agents designed to extend and augment the work of Development, SRE, and Security teams. This fusion of deterministic AI, agentic AI, and specialized domain agents enables Dynatrace Intelligence to detect anomalies, predict issues, identify root causes, run complementary investigations, and plan and execute corrective actions—ultimately enabling auto‑prevention, auto‑remediation, and auto‑optimization.

Dynatrace already observes customers’ digital services, end‑user experiences, and AI stacks automatically, making these domain agents immediately impactful. As examples, for developers, Dynatrace detects rising mobile‑app crashes, analyzes the code paths, and produces an immediate fix suggestion—turning what used to take hours into seconds. For security teams, Dynatrace continuously monitors emerging threats and instantly checks the environment for related vulnerabilities or indicators of compromise, helping teams respond proactively before attackers can act.

New Assist Agents simplify the adoption and everyday use of Dynatrace, while Agentic Workflows empower customers to build their own agents.

Dynatrace Intelligence is here not only to remediate symptoms—like a production infrastructure overload—but also to detect the underlying root cause and generate actionable plans to fix the source of the issue.

Dynatrace Intelligence Marketecture

AI-powered observability from Dynatrace enables the next generation software delivery life cycle process with AI engineering, AI operations, agentic SRE, and AI business analytics, through real-time facts from production systems.

Already today, Dynatrace Intelligence agents collectively power a wide range of real‑world use cases, from mobile app crash inspection and front‑end error explanation to infrastructure optimization, Kubernetes operations, and security‑context insights. They also accelerate tasks like anomaly and log‑pattern analysis, vulnerability validation, timeseries analysis, and even dashboard creation, with extensible workflows that let teams expand these capabilities as their needs grow.

Dynatrace Intelligence also coordinates a bi-directional interaction with the agentic ecosystem like AWS Kiro, GitHub Copilot, ServiceNow, Azure SRE agent, Atlassian Rovo and many others – e.g. submitting tickets, invoking a coding agent, assessing the risk of a new deployment, adjusting infrastructure, informing business workflows. It also responds when other systems request, for example, incident details, performance insights, business impact analysis, or resilience risk assessments.

Along with Dynatrace Intelligence, Dynatrace leads customers on a journey toward fully autonomous operations.

Journey to fully autonomous operations

Each organization goes through their own maturity stages and pace on the journey to autonomous operation; for the majority it can look like this:

Automated – a stage when an organization’s digital system executes pre‑defined workflows and actions automatically (based on AI‑generated answers) to support both reactive and predictive operations. I can say that currently, many organizations are striving to get to this stage or are already in it.

Digital systems must be automatable and observable. To validate this, workflows must be testable:

  • Can you automate it?
  • Can you observe it?
  • Can you understand its behavior in real time?

Organizations don’t need to automate every single workflow. Once it’s possible to confirm that key use cases are both automatable and observable, organizations can progress to the next maturity stage.

Supervised Autonomous – in this stage, AI generates execution‑ready action plans with clear reasoning and acts only after human oversight and approval. In a “crawl‑walk‑run” approach, organizations start with small, repetitive tasks that require agentic AI instead of hard-coded workflows.

Key principles in evaluating ability for supervised autonomous operations:

  • Reliability: For reliable decisions and actions, rely as much as possible on deterministic AI and analytics, and leverage generative AI for common sense and learned expertise.
  • Transparency: Allow people to set goals and guardrails for AI, validate reasoning, review knowledge graphs, improve real-time feedback loops, and assess explanations to build trust.
  • Feedback loop: Agentic AI relies on accurate factual inputs to create actionable plans, as well as precise real-time feedback to investigate plan details, refine them and validate execution.

Once the digital system consistently performs with human‑like review discipline, the organization is ready to move toward fully autonomous operations.

Fully Autonomous – in the future, fully autonomous stage, Dynatrace Intelligence acts independently to fulfill business goals and autonomously operates a wide range of aspects in successfully delivering software that end-users expect, requesting human input only when necessary. As much as Dynatrace uses AI to observe other AI within the cloud- and AI-native services customers run, it also continuously observes itself—to self-optimize, ensure compliance, and provide insights that help people set the goals. People still play a crucial role: they review outcomes, adjust goals, and refine instructions. As a result, organizations deliver and operate software with higher resilience, happier customers, and lower cost.

The fusion of deterministic AI and agentic AI sets Dynatrace apart by providing a reliable agentic AI-powered observability. It is an AI that observes other AI and helps organizations to build more resilient applications and better customer experiences.

The post Dynatrace Intelligence at the core of autonomous operations appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-intelligence-at-the-core-of-autonomous-operations/feed/ 0
Autonomous operations hits an inflection point: New agentic AI report reveals what’s fueling scale (and blocking it) https://www.dynatrace.com/news/blog/agentic-ai-report-reliable-autonomous-operations/ https://www.dynatrace.com/news/blog/agentic-ai-report-reliable-autonomous-operations/#respond Thu, 22 Jan 2026 13:48:03 +0000 https://www.dynatrace.com/news/?p=72552 Agentic AI report reveals the need for reliable autonomous operations

A new study of 919 leaders shows how organizations are adopting agentic AI and where they’re facing challenges on the path to autonomous operations. As enterprises scale from pilots to production, they need strong guardrails and real‑time observability.

The post Autonomous operations hits an inflection point: New agentic AI report reveals what’s fueling scale (and blocking it) appeared first on Dynatrace news.

]]>
Agentic AI report reveals the need for reliable autonomous operations

Agentic AI is accelerating into the enterprise faster than many leaders expected, bringing with it unprecedented complexity. Unlike traditional machine‑learning systems, agentic architectures combine goal‑directed reasoning, multi‑step autonomy, and real‑time adaptation across a wide variety of applications. This variability creates exponential interaction paths and the potential for unpredictable behaviors and downstream consequences that traditional monitoring simply can’t capture.

As organizations move from pilots toward autonomous operations, a clear trend is emerging: without guardrails, strategic human oversight, and a real‑time observability control plane, agentic systems face barriers to operating reliably at scale.

The Pulse of Agentic AI 2026 study—based on 919 global leaders responsible for agentic AI development and implementation—reveals how enterprises are adopting agentic AI, where they’re encountering barriers, and why observability is becoming foundational for building safe and reliable autonomous systems.

Agentic AI is rapidly expanding beyond ITOps

Although agentic AI is most established in IT operations, system monitoring, DevOps, cybersecurity, and software engineering, it’s expanding quickly into nearly every domain.

72% use AI agents for IT operations and DevOps 74% expect agentic AI budgets to increase in the next year

Key data points show:

  • 72% use agentic AI in ITOps and DevOps, followed by software engineering (56%) and customer support (51%).
  • Externally exposed use cases—product personalization, sales engagement, digital services—are the fastest‑growing over the next five years.
  • 74% expect budget increases in the next 12 months, often by an additional $2–5M or more.

Agentic systems gain traction first in domains where quick response is imperative, such as those that demand reliability and controlled automation. Observability and deterministic guardrails must therefore be foundational, not optional.

Even as customer‑facing use cases rise, organizations prioritize agentic AI in measurable, repeatable workflows with strong ROI, such as ITOps, data processing, reporting, and cybersecurity. Value and risk scale together, and the only way to manage both is through real‑time, end‑to‑end visibility into agent behavior.

Autonomous operations are growing—but hitting barriers

Organizations are no longer just experimenting. Portfolios are expanding quickly:

  • 72% have 2–10 projects; 26% have 11–21+.
  • 44% have agentic AI in production for select departments.
  • 23% have enterprise‑wide integration in some areas.
44% have projects in broad adoption in select departments 23% have projects in mature, enterprise-wide integration

Yet progress is uneven. The bottleneck is establishing trust in production‑level autonomy.

Top blockers include:

  • Security, privacy, and compliance concerns (52%)
  • Technical challenges in managing and monitoring agents at scale (51%)
  • Difficulty defining when agents act autonomously vs. require human approval (45%)
  • Limited real‑time visibility to trace and troubleshoot behavior (42%)

Organizations aren’t struggling with ideas—they’re struggling with control. Without deterministic guardrails, transparent model behavior, and real‑time signals showing what agents are doing and why, teams can’t safely operationalize autonomy.

Building trust requires incremental progression: human‑in‑the‑loop models, supervised autonomy, and phased functional expansion, all enabled by observability.

Trust and human oversight are intentional—and enduring

Despite enthusiasm for fully autonomous agents, human oversight remains central:

69% of agentic AI decisions are currently verified by a human
  • 69% of agentic AI decisions are verified by a human.
  • Top validation methods include data‑quality checks, human review, drift detection, and logs/traces.
  • Only 13% rely exclusively on fully autonomous agents, but 64% combine supervised and autonomous models.

Organizations are building human-AI partnerships, not replacements. In fact, in the long term, respondents expect a 60/40 human‑in‑the‑loop balance for business applications and 50/50 for IT and customer‑support functions.

Two insights stand out:

  1. Because agentic AI is probabilistic, enterprises depend on human judgment for high‑risk validation.
  2. Observability supplies the factual ground truth that makes this oversight effective.

As organizations scale, human involvement becomes more strategic—guiding goals and accountability while AI handles repeatable or time‑sensitive execution.

Reliability and resilience define success

To measure agentic AI success, organizations prioritize real‑time decision‑making, performance, efficiency, and reliability, for example:

60% use technical performance as their #1 agentic AI success measurement 44% use manual methods to review communication flows among AI agents
  • Technical performance is the top metric (60%)
  • Developer and operational efficiency follow
  • Customer satisfaction and business outcomes come next
  • Compliance and security are rising, especially in large enterprises

Still, 44% manually review inter‑agent communication flows—a clear scaling limitation.

Agentic systems are inherently interconnected. A performance regression or hallucination in one agent can cascade downstream into applications, user experiences, or security posture. As a result, resilience and rapid recovery—not just efficiency—must be built into agentic systems.

Doing so requires:

  • Observability signals that detect anomalous or unexpected actions
  • Real‑time tracing of inter‑agent communication
  • Automated risk detection informed by factual telemetry
  • Deterministic guardrails preventing stochastic failures from propagating

Reliability and security are no longer separate concerns—they’re inseparable in autonomous systems.

Observability is a control plane for agentic AI

The study’s most strategic finding: observability is shifting from a supporting function to the control plane for agentic AI.

Usage is already broad:

69% use observability in the implementation phase of the agentic AI lifecycle 57% use observability in operationalization 54% use observability in operationalization
  • 69% use observability during implementation
  • 57% in operationalization
  • 54% during development

But gaps remain in transparency, real‑time visibility, risk detection, and linking signals to business outcomes.

Because agentic behavior can’t be fully tested in advance, teams need real‑time observability to monitor performance in production and respond quickly to anomalies. Traditional monitoring tools can’t explain why an agent took an action, detect hallucinations in real time, or trace downstream impact.

A modern observability control plane must:

  • Blend deterministic telemetry with probabilistic model insights
  • Standardize semantic conventions and agent‑action signals
  • Link behavior to business outcomes
  • Detect and correct anomalies instantly
  • Keep agents aligned to shared, real‑time facts
  • Maintain clear human accountability and governance

This is the foundation organizations need to progress from supervised autonomy to reliable, production‑grade autonomous operations.

The path to operationalizing agentic AI

Autonomous operations will redefine enterprise technology. But success requires treating autonomy as a maturity journey, not a leap:

  • Start with preventive and recommendation‑driven workflows
  • Build trust through human‑in‑the‑loop models
  • Harden services, signals, and data paths
  • Use observability to detect anomalies and validate actions
  • Scale autonomy gradually—with transparency and governance

The message of the 2026 research is clear: the future is autonomous, but limited visibility is hindering reliability and control. Scaling agentic AI requires an observability‑based control plane that grounds probabilistic agent behavior in deterministic, real‑time facts.

For deeper segmentation, maturity criteria, fuller KPI breakdowns, and several stage-specific observability priorities, download the full report.

The post Autonomous operations hits an inflection point: New agentic AI report reveals what’s fueling scale (and blocking it) appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/agentic-ai-report-reliable-autonomous-operations/feed/ 0
Flexible vendor-agnostic log forwarding with OpenPipeline https://www.dynatrace.com/news/blog/flexible-vendor-agnostic-log-forwarding-with-openpipeline/ https://www.dynatrace.com/news/blog/flexible-vendor-agnostic-log-forwarding-with-openpipeline/#respond Fri, 16 Jan 2026 19:41:14 +0000 https://www.dynatrace.com/news/?p=72501 Flexible vendor-agnostic log forwarding

Log forwarding includes metrics, spans, and events Logs are a core pillar of observability, and in many organizations, logs serve at least a dual purpose. They drive day-to-day troubleshooting, root cause analysis, security investigations, and many other use cases. Logs also need to be available for compliance audits, Business Intelligence (BI) analysis, and traditional Security […]

The post Flexible vendor-agnostic log forwarding with OpenPipeline appeared first on Dynatrace news.

]]>
Flexible vendor-agnostic log forwarding

Log forwarding includes metrics, spans, and events

Logs are a core pillar of observability, and in many organizations, logs serve at least a dual purpose. They drive day-to-day troubleshooting, root cause analysis, security investigations, and many other use cases.

Logs also need to be available for compliance audits, Business Intelligence (BI) analysis, and traditional Security Information Event Management (SIEM) solutions. The challenge is doing this without relying on bolted-on forwarders or rigid, costly integrations.

With the introduction of the new Dynatrace OpenPipeline capability, we’re providing you with the freedom to precisely define your organization’s forwarding, paired with flexible and unified log collection and processing.

OpenPipeline capability
Figure 1. OpenPipeline provides rich configuration and customization options for each forwarding setup.

Update – June 2026

Data egress will be in General Availability starting June 17, 2026.

OpenPipeline now supports forwarding for all data types ingested into Dynatrace, including metrics, spans, and events, giving you a single, unified pipeline for your entire observability dataset. The same principles described below apply across the board: forward before or after processing, export in open formats, and route to the storage or third-party destination of your choice, regardless of signal type.

Read on for the full details of how OpenPipeline forwarding works.

Centralized log collection and processing

OneAgent, Cribl, OpenTelemetry, AWS Firehose, journalD, Logstash, and Fluent Bit … Have you heard of any of these tools? Maybe you’re using all of these at the same time.

With Dynatrace, you have the choice of using one or all of these tools simultaneously. Dynatrace OpenPipeline allows you to process and transform them in the same way, regardless of their source, and to move seamlessly from one shipping method to another. Just as our customer, United Wholesale Mortgages, successfully consolidated their tools while moving logs off Splunk.

Forwarding logs un/processed

Depending on your industry or geography, you might be required to retain logs in an unprocessed state on third-party-managed storage, just as you might have used tape backups in the old days. With Dynatrace OpenPipeline, you can configure a forward-before-process to retain log events in an unaltered state.

OpenPipeline flow diagram from ingest to forward
Figure 2. OpenPipeline flow diagram from ingest to forward

Alternatively, you can first send log events to your defined processing pipeline, extract or convert the logs into metrics, and even drop unnecessary details before determining what to forward-after-processing to your object storage.

The last step in the pipeline is to decide whether to route log events to a Dynatrace Grail® bucket for retention or drop them.

Overcome vendor lock-in with precise, open forwarding

While other log management solutions provide proprietary log archive formats, you can overcome these challenges with Dynatrace and log forwarding for data archiving.

The Dynatrace OpenPipeline log forwarding capability delivers log events in a standardized and compressed NDJSON (Newline Delimited JSON) (.JSON.GZ) archive format, for which you can define the details, such as segmentation and filename prefixes.

Residing on destinations like AWS S3 or Azure Blob storage, other 3rd party solutions such as Microsoft Sentinel, Snowflake, or Tibco Spotfire can easily leverage these archives for use cases not covered by your observability platform.

Cost-effective long-term retention for logs

Dynatrace Grail offers cost-effective log retention for up to 10 years, accessible at any time with the same speed as on day one of ingestion, supporting a daily log ingest volume of 1 PB per tenant.

Many of our customers already retain years’ worth of logs today, always accessible without the headaches of re-ingestion, re-indexing, or storage tiering and scaling concepts.

Log retention configuration for a new bucket, retaining logs for 10 years
Figure 3. Log retention configuration for a new bucket, retaining logs for 10 years

Still, there are use cases where aged logs become low-value, low-touch data.

You may want to keep certain log events available in Grail for the first 60 or 90 days for troubleshooting purposes, and then store a copy in an archive for years.

NDJSON-exported log archives are an ideal file format for further increasing cost-effectiveness and are ready for handover to other teams or parties for investigation, thanks to the standardized JSON.GZ format. These are especially valuable for use cases like:

  • Reviewing access logs during a post-mortem security incident
  • Providing transaction logs to fulfill a court order
  • File access audit-trial review

Tool consolidation with OpenPipeline and log forwarding

Log archiving presents an opportunity and an integration strategy for solutions that don’t yet integrate seamlessly with Dynatrace.

More importantly, OpenPipeline provides you with processing and transformation capabilities, volume control, and a reliable cadence for transmitting your log events.

You no longer have to rely on custom-defined scripts for the splitting and shipping of your logs. And there are no worries about truncating large log events; Dynatrace supports up to 10 MB.

Forget about all the retransmission headaches with custom log shippers when a connection breaks; all this is baked natively into Dynatrace components such as OneAgent, ActiveGate, and others.

Start leveraging OpenPipeline and log forwarding today and retire your third-party forwarding tool stack.

For complete details, go to Log Forwarding in Dynatrace Documentation.

State of Log Management 2026

Download the report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

The post Flexible vendor-agnostic log forwarding with OpenPipeline appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/flexible-vendor-agnostic-log-forwarding-with-openpipeline/feed/ 0
Write the future: Create your own agentic workflows https://www.dynatrace.com/news/blog/write-the-future-create-your-own-agentic-workflows/ https://www.dynatrace.com/news/blog/write-the-future-create-your-own-agentic-workflows/#respond Thu, 08 Jan 2026 08:00:10 +0000 https://www.dynatrace.com/news/?p=72347 Agentic workflows with Davis CoPilot

Imagine commissioning le Carré and Fleming to build your perfect undercover agent: quietly embedded in the system you’re watching. You hand in your mission brief, which includes the target, objective, and behaviors to track. Your agent observes without drawing attention, reporting insights back to you. On cue, the information flow you’ve carefully orchestrated turns signals […]

The post Write the future: Create your own agentic workflows appeared first on Dynatrace news.

]]>
Agentic workflows with Davis CoPilot

Imagine commissioning le Carré and Fleming to build your perfect undercover agent: quietly embedded in the system you’re watching. You hand in your mission brief, which includes the target, objective, and behaviors to track. Your agent observes without drawing attention, reporting insights back to you. On cue, the information flow you’ve carefully orchestrated turns signals into actionable intelligence that helps pre-empt risk.

Dynatrace doesn’t write spy fiction. However, even better, Dynatrace now lets you write your own smart agentic workflows that deliver intelligent reports and react to changes in your environment based on your objectives.

Adding generative AI to your workflow

Using the power of gen AI, Davis CoPilot® transforms your workflows into agentic instructions. Davis CoPilot lets you explore data using conversational language, translating complex data and queries into summaries, and provides intelligent recommendations across Dynatrace.

Integrated into Dynatrace Workflows, Davis CoPilot is your toolkit for building smart automations, bringing the power of generative AI into your mission-critical workflows.

Build conversational automation that adjusts to live data based on your instructions, sending summaries of your crash logs directly to Slack
Figure 1. Build conversational automation that adjusts to live data based on your instructions, sending summaries of your crash logs directly to Slack.

By embedding Davis CoPilot in your workflows, you can associate any automation with any number of conversational automations. Your workflows can even perform actions autonomously when combined with precise Davis® AI forecasting, for instance, scaling resources based on forecast demands.

In real time, these workflows monitor live data, summarize critical issues, and identify remediation paths or emerging threats. When scheduled, these smart workflows help you outsource routine tasks, such as alerting stakeholders of costly queries.

Let’s look at some examples of how these smart workflows can help you in your daily work.

Build agentic workflows that respond to critical events

Proactive guiding through complex problem remediation

Let’s assume you want to build an automation that cuts through alert noise and analyzes a problem as it occurs, guiding you through the remediation. When a new problem is detected, Davis CoPilot extracts the problem details, summarizes the situation, and provides tailored remediation guidance. By embedding it into a smart workflow, you can select your preferred automation to automatically syndicate this information, populate a ticket in ServiceNow or Jira, or post it to a dedicated Slack channel.

See how you can set up a workflow automation that automatically sends summaries and remediation guidance when a new problem is detected.

Monitor emerging threats to help you orchestrate a response

Next, you can build an agentic workflow that helps you monitor emerging threats and assess their risk to your environment as vulnerabilities are detected in your tenant. In plain language, you instruct your agent to extract IOCs, query security events in your environment, and correlate them with observability data in your environment. Information provided by the external threat feed is automatically matched against the live context in your tenant. The agent has now collected all the necessary information and provides a reliable risk assessment, along with a plan to orchestrate a response, directly in your Slack channel, ensuring around-the-clock visibility and a rapid response.

With a single workflow, you can monitor emerging security events as they occur, understand their impact, and determine the next steps.
Figure 2. With a single workflow, you can monitor emerging security events as they occur, understand their impact, and determine the next steps.
Example of a tailored and contextual analysis delivered to Slack as the issue arises
Figure 3. Example of a tailored and contextual analysis delivered to Slack as the issue arises

Write the future: Build an agentic workflow that autonomously auto-scales your resources

Dynatrace helps you build agents that reason autonomously. The key is to deliver data as precise as Dynatrace forecast capabilities. In this example, we linked the power of Davis AI to forecast demand, with generative AI and GitHub automations. Davis AI predicts the number of resources the hyperscaler infrastructure will need based on forecasted demand. When Davis AI notices a scaling need, Davis CoPilot interprets the data and autonomously edits the manifest using the GitHub automation. Giving you one end-to-end workflow that automatically scales resources up or down based on forecasted needs. To see this in action, watch how this workflow autonomously edits a manifest based on Davis AI suggestions to auto-scale a Kubernetes cluster.

Schedule agentic workflows to optimize routine tasks

Do you feel like sleeping in a little later? Maybe stretching your lunch break a little longer? Running that extra hill without sacrificing your productivity? Scheduling Davis CoPilot into your smart workflow is a great way to automate recurring tasks and save time.

Build an automation that predicts resource consumption

A recurring challenge for SREs is analyzing the full environment to predict future bottlenecks or over-resourcing and continuously translating the data to update stakeholders. Even with great observability in place, you need to ensure that you interpret the data and make timely decisions to inform future provisioning.

By combining Davis AI forecasting automation with Davis CoPilot, you can build an agent that answers key questions, such as which workloads are most resource-intensive, which resources show the most variance, and which require frequent scaling. This automation is capable of highly reliable forecasts, even when data points are limited. The automation interprets the data and emails actionable recommendations directly to you and anyone else who needs to stay informed.

To see this in action, watch the section of this video that explores predicting resource consumption.

Smart workflows that optimize query costs

Scheduling tasks can even help you keep costs lean and efficient. For admins or budget owners, staying within financial limits while maintaining performance is a constant challenge. In this example, we built a smart workflow that identifies the top 20 most expensive queries from the last 24 hours. Davis CoPilot analyzes each query and sends optimization recommendations directly to the query authors via email.

Smart workflow leveraging Davis CoPilot to recommend query optimizations tailored to your tenant
Figure 4. Smart workflow leveraging Davis CoPilot to recommend query optimizations tailored to your tenant
Example of an optimization suggestion delivered to the inbox of the query author
Figure 5. Example of an optimization suggestion delivered to the inbox of the query author

To implement this yourself, tailored to the most expensive queries executed on your tenant, go to our documentation

Conclusion: Adapt your workflows to any stage of your automation journey

These are just a few examples; the applications for it are endless. We’ve designed this workflow action to cater to your organization’s automation appetite. You may want to transform how you keep business stakeholders informed about what’s happening in your environment, leveraging Dynatrace’s highly accurate insights, which are translated into plain language and actionable next steps.

Alternatively, you may be ready to transition towards autonomous operations, where automation not only supports but also acts in a controlled and reliable manner. Davis CoPilot embedded into your workflows opens the door to your agentic journey.

Start your agentic journey and join the Davis CoPilot for Workflows Preview

Davis CoPilot for Workflows is available as a Preview. Sign up now and see how generative intelligence embedded into your workflows transforms your automation. Today, it helps you react faster, optimize more effectively, and collaborate seamlessly. Tomorrow, it will go even further: anticipating needs, orchestrating actions, and enabling truly autonomous reasoning.

Gain efficiency and have your agentic workflows do the work for you!

The post Write the future: Create your own agentic workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/write-the-future-create-your-own-agentic-workflows/feed/ 0
Reliable enterprise automation at scale: Accelerate the innovation loop with Dynatrace Workflows https://www.dynatrace.com/news/blog/reliable-enterprise-automation-at-scale-accelerate-the-innovation-loop-with-dynatrace-workflows/ https://www.dynatrace.com/news/blog/reliable-enterprise-automation-at-scale-accelerate-the-innovation-loop-with-dynatrace-workflows/#respond Mon, 05 Jan 2026 15:00:57 +0000 https://www.dynatrace.com/news/?p=72335 logs and traces

Enterprise automation rarely fails because teams lack ideas. It fails because scaling automation safely is challenging, and conventional workflows often remain static after they ship, while systems and conditions continue to evolve. The new update to Dynatrace Workflows helps teams accelerate their innovation loop and run automation like production software. Five improvements make it easier […]

The post Reliable enterprise automation at scale: Accelerate the innovation loop with Dynatrace Workflows appeared first on Dynatrace news.

]]>
logs and traces

Enterprise automation rarely fails because teams lack ideas. It fails because scaling automation safely is challenging, and conventional workflows often remain static after they ship, while systems and conditions continue to evolve.

The new update to Dynatrace Workflows helps teams accelerate their innovation loop and run automation like production software. Five improvements make it easier to iterate and operate at scale: workflow drafts, sub-workflows, approval requests, persistent execution data, and real-time notifications. Together, they support safer change, reusable standards, built-in governance, and data-driven refinement.

The innovation loop in enterprise automation

The innovation loop is a repeatable cycle for quickly and safely improving enterprise automation that treats workflows as living systems that evolve with your environment, self-improving as systems, teams, and requirements change, rather than remaining static:

  • Teams start by designing or updating a workflow and validating it safely before it goes live.
  • Next, they standardize the parts that work into reusable building blocks so future workflows are faster to create and more consistent.
  • Then they add governance where it matters, such as approval checkpoints for high-impact actions, so speed does not bypass policy.
  • After execution, they observe what actually happened using execution records and signals, including errors, latency, and outcomes, to understand reliability and impact.
  • Finally, they refine the workflow based on those findings by tuning logic, tightening controls, improving reuse, and adjusting notifications.

The result is a compounding impact: faster delivery over time, fewer incidents at scale, automation that integrates changing circumstances, and greater organizational trust, as transparency, auditability, and tight feedback loops back improvements.

How Dynatrace Workflows accelerates the innovation loop

Dynatrace Workflows helps you accelerate your innovation loop with powerful new features designed for adaptability, collaboration, transparency, and control. Each feature is designed to empower teams to safely iterate on automation in dynamic enterprise environments.

Faster iteration with safer testing and experimentation with workflow drafts

The new workflow drafts feature allows teams to create, edit, and refine workflows in a draft state, allowing them to validate logic and configuration without triggering live automation. Workflows that remain in draft mode don’t consume workflow hours.

Workflow drafts support the first step of the innovation loop: iterate safely. Teams can make changes, review them with peers, and confirm the workflow is ready before anything runs in production.

Key benefits

  • Reduce risk by testing configuration changes without production impact.
  • Allow faster iteration and cleaner change management.
  • Improve collaboration during the design and review stages.
  • Control costs with included workflow-hour consumption for draft-only workflows.
Create drafts to refine and safely test your workflows.
Figure 1. Create drafts to refine and safely test your workflows.

Standardize reusable blocks for reliability and speed with sub-workflows

Sub-workflows allow a workflow to be executed as a task inside another workflow, so teams can package repeated logic into reusable sub-workflows and compose larger automations from consistent building blocks.

This supports the “standardize” step of the innovation loop. When teams standardize and reuse proven components, automation becomes more consistent, easier to maintain, and quicker to expand across teams and environments.

Key benefits

  • Enhance reliability and consistency by reusing proven workflow components.
  • Speed up delivery by composing workflows from standard blocks.
  • Simplify maintenance by updating shared logic in a single location.
  • Keep complex automations understandable by breaking them down into smaller, manageable units.

Example use cases

  • Notification workflows: Automatically send notifications based on incident severity or timing.
  • Data processing workflows: Run filter/transform/store stages as distinct sub-workflows, processing large datasets in stages.
  • Incident response workflows: Automate repetitive tasks as reusable steps, such as creating tickets, logging incidents, notifying stakeholders, and triggering remediation steps.
Include a validated component in your workflows.
Figure 2. Include a validated component in your workflows.

Supervised human-in-the-loop control where it matters with approval requests

With Dynatrace Workflows, you can now insert an approval request into a workflow and pause execution until a human reviews and approves the next step.

This feature supports the “governance” step in the innovation loop. Approvals turn your workflows into supervised autonomous operations. This allows the automation of more processes while ensuring strict compliance with enterprise requirements for governance, risk management, and accountability, particularly for high-impact actions.

Key benefits

  • Add explicit human control at high-risk decision points.
  • Reduce errors in sensitive operations (changes, remediations, escalations).
  • Ensure compliance with organizational standards.
  • Increase trust in automation across security, ops, and platform teams.
  • Minimize operational risks.
Add human control to any of your workflow steps.
Figure 3. Add human control to any of your workflow steps.

End-to-end visibility with an audit trail and persistent execution data

Dynatrace Workflows now persists workflow execution data in DQL. All workflow activities are recorded as system events in the dt.system.events table, creating a queryable execution record for tasks, states, and outcomes as a comprehensive audit trail.

Persisting execution data supports the “observe” step of the innovation loop. When workflow execution is captured as data, teams can troubleshoot failures, provide transparency, prove what happened, and improve automation based on evidence rather than anecdotes.

Key benefits

  • Provide an audit trail for enterprise traceability and review.
  • Monitor workflow health and reliability through execution outcomes.
  • Speed up troubleshooting by pinpointing where failures occur.
  • Optimize automation using DQL-driven actionable insights (error rates, trends, hotspots, and more)

Example: Workflow health overview

This example demonstrates how to use execution data to support users in identifying those workflows that need improvement.

The following DQL query…

fetch dt.system.events
| filter event.kind == "WORKFLOW_EVENT" and event.provider == "AUTOMATION_ENGINE" and event.type =="TASK_EXECUTION"
| filter dt.automation_engine.state == "ERROR"
| summarize count = count() , by:{dt.automation_engine.workflow.title}

…counts task execution errors per workflow. Using a honeycomb visualization, it displays error distribution by workflow and provides visibility into those workflows that fail too often.

Audit and observe your workflows with persistent execution data.
Figure 4. Audit and observe your workflows with persistent execution data.

Instant alerts for failures and workflow changes with real-time notifications

Easily set up and customize real-time notifications by email for all workflow errors or changes made by others, ensuring you’re always informed.

These real-time notifications support the “respond and refine” step of the innovation loop. Proactive alerting allows you to manage workflows effectively, minimize disruptions, and maintain seamless operations.

Key benefits

  • Get immediate visibility into failures to address potential issues before they escalate.
  • Detect workflow changes early to reduce drift and surprises.
  • Improve operational coordination across teams that share workflows
  • Ensure uninterrupted operations with timely, actionable alerts.
See notifications in real time and get immediate visibility into your workflow changes.
Figure 5. See notifications in real time and gain immediate visibility into your workflow changes.

Empower your team: Build smarter, safer, and scalable workflows today

Scaling automation today is hard, changes are risky to test, automation becomes outdated and buggy as soon as one part of the organization makes changes, and when something inevitably breaks, you’re left hunting for answers across scattered logs and screenshots. But it doesn’t need to be like this.

Are you ready to make things easier, run your workflows like production software, and accelerate your innovation loop? Here’s how you get started:

Get started building smarter, safer, dynamic, and scalable workflows today!

The post Reliable enterprise automation at scale: Accelerate the innovation loop with Dynatrace Workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/reliable-enterprise-automation-at-scale-accelerate-the-innovation-loop-with-dynatrace-workflows/feed/ 0
Six observability predictions for 2026 https://www.dynatrace.com/news/blog/six-observability-predictions-for-2026/ https://www.dynatrace.com/news/blog/six-observability-predictions-for-2026/#respond Wed, 17 Dec 2025 13:55:29 +0000 https://www.dynatrace.com/news/?p=72228 Dynatrace predictions 2026

Digital systems will continue to grow in scale and complexity in 2026, driven by the rapid adoption of agentic AI, unified telemetry, and cloud-native delivery models. These shifts will influence how organizations understand system behavior, prepare for autonomy, and maintain reliability in environments that change in real time. The insights that follow highlight the trends […]

The post Six observability predictions for 2026 appeared first on Dynatrace news.

]]>
Dynatrace predictions 2026

dynatrace observability predictions 2026

Digital systems will continue to grow in scale and complexity in 2026, driven by the rapid adoption of agentic AI, unified telemetry, and cloud-native delivery models. These shifts will influence how organizations understand system behavior, prepare for autonomy, and maintain reliability in environments that change in real time. The insights that follow highlight the trends executives should monitor most closely, along with the conditions that will determine whether AI-driven operations deliver reliable, transparent, and resilient outcomes.

Key insights for executives

  • Complexity will surge with agentic AI: Digital ecosystems are already complex, but agentic systems are introducing an exponential leap. Each new agent brings its own logic, behavior, and interactions, often acting independently, sometimes unpredictably. Without visibility into how agents interact or what decisions they make, organizations risk losing control over their systems. Guardrails, oversight, and end-to-end observability will be essential to avoid chaos and maintain reliability as the complexity of this new AI layer accelerates system behavior and reshapes the digital environment.
  • Autonomous operations will depend on maturity, not ambition: Organizations will not move directly to full autonomy. They will progress through preventive operations and recommendation-driven workflows before adopting supervised autonomy, the final step before full autonomy. AI-assisted automation is where the foundation is built, because this stage forces organizations to expose and harden the services, data sources, and contextual signals that AI depends on. Autonomy is only possible when these components are accessible in real time, performant, and understood in context. With this foundation in place, supervised autonomy will reliably prepare the environment for full autonomous operations.
  • Resilience will become a primary measure of digital operations: Customers expect systems to remain available and secure even under stress, and leaders will treat reliability and security as a single requirement. Early detection and rapid recovery will be essential, because failures spread faster across these interconnected systems. As a result, organizations will need unified visibility to protect customer experience and revenue.
  • Reliable AI requires strong deterministic foundations: AI can only act dependably when its inputs are accurate, contextual, right-sized, and correctly interpreted. High-quality information must be available in real time and understood in a context-aware view of the broader system. Because large language models can’t reason over raw telemetry at scale, enterprises need mechanisms that distill massive data streams into concise, meaningful context and graph-based representations that show how systems and signals relate. Leaders should prioritize data quality, contextual integrity, and correct interpretation to ensure AI decisions remain reliable and useful.
  • Human supervision will remain essential in AI-enabled operations: AI will take on more execution, but humans will continue to set goals, define boundaries, and ensure accountability. Leaders should redesign roles so that human judgment guides the system while AI handles repeatable or time-sensitive tasks.
  • AI will become a standard component of newly developed digital services: AI workloads, pipelines, and operational practices will merge with existing cloud development processes, and executives should prepare for closer alignment among AI engineering, platform, SRE, and security teams to support consistent reliability and performance.

Prediction 1: Agentic AI triggers a new era of system complexity

a large curor in a field of connected dots representing Dynatrace observability prediction 1

Agentic AI is introducing a new level of system interaction. It’s more powerful, but exponentially harder to manage. As agents coordinate tasks, exchange context, and trigger downstream actions, even well-architected digital environments can spiral into unpredictable behavior. Most organizations are not ready for this shift. Without strong observability and consistent governance, these systems will become increasingly difficult to understand and control.

Think of each AI agent acting autonomously based on instructions and input from not only humans but plenty of first-and third-party agents. A single customer interaction might set off hundreds of background conversations among agents, each taking its own initiative. Roles shift depending on the situation, and some agents may direct others.

Common scenarios show how this plays out. When a vehicle detects an issue, task-specialized agents may check customer information, evaluate service options, estimate timelines, and coordinate a resolution. A travel assistance agent might do something similar, reaching out to agents that compare flights, check loyalty benefits, book transportation, and adjust plans in real time. In both cases, many agents work behind the scenes toward a single outcome, and the interactions between them can multiply in unpredictable ways. Every agent still reports to a human or another agent, and accountability remains with human supervision. This exponential growth in agent-to-agent communication can’t be managed without observability.

Organizations that adopt agentic AI without unified context and clear guardrails will face escalating costs, unpredictable behavior, and higher risk. The challenge is not just improving individual models, but managing the web of autonomous interactions that unfold in real time. In this next phase, observability is no longer a support function: It becomes the foundation for safe, scalable, and governable agentic ecosystems.

Prediction 2: The path to autonomous operations requires several maturity steps

a consecutive series of green boxes leading to a larger green box

Enterprises will take meaningful steps toward autonomous operations. Maturity, not ambition, will determine who succeeds. AI cannot act independently until the underlying systems, automation, and processes are stable, observable, and well-understood. Agentic systems are coming, but first, the groundwork must be solid. Earlier stages of automation are essential, because they surface the gaps in data access, service performance, and contextual signals that AI depend on. Only after those components are reliable and available in real time will supervised and autonomous operations take hold.

Most enterprises will follow a progression: they will start by ensuring their digital systems are fully automated, with runbooks, APIs, and interfaces in place to support reliable execution. This foundation enables predictive operations, where issues can be identified and remediated before they affect end users. From there, organizations can introduce supervised autonomous operations, using agentic automation with human oversight to build confidence and operational trust. As maturity increases and these systems consistently perform as expected, enterprises can progress naturally toward fully autonomous operations.

The journey toward fully autonomous operations will be gradual. Organizations that invest now in preventive workflows and recommendation-driven automation will be best positioned to introduce autonomous capabilities safely and responsibly.

Prediction 3: Resilience becomes the new benchmark for operational excellence

A red sextagon containing an alert icon representing Dynatrace observability prediction 3

Resilience will become the defining measure of digital performance. As systems become more distributed and interconnected, small faults can spread quickly across applications, cloud regions, payment systems, and third-party services. Leaders won’t treat reliability, availability, security, and observability as separate practices. They will view them as a single requirement: the ability of a system to absorb disruption, recover quickly, and maintain a consistent customer experience under stress.

Independent research we commissioned with FreedomPay shows why this shift is accelerating. The findings reveal how fragile digital ecosystems have become and how quickly technical failures turn into customer disruption and financial loss. In the United Kingdom, payment outages put an estimated £1.6 billion in annual revenue at risk. In France, the figure rises to €1.9 billion. A single service issue can ripple across connected systems and channels, showing how tightly coupled modern operations have become.

Customers feel these failures immediately. Patience begins to drop within the first few minutes, and many customers leave the transaction if the issue persists for more than fifteen minutes. Yet the average outage lasts more than an hour, which means most of the damage has already occurred. Nearly one in three customers say a single incident is enough to reduce their trust in a business, with younger digital native consumers even more likely to leave.

This environment requires a unified approach to resilience. Organizations need shared visibility into how services behave, how failures propagate, and how recovery affects the customer journey. Resilience will be measured by how systems respond under stress, not just how they perform when digital services run as expected.

Prediction 4: Reliability becomes the foundation of AI progress

a series of dots containing robot icons appear along a time continuum swoosh representing Dynatrace observability

Organizations will prioritize building foundations that make AI systems consistently reliable. The next phase of AI progress will depend as much on deterministic grounding and factual signals as on the generative power of stochastic models. Enterprises are recognizing that creativity alone is insufficient. Reliable AI requires both structured inputs and mechanisms that ensure outputs remain trustworthy.

Agentic systems add a new layer of complexity. As agents coordinate tasks, exchange context, and initiate downstream actions, even a small misunderstanding can propagate across the system. Greater capability amplifies this effect because a powerful agent can accelerate outcomes while also accelerating an error. This is how hallucination emerges at system scale—not from a single faulty model, but from inaccuracies that compound across agent interactions. Deterministic grounding and end-to-end observability prevent that inaccuracy by ensuring agents act on the same factual signals and remain accountable to the human operator.

A common scenario shows what this looks like. A vehicle detecting a problem may trigger agents that review customer data, vehicle status information, identify service locations, evaluate schedules, estimate travel time, and plan the full resolution workflow. In each case, many agents collaborate behind the scenes to produce a single outcome. Organizations that want transparent and dependable AI outcomes will prioritize deterministic guardrails, enabling agentic systems to behave safely, act predictably, and collaborate with clarity.

Prediction 5: AI will scale, but human supervision will remain essential

An outline of a person inside a circle at the center surrounded by robot icons representing Dynatrace observability prediction 5

In the next year, agentic AI growth will lead to a new operating model where humans define goals, and AI performs well-defined execution. As systems gain more context and become capable of coordinated action, the human role will shift from performing tasks to setting direction, providing instructions, and ensuring oversight. Organizations will rely on AI to analyze relationships, identify risks, and initiate safe actions, while humans remain accountable for outcomes and cross-domain judgment.

Agentic AI will behave much like a high-speed intern. When given clear goals, good tools, and instructions, and the right context, it will deliver results at a speed that is difficult for teams to match manually. But it will still require guidance. Humans will define the aim, interpret trade-offs, and make decisions where intent is unclear or results are ambiguous. If something goes wrong, accountability stays with the human operator, not the system.

This operating model will help teams manage complexity more predictably. AI will take on repetitive or time-sensitive tasks, and humans will focus on strategic decisions and system-level understanding. Growth in the agentic era will come from organizations that combine human judgment with AI-driven execution in a way that is transparent, governed, and aligned to business objectives.

Prediction 6: AI and cloud teams will converge

A series of robot icons interspersed and interconnected with cloud icons superimposed over an infinity loop representing DevOps

AI will stop operating as an isolated discipline and will become a normal component of cloud-native software delivery. Teams will integrate AI into digital services the same way they integrate databases or other core systems. As a result, AI engineering, cloud engineering, SRE, and security will converge into a shared operating model with common pipelines, shared SLOs, and unified accountability for the full lifecycle of AI-enabled services.

This shift reflects how modern software already behaves. AI features influence cost, latency, behavior, and compliance, and these effects span the entire stack. They can’t be monitored or governed in isolation. To operate reliably in production, AI must run within the same workflows, guardrails, and delivery pipelines used for the rest of the cloud-native system.

End-to-end observability becomes essential because what matters is the complete outcome for the user. The guidance agents receive, the actions they take, the database calls they trigger, and the costs they incur all contribute to the overall user experience. Observability must follow all of these signals together and treat AI components, application logic, and cloud infrastructure as one interconnected system. This removes the distinction between “AI observability” and traditional telemetry and creates a unified view that aligns to how customers experience the service.

Organizations that adopt this model will treat AI as a first-class software component. Central teams will define use cases, establish common stacks, and ensure compliance, while product teams will build AI directly into their delivery pipelines. This practical convergence will allow enterprises to operate AI-driven services with the same discipline and predictability as any other cloud-native system.

The post Six observability predictions for 2026 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/six-observability-predictions-for-2026/feed/ 0
Maximizing cloud efficiency: Driving cost optimization and sustainability with Dynatrace https://www.dynatrace.com/news/blog/optimize-cloud-cost-cloud-sustainability-with-dynatrace/ https://www.dynatrace.com/news/blog/optimize-cloud-cost-cloud-sustainability-with-dynatrace/#respond Thu, 26 Jun 2025 19:39:21 +0000 https://www.dynatrace.com/news/?p=69592 Dynatrace for Executives: Cloud cost optimization & sustainability

As organizations embrace a cloud- and AI-native future, the pressure to control infrastructure spending while meeting sustainability goals intensifies. As a CTO, I want my investments to go into people—building strong, innovative development teams—rather than overspending on cloud resources that don’t deliver business value. This is where Dynatrace plays a crucial role: helping organizations optimize […]

The post Maximizing cloud efficiency: Driving cost optimization and sustainability with Dynatrace appeared first on Dynatrace news.

]]>
Dynatrace for Executives: Cloud cost optimization & sustainability

As organizations embrace a cloud- and AI-native future, the pressure to control infrastructure spending while meeting sustainability goals intensifies. As a CTO, I want my investments to go into people—building strong, innovative development teams—rather than overspending on cloud resources that don’t deliver business value.

This is where Dynatrace plays a crucial role: helping organizations optimize cloud costs while advancing sustainability goals and enabling AI innovation. These priorities are no longer at odds; instead, they go hand in hand.

Key insights for executives

  • AI innovation and sustainability goals can go hand in hand. As AI workloads surge—projected to exceed 50% of cloud compute by 2028—organizations must balance innovation with cost and environmental impact. Dynatrace enables both by optimizing cloud usage in real time.
  • AI-powered observability empowers teams to cut waste, reduce emissions and align spending with business value. By providing deep insights into idle resources, inefficient architecture, and energy-heavy workloads.
  • Dynatrace improves efficiency and supports sustainability goals by dynamically scaling resources based on real-time data demand and business goals.

The true cost of cloud and AI

The rapid spread of AI—including LLMs, agentic AI systems, and coding assistants—and the shift to dynamic, multicloud environments have created yet more layers of complexity. AI workloads are compute-intensive. In fact, one of our customers in the banking sector shared that GenAI tasks cost five times more than traditional cloud workloads. Gartner® predicts that by 2028, more than 50% of cloud compute will be AI-related, up from just 10% in 2023.1

The International Energy Agency – Electricity 2024 report stated that when comparing the average electricity demand of a typical Google search (0.3 Wh of electricity) to OpenAI’s ChatGPT (2.9 Wh per request), and considering 9 billion searches daily, this would require almost 10 TWh of additional electricity in a year. That’s enough to power approximately 3 million households—or all private households in London—with energy for a year.

This growth brings significant environmental and financial implications. Yet, most organizations are beholden to opaque carbon footprint multipliers calculated by the cloud provider, which is often insufficient for actionable insights. Similarly, traditional cost reporting tools lack the depth and runtime visibility needed to drive meaningful optimization.

Dynatrace fills that gap, combining real-time observability, AI-powered insights, and topology-aware mapping to bring deep clarity into both cost and carbon impact.

Four steps to smarter cloud cost and energy management

1. Eliminate waste from idle or underutilized resources

Much of today’s cloud waste stems from overprovisioning and forgotten instances, especially in development and AI workloads. Dynatrace automatically detects underutilized or idle resources across the environments and surfaces insights that can drive decisions whether to shut down or re-size them, reducing both spend and carbon footprint.

Smartscape® automatic discovery and topology mapping adds unique value here, showing not just what’s idle, but whether it’s tied to business-critical processes or genuinely redundant.

2. Align cloud consumption to business value

Executives need more than cost data—they need to understand the why behind consumption. Dynatrace connects cloud utilization directly to applications, users, and business processes, enabling teams to assess whether resources are delivering real business value.

By linking costs to outcomes, organizations can prioritize what to keep, right-size what’s inefficient, and decommission what’s no longer serving a purpose.

3. Optimize architecture and energy efficiency

Most organizations are already taking basic steps like using contracted discount reserved instances or more flexible on-demand spot instances. The next level is architectural and source code optimization, such as green architecture and green coding. Dynatrace helps identify inefficient data flows, underperforming services, and high-cost cross-region transfers.

These insights enable teams to apply green coding techniques, reduce energy-hungry compute patterns, and bring data flows closer to where they’re needed, cutting both cost and carbon emissions.

4. Enable smart, automated orchestration

Finally, Dynatrace has a clear vision to make operations more autonomous. Its predictive, AI-driven orchestration of cloud resources enables teams to automatically scale resources up or down based on real-time demand, user behavior, and business impact.

However, autoscaling based on cloud metrics alone can’t ensure a great user experience or cost efficiency. Dynatrace links infrastructure and deep application observability to user-facing outcomes, allowing for smarter scaling that adapts dynamically to seasonal spikes, new product launches, or unexpected load while eliminating idle time and energy waste.

Accelerating sustainable innovation

Sustainability is now a strategic lever, not just a compliance checkbox. It resonates with environmentally conscious customers and a new generation of employees who want to work for conscientious companies.

By using Dynatrace Cost & Carbon Optimization and full-stack observability, organizations can:

  • Gain real-time, fine-grained insights into the energy and carbon impact of workloads
  • Make carbon reporting actionable and automatable instead of superficial
  • Build a more efficient, resilient, and future-proof cloud environment
Dynatrace Carbon Impact & Optimization dashboard
Figure1: Dynatrace Cost & Carbon Impact homepage

Imagine your cloud-native teams rapidly scaling up environments to test the scalability of new AI features, leading to a 40% spike in compute usage. Without visibility, one wouldn’t notice that this test left over idle or oversized instances, quietly driving up both cloud costs and carbon emissions. Now imagine having real-time insights from Dynatrace that reveal 200 idle instances across non-critical environments, costing thousands monthly and consuming unnecessary energy. Dynatrace AI leverages Smartscape® real-time topology to know automatically which instances can be confidently decommissioned or right-sized—cutting waste, aligning spend to business value, and advancing your sustainability goals.

The bottom line: Intelligent clouds mean a more sustainable planet

Organizations today must move beyond basic FinOps or simple sustainability checklists. The future lies in intelligent, self-optimizing clouds that balance performance, cost, and sustainability in real time.

Dynatrace empowers executives to realize this vision—transforming cloud environments into engines of innovation that are efficient, responsible, and aligned with business and environmental goals.

Follow along the new “Dynatrace for Executives” blog series. I’m diving deeper into each of the nine executive use case areas to help you unlock the potential of Dynatrace.
Want to learn more about all nine use cases? See the overview on the homepage.

1 Gartner Press Release, “Gartner IT Symposium/Xpo 2024 Orlando: Day 3 Highlights,” October 23, 2024, https://www.gartner.com/en/newsroom/press-releases/2024-10-23-gartner-it-symposium-xpo-2024-orlando-day-3-highlights.

GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.

The post Maximizing cloud efficiency: Driving cost optimization and sustainability with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/optimize-cloud-cost-cloud-sustainability-with-dynatrace/feed/ 0
Powerful exploratory analytics for AI-driven insights https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/ https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/#respond Tue, 04 Feb 2025 16:00:42 +0000 https://www.dynatrace.com/news/?p=67543 Problem alert dashboard

The Dynatrace platform empowers Operations, SRE, and DevOps teams to maintain high software quality, security, and reliability, allowing organizations to innovate and scale confidently. By leveraging Davis® AI with enhanced predictive analytics and automated workflows, Dynatrace simplifies issue detection and resolution, reduces MTTR, and enables proactive incident prevention.

The post Powerful exploratory analytics for AI-driven insights appeared first on Dynatrace news.

]]>
Problem alert dashboard


Deploying and safeguarding software services has become increasingly complex despite numerous innovations, such as containers, Kubernetes, and platform engineering. Recent global IT outages, such as the CrowdStrike incident, remind us how dependent society is on software that works perfectly.

Organizations must balance many factors to stay competitive.
Figure 1. Organizations must balance many factors to stay competitive.

Organizations strive to strike a delicate balance between cost, time to market, and innovation. This challenge is more pressing than ever as businesses seek to stay competitive while ensuring their software remains robust and secure.

This necessitates a comprehensive platform that empowers enterprises to understand IT and software within the broader context of their business operations, giving them confidence that their software and IT infrastructure are reliable.

Scale with confidence: Leverage AI for instant insights and preventive operations

Using Dynatrace, Operations, SRE, and DevOps teams can scale efficiently while maintaining software quality and ensuring security and reliability. Its AI-driven exploratory analytics help organizations navigate modern software deployment complexities, quickly identify issues before they arise, shorten remediation journeys, and enable preventive operations.

We’ve added numerous enhancements to our platform, leveraging advanced AI and automation for smarter software observability.

In this blog post, we show you how to

  • Get AI-driven insights directly on your operations dashboards
  • Improve MTTR with AI-assisted problem analysis and logs and traces in context
  • Leverage Gen AI through Davis CoPilot to get insights into root causes
  • Automate remediation of AI-detected problems with simple workflows
  • Adopt Preventive Operations with AI forecasting and automated action

Get AI-driven insights directly on your operations dashboards

A high-level, customizable view of your data is crucial in modern software operations. Dynatrace Dashboards, powered by Grail™ data lakehouse and Davis® AI, offer precisely that. They provide a comprehensive overview, seamlessly integrating health and problem-related information into a single view. You can chart your topology across data silos alongside all alerts, events, and problems using honeycomb tiles, which offer convenient drill-downs into the problem-debugging user flow.

Dynatrace ensures that context is seamlessly integrated into the platform, thus simplifying complexity for you as a user when analyzing issues and allowing you to focus on what truly matters. AI-driven analytics transform data analysis, making it faster and easier to uncover insights and act. This approach not only improves user experiences, it ensures that critical insights are accessible to both experts and novices. By simplifying remediation journeys and extending features to more user groups, Dynatrace enables results across all teams.

The new Problems dashboard, including rich honeycomb visualization, helps you focus on what’s important, turning technical data into a visual story.
Figure 2. The new Problems dashboard, including rich honeycomb visualization, helps you focus on what’s important, turning technical data into a visual story.

When a truly important issue stands out, the next step is refinement. With a few clicks, you can segment and filter your data to focus on specific applications, assignment groups, or regions. Directly mapping and surfacing ownership information within data segments accelerates incident assignment notifications and triggers automatic remediations.

Utilize the comprehensive filter functionality to update your dashboards dynamically.
Figure 3. Utilize the comprehensive filter functionality to update your dashboards dynamically.

If you see an issue or need to look closely at a specific application where an issue was identified, simply select the element to be seamlessly directed to the Problems app. There, you can dig deeper while continuing to focus on your selected segment. This tight integration, following a golden thread of insights, ensures that you’re more productive. To experience the possibilities of AI-empowered dashboards, try our example dashboard on the Dynatrace Playground.

Improve MTTR with AI-assisted problem analysis, logs, and traces in context

The Problems app delivers opinionated AI-assisted problem analysis optimized for Operations and Site Reliability Engineers (SREs) and developers. According to IDC, guiding users visually and automatically surfacing all critical details enables a 56% faster mean time to repair (MTTR) for critical incidents.

When a large-scale incident occurs, follow the red flag that Davis AI uses to identify the root cause, pinpoint all relevant details, and visually reproduce the details in charts, highlighting the affected deployment.

Analyze the root cause in the Problems app.
Figure 4. Analyze the root cause in the Problems app.

Besides identifying the root cause, Davis AI also automatically connects all relevant log lines. Logs are invaluable for identifying further insights and detecting fundamental flaws, such as process crashes or exceptions. With a single click in Problems, all incident logs are surfaced automatically. But we don’t stop there, Dynatrace also seamlessly integrates relevant trace data, offering full visibility into even complex, microservices-based architectures.

By providing these end-to-end insights, Dynatrace and Davis AI empower SREs, developers, and architects to quickly dive deep into an incident’s details, including all relevant logs and traces. Using this context, they can effectively focus on fixing and remediating code-level issues, significantly improving MTTR, and ensuring that critical incidents are resolved swiftly and efficiently.

Leverage GenAI via Davis CoPilot for insights into root causes

Dynatrace offers precision tools for domain experts to solve complex problems and dig deeper into their data. While product owners often focus on the intricate technical details of an incident, they often prefer a quick summary of what happened and what caused it. The soon-to-be-globally available Davis CoPilot™ bridges this gap by summarizing problems and their root causes and suggesting remediation steps based on these insights.

You’re not limited to one problem; Davis CoPilot can simultaneously analyze multiple problems, draw conclusions about their relationships, identify the common root cause, and propose corrective steps. Instead of relying on a team of experts and waiting hours for insights, Davis CoPilot helps you identify similarities and draw relevant conclusions independently and efficiently.

The use of generative AI adds significant value by augmenting Dynatrace-detected technical root causes with knowledge from the global tech community. Generative AI can access and synthesize vast amounts of information from various sources, providing a broader context and deeper insights. This ensures that your teams benefit from the latest advancements and solutions, enhancing their ability to resolve issues effectively and efficiently.


Dynatrace Problems App - Explain Problems video

Gain a better understanding of root causes with Davis CoPilot
Figure 5. Gain a better understanding of root causes with Davis CoPilot

Automate remediation of AI-detected problems with simple workflows

To automatically remediate Davis AI-detected problems, Dynatrace leverages powerful Workflows. Dynatrace workflows can be triggered by any problem or alerting event, automating domain-specific tasks to take remedial actions.

For example, workflows can scale up capacity to adapt to demand or automatically restart a service in case of a crash. With a large catalog of available workflow actions, you can react efficiently to AI-detected problems, reducing mean time to repair (MTTR) by automatically remediating issues.

But you can do much more with it: The recently introduced Simple Workflows, which are included in your Dynatrace subscription with no extra cost, offer greater flexibility and power than standard notifications. You can use the same mechanisms and trigger types to notify your developer team via Slack, create a JIRA issue, or send a PagerDuty alert.

This ensures that your operations, SRE, and DevOps teams can focus on more strategic tasks while the system handles routine problem resolutions. Automation enhances operational efficiency and ensures that your systems remain robust and reliable, even in the face of unexpected issues.

Easily set up automated remediation with the new Simple Workflows.
Figure 6. Easily set up automated remediation with the new Simple Workflows.

Adopt Preventive Operations with AI forecasting and automated action

Going beyond reactive problem detection, analysis, and remediation, Dynatrace can also leverage predictive AI to anticipate and avoid critical situations before they occur. Using Davis AI forecast, you can easily predict future capacity demands. Combining this knowledge with workflows allows you to take proactive measures to ensure system stability and performance.

Let’s have a look at a concrete example:

It’s easy to predict key indicators of your application, such as order levels or service request counts. Once load and demand rise and Davis AI identifies a potential future issue in your infrastructure setup, Davis CoPilot can automatically generate an updated Kubernetes configuration script for you and automatically upscale the environment to meet future demand. This ensures that your system scales appropriately to handle the anticipated demand, preventing incidents before they occur and eliminating the need to generate a problem.

That’s what we call Preventive Operations. Instead of sending an alert and notifying people, Dynatrace simply fixes the issue. According to Gartner’s Analytics Maturity Model, using predictive AI can significantly reduce the likelihood of incidents by taking preemptive action and remediation.

Start using Davis AI to analyze your environments and predict and address potential issues in advance. This will empower your teams to avoid potential problems and ensure a smooth, uninterrupted user experience.

Initiate automated, corrective action before an issue occurs
Figure 7. Initiate automated, corrective action before an issue occurs.

Tackle business challenges with confidence

Ensure your software runs securely and reliably with Dynatrace and Davis AI.

Dynatrace and Davis AI support you by running your software securely and reliably. This includes advanced root cause analysis, deep insights into detected issues, and corrective actions—whether manual or automatic—to prevent outages before they occur.

Get started

For more information, have a look at our documentation or explore the available resources on the Dynatrace Playground to experience some of these enhancements first-hand:

The post Powerful exploratory analytics for AI-driven insights appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/feed/ 0
Davis CoPilot expands: Get answers and insights across the Dynatrace platform https://www.dynatrace.com/news/blog/davis-copilot-expands-get-answers-and-insights-across-the-dynatrace-platform/ https://www.dynatrace.com/news/blog/davis-copilot-expands-get-answers-and-insights-across-the-dynatrace-platform/#respond Tue, 04 Feb 2025 16:00:17 +0000 https://www.dynatrace.com/news/?p=67510 Davis CoPilot

We’re excited to announce that Davis CoPilot Chat is now available across the Dynatrace platform. Davis CoPilot™, launched in October 2024 to support Dynatrace users with access to their data, now extends across the platform, streamlining user onboarding and providing comprehensive support and contextual insights from various Dynatrace® Apps. With the new Davis CoPilot conversational […]

The post Davis CoPilot expands: Get answers and insights across the Dynatrace platform appeared first on Dynatrace news.

]]>
Davis CoPilot


Update: We’ve launched Dynatrace Assist, our next-generation AI chat that goes far beyond answering questions.
Dynatrace Assist is the evolution of Davis CoPilot®.

We’re excited to announce that Davis CoPilot Chat is now available across the Dynatrace platform. Davis CoPilot™, launched in October 2024 to support Dynatrace users with access to their data, now extends across the platform, streamlining user onboarding and providing comprehensive support and contextual insights from various Dynatrace® Apps. With the new Davis CoPilot conversational interface, users can leverage natural language to quickly get answers to their questions, making it easier than ever for users to interact with Dynatrace.

Intuitive access to information boosts team productivity

We understand that taking advantage of the numerous features and functionalities offered by platforms like Dynatrace can be challenging. To help you navigate this and boost your efficiency, we’re excited to announce that Davis CoPilot Chat is now generally available (GA). This new feature provides information and guidance exactly when and where you need it, making your Dynatrace experience smoother and more efficient.

Davis CoPilot can be accessed anytime directly from the Dock.

Davis CoPilot leverages the power of generative AI to answer your questions through a globally accessible chat interface. We’re proud to say that Davis CoPilot is multilingual: you can ask questions and get answers in many different languages, including French, Spanish, German, Portuguese, Chinese, Japanese, and, of course, English. Davis CoPilot provides immediate, accurate responses, eliminating the need for extensive searches and reducing dependency on support channels. This makes knowledge more readily available and boosts productivity and user experience for both new and experienced users.

Davis CoPilot Chat follows our recent announcement of the general availability of Quick Analysis in Notebooks and Dashboards, which makes data accessible to technical and non-technical users alike. This means you can interact with data stored in the Dynatrace Grail™ data lakehouse just by using natural language.

Simplify onboarding and quickly find what you’re looking for with Davis CoPilot

You can start using the Davis CoPilot conversational interface immediately. Simply enable Davis CoPilot and assign the relevant user permissions, and the Davis CoPilot button will appear in the Dock.

Start a new conversation with Davis CoPilot Chat by selecting it in the Dock or by pressing CTRL/CMD + I and entering your question.

Davis CoPilot is great for guiding new and occasional users
Figure 2. Davis CoPilot is great for guiding new and occasional users

New users can quickly get up to speed with Dynatrace by asking Davis CoPilot for help with basic commands, setup instructions, and troubleshooting tips. This reduces the learning curve and enables new users to become productive faster. The conversational interface provides step-by-step guidance, making the onboarding process smoother and more efficient.

If you’re already familiar with Dynatrace, you can rely on Davis CoPilot to provide detailed explanations for a wide range of expert questions related to exploring new use cases, advanced configuration topics, and building custom apps.

Here are some examples of questions you can ask Davis CoPilot:

  • Onboarding: How do we start sending OpenTelemetry data to Dynatrace?
  • Understanding Dynatrace: What is the difference between an event and a problem in Dynatrace?
  • Exploring Dynatrace solutions: How can we comply with the Digital Operational Resilience Act (DORA) using Dynatrace?
  • Configuring your environment: How do I set up an alert based on an anomaly detector?
  • Developing custom apps: How can I import external table data and visualize it using the Dynatrace App Toolkit?

Get contextual assistance at the press of a button

Davis CoPilot seamlessly integrates into our use-case-specific Dynatrace Apps, offering you contextual insights and guidance at the press of a button. While we plan to release additional contextual app integrations in the coming months, several will be available a few weeks after launch, allowing Davis CoPilot to provide you with insights into:

  • Kubernetes warning signals
  • Individual problem details and the relationships between problems
  • Database performance optimization

Simplify Kubernetes: Davis CoPilot decodes warning signals

Understanding the background and root cause of warnings often requires in-depth subject matter expertise. That’s why we integrated Davis CoPilot into Kubernetes. Instead of manually looking up error messages, Davis CoPilot translates warning signals into clear, understandable language. In addition, Davis CoPilot offers a list of typical root causes and related remediation steps. This way, newcomers can quickly become proficient, and experts can elevate their expertise to hero status.

Davis CoPilot provides contextual guidance for Kubernetes warning signals
Figure 3. Davis CoPilot provides contextual guidance for Kubernetes warning signals

Problems demystified: Davis CoPilot provides insights into root causes

In Problems, Davis CoPilot provides clear summaries of problems, their root causes, and the suggested remediation steps. Davis CoPilot explains individual issues in clear language from the problem details page and can perform a comparative analysis when multiple problems are selected from the list view. This helps you identify common root causes and propose corrective steps without relying on a team of experts and waiting for hours for critical insights. If you want to learn more, have a look at Wolfgang Beer’s latest blog post and learn more about recent advancements in the Problems app.

Davis CoPilot explains problems in clear language
Figure 4. Davis CoPilot explains problems in clear language

Optimize database performance: Understand query execution plans

Query execution plans provide detailed information on how a database will execute an SQL query. While these provide the raw data on how to improve query performance and reduce resource consumption, they require expert knowledge to read and interpret. Now, in Databases, Davis CoPilot can provide natural language explanations of execution plans, breakdowns of relevant details, and recommendations on how to improve statement performance. This gives non-expert database users, such as developers, the knowledge they need to optimize their application performance and database utilization.

Davis CoPilot explains query execution plans
Figure 5. Davis CoPilot explains query execution plans

Empower your teams with Davis CoPilot today

The launch of Davis CoPilot Chat marks the second milestone of our journey. We’re committed to continuously enhancing the assistant’s capabilities with upcoming features, including query explanations, workflow actions, and troubleshooting guides.

Get started with Davis CoPilot today and transform how you and your teams interact with Dynatrace:

Thanks for joining us on this exciting journey. We look forward to your feedback and to seeing how Davis CoPilot helps your teams achieve their goals.

Davis CoPilot Chat, as well as the Dynatrace Apps integrations mentioned in this blog post, will be available starting with the release of Dynatrace SaaS version 1.307.

The post Davis CoPilot expands: Get answers and insights across the Dynatrace platform appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/davis-copilot-expands-get-answers-and-insights-across-the-dynatrace-platform/feed/ 0
Advancing AIOps: Preventive operations powered by Davis AI https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/ https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/#respond Tue, 04 Feb 2025 16:00:06 +0000 https://www.dynatrace.com/news/?p=67673 Davis AI alerts

The 2024 CrowdStrike incident demonstrated our societal vulnerabilities to IT outages. A faulty software update caused widespread issues, impacting critical services globally, including airlines, banks, hospitals, and public safety systems. Despite recent advancements such as containers, Kubernetes, and platform engineering, it’s evident that managing enterprise software services has become increasingly complex. IT operations must be prepared to quickly address and mitigate disruptions, ensuring business continuity and minimizing damage.

The post Advancing AIOps: Preventive operations powered by Davis AI appeared first on Dynatrace news.

]]>
Davis AI alerts

AI, especially AIOps, has emerged as a pivotal solution, promising to avoid downtime. The 2024 State of AI Report highlights this trend, with 89% of technology leaders anticipating that AI will significantly enhance incident response by learning to automate and optimize various tasks, such as performance monitoring and workload scheduling.

Blue screens of death at LGA airport due to the July 2024 CrowdStrike outage. (Source: Wikimedia Commons.)
Figure 1. Blue screens of death at LGA airport due to the July 2024 CrowdStrike outage. (Source: Wikimedia Commons.)

AIOps can identify and address potential issues before they become major incidents by learning from history and analyzing large amounts of data in real time. This approach improves operational efficiency and resilience, though it’s not without flaws. The complexity of IT environments and the changing nature of threats necessitate human oversight and ongoing adjustment of AIOps systems to handle unforeseen challenges and ensure optimal performance. Additionally, predictions based on historical data are reactive, solely relying on past information to anticipate future events, and can’t prevent all new or emerging issues. This limitation highlights the importance of continuous innovation and adaptation in IT operations and AIOps strategies.

“The shift from reactive to preventive operations represents the next evolution in AIOps.”
Bernd Greifeneder, CTO Dynatrace

When Dynatrace set out with Davis® AI over 10 years ago, pioneering AI-driven operations, we focused initially on problem identification before moving on to problem remediation. The next milestone in enhancing the capabilities of Davis AI—another pioneering step forward in AI-driven operations—is outright problem prevention. In this blog post, we explain how the unique combination of causal, predictive, and generative AI—augmented by the latest Davis AI advancements—is transforming how Dynatrace customers manage and optimize their IT infrastructure.

Automatic root cause detection

Modern, complex, and distributed environments generate a substantial number of events. This necessitates additional requirements such as minimizing the total number of issues, eliminating false positives, and conducting accurate root cause analysis.

Dynatrace has a longstanding reputation for accurately analyzing root causes and identifying related events. While other methods typically rely on mere correlation and historical data analysis, we’ve further enhanced our capabilities by implementing causational analysis, which leverages contextual information automatically gathered during data ingestion and processing in addition to historical data analysis. This is achieved using Dynatrace Grail™, our causational data lakehouse, which unifies all data in an always-up-to-date topology model. By applying causal AI to incoming data in real time, Davis instantly learns and continuously adapts to new information. This facilitates more precise root cause analysis and anomaly detection, including identifying seasonal anomalies and establishing auto-adaptive thresholds.

Root cause analysis with the Problems app
Figure 2. Root cause analysis with the Problems app

When applying this Davis root cause detection within our own IT environment, Davis effectively filters out over 99.9% of incoming data noise, condensing hundreds of thousands of daily system events into no more than four or five incidents that require attention from our IT operations team.

These algorithms are not limited to monitoring IT environments. At our February 2025 Dynatrace Perform session on exploratory analytics with AI-driven insights, the Performance Engineering Lead of XXXLutz—one of the world’s largest furniture retailers operating more than 370 stores across Europe—explains how XXXLutz utilizes Davis AI to proactively identify critical order drops, allowing them to respond quickly and effectively to changing market conditions and ensuring that their business remains agile and responsive to the needs of their customers.

Problem journey and reactive remediation

At the core of Dynatrace problem remediation stands the Problems app—an optimized view into opinionated insights, details, and context of each detected issue—for Operations, SREs, and developers. It filters billions of log lines, including the topology of each incident and its affected entities, for efficient problem triaging and troubleshooting, resulting in a 56% faster mean time to repair (MTTR) for critical incidents.

With the latest release, we drive this further by improving the automatic connection of relevant log and trace data for further drill down, presenting the full context of an issue in a single view. This provides comprehensive visibility into even complex architectures, simplifying the process of examining relevant details and addressing code-level issues, reducing 100 clicks and manual filtering to a single click with no loss of context.

Comparative analysis of multiple problems with Davis CoPilot
Figure 3. Comparative analysis of multiple problems with Davis CoPilot

By utilizing Davis CoPilot™, you can conduct comparative analyses of multiple issues, obtain natural language summaries of individual problems, and receive contextual recommendations along with specific remediation steps.

You can also link troubleshooting guides created in Notebooks to remediated issues, thereby building an intelligent knowledge base. Davis automatically connects additional documents as well as stored workflows. So the next time a similar problem arises, Davis brings up related guides, enabling teams to learn from previous experiences and reducing the risk of knowledge loss.

Harness your collective knowledge by connecting troubleshooting guides
Figure 4. Harness your collective knowledge by connecting troubleshooting guides

Please refer to our recent blog posts for more information on utilizing Problems for AI-driven insights and the latest Davis CoPilot advancements.

Automating the remediation

While obtaining comprehensive insights is beneficial, true transformation occurs through the use of tools that automatically execute remediation steps. To implement these “AI-driven operations,” it’s essential to forecast future requirements, including capacity demands, potential system failures, and security incidents.

Traditional forecasting engines typically depend on historical data, stored in metrics. In contrast, Davis AI generates real-time predictions, facilitating proactive operations. This capability is due to Davis’s ability to process raw data, such as logs, for forecasting, leveraging Grail to execute previously unattainable queries.

Consider the following scenario: You begin by retrieving and analyzing logs to identify relevant values for automation. Once this task is complete, you proceed to your pipelining tool to configure ingestion rules that extract these values into metrics and then wait several weeks for your prediction engine to generate alerts that can serve as triggers for your workflows.

However, when utilizing Dynatrace with its integrated anomaly detection and forecasting capabilities, you gain the advantage of schema-less data analysis and the ability to process any raw data into time series in real time. This significantly reduces the time required to establish AIOps workflows from several weeks to less than 30 minutes.

Preventive operations

The complexity of modern software environments makes it challenging to determine a service’s reliability solely through testing. It’s impractical to emulate scenarios such as generating a million tickets to assess performance capabilities. This necessitates real-time insights and operations rather than reactive problem-solving or raising alerts to notify personnel.

Preventive operations address this need by enabling proactive corrective actions before issues arise, akin to predictive maintenance. AI-supported anomaly detection identifies parameters that deviate from the norm, allowing for automatic configuration adjustment to mitigate potential problems preemptively.

Dynatrace offers the only unified, AI-powered platform for all data, all teams, and all possibilities.
Figure 5. Dynatrace offers the only unified, AI-powered platform for all data, all teams, and all possibilities.

Davis CoPilot combines the “power of three”:

  • Davis causal AI for identifying anomalies and root cause analysis
  • Davis predictive AI for precise forecasting and determining when to take action
  • Generative AI capabilities that perform actions beyond simply sending notifications or restarting services

In this way, Dynatrace extends AIOps beyond traditional IT operations tasks and addresses complex scenarios, including security use cases such as threat observability. Consider the following real-world example:

At Dynatrace, we log all failed login attempts. We can predict potential threats when abnormal patterns are identified and raise a security event by utilizing seasonal baselining. The subsequent workflow involves checking the IP address and generating a threat score. Upon reaching a certain threshold, a new ruleset is automatically added to the web application firewall. This entire process is fully automated, running before a problem even occurs, significantly reducing the response time from over an hour to a fraction of a second.

In another instance, automatic log pattern analysis crawling our application logs decreased the number of bugs in the production environment by 15% and freed up time previously spent on log analysis and triaging (in pre-prod), equivalent to 17 full-time employees. Consequently, these 17 developers can now dedicate their efforts to adding more value to Dynatrace.

Summary

The State of AI report states that over 88% of technology leaders anticipate AI will enhance incident responses and improve their teams’ ability to predict and proactively resolve service-affecting issues.

With Dynatrace, organizations are prepared to evolve their ITOps and SRE departments from troubleshooting to prevention, getting proactive with forecasting, and utilizing generative AI instead of purely focusing on history-focused root cause analysis.

Start your preventive operations journey with smart automation and auto-remediation that prevents larger issues.

Are you interested in gaining more insights?

The post Advancing AIOps: Preventive operations powered by Davis AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/advancing-aiops-preventive-operations-powered-by-davis-ai/feed/ 0
Better dashboarding with Dynatrace Davis AI: Instant meaningful insights https://www.dynatrace.com/news/blog/better-dashboarding-with-dynatrace-davis-ai/ https://www.dynatrace.com/news/blog/better-dashboarding-with-dynatrace-davis-ai/#respond Tue, 21 Jan 2025 21:38:07 +0000 https://www.dynatrace.com/news/?p=67370 abstract image showing connected dots and waves representing MCP best practices for agentic AI

Discover the value of Davis® AI when working with dashboards for observability, security, or business use cases. Quickly spot anomalies by activating Davis AI on any numeric time series chart data. Stay ahead with visual, AI-powered forecasting, or get new insights into your data with just a few clicks by leveraging Davis CoPilot™.

The post Better dashboarding with Dynatrace Davis AI: Instant meaningful insights appeared first on Dynatrace news.

]]>
abstract image showing connected dots and waves representing MCP best practices for agentic AI

Ensuring smooth operations is no small feat, whether you’re in charge of application performance, IT infrastructure, or business processes. Chances are, you’re a seasoned expert who visualizes meticulously identified key metrics across several sophisticated charts. Your trained eye can interpret them at a glance, a skill that sets you apart.

However, your responsibilities might change or expand, and you need to work with unfamiliar data sets. The market is saturated with tools for building eye-catching dashboards, but ultimately, it comes down to interpreting the presented information. This is where Davis AI for exploratory analytics can make all the difference.

Activate Davis AI to analyze charts within seconds
Figure 1. Activate Davis AI to analyze charts within seconds

Davis AI can help you expand your dashboards and dive deeper into your available data to extract additional information. Our customers value the nearly unlimited possibilities for querying and joining data on the Dynatrace platform, with the option of instant, real-time visualization of query results. Whether you’re an expert or an occasional user, our recently launched Davis CoPilot will enable you to get instant results without the need to write complex queries yourself. Have a look at our recent Davis CoPilot blog post for more information and practical use cases.

If you’ve already created your dashboards, now is the time to use Davis AI to identify anomalies or predict future trends without restricting use cases.

Leverage Davis AI for anomaly detection and instant insights

“My chart shows a peak at 8:00 AM. Do I need to investigate this further?” You might be regularly confronted with this or similar questions. Davis AI machine learning capabilities will help you identify actual anomalies within seconds, enabling you to focus resources on issues that matter.

Based on your requirements, you can select one of three approaches for Davis AI anomaly detection directly from any time series chart:

  • Auto-Adaptive Threshold: This dynamic, machine-learning-driven approach automatically adjusts reference thresholds based on a rolling seven-day analysis, continuously adapting to changes in metric behavior over time. For example, if you’re monitoring network traffic and the average over the past 7 days is 500 Mbps, the threshold will adapt to this baseline. An anomaly will be identified if traffic suddenly drops below 200 Mbps or above 800 Mbps, helping you identify unusual spikes or drops.
  • Seasonal Baseline: Ideal for metrics with predictable seasonal patterns, this option leverages Davis AI to create a confidence band based on historical data, accounting for expected variations. For instance, in a web shop, sales might vary by day of the week. Using a seasonal baseline, you can monitor sales performance based on the past fourteen days. An anomaly is identified if sales on a Friday are significantly lower than on previous Fridays, indicating a potential issue.
  • Static Threshold: This approach defines a fixed threshold suitable for well-known processes or when specific threshold values are critical. For example, if you have an SLA guaranteeing 95% uptime, you can set a static threshold to alert you whenever uptime drops below this value, ensuring you meet your service commitments.

Davis AI is particularly powerful because it can be applied to any numeric time series chart independently of data source or use case.

The following example will monitor an end-to-end order flow utilizing business events displayed on a Dynatrace dashboard. By leveraging Davis AI anomaly detection, we can identify potentially fraudulent behavior by activating anomaly detection on the Average order size chart. As shown in the chart below on the lower left, most values fall within the band of acceptable response time (highlighted in green), with only one spike occurring at 5:00 AM. Since this spike was outside the expected range, an anomaly was identified.

Apply Davis AI anomaly detection to detect fraudulent behavior in a business process
Figure 2. Apply Davis AI anomaly detection to detect fraudulent behavior in a business process
  • Application Observability: Identify unexpected error rate increases in application performance, helping pinpoint and resolve issues quickly.
  • Digital Experience Management: Monitor user interaction patterns to spot anomalies in website or app performance that could affect user experience, such as slow page load times.
  • FinOps: Track irregularities in cloud spending or resource usage, enabling cost optimization and preventing budget overruns.

Davis AI forecast analysis predicts future numeric values of any time series. It can even process external datasets or the results of any data query if it can be displayed as a numeric time series, such as occurrences over time.

The forecast is created instantly, even for large data sets, and updates dynamically whenever filter settings are changed.

In application performance management, acting with foresight is paramount. Maintaining reliability and scalability requires a good grasp of resource management; predicting future demands helps prevent resource shortages, avoid over-provisioning, and maintain cost efficiency.

On this SRE dashboard, we utilize Davis AI to forecast and visualize future resource utilization:

SRE dashboard monitoring the four golden signals and forecasting resource utilization
Figure 3. SRE dashboard monitoring the four golden signals and forecasting resource utilization

Other potential applications for forecasting include:

  • Kubernetes: Forecasting helps dynamically scale Kubernetes clusters by predicting future resource needs. This ensures optimal resource utilization and cost efficiency. Forecasting can identify potential anomalies in node performance, helping to prevent issues before they impact the system.
  • Business: Using information on past order volumes, businesses can predict future sales trends, helping to manage inventory levels and effectively plan marketing strategies.

AIOps: Utilize Davis AI to predict and prevent

Utilizing the Dynatrace AutomationEngine, Davis AI forecasting capabilities can even trigger automated actions. One of our customers’ SRE teams needed to increase disk space to avoid ongoing over- and under-provisioning, which was time-consuming and annoying. Now, with Davis AI forecasting capabilities, the target disk size is predicted automatically, and an automated task for disk resizing is triggered when necessary.

If you want to further explore the possibilities for prediction and prevention management with Dashboards, have a look at our example dashboard in the Dynatrace Playground.

Prevent incidents through predictive maintenance and capacity management
Figure 4. Prevent incidents through predictive maintenance and capacity management

Experience Davis AI in action

To experience the possibilities of Davis AI, look at this short introduction video by Andreas Grabner:
How to chart and forecast any data point

To explore the depth of functionality of Dynatrace Dashboards yourself and get first-hand experience, try out the app in the Dynatrace Playground.

The post Better dashboarding with Dynatrace Davis AI: Instant meaningful insights appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/better-dashboarding-with-dynatrace-davis-ai/feed/ 0
How to implement an AIOps strategy at scale https://www.dynatrace.com/news/blog/how-to-implement-an-aiops-strategy-at-scale/ https://www.dynatrace.com/news/blog/how-to-implement-an-aiops-strategy-at-scale/#respond Fri, 20 Dec 2024 15:42:15 +0000 https://www.dynatrace.com/news/?p=67149 AIOps strategy

Imagine a day when your IT team resolves critical incidents before users even notice them. In today’s multicloud world, complexity is the norm. But what if an AIOps strategy could transform that complexity into your organization’s greatest advantage? IT operations teams must monitor, maintain, and optimize a broader mix of applications, infrastructure, and technologies, with […]

The post How to implement an AIOps strategy at scale appeared first on Dynatrace news.

]]>
AIOps strategy

Imagine a day when your IT team resolves critical incidents before users even notice them. In today’s multicloud world, complexity is the norm. But what if an AIOps strategy could transform that complexity into your organization’s greatest advantage?

IT operations teams must monitor, maintain, and optimize a broader mix of applications, infrastructure, and technologies, with newer digital business solutions only adding to the strain. To manage it, organizations are turning to AI to help automate tasks such as anomaly detection, root-cause analysis, and incident response. But some IT operations teams are taking this approach a step further, applying multiple forms of AI to accelerate and enhance all business operations.

As AI is evolving, it’s important for organizations to understand the different types of AI and the keys to implementing them to achieve an AIOps strategy that moves from reactive to predictive problem solving.

The three kinds of AI that are key to a successful AIOps strategy

AI has evolved beyond the traditional correlation and probability-based approaches. Now, organizations turn to multiple forms of AI, such as causal, generative, and predictive AI, to manage cloud environments, secure data, and improve business decision-making. Therefore, implementing a successful AIOps strategy requires a deeper understanding of these types of AI.

Causal AI: Think of a global retail chain instantly pinpointing the root cause of a checkout slowdown across thousands of stores. Causal AI uses real-time, contextual data and causal dependencies for precise root-cause analysis and issue prevention. This establishes a business environment safeguarded by automated health monitoring and risk remediation.

Predictive AI: Imagine anticipating server outages hours before they occur, allowing seamless customer experiences. Predictive AI analyzes data patterns and trends, using statistical algorithms and other advanced machine learning techniques to anticipate future system behavior. This means technical users and business leaders can be much more proactive and innovative in the face of constant change.

Generative AI: Consider using generative AI to automate repetitive tasks, freeing up teams for innovation. Generative AI trains on large and diverse data sources, boosting business productivity and efficiency.  But when used in combination with causal and predictive AI and trained on real-time, high-fidelity observability data, generative AI can help organizations accelerate productivity and automate workflows.

The key strategy is using GenAI in conjunction with causal and predictive AI to understand your data and environment using natural conversation, and not as a way to correlate disparate events or anticipate the future. GenAI alone has limitations in these areas.

Implementing AIOps at scale

The multicloud complexity challenge demands a new approach—one that moves beyond toolchains that are stitched together. That’s where Davis AI™ comes in, combining the strengths of three AI capabilities in one: predictive, causal, and generative AI. Davis provides advanced analytics and proactive problem-solving to deliver out-of-the-box insights. Additionally, Davis doesn’t require extensive configuration or integrations. As a result, IT teams and business decision-makers can take immediate advantage of AI and scale it across the organization.

This power-of-three AI and observability approach democratizes the value of complex data analytics for both technical and nontechnical users. By unifying operations, security, development, and business teams with a single, complete, real-time view of activity across every cloud, app, and system, teams have an intuitive, self-service observability environment for problem-solving security and business events. Unified observability and AI put users in control of finding the answers they need from data, without needing to be experts in query languages.

Take your AIOps strategy from reactive to proactive

In addition to unifying teams, Dynatrace enables organizations to cut down on redundant observability tools and vendors, simplifying the overall view and improving insight coherency. And with new standards and capabilities in AI, teams can scale these efforts across the business for a more streamlined, inclusive, and collaborative business approach to resolving operational issues and driving your business forward.

Tool sprawl is an obvious problem for organizations. But choosing the wrong end-to-end observability platform can make things worse—and more expensive. Contact us today to request a demo of the AI-powered unified observability and security platform from Dynatrace and take your first step toward eliminating tool sprawl.

eBook: Developing an AIOps strategy for cloud observability

Download our free eBook to learn the best practices for developing an AIOps strategy that drives efficiency, innovation, and better business outcomes

The post How to implement an AIOps strategy at scale appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-to-implement-an-aiops-strategy-at-scale/feed/ 0
Five observability predictions for 2025 https://www.dynatrace.com/news/blog/observability-predictions-for-2025/ https://www.dynatrace.com/news/blog/observability-predictions-for-2025/#respond Fri, 13 Dec 2024 17:01:13 +0000 https://www.dynatrace.com/news/?p=67064 observability predictions for 2025

Rapid AI advancements, evolving regulations, and sustainability pressures make 2025 pivotal for observability, driving innovation and resilience.

The post Five observability predictions for 2025 appeared first on Dynatrace news.

]]>
observability predictions for 2025

As the digital world grows more complex, 2025 will bring a tipping point for organizations navigating increasingly dynamic and interconnected IT environments. Observability, long a cornerstone of IT operations, will take on transformative new roles. Driven by rapid advances in AI, evolving regulatory frameworks, and mounting sustainability pressures, observability will no longer be a passive diagnostic tool. Instead, it will lead to proactive, automated, and intelligent operations.

These prediction themes for 2025 outline how observability will evolve to meet the needs of a rapidly changing landscape. From new standards for automation and security convergence to redefining sustainability in IT, these shifts represent not just technological advancements but paradigm changes in how organizations operate, innovate, and compete.

Key insights for executives

  • Adopt preventive observability to stay ahead of disruptions. Shift from reactive to proactive IT management by leveraging AI-driven systems that autonomously predict and prevent issues before they become a problem, ensuring uninterrupted operations and enhanced customer satisfaction.
  • Integrate observability and security for continuous compliance. Simplify regulatory adherence and enhance resilience by implementing platforms that automate compliance reporting and proactively mitigate risks, safeguarding your reputation and bottom line.
  • Observability becomes mandatory for any serious sustainability strategy in IT. Instead of just reporting sustainability, leverage observability tools to optimize energy usage and reduce carbon footprints, achieving sustainability goals while lowering operational costs and meeting regulatory expectations.
  • Ensure trust in AI with robust observability. Monitor and validate AI-driven decisions with observability platforms that enforce ethical standards and prevent errors, building stakeholder trust and ensuring automation aligns with business objectives.
  • Leverage AIOps to enable preventive operations and boost agility. Replace reactive workflows with AI-powered, observability-driven systems to predict and resolve issues proactively, reducing costs, increasing efficiency, and accelerating time to market.

Here are five ways observability will shape the future, starting in 2025.

Prediction #1: Observability shifts from reactive to preventive

observability prediction: Observability shifts from reactive to preventive

Preventive observability will move beyond siloed systems into interconnected, autonomous ecosystems, redefining how organizations ensure reliability and resilience. These ecosystems will function seamlessly across distributed environments, leveraging AI to understand the real-time context of digital services. This capability enables automation to predict and prevent issues before they occur, leading to an era of foresight and collaboration.

Unlike earlier AIOps approaches that struggled to deliver due to limited contextual understanding, this new generation of AI-powered observability integrates insights from across systems to identify root causes, predict cascading failures, and act autonomously in real time.

Consider these examples:

  • A logistics company could leverage preventive observability to identify potential bottlenecks in supply chains and reroute shipments before delays occur.
  • In healthcare, observability could predict system slowdowns during critical periods, ensuring seamless patient care.
  • Financial services organizations could use it to preempt system outages during peak trading hours, protecting both customers and market stability.
  • During the holiday season, an e-commerce platform anticipating a traffic surge could use preventive observability to predict slowdowns or overloads, proactively scale resources, optimize performance, and balance cloud costs.

Imagine a future where operational insights are shared seamlessly across industries—whether logistics, healthcare, financial services, or online services—creating a level of interconnectedness that will drive operational excellence.

In these autonomous ecosystems—built on hybrid and multicloud environments—preventive observability automates the complex task of orchestrating distributed systems. By predicting and resolving issues before they impact operations, organizations can ensure service availability, minimize downtime, and reduce operational overhead. This proactive, context-aware approach will soon become the industry standard.

Prediction #2: Observability and security converge around continuous compliance

prediction for 2025: Observability and security converge around continuous compliance

In 2025, compliance will no longer be a static exercise—particularly for European and globally operating financial services companies—with more changes to follow. Continuous compliance will evolve into a real-time dynamic system driven by security standards and regulatory frameworks like the following:

  • EU’s Digital Operational Resilience Act (DORA) and Network and Information Security Directive 2 (NIS2),
  • Bank of England’s Operational Resilience Policy in the UK
  • Australia’s CPA 230
  • Hong Kong’s Monetary Authority Operational Resilience Framework
  • The Federal Reserve Regulation HH in the United States

This shift adds to the growing need for observability and security to converge, providing organizations with unified insights to address compliance, reduce redundant data collection, and strengthen threat detection and incident response.

For example, AI systems will continuously monitor threat exposure to assess risks and prepare configuration adjustments. These adjustments can be reviewed and approved by humans or applied automatically, ensuring organizations maintain compliance without disrupting operations. This human-in-the-loop approach is essential for maintaining accountability, particularly when regulatory violations must be reported to governing institutions or national competent authorities (NCAs).

Continuous compliance will replace periodic audits with automated systems that monitor, analyze, and alert on regulatory adherence. By integrating observability and security, organizations gain the additional context needed to qualify violations, track interdependencies across systems, and address vulnerabilities proactively. For instance, a financial services organization could leverage a unified observability and security platform that uses AI to identify potential compliance risks, such as service-level violations or third-party software vulnerabilities, and implement automated remediations to maintain adherence to stringent regulations.

The convergence of observability and security offers more than just regulatory benefits—it equips organizations to combat increasingly sophisticated cyber threats. Observability widens the lens through which security professionals view and analyze data, delivering the context necessary to enhance resilience and reduce costs. By integrating observability into security strategies, organizations can foster the trust needed to operate confidently in an era of heightened risk.

Prediction #3: Observability is mandatory for any serious IT sustainability strategy

observability predictions: Prediction #3: Observability is mandatory for any serious IT sustainability strategy

Sustainability will take center stage in 2025, as organizations face growing energy demands from cloud environments and increasingly AI-driven operations.

Observability platforms will become essential for monitoring and optimizing the energy consumption of AI workloads, identifying inefficiencies, and enabling intelligent workload distribution. As an added benefit, optimizing energy efficiency through observability not only lowers operational costs, but also aligns with sustainability commitments set by cloud providers. This approach ensures businesses stay competitive as energy costs rise and sustainability regulations tighten.

For example, a global retailer could leverage observability to track energy efficiency across its data centers. With platforms offering discovery and automatic topology mapping, a team can easily identify underutilized resources, revealing significant opportunities for architectural optimizations—often referred to as “green coding.” Such platforms can also enable smart orchestration for dynamic resource utilization. By embracing these strategies, the retailer could significantly reduce energy consumption and operational costs while fulfilling its environmental commitments. This evolution redefines IT’s role from a traditional cost center to a strategic enabler of sustainability.

As energy-intensive AI workloads become the norm and new sustainability quotas gain traction through regional mandates, like the EU’s Green Deal and Corporate Sustainability Reporting Directive (CSRD), sustainability will no longer be optional. Organizations that fail to integrate sustainability into their IT strategies risk non-compliance, reputational harm, and rising costs. Observability plays a leading role in this transformation by delivering the detailed insights necessary to optimize operations—not just report on sustainability. This shift helps businesses strike a sustainable balance between innovation and environmental stewardship.

Prediction #4: AI observability becomes indispensable for AI-driven services

Predictions for 2025: AI observability becomes indispensable for AI-driven services

In the evolution of digital transformation, the rise of AI-based services introduces new complexities that make observability more critical than ever. With observability, teams will be able to build and operate new AI-powered digital services for performance and reliability, keeping cost, AI drift, user experience, and transparency in mind. These capabilities will give organizations the confidence to deploy AI technologies at scale.

As businesses deploy AI-driven services for predictive maintenance, financial forecasting, or cybersecurity, observability platforms will go beyond monitoring system performance to include visibility into AI queries. With this clarity, organizations can identify potential errors, correct biases, and ensure decisions align with both business goals and ethical standards.

In 2025, observability will play an even larger role, as organizations increasingly rely on AI to power critical services. By providing end-to-end visibility and actionable insights, observability will empower businesses to confidently scale AI systems, maintaining accountability, reducing risks, and building trust. Therefore, observability will no longer be an optional enhancement, but a mandatory component for delivering safe and effective AI-driven services.

Prediction #5: AIOps is dead, long live AIOps!

observability predictions: AIOps is dead, long live AIOps!

AIOps has promised to transform operations for some time, but its fragmented technologies and limited context have prevented it from achieving its full potential. In 2025, AIOps will finally deliver on its promise, fueled by advances that enable AI systems to effectively communicate and collaborate. This evolution will redefine how organizations manage IT and business operations, setting a new standard for preventive operations.

The breakthrough comes from composing diverse AI techniques to work together toward common goals. Through interconnected AI systems—some specialized in prediction, others in precision processing of context, and others in suggesting remediations—AIOps will deliver intelligent automation and real-time root-cause analysis. These systems will predict disruptions, resolve issues before they escalate, and ensure business continuity. By integrating these capabilities, AI can provide deeper insights and take more precise actions by analyzing data in context and learning continuously from operational feedback.

For example, an enterprise managing complex, distributed environments could leverage this advanced AIOps approach to proactively address potential capacity bottlenecks during peak demand, ensuring smooth customer experiences while minimizing costs. Beyond IT, this approach will help organizations address broader challenges, such as forecasting supply chain risks or adapting to shifting market conditions.

The resurgence of AIOps will redefine industry benchmarks for efficiency and resilience. By automating tasks, reducing operational overhead, and enabling faster time-to-market, organizations will achieve unprecedented agility. To succeed, businesses must invest in advanced observability solutions and ensure teams are equipped to unlock the full potential of AI-driven operations.

The post Five observability predictions for 2025 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/observability-predictions-for-2025/feed/ 0
Helping customers unlock the Power of Possible https://www.dynatrace.com/news/blog/helping-customers-unlock-the-power-of-possible/ https://www.dynatrace.com/news/blog/helping-customers-unlock-the-power-of-possible/#respond Tue, 29 Oct 2024 18:32:13 +0000 https://www.dynatrace.com/news/?p=66365 Abstract image depicting unlocking business potential with Dynatrace using power dashboarding

At Dynatrace, The Power of Possible is about uncovering insights that move business forward. Discover how our customers unlock business potential with Dynatrace.

The post Helping customers unlock the Power of Possible appeared first on Dynatrace news.

]]>
Abstract image depicting unlocking business potential with Dynatrace using power dashboarding

We are in the era of data explosion, hybrid and multicloud complexities, and AI growth. In this dynamic landscape, imagine understanding your digital environment every day—what’s working, what’s not, what may have an issue, and more importantly, how to solve it? Picture gaining insights into your business from the perspective of your users. What new possibilities would open up for your organization?

This isn’t about imagining what’s possible just because it’s possible. It’s about uncovering insights that move business forward. This is what companies like BT and TD Bank have achieved by leveraging Dynatrace.

What’s behind it all? The Dynatrace platform automatically captures and maps metrics, logs, traces, events, user experience data, and security signals into a single datastore, performing contextual analytics through a “power of three AI”—combining causal, predictive, and generative AI. Dynatrace analyzes billions of interconnected data points to deliver answers, not just data and dashboards sending signals without a path to resolution.

With over 2.5 quintillion bytes of data generated daily, managing this influx has far surpassed human capacity. Dynatrace transforms this unstructured data into a strategic advantage, processing it automatically—no manual tagging required.

Unlike traditional observability tools that merely bark signals through dashboards without real context, the Dynatrace AI-driven approach goes deeper. It delivers precise, actionable insights to help you understand what’s truly happening in complex environments and resolve issues in real time. It empowers teams to act proactively rather than reactively. And it enables executives to have unprecedented insight into how user experiences, applications and underlying infrastructure health can power their business.

Let’s explore how leading organizations have harnessed the power of end-to-end observability—to reduce costs, drive innovation and acceleration, and deliver exceptional experiences for their customers. You’ll see how a clear line of sight across your entire technology stack can be transformative and learn how to apply these lessons to your own business.

Breaking down complexity with unified observability

In today’s landscape, it’s imperative to know the health of your digital business. But many enterprises face escalating complexity across their digital environments, leading to visibility gaps and inefficiencies using their traditional toolsets. BT, the UK’s largest mobile and fixed broadband provider, faced this challenge when managing multiple monitoring tools across different teams. By consolidating their tools into Dynatrace, they were able to reduce outage times and digital incidents by 50%. This meant better service reliability, reduced costs, and less time spent on incident management—enabling their teams to focus on innovation.

Leveraging AI-driven automation to innovate faster

As its business expanded, TD Bank encountered challenges with frequent IT incidents and slow resolution times, which were creating inefficiencies in their operations. To improve this, they turned to Dynatrace for AI-driven automation to accelerate problem detection and resolution.

By automating root-cause analysis, TD Bank reduced incidents, speeding up resolution times and maintaining system reliability. The result? More time for teams to focus on developing new services and improving customer experience, all while keeping operational costs under control. This ability to innovate faster has given TD Bank a competitive edge in a complex market.

The Power of Possible: Uncovering insights that move business forward

These stories illustrate how leading organizations leverage Dynatrace to achieve real business impact. For BT, simplifying their observability strategy led to faster issue resolution and reduced costs. TD Bank’s AIOps capabilities meant more proactive service management, keeping customers satisfied.

At Dynatrace, we believe that anything is possible for our customers, and we are committed to helping them move beyond simply managing their digital businesses. Our goal is to empower customers with insights that enhance user experiences and drive business growth. As organizations like BT and TD Bank navigate the complexities of modern IT environments, we are humbled to be their trusted partner, helping to turn data into action and transform challenges into opportunities.

Ready to see how Dynatrace makes the impossible possible? Watch the replay of The Power of Possible to learn more.

Want to go deeper? Learn more about the innovations driving these outcomes, and build a case to understand the ROI Dynatrace could deliver for your organization:

Ready to explore how end-to-end observability can unlock new possibilities for your business?

Contact us today to get started.

The post Helping customers unlock the Power of Possible appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/helping-customers-unlock-the-power-of-possible/feed/ 0
Don’t just react: How executives can predict and prevent outages to maximize availability https://www.dynatrace.com/news/blog/dynatrace-for-executives-improved-availability/ https://www.dynatrace.com/news/blog/dynatrace-for-executives-improved-availability/#respond Thu, 03 Oct 2024 14:15:16 +0000 https://www.dynatrace.com/news/?p=65679 Dynatrace for Executives: Improved Availability

I’ve seen firsthand the sleepless nights and high-stress environments that come with keeping digital services up and running in production. The stakes are high, and the pressure to deliver fast while maintaining uptime and preventing outages is relentless. With Dynatrace, executives can now benefit from predicting and preventing issues before customers are impacted and reducing […]

The post Don’t just react: How executives can predict and prevent outages to maximize availability appeared first on Dynatrace news.

]]>
Dynatrace for Executives: Improved Availability

I’ve seen firsthand the sleepless nights and high-stress environments that come with keeping digital services up and running in production. The stakes are high, and the pressure to deliver fast while maintaining uptime and preventing outages is relentless. With Dynatrace, executives can now benefit from predicting and preventing issues before customers are impacted and reducing the need to react. And when outages do occur, Dynatrace AI-powered, automatic root-cause analysis can also help them to remediate issues as quickly as possible. The end goal, of course, is to optimize the availability of organizations’ software.

Key insights for executives:

  • Predict and prevent outages before they happen with a unique combination of causal, preventive, and generative AI—enhanced by agentic AI for autonomous action
  • Remediate faster with automatic root cause analysis fueled by deterministic AI
  • Prioritize incidents based on customer impact insights from end-to-end traces
  • Automate to scale proactively and self-heal systems before customers are impacted

I realized that automating root-cause analysis requires a comprehensive approach: observing end-to-end and full-stack with deep insights, unifying all data in real time with up-to-date topology, and applying causal AI that learns instantaneously to handle cloud-native dynamics. Combining multiple types of AI made Dynatrace even stronger, and enables auto-optimize, auto-prevention, and auto-remediation all in one. However, the ultimate goal goes beyond technical excellence. For executives, the real business need is understanding customer impact—which is why it never made sense to me to just monitor servers but make end-to-end observability essential. This is what we uniquely solved for our customers with Dynatrace.

Respond to issues before they impact your customers

For executives, IT outages are a major headache. For issues that can’t be prevented in the first place, the next best option is to resolve issues faster than customers notice. Being faster, however, requires automation.

As the name Dynatrace suggests, dynamic tracing is at the heart of what we do. Dynatrace traces end-user interactions deep into the full stack of server-side activity to understand dependencies, allowing the platform to quantify the impact, qualify the situation, and prioritize actions. A power-of-three approach to AI, complimented with agentic AI, fuels automatic root-cause analysis to pinpoint the culprit amongst millions of service interdependencies and lines of code faster than humans can grasp.

Cloud technology complexity with billions of dependencies has outgrown human ability to manage and requires AI to analyze and comprehend. Dynatrace AI increases efficiency by magnitudes and prevents alert storms. This means you can avoid finger–pointing and war rooms, and dev teams’ productivity and happiness improve, eliminating business risk alert fatigue. Session replay capabilities provide visual proof and incident context so that teams can more easily understand and act upon the root cause. Automatic root cause analysis with Dynatrace can ultimately reduce mean time to repair (MTTR) by 90% or more.

Dynatrace is widely recognized for its causal, predictive, and generative AI capabilities, which can predict and prevent issues and automatically identify root causes, maximizing availability.

As responsibilities shift left due to the increased use of cloud-native technologies, development teams take more control over production deployments. While I am excited that the people who create software are also responsible for it—in contrast to “throw over the wall” approaches—it poses consistency and compliance challenges in larger organizations. That’s why we have Dynatrace extended (not shifted) to the left to address both needs: developers have easy and safe access to staging and production deployments while central SRE and DevOps teams have the scalable and automatic observability they need to remain compliant, consistent, and resilient. Finally, a standardized approach to observability coupled with self-service for departmental users reduces tool sprawl and complexity.

Gone are the days when executives could afford for their teams to stare at dashboards 24/7 to manually interpret data and act on runbooks. By unifying observability data and applying advanced AI, Dynatrace progresses to a new generation of AIOps that can predict and prevent issues and leverage automation for self-healing.

Predict and prevent outages with AI

The 2025 State of Observability Report found that 100% of business leaders are now using AI in their operations, with the top anticipated benefits being real-time anomaly detection (41%) and improved detection and response to security risks (37%). In this journey, many organizations have investigated AIOps tools to improve pattern analysis and noise reduction, as most of these solutions provide only correlation, not true causation. Even worse, the idea that such systems learn from past outages is flawed, as training would require thousands of production outages that no executive can afford.

So, to truly predict and prevent issues, the complexity of systems must be captured instantaneously and continually assessed in full context, through AI that maps causation in real-time. Dynatrace addresses this need with causal, predictive, and generative AI capabilities in a single framework. This approach eliminates the need for learning from past outages and enables a highly automated software delivery process, maximizing resilience.

Moreover, along with the maturity of the market to use agentic AI and AI overall, the ability to make use of Dynatrace capabilities has expanded too – since we have pioneered automation of operations. On this path, we address the skepticism in AI usage by fusing deterministic AI and agentic AI, making AI more reliable. And we see executives are more willing to drive steps towards more proactive automation and open the doors to autonomous operations.

IT teams can also embed quality gates into their workflows so they continually meet the thresholds for user experience defined through service-level objectives (SLOs). As a result, they can predict capacity demands based on seasonal patterns and use causal dependencies to automatically capture and prevent problems as they emerge.

Improving availability to meet ever-growing customer expectations requires high grades of automation for scale, agility, and resilience. This includes auto-scaling, overload protection, auto-remediation, auto-rollback, auto-quality-gating, and more. Eventually, the goal is to arrive at self-healing through autonomous cloud operations.

Therefore, platform engineering emerges as a discipline for a holistic approach to software, infrastructure and delivery, with a relentless aim to automate. Automation, however, should not be done in isolation of tech. It needs to execute in the context of the business, which requires insights into business-impacting metrics including end-user experiences, public API call success rates, learning from seasonal changes, and strategic business considerations such as cost vs. performance goals.

That’s where observability from Dynatrace goes far beyond “observing systems.” Dynatrace observability provides AI, analytics, and automation that integrates with platform engineering, continuous delivery, and automated operations. This greatly offloads DevOps, SRE, and operations teams from manual tasks and allows them to shift their work to automation tasks. Note that the work doesn’t get reduced. The key benefit is increased availability and security, faster software delivery, improved productivity, and cloud cost optimization.

New certification and security legislation projects, such as the Digital Operational Resilience Act (DORA) in Europe, are emphasizing the heightened expectations for digital systems availability. DORA further requires continuous compliance and the ability to report on the status, placing a heavy burden on organizations. This is where Dynatrace provides additional help and automation with the new Compliance Assistant app.

Likewise, since availability is affected by not only technical issues but also security threats, observability, and cloud security must converge to minimize availability issues. That is where Dynatrace AI and analytics—on top of unified observability and security data—raise the bar to proactively prevent problems and remediate them faster.

Want to learn more about all nine use cases? See the overview on the homepage.
In case you missed it, we hosted a must-see streaming event unveiling the innovations that are powering a new era of possibility for customers all over the world. Watch the on-demand recording now.

The post Don’t just react: How executives can predict and prevent outages to maximize availability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-for-executives-improved-availability/feed/ 0
AIOps strategy unlocks new possibilities for automation, customer satisfaction https://www.dynatrace.com/news/blog/aiops-strategy-unlocks-new-possibilities-for-automation-customer-satisfaction/ https://www.dynatrace.com/news/blog/aiops-strategy-unlocks-new-possibilities-for-automation-customer-satisfaction/#respond Tue, 01 Oct 2024 14:36:07 +0000 https://www.dynatrace.com/news/?p=65865 How to implement an AIOps strategy at scale

From managing complex IT environments to ensuring seamless customer experiences, the demands on IT departments have never been greater. To manage these complexities, organizations are turning to AIOps, an approach to IT operations that uses artificial intelligence (AI) to optimize operations, streamline processes, and deliver efficiency. One Dynatrace customer, TD Bank, placed Dynatrace at the […]

The post AIOps strategy unlocks new possibilities for automation, customer satisfaction appeared first on Dynatrace news.

]]>
How to implement an AIOps strategy at scale

From managing complex IT environments to ensuring seamless customer experiences, the demands on IT departments have never been greater. To manage these complexities, organizations are turning to AIOps, an approach to IT operations that uses artificial intelligence (AI) to optimize operations, streamline processes, and deliver efficiency. One Dynatrace customer, TD Bank, placed Dynatrace at the center of its AIOps strategy to deliver seamless user experiences.

Why AIOps?

AI for IT operations (AIOps) uses AI for event correlation, anomaly detection, and root-cause analysis to automate IT processes. It plays a crucial role in managing complex multicloud environments by streamlining operations and enhancing efficiency, reducing costs, and driving innovation.

Paired with an observability platform, AIOps identifies and helps to resolve cloud application performance and security issues, preventing problems before they disrupt operations. Its adoption is growing rapidly, driven by the explosion of data complexity that accompanies modern cloud IT environments. Valued at $17 billion annually, the AIOps market reflects its importance as large companies increasingly integrate AIOps and digital experience monitoring tools, with adoption expected to rise significantly in the coming years.

As a leader and trailblazer in the AIOps space, Dynatrace uses AI-powered root-cause analysis to provide precise actionable insights, enabling businesses to automate operations across the enterprise.

TD Bank adopts an observability-based AIOps strategy

As one of the 10 largest banks in the U.S., with $1.4 trillion in assets and 27 million customers, TD Bank places customers at the center of everything it does. As its enterprise monitoring team modernized the bank’s digital ecosystem from legacy on-premises data centers to a hybrid multicloud environment, TD Bank faced significant challenges with complexities its traditional monitoring tools couldn’t handle.

TD Bank’s modernized technology stack became increasingly intricate, leading to operational inefficiencies. The bank had accumulated multiple monitoring tools, each providing fragmented insights. This disjointed approach made it difficult to collaborate effectively and resolve issues promptly.

The Dynatrace unified observability platform provided TD Bank with a single source for answers, offering end-to-end visibility across its entire technology stack. Using Dynatrace at the center of its AIOps strategy, the TD Bank team reduced the number of IT incidents they were experiencing, improving customer trust.

Faster responses for greater reliability

A standout feature of Dynatrace is its ability to deliver rapid and precise answers. For TD Bank, this meant significantly reducing the time to identify and resolve transaction failures. With AI-driven certainty, the bank could instantly pinpoint the root cause of issues, leading to a 25% increase in proactive incident identification and a 20% faster response rate. This efficiency translated to a dramatic reduction in the transaction failure rate, from 0.16% to just 0.06%.

Cost optimization and efficiency

Using Dynatrace, TD Bank was able to consolidate its observability tools and achieve substantial cost savings. With its platform-based approach to end-to-end observability, Dynatrace enabled the bank to eliminate up to seven redundant monitoring solutions, reducing infrastructure and licensing costs by up to 45%. Beyond cost savings, this consolidation freed up TD Bank’s teams to focus on innovation rather than routine maintenance, driving further efficiency.

Enhanced customer satisfaction

For TD Bank, customer satisfaction is paramount. With the efficiencies stemming from Dynatrace AI capabilities, the bank reduced customer irritants by more than 60% and sped up issue resolution by 20%. With precise answers, TD Bank’s teams can quickly understand issues and resolve customer calls , enhancing the overall customer experience and building trust in the bank’s digital services.

How Dynatrace delivers on AIOps

AI-powered root-cause analysis

At the heart of Dynatrace AIOps capabilities is its power-of-three AI engine, Davis®. Using causal, predictive, and generative AI, Davis delivers detailed insights into issues, including their root cause and impact. For TD Bank, the technology has been instrumental in quickly identifying and resolving issues, ensuring minimal disruption to customer services.

Automated remediation

By automating routine tasks and responses to common issues, Dynatrace helps businesses like TD Bank achieve zero-touch operations. This automation not only improves efficiency but also ensures teams can address critical issues promptly, minimizing downtime.

Predictive analytics

Dynatrace AI-driven predictive analytics provide foresight into potential issues before they occur. For enterprises, this means staying ahead of the curve, preventing disruptions, and ensuring seamless operations. TD Bank has used these capabilities to anticipate and mitigate risks, ensuring a smooth banking experience for its customers.

The broader effect of AIOps: Transforming IT operations

AIOps is not just a tool; it’s a transformation strategy. By integrating AI into IT operations, businesses can achieve unparalleled efficiency, agility, and resilience. It is clear that the future of IT operations lies in AI, and Dynatrace is leading the charge. With its deep-rooted AI expertise and innovative AIOps platform, Dynatrace is transforming the way businesses operate. The success TD Bank has achieved demonstrates how Dynatrace helps unlock the hidden value in customer data and maximize the tangible benefits of AI-driven operations.

Discover the power of Dynatrace to unlock the future of IT operations and transform your business. Sign up for a free trial today and experience the difference Dynatrace AI can make.

The post AIOps strategy unlocks new possibilities for automation, customer satisfaction appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/aiops-strategy-unlocks-new-possibilities-for-automation-customer-satisfaction/feed/ 0
Why 85% of AI projects fail and how Dynatrace can save yours https://www.dynatrace.com/news/blog/why-ai-projects-fail/ https://www.dynatrace.com/news/blog/why-ai-projects-fail/#respond Wed, 03 Jul 2024 14:20:58 +0000 https://www.dynatrace.com/news/?p=64570 Generative AI poised to have an impact by automating software development. And why AI projects fail

Artificial Intelligence (AI) has the potential to transform industries and foster innovation. However, navigating the path to successful AI deployments can be quite challenging, leaving many organizations to wonder why their AI projects fail. Why AI projects fail According to one Gartner report, a staggering 85% of AI projects fail. Several factors contribute to this […]

The post Why 85% of AI projects fail and how Dynatrace can save yours appeared first on Dynatrace news.

]]>
Generative AI poised to have an impact by automating software development. And why AI projects fail

Artificial Intelligence (AI) has the potential to transform industries and foster innovation. However, navigating the path to successful AI deployments can be quite challenging, leaving many organizations to wonder why their AI projects fail.

Why AI projects fail

According to one Gartner report, a staggering 85% of AI projects fail. Several factors contribute to this high failure rate, including poor data quality, lack of relevant data, and insufficient understanding of AI’s capabilities and requirements. These issues underline the importance of robust data management and precise strategic planning for AI projects, including cloud-based models and LLMs.

The data challenge in AI projects

Data is the lifeblood of AI and machine learning (ML) projects. Without robust data, AI models struggle to produce accurate and reliable results. A NewVantage survey from 2024 highlights this issue, with 92.7% of executives identifying data as the most significant barrier to successful AI implementation. Moreover, a Vanson Bourne survey reveals that 99% of AI and ML projects encounter data quality issues. These statistics underscore the critical need for effective data management and monitoring solutions.

The role of data observability

Data observability refers to the ability to monitor and understand the state of data systems. It involves tracking data quality, lineage, and performance across data pipelines.

Organizations can ensure the success of an AI project by monitoring the data’s freshness, volume, distribution, schema, and lineage. Dynatrace offers comprehensive data observability features designed to significantly enhance the success of AI projects, for example:

  1. Data freshness monitoring. Monitoring data freshness helps to ensure that your data-driven decisions are based on the most current information. When making decisions, relying on the freshest and most relevant information available is crucial.
  2. Volume monitoring. By monitoring data volume, you can receive alerts about unexpected changes in volume, which can signal potential issues.
  3. Distribution monitoring. This feature helps identify anomalies, patterns, and outliers in your data, enabling you to detect any irregularities.
  4. Schema monitoring. This feature detects and flags unanticipated changes in data structures so you can stay informed about any unexpected adjustments.
  5. Lineage tracking. This feature offers a clear view of the origins and impacts of data, which is essential for troubleshooting and optimizing data flows. Understanding the origin and impact of data both upstream and downstream, also known as data lineage, provides valuable context for comprehending the data’s journey and reliability.

AI Observability: Ensuring model performance and reliability

However, we can’t just stop at data observability. As already mentioned, data isn’t the sole reason why AI projects fail. Insufficient understanding of AI’s capabilities and requirements is also part of it.

AI observability involves monitoring and understanding AI models and systems to ensure they perform as expected in real-world applications. It includes:

  • Performance monitoring. Tracking metrics like accuracy, precision, recall, and token consumption.
  • Explainability. Understanding how models make decisions to ensure transparency and interpretability.
  • Drift detection. Identifying shifts in model performance that indicate issues with accuracy or relevance.

The following diagram shows how AI Observability can detect and resolve degradation in an AI model’s performance:

AI observability workflow combats why AI projects fail

Comprehensive monitoring with data and AI observability

Data and AI observability work together to provide a holistic view of the system. Data observability concentrates on the pipeline and infrastructure, while AI observability delves into the model’s performance and results.

Feedback mechanisms

Effective AI observability relies on feedback loops, which rely on data observability. Identifying data drift, for instance, necessitates a profound comprehension of the data pipeline and its lineage.

Swift issue resolution

Discrepancies in AI performance are often rooted in problems within the data pipeline. Data observability aids in tracking these issues to their source, enabling the swift identification and resolution of problems affecting AI systems.

Compliance and governance assurance

Ensuring compliance with regulations and governance standards is crucial. Data observability ensures data management aligns with industry standards, while AI observability ensures that models adhere to ethical guidelines and regulations.

Ensuring AI excellence: How Dynatrace enhances model reliability

Dynatrace equips organizations with the necessary tools for data and AI observability, helping build sustainable AI applications. By providing comprehensive monitoring and understanding of data systems and AI models, Dynatrace ensures that AI projects are based on high-quality data and that AI models perform reliably and transparently in real-world applications.

Watch the Ensure AI Project Success with AI Observability webinar for a demo and more!

The post Why 85% of AI projects fail and how Dynatrace can save yours appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/why-ai-projects-fail/feed/ 0
Dynatrace completes 2024 FedRAMP Moderate reauthorization with Rev.5 transition https://www.dynatrace.com/news/blog/dynatrace-completes-2024-fedramp-moderate-reauthorization/ https://www.dynatrace.com/news/blog/dynatrace-completes-2024-fedramp-moderate-reauthorization/#respond Wed, 26 Jun 2024 16:45:46 +0000 https://www.dynatrace.com/news/?p=64466 FedRAMP

Dynatrace completed its FedRAMP Moderate reauthorization with the transition from Rev.4 to Rev.5. This achievement underscores our ongoing commitment to providing secure and reliable solutions for U.S. government agencies.

The post Dynatrace completes 2024 FedRAMP Moderate reauthorization with Rev.5 transition appeared first on Dynatrace news.

]]>
FedRAMP

Understanding FedRAMP Moderate and transition to Rev.5

FedRAMP (Federal Risk and Authorization Management Program) is a government program that provides a standardized approach to security assessment, authorization, and continuous monitoring for cloud products and services for U.S. state and federal agencies. The FedRAMP Moderate baseline is designed to protect sensitive data that, if compromised, could seriously adversely affect operations, assets, or individuals.

What you need to know

FedRAMP Revision 5 (Rev.5) prioritizes customizing security controls for specific risks. It aligns with Cloud Service Providers, providing a baseline while allowing tailored adjustments for individual federal agencies. FedRAMP adopts a threat-based approach to security controls. Rather than a one-size-fits-all model, this methodology tailors controls based on specific threats and risks. The result is a more effective security posture.

FedRAMP increased emphasis on privacy, which takes center stage in Rev.5, including:

  • Configuration Change Control and CM-4 – Impact Analysis now requires privacy impact analysis for configuration changes.
  • Role-based training requires privacy training alongside security training.
  • System Backup now requires the backup of privacy-related system documentation.

FedRAMP Rev.5 has some notable changes in control families and controls, such as:

  • SR – Supply Chain Risk Management is a new addition to the Rev.5 control family that more comprehensively addresses the risks associated with acquiring, developing, and maintaining information systems and components associated with third-party and vendor services, products, and supply chains.
  • Public Disclosure Program, which requires a reporting channel for the public to notify Cloud Service Providers of vulnerabilities.

FedRAMP assessments for Moderate and High systems now require an annual Red Team exercise (in addition to the previously required penetration tests). These exercises go beyond penetration testing by targeting multiple systems and potential avenues of attack. They help organizations understand risks, improve processes, and boost security readiness.

Dynatrace for U.S. government

Dynatrace for the U.S. government enables federal agencies to accelerate cloud adoption while ensuring compliance with stringent security standards. It provides deep, AI-powered insights across the entire digital ecosystem, facilitating proactive resolution, enhanced collaboration, and streamlined operations. This robust solution supports the U.S. government’s mission-critical applications by optimizing performance, reducing costs, and driving innovation, ultimately leading to improved service delivery for the public.

We continue to invest in our security infrastructure, refine our processes, and expand our capabilities to meet the evolving needs of U.S. government clients. By maintaining our FedRAMP Moderate status and continuously enhancing our offerings, including our commitment to achieve FedRAMP High, Dynatrace remains a trusted partner for U.S. government agencies seeking reliable and secure cloud solutions.

Learn more about how Dynatrace helps modernize U.S. government agencies with automatic and intelligent observability.

The post Dynatrace completes 2024 FedRAMP Moderate reauthorization with Rev.5 transition appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-completes-2024-fedramp-moderate-reauthorization/feed/ 0
Learn how to create a Davis AI anomaly detector on Grail https://www.dynatrace.com/news/blog/create-a-davis-ai-anomaly-detector-on-grail/ https://www.dynatrace.com/news/blog/create-a-davis-ai-anomaly-detector-on-grail/#respond Tue, 11 Jun 2024 19:22:55 +0000 https://www.dynatrace.com/news/?p=64325 Davis AI Fetch logs

From working with Dynatrace Notebooks, you know that exploratory analytics are crucial for uncovering the narratives within your organization’s data. By leveraging visual data analytics and collaboration input from development, security, and business teams, such insights become transparent, enabling immediate understanding and action on the implications for your business. Further, it’s essential to take automated actions to proactively use anomaly detection to determine if your business is at risk. Such anomaly detection should be implemented in straightforward steps, as described in this blog post.

The post Learn how to create a Davis AI anomaly detector on Grail appeared first on Dynatrace news.

]]>
Davis AI Fetch logs
Update: We’ve enhanced anomaly detection on Grail with Dynatrace Intelligence, enabling smarter, AI-powered insights and automated actions across the Dynatrace platform.
Dynatrace Intelligence builds on Davis AI®, advancing how teams detect, analyze, and respond to anomalies in their data.

Dynatrace Grail™ data lakehouse provides contextual analytics across unified observability, security, and business data. It allows you to query and combine data anytime using the Dynatrace Query Language (DQL). This enables exploratory data analysis and the ability to collaborate visually on the results with your colleagues.

Anomaly detection in Notebooks

You likely encounter “why” questions in your daily work. Why did we have an outage? Why did the system behave differently? Why did I receive an alert? These questions can be effectively investigated in Dynatrace Notebooks, where you can easily compile the necessary data and break it down into a time series. However, in the time series example below, we must determine whether the number of access attempts to our example Travel Mobile app is normal or abnormal.

Figure 1: Generated time series based on access logs in Notebooks
Figure 1: Generated time series based on access logs in Notebooks

In many cases, it’s evident, based on your past experiences looking at time series data, whether or not something is an anomaly. But how can you automate your expertise? Such automation could ensure that you and your colleagues don’t have to manually monitor time series to identify whether or not they include anomalies.

Davis® AI provides such automated anomaly detection out of the box. Still, your business requires the flexibility of Davis AI to detect anomalies based on your specific requirements, for example, to automatically generate a Davis problem based on a detected anomaly. For this purpose, we provide the Davis AI Analyzer, which allows you to select a specific analyzer. Three anomaly detection analyzers are available, each equipped with unique mechanisms to detect anomalies in your data that significantly deviate from the norm.

One unique feature of the Davis AI Analyzer is that it works on any time series, regardless of its origin—whether generated with makeTimeseries from events, business events, logs, or other sources or the joining of different time series. As you can see in the screenshot below, Davis AI Analyzer gains the full power of DQL, making Davis anomaly detection even more flexible and stronger than ever. This power can be easily experienced by selecting the desired Davis anomaly detection analyzer in Notebooks or Dashboards.

Figure 2: Using the seasonal baseline anomaly detection analyzer in Notebooks.
Figure 2: Using the seasonal baseline anomaly detection analyzer in Notebooks.

By selecting the seasonal baseline analyzer, Davis AI recognizes that the number of attempted accesses to the app in this example doesn’t deviate from the norm based on the past data during the same period. The time series falls within the seasonal green confidence band. A potential alert would be visually simulated if the time series fell outside this band.

This anomaly detector observes the number of attempted accesses per minute and triggers an event when anomalies are detected. You can create a similar Davis anomaly detector in a few simple steps.

Automate your experience with Davis Anomaly Detection

In Notebooks, select open with and choose Davis Anomaly Detection; all settings required for creating an anomaly detector will be carried over.

Create a new anomaly detector in Davis Anomaly Detection.
Figure 3: Create a new anomaly detector in Davis Anomaly Detection.

The new anomaly detector is created in four steps; the first two steps are carried over automatically from Notebooks. Let’s start with the most straightforward step, Get started, where you define a title for your anomaly detector and a description for the configuration.

The next two steps, as mentioned, have already been prefilled from Notebooks. In the Configure your query step, you’ll find the DQL query you predefined, and in the Customize parameters step, you’ll find your selected anomaly detection analyzer. The last significant step, the Create an event template step, remains. Here, you can define the template for your event and describe all essential information for the subsequent process.

Define the description and properties in the event template.
Figure 4: Define the description and properties in the event template.

What makes this template exceptional is that you can use {placeholder} hints to add additional context to the text about the event. For example, the value of the violation or the source entity where the anomaly was detected. This ensures that all essential information about the event is immediately visible to the Site Reliability Engineer (SRE). After completing all four steps, we can create the Davis anomaly detector by selecting Create. The anomaly detector will automatically monitor your defined time series every minute and trigger your specified event upon detection of an anomaly.

The new anomaly detector is now listed in Davis Anomaly Detection. Here, you’ll find all anomaly detector configurations, and you can filter them according to your specific criteria. Additionally, you can expand this table with extra information about the configurations, such as when the anomaly detectors were last modified.

Overview of anomaly detectors available within Davis Anomaly Detection.
Figure 5: Overview of anomaly detectors available within Davis Anomaly Detection.

Of course, you always have the option to reopen an anomaly detector directly in Notebooks, where all configuration settings are carried over. You also have the option to display a preview of your anomaly detector directly in Davis Anomaly Detection.

Figure 6: Visualize your custom anomaly detectors in Notebooks without leaving Davis Anomaly Detection.
Figure 6: Visualize your custom anomaly detectors in Notebooks without leaving Davis Anomaly Detection.

The exciting challenge is finding answers to your everyday “why” questions using Grail and DQL analytics capabilities. If the answer is successfully identified in a time series and you want to automate the result with anomaly detection, this can be done in just a few steps. We recommend you explore the new Davis Anomaly Detection analyzer in Notebooks; we’re confident you’ll quickly discover its many uses.

Try out Davis Anomaly Detection

Want to know more? Check out the following video, in which Andreas Grabner and I collaborated on a new episode of the Dynatrace Observability Clinic. Here, we share a live introduction to Anomaly Detection based on DQL.

We also recommend watching the exciting use case for Anomaly Detection and the 5 Pillars of Data Observability.

What’s next

Davis Anomaly Detection is automatically enabled for all Dynatrace SaaS environments with the release of Dynatrace version 1.291. No effort is needed from your side. We’re, of course, highly interested in your feedback. So, please head to the Dynatrace Community and share your suggestions and product ideas to help us continuously improve Dynatrace Anomaly Detection.

Are you interested in learning more? In Dynatrace Documentation, you can learn more about Davis Anomaly Detection and how to use anomaly detection within Notebooks, or look at our Playground, where you can explore practical examples of how to utilize Davis AI Analyzer in your anomaly detection.

See examples of using Davis AI to detect anomalies. Visit Dynatrace Playground.

The post Learn how to create a Davis AI anomaly detector on Grail appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/create-a-davis-ai-anomaly-detector-on-grail/feed/ 0
Generative AI poised to have impact by automating software development, report says https://www.dynatrace.com/news/blog/generative-ai-poised-to-have-impact/ https://www.dynatrace.com/news/blog/generative-ai-poised-to-have-impact/#respond Mon, 22 Apr 2024 16:07:36 +0000 https://www.dynatrace.com/news/?p=63739 Generative AI poised to have an impact by automating software development. And why AI projects fail

According to recent research from TechTarget’s Enterprise Strategy Group (ESG), generative AI will change software development activities, from quality assurance to debugging to CI/CD pipeline configuration. Many organizations are turning to generative artificial intelligence and automation to free developers from manual, mundane tasks to focus on more business-critical initiatives and innovation projects. Therefore, it’s no […]

The post Generative AI poised to have impact by automating software development, report says appeared first on Dynatrace news.

]]>
Generative AI poised to have an impact by automating software development. And why AI projects fail

According to recent research from TechTarget’s Enterprise Strategy Group (ESG), generative AI will change software development activities, from quality assurance to debugging to CI/CD pipeline configuration.

Many organizations are turning to generative artificial intelligence and automation to free developers from manual, mundane tasks to focus on more business-critical initiatives and innovation projects. Therefore, it’s no surprise that generative AI is poised to have a massive impact by automating software development tasks today and in the near term, according to ESG’s data.

In the research, “Code Transformed: Tracking the Impact of Generative AI on Application Development,” sponsored by Dynatrace, findings indicate that AI and automation are already having a major impact on how developers are working today.

Weighing the pros and cons of automating software development

AI-enabled development can eliminate manual effort and free developers’ time to engage in more strategic, high-level code development. Software development tasks include testing and quality assurance (QA), security, coding, debugging, CI/CD pipeline configuration, and documentation.

On the whole, survey respondents view AI as a way to accelerate software development and to improve software quality. According to the survey, 79% of respondents say AI is already helping to reduce time spent on manual tasks.

At the same time, 75% of respondents say it has taken longer than expected to derive value from AI initiatives related to automating CI/CD pipelines.

What are continuous integration and continuous delivery?

Continuous integration (CI) is a software development practice that streamlines the process of creating software within an organization.

Continuous delivery (CD) enables DevOps teams to develop and deliver complete portions of software to repositories in short, controlled cycles.

How AI is reshaping application development

The ESG report explains how three types of AI are reshaping the app development ecosystem:

  • Generative AI leverages large language AI models to create new outputs. These help teams with data augmentation, anomaly detection, simulation, and documentation, among other areas.
  • Predictive AI uses data collection, algorithm assignment, and model training for user behavior prediction, demand forecasting, fraud detection, and quality control, among other areas.
  • Causal AI models the cause-and-effect relationship between variables to help with areas that include personalization, testing, optimization, policy impact assessment, and more.

AI influences QA and container orchestration

Organizations are also using AI for myriad testing and QA activities, including error detection and debugging (44%), among other tasks.

Generative AI is also becoming key to container orchestration—a process that automates the deployment and management of containerized applications and services at scale.

Organizations use generative AI for myriad use cases involving container orchestration, according to the research, such as automated remediation (35%).

Reaping the rewards of generative AI and automation

Ultimately, the report indicates that IT operations (57%) and product development (42%) stand to benefit most from generative AI.

The trend in automating software development tasks, therefore, stands to benefit the stewards of IT systems and product innovation—two central locations of organizational growth and risk mitigation—today and in the future.

Source: Enterprise Strategy Group, a division of TechTarget, Inc. Research Report, Code Transformed: Tracking the Impact of Generative AI on Application Development, February 2024.

The post Generative AI poised to have impact by automating software development, report says appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/generative-ai-poised-to-have-impact/feed/ 0
Efficient SLO event integration powers successful AIOps https://www.dynatrace.com/news/blog/efficient-slo-event-integration-powers-successful-aiops/ https://www.dynatrace.com/news/blog/efficient-slo-event-integration-powers-successful-aiops/#respond Fri, 05 Apr 2024 22:05:41 +0000 https://www.dynatrace.com/news/?p=63540 SLO metrics graphic

Across all industries—from health care and finance to retail and professional services—Service Level Objectives (SLOs) help teams work toward a common goal.

The post Efficient SLO event integration powers successful AIOps appeared first on Dynatrace news.

]]>
SLO metrics graphic

This blog post is for both novice and seasoned audiences alike. If you’re interested in the relevance and utility of SLOs and how they might be helpful for you, you’ve come to the right place.

The first part of this blog post briefly explores the integration of SLO events with AI. The second part delves into the approach to be adopted with SLOs based on a client’s context and environment, considering whether SLAs (Service Level Agreements) have already been historically defined. Finally, this post addresses the practical aspects, focusing on cultivating the right leading indicator instincts for problem detection through SLOs and extracting their full value to meet business and functional requirements.

Strategic approach to root cause detection

Two frequently used SLOs are the Apdex score and failure rate. For a more proactive approach and to gain further visibility, other SLOs focusing on performance can be implemented.

Every problem identified through Dynatrace Davis® AI indicates an issue with potential user impact. However, understanding the precise impact on the end user can be challenging. For instance, consider how fine-tuning failure detection can provide insights for comprehensive understanding.

Davis AI uses Dynatrace Smartscape® topology to map out the application chain. As a result, the AI operates on the related events and detection parameters (threshold, period, analysis interval, frequent detection, and so on) an issue represents. Using an SLO, you can target the desired entity and map the incoming and outgoing interactions against it.

For example, envision an apple tree that drops an apple. Can the AI discern exactly where and when the apple dropped? And what caused the apple to fall? Was it the wind, a natural progression, a leaping cat, or something else entirely? Likewise, what is the result of the apple falling? Did it split open on the ground or dent a car?

These conditions are pivotal for pinpointing the root cause and facilitating an expansion of the AI’s detection capabilities by incorporating statistical and forecasting models. Consequently, we recognize the value in augmenting AI capabilities through a Dynatrace generative AI model that incorporates user feedback, enhancing its functional dimension.

Further in this article, we will explore which empirical approach to adopt in order to best align with the service-level objectives (SLOs). Specifically, we can initially divide the approach into two parts. We either know the entities to apply the SLO, or we don’t.

Which approach should you take?

SLO methodology is based on empirical approaches with some catalyzers. The approach we will adopt for implementing SLOs follows the diagram illustrated below. Either you have no immediate functional or business visibility (path 1), and you will implement your initial SLOs, or you have business assistance (path 2) and can strategically position your SLOs in key areas.

There are two possible scenarios: You either have defined SLAs, or you do not.

SLO methodology approach diagram

This approach can be accurately validated through empirical means. From my experience, a month of monitoring is the optimal duration to gain statistically significant insights into “how my entity behaves with the configured SLO.” Consequently, by understanding the intricacies involved, this enables us to swiftly confirm or refute the functional/business case of the business teams involved, including project managers, top management, developers, and tech leads.

Path 1: SRE is siloed and lacks familiarity with the instrumented application.

The suggested approach involves utilizing Frontend SLOs to swiftly draw attention to business-related issues. This method generates interest and prompt action by promptly highlighting pertinent business concerns.

Next, a pragmatic approach involves examining the backend, focusing on Service type entities prominently exposed to the frontend (for example, Apache Tomcat in a Linux environment). Additionally, meaningful functional names should be considered, and the number of calls (throughput) should be analyzed for applying Service Level Objectives (SLOs).

SLO in Smartscape view

Path 2 (easier): I’m familiar with my application and the development team.

The recommendation emphasizes promptly adopting Backend Service Level Objectives (SLOs). In today’s landscape, we lack a clear understanding of properly creating frontend SLOs (for example, RUM application type entities) based on key user actions. Often, businesses find themselves in disagreement or simply lost. The most user-friendly and effective approach to pinpointing weaknesses involves establishing a “Service SLO” (dt.entity.service, the “backend”). Given that the implementation of backend SLOs is generally better understood by the business/development and allows for a straightforward and rapid placement of SLOs on services, controllers, or APIs that are generally comprehensible to the in-house developers of the client. That said, the guiding thread, predominantly under the client’s control, remains the backend perspective. In other words, where the application code resides.

However, it’s essential to exercise caution: Limit the quantity of SLOs while ensuring they are well-defined and aligned with business and functional objectives. For example, in a specific application type like e-commerce, select significant Service names (such as Basket, CardPayment, or UserController).

When the SLO status converges to an optimal value of 100%, and there’s substantial traffic (calls/min), BurnRate becomes more relevant for anomaly detection.

What characterizes a weak SLO?

Let’s assume we created a service-availability SLO, monitoring the request failure count against the overall request counts. Now, let’s try to explain what the reasons for no AI detection could be related to this SLO event

AI does not correctly detect that two relevant criteria are involved in disadvantaging AI detection:

When the status of the SLO falls below the specified target by a certain threshold (for example, below 90%), it signifies a failure rate event, indicating an average error rate of 10%.

This implies that when the status is unfavorable, implementing sophisticated alerting methods like error budget burn rate alerting presents challenges and is therefore not applicable. Bad status indicates constant violations of the threshold, akin to the state of a broken door that requires fixing before defining the Service Level Objective (SLO).

For instance, if we maintain an average SLO status of 80%, indicating an average error rate of 20%, utilizing the error budget burn rate in this scenario might be challenging. A simple ratio of 2 implies a 40% error rate, a situation that rarely occurs unless there’s a service interruption leading to observed consequences.

Therefore, these considerations outline two types of coupling for the SLO: Rapid adoption of backend SLOs and cautious application of error budget burn rate alerting, aligning with the traffic and error rate observations for effective anomaly detection.

When there isn’t enough traffic (requests/min) for an SLO

Detection becomes sporadic, resulting in insufficient data points for establishing the sampling interval necessary for generating “Event duration.”

This necessitates a subsequent description of how we should apply default transformation in this specific case, only if this SLO is deemed worthy of tracking.

SLO tracking in Dynatrace screenshot

Strong SLOs

With a view to becoming highly sensitive to detection and adopting a proactive approach, we need to fulfill this condition. If we have constant traffic (meaning sufficient data points to feed the entity) and an averaged baseline, thus exhibiting a regular trend, then it would be easier to detect significant deviations from the behavior of the entity. Therefore, configuring SLOs for this type of entity is considered “strong”, and the error budget burn rate makes sense in detecting these deviations. This proactive stance allows us to maximize our focus on the actual impact experienced by the end user in real time. See the following example with BurnRate formula for Failure rate event.

Error budget burn rate = Error Rate / (1 – Target)

Error budget burn rate monitoring in Dynatrace screenshot

Best practices in SLO configuration

To detect if an entity is a good candidate for strong SLO, test your SLO. SLOs must be evaluated at 100%, even when there is currently no traffic.

If the targeted entity is validated as relevant for the SLO and occasionally experiences a lack of traffic, then using the “default” transformation in the SLO expression is advisable to prevent misleading SLO status!

Use the default transformation.

This example shows the SLI of a performance SLO, targeting response times/loading times of a key-user-action/DOMload.

((
(builtin:apps.web.action.domInteractive.load.browser
:splitBy("dt.entity.application_method")
:avg
:filter(and(or(in("dt.entity.application_method", entitySelector("type(~"APPLICATION_METHOD~"),entityId(~"APPLICATION_METHOD-XXX~")")))))
:partition("perf",value("good",lt(100))) 
:splitBy()
:count
:default(0))
/
(builtin:apps.web.action.domInteractive.load.browser
:splitBy("dt.entity.application_method")
:auto
:filter(and(or(in("dt.entity.application_method", entitySelector("type(~"APPLICATION_METHOD~"),entityId(~"APPLICATION_METHOD-XXX ~")")))))
:splitBy("dt.entity.application_method")
:count
:default(1)
))
:default(1)
*
(100))

The outer default(1) is employed in this context to specify how any data should be handled. This assumes that if there’s an absence of data, the Service Level Objective (SLO) must be assessed at 100%. This approach prevents the truncation of the SLO health state in cases where no data is received, ensuring that the SLO status remains intact even without data.

  • Data Explorer “test your Metric Expression” for info result coming from the above metric.
    Following the previous metric (above) used for the SLO, the threshold employed is an average of 100 ms for the Key Performance Indicator (KPI) of DOM Interactive.
    DOM Interactive KPI in Dynatrace screenshot
    If this threshold is exceeded during THE test within Data Explorer, I will have the tendency or a preemptive glimpse of the impending alert. This allows me to anticipate the threshold and future adjustments to the alert taken by the SLO. In other words, the peaks here indicate that the condition of the SLO is met, resulting in a status of 100%, meaning we adhere to the set threshold. Outside of these peaks, the threshold is violated! On the other hand, if the threshold is violated, it decreases the error budget.Spikes indicate that the conditions for meeting the Service Level Objective (SLO) are fulfilled. However, the gap and space between these impactful SLO statuses, on the other hand, contribute to an increase in the error budget.
  • Service type (General Parameters Exceptions / mute request)

To maintain SLO integrity, we employ various strategies, such as muting requests, ignoring certain elements, or enforcing exceptions only when a problem is identified. This approach ensures that SLO degradation is prevented unnecessarily, with the aim of preserving performance until a patch release can be implemented.

As mentioned earlier, it is necessary to have a minimum observation period to determine the SLO behavior of a targeted entity. It is reiterated that the positioning of this SLO includes the selection of strategic factors (entity positioning within the application architecture = ServiceFlow, sufficient traffic in requests per minute, etc).

During this observation phase, business stakeholders will quickly validate this case and thus define a waiting period for the implementation of the application fix. It is understood that during this waiting period, we will mute the actual exception/error impacting the health of the SLO.

Caution is advised, as by doing this, it should not come as a surprise to observe a better health state (score) of the SLO. This is normal because we have bypassed a parasitic part constituted by the exceptions and errors that actually cause user impact.

Find below the way to detect potential user-impacted cases thanks to SLO.

Error burn rate events in Dynatrace screenshot
Open problem with entity included into SLO calculation in Dynatrace screenshot
Root cause found by AI thanks to an SLO in Dynatrace screenshot

Details of the root cause

Details of a root cause in Dynatrace screenshot

The developer deems it appropriate to either exclude or designate this error as acceptable during the patch release to prevent being overwhelmed with false positive alerts. Most importantly, this should not impede the health status of the SLO, as this is a recognized issue.

Payment Controller settings in Dynatrace screenshot

  • The metric selector must be split by service
(100)*(builtin:service.errors.server.successCount:splitBy("dt.entity.service"))/(builtin:service.requestCount.server:splitBy())

Splitting with the appropriate entity dimension enables the AI to more accurately pinpoint the created Service Level Objective (SLO). Consequently, when triggering an alert, this approach facilitates drilling down into the affected entity, thereby enabling a deeper understanding of the root cause behind the issue.

  • Smart alerting approach: BurnRate/ Status and AlertingProfile

E-Commerce Use Case

In this instance, upon reviewing the week’s alerts, it’s evident that parasitic noise has significantly decreased due to implementing good practices. There were 441 issues without intelligent configuration, whereas with intelligent configuration, only 29 were identified as potentially impacting the end user.

Problem overview in Dynatrace screenshot

Error burn rate problems in Dynatrace screenshot

The kinetics of the impact on the Service Level Objective (SLO) are represented by the red arrow, illustrating the concept of error budget burn rate, which indicates how rapidly the error rate escalates.

Concept of error budget burn rate

Mean time to recovery (MTTR) diagram

Based on the IT incident indicators, our MTTD is 3 min for this event below, degrading the SLO.

Service problems found in Dynatrace screenshot

For more info on how to configure and use BurnRate, see SLO monitoring alerting on SLOs error budget burn rates.

Conclusion

An effective Service Level Objective (SLO) holds more value than numerous alerts, reducing unnecessary noise in monitoring systems. The crucial final step is to highlight strong signals that genuinely impact the end user. Validating and integrating this approach into an intelligent ITSM tool like ServiceNow, for example, will optimize service management by aligning IT services with business needs, ensuring efficient delivery and improved user experiences.

Interested in learning more? Contact us for a free demo.

The post Efficient SLO event integration powers successful AIOps appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/efficient-slo-event-integration-powers-successful-aiops/feed/ 0