Wolfgang Beer | Dynatrace news https://www.dynatrace.com/news/blog/author/wolfgang-beer/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Tue, 23 Jun 2026 07:00:25 +0000 en hourly 1 Build trust with Dynatrace AI-driven root cause and impact analysis https://www.dynatrace.com/news/blog/build-trust-with-dynatrace-ai-driven-root-cause-and-impact-analysis/ https://www.dynatrace.com/news/blog/build-trust-with-dynatrace-ai-driven-root-cause-and-impact-analysis/#respond Tue, 13 Jan 2026 16:14:34 +0000 https://www.dynatrace.com/news/?p=72406 Problem alert dashboard

For 13 years, Dynatrace® AI has successfully delivered fully automated root cause and impact analysis to hundreds of thousands of operations team members worldwide. Dynatrace root cause analysis is the most mature and fully automated solution available for analyzing application, service, and cloud infrastructure-related incidents, having proved its value in large-scale environments over many years.

The post Build trust with Dynatrace AI-driven root cause and impact analysis appeared first on Dynatrace news.

]]>
Problem alert dashboard

Automating the analysis of petabytes of traces, metrics, and logs during critical incidents delivers high value to operations teams and helps significantly speed up Mean Time to Repair.

Build trust in AI by surfacing all the logical facts

What’s equally important is working with a highly transparent data platform and leveraging the unique context that Dynatrace Smartscape® topology delivers, so that every single AI reasoning step can be reliably explained and proven. The transparent surfacing of all relevant information and necessary facts, such as the real-time, trace, and topology-induced impact tree, helps you learn about the facts and gain trust in agentic AI when making critical decisions in stressful incident situations.

The new Visual Resolution Path of each detected problem is based on deterministic logic, rather than probabilistic and less reliable models, such as potentially hallucinated results from generative AI. The Visual Resolution Path is closely tied to your configured alerts and the underlying Smartscape topology.

The new incident summary, along with the Visual Resolution Path, surfaces metrics and timings for all affected frontend and downstream service dependencies. Additionally, the incident summary now displays new metrics, including affected business flows, which link the incident to critical business processes.

The automatically derived root cause is displayed along with the problem details and an impact graph, as well as a summary of what happened and what caused the cascade of events.

Figure 1. The Visual Resolution Path visually traces the service flow and dependencies of the root cause analysis.
Figure 1. The Visual Resolution Path visually traces the service flow and dependencies of the root cause analysis.

View all actions in context

Each finding includes context-aware follow-up actions based on the incident you’re viewing.

The same context-aware actions are displayed in both the incident summary and the full Smartscape view.

Each Smartscape node displays context-based actions, saving you valuable time when your teams are frantically searching for incident logs or need to view failing traces.

Figure 2. Opinionated drill-downs speed up problem resolution
Figure 2. Opinionated drill-downs speed up problem resolution

Review incident timing

Reviewing incident timing and configured alerts for individual Smartscape nodes helps teams understand their AI’s reasoning and fosters trust throughout the incident response process.

Figure 3. Understand AI reasoning and foster trust throughout incident response.
Figure 3. Understand AI reasoning and foster trust throughout incident response.

Transparent alert notifications and remediation automation

To manage your incident, you typically need to know what happened, what caused the issue, and what was impacted. It’s also crucial to understand immediately which automated measures were taken to mitigate the incident.

The process might look like this: a set of simple alert notifications is sent to your team’s Slack channels, triggered by the detection of the problem. This is followed by an automated ServiceNow ticket workflow, which initiates the automated remediation workflow to mitigate the detected root cause.

The Dynatrace Automation Engine can also orchestrate calls to external AI agents, allowing operations teams to engage AWS or Azure SRE agents to identify potential cloud resource misconfigurations. In some cases, these agents can fix them instantly.

To maintain oversight of what is triggered and what your automation measures are doing, the Problems app provides a clear view of remediation actions, along with their timing and the outcome of each action.

This workflow execution information is stored in the Grail® data lakehouse, allowing teams to build dashboards and notebooks from it. The Problems app surfaces this information exactly where your team expects to find it during incident response.

Figure 4. The Problems app provides a clear view of remediation actions, their timing, and whether they were successful or unsuccessful, enabling you to track your automations.
Figure 4. The Problems app provides a clear view of remediation actions, their timing, and whether they were successful or unsuccessful, enabling you to track your automations.

Customize the layout to suit your needs

The new problem overview helps you focus on what matters most, whether that’s the root cause, impact, graph, or automation actions taken. You can rearrange the layout to match your priorities, and Dynatrace will remember your setup as a per-user preference for future incidents.

Have confidence in reported root causes when every minute counts

Reacting quickly in critical situations, such as large-scale outages, is crucial for operations teams. With AI-driven, fully automated analysis of all your incoming traces, logs, and alerts, teams can make better-informed decisions about root cause and impact when time matters most. By exposing the context and decision logic behind AI root cause analysis, Dynatrace adds the transparency needed to build trust and confidently guide remediation.

The new Dynatrace problem root cause and impact overview, along with the Visual Resolution Path, helps ensure that your AI follows the logically correct dependencies and builds trust in your root cause analysis.

Speed up incident response and get the full context of each incident’s root cause.

The post Build trust with Dynatrace AI-driven root cause and impact analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/build-trust-with-dynatrace-ai-driven-root-cause-and-impact-analysis/feed/ 0
Remediation intelligence: Accelerate MTTR with AI-powered context and knowledge https://www.dynatrace.com/news/blog/remediation-intelligence-accelerate-mttr-with-ai-powered-context-and-knowledge/ https://www.dynatrace.com/news/blog/remediation-intelligence-accelerate-mttr-with-ai-powered-context-and-knowledge/#respond Wed, 13 Aug 2025 15:38:30 +0000 https://www.dynatrace.com/news/?p=70391 Remediation intelligence

There’s a hidden barrier to deep understanding of your revenue, performance, and bottom line that has nothing to do with tools or telemetry. It’s organizational knowledge, the remediation know-how that’s scattered across hundreds of documents, private notebooks, dashboards, and in the minds of your engineers. This implicit, loosely documented knowledge is invisible to machines and rarely timely for humans. So, when a high-priority incident hits, this knowledge gap fills your war rooms with engineers tasked with solving problems they don’t own.

The post Remediation intelligence: Accelerate MTTR with AI-powered context and knowledge appeared first on Dynatrace news.

]]>
Remediation intelligence

Delivering reliable, business-critical applications to production is more complex than ever. The growing complexity and granularity of modern software systems and the trend to shifting more and more responsibilities to development teams (shift left) lead to increased pressure on your development teams.

In fact, traditional development teams now have a wider set of responsibilities; they must be highly skilled and educated in a multitude of domains. The responsibilities of these teams now range from specification, planning, testing, risk assessment, cost estimates, and deployment, to load testing, UI testing, integration testing, and on-call responsibilities for their software services.

While development teams need to be literate in all new technology stacks, cloud resources, and quality assurance methods, they also need to work with numerous tools.

The invisible bottleneck in your remediation process

The whole shift-left trend has pushed operational responsibility closer to development, expanding workloads to include on-call rotations, more frequent deployments, and growing expectations around uptime. When incidents occur, dozens of engineers are dragged into war rooms to perform analysis of the underlying root causes.

Reducing the Mean Time to Repair (MTTR) is essential to business continuity and customer satisfaction. Without access to the right knowledge at the right time, even skilled teams lose momentum. In high-pressure situations caused by critical incidents, it becomes more important than ever to ensure that all relevant information is shared with every role involved. This requires the most automated and intelligent methods available for effectively collecting, analyzing, and distributing information.

The on-call engineer’s journey

Take Omar, an SRE; an automated voice jolts him awake to summon him into a war room in the middle of the night after a routine update caused a spike in failed requests for a cloud-based payment service. It’s a P1 incident. With each passing minute, merchants are losing value, support requests surge, and customers are complaining. Dozens of caffeinated engineers are already in the war room.

Logs point to timeouts, but this is just a symptom; the real problem is somewhere else. Reading every message and document would take hours, so Omar scans for summaries, key findings, and any mention of his team’s services. Several hypotheses have already been tested, and one points to a potential issue in a backend service Omar’s team owns. Meanwhile, customer complaints are beginning to surface from other time zones. The payment service is failing, and customer success managers are growing increasingly anxious.

And while a similar outage has happened before, Omar is not able to find any documentation or insights into how to remediate the issue.

Why organizational knowledge doesn’t scale

When remediation history lives in documents, scattered across teams, formats, and platforms, engineers waste time searching instead of solving. Even well-documented incidents don’t prevent recurrence if they’re disconnected from future incidents. Even centralized platforms like Backstage don’t help if they can’t surface the right guidance at the right time.

Without a way to systematically identify, reuse, and scale this knowledge, it remains reactive. That’s not just inefficient, it’s a blocker to building intelligent automation and truly preventative operations. This is the hidden obstacle, silently inflating your MTTR, buried knowledge that costs time, delays response, and drains focus from what really matters. This is what remediation intelligence solves.

Introducing remediation intelligence

Dynatrace has a long history of providing DevOps teams with AI-driven tools for anomaly detection, root cause identification, and incident impact assessment in complex application environments. Over the past decade, it has contributed to reducing mean time to resolution (MTTR) by learning application behavior and analyzing dependencies in real time.

Building on this foundation, Dynatrace launched remediation intelligence, which adds an additional element to the incident response process from alert through resolution. It assists engineers during remediation by combining Davis® AI root cause and impact analysis with input from global community knowledge and internal expertise. It integrates data such as logs, metrics, traces, and topological context into a single view and offers support for documenting post-incident reviews.

Figure 1. The problems page displays all important information, allowing you to directly access all incident-relevant error logs.
Figure 1. The problems page displays all important information, allowing you to directly access all incident-relevant error logs.

Close the knowledge gap:  Embedded troubleshooting knowledge

What truly sets Dynatrace remediation intelligence apart is its ability to proactively surface relevant internal knowledge at the moment it’s needed most. It adds an AI-guided assistive layer to the Problems app that brings implicit, organizational knowledge directly into the flow of incident response. Once a problem is detected, Davis AI scans the historical data, surfacing past remediation playbooks, troubleshooting dashboards, and notebooks that were used to resolve similar issues.

With troubleshooting guides, we introduce a context-aware guidance system, built on Davis AI, that connects current incidents with prior resolution paths. It makes organizational knowledge queryable, remediation patterns reusable, and every responder effective, even when they’re solving an unfamiliar issue.

Figure 2. Review related documents from similar past incidents.
Figure 2. Review related documents from similar past incidents.

Remediation intelligence surfaces the most relevant remediation insights

When an incident occurs, Dynatrace excels at automatically analyzing and surfacing technical insights. It collects and organizes all relevant signals—logs, metrics, traces, and topology—into a single coherent problem. No fragmented alerts. No disconnected symptoms. Just one structured, AI-curated incident view. In parallel, Dynatrace AI scans all documents marked as troubleshooting-relevant. Using advanced semantic search and vector embeddings, it ranks and surfaces the most relevant past incidents, dashboards, notes, and postmortems, based on their similarity to the current problem. This is not just keyword matching; Dynatrace understands patterns, failure modes, and system relationships, surfacing ranked, high-similarity incidents in the problem view.

Figure 3. Example of a troubleshooting guide
Figure 3. Example of a troubleshooting guide

From observability to trusted automation: The power of context-aware AI

The future of resilient, self-healing systems lies in the seamless integration of observability, AI, and organizational knowledge. When these elements come together, they form the foundation for trusted automation—a system that not only reacts to incidents but learns from them, adapts, and eventually prevents them altogether.

At the heart of this vision is context. Effective auto-remediation depends on the ability to precisely identify the root cause of an issue and understand its broader impact across the application stack. But automation doesn’t stop at detection. By capturing and integrating the remediation strategies used by engineers, Dynatrace builds a living knowledge base. This organizational knowledge, when combined with AI-driven root cause analysis, allows the system to replicate proven remediation paths and suggest next steps with increasing accuracy.

With all relevant data and insights unified in a single platform, engineers gain a single pane of glass view into their systems. This not only streamlines manual remediation efforts but also lays the groundwork for flexible, context-aware auto-remediation. Each incident you resolve fuels the knowledge. Over time, as the system learns, it evolves from reactive automation to proactive incident prevention—anticipating issues before they escalate and taking preemptive action.

This is the vision Dynatrace is delivering: a future where engineers can trust automation not just to respond, but to understand, learn, and improve—turning every incident into a step toward greater system intelligence and reliability.

Empower your teams: turn hard-won operational insights into scalable remediation power

Dynatrace remediation intelligence is ready to work for you today. To start benefiting, opt into Davis CoPilot®, the Dynatrace generative AI assistant. Once turned on, you’ll need to configure Davis CoPilot to learn from a curated set of Dynatrace documents—specifically, Notebooks and Dashboards that are either created directly from detected problems or clearly labeled with the prefix [TSG] in their titles (short for Troubleshooting Guide).

Davis CoPilot will analyze and learn from your team’s historical remediation efforts, capturing valuable insights and strategies, allowing it to proactively suggest relevant documentation and guidance when similar incidents are detected in the future, helping your engineers respond faster and more effectively.

Figure 4. Turn on Davis CoPilot in Settings.
Figure 4. Turn on Davis CoPilot in Settings.

All data uploaded to Dynatrace remains strictly private. All data and remediation insights remain within your tenant. Davis CoPilot treats your documents as strictly confidential and never shares or transfers this information outside your Dynatrace environment. Your team’s knowledge stays private, secure, and entirely under your control, while still powering smarter, more context-aware automation.

Learn more about document suggestions and Dynatrace remediation intelligence in our documentation, and read about discovering relevant troubleshooting guides and how to create new ones.

Don’t let organizational knowledge stay buried. Make it actionable. Make it scalable.

The post Remediation intelligence: Accelerate MTTR with AI-powered context and knowledge appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/remediation-intelligence-accelerate-mttr-with-ai-powered-context-and-knowledge/feed/ 0
Powerful exploratory analytics for AI-driven insights https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/ https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/#respond Tue, 04 Feb 2025 16:00:42 +0000 https://www.dynatrace.com/news/?p=67543 Problem alert dashboard

The Dynatrace platform empowers Operations, SRE, and DevOps teams to maintain high software quality, security, and reliability, allowing organizations to innovate and scale confidently. By leveraging Davis® AI with enhanced predictive analytics and automated workflows, Dynatrace simplifies issue detection and resolution, reduces MTTR, and enables proactive incident prevention.

The post Powerful exploratory analytics for AI-driven insights appeared first on Dynatrace news.

]]>
Problem alert dashboard


Deploying and safeguarding software services has become increasingly complex despite numerous innovations, such as containers, Kubernetes, and platform engineering. Recent global IT outages, such as the CrowdStrike incident, remind us how dependent society is on software that works perfectly.

Organizations must balance many factors to stay competitive.
Figure 1. Organizations must balance many factors to stay competitive.

Organizations strive to strike a delicate balance between cost, time to market, and innovation. This challenge is more pressing than ever as businesses seek to stay competitive while ensuring their software remains robust and secure.

This necessitates a comprehensive platform that empowers enterprises to understand IT and software within the broader context of their business operations, giving them confidence that their software and IT infrastructure are reliable.

Scale with confidence: Leverage AI for instant insights and preventive operations

Using Dynatrace, Operations, SRE, and DevOps teams can scale efficiently while maintaining software quality and ensuring security and reliability. Its AI-driven exploratory analytics help organizations navigate modern software deployment complexities, quickly identify issues before they arise, shorten remediation journeys, and enable preventive operations.

We’ve added numerous enhancements to our platform, leveraging advanced AI and automation for smarter software observability.

In this blog post, we show you how to

  • Get AI-driven insights directly on your operations dashboards
  • Improve MTTR with AI-assisted problem analysis and logs and traces in context
  • Leverage Gen AI through Davis CoPilot to get insights into root causes
  • Automate remediation of AI-detected problems with simple workflows
  • Adopt Preventive Operations with AI forecasting and automated action

Get AI-driven insights directly on your operations dashboards

A high-level, customizable view of your data is crucial in modern software operations. Dynatrace Dashboards, powered by Grail™ data lakehouse and Davis® AI, offer precisely that. They provide a comprehensive overview, seamlessly integrating health and problem-related information into a single view. You can chart your topology across data silos alongside all alerts, events, and problems using honeycomb tiles, which offer convenient drill-downs into the problem-debugging user flow.

Dynatrace ensures that context is seamlessly integrated into the platform, thus simplifying complexity for you as a user when analyzing issues and allowing you to focus on what truly matters. AI-driven analytics transform data analysis, making it faster and easier to uncover insights and act. This approach not only improves user experiences, it ensures that critical insights are accessible to both experts and novices. By simplifying remediation journeys and extending features to more user groups, Dynatrace enables results across all teams.

The new Problems dashboard, including rich honeycomb visualization, helps you focus on what’s important, turning technical data into a visual story.
Figure 2. The new Problems dashboard, including rich honeycomb visualization, helps you focus on what’s important, turning technical data into a visual story.

When a truly important issue stands out, the next step is refinement. With a few clicks, you can segment and filter your data to focus on specific applications, assignment groups, or regions. Directly mapping and surfacing ownership information within data segments accelerates incident assignment notifications and triggers automatic remediations.

Utilize the comprehensive filter functionality to update your dashboards dynamically.
Figure 3. Utilize the comprehensive filter functionality to update your dashboards dynamically.

If you see an issue or need to look closely at a specific application where an issue was identified, simply select the element to be seamlessly directed to the Problems app. There, you can dig deeper while continuing to focus on your selected segment. This tight integration, following a golden thread of insights, ensures that you’re more productive. To experience the possibilities of AI-empowered dashboards, try our example dashboard on the Dynatrace Playground.

Improve MTTR with AI-assisted problem analysis, logs, and traces in context

The Problems app delivers opinionated AI-assisted problem analysis optimized for Operations and Site Reliability Engineers (SREs) and developers. According to IDC, guiding users visually and automatically surfacing all critical details enables a 56% faster mean time to repair (MTTR) for critical incidents.

When a large-scale incident occurs, follow the red flag that Davis AI uses to identify the root cause, pinpoint all relevant details, and visually reproduce the details in charts, highlighting the affected deployment.

Analyze the root cause in the Problems app.
Figure 4. Analyze the root cause in the Problems app.

Besides identifying the root cause, Davis AI also automatically connects all relevant log lines. Logs are invaluable for identifying further insights and detecting fundamental flaws, such as process crashes or exceptions. With a single click in Problems, all incident logs are surfaced automatically. But we don’t stop there, Dynatrace also seamlessly integrates relevant trace data, offering full visibility into even complex, microservices-based architectures.

By providing these end-to-end insights, Dynatrace and Davis AI empower SREs, developers, and architects to quickly dive deep into an incident’s details, including all relevant logs and traces. Using this context, they can effectively focus on fixing and remediating code-level issues, significantly improving MTTR, and ensuring that critical incidents are resolved swiftly and efficiently.

Leverage GenAI via Davis CoPilot for insights into root causes

Dynatrace offers precision tools for domain experts to solve complex problems and dig deeper into their data. While product owners often focus on the intricate technical details of an incident, they often prefer a quick summary of what happened and what caused it. The soon-to-be-globally available Davis CoPilot™ bridges this gap by summarizing problems and their root causes and suggesting remediation steps based on these insights.

You’re not limited to one problem; Davis CoPilot can simultaneously analyze multiple problems, draw conclusions about their relationships, identify the common root cause, and propose corrective steps. Instead of relying on a team of experts and waiting hours for insights, Davis CoPilot helps you identify similarities and draw relevant conclusions independently and efficiently.

The use of generative AI adds significant value by augmenting Dynatrace-detected technical root causes with knowledge from the global tech community. Generative AI can access and synthesize vast amounts of information from various sources, providing a broader context and deeper insights. This ensures that your teams benefit from the latest advancements and solutions, enhancing their ability to resolve issues effectively and efficiently.


Dynatrace Problems App - Explain Problems video

Gain a better understanding of root causes with Davis CoPilot
Figure 5. Gain a better understanding of root causes with Davis CoPilot

Automate remediation of AI-detected problems with simple workflows

To automatically remediate Davis AI-detected problems, Dynatrace leverages powerful Workflows. Dynatrace workflows can be triggered by any problem or alerting event, automating domain-specific tasks to take remedial actions.

For example, workflows can scale up capacity to adapt to demand or automatically restart a service in case of a crash. With a large catalog of available workflow actions, you can react efficiently to AI-detected problems, reducing mean time to repair (MTTR) by automatically remediating issues.

But you can do much more with it: The recently introduced Simple Workflows, which are included in your Dynatrace subscription with no extra cost, offer greater flexibility and power than standard notifications. You can use the same mechanisms and trigger types to notify your developer team via Slack, create a JIRA issue, or send a PagerDuty alert.

This ensures that your operations, SRE, and DevOps teams can focus on more strategic tasks while the system handles routine problem resolutions. Automation enhances operational efficiency and ensures that your systems remain robust and reliable, even in the face of unexpected issues.

Easily set up automated remediation with the new Simple Workflows.
Figure 6. Easily set up automated remediation with the new Simple Workflows.

Adopt Preventive Operations with AI forecasting and automated action

Going beyond reactive problem detection, analysis, and remediation, Dynatrace can also leverage predictive AI to anticipate and avoid critical situations before they occur. Using Davis AI forecast, you can easily predict future capacity demands. Combining this knowledge with workflows allows you to take proactive measures to ensure system stability and performance.

Let’s have a look at a concrete example:

It’s easy to predict key indicators of your application, such as order levels or service request counts. Once load and demand rise and Davis AI identifies a potential future issue in your infrastructure setup, Davis CoPilot can automatically generate an updated Kubernetes configuration script for you and automatically upscale the environment to meet future demand. This ensures that your system scales appropriately to handle the anticipated demand, preventing incidents before they occur and eliminating the need to generate a problem.

That’s what we call Preventive Operations. Instead of sending an alert and notifying people, Dynatrace simply fixes the issue. According to Gartner’s Analytics Maturity Model, using predictive AI can significantly reduce the likelihood of incidents by taking preemptive action and remediation.

Start using Davis AI to analyze your environments and predict and address potential issues in advance. This will empower your teams to avoid potential problems and ensure a smooth, uninterrupted user experience.

Initiate automated, corrective action before an issue occurs
Figure 7. Initiate automated, corrective action before an issue occurs.

Tackle business challenges with confidence

Ensure your software runs securely and reliably with Dynatrace and Davis AI.

Dynatrace and Davis AI support you by running your software securely and reliably. This includes advanced root cause analysis, deep insights into detected issues, and corrective actions—whether manual or automatic—to prevent outages before they occur.

Get started

For more information, have a look at our documentation or explore the available resources on the Dynatrace Playground to experience some of these enhancements first-hand:

The post Powerful exploratory analytics for AI-driven insights appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/powerful-exploratory-analytics-for-ai-driven-insights/feed/ 0
Transform your operations with Davis AI root cause analysis https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/ https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/#respond Tue, 08 Oct 2024 19:03:46 +0000 https://www.dynatrace.com/news/?p=66044 root cause analysis

Complexity is ever-increasing in today’s fast-paced world of software deployments and cloud infrastructure. This is why Davis® AI root cause analysis is an indispensable tool for Operations, Site Reliability, and DevOps teams.

The post Transform your operations with Davis AI root cause analysis appeared first on Dynatrace news.

]]>
root cause analysis

Without AI-assisted observability tooling, the productivity of operations teams drops, leading to a dramatic increase in Mean Time to Repair (MTTR) and a significant rise in the personnel needed to manage critical incidents. In an era dominated by automated, code-driven software deployments through Kubernetes and cloud services, human operators simply can’t keep up without intelligent observability and root cause analysis tools.

Modern observability has evolved from simple metric telemetry monitoring to encompass a wide range of data, including logs, traces, events, alerts, and resource attributes. Dynatrace Root Cause Analysis (RCA) seamlessly integrates all this information, providing crucial analysis to remediate incidents in real time.

Problem feed for fast triage and remediation of AI-detected problems.
Figure 1. Problem feed for fast triage and remediation of AI-detected problems.

By offering root cause analysis on top of the highly flexible Grail™ data lakehouse, Dynatrace empowers SRE and operations teams to further reduce MTTR. Direct access to the underlying data allows the automatic RCA analysis to eliminate data silos and to dive deep into every aspect of the collected incident data.

Unlike generic DIY query frontends, the Dynatrace Problems app is a tailor-made solution for efficiently supporting operations use cases. This approach ensures that your operation teams have all the tools they need to manage modern software deployments.

Transform your operations today with the new Problems app and stay ahead in the ever-evolving software and cloud infrastructure landscape.

Rapid response to critical incidents

Operations teams can quickly focus on incoming Davis AI-detected and -analyzed problems by referring to the problems feed.

The problem feed is designed to prioritize active issues, ensuring they always appear at the top, regardless of how long they’ve been ongoing. This default sorting strategy, which uses time as a secondary criterion, guarantees that Operations teams never overlook an active problem, no matter which primary filter is applied.

You can focus on your domain using the filter bar at the top, the quick filters on the side, or both. The chart feature allows for quick analysis of problem peaks at specific times.

Operations teams will appreciate the ability to sort problems by duration and the number of affected entities. This aids in assessing Davis-detected root causes and prioritizing remediation efforts. The native multi-select feature lets users open a filtered group of problems simultaneously, facilitating quick comparisons and detailed analysis.

Streamline deployment insights with AI-generated summaries

Every second counts during wide-scale incidents affecting large parts of your production systems. This is why precisely showing the root cause ultimately helps to speed up problem resolution.

You can multi-select a cohort of active problems, select Show detail, and review all critical problem details, including preview charts and event details, without losing the context of your problem feed.

The new problem experience transparently displays all the available details, with prominently displayed root-cause markers to precisely guide your attention.

In the realm of cloud infrastructure management, having a clear and concise view of your deployment’s health is crucial. Our dedicated deployment perspective offers just that, showcasing the hierarchy of affected and related infrastructure components. The root cause of any issue is prominently marked with a root-cause badge, making it easy to identify and address problems swiftly.

This perspective not only highlights the affected cloud regions but also provides a quick summary of the Kubernetes context where your workloads encountered failures. Gone are the days of clicking and navigating through multiple dashboards. Instead, you receive an AI-generated summary as an affected deployment architecture diagram.

This diagram, akin to a UML (Unified Modeling Language) deployment diagram, offers a familiar representation for software architects, ensuring they can quickly grasp the situation and take necessary actions. By streamlining the visualization of deployment issues, we empower teams to resolve problems more efficiently and maintain optimal performance.

To save time, the root-cause component is preselected, and all the details of the root cause are displayed on the right, along with charts showing the detected breaches from learned normal behavior.

You can review each individual finding on all problem-affected entities by selecting the individual deployment components or by switching to the detailed event perspective, which shows all the single events that the root cause analysis collected into a single problem.

Confirm the AI-detected root cause and review the deployment context.
Figure 2. Confirm the AI-detected root cause and review the deployment context.

In addition to using markers for swift root cause analysis, operations teams often seek to attach valuable remediation hints and playbooks for familiar scenarios.

By implementing a flexible event tagging mechanism, event sources and detectors can be easily customized to include additional custom event properties. This allows for markdown-formatted event description text that can contain remediation links, as illustrated in the screenshot below.

Root cause remediation hints as markdown links
Figure 3: Root cause remediation hints as markdown links

The Dynatrace Semantic Dictionary helps identify the semantics of well-known event properties and provides convenient platform intents. For instance, entity links (dt.entity.*) or links to the responsible settings entry (dt.settings.object_id) that detected and opened an event can be included. These settings links save valuable time when adjusting detection sensitivity for thresholds or baselines. Additionally, the event setting property can be utilized in a DQL query to create a table of the top-triggering configurations or to automate settings changes using an automation workflow.

Quick access to incident logs

The seamless integration of logs powered by Dynatrace Grail™ data lakehouse with Davis AI root cause analysis is a game changer for modern operation teams, as it offers a quick summary of all incident-relevant logs.

The Dynatrace root cause engine already combines all incident-relevant information to recommend log queries, which saves a lot of navigation time and completely eliminates the need to manually identify complex log filters.

A single click on the Problem details log perspective immediately surfaces all relevant logs related to the given incident, as shown below.

Failure rate increase logs
Figure 4.
100 errors and warnings of failure rate logs
Figure 5.

Within this view the Operations team can further refine the query or adapt the filters and open a notebook to persist the log findings for critical post-mortem documentation purposes.

Root cause analysis in a user-focused context

Most modern application stacks are deployed through Kubernetes, making it essential for operations teams to focus on Kubernetes clusters, cloud resources, and workloads of critical services.

Since operations engineers prefer not to switch contexts, a consistent root-cause experience is provided regardless of where the user journey begins.

Whether you start your remediation journey within the Infrastructure & Operations app or the Kubernetes app, you receive the same root-cause information without needing to navigate between different apps. This seamless embedding of root-cause information into the current context saves valuable time during incident remediation.

Root cause shown in context of the Infrastructure & Operations context.
Figure 6. The root cause is shown in the context of Infrastructure & Operations.
CPU throttling root cause shown in Kubernetes context.
Figure 7. CPU throttling root cause shown in Kubernetes context.

Notify and automate to speed up remediation

The Problems app features a global problem indicator that is always visible within the Dock to capture your attention. This indicator shows whether there are active problems within the environment. You can personalize this number by selecting and saving a problem filter within the problem feed, as demonstrated below. The saved default filter is then automatically applied to the global problem indicator, reducing the number of active problems for the user.

Select Alerting (bell icon) to set up alerts related to filtered problems and configure email addresses for notification recipients.

The email payload and the use of an email address for notifications are preset, allowing for a personalized notification setup, as shown below.

Save the personal default filter and set up email notifications.
Figure 8. Save the personal default filter and set up email notifications.
Find the global problem indicator in the Dock.
Figure 9. Find the global problem indicator in the Dock.

You can take a further step towards answer-driven automation and use the detected Davis problem event to trigger workflow automation. Automatically remediate an issue using our no-code workflow actions for collaboration (for example, Slack, Microsoft Teams, ServiceNow, Pagerduty) and remediation (for example, AWS, Red Hat Ansible, Kubernetes).

The introduction of a filterable global problem indicator ensures that Operations teams remain focused on active problems within the environment, even while exploring data in Notebooks or Dashboards.

In future updates, the Problems app will support multiple named filters and introduce Segments as the primary method for using and sharing numerous predefined filters among operations teams.

Outlook

The newly released Problems app enhances transparency by providing detailed AI-detected root-cause information. It also offers convenient deployment and architectural visualizations, along with a log perspective, to help operations teams reduce Mean Time to Repair (MTTR).

In future updates, we aim to support the ability to acknowledge and label incoming problems, improving team coordination. Additionally, plans include a visual representation of the application map, direct propagation of information such as application IDs into the problem feed, and support for segments to filter the problem feed.

Summary

For over a decade, Dynatrace has been at the forefront of integrating AI into incident analysis, particularly through Davis root cause analysis.

Davis is now essential for Operations, Site Reliability, and DevOps teams, helping them to navigate the complexities of modern software deployments and cloud infrastructure.

Without Davis, the productivity of these teams would plummet, leading to longer Mean Time to Repair (MTTR) and increased staffing needs to handle critical incidents.

In today’s automated deployments and cloud services, traditional observability tools fall short, unable to keep pace with the intelligence needed for effective root cause analysis.

Modern observability encompasses various data sources, from metrics to logs and events, requiring intelligent tools like Davis to seamlessly integrate and analyze this information in real time. By providing Davis on top of the flexible Grail data lakehouse, Dynatrace empowers teams to swiftly reduce MTTR by accessing and previewing incident data comprehensively.

The Davis Problems app streamlines triage, allowing teams to swiftly focus on AI-detected issues. Its intuitive interface simplifies problem resolution.

Try out the new Problems app in the Dynatrace Playground.

The post Transform your operations with Davis AI root cause analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/transform-your-operations-with-davis-ai-root-cause-analysis/feed/ 0
Enhanced AI model observability with Dynatrace and Traceloop OpenLLMetry https://www.dynatrace.com/news/blog/enhanced-ai-model-observability-with-dynatrace-and-traceloop-openllmetry/ https://www.dynatrace.com/news/blog/enhanced-ai-model-observability-with-dynatrace-and-traceloop-openllmetry/#respond Mon, 04 Dec 2023 18:32:01 +0000 https://www.dynatrace.com/news/?p=60953 Enhancing AI model observability

In the rapidly evolving landscape of artificial intelligence, ensuring your AI model’s optimal performance, reliability, security, and user trust is paramount. This blog post explores how combining the Dynatrace full stack observability platform and Traceloop's OpenLLMetry OpenTelemetry SDK can seamlessly provide comprehensive insights into Large Language Models (LLMs) in production environments. Observing AI models enables you to make informed decisions, optimize performance, and ensure compliance with emerging AI regulations.

The post Enhanced AI model observability with Dynatrace and Traceloop OpenLLMetry appeared first on Dynatrace news.

]]>
Enhancing AI model observability

“Engineers today lack an easy way to track the tokens and prompt usage of their LLM applications in production. By using OpenLLMetry and Dynatrace, anyone can get complete visibility into their system, including gen-AI parts with 5 minutes of work.”

Nir Gazit, CEO and Co-Founder Traceloop

Why AI model observability matters

The adoption of LLMs has surged across various industries, particularly since the introduction of OpenAI’s GPT model. While these models yield impressive results, the challenge of maintaining their operation within defined boundaries has increased.

AI model observability plays a crucial role in achieving this by addressing these key aspects:

  1. Model performance and reliability: Evaluating the model’s ability to provide accurate and timely responses, ensuring stability, and assessing domain-specific semantic accuracy.
  2. Resource consumption: Observing computational resource availability and saturation, whether deployed in cloud-native environments like Kubernetes or CPU-enabled servers.
  3. Data quality and drift: Monitoring the quality and characteristics of training and runtime data to detect significant changes that might impact model accuracy.
  4. Explainability and interpretability: Providing information on model versions, parameters, and deployment schedules, which is essential for interpreting and understanding model answers.
  5. Security and compliance: Actively preventing security threats at both the application and model levels to ensure responsible and compliant AI usage.

The challenge of AI model observability

One challenge in AI model observability is the diverse tooling landscape required to gain critical insights. OpenTelemetry has become a standard for collecting traces, metrics, and logs. However, seamless support for various SDKs and AI model frameworks, such as LangChain and Pinecone, remains essential.

Combining Dynatrace with Traceloop’s OpenLLMetry addresses the heterogeneity challenge by supporting a range of popular LLMs, prompt engineering, and chaining frameworks. OpenLLMetry, an open source SDK built on OpenTelemetry, offers standardized data collection for AI Model observability.

How OpenLLMetry works

OpenLLMetry supports AI model observability by capturing and normalizing key performance indicators (KPIs) from diverse AI frameworks. Utilizing an additional OpenTelemetry SDK layer, this data seamlessly flows into the Dynatrace environment, offering advanced analytics and a holistic view of the AI deployment stack.

Given the prevalence of Python in AI model development, OpenTelemetry serves as a robust standard for collecting observability data, including traces, metrics, and logs. While OpenTelemetry’s auto-instrumentation provides valuable insights into spans and basic resource attributes, it falls short in capturing specific KPIs crucial for AI models, such as model name, version, prompt and completion tokens, and temperature parameters.

OpenLLMetry bridges this gap by supporting popular AI frameworks like OpenAI, HuggingFace, Pinecone, and LangChain. Standardizing the collection of essential model KPIs through OpenTelemetry ensures comprehensive observability. The open source OpenLLMetry SDK, built atop OpenTelemetry, enables thorough insights into your Large Language Model (LLM) applications.

As the collected data seamlessly integrates with your Dynatrace environment, you can analyze LLM metrics, spans, and logs in the context of all traces and code-level information. Maintained under the Apache 2.0 license by Traceloop, OpenLLMetry is a valuable asset for product owners, providing a transparent view of AI model performance.

The diagram below illustrates how OpenLLMetry captures and transmits AI model KPIs to your Dynatrace environment, empowering your business with unparalleled insights into your AI deployment landscape.

Enhancing AI model observability

Dynatrace OneAgent® is perfectly capable of automatically injecting and tracing code-level information for many technologies, such as Java, .NET, Golang, and NodeJS. However, Python models are trickier.

In the Dynatrace web UI, you can track your AI model in real time, examine its model attributes, and assess the reliability and latency of each specific LangChain task, as demonstrated below.

LangChain task distributed traces in Dynatrace screenshot

The captured span by Traceloop automatically displays vital details, including the mode utilized by our LangChain model gpt-3-5-turbo, the model’s invocation with a temperature parameter of 0.7, and the utilization of 53 completion tokens for this individual request.

LangChain task distributed traces in Dynatrace screenshot

With the growth of AI, maintaining transparency is essential

Observing AI models like Large Language Models (LLMs) in production is crucial for enhancing performance, reliability, security, and user trust. This includes the monitoring of AI-related costs to ensure they remain within acceptable margins. The Dynatrace platform, coupled with Traceloop’s OpenLLMetry OpenTelemetry SDK, offers comprehensive visibility from model inception to completion.

As AI adoption grows, maintaining transparency is essential for regulatory compliance. While Dynatrace automates tracing for various technologies, Python-based AI models require OpenTelemetry. OpenLLMetry bridges this gap, supporting popular AI frameworks and vendors to ensure standardized data collection. OpenLLMetry provides an open source SDK for LLM observability, seamlessly integrating with Dynatrace for in-depth analysis.

References

The post Enhanced AI model observability with Dynatrace and Traceloop OpenLLMetry appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enhanced-ai-model-observability-with-dynatrace-and-traceloop-openllmetry/feed/ 0
Automate predictive capacity management with Davis AI for Workflows https://www.dynatrace.com/news/blog/automate-predictive-capacity-management-with-davis-ai-for-workflows/ https://www.dynatrace.com/news/blog/automate-predictive-capacity-management-with-davis-ai-for-workflows/#respond Tue, 11 Jul 2023 20:19:17 +0000 https://www.dynatrace.com/news/?p=58574 predictive capacity management

>> Scroll down to see predictive capacity management in action (14-second video) Our recent blog post, Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI, showed how Dynatrace Notebooks are used to predict the future behavior of time series data stored in Grail™. This follow-up post introduces Davis® AI for […]

The post Automate predictive capacity management with Davis AI for Workflows appeared first on Dynatrace news.

]]>
predictive capacity management
>> Scroll down to see predictive capacity management in action (14-second video)

Our recent blog post, Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI, showed how Dynatrace Notebooks are used to predict the future behavior of time series data stored in Grail™. This follow-up post introduces Davis® AI for Workflows, showing you how to fully automate prediction and remediation of your future capacity demands. The anticipation of future capacity demands makes it possible to completely avoid critical outages by notifying you days in advance, well before incidents arise.

Predictive capacity management starts within a Dynatrace Notebook, where the operations team explores important capacity indicators, such as the percentage of free disks, as shown below.

Figure 1. Example forecast of remaining disk capacity with upper/lower bounds and an anticipated value.
Figure 1. Example forecast of remaining disk capacity with upper/lower bounds and an anticipated value.

After exploring and selecting the most important capacity indicators for your environment, a workflow triggers forecast reporting at regular intervals. The example workflow below is triggered every Monday at 8:00 AM to provide a capacity report for all the disks that will likely run out of space within the next week.

Figure 2. Over of the predict disk capacity workflow
Figure 2. Predict disk capacity workflow

Define the forecast

The workflow uses the Davis for Workflows action to automatically trigger a forecast for a selected set of disks. The forecast operation is selected within the Davis action, and a DQL query is used to specify the set of disks and the capacity indicator metric that should be predicted. Note that you can use any time series data you can fetch from Grail using DQL within the forecast action.

While this example uses the metric dt.host.disk.free, you can choose any kind of capacity metric, such as host CPU, memory, or network load—you can even extract a metric value from a given log line.

The forecast is trained on a relative timeframe (for example, the last seven days) which is specified in the configured DQL query. The DQL query example below trains forecasting on a relative timeframe of the last seven days:

timeseries avg(dt.host.disk.free), by:{dt.entity.host, dt.entity.disk}, bins: 120, from:now()-7d, to:now()

The configuration below shows that a forecast horizon of 100 data points is requested, which means that 100 additional predicted points will expand the initially fetched 120 data bins of the source DQL query. This predicts one week into the future.

Figure 3. Detail of the forecasting workflow step
Figure 3. Detail of the forecasting workflow step

The prediction action returns all its forecasted time series lines, which can include hundreds or even thousands of individual disk predictions.

Evaluate the forecast results

Within the following TypeScript action, each disk prediction is tested against a threshold to determine if the disk will run out of space in the next week. The TypeScript code snippet below is responsible for checking for threshold violations and for preparing all the violations in a result object for subsequent actions to follow up on:

Figure 4. Evaluating the results with a custom TypeScript action
Figure 4. Evaluating the results with a custom TypeScript action

The TypeScript action returns a custom object that uses a Boolean flag (violation) to tell the follow-up actions about violations and an array of all the violation details (violations).

const predictionSummary = { violation: false, violations: new Array<Record<string, string>>() };

Tip: Download the TypeScript template from our documentation.

Trigger remediation actions

A collection of remediation actions can be used to follow up on predicted capacity shortages. In this example, two parallel actions are defined. One action sends out an email notification; the other raises a Davis problem for each violating disk. All remediation actions use the Boolean violation flag of the previous workflow action to avoid invocations when there are no violations.

Here you can see the invocation condition used in the follow-up actions that control the invocation.

Figure 5. Conditional execution
Figure 5. Conditional execution

Raise events in case of disk capacity shortage!

A TypeScript remediation action is used to iterate through all the predicted disk shortages and to raise individual alarm events. Each alarm event has custom event properties that can be used to deliver further details about the situation and to further identify the disk or host.

Figure 6. Create an alarm event for predicted shortages.
Figure 6. Create an alarm event for predicted shortages.

Tip: Download the TypeScript template from our documentation.

Review all Davis-predicted capacity problems

Navigating to the Davis problems feed, the operations team can review all the predicted disk capacity shortages. Remember, raising events and problems is an optional remediation step that can be skipped entirely by directly sending emails or Slack messages to the responsible teams.

The creation of alerting events within this workflow example highlights the flexibility and power of the Dynatrace AutomationEngine combined with the analytical capabilities of Davis AI and Grail.

Figure 7. List of events created by the workflow.
Figure 7. List of events created by the workflow.

Summary

The combination of Davis AI forecasts with Dynatrace AutomationEngine and Grail opens the door for many valuable use cases—anticipative management of capacity being the most prominent of these. Predicting future capacity shortages for thousands of disks or hosts allows operations teams to anticipate critical situations weeks before incidents occur. The flexibility and power of the Dynatrace AutomationEngine allow operations teams to react to detected shortages flexibly and to customize and implement their remediation flows.

You can install Davis® for Workflows via the Dynatrace Hub. As a starting point for implementing your own anticipative capacity management workflow, you can download all the TypeScript code used in this example from our documentation:

For full details, see Davis AI analysis in workflows documentation.

Predictive capacity management in action (14-second video)

The post Automate predictive capacity management with Davis AI for Workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automate-predictive-capacity-management-with-davis-ai-for-workflows/feed/ 0
Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI https://www.dynatrace.com/news/blog/dynatrace-automatically-monitors-openai-chatgpt-for-companies-that-deliver-reliable-cost-effective-services-powered-by-generative-ai/ https://www.dynatrace.com/news/blog/dynatrace-automatically-monitors-openai-chatgpt-for-companies-that-deliver-reliable-cost-effective-services-powered-by-generative-ai/#respond Wed, 07 Jun 2023 17:07:42 +0000 https://www.dynatrace.com/news/?p=58130 Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI

This blog post looks at how Dynatrace automatically collects OpenAI/GPT model requests and charts them within Dynatrace, as well as how abnormal service behavior can be used to identify slowdowns in OpenAI/GPT requests as the root cause of large-scale issues. Both functionalities have been part of the Dynatrace platform for a couple of years already, and so have withstood the challenges of customer usage.

The post Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI appeared first on Dynatrace news.

]]>
Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI

AI observability is becoming imperative as businesses in all sectors are introducing novel approaches to innovate with generative AI in their domains. Advanced AI applications using OpenAI services don’t just forward user input to OpenAI models; they also require client-side pre- and post-processing. A typical design pattern is the use of a semantic search over a domain-specific knowledge base, like internal documentation, to provide the required context in the prompt. This is achieved by using OpenAI services to compute numerical representations of text data that ease the computation of text similarity, called “embeddings,” for the documents as well as for the user input.

Furthermore, tools like LangChain leverage large language models (LLM) as one of their basic building blocks for creating AI agents (think of AI agents as APIs that perform a series of chat interactions that target a desired outcome) which perform complex and potentially large queries against an LLM like GPT-4. They then connect to third-party services such as online calculators, web search, or flight status information to combine real-time information with the power of an LLM.

One of the crucial success factors for delivering cost-efficient and high-quality AI-agent services following the approach described above is using AI observability to closely observe their cost, latency, and reliability.

Dynatrace enables enterprises to automatically collect, visualize, and alert on OpenAI API request consumption, latency, and stability information in combination with all other services that are used to build AI applications. This includes OpenAI as well as Azure OpenAI services, such as GPT-3, Codex, DALL-E, or ChatGPT.

AI observability example: OpenAI token consumption

Our example dashboard below visualizes OpenAI token consumption. It shows critical SLOs for latency and availability, as well as the most important OpenAI generative AI service metrics, such as response time, error count, and the overall number of requests.

AI observability dashboard showing OpenAI service health and performance
With these latency, reliability, and cost measurements in place, your operations team can now define their own OpenAI dashboards and SLOs.

Dynatrace OneAgent® discovers, observes, and protects access to OpenAI automatically, with no manual configuration, revealing the full context of used technologies, service interaction topology, security-vulnerability analysis, and the observability of all metrics, traces, logs, and business events in real time.

How Dynatrace traces OpenAI model requests

Let’s use a simple NodeJS example service to show how Dynatrace OneAgent automatically traces OpenAI model requests. OpenAI offers an official NodeJS language binding that allows the direct integration of a model request by adding the following lines of code to your own NodeJS AI application:

const { Configuration, OpenAIApi } = require("openai");

const configuration = new Configuration({

apiKey: process.env.OPENAI_API_KEY

});

const openai = new OpenAIApi(configuration);

const response = await openai.createCompletion({

model: "text-davinci-003",

prompt: "Say hello!",

temperature: 0,

max_tokens: 10,

});

Once the AI application is started on a OneAgent-monitored server, the application is automatically detected, and the traces and metrics for all outgoing requests are collected. OneAgent automatic injection of monitoring and tracing code works not only for the NodeJS language binding but also when using the raw HTTPS request in NodeJS. While OpenAI offers official language bindings only for Python and NodeJS, there is a long list of community-provided language bindings.

OneAgent can automatically monitor all C#, .NET, Java, Go, and NodeJS bindings. However, we recommend following the OpenTelemetry approach to monitoring Python with Dynatrace.

The screenshot below shows the traces that OneAgent collects, along with all the latency and reliability measurements for each of the outgoing GPT model requests.

Traces that OneAgent collects, along with all the latency and reliability measurements for each of the outgoing GPT model requests in Dynatrace screenshot

Dynatrace further refines the OpenAI calls by automatically splitting specific services for the OpenAI domain, as shown below.

General Settings for OpenAI calls in Dynatrace screenshot

Once this is done, the Dynatrace Service Flow shows the flow of your requests, starting with your NodeJS service and calling the OpenAI model, as shown below.

AI observability service flow for conversastionService

As shown in the example above, Dynatrace OneAgent automatically collects all latency and reliability-related information along with all the traces showing how your OpenAI requests traverse your service graph.

The seamless tracing of OpenAI model requests allows operators to identify behavioral patterns within their AI service landscape and to understand the typical load situation of their infrastructure.

This AI observability knowledge is essential for further optimizing the performance and cost of services.

By adding some lines of manual instrumentation to a NodeJS service, cost-related measurements are also picked up by OneAgent, collecting the number of OpenAI conversational tokens used.

Observing OpenAI request cost

Each request to an OpenAI model, such as text-davinci-003, gpt-3.5-turbo, or GPT-4 reports back how many tokens were used for the request prompt (the length of your text question) and how many tokens the model generated as a response.

OpenAI customers are billed based on the total number of tokens consumed by all the requests they make. By extracting these token measurements from the returning payload and reporting them through Dynatrace OneAgent, users can observe token consumption across all OpenAI-enhanced services in their monitoring environment.

Here is the instrumentation used to extract the token count from the OpenAI response and to report the three measurements to the local OneAgent:

function report_metric(openai_response) {

var post_data = "openai.promt_token_count,model=" + openai_response.model + " " + openai_response.usage.prompt_tokens + "\n";

post_data += "openai.completion_token_count,model=" + openai_response.model + " " + openai_response.usage.completion_tokens + "\n";

post_data += "openai.total_token_count,model=" + openai_response.model + " " + openai_response.usage.total_tokens + "\n";

console.log(post_data);

var post_options = {

host: 'localhost',

port: '14499',

path: '/metrics/ingest',

method: 'POST',

headers: {

'Content-Type': 'text/plain',

'Content-Length': Buffer.byteLength(post_data)

}

};

var metric_req = http.request(post_options, (resp) => {}).on("error", (err) => { console.log(err); });

metric_req.write(post_data);

metric_req.end();

}

After adding these lines to your NodeJS service, three new OpenAI token consumption metrics are available in Dynatrace, as shown below.

OpenAI token consumption metrics available in Dynatrace screenshot

Davis AI automatically detects ChatGPT as the root-cause

One of the superb features of Dynatrace is Davis® AI, which automatically learns the typical behavior of monitored services. Once an abnormal slowdown or increase of errors is detected, Davis AI triggers root cause analysis to identify the cause.

Our simple example of a NodeJS service entirely depends on the ChatGPT model response. So, whenever the latency of the model response degrades or the model request returns an error, Davis AI automatically detects it.

In the example below, Davis AI automatically reported a slowdown of the NodeJS prompt service and correctly detected the OpenAI generative service as the root cause of the slowdown.

Davis AI automatically reportes a slowdown of the NodeJS prompt service and correctly detected the OpenAI generative service as the root cause of the slowdown

The Davis problem details page shows all affected services for which the OpenAI generative service was the root cause of the slowdown, along with the ripple effects of the slowdown.

The problem details also list all Service Level Objectives that were negatively impacted by the slowdown.

List of all Service Level Objectives that were negatively impacted by a slowdown

AI observability with Dynatrace brings peace of mind when using OpenAI models

The massive popularity of generative AI cloud services, such as OpenAI’s GPT-4 model, is forcing companies to rethink and redesign their existing service landscapes. Integrating generative AI into traditional service landscapes comes with all kinds of uncertainties. Using AI observability from Dynatrace to observe OpenAI cloud services helps you gain cost transparency and ensure the operational health of your AI-enhanced services.

Also, full transparency and observability of AI services will play a significant role in upcoming AI regulations at a national level and for risk assessments within your own company.

For further details, you can view the full source of the NodeJS service on GitHub.

The post Dynatrace automatically monitors OpenAI ChatGPT for companies that deliver reliable, cost-effective services powered by generative AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-automatically-monitors-openai-chatgpt-for-companies-that-deliver-reliable-cost-effective-services-powered-by-generative-ai/feed/ 0
Davis AI: Your personal interactive troubleshooting assistant  https://www.dynatrace.com/news/blog/davis-ai-your-personal-interactive-troubleshooting-assistant/ https://www.dynatrace.com/news/blog/davis-ai-your-personal-interactive-troubleshooting-assistant/#respond Wed, 24 May 2023 08:59:02 +0000 https://www.dynatrace.com/news/?p=57846 business observability

When Dynatrace started reinventing cloud-native service tracing and observability ten years ago, it was already clear that human operators were overwhelmed with traditional monitoring systems' massive raw data inflow. Besides being unable to watch that amount of telemetry data on dashboards, classic operations teams were also blown away by the sheer number of alerts they received 24/7 from hundreds of different monitoring tools. 

The post Davis AI: Your personal interactive troubleshooting assistant  appeared first on Dynatrace news.

]]>
business observability

Update: We’ve expanded AI-powered dashboarding with Dynatrace Intelligence, delivering smarter insights, forecasting, and anomaly detection across the Dynatrace platform.
Dynatrace Intelligence is the evolution of Davis AI®, improving how users interact with and understand observability data.

With the introduction of Davis® root-cause detection, Dynatrace reduced the amount of single-alert spam that arises when large-scale incidences occur. Instead of immediately firing off an alert for all raw events, the Davis root-cause engine follows each violating service’s causal relationships. By automatically following the causal direction of the topology between services and their underlying infrastructure, Davis collects all raw events that belong to the same root cause and then notifies you by raising a problem.

With interactive problem mode, Dynatrace introduces a new, powerful troubleshooting assistant. This blog post explains how Davis can help reduce your MTTR (mean time to resolve) using interactive user guidance that retains context when drilling deeper into problem analysis.

Davis problem analysis
Select any entry in the side panel to navigate to the corresponding metric, in context.

Faster remediation through precise root cause analysis

Once Davis identifies a problem, a Problem overview page is created, which shows a comprehensive management summary of what happened (impact) and the root cause of the problem. DevOps teams use this page to quickly identify and remediate unexpected incidences.

Usually, the journey doesn’t stop here. When the DevOps team has finished their work, software experts must investigate the underlying software stack. They need to analyze all relevant information that Davis found along the deployment stack to avoid such problems in the future. When navigating to the underlying service—identified as the root cause—the problem detail page opens with retained problem context, which includes:

  • Date and time of the current problem, so you don’t need to manually adapt the date and time on each page in the analysis journey.
  • A side panel that interactively informs you about all problem-related information for the relevant service.
  • Davis highlights all relevant problem information on each page you navigate to.

The screenshot below shows how Davis interactively guides you by highlighting all the relevant information with red and yellow markers (on the left side) while showing a list of AI root-cause findings in the side panel on the right (if the Davis side panel is closed, an icon is displayed on the right-hand panel so you can re-open it).

AI root-cause findings

Davis highlighting detected problems in side panel

Optimize your software stack using Davis interactive problem mode

Watch out for red and yellow markers in the navigation section headers—these indicate that Davis has found information related to the problem.

The red marker highlights events and their duration, whereas the yellow marker indicates metric anomalies where suspicious metric change points were found during the problem analysis. The yellow metric change points highlight a point in time, while the red markers represent event durations.

If you select one of the markers (either directly or via the side panel), you can view additional information, such as the timeframe and duration.

Davis AI change point and event markers

Davis AI change point (in yellow on the left) and event duration (in red on the right) markers

Meeting SLO requirements

In addition to providing context to detected problems, Davis also supports you when spikes are detected in connected SLOs (Service Level Objectives). Via the dedicated SLO button in the top bar, service-level objectives relating to the selected service can be reviewed immediately without losing context.

Spikes can easily be investigated by selecting a timeframe and clicking Analyze. Davis instantly collects all connected signals and provides relevant, contextual information. Watch the following video for examples of how the interactive problem mode helps identify SLO-relevant issues.

Davis SLO analysis
Review related Service Level Objectives (SLOs)

Summary

Davis problem detection and root cause analysis is essential for modern AIOps (Artificial Intelligence for IT Operations) and DevOps to minimize the MTTR. Real-time insights are crucial for quickly triaging unexpected incidents and remediating them in a timely manner.

Davis interactive problem mode guides you through all the detailed problem-related information and marks problems visually to make them easier to understand. It also seamlessly integrates user-defined SLOs, including leveraging Davis AI for analyzing SLO degradations, which saves precious time during critical incidents. You no longer need to leave the context of your page when using the side panel for navigational help to dig through all relevant findings and SLOs discovered during root cause analysis.

We’re, of course, highly interested in your feedback! We encourage you to try the interactive problem mode and share your feedback and product ideas via the Dynatrace Community. Every message we receive helps us to continuously improve the Dynatrace platform.

The post Davis AI: Your personal interactive troubleshooting assistant  appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/davis-ai-your-personal-interactive-troubleshooting-assistant/feed/ 0
Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI https://www.dynatrace.com/news/blog/stay-ahead-of-the-game-forecast-it-capacity-with-dynatrace-grail-and-davis-ai/ https://www.dynatrace.com/news/blog/stay-ahead-of-the-game-forecast-it-capacity-with-dynatrace-grail-and-davis-ai/#respond Wed, 26 Apr 2023 12:03:45 +0000 https://www.dynatrace.com/news/?p=57280 Cloud observability graphic

Anticipatory management of cloud resources within highly dynamic IT systems is a critical success factor for modern companies. Operators need to closely observe business-critical resources such as storage, CPU, and memory to avoid outages that are driven by resource shortages.

The post Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI appeared first on Dynatrace news.

]]>
Cloud observability graphic
>> Scroll down to see Davis Forecasting in action

Traditionally, cloud-resource management is done by collecting telemetry data for critical-capacity resources and configuring multi-level reactive alerting (warnings, errors, and critical errors) for those resources.

While the traditional approach to cloud-resource management might have been acceptable in the past, it doesn’t scale up to address the requirements of modern cloud environments. Highly dynamic services are deployed to the cloud globally, where resources are requested and deployed on demand. The end result of this global scale is that—without the right tools—operators are completely lost in alert storms.

Some of our customers run tens of thousands of storage disks in parallel, all needing continuous resizing. This can lead to hundreds of warnings and errors every week. The most annoying aspect of this, according to the operations teams, is that alerts are often sent after business hours, including on weekends.

Disk alert storm
Figure 1. In the past, disk alerting was triggered using static capacity thresholds.

One effective capacity-management strategy is to switch from a reactive approach to an anticipative approach: all necessary capacity resources are measured, and those measurements are used to train a prediction model that forecasts future demand.

The example below shows how the reactive approach can be transformed into a scheduled, predictive capacity management model. With this approach, when available capacity of a critical resource is forecast to soon fall below acceptable levels, the operations team is notified with a single report, sent during business hours, well in advance of the resource actually experiencing a resource shortage.

Disk forecast report
Figure 2. Capacity planning with Dynatrace Davis® AI forecasts actual usage rather than waiting for thresholds to be exceeded before alerts are sent.

Use Grail and Davis to predict capacity demands

The Dynatrace Query Language (DQL) allows you to analyze all data that’s stored for your environment within the Dynatrace Grail™ data lakehouse. With the newly introduced Notebooks, you can use DQL for exploratory analysis of any capacity-related telemetry.

The screenshot below shows a section of an example notebook that plots the average percentage of free disk space over the last 7 days.

Disk capacity notebook
Figure 3. Example notebook showing average percentage of free disk space during past 7 days.

Select any line in the chart to display the available actions for that line. For example, you can request that Davis AI forecast the future of any given time series.

Disk capacity notebook forecast
Figure 4. Select Filter and forecast for any chart line to start Davis Forecast.

Davis AI analyzes the selected time series, automatically chooses the best prediction model based on the characteristics of the time series, and then trains a prediction model. After the training is finished, the notebook chart shows a probabilistic forecast of the given time series with upper and lower bounds as well as the predicted value.

The example result of the trained prediction model on this disk capacity measurement shows that the lower bound of the prediction (worst case scenario) will fall below 6% free disk capacity in early April.

Given this information, the operations team can anticipate that they need to resize this disk before early April (during business hours, of course).

Disk capacity notebook forecast result
Figure 5. Example forecast of remaining disk capacity with upper/lower bounds and an anticipated value.

While this notebook focuses on only one disk, Davis Forecast can learn and predict the future capacity needs of thousands of individual disks in parallel. For example, in our own cloud infrastructure at Dynatrace, we track over 8,000 disks that require periodic resizing. By running a scheduled, weekly forecast, our cloud automation teams avoid reactive alerts sent outside of business hours.

Reactive alerts as last line of defense

Of course, unexpected things still happen and a weekly forecast can’t, for example, anticipate a customer onboarding 5,000 OneAgents on a Sunday morning. Therefore, reactive alert conditions remain in place, to ensure that alerts are still sent if needed during unexpected events. Scheduled forecasts do not replace these reactive alerts, rather they serve as a last line of defense for unforeseen situations.

AutoML detects seasonality and chooses the best prediction model

Dynatrace invested significantly into simplifying prediction and forecasting for you. You can use any DQL query that yields a time series to train a prediction model. This AutoML approach analyzes the statistical characteristics of any time series (variance, seasonality, trend, and noise) to determine the best prediction model.

The AutoML approach also helps when you need to automate your forecasts and the underlying metric characteristics change over time.  Please see Davis Forecast analysis documentation to learn more about our AutoML approach and which algorithms are used within the Davis Forecast service.

So far, we’ve only discussed resource consumption measurements, which by nature show more linear changes than seasonal characteristics. Now let’s see how the AutoML approach selects the best suitable method for a time series that has seasonal behavior. The example below shows that Davis Forecast automatically detects any given seasonality, independent of the noise level, and correctly returns a probabilistic forecast.

Probabilistic seasonal forecasting with AutoML
Figure 6. Probabilistic seasonal forecasting with AutoML

Automate your Davis capacity forecasts

While the operations team could regularly check this notebook, see what Davis Forecast anticipates as the upcoming capacity need, and then proactively resize all disks that will run out of space the following week, the better option is to use the newly introduced AutomationEngine to schedule an automated weekly Davis prediction workflow. With this approach, an automated workflow can automatically run a forecast for all disks, check against a critical capacity limit, and notify the operations team with a list of the disks that need their attention. Below is an example of such a workflow that includes a Davis Forecast action and a notification email action. The new forecasting capabilities together with Dynatrace AutomationEngine and the Workflows app, allow you to automate any predictive analytics in a few simple steps.

Davis Forecast as part of a workflow automation
Figure 7. Use Davis Forecast as part of a workflow automation.

This use case can be spun even further: so far, we’ve automated the forecast and introduced reporting for the disks that need to be resized. Once the operations team becomes familiar with the anticipatory approach, full automation can easily be configured by adding additional action steps within the existing workflow, for example, automated provisioning of new disk space.

Summary

Davis Forecast provides a powerful mechanism on top of the Grail data lakehouse that enables organizations to switch from reactive strategies to more proactive anticipative strategies. Such predictive approaches help avoid outages and reactive alert storms outside business hours.

By offering a standard forecast mechanism on top of the powerful DQL query language, Dynatrace opens predictive analytics for any kind of anticipative use cases, including the business-critical topic of predictive capacity management. These new analytics capabilities can be used as part of your exploratory analytics in Notebooks, as a step within workflows, or as part of your custom app–addressing your specific business needs with Dynatrace AppEngine.

See Davis Forecasting in action

Check out the “Forecasting with Dynatrace” Observability Clinic, where Linda Gratzer, Andreas Grabner, and Bernhard Kepplinger dig deeper into the topic of forecasting and the data science behind it, and also share a live demonstration.

For further details, have a look at our Davis AI Forecast Analytics documentation, or watch the recording of my Perform breakout session, Easy forecasting and predictive analytics with Davis AI.

We are of course highly interested in your feedback! We encourage you to try out Davis Forecasting and then head over to the Dynatrace Community and share your suggestions and product ideas, to help us continuously improve the Dynatrace platform.

The post Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/stay-ahead-of-the-game-forecast-it-capacity-with-dynatrace-grail-and-davis-ai/feed/ 0
Leverage Dynatrace AIOps in GitHub CI pipelines to prevent critical incidents https://www.dynatrace.com/news/blog/leverage-dynatrace-aiops-in-github-ci-pipelines-to-prevent-critical-incidents/ https://www.dynatrace.com/news/blog/leverage-dynatrace-aiops-in-github-ci-pipelines-to-prevent-critical-incidents/#respond Wed, 02 Mar 2022 15:39:52 +0000 https://www.dynatrace.com/news/?p=49143 Business analytics graphic

The newly released Dynatrace GitHub Action enables DevOps teams to fully leverage all available observable information from CI pipelines. By leveraging CI pipeline information within your observability platform, your DevOps teams can closely monitor the health of all pipelines and react faster when critical incidents are detected. This makes deployment outages much easier to detect and remediate.

The post Leverage Dynatrace AIOps in GitHub CI pipelines to prevent critical incidents appeared first on Dynatrace news.

]]>
Business analytics graphic

GitHub Actions offer a great way to automate, customize, and execute your software development workflows. A fully automated software release pipeline helps you release new functionally faster and frees up precious developer resources to focus on innovation. GitHub workflows can be quite complex, executing dozens of individual build, test, and deployment steps. By dropping a Dynatrace GitHub Action into your GitHub CI/CD workflow you gain observability and real-time insights into the performance of your pipeline. The backflow of CI/CD workflow information also helps your DevOps teams quickly find faulty software deployments and react quickly to prevent and remediate critical outages.

What is a GitHub Action?

GitHub Actions are configurable workflow steps that can be used to perform any kind of automation task. Examples of such automation tasks include checking out a repository, building some software, signing the resulting binary, and, finally, copying and deploying the resulting software artifacts to a cloud runtime environment.

A series of such actions can be combined into a GitHub workflow that can be triggered whenever new code is committed or when a new release tag is pushed to a repository. Multiple GitHub Actions are available within the GitHub Marketplace:

GitHub Actions
The use of GitHub Workflows is straight forward as their configurations are stored in your GitHub repository.

With a single click, you can select a GitHub Action that begins building a Docker container from your repository and automatically deploys the resulting image to Docker Hub. The workflow is configured to run automatically whenever any push or pull request is made on the main branch, or when the workflow is triggered manually.

As you can see below, all that’s needed to automatically build and deploy a container on Docker Hub is a 40-line YAML workflow configuration file:

GitHub Actions

Dynatrace GitHub Action

The purpose-built Dynatrace GitHub Action is available on the GitHub Marketplace in the monitoring category. It’s useful for seamlessly observing all your GitHub workflows. Simply drag and drop the Dynatrace Action into your CI pipeline and collect all your relevant metrics and events during each of the execution steps.

GitHub Actions

This GitHub Action is part of the Dynatrace Open Source Initiative, which maintains and contributes to numerous open source projects.

What‘s the value of observing your CI/CD pipeline in Dynatrace?

Collecting insights about your CI/CD pipeline comes with many benefits. Collecting statistics about the execution of your build and deployment automation workflow helps you evaluate the overall quality of your pipeline.

Counting the number of failing builds versus the number of successful builds helps you understand why and when code commits break your pipeline. Counting the number of failing integration tests helps to inform the responsible dev teams early so that they can prevent critical issues from reaching production.

A Service Level Objective (SLO) can be defined to help keep track of the quality of your build and test pipeline, as shown below:

Service Level Objective (SLO)
By closely monitoring your CI/CD pipeline health in Dynatrace, you can react early if the quality of the pushed code decreases.

Integrate workflow information within Dynatrace Davis root-cause detection

Besides collecting continuous health information about all your build pipelines, it’s also crucial that you have the ability to quickly identify when a faulty deployment is the root cause of a large-scale incident in production.

By seamlessly feeding your pipeline health metrics and event information back to your Dynatrace monitoring environment, Dynatrace Davis AIOps can pick up and evaluate that information in case of detected incidents.

One use-case is to send all relevant deployments to the monitoring platform and to attach that information to all affected services. Below you can see a Dynatrace GitHub Action configuration that counts the number of failing and successful builds. The Action sends a deployment event to Dynatrace with each pipeline execution.

Dynatrace GitHub Action configuration

Each of the executions now automatically informs Dynatrace of the new deployment, attaching the information to all services that are named ‘loginService’. The relevant information is then shown on the service’s overview page in the Dynatrace web UI. In case of an incident, the Davis AIOps engine automatically picks up the metric and identifies the deployment event as the root cause of the issue.

See the example below which shows that Davis AIOps has identified a new service deployment triggered by the GitHub CI pipeline as the root cause of a slowdown in production.

Root cause dashboard

Root cause dashboard

New to Dynatrace?

If you haven’t used Dynatrace yet, try it out by starting your Dynatrace free trial today.

The post Leverage Dynatrace AIOps in GitHub CI pipelines to prevent critical incidents appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/leverage-dynatrace-aiops-in-github-ci-pipelines-to-prevent-critical-incidents/feed/ 0
Dynatrace AI predicts SLO violations and pinpoints root causes proactively https://www.dynatrace.com/news/blog/dynatrace-ai-predicts-slo-violations-and-pinpoints-root-causes-proactively/ https://www.dynatrace.com/news/blog/dynatrace-ai-predicts-slo-violations-and-pinpoints-root-causes-proactively/#respond Mon, 28 Jun 2021 23:07:44 +0000 https://www.dynatrace.com/news/?p=45137 SLOs

Dynatrace enables Site Reliability Engineering (SRE) teams to proactively ensure the highest service quality levels. Davis, the Dynatrace AI engine, identifies potential contributors to SLO violations in real time, before thresholds are breached. Dynatrace pinpoints the root causes of problems and their impact on SLOs.

The post Dynatrace AI predicts SLO violations and pinpoints root causes proactively appeared first on Dynatrace news.

]]>
SLOs

In modern service landscapes, Service-level-objectives (SLOs) are the chosen methodology of Site Reliability Engineering (SRE) teams for ensuring the high quality of delivery of their digital services. There is however a major challenge faced by many SRE teams: how to catch relevant degradations early, before a long-term SLO shows an unhealthy state.

Are you still “reacting to bad numbers”?

SLOs with an observation period of, for example, one week, are of course not overly affected by short-lived outliers. However, such observation periods come with a disadvantage: incidents can pile up and there is a delay between those incidents and the corresponding health metrics ultimately dropping low enough to trigger a warning.

Teams who are primarily reactive in their approach therefore use SLOs to decide when the state of a system has become so bad that it requires intervention.

Some SRE teams counter this by defining the same SLOs for different observation periods to reduce the reaction times in case of incidents. Many teams use three different levels of observation periods, one for strategic decisions, one for tactical decisions, and one short period for catching incidents. These redundancies can of course create additional efforts and complexity.

Error budgets and the tracking of their burn rates offer a much better approach, however without extensive manual effort, this approach still leaves two questions open:

  • How can I detect anomalies early, before they impact your SLOs?
  • To facilitate fast remediation, how can I quickly identify the root causes of emerging issues that have massive potential SLO impact

Most monitoring tools offer only a single SLO metric. However, watching a single SLO health metric and error budget drop doesn’t provide much in the way of answers; it only confirms the obvious—that your SLO is unhealthy. In the best case scenario, to answer the above questions, you need experts to conduct manual investigation and interpret the data for you. In the worst case scenario, the nature of today’s dynamic and heterogenous environments renders such manual investigation impossible.

Dynatrace proactively helps Site Reliability Engineers keep their SLOs healthy

Dynatrace Davis, our AI-engine, offers a unique feature that overcomes the fundamental challenge of reacting quickly enough, even within strategic observation periods. Davis notifies you when any of your SLOs are at risk, before any metrics turn red.

This works out-of-the-box because Dynatrace understands how all your application and infrastructure components depend on each other. In this way, Davis can link defined SLOs to those anomalies that present potential negative impact.

Davis AI predicts if future SLO health is at risk

Let’s look at an example where an SLO was defined for the stability of a frontend service that shows a perfect 100% SLO health status:

Dynatrace screenshot SLO status

Notice in the above SLO tile that Davis has displayed a critical error indicator to inform the SRE team about an ongoing incident within the service topology that the SLO covers. Even though the SLO still shows perfect 100% health, Davis AI is proactively predicting that the future SLO health is at risk.

Dynatrace AI pinpoints the root causes of SLO-impacting incidents

Further, a single click on this tile displays all active incidents along with the potential negative impact on the future health of the SLO.

Dynatrace screenshot What's the root-cause

A drill-down from an unhealthy SLO takes you to a filtered feed of detected problems that are the root causes of these incidents. This precise AI-assisted identification of root causes saves valuable time for SRE and DevOps teams during critical service outages, instead of just showing a single, isolated health metric.

Get up and running in under a minute with SLO templates

Service-level-objectives consist of carefully selected Service-level-indicators (SLIs) which provide a quantitative measure of each aspect of the service level. Typically, an SRE team spends a good amount of time selecting the best indictor metrics for their given services, which then leads to well-defined SLOs that reflect the service quality.

The Google Site Reliability Engineering page is a great read for understanding and embracing the idea of defining SLOs for reliable global IT services.

However, getting started with SLOs in Dynatrace is even easier.

We offer a collection of best-practice SLO definitions for various use cases beyond the observability domain; simply choose one of the predefined SLO templates that Dynatrace provides out-of-the-box.

For example, you can measure the quality of service of your mobile app offering. Dynatrace offers a mobile crash-free users SLO template that you can use to create a best-practice SLO for measuring the reliability and stability of your mobile apps.

Dynatrace screenshot Add new SLO

Once you’ve defined your business-critical SLOs, you’re all set. Davis will then automatically analyze your SLOs continuously and provide a truly proactive approach to SLOs.

Get started with SLOs

Proactive SLO impact analysis is available with the release of Dynatrace version 1.220. If you’re new to Dynatrace, you can experience the magic yourself by starting a Dynatrace free trial.

Learn more

If you want to learn more about SLIs/SLOs, here are a few resources that we recommend:

You can check out the session recording to find all the details.

The post Dynatrace AI predicts SLO violations and pinpoints root causes proactively appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-ai-predicts-slo-violations-and-pinpoints-root-causes-proactively/feed/ 0
Use the Davis® AI to detect outages within your custom data streams https://www.dynatrace.com/news/blog/use-the-davis-ai-to-detect-outages-within-your-custom-data-streams/ https://www.dynatrace.com/news/blog/use-the-davis-ai-to-detect-outages-within-your-custom-data-streams/#respond Fri, 05 Mar 2021 19:43:51 +0000 https://www.dynatrace.com/news/?p=43017 StatsD, Telegraf, and Prometheus observability

In today’s complex IT environments, the sheer volume of data created makes it impossible for humans to monitor, comprehend, or troubleshoot problems before they impact the experience of your end users. Dynatrace Davis® AI has proven over the past four years that a fully automated approach to problem analysis is the only valid approach—especially in […]

The post Use the Davis® AI to detect outages within your custom data streams appeared first on Dynatrace news.

]]>
StatsD, Telegraf, and Prometheus observability

In today’s complex IT environments, the sheer volume of data created makes it impossible for humans to monitor, comprehend, or troubleshoot problems before they impact the experience of your end users.

Dynatrace Davis® AI has proven over the past four years that a fully automated approach to problem analysis is the only valid approach—especially in highly dynamic, web-scale cloud environments where manual root cause analysis is impossible. Davis automatically analyzes and alerts on many important anomalies that can occur within your IT environment, such as host, process or service outages.

Still, you might have use cases that rely on important custom data streams. For example, you might be using:

  • any of the 60+ StatsD compliant client libraries to send metrics from various programming languages directly to Dynatrace;
  • any of the 200+ Telegraf plugins to gather metrics from different areas of your environment;
  • Prometheus, as the dominant metric provider and sink in your Kubernetes space.

But how can you ensure that these data sources are always up and running? How can you ensure that the measurements are always available and healthy?

In order to detect all kinds of availability issues, you need AI-powered alerting for your third-party data sources, too.

Let the Davis AI automatically detect outages within your custom data streams

Recently, we simplified StatsD, Telegraf, and Prometheus observability by allowing you to capture and analyze all your custom metrics. By automatically feeding these captured metrics into our Smartscape topology model and Davis AI, Dynatrace eliminates the need for manual maintenance of hundreds of alerts, thanks to our trusted auto-adaptive baseline engine.

Dynatrace Hub

Now you can:

  • Alert on the outage of a custom data source

Alert whenever your custom data source stops sending measurements for whatever reason. While Dynatrace OneAgent has resilience, built-in health checks, and reports automatically on unavailability, your custom data sources will not offer the same self-monitoring capability.

  • Alert on expected but missing measurements

This use case might sound similar the first one, but it solves a completely different purpose. Here the data source is up and running but you’re not receiving the expected measurements. Say, for example, that you expect the count of executed batch jobs to be sent every 10 minutes, but you haven’t received a count during the last 30 minutes. As this is suspicious, you need Davis to report on the situation.

  • Alert on unhealthy metric states and missing data

This use case is a combination of the first two, along with alerting on low or high levels of the expected measurements. An example here is if you report the CPU usage of a SNMP network device through your Telegraf agent, and you want to receive an alert whenever the CPU usage reaches a critical level or finally when the device is gone, but no measurements are coming. So, this alert is a combination of unhealthy metric state together with missing measurements.

Now let’s see how you can detect and alert on those use cases listed above within your own Dynatrace environments.

Easily alert on the outage of a custom data stream

Let’s assume that you are using our convenient metric ingest protocol through the OneAgent ingest channel, or through the Dynatrace REST API. The metric source in this case is either a third-party agent such as Telegraf or your own script written in any of your favorite languages.

In this example, we send measurements from a Synology NAS to a Dynatrace monitoring environment to check the health and activity of the network disk.

For this purpose, we embed a Telegraf agent into a Docker container and run it directly on Synology NAS to continuously report selected SNMP measurements. As Synology offers an integrated Docker runtime within their network disks, it doesn’t take much effort to turn the network disk into a custom data source for Dynatrace.

The screenshot below shows the Dynatrace Metrics browser, filtered by the metrics that the Synology custom data source reports:

Synology metrics dashboard

Charting the ssCpuUser metric shows that a reliable and continuous stream of measurements is sent directly by the Synology network device, as shown below:

Synology CPU dashboard

The alert on an outage of the data source of course relies on a reliable continuous data stream, as you simply can’t alert on a data source that reports measurements in a sparse and unregular manner.

As our data source is expected to report back measurements to Dynatrace every 10 seconds, it is the perfect source for alerting on outages.

To detect an outage of the Synology disk, navigate to Settings > Anomaly detection > Custom events for alerting. Then create a new alerting config and select your own metric as input, as we’ve done for the Synology CPU metric below:

Synology configuration dashboard

Here, the individual metric dimensions that can be used to filter the alerting to a specific measurement are also shown. For example, if you monitor 100 Synology NAS devices, the serial number dimension will show 100 different numbers, one for each of your devices.

Mind that without a filter, you will alert on all one hundred NAS devices, so use the filter to reduce the alerting scope of your configuration.

The next step is to enable the new Alert if data is missing…. option within the individual monitoring settings. Davis automatically suggests a reasonable CPU threshold value for your device along with an observation period of 3 out of 5 minutes, as shown below:

Synology configuration dashboard

Define your alert message

The final step in the configuration is to define the text message and event description that will be sent whenever the Synology NAS stops sending data.

As shown below, we’ve defined the event name to be Synology NAS outage, along with a Davis severity level of Availability and added the violating Synology model name in the event description:

Synology configuration dashboard

Use dimension placeholders to customize alert messages

Another newly introduced feature is the selection of a metric dimension within the alert message. In our example, we’ll use the {dims:modelname} placeholder because it contains the human readable name of our Synology disk.

Whenever one of the Synology NAS devices experiences an outage, its name will show up automatically in the alert message.

In general, the placeholder suggestion proposes all the available metric dimensions, where you can select the most important ones for your alert message.

In case you would like to get the full set of metric dimensions of the violating measurement, you have to use the unfiltered {dims} placeholder.

See the possible placeholders for the Synology example below:

Synology configuration dashboard

Limits safeguard the health of your monitoring infrastructure

Within your Dynatrace environment there are specific limits in terms of possible alerting configurations that help to safeguard the health of your monitoring infrastructure.

We classify metric alerting into what we call Basic metric queries, where the number of actively monitored metric dimensions is not limited and Advanced metric queries, where we must apply a technical limit of 100,000 dimensions per environment.

Most of the simple metric threshold models are under the category of basic metric queries, as their query overhead is negligible.

Whenever you switch to Alert on missing data, you’re switching from a basic metric query to an advanced metric query as Dynatrace has to actively check for the existence of data in each minute.

You see the query category in the preview, along with the used quota within your environment, as shown here:

Synology configuration dashboard

Last step: Testing the outage incident

Finally, now that we’ve configured the outage detection for Synology NAS, we can turn off the SNMP monitoring and the appropriate outage report shows up in Dynatrace, as you can see here:

Synology alert dashboard

Summary

Opening the Dynatrace Software Intelligence platform for all kinds of third-party data ingest has opened the door for interesting use-cases, such as the monitoring of network devices or the reporting of custom metrics.

By giving Davis the capability to actively detect the outage of third-party data sources, such as Telegraf, StatsD or SNMP, and to report on missing but expected data, your Dynatrace environment becomes even more powerful in terms of detecting all kinds of availability issues.

Seeing is believing

New to Dynatrace? Try it out by starting your free trial today.

The post Use the Davis® AI to detect outages within your custom data streams appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/use-the-davis-ai-to-detect-outages-within-your-custom-data-streams/feed/ 0
Leverage the new Problems API to resolve Dynatrace-detected issues faster in your third-party tools https://www.dynatrace.com/news/blog/leverage-the-new-problems-api-to-resolve-dynatrace-detected-issues-faster-in-your-third-party-tools/ https://www.dynatrace.com/news/blog/leverage-the-new-problems-api-to-resolve-dynatrace-detected-issues-faster-in-your-third-party-tools/#respond Tue, 03 Nov 2020 15:28:41 +0000 https://www.dynatrace.com/news/?p=40632 Root cause; root cause analysis

Our new Problems REST API v2 fully delivers the Davis® AI power for third-party tools, allowing you to resolve Dynatrace detected problems faster in your third-party tool of choice.

The post Leverage the new Problems API to resolve Dynatrace-detected issues faster in your third-party tools appeared first on Dynatrace news.

]]>
Root cause; root cause analysis

Dynatrace v2 APIs transform your entire organization by making it easy to get started with monitoring automation and to solve your business problems with data-driven answers.

A few months ago we wrote about how you can scale your API operations with our version 2 APIs, by showing off the Dynatrace Metrics API v2 and the Monitored entities API v2. Today we’re happy to announce a further API that will make your life easier: our brand new Problems REST API v2.

Leverage the power of the Davis® AI with the new Problems REST API v2

Davis, our radically different AI causation engine, automatically processes billions of dependencies and pinpoints the root cause of performance issues with unmatched precision. To achieve this, Davis uses powerful AIOps capabilities such as automatic impact assessment for detected problems and the grouping of all critical raw information related to each incident.

We know how crucial this information can be when it comes to reporting, analyzing long-term trends, or identifying problem entities in your infrastructure. To allow you to resolve Dynatrace detected problems faster in your third-party tool of choice, but also to build feature-rich analysis and reporting use cases on top of Dynatrace Davis-analyzed results, our new Problems REST API v2 fully delivers the Davis AIOps power for third-party tools.

We deliberately chose an API-first approach in designing our own new web UI for the Problems list view to ensure that external integrations receive the same level of expressiveness through the public REST API. All filters and problem meta-information that are available in our own problem list are now seamlessly available through the public Problems REST API v2. Among other use cases, the this will allow you to:

Read on below to understand the benefits of the new API and possible use cases for leveraging its newly added capabilities.

Analyze and understand long-term trends by easily paging through millions of detected problems and jumping to specific pages

You might want to quickly and efficiently page through the millions of Dynatrace Davis-detected events and problems in an external UI rather than scrolling through them in one long list.

To provide better accessibility and the ability to track results on a per-page basis, the new Problems v2 and Events REST APIs allow you to page through a huge query result—you can load more results and access specific pages within these results. The paging follows a cursor-like approach, which means that, in your first query, you specify the page size of each individual result and then use the returned cursor to navigate from the first result page to the last.

A typical use case is the display of problems in an external UI and offering a paged or lazy-loading approach that allows the user to load more results on demand. Let’s say that you want to page through all the problems that were detected within the last year and return the result in page sizes of 50 problems per request. See the example request below for querying the first page:

https://YOUR_ENV.live.dynatrace.com/api/v2/problems?from=-1y&pageSize=50

The result shows the first 50 problems ordered in the sequence in which they were detected, the total count of the problems (6336), and the cursor (nextPageKey) to get the next page.

{
"totalCount": 6336,
"pageSize": 50,
"nextPageKey": "AQANMTU1MTUyNzMzNzEyNAEADTE1ODMxNDk3MzcxMjQBABwxNzA4Zj",
"problems": [ … ] }

Note: Each incremental result returns its own nextPageKey cursor, which you must include to get to the next page.

A typical error here is to mistakenly only use the first nextPageKey cursor, which lands the user in an endless loop that continuously returns to the first page. To avoid this, request the next page by including the nextPageKey value from the previous result, as shown below:

https://YOUR_ENV.live.dynatrace.com/api/v2/problems?nextPageKey= AQANMTU1MTUyNzMzNzEyNAEADTE1ODMxNDk3MzcxMjQBABwxNzA4Zj

Use powerful filters to focus on the problems you’re most interested in

Instead of paging through a huge number of problems or events, you can now use an efficient query to target the issue you’re most interested in.

You can achieve this with the same consistent query approach that we follow in all the API v2 endpoints: by using the entitySelector as well as an endpoint and domain-specific selector, namely, the problemSelector.

By using the entitySelector, you can narrow down the query on the topology, for example, by querying problems on a specific host. The domain-specific problemSelector allows you to further narrow down the query by using problem-related attributes, such as whether a problem is in an “open” or “closed” state. The concept behind selectors is to foster interplay between endpoints. This means that the entitySelector in the Problems v2 and Events endpoints can also be used in the Monitored entities v2 endpoints to select the same subset of entities.

Problem and entity selectors in the API Explorer

Let’s try a simple example query for problems that occurred on hosts within a management zone named PROD. To do this, we’ll use the entitySelector as shown below:

https://YOUR_ENV.live.dynatrace.com/api/v2/problems?from=-1y&entitySelector=type(“HOST”),mzName(“PROD”)

Note: This query is focused first on the topology and then on the problems that occurred within this topological section. You can also try to use the management-zone filter in the problemSelector to get all problems that were detected within the management zone rather than within a section of the topology.

By using the problemSelector, you can further refine the query with problem-related criteria such as problem severity level or problem status.

Define SLOs and KPIs for your services by fetching root cause details across the Problems, Metrics, and Events API endpoints

Davis detects incidents in your monitoring environment, analyzes the relevant topology, and collects all available information that indicate the ultimate root cause component. Each data hint that leads Davis in the correct direction of a root cause is called “evidence.” These hints are exposed through the new Problems REST API v2.

Root cause evidence can be manifold: baseline violation events, non-metric events (for example, process crashes), information events (for example, a deployment), or change points detected on any of the analyzed metrics. All this information is now exposed through the new Problems REST API v2, which enables further reporting.

One use case is to automatically fetch all metric-based root cause evidence for the past week and to check which metrics are the most “interesting” in this regard. Such information can then be used in Keptn to further define SLOs and KPIs for your services.

See the example below of metric-based evidence that was automatically detected by Davis during its causation run. This includes the metric name and metric identifier, which can be used seamlessly in the Metrics API v2:

Root cause evidence based on a metric returned by Davis

See the example below of another type of evidence that’s based on an event that Davis detected on an affected topological node:

Root cause evidence based on an event returned by Davis

Again, as before, you can use the event ID in the evidence returned by a problem to query that event using the Events API.

Easily identify the problems that affect most of your real users by accessing impact-related information

Besides the root cause, Davis AIOps impact analysis can play a major role in your reporting use cases for third-party tools. Once an incident is detected on an entity that also shows incoming transactions, Davis follows the backtrace of those transactions and identifies the entry points (that is, the application and services where those transactions originate). See the business-impact analysis example below.

Problem impact analysis details

With our newly introduced Problems API v2 endpoints, external integrations can now access the same impact-related information (shown below). With this information you can, for example, check which problems affected the most real users.

Summary

With the new Problems REST API v2, powerful Davis AIOps capabilities such as the numerous filtering possibilities, paging, and export of impact and root cause information are all seamlessly exposed to pave the way for building feature-rich external analysis and reporting use cases on top of Davis analysis results from Dynatrace.

Seeing is believing

If you’re new to Dynatrace, be sure to sign up for the Dynatrace free trial. If you’re already a Dynatrace customer, sign in to your account and experience how you can boost your external analyses and reports by leveraging these powerful new Davis AIOps capabilities.

The post Leverage the new Problems API to resolve Dynatrace-detected issues faster in your third-party tools appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/leverage-the-new-problems-api-to-resolve-dynatrace-detected-issues-faster-in-your-third-party-tools/feed/ 0
Intelligent, context-aware AI analytics for all your custom metrics https://www.dynatrace.com/news/blog/intelligent-context-aware-ai-analytics-for-all-your-custom-metrics/ https://www.dynatrace.com/news/blog/intelligent-context-aware-ai-analytics-for-all-your-custom-metrics/#respond Wed, 07 Oct 2020 16:02:08 +0000 https://www.dynatrace.com/news/?p=40174 Dynatrace Dashboard

For custom metrics ingested into Dynatrace via open API interfaces like StatsD, Telegraf, and Prometheus, you can now take advantage of the full power of Davis AI topology-aware anomaly detection and alerting.

The post Intelligent, context-aware AI analytics for all your custom metrics appeared first on Dynatrace news.

]]>
Dynatrace Dashboard

Dynatrace recently opened up the enterprise-grade functionalities of Dynatrace OneAgent to all the data needed for observability, including metrics, events, logs, traces, and topology data. Our breakthrough in augmenting open API interfaces like StatsD, Telegraf, and Prometheus now allows customers to feed third-party metrics into Dynatrace and map those metrics into our real-time Smartscape topology.

As your organization moves beyond the myriad of out-of-the-box technologies that are offered by Dynatrace and you begin to stream in third-party metrics, you need to apply the full power of the Dynatrace Davis® AI causation engine to these ingested metrics—think dependency detection, topology visualization, anomaly detection, auto-baselines, root cause analysis, or even business-impact analysis. This is exactly what Dynatrace now delivers.

Davis topology-aware anomaly detection and alerting for your custom metrics

We’re happy to announce that with the latest Dynatrace release, you can leverage the full power of Dynatrace Davis AI to detect and receive alerts on anomalies in your custom metrics. This allows you to:

  • Use auto-adaptive baselines for all your custom metrics.
  • Seamlessly report and be alerted on topology-related custom metrics.
  • Seamlessly report and be alerted on non-topology-related custom metrics, using Dynatrace as a metric database.
  • Convert non-topological custom metrics into topological metrics on the fly simply by adding semantic links to Smartscape topology.

Topology and non-topology metrics—what’s the difference?

Before diving deeper into anomaly detection for all custom metrics, let’s review the fundamentals. What’s the difference between topology and non-topology metrics in Dynatrace?

  • Topology metrics are related to specific entities in your Smartscape topology (for example, the number of successful and failed batch jobs processed by a host).
  • Non-topology metrics are not related to any Smartscape entity (for example, a retailer’s revenue numbers per store). Instead, the metric is related to the monitored environment as a whole.

Smartscape auto-detected topology is an important differentiator of the Dynatrace Software Intelligence Platform as compared to any other legacy monitoring solution. The Smartscape entity model plays an important role for Davis AI, as all built-in metrics are automatically linked to context-rich entities such as hosts, disks, processes, or services.

A topological link to an entity only makes sense, of course, if the measurement that’s sent to Dynatrace has a semantic relationship to that entity. This means that if a measurement is sent for a host, it must be logically linked to that specific host. The same is true for measurements that are sent for services or applications.

Choose your custom metric type

While, in the past, it was only possible to stream third-party metrics into Dynatrace through a custom device API (i.e., an entity), Dynatrace now also supports use cases for reporting, charting, and alerting on non-topological metrics.

When streaming custom metrics into your Dynatrace monitoring environment, you can now specify whether or not a metric has a topological relationship. Either way, you’re now able to seamlessly report, chart, and alert on these metrics, which fulfills a wide array of use cases across your organization that rely on time series metrics and alerting.

Now let’s take a look at anomaly detection for topology-related and non-topology-related custom metrics in action!

Let’s assume that you have an existing OneAgent instance running on a host and you want to stream measurements for the number of successful and failed batch jobs into your Dynatrace monitoring environment.

OneAgent comes with a new metric ingest channel already enabled. You can use a simple curl command to pipe these metrics into Dynatrace. Representative incoming measurements for each are shown below:

$ curl -d "batchjobs.execution.successes,jobname=payslip 5" http://127.0.0.1:14499/metrics/ingest
$ curl -d "batchjobs.execution.fails,jobname=payslip 1" http://127.0.0.1:14499/metrics/ingest

Ingest data via OneAgent rather than our REST API

Ingesting custom metrics through the OneAgent channel comes with a two major benefits as compared to using the same channel via the REST API:

  • Unlike the REST ingest channel, you don’t need an API token; OneAgent handles the secure connection for you.
  • Each OneAgent instance is already aware of the topology that it’s reporting on, so information about related hosts is automatically added to your ingested metrics.

Once you begin sending the two metrics through the OneAgent channel, they will automatically appear within the metric picker (shown below).

Metric picker for a custom chart showing ingested topology metrics

As mentioned above, each OneAgent instance adds its own topological information to each measurement sent to Dynatrace. You can see this in the image below where metrics have been split by host. This metric dimension was automatically added by OneAgent.

The job name is another dimension automatically reported by OneAgent for the two batch job metrics in this example. All metric dimensions, whether you report them or OneAgent adds them to enrich the topological information, can be transparently filtered and drilled into, as shown below.

Ingested custom metric with topological dimensions

Now let’s assume that we want Davis to trigger an alert whenever an anomaly is detected in the number of failed batch jobs. For this, we go to Settings > Anomaly detection > Custom events for alerting where we can select the metric for the number of failed batch jobs (batchjobs.execution.fails) using the metric picker.

Choose your monitoring strategy (i.e., either a Static threshold or an Auto-adaptive baseline), and define the event title and description for the resulting alert.

Set up alerts for ingested custom topology metrics

Our latest innovation for detecting anomalies in metrics, topology-aware Davis-AI auto-adaptive baselining, is unique in that it adapts to changing metric behavior over time, thereby helping you to avoid false-positive alerts.

Once configured, this event will be raised whenever an anomaly is detected in the number of failed batch jobs. As the metric is topology aware, it has a logical link to the host it is reported for. The event will be raised on this host.

Davis AI root cause detection is triggered based on your chosen event Severity level. Refer to Dynatrace Help to learn about which severity levels trigger Davis and which raise problems.

Now let’s see how Dynatrace ingests and alerts on non-topological metrics, which don’t have logical relationships with Smartscape entities.

Because you can now seamlessly report non-topological metrics, you can now use Dynatrace as a metric database. This gives you all the benefits of a metric storage system, including exploring and charting metrics, building dashboards, and alerting on anomalies.

Let’s take the example of a globally distributed retailer that collects revenue measurements every minute for all its shops worldwide. Revenue per shop isn’t really connected to any topological Smartscape entity, so we skip the association and simply stream the metric into Dynatrace.

Each shop sends its revenue measurements enriched with information about its region, country, and city.

See sample measurements below as they stream into Dynatrace (you can find the complete example on Github).

business.shop.revenue,country=us,region=useast,city=Charlotte,store=shop1 60
business.shop.revenue,country=us,region=useast,city=Jacksonville,store=shop2 24
business.shop.revenue,country=us,region=useast,city=Indianapolis,store=shop3 47
business.shop.revenue,country=us,region=useast,city=Columbus,store=shop4 44
business.shop.revenue,country=us,region=useast,city=NewYork,store=shop5 65
business.shop.revenue,country=us,region=uswest,city=SanFrancisco,store=shop6 95
business.shop.revenue,country=us,region=uswest,city=Seatle,store=shop7 83
business.shop.revenue,country=us,region=uswest,city=SanDiego,store=shop8 100
business.shop.revenue,country=us,region=uswest,city=Portland,store=shop9 100
business.shop.revenue,country=us,region=uswest,city=Anaheim,store=shop10 106

Once all the shops begin reporting their revenue, you can explore the data by slicing and dicing it based on multiple dimensions. The image below shows shop revenue by city.

Now, let’s set up a basic alert for one of these metrics: Go to global Settings > Anomaly detection > Custom events for alerting.

Choose the business.shop.revenue metric and select a dimension value, such as Anaheim, that you want to be alerted on if revenue drops for that city. Here, too, you can select a threshold (Monitoring strategy) and provide a name and description for the alert. Note in the image below that a Static threshold of 100 has been configured to immediately force an alert (for demonstration purposes).

Here’s the completed alert configuration for shops located in Anaheim:

Non-topological metric alert configuration

This intentionally low threshold generates an event and an alert within a few minutes, as shown below. With non-topological metrics, you have all the benefits in Dynatrace that a metric storage system can provide, such as exploring and charting metrics, building dashboards, and alerting on anomalies.

It’s important to note that the alert above was raised at the environment level, as no topological entity is linked to the incoming business metric.

What if a topological connection makes sense after all?

That’s easy! Say that, after some time, you discover that a topological connection to an entity (such as an application or a purchase service) makes perfect sense for a business metric. The benefit of the new Dynatrace metric ingestion functionality is the flexibility you get in adding semantic links to Smartscape topology on the fly. You can do this by adding dimensions to your measurements.

For example, let’s assume that you want to link all the business measurements in the example above to an existing application, easyTravel.

We can link a measurement to an application (by ID) in the job/script sending the metric to Dynatrace. This simply adds the reserved dimension dt.entity.application to the metric stream. Depending on your use case, you might want to link different applications to each individual shop’s business measurement or use one application for all, as we’ve done in the example below. The added dimension is highlighted in each measurement.

business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=Charlotte,store=shop1 73
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=Jacksonville,store=shop2 60
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=Indianapolis,store=shop3 90
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=Columbus,store=shop4 42
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=NewYork,store=shop5 42
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=SanFrancisco,store=shop6 106
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=Seatle,store=shop7 122
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=SanDiego,store=shop8 137
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=Portland,store=shop9 98
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=Anaheim,store=shop10 120

Now, if we set up the same custom alert as shown above, we’ll get a topologically enriched event and alert. The event will therefore be raised with reference to the specific application instead of the entire environment.

This also means that if you choose a dedicated event severity level, your configured event will be fully Davis enabled and trigger root cause detection on the auto-discovered topology.

See the Custom info metric event below that was raised for the easyTravel application based on an anomaly that was detected in the business metric for stores in Anaheim.

Info-level event for an application connected to an ingested custom metric

If you selected the Error severity level in the alert configuration, you will also get an alert and Davis root cause analysis will be triggered for your connected easyTravel application, as shown below:

Problem generated with Error severity level for ingested custom metric

Summary

Dynatrace has achieved a breakthrough in augmenting open API interfaces like StatsD, Telegraf, and Prometheus by allowing you to feed third-party data from these sources into Dynatrace and map the metrics into real-time Smartscape topology.

As you stream in third-party metrics, you now have the full power of Davis AI on these metrics—topology visualization, anomaly detection, auto-baselining, root cause analysis, and even business-impact analysis. You can either ingest these metrics with no topological connections (using Dynatrace as a metric storage system) or you can enrich the incoming metrics with semantic links to your autodiscovered Smartscape topology model.

With this advancement, Dynatrace is now the data-to-answers-to-actions processing engine of choice that relieves you of the burden of manual health and performance analysis. By leveraging automation over existing data sources, Davis AI enables proven, state-of-the-art AIOps, including auto-remediation workflows.

The post Intelligent, context-aware AI analytics for all your custom metrics appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/intelligent-context-aware-ai-analytics-for-all-your-custom-metrics/feed/ 0
Dynatrace innovates again with the release of topology-driven auto-adaptive metric baselines https://www.dynatrace.com/news/blog/dynatrace-innovates-again-with-the-release-of-topology-driven-auto-adaptive-metric-baselines/ https://www.dynatrace.com/news/blog/dynatrace-innovates-again-with-the-release-of-topology-driven-auto-adaptive-metric-baselines/#respond Thu, 23 Jul 2020 18:15:21 +0000 https://www.dynatrace.com/news/?p=38860 Root cause; root cause analysis

With Dynatrace version 1.198, all Dynatrace metrics (built-in as well as custom) can now use auto-adaptive baselines rather than static thresholds for root cause analysis and alerting. Auto-adaptive baselines adapt to changing metric behavior over time, thereby avoiding false-positive alert spam.

The post Dynatrace innovates again with the release of topology-driven auto-adaptive metric baselines appeared first on Dynatrace news.

]]>
Root cause; root cause analysis

While most monitoring solutions still rely on manual thresholds or baselines to identify the root causes of detected issues, Dynatrace Davis® AI has proven over the past four years that a fully automated approach to problem analysis is the only valid approach—especially in highly dynamic, web-scale cloud environments where manual root cause analysis is impossible. With the advent and ingestion of thousands of custom metrics into Dynatrace, we’ve once again pushed the boundaries of automatic, AI-based root cause analysis with the introduction of auto-adaptive baselines as a foundational concept for Dynatrace topology-driven timeseries measurements.

Dynatrace, as the leading software intelligence platform (Gartner Magic Quadrant), has been known for many years for its unique auto-adaptive, multidimensional baselining in the APM space.

Today we’re happy to announce, that with the release of Dynatrace version 1.198 (SaaS and Managed), auto-adaptive baseline extends beyond application performance (APM) metrics to include thousands of infrastructure and cloud metrics as well.

AI-powered root cause analysis via auto-adaptive baselines for more than classic APM metrics

Auto-adaptive baselining represents a dynamic approach to baselining where the reference value that is used to detect anomalies changes over time. The main advantage of this over static thresholds is that the reference values dynamically adapt over time and no threshold need be determined up front. In many cases, metric behavior changes over time. To avoid the effort of manually adjusting thousands of static thresholds, adaptive baselines are used.

Static vs. auto-adaptive baselining

Below is a typical example, where an adaptive baseline has a clear advantage over a statically defined threshold. The chart below shows the measured disk write times in milliseconds for a given disk.

Disk write time is a volatile metric that spikes depending on the amount of write pressure the disk faces. Of course, we could define a static threshold for each disk within the IT system. For this example, we’ll set the static threshold to 20 milliseconds at the start of the chart measurement period. Once defined, you see in the chart below that the usage behavior of the disk changes slightly and that a higher threshold would produce fewer false-positive alerts.

Disk write times with static baseline

The auto-adaptive baseline however automatically adapts its reference threshold daily, considering the measurements of the previous week. In the example below, this means that the threshold increased after the metric changed its typical level, and a new reference value is used for alerting:

Auto-adaptive baseline for disk write times

As you can see, the key benefit of an auto-adaptive baseline is that it adapts to changing metric behavior over time, thus avoiding false-positive alert spam.

Seamlessly integrate auto-adaptive baselines into all your metrics

This release extends auto-adaptive baselines to the following generic metric sources, all in the context of Dynatrace Smartscape topology:

  • Built-in OneAgent infrastructure monitoring metrics (host, process, network, etc.)
  • Built-in OneAgent service metrics (request duration, errors, resource consumption, etc.)
  • Calculated service/DEM metrics (revenue numbers, conversions, event counts, etc.)
  • Custom log metrics
  • Synthetic monitor metrics
  • Cloud platform metrics (AWS, Azure, Kubernetes, etc.)
  • VMware integration metrics
  • Your own ingested OneAgent extension metrics
  • Your own ingested ActiveGate extension metrics
  • Your own API ingested custom metrics

The screenshot below shows all the different metric categories for which you can opt into auto-adaptive baselining:

Metric categories for adaptive baselines

FAQs

How does Dynatrace calculate auto-adaptive baselines?

To adapt to changing metric behaviors over time, adaptive baselines must “learn” a new baseline value at regular intervals, based on historical data. When updating the reference value, the data of the last seven days is evaluated. Measurements for each minute are used to calculate the 99th percentile of all measurements, which results in our baseline.

The inter-quantile range between the twenty-fifth and seventy-fifth percentiles (25th–75th) is then used as the signal fluctuation, which can be added to the baseline. By using the number of signal fluctuations (n x signal fluctuations) parameter, you can control how often that inter-quantile range is added on top of the baseline, which results in the actual alerting threshold (see the screenshot of the settings page below).

Important parameters for this baseline model are the sliding time window (the “evaluation window”) that is used to compare the current incoming measurements against the baseline and the number of signal fluctuations that’s used to adjust the sensitivity of alerting. By default, to raise an event, any three minutes out of a sliding window of five minutes must violate your baseline-based threshold. The sliding window can be changed to a maximum of 60 minutes to adjust the sensitivity of alerting in order to avoid events over shorter periods of time.

Calculation of alerts based on a sliding look-back window

The number of signal fluctuations can also be used to adjust the sensitivity of alerting based on the calculated baseline. n times the normal signal fluctuation, defined as the timeseries variance over the past seven days, is added to the learned baseline reference value.

How can I define an auto-adaptive baseline alert to trigger Davis?

Auto-adaptive baselines are seamlessly integrated into the custom event settings of your Dynatrace environment.

Navigate to Settings > Anomaly detection > Custom events for alerting.

Select your metric using the standard metric picker. Filter by metric dimensions and then choose Auto-adaptive baseline as your monitoring strategy:

Auto-adaptive baseline setting for custom events for alerting

Both monitoring strategies (static and adaptive baselines) offer a convenient alert preview that allows you to review the potential number of alerts when using the given settings.

The number of signal fluctuations and the sliding evaluation window for alerting allow you to further fine-tune alerting sensitivity.

Configuration settings also include meta-information such as the title of the resulting event, a textual description, and a severity level.

The following table summarizes the semantics of all available event severities (which severity types trigger a problem and which severities are analyzed by Davis).

Severity Problem raised? Davis analyzed? Semantic
Availability Yes Yes Reports any kind of severe component outage.
Error Yes Yes Reports any kind of degradation of operational health due to errors.
Slowdown Yes Yes Reports a slowdown of an IT component.
Resource Yes Yes Reports a lack of resources or a resource-conflict situation.
Info No Yes Reports any kind of interesting situation with a component, such as a deployment change.
Custom alert Yes No Triggers an alert without causation and Davis AI involvement.

Can I use this feature in my own automation scripts?

The complete configuration of custom events for alerting along with its auto-adaptive baseline monitoring strategy can be configured using our Configuration API. See the relevant Anomaly detection (Metric events) API endpoints below.

Anomaly detection (metrics) API endpoints

Do I require a special license to use auto-adaptive baselines?

Auto-adaptive baselines are shipped as core functionality within your Dynatrace monitoring environment, which means that they are not bound to any additional cost or license.

In Dynatrace environments, a technical limit of up to 100 baseline configurations is enforced. Each configuration can baseline up to 100 metric dimensions. Whenever a configuration reaches that limit, an alert is raised on the environment level and the related configuration is disabled.

Summary

With the advent and ingestion of thousands of custom metrics into Dynatrace, we’ve again pushed the boundaries of automatic, AI-based root cause analysis and introduced auto-adaptive baselines as a foundational concept.

With this release, you can define an auto-adaptive baseline on any kind of metric, whether it’s a built-in OneAgent metric or one of your own custom ingested metrics. Auto-adaptive baselines are a great monitoring strategy for triggering the Davis AI to provide deep root-cause analysis. Also, they enable infrastructure monitoring use cases where static thresholds would trigger too many false-positive alerts due to changes in metric behavior.

For more details, see our Auto-adaptive baselining for custom metric events documentation.

The post Dynatrace innovates again with the release of topology-driven auto-adaptive metric baselines appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-innovates-again-with-the-release-of-topology-driven-auto-adaptive-metric-baselines/feed/ 0
Massively automate enterprise operations using Dynatrace information discovered with the Monitored entities API v2 https://www.dynatrace.com/news/blog/massively-automate-enterprise-operations-using-dynatrace-information-discovered-with-the-monitored-entities-api-v2/ https://www.dynatrace.com/news/blog/massively-automate-enterprise-operations-using-dynatrace-information-discovered-with-the-monitored-entities-api-v2/#respond Fri, 19 Jun 2020 15:25:26 +0000 https://www.dynatrace.com/news/?p=38071 Automation graphic

Dynatrace is happy to announce the improved Monitored entities REST API v2. Use it to flow Dynatrace context information into third-party integrations and massively automate business-critical processes like deployment pipelines, CI toolchains, and cloud migration projects.

The post Massively automate enterprise operations using Dynatrace information discovered with the Monitored entities API v2 appeared first on Dynatrace news.

]]>
Automation graphic

The world is moving faster than ever. To drive productivity and boost efficiency, enterprises are increasingly automating their critical business processes. At the same time, navigating modern IT infrastructures can seem like stumbling through a maze, as they’re becoming overly complex, containing hundreds of technologies and billions of dependencies. So how can your Operations teams automate processes such as CI/CD deployment pipelines, configuration changes of ITSM tools, or the simple generation of a monthly architectural topology report if they’re struggling to find critical information about your highly complex environments? With Dynatrace, all this can be done easily.

One of the many ways our customers can cut through this ever-growing complexity is by using the Dynatrace Monitored entities API v2. This version 2 API helps you integrate Dynatrace context information into all kinds of automation workflows, from cloud migration projects, through deployment pipelines, and CI tool chains.

Read on to learn more about the improved entity API and see the implementation of a few use cases.

Automate your CI toolchain, migrate to the cloud, and more with the Dynatrace Monitored entities API v2

With its patented Smartscape technology, Dynatrace offers a unique way of visualizing all your running IT services and their critical relationships. As Dynatrace OneAgent discovers all the entities and dependencies within your application environment, Smartscape simultaneously builds an interactive map of how everything is interconnected. Every service is overlaid with real-time performance and stability measurements, which provide critical information for your Operations teams.

By leveraging the Monitored entities API, you can empower third-party integrations to use and integrate Dynatrace context information into all kinds of automation workflows. A few of the customer use cases we’ve seen so far include:

  • Accelerating cloud migration by measuring the scale of complex applications
  • Automatically keeping track of business-critical services in the existing CMDB (ServiceNow, Atlassian, and others)
  • Making design decisions based on size and critical dependencies of applications

Smartscape - hosts

What’s new in the Monitored entities v2 REST API (Early Adopter release)

To further enable the efficient use of Smartscape-provided information with external integrations, we’re happy to announce improved Monitored entities API endpoints. The following major features are available in this v2 REST API:

  • One single endpoint to fetch information about any exported entity type

By having a single endpoint, querying for different entity types such as hosts, services, processes, and applications becomes much simpler.

  • More entity types available for export

With this improvement, you can now also export disk, data center, and Docker components to your third-party CMDB.

  • New API endpoint for getting the schema (type) of a given component type

This allows your integrations to automatically check for new attributes and types.

  • Consistent pagination, entity and timeframe selection across all v2 API endpoints

The consistent use of pagination and entity and timeframe selection across all the v2 API endpoints allows for scale within your web-scale monitoring environments.

  • Full control over the resulting payload and information

By having full control over the resulting payloads and information, you can minimize network traffic at your data centers when fetching Dynatrace data.

  • Common UI backlink

We’ve added a single navigation URL for redirections back to entity pages within the Dynatrace web UI, so that you can build your own reporting UIs with direct backlinks to the relevant Dynatrace entities.

Here’s the family of entity API endpoints that are newly available in the version 2 REST API:

Monitored entities endpoints in the API Explorer

Monitored entities API use cases

The examples below guide you through the implementation of use cases featuring the Monitored entities API. We start with a simple example and then iteratively implement more complex use cases.

Get an initial overview of the types of entities available

This first example examines the different kinds of entities that are exposed by the entity API and the kinds of properties and relationships each entity type has. The following endpoints provide overviews of the Dynatrace entity model in an automated fashion.

Entity types GET endpoints

As an example, we’ll call the /api/v2/entityTypes endpoint to get the list of entity types that are available. The resulting JSON payload, as shown below, returns 22 individual entity types along with all their properties. The list endpoint exposes many entity types, such as hosts, applications, services, processes, custom devices, disks, VCenters, and more.

JSON response showing all entity types

If we focus on a specific entity type, such as hosts, in /api/v2/entityTypes/{type}, we can see all the relevant properties, tags, and management zone information that Dynatrace collects for the hosts and the relationships that these hosts have with other parts of the topology:

JSON payload for a specific entity type - hosts

Get a filtered list of processes

Now let’s switch to the endpoints that expose instances of identified entities monitored by Dynatrace within your IT environment.

Use the /api/v2/entities endpoint with entitySelector=type(PROCESS_GROUP_INSTANCE) to receive the list of processes running in your environment:

Get instances of specific entities using the entity selector

The API call immediately returns all processes running in the Dynatrace environment. As the total number of monitored processes is 50593, the API call does not return everything in a single result, instead returning the first page (the default size is 100 time series per page, which can be changed) of processes along with a nextPageKey that you can use to get the next page of results.

The resulting payload returns the entityId and displayName for each process by default.

It’s likely that you’re not interested in all 50,000+ processes and just want to focus on a subset instead.

The entitySelector request parameter can again be used to refine your query for processes. For example, you can modify the query for processes within a given management zone:

entitySelector=type(PROCESS_GROUP_INSTANCE),mz(Easytravel)

In addition, let’s filter for processes that have the loadtest tag assigned and where the displayName contains the string easy:

entitySelector=type(PROCESS_GROUP_INSTANCE),mz(Easytravel),tag(loadtest),entityName(easy)

This immediately filters the 50,000+ processes down to a reasonable number of five processes (see below).

Get entity instances using entity selector for display name, tag, and management zone

Take control of the resulting payload

Next, you can select which information you’d like to include in the results payload.

By default, the API only returns the entityId and displayName to keep the payload as small as possible. Use the fields request parameter to control what else you’d like to get in the results payload.

Say you want to get the tags for each process. Simply add fields=+tags to your API call. Here’s what the result looks like after requesting tags with the fields parameter:

Entity instance results modified with the fields parameter

You can add multiple attributes with the fields parameter using the syntax shown below:

fields=+tags,+properties

properties represents grouped attributes that can be added at a group level, as shown above. If the API consumer is interested in a subset of properties, you can add those individually, as shown below:

fields=+tags,+properties.appVersion

The resulting payload:

Use relationships to query entities

In many cases, API consumers might want to filter a request based on a given relationship with either a parent entity or a dependent entity within the monitored topology. Popular examples of this are fetching all the disks of a given parent host or fetching all processes that run on a given host.

The newly introduced relationship query functionality via the entitySelector allows you to follow each relationship type, incoming or outgoing, of a given entity.

Relationships are generally split into two major categories (from and to). These categories denote the direction in which a relationship applies to an entity. In most cases, the direction of the relationship comes with a specific semantic. If you query all the services that service A calls, this is a “from (service A)” relationship to other services. When you query all the services that call service A, this a “to (service A)” relationship.

The example below shows how to fetch all processes running on a given host:

entitySelector=type(PROCESS_GROUP_INSTANCE),toRelationship.isProcessOf(entityId(HOST-3052A99AAD944176))

If you want to select an entity by name rather than by its unique ID, you can do so using entityName instead of the entityId query along with the type query, as shown below:

entitySelector=type(PROCESS_GROUP_INSTANCE),toRelationship.isProcessOf(type(HOST),entityName(myHost))

Relationship query functionality also enables API consumers to follow a service call relationship in the same way they would query all the disks or processes of a given host.

How to find an entity within the Dynatrace Web UI

Many integrations would like to offer backlinks to dedicated entity pages in Dynatrace to save users time as they switch to and from Dynatrace. So this section of the blog post is dedicated to the related question we hear a lot from our user community: “How do you find an entity’s Dynatrace web UI page if you’ve retrieved its unique Dynatrace ID using the API?”

As Dynatrace changes and improves at a fast pace, the UI structure and navigation paths can change over time, which is cumbersome if you want to offer a consistent backlink to a given type of entity.

So instead of offering backlinks as part of the API, we’ve decided to provide a common redirection address format for all entities. You can use the common entity navigation URL within your Dynatrace environment to jump to any entity simply by using its unique ID.

<YOUR_DYNATRACE_ENVIRONMENT_URL>/ui/nav/<ENTITY_ID>

To give you a concrete example, let’s assume your Dynatrace environment is https://d2334sdfssdf.live.dynatrace.com. If you want to visit the UI page of your host with the unique ID HOST-AF82014F952CF161, navigate to this URL, which is based on the standard format described above:

https://d2334sdfssdf.live.dynatrace.com/ui/nav/HOST-AF82014F952CF161

The same is true for any other entity type, such as a service:

https://d2334sdfssdf.live.dynatrace.com/ui/nav/SERVICE-3282014F952CF163

So, no matter how the Dynatrace UI may change in the future, your backlinks will always lead to the most appropriate Dynatrace UI pages for your entities. In case there’s no UI page, the entity navigation URL will lead you to the next closest page so you can find the entity. In the case of a disk, for example, you’ll navigate to the disk’s host page.

Summary

Modern enterprises depend more than ever on highly automated processes to drive productivity, boost efficiency, and stay competitive. Dynatrace provides all the necessary API power to further automate your toolchains and processes to help your teams handle complexity.

The newly introduced Monitored entities API version 2 offers a lot of improvements over its predecessor. It further empowers third-party integrations to use and integrate Dynatrace context information into all kinds of automation workflows such as your deployment pipeline, CI toolchain, or a simple monthly architecture-topology report.

Relevant material

The post Massively automate enterprise operations using Dynatrace information discovered with the Monitored entities API v2 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/massively-automate-enterprise-operations-using-dynatrace-information-discovered-with-the-monitored-entities-api-v2/feed/ 0
Get quick alerts and avoid false positives with the new baseline setting https://www.dynatrace.com/news/blog/get-quick-alerts-and-avoid-false-positives-with-the-new-baseline-setting/ https://www.dynatrace.com/news/blog/get-quick-alerts-and-avoid-false-positives-with-the-new-baseline-setting/#respond Thu, 26 Mar 2020 18:18:52 +0000 https://www.dynatrace.com/news/?p=36432 Dynatrace employees

Our newly introduced service baseline settings allow you to adapt alerting based on your own needs and according to each service's criticality.

The post Get quick alerts and avoid false positives with the new baseline setting appeared first on Dynatrace news.

]]>
Dynatrace employees

We’re happy to announce that with Dynatrace version 1.189, you can give your baselining routines more time to evaluate short-lived performance conditions. This avoids annoying false-positive alerts on short spikes while still alerting you to conditions that require your attention. Read on for an example and description of our new baselining functionality.

What do stock markets and monitoring alerts have in common?

Hint 1: There comes a point in time when a decision must be made.

Hint 2: In retrospect, it’s easy to see if you should have bought or sold (i.e., been alerted or not).

While stock market decisions are based on market observation, alerts and decisions in software monitoring are mostly based on learned baselines. Both in stock markets and in software monitoring, the observation period and the point at which decisions are made are crucial.

How does a longer observation period help avoid false-positive service-baseline alerts? The screenshot below shows a monitored service that suddenly experiences a much higher error rate compared to the learned baseline.

Increasing error rate for a monitored service

Imagine that you need to make a decision as to whether you need to wake the Ops team in the middle of the night. Let’s say that you decide to wait another five minutes to monitor the situation. Waiting an additional five minutes reveals that the issue was just a short hiccup caused by a client that used an outdated service parameter. The situation quickly settles back into the learned baseline, as shown below:

Monitored service with temporary increase in error rate

As you can see, efficient monitoring is a matter of balance between quick reaction when necessary and avoiding overreacting to short hiccups.

Give your baselining more observation time to avoid false positives

Dynatrace monitors all your services with an out-of-the-box automatic baseline, which immediately learns each service’s typical behavior and alerts on abnormal situations. The automatic-baseline approach has many benefits, such as immediate insights into all dynamic microservices.

Our newly introduced service baseline settings allow you to adapt alerting based on your own needs and according to each service’s criticality. By default, Dynatrace service baselining evaluates the criticality of each situation by taking into account statistical confidence within one to five-minute intervals. This means that Dynatrace alerts more quickly when an error spike occurs in a high-traffic service (compared to a low-traffic service where statistical confidence is lower).

But it’s difficult, or even impossible, for Dynatrace to automatically detect how critical a service is to the success of your digital business. While you might want to wait a bit longer before alerts are sent out for non-critical, low-load services, you might want to receive immediate alerts for changes in the performance of your most critical services, even if you know that such alerts have a high rate of false positives.

See the Anomaly detection for services baseline configuration settings page below where we’ve added two new settings:

Service baseline configuration screen

Here are some pointers on using the new settings:

  • By default, the minimum observation period in Dynatrace is one minute. To avoid false-positive alerts on your services, you can add more time.
  • By setting the configuration to 15 minutes globally for all your services, you can avoid false-positive alerts on short spikes.
  • The configuration is available at the global level as well as the service level. This allows you to override global settings with a stricter setting for critical services.
  • This setting distinguishes between slowdown alerts and error detection, so that you can choose an individual strategy for each.

Close issues sooner with shorter event timeouts

If a problem has a long timeout (the time a problem stays open before being dismissed), for example, 15 minutes, in a low-traffic situation, you can’t really suppress short spikes because all spikes will be reported as 15-minute duration problems.

To improve this, in the latest release, we’ve reduced the default timeout for low-load events from 15 minutes down to 5 minutes. The rationale behind this is that low-load events may consist of no more than a handful or errors at best; so it makes sense to reduce the problem timeout. As a result, you’ll no longer see low-load events that are kept open for more than five minutes following any detected spike. The reason for having a timeout period at all is to avoid reopening events when multiple spikes follow in sequence.

The screenshot below illustrates the timeout reduction for better understanding:

Low-load event timeout reduction

Summary

Efficient monitoring takes place when you can maintain a balance between quick alerting on critical situations and eliminating false positives.

It might sound simple, but it’s tremendously tricky to distinguish between a critical service alert and a false-positive situation. Increasing the observation window can significantly reduce alert spam on all your non-critical services, while preserving quick alerts for a handful of highly critical services.

Just like in the stock market where you need to make the decision to buy or sell in a falling market, the decision must be made at a single point in time. Having the option to observe the market for an extra day or week makes it easier to see when the decision should be made.

The post Get quick alerts and avoid false positives with the new baseline setting appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/get-quick-alerts-and-avoid-false-positives-with-the-new-baseline-setting/feed/ 0
Additional IP addresses for public alert notifications (customer action required) https://www.dynatrace.com/news/blog/additional-ip-addresses-for-public-alert-notifications-customer-action-required/ https://www.dynatrace.com/news/blog/additional-ip-addresses-for-public-alert-notifications-customer-action-required/#respond Thu, 16 Jan 2020 17:17:02 +0000 https://www.dynatrace.com/news/?p=35140 Dynatrace employee

Alerting is a crucial aspect of monitoring your web application. Alerts warn you when your web application doesn't meet performance standards so you can respond quickly and fix the issue. To scale up and further guarantee the delivery of all your Dynatrace alerts, we're adding new public IP addresses.

The post Additional IP addresses for public alert notifications (customer action required) appeared first on Dynatrace news.

]]>
Dynatrace employee

Please note that this configuration change will affect all our Dynatrace public alert notification IP addresses.

Your action is required: Add new public IP addresses to Allow list

Our public IP addresses are used by Dynatrace SaaS to send out Dynatrace-detected problem notifications through your configured alerting channels (Webhook, ServiceNow, etc).

Action required: If you’ve defined a firewall rule to allow incoming alerts from the Dynatrace public IP addresses, please also add our newly introduced public IP addresses to the firewall rule.

Am I affected?

This configuration change affects all Dynatrace SaaS users who’ve configured any alerting channel in Dynatrace other than email (which isn’t affected).

When will the change be effective?

The change in IP addresses will take place in the first week of March 2020! To avoid an alert outage, you must add the new IP addresses to your firewall allow list before March 2020.

What can happen if I don’t change my firewall configuration?

If you don’t add the new Dynatrace public IP addresses to your firewall, all incoming Dynatrace alerts will be blocked by your own firewall.

Ok, what should I do?

To be on the safe side, simply add the current as well as the newly introduced public Dynatrace IP addresses for your environment region to your firewall’s allowed addresses list.

Current public IP addresses

Currently, Dynatrace uses the following public IP addresses in the different AWS regions to send out notifications to your systems whenever a problem is detected:

North Virginia [us-east-1]:

  • 52.0.97.215

Oregon [us-west-2]:

  • 52.26.178.159

Ireland [eu-west-1]:

  • 54.154.238.149

Sydney [ap-southeast-2]:

  • 54.66.240.72

Newly added public IP addresses (these will replace the current ones for each region)

North Virginia [us-east-1]

  • 54.147.40.70
  • 54.236.238.11
  • 3.231.89.157

Oregon [us-west-2]

  • 100.21.193.142
  • 100.21.71.51
  • 35.163.122.249

Ireland [eu-west-1]

  • 34.253.6.125
  • 34.248.38.224
  • 52.19.116.240

Sydney [ap-southeast-2]

  • 13.210.148.174
  • 3.105.21.198
  • 52.62.189.123

Additional ways we’re announcing this news

This important message is also shown within in-product notifications for all environment admins, sent out to your Dynatrace One representative, and reported in advance in our regular release notes.

The post Additional IP addresses for public alert notifications (customer action required) appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/additional-ip-addresses-for-public-alert-notifications-customer-action-required/feed/ 0
Leverage the power of Davis AI with custom time-series events for your specific use cases https://www.dynatrace.com/news/blog/leverage-the-power-of-davis-ai-with-custom-time-series-events-for-your-specific-use-cases-preview/ https://www.dynatrace.com/news/blog/leverage-the-power-of-davis-ai-with-custom-time-series-events-for-your-specific-use-cases-preview/#respond Wed, 04 Sep 2019 16:37:06 +0000 https://www.dynatrace.com/news/?p=33350 Delivering excellent digital experience for customers in a complex digital world

By defining metric-based custom events, you can leverage the power of Davis AI for your specific use cases.

The post Leverage the power of Davis AI with custom time-series events for your specific use cases appeared first on Dynatrace news.

]]>
Delivering excellent digital experience for customers in a complex digital world

Dynatrace Davis® automatically analyzes abnormal situations within your IT infrastructure and reports all relevant impacts and root causes.

Davis relies on a wide spectrum of information sources, including a transactional view of your services and applications and the monitoring of all events that are raised on individual nodes within your Smartscape topology map.

There are two main sources for individual events in Dynatrace: (1) Events that are triggered by a series of measurements (i.e, metric-based events) and events that are independent of any metric (for example, process crashes, deployment changes, and VM motion events).

This blog post focuses on the definition of events that are triggered by measurements (i.e, metric-based events) within your Dynatrace monitoring environment.

Create custom events for time-series measures that trigger Davis problem analysis or provide context for Davis-analyzed situations

By defining metric-based events, you can leverage the power of Davis AI for your specific use cases. For example, you can:

  • Define your own service metric for revenue generation and define a metric-based event that will alert you if your generated revenue drops below a critical level within a given observation period.
  • Annotate any of your hosts, applications, or services with an info-level event that indicates an interesting metric level, such as an unusually high number of users or service calls. Info-level events don’t trigger alerts. They are however monitored and reported by Davis in case a related problem is detected.
  • Collect your own custom metric, such as the number of reports processed and written to a local folder, and raise an event if the number of processed reports falls below a specific threshold.
By selecting the severity of an event, you define whether or not a problem should be raised by Davis AI or if you’re fine with just an alert or an info event being logged. Dynatrace even simulates the results of your provided settings with real data pulled from the last seven days, showing you how many alerts and what types of alerts will be raised with your new custom settings.
In this blog post, you’ll see how Davis analyzes problems and how you can define custom events that can either trigger deeper Davis AI analysis or simply contribute additional contextual information that Davis AI can use for related analysis tasks.

How to set up custom events

Custom metric events are configured globally at the environment level. This means that an event, once defined and raised, will be visible to all Dynatrace users within your environment.

Open Settings > Anomaly detection > Custom events for alerting to define a new custom metric event, as shown below.

Where to define a custom event for alerting

Click Create custom event for alerting to open the configuration screen that allows you to define a metric-based event in a step-by-step process.

Choose the title and severity of your event

The event title is a short, human-readable string that describes the situation in an easily comprehensible way. Examples here could be, “High network activity” or “CPU saturation.” Event titles are used throughout the Dynatrace UI as well as within problem alerting.

Custom event title

After assigning a proper name for your custom event, you must specify its severity. The severity of an event determines if a problem is raised or not and if Davis AI should try to find a root cause of the event.

Severity levels for custom events

The following table summarizes the semantics of all available event severities that trigger problems and are analyzed by Davis:

Severity Problem raised? Davis analyzed? Description
Availability Yes Yes Indicates an outage situation
Error Yes Yes Indicates a degradation of operational health
Slowdown Yes Yes Indicates a slowdown of IT services.
Resource Yes Yes Indicates a lack of resources
Info No Yes Informational only
Custom alert Yes No Alert without correlation and Davis logic

Read more about built-in events and their severity levels in our Event types documentation.

In the example below, we’ll call a new event called Critical network packet loss that has the severity level Error:

Example custom event title and severity level

Choose the metric to monitor

One of the most essential aspects of configuring a new metric-based event is the selection of the metric and, optionally, the metric dimension that is to be monitored.

The metric picker offers a structured list of all available metrics within your Dynatrace monitoring environment. The structure of the metric categories is the same as that used when creating a custom chart on a dashboard or when requesting a metric through the metric API. For more details, see our blog post about metric selection in custom dashboarding and when using the Dynatrace API.

Within the metric picker, you either search for a metric in all categories or you select a category and search within it. The image below shows a metric search across all categories.

Searching for a metric in the metric picker

Alternatively, you select a category and browse the metrics that Dynatrace offers in that category:

Metric categories

For our example event, we’ll select the Network interface sent packets dropped on host network metric, as shown below. We don’t want to filter for a specific metric dimension; in this case, the dimension represents all network interfaces on the selected hosts.

Choosing a custom event metric and dimensions

Define the event scope

The next step is to define the event’s scope as a subset of all the hosts you’re monitoring within your environment.

After selecting a metric, all entities that supply that metric are counted and shown in the preview section. In our monitoring environment, there are more than 100 hosts that supply the selected metric.

The preview presents a maximum of 100 entities. We don’t recommend defining a shared threshold on a huge and heterogeneous collection of entities as this typically results in a high number of alerts.

Choosing entities for alerting scope

Use convenient scope filters such as host group, management zone, name, and tag filters depending on how you organize your entities. In many cases, a naming convention used for all your hosts along with a dedicated name filter is effective for defining alerting scope, as shown below:

Adding a rule-based name filter to define alerting scope

When the number of entities within your alerting scope is fewer than 100, Dynatrace offers a convenient preview of how many alerts would’ve been triggered over the last 12 hours based on the defined scope. Alternatively, you can select and analyze the last day or the last seven days to see how many alerts you would’ve received with historic measurements. The configuration page even provides a baseline threshold suggestion for the selected group of hosts:

In the example above, the configured threshold of 17.8 packet errors per second still results in the quite high number of 65 events during the last 12 hours. To reduce alert spam, you can either increase the threshold baseline to 19 errors per second and/or specify a larger sliding window of 5 out of 10 minutes. Let’s see how that change works in our example:

The last step before you go live with your newly created event is to review the description message. Use the four placeholders {metricname}, {severity}, {alert_condition}, and {threshold} to fill the text message with the actual values.

Configure a description message

Configuring the custom event's description message

Once your event is triggered, this description message will read, for example, as “The Network interface sent packets dropped on host value of 24 was above your custom threshold of 19.” Adapt the event description to your own needs and then save the custom event definition.

See the screenshots below of how our configured event is visualized in the Dynatrace problem feed:

Custom event in Problems page

Problem details for the custom event

Summary

Our improved configuration workflow for custom event alerting offers a lot of power in terms of defining additional metric-based events for your Dynatrace environment. Additional flexibility in selecting event severity adds the power to decide if a problem should be raised, if the Davis AI should look at it, or if you are fine with just a single alert or info event.

The newly introduced threshold baseline recommendation helps to define reasonable thresholds. A dry run with data taken from the last seven days helps you identify the number of alerts that would’ve been raised with specific custom settings.

The post Leverage the power of Davis AI with custom time-series events for your specific use cases appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/leverage-the-power-of-davis-ai-with-custom-time-series-events-for-your-specific-use-cases-preview/feed/ 0
Integrate Dynatrace more easily using the new Metrics REST API https://www.dynatrace.com/news/blog/integrate-dynatrace-more-easily-using-the-new-metrics-rest-api/ https://www.dynatrace.com/news/blog/integrate-dynatrace-more-easily-using-the-new-metrics-rest-api/#respond Fri, 28 Jun 2019 15:37:16 +0000 https://www.dynatrace.com/news/?p=32669 Digital-Transformation

As a full stack monitoring platform, Dynatrace collects a huge number of metrics for each OneAgent monitored host in your environment. Depending on the types of technologies you’re running on your individual hosts, the average number of metrics is about 500 per computational node. Besides all the metrics that originate from your hosts, Dynatrace also […]

The post Integrate Dynatrace more easily using the new Metrics REST API appeared first on Dynatrace news.

]]>
Digital-Transformation

As a full stack monitoring platform, Dynatrace collects a huge number of metrics for each OneAgent monitored host in your environment. Depending on the types of technologies you’re running on your individual hosts, the average number of metrics is about 500 per computational node.

Besides all the metrics that originate from your hosts, Dynatrace also collects all the important key performance metrics for services and real-user monitored applications, as well as cloud platform metrics from AWS, Azure, and Cloud Foundry.

All told, there are thousands of metric types that can be charted and that Davis® automatically analyzes and alerts on within your Dynatrace environment. With the introduction of our new metric API version it gets even easier to fetch an API report of all the abnormal metric behavior that Davis AI detects.

Our API-first design promotes automation

Many use cases within your software development and delivery pipelines depend on the real-time metrics that your Dynatrace environment collects. One example is the automatic check of monthly load-test results for performance reporting based on Dynatrace synthetic tests.

The Dynatrace REST API endpoint /api/v1/timeseries has long enabled API consumers to ingest individual metrics for the implementation of external use-cases. With Dynatrace version 1.172, an updated of our metrics API endpoint (version 2) is now available. The latest version is based on our improved metrics framework, which provides:

  • A logical tree structure for all available metric types
  • Globally unique metric keys that better integrate over multiple Dynatrace environments
  • Flexibility to extend Dynatrace and better fit it to your specific business cases

Discover the new features and enhancements of the Metrics API

Following is an overview of the new enhancements, along with some practical examples for each.

New metric identifiers and structure

One of the most important aspects of the newly introduced metrics API endpoint is the logical structure for all Dynatrace metrics. All existing metrics have received new metric identifiers that can be fetched using the metric descriptor endpoint. The previous metric identifiers can still be used in version 1 of the timeseries API, but the identifiers aren’t valid with version 2 of the Metrics API.

Split of metric descriptor from series endpoints

In the previous version of the Timeseries API endpoint, metric descriptors and timeseries data were mixed into a single endpoint, which led to some confusion.

The new version of the Metrics API splits metric descriptors and timeseries data endpoints into separate endpoints:

  • /api/v2/metrics/series
  • /api/v2/metrics/descriptors

To query all available metric descriptors, call HTTP GET /api/v2/metrics/descriptors to receive the list of all defined metric types within your monitoring environment. For convenience purposes, you can also receive the information as a CSV list or as JSON payload. See an example below:

Screenshot Dynatrace Environment API

The complete list of Dynatrace metrics has been rearchitected based on a logical tree structure that helps you query related metrics.

Use an asterisk (*) as a wildcard character to query for all host and disk related metric definitions, as shown below:

Screenshot Dynatrace

Efficient query strings for selecting one or more metrics

One of the most requested features for the Metrics API was the ability to fetch more than one type of metric within a single API request. Version 2 of the metric series endpoint allows you to query multiple metric types across a set of entities using convenient query strings.

The example below shows a typical API call to request a CPU metric as well as a memory metric for all hosts within an environment:

Screenshot Dynatrace API call

The resulting timeseries is split into the memory and CPU measurements:

Screenshot Dynatrace timeseries

Use query strings to select metric percentiles

The newly introduced query selector string provides a lot more power than just the ability to select metrics. The query string also allows you to transform results by applying an aggregation function (for example, to get a percentile, as shown in the example below).

Screenshot Dynatrace aggregation function

Other useful aggregation functions are sum, avg, min, and max.

Select the resolution of your metrics

Many use cases demand a specific resolution for fetched timeseries, such as hourly resolution, minute resolution, or even the total value during a given time period.

The example below shows how to fetch a CPU metric at hourly resolution:

Screenshot Dynatrace fetch a CPU metric

Need a metric in CSV format?

While automation use cases often require programmatically processable results in JSON format, there is also often a requirement for the direct import of Dynatrace data into Tableau or Excel.

The new API endpoint supports JSON as well as CSV format by specifying the HTTP header Content-Type’ : ‘application/csv.

This convenient feature allows for immediate and seamless import of data into Excel and third-party reporting tools.

The option to export CSV format through the Dynatrace API allows you to directly create an Excel data source that can be updated whenever required by performing a simple refresh. In this way, your Excel charts can remain up to date based on the latest Dynatrace delivered KPIs, as shown below:

curl -X GET "https://<YOUR_URL>/api/v2/metrics/series/builtin:service.response.client:percentile(50)?resolution=h&from=now-2w&to=now" -H "accept: text/csv; header=present; charset=utf-8" -H "Authorization: Api-Token <YOUR_TOKEN>"

Here’s an example set of results in a CSV export that can easily be processed by Excel:

Screenshot Dynatrace CSV export

Convenient date and time selector

Compared to the previous version of our Timeseries API, version 2 offers more convenient date and time selection. The newly introduced date time format for specifying the from and to parameters allows you to choose between millisecond timestamps, human-readable date and time format, and relative times, such as last 2 days.

The following examples show how much easier the selection of a timeseries timespan is with version 2 of the API endpoint:

Receive a metric for the last week:

from=now-2w&to=now-1w

Receive a metric for the timespan between two given dates in the past:

from=2019-12-21T05:57:01.123+01:00&to=2019-12-22T05:57:01.123+01:00

Receive the timestamp in UTC millisecond format:

from=1558963489634

Receive a metric for a given date and time aligned with a specific URC timezone:

from=2019-12-21T05:57:01.123+01:00

In upcoming releases, we’ll begin using this date time format as a consistent means of selecting dates within all our API endpoints.

Flexible filters for focusing on relevant results

One of the most important features of web-scale monitoring solutions like Dynatrace is the ability to efficiently apply filters to reduce result sets.

Efficient filters allow you to reduce the network bandwidth necessary to transfer data as well as to speed up result queries in general.

Whenever possible, use filters to focus your result set on the entities within Dynatrace that you’re interested in, instead of fetching all available data.

While in the past multiple parameters were used to apply filters, in version 2 of our Timeseries API, a combined query string is used.

Some of the most important filter queries are shown within the examples below:

Filter by a given management zone with name easyTravel or use multiple management zone filters:

scope=mz(easyTravel)

scope=mz(easyTravel, production)

Filter by given entities:

scope=entity(HOST-123445678, HOST-23456789)

Filter by entity names:

scope=entity(easy)

Combine multiple filters:

scope=mz(easyTravel, production),name(easy)

Summary

The completely reworked new version of the Dynatrace Metrics REST API comes with many improvements that have been requested by Dynatrace customers over the last 5 years.

The new version offers a lot more convenience in building metric queries and it allows you to efficiently filter huge result sets.

The Metrics API introduces some necessary limits that help you to keep your metric requests stable and not trigger huge queries that could result in Gigabytes of result size.

The redesigned consistent structure of all available metrics within Dynatrace is aligned with recent improvements in custom charting and dashboarding. Custom charting, including API-based charting, and dashboarding now deliver a consistent view of all available metrics. The goal is to offer the same query and filtering features seamlessly across both the Dynatrace web UI and the REST API.

These recent changes further position Dynatrace as a platform that offers the same power in its integration APIs as it does in its web UI.

With the introduction of version 2 of our Timeseries endpoint, we will no longer introduce new features in version 1 of the endpoint. The version 1 endpoint will continue to be supported. We’ll announce a discontinuation date for version 1 at a later time.

Last, but not least, please stay up to date with our progress and changes with the Dynatrace APIs by reviewing our API release notes on a regular basis.

Start a free trial!

Dynatrace is free to use for 15 days! The trial stops automatically, no credit card is required. Just enter your email address, choose your cloud location and install our agent.

The post Integrate Dynatrace more easily using the new Metrics REST API appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/integrate-dynatrace-more-easily-using-the-new-metrics-rest-api/feed/ 0
Ability to manually close problems when justified https://www.dynatrace.com/news/blog/ability-to-manually-close-problems-when-justified/ https://www.dynatrace.com/news/blog/ability-to-manually-close-problems-when-justified/#respond Fri, 09 Nov 2018 17:22:42 +0000 https://www.dynatrace.com/news/?p=29255 Problem dashboard

Within Dynatrace, all problem detection is done automatically. The lifespan of each detected problem is automatically managed by Dynatrace. This means that Dynatrace detects when an unhealthy component recovers and automatically closes the corresponding problem once it receives a resolved status. This is a helpful feature, but sometimes you may want to manually close a […]

The post Ability to manually close problems when justified appeared first on Dynatrace news.

]]>
Problem dashboard

Within Dynatrace, all problem detection is done automatically. The lifespan of each detected problem is automatically managed by Dynatrace. This means that Dynatrace detects when an unhealthy component recovers and automatically closes the corresponding problem once it receives a resolved status. This is a helpful feature, but sometimes you may want to manually close a problem for a good reason. With the latest Dynatrace release, you can now manually close problems and add a closing reason within a textual comment.

Every problem that’s detected within Dynatrace automatically closes when either the infrastructure recovers or a specific timeout period expires. In most cases, the responsible Ops team ensures that remediation is triggered in a timely manner and that the system is brought back to a healthy state so as to minimize the business impact on customers.

There are some corner cases where Dynatrace problems are raised and, for good reason, no remediation action is required.

One such example is the unexpected shutdown of a VM or host that is no longer used. Dynatrace detects the situation and alerts by raising a ‘Host unavailable’ problem. As the host is no longer used, it’s not expected that the host will become available again. In such cases, the problem remains open until the timeout automatically closes it.

This is a perfect situation where the responsible operator can push the Close problem button to manually close and resolve the issue. Additionally, the operator must add a textual comment explaining why the problem was closed, for future reference.

The screenshot below shows the close problem action as well as a text comment that contains the closing remarks.

close problems manually

close problems manually

A problem’s comment feed shows who closed the problem, when, and why. The problem comment feed can be used as a reference for understanding what’s going on within the remediation workflow and to review relevant findings.

The comment feed can also be used to show if a Jira ticket was automatically created or if a remediation action was triggered, as Andi Grabner explains in his blog post, Atlassian Connect(-ing) DevOps Tools JIRA, xMatters and Dynatrace.

close problems manually

By introducing a way to manually close problems, Ops teams can now disregard problems that they’ve already acknowledged and that are awaiting the recovery of a component.

The closing reason is preserved within the problem’s comment feed for further reference and can be used by third-party integrations through the Problem Comments API.

The post Ability to manually close problems when justified appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ability-to-manually-close-problems-when-justified/feed/ 0
Push Dynatrace-detected problems to your Microsoft Teams channels https://www.dynatrace.com/news/blog/push-dynatrace-detected-problems-to-your-microsoft-teams-channels/ https://www.dynatrace.com/news/blog/push-dynatrace-detected-problems-to-your-microsoft-teams-channels/#respond Wed, 05 Sep 2018 17:14:50 +0000 https://www.dynatrace.com/news/?p=28080 Microsoft teams

Microsoft Teams is one of the most popular enterprise team communication platforms these days. Team communication platforms are widely used by DevOps teams to notify members about critical situations, discuss remediation actions, and trigger follow-up actions. Dynatrace problem details can be fed into your Microsoft Teams channels so that your teams are always aware of […]

The post Push Dynatrace-detected problems to your Microsoft Teams channels appeared first on Dynatrace news.

]]>
Microsoft teams

Microsoft Teams is one of the most popular enterprise team communication platforms these days. Team communication platforms are widely used by DevOps teams to notify members about critical situations, discuss remediation actions, and trigger follow-up actions.

Dynatrace problem details can be fed into your Microsoft Teams channels so that your teams are always aware of potential risks within your applications, services, and infrastructure. Integrating a Microsoft Teams channel with Dynatrace gives your teams the ability to discuss incidents, evaluate solutions, and link to similar problems while remaining up to date regarding problem states and severity.

Integrate Dynatrace with Microsoft Teams

To set up an integration between Dynatrace and Microsoft Teams

  1. Within Microsoft Teams, open the Store menu.
  2. Search for and select Incoming Webhook.
    Microsoft Teams Store menu
  3. Click Next.
  4. Define a name for your Microsoft Teams connector.
  5. Click Create.
  6. Copy the unique connector URL as shown below and click Done.
    Unique connector URL in Teams
  7. Within Dynatrace, go to Settings > Integration > Problem notifications.
  8. Click Set up notifications.
  9. Select the Custom integration option.
  10. Paste your Microsoft Teams connector URL as the Dynatrace destination URL (Webhook URL field).
  11. Enter the Microsoft Teams-specific JSON payload in the format:
    {"title":"{ProblemTitle}","text":"{ProblemDetailsHTML}","themeColor":"EA4300"}
    The connector payload format can be completely customized. To read more about the Microsoft Teams connector format, please refer to this Microsoft help page.
    MS Teams integration in Dynatrace
  12. Save your integration.

Once you’ve successfully set up a connection between Dynatrace and Microsoft Teams, you’ll receive all Dynatrace-detected problems directly within your Teams channels, as shown in the example below:

Dynatrace problems delivered to MS Teams

The post Push Dynatrace-detected problems to your Microsoft Teams channels appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/push-dynatrace-detected-problems-to-your-microsoft-teams-channels/feed/ 0
Trigger Dynatrace problem alerts from external sources using Events API https://www.dynatrace.com/news/blog/trigger-dynatrace-problem-alerts-from-external-sources-using-events-api/ https://www.dynatrace.com/news/blog/trigger-dynatrace-problem-alerts-from-external-sources-using-events-api/#respond Tue, 21 Aug 2018 05:44:51 +0000 https://www.dynatrace.com/news/?p=27775 Problem dashboard

Dynatrace OneAgent is the most powerful tool available for full-stack monitoring of events in your software ecosystem. OneAgent provides alerting on more than 100 different problem types. The severity levels of detected and intelligently correlated problem types include resource contentions, slowdowns, errors, availability issues, and more. Although built-in intelligence enables Dynatrace to raise alerts about […]

The post Trigger Dynatrace problem alerts from external sources using Events API appeared first on Dynatrace news.

]]>
Problem dashboard

Dynatrace OneAgent is the most powerful tool available for full-stack monitoring of events in your software ecosystem. OneAgent provides alerting on more than 100 different problem types. The severity levels of detected and intelligently correlated problem types include resource contentions, slowdowns, errors, availability issues, and more.

Although built-in intelligence enables Dynatrace to raise alerts about all kinds of critical situations, some use cases still require external tools to trigger alerts within Dynatrace.

Generate events and trigger alerts with third-party tools

Until now, the Dynatrace Events API only offered the push of informational events into Dynatrace to enrich root-cause data—it wasn’t able to trigger new problems. With the release of Dynatrace Saas version 1.151, the Dynatrace Events API endpoint is no longer limited to informational events. Third-party API clients and tools can now create all types of severity events and thereby trigger the creation and correlation of alerts for new problems.

For example, imagine that you need to receive an alert from your UPS notifying you that one of the server rooms in your data center has suffered a severe power outage. Or in the unusual case that you need to use a legacy network monitoring system to alert you of network problems, say you may want to push these events into Dynatrace to take advantage of the intelligent problem correlation Dynatrace provides.

Note: Unbound events can’t be sent to Dynatrace as they can’t be correlated with auto-detected topology.

In the example of a power outage in a specific room within your data center, such an event would provide valuable root-cause information, even if Dynatrace has already alerted the outage of several hosts through standard availability alerting.

External alerts via the Events API

Now let’s take a look at how such an external alert can be triggered through the Dynatrace Events API endpoint. We’ll choose an availability event, as power outages typically cause severe problems in traditional infrastructure.

As you can see in the example JSON event payload below, we’ll target our event at all hosts that have the tag room23. This will enable us to focus on components that will be affected by a power outage in server room 23.  The event payload includes a list of key-value properties that provide details about the detected incident.

It’s important to note that all external events must target at least one Smartscape component, which can be a host or other component, such as a service, application, or even a single process.

This HTTP POST call opens an availability problem for the example power outage in server room 23:

HTTP POST https://<YOUR_ENVIRONMENT>/api/v1/events/?Api-Token=<YOUR_API_TOKEN>

Payload:


{

  "title": "Power outage",

  "source" : "Power control monitoring",

  "description" : "A power outage was detected in server room 23",

  "eventType": "AVAILABILITY_EVENT",

   "attachRules":{

                               "tagRule":[{

               "meTypes":["HOST"],

            "tags":["room23"]

        }]

  },

  "customProperties":{

    "Power out time": "12:00"

  }

}

The problem generated in this example automatically closes after 15 minutes if no other active events are correlated with it. If you need to overwrite the standard timeout of your custom event, use the timeoutMinutes attribute to set a different timeout period.

The Events API also allows you to refresh the event while it’s still open. Say you set a timeout of 30 minutes but discover that the incident is ongoing. You can send the same event again and it will refresh for another 30-minute duration. This refresh mechanism allows external event sources to keep events open while incidents are ongoing.

Problem creation in Dynatrace

Once the event is accepted by the Events API endpoint, all your hosts receive your custom event and open a problem, as shown below.

External event problem

Clicking the availability problem reveals the custom event name along with the custom key-value properties we provided with the event payload.

Exernal event problem details

The newly introduced event types greatly enhance your options for seamlessly integrating external event sources into Dynatrace problem correlation. The additional event types can be used by third-party tools. They can also be pushed from cloud deployment platforms.

To implement your own use cases, please refer to our Events API help page and read more about how to push external events to Dynatrace.

To learn more about the different event severity levels and the difference between single events and correlated problems, refer to our event type help page.

The post Trigger Dynatrace problem alerts from external sources using Events API appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/trigger-dynatrace-problem-alerts-from-external-sources-using-events-api/feed/ 0
Announcing the Dynatrace API explorer & OpenAPI specification https://www.dynatrace.com/news/blog/announcing-the-dynatrace-api-explorer-and-openapi-specification/ https://www.dynatrace.com/news/blog/announcing-the-dynatrace-api-explorer-and-openapi-specification/#respond Fri, 29 Jun 2018 17:58:50 +0000 https://www.dynatrace.com/news/?p=26535 D1

It’s been some time since Dynatrace was announced as a new member of the OpenAPI specification consortium. Since then, we’ve worked hard to enhance each of our Dynatrace API endpoints by providing a formal and machine processable OAS specification. The OpenAPI Specification (OAS) is a specification for machine-readable interface files that are used to describe, produce, consume, […]

The post Announcing the Dynatrace API explorer & OpenAPI specification appeared first on Dynatrace news.

]]>
D1

It’s been some time since Dynatrace was announced as a new member of the OpenAPI specification consortium. Since then, we’ve worked hard to enhance each of our Dynatrace API endpoints by providing a formal and machine processable OAS specification. The OpenAPI Specification (OAS) is a specification for machine-readable interface files that are used to describe, produce, consume, and visualize RESTful web services. There are a variety of tools available that can generate code, documentation, and test cases based on a given interface file.

One major enhancement that accompanies a machine-readable API specification is the generation of an API explorer. Typically, API explorers allow you to review all existing endpoints and directly try out those endpoints.

Because an OAS specification is now automatically included for each Dynatrace REST endpoint, you’re assured that all recent changes are automatically reflected within the new Dynatrace API explorer.

There are numerous use cases for OAS specifications beyond the generation of an API explorer. Popular use cases for OAS specifications include:

  • Automatic generation of API documentation and an API explorer
  • Automatic check for new endpoints and deprecations
  • Automatic generation of language bindings (for example, Python, Go, Java, and .NET)
  • Automatic generation of API tests

Dynatrace API explorer

Once you’ve learned what an API specification provides, it’s time to review the benefits of the newly introduced Dynatrace API explorer.

To access the Dynatrace API explorer

  1. From the navigation menu, select SettingsIntegrationsDynatrace API.
  2. Click the Dynatrace API explorer link at the top of the page.
    OpenAPI and Dynatrace API explorer

The Dynatrace API explorer lists all available API endpoints for the given Dynatrace environment. You’ll recognize the traditional API endpoints for pulling time-series or topological information from your monitored environment. Just below the title, you’ll find a link to the raw OAS specification, which can be used with any OAS-compatible tool (for example, Swagger).

OpenAPI and Dynatrace API explorer

Global unlock of API endpoints

When you open one of the given API endpoints, a dialog appears with information about the API tokens that secure the endpoint you’ve selected. Each endpoint demands a specific type of token. By clicking the global Authorize button (see callout in the example below), you can see which token scopes are necessary for all API endpoints.

OpenAPI and Dynatrace API explorer

By entering your personal API token into the global Available authorizations dialog, you unlock all related API endpoints. Once you’ve entered your API token, you can directly execute API calls within the API explorer. The Dynatrace API call example below retrieves the cluster release version. A click on the Try it out button (not shown) opens the parameter section of the selected API endpoint, where you can enter additional parameters and modify the request payload before executing it by clicking the Execute button.

The cluster version API endpoint is a simple HTTP GET call that demands no additional parameters or payload. A click on the Execute button returns the actual cluster release version (see example below).

The API explorer also transforms your requests into complete cURL commands that can be used directly within your bash or command line shell to run your API requests.

OpenAPI and Dynatrace API explorer

Beyond creating and executing API calls directly within the Dynatrace API explorer, the OAS specification of our Dynatrace API also allows you to quickly generate any required language bindings.

Auto-generated cURL commands also enable flexible combinations of multiple API calls, which you can use to create bash automation scripts or integrate Dynatrace API calls into third-party tools, such as CI build chains.

General availability, beta, or early access?

As Dynatrace follows a bi-weekly release cycle, the Dynatrace API is continuously enhanced with each new release. The API is updated with version numbers so that existing API endpoints won’t break. However new endpoints and attributes can be added to existing API versions.

We use annotations within the OpenAPI specification to inform our API consumers about changes in the existing API (for example, newly introduced API endpoints that are still in beta or early access stage).

Each endpoint that’s still in beta or early access is marked within the Dynatrace API explorer. See the example below which shows what beta annotation looks like in the OAS specification and how this is presented within the API explorer:

OpenAPI and Dynatrace API explorer

OpenAPI and Dynatrace API explorer

Introduction of a consistent OAS API specification across all Dynatrace platform APIs is a major step toward enabling seamless integration of Dynatrace monitoring intelligence into your enterprise toolchains. The auto-generated API explorer further simplifies the execution of API calls against your monitoring environment and ensures that you always have current state information about the API version you have deployed—all directly from our production code.

The post Announcing the Dynatrace API explorer & OpenAPI specification appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/announcing-the-dynatrace-api-explorer-and-openapi-specification/feed/ 0