IT operations | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Tue, 08 Apr 2025 07:46:51 +0000 en hourly 1 Fueling the next wave of IT operations: Modernization with generative AI https://www.dynatrace.com/news/blog/fueling-the-next-wave-of-it-operations/ https://www.dynatrace.com/news/blog/fueling-the-next-wave-of-it-operations/#respond Fri, 29 Mar 2024 16:06:52 +0000 https://www.dynatrace.com/news/?p=63270 How generative AI is fueling IT operations modernization

As IT operations teams face increasing pressure to enable digital transformation and more, generative AI is a key enabling technology that can help and improve outcomes.

The post Fueling the next wave of IT operations: Modernization with generative AI appeared first on Dynatrace news.

]]>
How generative AI is fueling IT operations modernization

At every organization, the digital landscape is evolving rapidly, presenting IT operations teams with unique challenges.

Teams require innovative approaches to manage vast amounts of data and complex infrastructure as well as the need for real-time decisions. Artificial intelligence, including more recent advances in generative AI, is becoming increasingly important as organizations look to modernize how IT operates.

As a result, organizations are turning to AI to automate tasks—from code development to incident response—to reduce manual effort and human error, and to boost workforce efficiency.

At the same time, challenges remain as organizations aim to become more automated. Some of these challenges involve basic tasks—such as data collection. Others involve introducing new threats as AI becomes more integrated into IT systems as a whole.

In this article, we explore recent survey data from Enterprise Strategy Group (ESG), sponsored by Dynatrace, on how organizations approach IT automation, as well as the benefits and challenges they encounter as they adopt it.

Unleashing automation and AI

According to recent ESG research, 85% of organizations are using, planning to use, or considering artificial intelligence, such as generative, causal, and predictive AI, in many of their functional areas, including IT operations. One could say that AI has moved beyond the “hype cycle” phase and entered a new phase of implementation.

A survey of 360 IT professionals at organizations in the U.S. and Canada involved with observability, IT service management, and IT automation technologies offers insight into the current status and future of AI in IT operations.

Three kinds of AI

The ESG report “Generative AI in IT Operations: Fueling the Next Wave of Modernization,” defines causal, generative, and predictive AI as follows:

Causal AI: A type of AI that analyzes real-time, context-rich data and causal dependencies to provide precise answers for issue prevention, deterministic root-cause analysis, and automated risk remediation.

Generative AI: A type of AI that uses an algorithm trained on large amounts of data collected from diverse sources to generate various types of content, including text, images, audio, and synthetic data. While ChatGPT and Google Bard are well-known examples of generative AI tools, several organizations are now utilizing proprietary, open source, or self-made generative AI large language models to help improve productivity, efficiency, and customer experiences.

Predictive AI: A type of AI that analyzes patterns, trends, and data using statistical algorithms and other advanced machine learning techniques to anticipate future behavior in systems.

AI in production

Sixty percent of respondents indicate generative AI is in production, 54% indicate causal AI is in production, and 53% indicate predictive AI is in production.

Generative AI awareness is most widespread and has an early adoption lead given the popularity of ChatGPT, Gemini, and similar tools on the consumer side, as well as the proliferation of generative AI-enabled natural language querying interfaces. As a result, many organizations are adopting it into production environments.

The heavy burden of collecting and correlating logs

Forty-five percent of respondents find collecting and correlating logs as burdensome or complex.

But organizations still wrestle with even the basics of log management. While respondents have made progress in terms of instrumentation,

This suggests there is ample opportunity for organizations to use a log management and analytics platform such as Dynatrace to ingest and analyze log data. Dynatrace Grail enables organizations to ingest data without predefining schema. Grail, alongside Dynatrace Davis AI, enables organizations to move beyond simple event correlation and to identify the root cause of problems in their applications and infrastructure.

The most likely beneficiaries of generative AI

The top three areas most likely to benefit from generative AI are IT operations (72%), cybersecurity (47%), and application development or DevOps (30%).

Organizations are turning to AI to automate manual tasks and see immediate benefits in IT operations, cybersecurity, and application development or DevOps. For IT operations, this means streamlining resource allocation, automating tasks, and enhancing incident response. For cybersecurity, it means detecting anomalies, strengthening defenses, and evolving alongside emerging threats. And for DevOps, it means accelerating DevOps processes, improving agility, and speeding time to market.

Security remains top of mind

Twenty-seven percent of respondents indicated security vulnerability is a top concern with integrating AI into IT operations.

Traditional and new challenges are emerging when integrating AI into IT operations. Therefore, it’s no surprise that 27% of those surveyed mention security vulnerability as a top concern when it comes to integrating AI into IT operations.

How generative AI improves IT operations metrics

Thirty-four percent of respondents whose organizations use or plan to use generative AI and subsequently measure or plan to measure its value indicate a 31% to 50% improvement in IT operations metrics from generative AI integration in 24 months.

The value of AI in operational acceleration carries tangible value above and beyond incremental features. This acceleration translates to a better return on assets, but it can also increase greenhouse gas emissions, complicating organizations’ ability to sustainably meet acceleration objectives.

To dive deeper into this research, download the free ebook, “Generative AI in IT Operations: Fueling the Next Wave of Modernization.”

Source: Enterprise Strategy Group, a division of TechTarget, Inc. Research Report, Generative AI in IT Operations: Fueling the Next Wave of Modernization, February 2024.

The post Fueling the next wave of IT operations: Modernization with generative AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/fueling-the-next-wave-of-it-operations/feed/ 0
Trace, diagnose, resolve: Introducing the Infrastructure & Operations app for streamlined troubleshooting https://www.dynatrace.com/news/blog/trace-diagnose-resolve-introducing-the-infrastructure-operations-app-for-streamlined-troubleshooting/ https://www.dynatrace.com/news/blog/trace-diagnose-resolve-introducing-the-infrastructure-operations-app-for-streamlined-troubleshooting/#respond Thu, 01 Feb 2024 14:00:55 +0000 https://www.dynatrace.com/news/?p=61492 Observability graphic

The new Dynatrace Infrastructure & Operations app provides ITOps and SRE teams with an up-to-date and comprehensive view of their monitored environments. The app offers a consolidated overview across data centers and all monitored hosts. The app helps users to quickly identify areas that require attention and drill down to the host level, where all necessary information is provided to quickly address any issue.

The post Trace, diagnose, resolve: Introducing the Infrastructure & Operations app for streamlined troubleshooting appeared first on Dynatrace news.

]]>
Observability graphic

Infrastructure and operations teams must maintain infrastructure health for IT environments. Traditional tools struggle with the intricacy of modern cloud services and containerized applications. These complex cloud environments obscure visibility and complicate troubleshooting, especially as teams take on the daunting task of pinpointing the exact root cause of issues.

The complex interconnections in cloud-based systems make it crucial to always have a topological overview to understand dependencies. Any problem, such as a simple software update overburdening a critical database, can cause a ripple effect that degrades the performance of dependent services or applications. For example, an unnoticed database strain could slow down the response time of a web frontend, resulting in poor user experience.

To overcome these complex issues, teams must quickly find root causes among numerous alerts and metrics.

Dynatrace automatically detects and analyzes problems

This is where Dynatrace sets itself apart, using Dynatrace Smartscape® and Davis® AI to transform IT operations. It provides accurate real-time insights to ensure operational integrity and optimal performance to meet an organization’s service-level objectives (SLOs).

Supported by Dynatrace Smartscape, Davis AI automatically detects and analyzes problems. Based on the topology model, detected dependencies, and thousands of events and metrics, Davis AI can pinpoint the origin of an issue.

However, small operations and SRE teams often deal with numerous concurrent issues managing vast IT ecosystems. Identifying a problem’s severity and prioritizing effectively can be challenging.

How do they know if it’s a $5 problem or a $1 million problem?

The Infrastructure & Operations app provides a comprehensive overview for effective prioritization

The new Infrastructure & Operations app provides situational awareness to help ops and SRE teams group and categorize problems efficiently based on their impact. This helps teams to anticipate problems, rather than just react to them, which fosters strategic, value-based decision making.

With the Infrastructure & Operations app ITOps teams can quickly track down performance issues at their source, in the problematic infrastructure entities, by following items indicated in red. This approach provides an immediate understanding of how entities such as hosts, processes, and their associated relationships contribute to the identified issue. Customers can access a real-time, granular view of their environment’s status.

Data center overview

Beginning with a comprehensive view of all interconnected data centers, Dynatrace Davis AI instantly recognizes and categorizes any issues. This allows you to pinpoint troubled data centers at a glance.

Figure 1. List view of all data centers, automatically sorted by Davis AI identified problems.
Figure 1. List view of all data centers, automatically sorted by Davis AI identified problems.

Focusing on a particular data center reveals a detailed list of all the monitored hosts. You can filter data centers based on their type and location, then sort the number of open problems.

Figure 2. Hosts view for a selected data center helps to quickly identify the most problematic hosts within the data center.
Figure 2. Hosts view for a selected data center helps to quickly identify the most problematic hosts within the data center.

The hosts page allows teams to quickly identify hosts requiring attention through straightforward sorting and filtering tools. Immediately, you can spot and understand issues with problematic hosts and seamlessly organize them according to vital health metrics, such as CPU load, available memory, disk capacity, and network connectivity, facilitating prompt and efficient issue resolution.

Furthermore, the sorting feature organizes hosts based on key health indicators, including CPU usage, memory capacity, disk space, and network performance, making it easy to quickly identify the most critical areas of concern.

Host details

Focusing on a specific host, you can see all used technologies with detailed status information and links to processes.

Figure 3. Host technologies in use with status information and links to processes.
Figure 3. Host technologies in use with status information and links to processes.

The process analysis page enables you to quickly identify problematic processes. Individual process metrics for each critical process and combined metrics make analysis more accessible and faster. The ability to sort by technology, CPU usage, and memory usage combined with process state display gives you complete process observability.

Figure 4. Host process analysis with interactive features.
Figure 4. Host process analysis with interactive features.

Gone are the days of toggling between multiple tools and sifting through disjointed data. Customers can now see all data required on the level of data centers and hosts in context, providing a comprehensive overview of how each one affects overall system health.

Start using the app now

Start enhancing your monitoring capabilities today by setting up Dynatrace OneAgent® or by integrating your cloud infrastructure. For the most granular metrics and network insights, OneAgent is the optimal choice.

Start using the Infrastructure & Operations app now to assess the health of your system. Embrace proactive monitoring with Dynatrace to keep your IT environment performing at its best.

What’s next?

Dynatrace is constantly improving the Infrastructure & Operations app to make it even more powerful. Some of the upcoming enhancements to the app include the following:

  • Topological presentation of problematic entities and their relationships
  • Mini dashboards at the data center level
  • More detailed host analytics, including OS services
  • More detailed networking observability tools

Embrace the opportunity to explore the Dynatrace Infrastructure & Operations application and let your voice be heard. Join us on our community channel and be a part of shaping the future of the Infrastructure & Operations app.

The post Trace, diagnose, resolve: Introducing the Infrastructure & Operations app for streamlined troubleshooting appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/trace-diagnose-resolve-introducing-the-infrastructure-operations-app-for-streamlined-troubleshooting/feed/ 0
What is AIOps? An insider’s guide to AI for ITOps — and beyond https://www.dynatrace.com/news/blog/what-is-aiops-2/ https://www.dynatrace.com/news/blog/what-is-aiops-2/#respond Mon, 17 Oct 2022 07:43:38 +0000 https://www.dynatrace.com/news/?p=39629 What is AIOps?

As organizations embrace automation instead of time-consuming, manual processes, many turn to artificial intelligence for IT operations, or AIOps. AIOps uses machine learning and artificial intelligence, or AI, to cut through the noise in IT operations — specifically incident management. But what is AIOps, exactly? Are all AI and AIOps approaches the same? And how […]

The post What is AIOps? An insider’s guide to AI for ITOps — and beyond appeared first on Dynatrace news.

]]>
What is AIOps?

As organizations embrace automation instead of time-consuming, manual processes, many turn to artificial intelligence for IT operations, or AIOps.

AIOps uses machine learning and artificial intelligence, or AI, to cut through the noise in IT operations — specifically incident management. But what is AIOps, exactly? Are all AI and AIOps approaches the same? And how can it support your organization?

What is AIOps?

According to Gartner, “AIOps combines big data and machine learning to automate IT operations processes, including event correlation, anomaly detection and causality determination.” A modern approach to AIOps serves the full software delivery lifecycle. It addresses the volume, velocity, and variety of data in complex multicloud environments with advanced AI techniques to provide precise answers and intelligent automation.

Most AIOps tools ingest pre-aggregated data from various technologies across the IT management landscape — including disparate observability tools — and conclude what is relevant for an analyst to focus on. But there are a few caveats to consider. We’ll discuss the current AIOps landscape and an alternative approach that truly integrates AI into the DevOps process.

How does AIOps work?

AIOps is distinct from other IT data collection solutions that process and make inferences from data. While most large organizations already have comprehensive data collection tools, they don’t provide the whole picture. Modern collection and monitoring tools often generate too much data for a human to parse and use, which is where AIOps can help.

AIOps uses AI methods to ingest, sort, and make inferences from data. A full AIOps pipeline often includes several algorithmic processes with different jobs:

  • One handles data ingestion and sorting.
  • One recognizes patterns.
  • One makes inferences from the patterns.

When combined, they significantly reduce alert fatigue and the data-sorting burden.

Equally important is AIOps’ ability to communicate information directly to the right teams. Additionally, AIOps often accompanies an increased focus on incident response automation. AI for IT operations aims to increase efficiency and observability throughout an organization. The building blocks of AIOps — machine learning algorithms and other AI processes — all help to accomplish that goal.

Two approaches to AIOps

There are two overarching AIOps approaches: traditional correlation-based AIOps and modern deterministic AIOps. This modern approach is also referred to as causal AI and uses fault-tree analysis.

Traditional AIOps

Traditional AIOps approaches are designed to reduce alerts and use machine learning models to deliver correlation-focused dashboards. These systems are often difficult to scale because the underlying machine-learning engine doesn’t provide continuous, real-time insight into an issue’s precise root cause. They require extensive training, and analysts must spend valuable time manually tuning the model and filtering out false positives.

Deterministic AI vs. statistical correlation-based AI

Modern AIOps using deterministic, causal AI

A modern AIOps solution, on the other hand, is built for dynamic clouds and software delivery lifecycle automation. It combines full stack observability with a deterministic, or causal, AI engine that can yield precise, continuous, and actionable insights in real-time. This contrasts stochastic (or randomly determined) AIOps approaches that use probability models to infer the state of systems. Only deterministic, causal AIOps technology enables fully automated cloud operations across the entire enterprise development lifecycle.

Is AIOps necessary?

Modern applications are built from hundreds or thousands of interdependent microservices distributed across multiple clouds, creating incredibly complex software environments. This complexity makes it difficult for IT pros to understand the state of these systems, especially when something goes wrong. While AIOps is often presented as a means to reduce the noise of countless alerts, it can do much more. A full-featured, deterministic AIOps solution fosters faster, higher-quality innovation; increased IT staff efficiency; and vastly improved business outcomes.

Humans can’t manually review and analyze the massive amount of data that a modern observability solution processes automatically. Typically, any approach that adds more visualizations, dashboards, and slice-and-dice query tools is more of an unwieldy bandage than a solution to the problem. Disparate interfaces still require manual intervention and analysis. In this way, traditional AIOps solutions have essentially become event monitoring tools.

How AI, observability, and analytics fit together

Modern IT strives for more capable automation, and AI is critical to achieving this goal. Continuous integration and continuous delivery processes provide smart pipelines for rolling out new features and services. Orchestration platforms, such as Kubernetes, are relieving operations teams from error-prone and mundane tasks related to keeping services up and running. This automation enables developers and operations teams to focus on innovation, rather than endless administrative tasks.

The challenges of traditional AIOps

Despite the AIOps benefits, such as improved time management and event prioritization, increased business innovation, enhanced automation, and accelerated digital transformation, correlation-based AIOps solutions have limitations.

AIOps based on correlation does not scale

With a machine learning approach, traditional AIOps solutions must collect a substantial amount of data before they can create a data set — i.e., training data — from which the algorithm can learn. Administrators can reinforce learning through rating and similar means, but it can take weeks or even months until this AI is calibrated to deliver insights into business-critical applications in production.

This approach is hardly “set and forget.” Modern applications undergo frequent changes, and their deployments are highly volatile, which implies an ever-changing data set. Traditional AIOps can’t scale up with frequent changes that occur within complex distributed applications.

Lost and rebuilt context

The second challenge with traditional AIOps centers on the data processing cycle. Traditional AIOps solutions are built for vendor-agnostic data ingestion. This means data sources typically come from disparate infrastructure monitoring tools and older-generation application performance monitoring solutions.

These tool sets first acquire one or more raw data types — such as metrics, logs, traces, events, and code-level details — at different levels of granularity. Then, they process them before finally creating alerts based on a predetermined rule — for example, a threshold, learned baseline, or certain log pattern.

Typically, machine learning can access only the aggregated events, which often exclude additional details. Now, the AI learns similar reoccurring clusters of incoming events for later classification of new events. With that data, it builds and rebuilds context — time- and metadata-based correlation — but has no evidence of actual dependencies. Integrations allow for the system to process more data, such as metrics. But those add more data sets without solving the cause-and-effect problem with certainty.

What are the key capabilities of a modern AIOps solution?

An AIOps solution should be comprehensive to save teams time and manual effort. Here are key capabilities an AIOps solution should provide.

Unified platform. A comprehensive, modern approach to AIOps is a unified platform that encompasses observability, AI, and analytics. This all-in-one approach addresses the complexity of identifying problems in systems, analyzing their context and broader business impact, and automating a response. The best solutions provide real-time, continuous insights into the state of systems and services that are critical to business operations. That way, businesses can focus on innovation rather than responding to inevitable problems with complex systems.

Topology mapping and distributed tracing. A truly modern AIOps solution should include topology-mapping capabilities, perform distributed tracing, and have strong integration capabilities. With strong topology mapping, users immediately gain a comprehensive visualization of all infrastructure, process, and service dependencies. A similarly important visibility requirement is distributed tracing, which should provide DevOps with fine-grained topology and telemetry data and metadata.

Full observability of Kubernetes environments. Kubernetes has abstracted resource management to such a high degree that the platform can be adopted across industries for a wide range of applications. But that adaptability brings complexity. AIOps is an increasingly essential part of DevOps in Kubernetes environments where reliability, scalability, and flexibility are key considerations.

Comprehensive integrations. Finally, integration is critical for the success of any modern IT solution. In addition to supporting fine-grained observability, AIOps solutions should support integration with existing security systems. Most often, the problem with existing security systems is not that they fail to work properly. Rather, it’s that they cannot be used properly due to alert fatigue and false-positive frequency.

Deterministic AI is key to AIOps success

Traditional AIOps is limited in the types of inferences it can make because it depends on metrics, logs, and trace data without a model of how systems’ components are structured. AIOps should instead use deterministic AI to fully map the topology of complex, distributed architectures to reach resolutions significantly faster.

By applying real-world AIOps use cases, businesses can harness the power of advanced analytics, machine learning, and automation to enhance monitoring, detect anomalies, and optimize performance. This transformative approach enables proactive problem resolution, improves efficiency, and empowers IT teams to deliver exceptional user experiences. Learn how AIOps can revolutionize your business by driving efficiency, reliability, and proactive decision-making.

In part two, “Applying real-world AIOps use cases to your operation,” discover how to achieve autonomous operations, and explore AIOps use cases, like applying AIOps to multicloud operations, development environments, and secure applications in real-time.

To learn more about how deterministic AI and observability can take your AIOps strategy to the next level, register for our on-demand webinar series, “AIOps with Dynatrace software intelligence” today.

The post What is AIOps? An insider’s guide to AI for ITOps — and beyond appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/what-is-aiops-2/feed/ 0