Apps and microservices Archives | Dynatrace news https://www.dynatrace.com/news/category/apps-and-microservices/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Fri, 03 Jul 2026 09:42:53 +0000 en hourly 1 What is APM? Application performance monitoring in a cloud-native world https://www.dynatrace.com/news/blog/what-is-apm/ https://www.dynatrace.com/news/blog/what-is-apm/#respond Thu, 05 Mar 2026 08:42:53 +0000 https://www.dynatrace.com/news/?p=37767 The benefits of Application Performance Monitoring.

What is APM? Application performance monitoring (APM) is the practice of tracking key software application performance metrics using monitoring software and telemetry data. Practitioners use APM to ensure system availability, optimize service performance and response times, and improve user experiences. Research firm Gartner defines APM this way: “Application performance monitoring is a suite of monitoring […]

The post What is APM? Application performance monitoring in a cloud-native world appeared first on Dynatrace news.

]]>
The benefits of Application Performance Monitoring.

What is APM?

Application performance monitoring (APM) is the practice of tracking key software application performance metrics using monitoring software and telemetry data. Practitioners use APM to ensure system availability, optimize service performance and response times, and improve user experiences.

Research firm Gartner defines APM this way:

“Application performance monitoring is a suite of monitoring software comprising digital experience monitoring (DEM), application discovery, tracing and diagnostics, and purpose-built artificial intelligence for IT operations.”

In today’s digital markets, ensuring an application is always operating at peak performance is essential. The connection between a user’s experience with the back-end services supporting its functions is not always clear, which is especially true with distributed cloud-native applications. To address these ambiguities, APM enhanced by AI greatly increases the efficiency of analyzing application dependencies and how each component affects the other.

APM provides insight into how users experience applications and where performance gaps lie. Mobile apps, websites, and business applications are some examples of front ends that APM can monitor, providing insight into an application’s user experience. APM also includes supporting elements, such as hosts, processes, services, the network, and logs, to foster additional understanding of application performance.

This GigaOm CxO Decision Brief discusses why APM and distributed tracing are essential for operating autonomous systems at enterprise scale.

Application performance monitoring vs. application performance management

In addition to application performance monitoring, APM also stands for application performance management.

While application performance monitoring focuses on specific metrics and measurements, application performance management is the wider discipline of developing and managing an application performance strategy. Both these terms refer to related technology and practices.

Observability vs. monitoring

True to its name, APM is about monitoring application performance and system health by capturing and displaying data that teams then analyze using various means. Monitoring focuses on individual metrics that can indicate specific problems.

Observability, on the other hand, ascertains a system’s internal state based on the data it generates, such as logs, metrics, and traces. With this additional granularity, observability can determine the root cause of problems and realize their effects. Using this telemetry data, observability captures the context of what’s happening across multicloud environments so teams can detect and resolve the underlying causes of issues.

The highly distributed nature of modern cloud environments requires that an effective APM solution take a holistic observability-based approach.

What does APM do?

APM has rapidly expanded to encompass a broad range of capabilities, technologies, and use cases to keep pace with cloud-native IT environments.

Continuous improvement

APM can assist teams in optimizing application performance, reliability, and response times by providing the necessary metrics and data for continuous improvement. By using APM to acquire key data points relating to application performance, such as user interaction patterns, application bottlenecks, and software issues, teams gain a greater understanding of where to concentrate efforts and resources for enhancing applications.

Cloud resource utilization

Moreover, APM assists in managing cloud spend and meeting sustainability goals by identifying where organizations can consolidate and optimize resource utilization and consumption.

Application security

APM also enhances application security, a crucial aspect of preserving business value and delivering secure user experiences, by identifying software vulnerabilities and abnormal or suspicious activities.

AI model performance

Amid the increasing prominence and necessity of AI-driven services, APM can also monitor the performance of AI models embedded in applications. AI model monitoring helps ensure that organizations can predict and control AI costs, performance, and data reliability.

Why do organizations need APM?

Every day, customers use apps to shop, work, stream shows and movies, connect to social media, and manage finances. When an app crashes, is slow to load, or doesn’t load at all, users become frustrated, which can cause the business to suffer brand damage or lose revenue. When an internal business application begins to falter, the company may also see reduced employee productivity.

Discover problems before they disrupt

By monitoring systems at the level of metrics, logs, and traces, an advanced observability-based APM solution can discover problems before they disrupt operations or cause outages. observability metrics can establish a performance baseline and detect variances that could lead to wider problems.

Determine the root cause of issues

If an issue does occur, digital teams often find it difficult to identify the root cause of an application performance problem. Causes can run the gamut, from coding errors to database slowdowns and hosting or network performance issues. Even a conflict with the operating system or the specific device being used to access the app can degrade an application’s performance. Observability-based APM can pinpoint and help teams to prioritize these issues.

Cut through cloud complexity

While modern applications such as mobile apps, websites, and business apps may seem simple on the surface, they’re highly complex. These apps comprise millions of lines of code. They include hundreds of interconnected digital services and open source solutions, and they run in containerized environments hosted across multiple cloud services. Without APM technologies, teams struggle to resolve the numerous problems that can arise, raising the likelihood of customers getting frustrated and abandoning the app altogether.

For this reason, a powerful APM solution based on end-to-end observability is necessary to properly maintain and optimize modern applications.

APM core features

APM encompasses many types of monitoring across the full IT stack, including the following, among others:

APM core features
APM core features
  • Infrastructure monitoring
  • Network monitoring
  • Database monitoring
  • Log monitoring
  • Container monitoring
  • Cloud monitoring
  • Serverless monitoring
  • Synthetic monitoring
  • End-user monitoring

Organizations often run dozens of individual monitoring tools at once, especially when they’re holding onto legacy applications and managing them using the tools they find most familiar.

Although individual tools may seem easier, especially to meet the needs of many teams, fragmented monitoring frequently creates problems. A single APM solution that takes a full-stack observability approach makes monitoring all these use cases easy and more reliable.

What are the benefits of APM?

Observability-based APM provides modern IT operations with numerous benefits, including the following.

Full-stack observability

As application infrastructures expand to encompass both on-premises and multicloud environments, organizations increasingly understand that only a full-stack observability approach can deliver comprehensive visibility into the root causes of issues, wherever they originate. Teams can monitor their entire infrastructure from end to end—encompassing everything from infrastructure health to application performance and even the end-user experience. With this visibility, teams can see all these components and understand the interdependencies among them, getting faster answers to key questions.

Continuous automation

Trying to manually maintain, configure, script, and source the volume of data in a cloud-native environment is beyond human capabilities. Therefore, organizations must continuously automate these tasks to ensure proper application performance. Processes including deployment, configuration, discovery, and updates require automation to keep pace with modern multicloud environments and user demands. For this reason, an APM solution that continuously informs and automates every touchpoint of the software development lifecycle (SDLC) and other business processes is crucial to maintaining efficiency.

AI assistance

AI assistance empowers teams by reducing manual or redundant work, allowing them to be more productive in areas of critical importance to the business. An observability-based APM solution that provides multi-tiered AI capabilities goes beyond just collecting data by using that data for real-time answers. A successful APM solution uses predictive, causal, and generative forms of AI in tandem to proactively resolve problems and improve performance without the need for extensive manual effort.

Cross-team collaboration

APM is a team sport, typically requiring the expertise of multiple teams. When organizations can depend on a unified observability-based platform as a single source of truth, teams can break down silos and achieve greater cross-team collaboration. When business, operations, application, and development teams are working from the same data sets, they can streamline communication and reach decisions quickly to resolve problems and optimize applications.

User experience and business impact

User experience is inextricably linked to business outcomes, whether the application is mobile app-to-user, IoT device-to-customers, or a web application behind the scenes. With intelligence into user sessions, including real user monitoring and session replay, teams can connect user experiences to application performance and business outcomes such as increased conversions, revenue, and completed customer journeys.

Synthetic monitoring also enables teams to proactively resolve issues and optimize applications by simulating artificial user interactions, which can ensure optimal user experience before any real issues occur.

With data-backed decisions, answers at the ready, and real-time visibility into user journeys and business key performance indicators (KPIs), organizations can consistently and more efficiently deliver ideal customer experiences across all their channels for better business outcomes.

The technical, operational, and business benefits of APM

APM provides specific benefits for technical, operations, and business teams.

The benefits of Application Performance Monitoring.
The technical and operational benefits of application performance monitoring

APM technical benefits

Business, operations, application, and development teams can expect several practical benefits from adopting APM practices and tools, including the following:

  • Increase application stability and uptime by AI-powered root-cause analysis and real-time answers
  • Reduce performance incidents with proactive alerting
  • Speed up and automate performance problem resolution
  • Accelerate and increase the quality of software releases with automated development and delivery processes

APM operational benefits

Long-time users also report that APM has given their organizations some unexpected but impactful advantages. Operational benefits include the following:

  • Increase collaboration across teams with a single source of truth
  • Boost confidence to make well-informed and impactful business decisions for teams across the organization with new insights and reliable intelligence
  • Increase efficiency and innovation for application, operations, and development teams for faster issue resolution
  • Bolster job satisfaction and higher employee retention among team members

APM business benefits

Those in the boardroom have just as much to gain from adopting APM solutions as those on the front lines of DevOps efforts. Business benefits include the following:

  • Reduce operational costs with greater automation and efficiency
  • Upgrade developer and operational productivity by facilitating cross-team collaboration
  • Improve customer experience by increasing understanding of end users and their preferences
  • Increase conversion rates from improved application stability and optimized user experience
  • Achieve sustainability goals with greater insight into the IT carbon footprint

However, modern cloud-native environments present challenges for APM solutions that require specialized capabilities to achieve these benefits.

Why do cloud-native applications make APM challenging?

Even though the benefits of APM are well established, the rise of complex cloud-native applications has made it more challenging for organizations to perform well and remain competitive.

Massive amount of telemetry data

For example, cloud-native apps generate far greater quantities of telemetry data because they are made up of myriad microservices that dynamically spin up and down in the background. Each of these microservices exists for a short period and generates its own telemetry data, adding to the overall signal noise. When this happens, it becomes more difficult to find the most important events taking place within an application’s infrastructure.

Distributed cloud architectures

What’s more, the distributed and dynamic nature of microservices often makes it difficult to pinpoint the root cause of issues without the assistance of a reliable AI engine. As a result, strong AI capabilities are a necessity for cutting through the noise and garnering meaningful answers relating to problem remediation and application optimization.

Heterogeneous data

Cloud-native apps also produce many kinds of data. Telemetry data from a serverless environment is quite different from a database or a virtual machine (VM), for example. But an organization still needs to centrally manage and make sense of all the information as it comes in.

Increased velocity

The velocity at which systems generate this data is another problem. When a cloud-native app includes many smaller microservices, data comes in at a much faster rate than with a monolithic application. All these factors add challenges that make traditional APM more difficult in a cloud-native application environment.

APM tools vs. APM platforms

Though often referred to as one in the same, both APM tools and APM platforms offer unique benefits that teams can apply based on an organization’s needs, use cases, and resource availability.

What are APM tools?

APM tools are software utilities that often focus on one specific aspect of application performance. Such point solutions can help identify specialized issues. Over time, however, organizations often find themselves using multiple APM tools that don’t necessarily integrate with one another or provide comprehensive insights into the application environment.

In response to the rapid influx of telemetry data, organizations can take one of two approaches when picking APM tools. By default or by design, different teams may deploy a combination of point solutions—tools that solve only one business problem. Conversely, they may choose a single platform that more fully encompasses the many layers and use cases within the application environment.

What is an APM platform?

An APM platform is a software system that provides a single integrated solution using AI and automation to deliver a precise, context-aware analysis of the application environment. Organizations can use an APM platform to continuously monitor the full stack for system degradation and performance anomalies.

With the deluge of telemetry data associated with cloud-native apps comes a profusion of performance monitoring tools and platforms. In response, organizations are turning to open source standards and tools, such as OpenTelemetry, to standardize how they instrument, generate, and collect telemetry data for analysis. However, to thoroughly analyze the data such tools gather, the tools often must be complemented by an observability-based APM platform. This combination can provide valuable insights into software performance and behavior across multiple cloud platforms and tools.

Point solutions can pose benefits at a local level and challenges at a macro level, while a platform approach embraces a modern vision of APM that demonstrates clear advantages at the local and macro levels. What’s more, a platform approach also streamlines and simplifies business and technical processes, enhancing collaboration across teams.

Benefits of individual APM tools

Individual APM tools are specialized to monitor specific components and provide advantages for those specific use cases. For example, some organizations use Grafana to consolidate their metrics visualizations in a single dashboard while others use Jaeger for its distributed tracing capabilities to gain better observability of their systems and troubleshoot performance issues. Both these tools are highly specialized for the environments to which they’re applied.

Teams focused on solving a specific, specialized issue, such as implementing a service mesh to help manage orchestration in their Kubernetes environment, turn to individual tools because they’re cost-effective and easy to implement.

Challenges of individual APM tools

Individual tools only provide a limited view of an organization’s application architecture. This limited visibility makes it harder to identify root causes of application performance issues, resulting in longer downtimes when problems arise. Further, they only provide a single view of the application architecture, often missing the “cause and effect” of performance problems—for example, increased CPU usage caused by a microservice failure. This limited visibility may result in unnecessary troubleshooting exercises and finger-pointing, not to mention wasted time and money.

Because the scope of these solutions is limited, they can also create silos in which teams may disagree on service-level objectives (SLOs) and metrics. This silo effect can lead to more inefficiency and blame-shifting, as teams rely on different sets of information.

APM as part of a larger multicloud observability strategy

Because APM has its roots in the era of monolithic applications before the rise of microservices, open source technologies, and cloud-native environments, some industry observers have argued that APM platforms lack the innovation and deep-dive capabilities required to keep up with bespoke point solutions and individual tools. This may be true for many traditional APM monitoring tools.

However, an observability-based platform such as Dynatrace can offer broad technological coverage across the full stack, including bespoke point solutions.

Purpose-built for cloud-native environments

By leveraging data capture for any type of application and APIs to ingest data, a cloud-native platform like Dynatrace can broaden its coverage to the entire hybrid-multicloud network. This provides a macro-level view across multiple environments to provide continuous discovery. Visibility extends to the applications running within these environments, providing proactive anomaly detection prioritized by business impact.

AI and continuous automation

Crucial capabilities of a modern APM platform include AI and continuous automation. With the explosion of observability data, a platform needs to automatically process billions of dependencies in real-time, continuously monitor the full stack, and deliver precise answers with root-cause determination. Dynatrace takes a power-of-three approach to AI that leverages predictive, causal, and generative AI to help organizations deliver the highest software performance and enable workflow automation.

Integration with cloud platforms

With the scale, diverse functionality, and dynamic nature of cloud platforms such as Amazon Web Services, Microsoft Azure, and Google Cloud Platform, successful APM solutions need to work immediately without configuration or model training. The Dynatrace platform provides complete observability out of the box for dynamic cloud environments, at scale and in context.

By going beyond metrics, logs, and traces with AI-powered analytics, Dynatrace provides automated and intelligent answers from data across the full stack, including entity relationships, user experience data, cloud services, and the latest open source standards, including OpenTelemetry.

The future of APM

APM has always played a crucial role in optimizing digital business, but the recent influx of rapid AI innovations has rendered unified, holistic APM especially crucial. AI introduces new monitoring challenges that require teams to understand their expanding applications to maintain efficient operations and a positive user experience. But achieving this level of understanding can be difficult due to the complexity of generative and agentic AI stacks. As a result, auto-discovery and AI-powered answers are essential to help resolve issues when they arise. Teams need a reliable way to cut through the noise with an efficient, unified approach to problem remediation and application optimization.

The increasing size and complexity of AI workloads have also grown past human ability to manage. For this reason, automating as many processes as possible (with an end goal of fully autonomous operations) is the most reliable way to meet the demands of modern AI workloads. However, autonomous operations are only as effective as the answers they rely on. These answers must be deterministic and exact, not based on causation. APM equipped with deterministic AI is key to facilitating effective autonomous operations, enabling meaningful action based on answers, not guesses.

Leading vendors in the APM market

Gartner names leading vendors in the APM and observability market in its annual Magic Quadrant report, giving APM users valuable insight into which solutions are best suited to their unique needs. Gartner positions each vendor into various quadrants on a graph, rating them according to their leadership position within the market and their completeness of vision.

Dynatrace was named a Leader in the 2025 Gartner® Magic Quadrant™ for Observability Platforms, positioning the company highest for Ability to Execute.

The post What is APM? Application performance monitoring in a cloud-native world appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/what-is-apm/feed/ 0
What is OpenTelemetry?  An open-source standard for logs, metrics, and traces https://www.dynatrace.com/news/blog/what-is-opentelemetry/ https://www.dynatrace.com/news/blog/what-is-opentelemetry/#respond Tue, 15 Jul 2025 14:43:50 +0000 https://www.dynatrace.com/news/?p=69968 OpenTelemetry and Dynatrace make a winning combination

OpenTelemetry is an open-source framework of tools, APIs, and SDKs that help analysts understand software performance and behavior. Also referred to as OTel, OpenTelemetry is rapidly solidifying its position as a fundamental tool in the world of observability. Born as an open-source project under the Cloud Native Computing Foundation (CNCF), OpenTelemetry provides a unified framework […]

The post What is OpenTelemetry?  An open-source standard for logs, metrics, and traces appeared first on Dynatrace news.

]]>
OpenTelemetry and Dynatrace make a winning combination


OpenTelemetry is an open-source framework of tools, APIs, and SDKs that help analysts understand software performance and behavior. Also referred to as OTel, OpenTelemetry is rapidly solidifying its position as a fundamental tool in the world of observability.

Born as an open-source project under the Cloud Native Computing Foundation (CNCF), OpenTelemetry provides a unified framework for generating, collecting, processing, and exporting telemetry data—including logs, metrics, and traces.

Using OpenTelemetry,  IT teams can instrument, generate, collect, and export telemetry data for analysis to better understand software performance and behavior. When OpenTelemetry debuted in beta in 2020, it replaced its predecessors, OpenTracing and OpenCensus.

OpenTelemetry enables observability

To appreciate what OTel does, it helps to understand observability. Traditionally speaking, observability is the ability to understand what’s happening inside a system from the knowledge of the external data it produces; usually logs, metrics, and traces.

But the data itself is only as good as what you can learn from and do with it. This definition from Hazel Weakly sums it up nicely:

“Observability is the ability to ask meaningful questions, get useful answers, and act effectively on what you’ve learned.”

Observability is important because the systems of today are exponentially more complex than the systems of ten, or even five years ago. The shift from monolithic to distributed IT architectures introduces many more moving parts and interactions to keep track of, sometimes leading to systems behaving in unpredictable ways. Observability helps you make sense of what’s happening so you can act on this information, and OpenTelemetry helps to enable observability.

By promoting consistency and interoperability, OpenTelemetry enhances observability practices and benefits the entire industry by streamlining and standardizing how everyone can collect and use data.

Since the project’s start, many vendors, including Dynatrace, have contributed to the project to make rich data collection easier and more consumable. In fact, Dynatrace is one of the top contributing organizations to OpenTelemetry.

Benefits of OpenTelemetry

Collecting application data is nothing new. However, the collection mechanism and format are rarely consistent from one application to another. This inconsistency can be a nightmare for developers and Site Reliability Engineers (SREs) who are just trying to understand the health of an application.

Most of the major observability vendors, including Dynatrace, support OTel. As a result, it has become the de facto standard for instrumenting cloud-native applications. What differentiates observability solutions from one another is what they do with your data to help you ask the right questions. Asking the right questions unlocks an elevated level of understanding, giving businesses the ability to accelerate growth, drive innovation, and deliver experiences customers love.

It’s akin to how Kubernetes became the standard for container orchestration. This broad adoption has made it easier for organizations to implement container deployments since they don’t need to build their own enterprise-grade orchestration platform. Using Kubernetes as the analog for what it can become, it’s easy to see the benefits it can provide to the entire industry.

To understand why observability and OTel’s approach to it are so critical, let’s take a deeper look at telemetry data itself and how it can help organizations transform how they do business.

What is telemetry data?

Telemetry is the process of gathering and transmitting signals (data) emitted by instrumentation code within a system’s components. Traces, metrics, and logs make up most of all telemetry data.

  • Traces result from following a process (for example, an API request or other system activity) from start to finish, showing how services connect. Keeping watch over this pathway is critical to understanding how your ecosystem works, if it’s working effectively, and if any troubleshooting is necessary. Traces consist of individual operations called spans, which include unique identifiers, such as operation name, timestamp, context, attributes, events, and status.
  • Metrics are numerical data points, either counts or measures, that systems can calculate or aggregate over time. Metrics originate from several sources, including infrastructure, hosts, and third-party sources. While logs may not always be accessible, most metrics are readily available via query. Timestamps, values, and even event names can preemptively uncover a growing problem that needs remediation.
  • Logs are important because you’ll naturally want an event-based record of notable anomalies across the system. Structured, unstructured, or plain text, these readable files can tell you the results of any transaction involving an endpoint within your multicloud environment. However, not all logs are inherently reviewable—a problem that’s given rise to external log analysis tools.

Telemetry data becomes observability data when, as noted above, you can “ask meaningful questions, get useful answers, and act effectively on that information.” Making sense of it all requires an observability backend.

How does OpenTelemetry work?

OTel consists of a few components as depicted in the following figure. Let’s take a high-level look at each one from left to right:

OpenTelemetry Components
OpenTelemetry Components (Source: Based on OpenTelemetry: beyond getting started)
  • Specification. Defines a standard telemetry data format and describes how to build instrumentation. This ensures that users have a similar experience, regardless of what language they’re using.
  • Data model. Defines fields for each signal and how they interact. Signals include traces, logs, and metrics.
  • API. Defines the methods used to instrument applications and serves as the entry point for instrumentation. Each language supported by OpenTelemetry has its own API implementation.
  • SDK. Implements the API and also determines how systems generate and correlate their telemetry. Each language supported by OpenTelemetry has its own SDK. Both the APIs and the SDKs are defined in the specification to ensure a consistent experience across implementations.
  • Collector. A vendor-neutral binary used to ingest, transform, and export data to one or more observability backends.
  • OpenTelemetry Protocol (OTLP). A vendor—and tool-agnostic specification for encoding—transmitting and delivering OpenTelemetry data. Telemetry data emitted by the SDK uses OTLP, and many observability backends now support ingesting data in the OTLP format. For those who do not, there are exporters available that convert data from OTLP to a tool-specific format. OTLP supports both HTTP and gRPC.
  • Observability backend. A system or tool where telemetry data collected by OpenTelemetry is sent, stored, and analyzed. It enables organizations to derive meaningful insights and make sense of telemetry data in a cohesive way.

Flexible API/SDK integration

You can decouple the API from the telemetry-generating code with minimal implementation. This decoupling allows your app or library to run with just the API package, without sending telemetry data to the backend. This setup acts as a placeholder until you’re ready to integrate an SDK. When ready, you can choose an SDK that best fits your needs, whether it’s the OpenTelemetry SDK, a vendor-specific one, or a custom-built SDK. This flexibility ensures you can add functionality without significant code changes.

What’s next for OpenTelemetry?

OpenTelemetry is maturing and is fast approaching its graduation as a CNCF project. Traces, logs and most parts of metrics are now considered generally available. The OpenTelemetry project’s goals extend well beyond its current offerings. Exciting initiatives are paving the way for even broader use cases, such as improving digital experiences and enabling detailed insights into application performance through code-level profiling. Let’s check out some highlights of OpenTelemetry’s exciting initiatives, all designed to take observability to the next level.

  • Digital Experience Monitoring. Developers and product teams will soon be able to gather telemetry data directly from user-facing applications. This allows organizations to identify where users face lags or issues, enhancing overall app performance.
  • Code-Level Profiling. OpenTelemetry is also evolving to profile application code in real-time. This provides deeper insights into how specific sections of code behave in production, helping engineers optimize critical parts of their applications.
  • AI Agents. Recently, there’s been an explosion in the need for monitoring AI systems. OpenLLMetry is donating their code to the OpenTelemetry project. If accepted, it will soon become an extension of the OpenTelemetry ecosystem.

How can I contribute to the OTel community?

If you’ve been curious about contributing to OpenTelemetry (OTel) but are unsure where to begin, there are plenty of ways to get involved. Whether you’re a newcomer or a seasoned practitioner, the OpenTelemetry community offers a variety of opportunities suited to different interests and skill levels.

Some ways you can contribute to the OpenTelemetry Project include the OpenTelemetry documentation, OpenTelemetry blog, End User SIG, OpenTelemetry Demo, or a language or component-specific Special Interest Group SIG). No matter how big or small your contributions, they make a difference and are deeply valued.

OpenTelemetry veteran, Adriana Villela, wrote a great article to help you get started.

How does Dynatrace contribute to the Otel community?

Dynatrace is an active member of the OpenTelemetry community. Dynatracers hold key leadership roles as project maintainers or approvers in the following groups:

  • OpenTelemetry Technical Committee
  • OpenTelemetry Specification (Metrics, Semantic Conventions)
  • OpenTelemetry for JavaScript
  • OpenTelemetry Collector
  • OpenTelemetry Demo project
  • OpenTelemetry End User SIG

In fact, Dynatrace has a team dedicated to contributing to OpenTelemetry and ensuring that the Dynatrace platform integrates smoothly with OpenTelemetry data.

Dynatrace and OpenTelemetry together can deliver more value

OpenTelemetry is a key enabler in the observability space, providing a unified framework for collecting telemetry data. However, to unlock the full potential of this data and turn it into actionable insights, a powerful platform to manage, analyze, and visualize it effectively is essential.

Dynatrace is purpose-built to enhance OpenTelemetry’s capabilities. Data plus context are critical to supercharging observability, and with Dynatrace, you’re not just collecting data; you’re gaining a deep understanding of how your systems work and how to optimize them. With seamless integration, advanced analysis across telemetry data, insights into business outcomes, and predictive analytics spanning your entire stack, Dynatrace turns your OpenTelemetry data into actionable intelligence to optimize your systems.

Explore how Dynatrace can transform your OpenTelemetry data into a powerful driver of innovation and business success. Learn more with this video series on getting started with Dynatrace and OpenTelemetry.


Dynatrace Can Do THAT with OpenTelemetry? video thumbnail

Want to explore on your own? Check out the Dynatrace playground.

Want to try Dynatrace with your OpenTelemetry data? Check out our free trial and walk through this Astronomy Shop demo to populate your own data.

The post What is OpenTelemetry?  An open-source standard for logs, metrics, and traces appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/what-is-opentelemetry/feed/ 0
CrowdStrike update: How Dynatrace helped customers recover in hours https://www.dynatrace.com/news/blog/crowdstrike-update-crisis-dynatrace-customers-recovery/ https://www.dynatrace.com/news/blog/crowdstrike-update-crisis-dynatrace-customers-recovery/#respond Wed, 31 Jul 2024 13:29:47 +0000 https://www.dynatrace.com/news/?p=65026 CrowdStrike update

The ability to recover quickly in case of a sudden IT outage is crucial for business resilience. This blog is part of a series that explores how organizations can maintain business resilience by having the right capabilities to recover quickly from an IT outage.

The post CrowdStrike update: How Dynatrace helped customers recover in hours appeared first on Dynatrace news.

]]>
CrowdStrike update

On July 19th, 2024, countless organizations had their operations disrupted by a routine software update from CrowdStrike, a popular cybersecurity software. The resulting outages wreaked havoc on customer experiences and left IT professionals scrambling to quickly find and repair affected systems.

A wide variety of companies and industries have suffered the effects of this incident, from delayed flights to disruptions in healthcare, insurance, and the financial industry. The ripple effects on the global supply chain have been equally significant. The crisis has emphasized the importance of having a strategy for maintaining stability and performance.

Time is of the essence in any crisis—so is having the right tools and capabilities. Although Dynatrace can’t help with the manual remediation process itself, end-to-end observability, AI-driven analytics, and key Dynatrace features proved crucial for many of our customers’ remediation efforts.

Following are some of the critical capabilities many of our customers relied on to recover from the CrowdStrike update crisis in just hours.

1. Real-time monitoring with out-of-the-box features

Real-time data and monitoring are crucial for maintaining situational awareness of IT environment stability and performance, especially during a crisis. Knowing what’s offline and which dependencies connect to mission-critical services is key to determining the impact of an incident and determining where to start remediation. Dynatrace offers various out-of-the-box features and applications to provide a high-density overview of system health for all hosts and related metrics in a single view.

Smartscape topology mapping

Dynatrace Smartscape® provides a dynamic, real-time visualization of the entire application topology so teams can quickly identify and address issues. Understanding application dependencies helps teams prioritize what to address first. For example, a good course of action is knowing which impacted servers run mission-critical services and remediating those first.

Problems application

The Problems application automatically identifies issues, collects the context behind them, and presents their root cause and impacts in a single view. Powered by Davis® AI, the app helps teams immediately see the problem’s duration, root cause, and business impact.

Dashboards and visualizations

Standard dashboards and visualizations also provide situational awareness out of the box. The following honeycomb visualization shows a healthy environment before an incident and after the incident starts. All the problems, offline hosts, databases, and failing services appear in red.

Honeycomb visualization: Before CrowdStrike outage
Systems before the outage
Honeycomb visualization: CrowdStrike outage in progress
Systems during the outage

Dynatrace OneAgent full-stack and Foundation and Discovery modes

Dynatrace OneAgent provides automatic discovery and monitoring. In addition to using OneAgent for full stack monitoring of the most critical applications, Dynatrace offers OneAgent Foundation and Discovery mode. This lightweight alternative provides full coverage of an environment in scenarios where teams need cost-effective yet comprehensive monitoring. Foundation and Discovery provide essential metrics and topology discovery, making it useful to quickly identify and recover affected hosts.

Together, these technologies enable organizations to maintain real-time visibility and control, swiftly mitigating the impact of incidents and efficiently restoring critical services. They also enable companies to measure the effectiveness of their remediation activities to ensure that recoveries proceed as expected.

How out-of-the-box monitoring features helped one company recover from the CrowdStrike outage within hours

A US-based pharmaceutical company recovered its most critical systems within hours of the CrowdStrike incident using out-of-the-box real-time monitoring features. Dynatrace automatically found the hosts that were unavailable or having problems. The key information displayed on the standard Dynatrace Problems app and the Infrastructure and Operations App became the basis of their team’s remediation plan. The company was back to normal business operations as other companies continued to struggle with recovering days after the initial software push.

2. Synthetic monitoring

Synthetic monitoring is a critical tool for ensuring application reliability and performance, especially during a crisis. By simulating user interactions and running tests from various locations worldwide, synthetic monitoring provides a comprehensive view of application performance and availability. This proactive approach allows organizations to detect and resolve issues early, optimize performance, and maintain a high-quality user experience.

Organizations can use synthetic monitoring to continuously monitor API endpoints and ensure that critical user journeys perform well and meet service level agreements (SLAs). This awareness can help reduce downtime and minimize disruption, enabling swift action if an incident occurs.

Many businesses rely on third-party services, such as payment processors, content delivery networks (CDNs), and ticketing systems to get through their day-to-day operations. Even if the business isn’t directly affected by a crisis, their third-party suppliers may be, which can still disrupt their operations.

How synthetic monitoring helped one company get an early warning of the CrowdStrike impact

During the recent CrowdStrike crisis, a US life insurance provider was impacted indirectly through its third-party ticketing system. The Dynatrace synthetic monitors they used to monitor the performance of this critical third-party application immediately detected the outage when the CrowdStrike software affected the vendor’s servers.

Dynatrace created a problem notification and Davis AI determined the root cause on the vendor’s side, hours before the company publicly announced the CrowdStrike incident affected it. This advanced warning allowed the life insurance provider to execute a contingency plan and course of action much earlier than if they had to wait for the third-party provider to notify them of the problem.

Synthetic monitoring view of the CrowdStrike outage

Synthetic monitoring view of the CrowdStrike outage
Synthetic monitoring views of the CrowdStrike outage

3. Dynatrace Query Language (DQL)

Dynatrace Query Language (DQL) is a structured syntax for exploring, querying, and processing observability data in Dynatrace. It allows users to chain commands together to filter, manipulate, and analyze data efficiently.

As a data query tool, DQL provides flexibility and customization, allowing organizations to tailor their investigations to meet specific needs, addressing unique challenges and optimizing performance.

Using DQL, investigators can find specific answers so they can quickly identify and respond to issues, which is crucial during crises like the CrowdStrike incident.

Dynatrace Notebooks is an interactive capability that enables users across the organization to collaborate using code, text, and rich media to build, evaluate, and share insights for exploratory analytics. This ability to track and collaborate on issue details is a crucial capability in a crisis.

How DQL and Notebooks helped companies pinpoint affected systems and prioritize remediation

During the CrowdStrike crisis, a North American telecommunications provider used DQL and a notebook to create custom charts filtered by application. This helped the company prioritize remediating its most critical servers first, restoring essential services promptly.

DQL and Notebooks investigation showing systems affected by the CrowdStrike outage
DQL and Notebooks investigation showing systems affected by the CrowdStrike outage

When a major US airline began experiencing the CrowdStrike outage, its IT team also used DQL and Dynatrace Notebooks to identify which systems were no longer forwarding logs back to Dynatrace. This newfound visibility enabled the kiosk management team to focus their recovery efforts effectively.

To see an example of how to use DQL to find when BSOD issues are being written to Windows system logs, see the blog Crowdstrike BSOD: Quickly find machines impacted by the CrowdStrike issue by Dynatrace Principal Solutions Engineer Josh Wood, Ph.D.

4. Real user monitoring to understand business impact

Real User Monitoring (RUM) offers comprehensive insights into user experiences across web, mobile, and custom applications. By capturing user sessions, RUM provides a detailed view of user journeys, helping businesses understand critical actions for conversions.

When an incident occurs, Dynatrace automatically generates a problem notification. Davis AI analyzes details from the front end to the backend to identify the root cause, severity, and impact of the issue.

Dynatrace RUM also tracks application downtime, enabling organizations to calculate the cost of business interruptions and account for lost revenue. RUM offers flexible options for tracking and reporting key information, such as conversion goals. Examples include successful checkouts, newsletter signups, or demo requests. By monitoring conversion rates, businesses can estimate expected revenue to better understand the financial impact of an incident and help organizations account for lost revenue.

How Dynatrace RUM helped one company identify the business impacts of the CrowdStrike outage

For example, during the CrowdStrike crisis, a North American mortgage provider received an alert for unexpected low traffic. The problem card helped them identify the affected application and actions, as well as the expected traffic during that period. This allowed them to prioritize remediation efforts on their most critical services.

Dynatrace RUM shows the user impact of the CrowdStrike outage
Dynatrace RUM shows the user impact of the CrowdStrike outage

5. Using SLOs to verify recovery

Service level objectives (SLOs) are essential for maintaining and enhancing the performance of applications and services, especially during and after a crisis. Dynatrace makes it easy to create, capture, and visualize SLOs in real time. Establishing and monitoring SLOs can play an instrumental role before, during, and after a crisis.

  • Before a crisis. Setting up SLOs for mission-critical services helps establish and maintain standards for availability and performance. Dynatrace AI continuously monitors these benchmarks, allowing teams to identify and address potential issues proactively.
  • During a crisis. SLOs provide real-time monitoring and immediate feedback on service performance. Dynatrace AI can quickly pinpoint the root cause of issues, enabling swift resolution and minimizing user impact.
  • After a crisis. SLOs ensure that application performance returns to the same standard of performance as before the incident. They play a crucial role in post-incident analysis, helping teams understand the business impact of the incident and implement improvements to prevent future occurrences.

By implementing Dynatrace SLOs, organizations can ensure robust performance management before, during, and after a crisis, leading to more resilient and reliable services.

Prepare for any crisis with observability and the right capabilities

The recent CrowdStrike crisis has highlighted the critical need for robust monitoring and observability tools. For organizations navigating disruptions, observability is crucial for rapid detection and remediation. Dynatrace offers comprehensive solutions with real-time data, synthetic monitoring, DQL querying capabilities, real user monitoring, and SLOs. These tools empower organizations to maintain stability, swiftly identify and resolve issues, and ensure the continuity of essential services.

The real-world examples of our customers demonstrate how Dynatrace monitoring solutions have enabled organizations to recover quickly and maintain operational stability, ultimately safeguarding their customer experiences and bottom lines. As we continue to lean heavily on technology for day-to-day operations, observability will be essential for navigating the complexities of an increasingly digital world.

Using Dynatrace, organizations can not only react fast to mitigate the immediate impacts of crises but also be proactive and build resilient IT infrastructure that’s prepared for future challenges.

Contact us to learn how you can gain the same situational awareness and responsiveness that enabled these customers to recover so quickly from the CrowdStrike outage.

To learn more about the recent CrowdStrike update outage and explore more resources to help you maintain business resilience, check out the resource center, Business Resilience through CrowdStrike and Beyond.

The post CrowdStrike update: How Dynatrace helped customers recover in hours appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/crowdstrike-update-crisis-dynatrace-customers-recovery/feed/ 0
What is observability? Not just logs, metrics and traces https://www.dynatrace.com/news/blog/what-is-observability-2/ https://www.dynatrace.com/news/blog/what-is-observability-2/#respond Wed, 26 Jun 2024 13:36:21 +0000 https://www.dynatrace.com/news/?p=39527

As organizations embrace cloud-native technologies, system architectures have dramatically increased in complexity and scale. Customer experiences are more important than ever, and as a result, IT teams face mounting pressure to track and respond to issues much faster. To address these challenges, teams are turning to observability solutions so they can proactively identify and resolve […]

The post What is observability? Not just logs, metrics and traces appeared first on Dynatrace news.

]]>

As organizations embrace cloud-native technologies, system architectures have dramatically increased in complexity and scale. Customer experiences are more important than ever, and as a result, IT teams face mounting pressure to track and respond to issues much faster. To address these challenges, teams are turning to observability solutions so they can proactively identify and resolve issues and automate workflows in their highly distributed and complex computing environments. But what is observability, and what do teams need to do it right?

What is observability?

In IT and cloud computing, observability is the ability to measure a system’s current state based on the data it generates, recorded as logs, metrics, and traces:

  • Logs record the details of an event
  • Metrics capture the numeric measurements used to quantify the performance and health of services
  • Traces track how services connect from end to end in response to requests

Observability has become more critical in recent years as cloud-native environments have gotten more complex, and the potential root causes for a failure or anomaly have become more difficult to pinpoint.

Because cloud services rely on a distributed and dynamic architecture, observability may also refer to the specific software tools and practices organizations use to interpret cloud performance data.

How observability works

Observability relies on telemetry derived from instrumentation that comes from the endpoints and services in your multicloud computing environments. In these modern environments, every hardware, software, and cloud infrastructure component and every container, open source tool, and microservice generates records of every activity. The goal of observability is to understand what’s happening across all these environments and among the technologies, so you can detect and resolve issues to keep your systems efficient and reliable and your customers happy.

Implementing observability

Organizations usually implement observability using a combination of instrumentation methods, including open source instrumentation tools, such as OpenTelemetry.

Many organizations also adopt an observability solution to help them detect and analyze the significance of events to their operations, software development life cycles, application security, and end-user experiences.

As teams begin collecting and working with observability data, they are also realizing its benefits to the business, not just IT.

Although some people may think of observability as a buzzword for sophisticated application performance monitoring (APM), there are a few key distinctions to keep in mind when comparing observability and monitoring.

Monitoring vs. observability: What’s the difference between monitoring and observability?

Is observability really monitoring by another name? In short, no. While observability and monitoring are related—and can complement one another—they are actually different concepts.

Monitoring

In a monitoring scenario, you typically preconfigure dashboards to alert you about performance issues you expect to see later. However, these dashboards rely on the key assumption that you’re able to predict what kinds of problems you’ll encounter before they occur.

Cloud-native environments don’t lend themselves well to this type of monitoring because they are dynamic and complex, which means you cannot predict what problems might arise in advance.

Observability

In an observability scenario, where teams have fully instrumented an environment to provide complete observability data, you can flexibly explore what’s going on and quickly figure out the root cause of issues you may not have been able to anticipate.

Traditionally, the industry defines observability as logs, metrics, and traces. In more complex cloud environments, however, observability must encompass more, including metadata, user behavior, topology and network mapping, and access to code-level details.

Observability pillars include logs, metrics, and traces.
Observability pillars include logs, metrics, and traces. Modern observability also includes metadata, user behavior, topology and network mapping, and code-level details.

Why is observability important?

In enterprise environments, observability helps cross-functional teams understand and answer specific questions about what’s happening in highly distributed systems. Observability enables you to understand what is slow or broken and what you need to do to improve performance. With an observability solution in place, teams can receive alerts about issues and proactively resolve them before they impact users.

Understanding “unknown unknowns”

Because modern cloud environments are dynamic and constantly changing in scale and complexity, teams neither know about nor can monitor most problems. Observability addresses this common issue of “unknown unknowns,” enabling you to continuously and automatically understand new types of problems as they arise.

Automating AIOps and DevSecOps

Observability is also a critical capability of artificial intelligence for IT operations (AIOps). As more organizations adopt cloud-native architectures, they are also looking for ways to implement AIOps, harnessing AI as a way to automate more processes throughout the DevSecOps lifecycle. By bringing AI to everything—from gathering telemetry to analyzing what’s happening across the full technology stack—your organization can have the reliable answers essential for automating application monitoring, testing, measuring service level objectives (SLOs), continuous delivery, application security, and incident response.

Optimizing user experiences

The value of observability doesn’t stop at IT use cases. Once you begin collecting and analyzing observability data, you have an invaluable window into the business impact of your digital services. This visibility enables you to optimize conversions, validate that software releases meet business goals, and prioritize business decisions based on what matters most.

When an observability solution also analyzes user experience data using synthetic and real-user monitoring, you can discover problems before your users do and design better user experiences based on real, immediate feedback.

Benefits of observability

Observability delivers powerful benefits to IT teams, organizations, and end users alike. Following are some of the use cases observability facilitates.

1 Application performance monitoring

Full end-to-end observability enables organizations to get to the bottom of application performance issues much faster, including issues that arise from cloud-native and microservices environments. Teams can also use an advanced observability solution to automate more processes, which increases efficiency and innovation among Ops and Apps teams.

2 DevSecOps and SRE

Observability is not just the result of implementing advanced tools but a foundational property of an application and its supporting infrastructure. The architects and developers who create the software must design it to be observed. Then DevSecOps and SRE teams can leverage and interpret the observable data during the software delivery lifecycle to build better, more secure, and more resilient applications.

3 Monitoring infrastructure, cloud, and Kubernetes environments

Infrastructure and operations (I&O) teams can leverage the enhanced context an observability solution offers for monitoring on-premises and cloud infrastructure and Kubernetes environments. This unified observability-based approach can improve application uptime and performance, cut down the time required to pinpoint and resolve issues, detect cloud latency issues, optimize cloud resource utilization, and improve administration of their Kubernetes environments and modern cloud architectures.

4 End-user experience

A good user experience can enhance a company’s reputation and increase revenue, delivering an enviable edge over the competition. By spotting and resolving issues well before the end-user notices and making an improvement before it’s even requested, an organization can boost customer satisfaction and retention. It’s also possible to optimize the user experience through real-time playback, gaining a window directly into the end-user’s experience exactly as they see it, so everyone can quickly agree on where to make improvements.

5 Business analytics

Business analytics enable organizations to combine business context with full stack application analytics and performance to understand real-time business impact, improve conversion optimization, ensure that software releases meet expected business goals, and confirm that the organization is adhering to internal and external SLAs.

6 DevOps and DevSecOps automation

DevSecOps teams can tap observability to get more insights into the apps they develop, and automate testing and CI/CD processes so they can release better quality code faster. This means organizations waste less time on war rooms and finger-pointing. Not only is this a benefit from a productivity standpoint, but it also strengthens the positive working relationships that are essential for effective collaboration.

These organizational improvements open the door to further innovation and digital transformation. And more importantly, the end-user ultimately benefits in the form of a high-quality user experience.

How do you make a system observable?

If you’ve read about observability, you likely know that collecting the measurements of logs, metrics, and distributed traces are the three key pillars to achieving success. However, observing raw telemetry from back-end applications alone does not provide the full picture of how your systems are behaving.

Neglecting the front-end perspective potentially skews or even misrepresents the full picture of how your applications and infrastructure are performing in the real world for real users. Extending the three-pillars approach, IT teams must augment telemetry collection with user-experience data to eliminate blind spots:

  1. Logs: Logs are structured or unstructured text records of discreet events that occurred at a specific time.
  2. Metrics: Metrics are the values represented as counts or measures that are often calculated or aggregated over a period of time. Metrics can originate from a variety of sources, including infrastructure, hosts, services, cloud platforms, and external sources.
  3. Distributed tracing: Tracing follows the activity of a transaction or request as it flows through applications and shows how services connect, including code-level details.
  4. User experience: User experience data extends traditional observability telemetry by adding the outside-in user perspective of a specific digital experience on an application, even in pre-production environments.

Why the three pillars of observability aren’t enough

Obviously, data collection is only the start. Simply having access to the right logs, metrics, and traces isn’t enough to gain true observability of your environment. Once you’re able to use that telemetry data to achieve the end goals of improving end-user experience and business outcomes, only then can you really say you’ve achieved the purpose of observability.

The importance of open source solutions

There are other observability capabilities organizations can use to observe their environments. Open source solutions, such as OpenTelemetry, provide a de facto standard for collecting telemetry data in cloud settings. These open source solutions enhance observability for cloud-native applications and make it easier for developers and operations teams to achieve a consistent understanding of application health across multiple environments.

The role of real-user monitoring (RUM) and synthetic testing

Organizations can also use real user monitoring to gain real-time visibility into the user experience, tracking the path of a single request and gaining insight into every interaction it has with every service along the way. Teams can observe this experience using synthetic monitoring or even view a recording of the actual session. These capabilities extend telemetry by adding in data for APIs, third-party services, errors occurring in the browser, user demographics, and application performance from the user’s perspective.

With real-user monitoring IT, DevSecOps, and SRE teams can not only see the complete end-to-end journey of a request but also access real-time insight into system health. From there, they can proactively troubleshoot areas of degrading health before they impact application performance. They can also more easily recover from failures and gain a more granular understanding of the user experience.

Don’t forget over-burdened teams

While IT organizations have the best of intentions and strategy, they often overestimate the ability of already overburdened teams to constantly observe, understand, and act upon an impossibly overwhelming amount of data and insights. Although there are many complex challenges associated with observability, the organizations that overcome these challenges will find it worth their while.

What are the challenges of observability?

Observability has always been a challenge, but cloud complexity and the rapid pace of change have made it an urgent issue. Cloud environments generate a far greater volume of telemetry data, particularly in microservices and containerized application environments. They also create a far greater variety of telemetry data than teams have ever had to interpret in the past. Lastly, the velocity with which all this data arrives makes it that much harder to keep up with the flow of information, let alone accurately interpret it in time to troubleshoot a performance issue.

Organizations also frequently run into the following challenges with observability.

1 Data silos

Multiple agents, disparate data sources, and siloed monitoring tools make it hard to understand interdependencies across applications, multiple clouds, and digital channels, such as web, mobile, and IoT.

2 Volume, velocity, variety, and complexity

It’s nearly impossible to get answers from the sheer amount of raw data collected from every component in ever-changing modern cloud environments, such as AWS, Azure, and Google Cloud Platform (GCP). This is also true for Kubernetes and containers that can spin up and down in seconds.

3 Manual instrumentation and configuration

When IT resources are forced to manually instrument and change code for every new type of component or agent, they spend most of their time trying to set up observability rather than innovating based on insights from observability data.

4 Lack of pre-production

Even with load testing in pre-production, developers still don’t have a way to observe or understand how real users will impact applications and infrastructure before they push code into production.

5 Wasting time troubleshooting

Application, operations, infrastructure, development, and digital experience teams are pulled in to troubleshoot and try to identify the root cause of problems, wasting valuable time guessing and trying to make sense of telemetry and come up with answers.

6 Multiple tools and vendors

While a single tool may give an organization observability of one specific area of their application architecture, that one tool may not provide complete observability across all the applications and systems that can affect application performance.

7 Inability to determine root-causes

Also, not all types of telemetry data is equally useful for determining the root cause of a problem or understanding its impact on the user experience. As a result, teams waste time digging for answers across multiple solutions and painstakingly interpreting the telemetry data when they could be applying their expertise toward fixing the problem right away.

However, with a single source of truth, teams can get answers and troubleshoot issues much faster.

The importance of a single source of truth

Organizations need a single source of truth to gain complete observability across their application infrastructure and accurately pinpoint the root causes of performance issues. When organizations have a single platform that can tame cloud complexity, capture all the relevant data, and analyze it with AI, teams can instantly identify the root cause of any problem, whether it lies in the application itself or the supporting architecture.

A single source of truth enables teams to do the following:

  • Turn terabytes of telemetry data into real answers rather than asking IT teams to cobble together an understanding of what has happened using snippets of data from disparate sources
  • Gain crucial contextual insights into areas of the infrastructure they might not have otherwise been able to see
  • Work collaboratively and accelerate the troubleshooting process further, which empowers the organization to act faster than it could by using traditional monitoring tools thanks to enhanced awareness

Making observability actionable and scalable for IT teams

To achieve observability, resource-constrained teams need to be able to collect and act upon a deluge of telemetry data in real time. Responding in real time enables teams to prevent business-impacting issues from propagating further or even occurring in the first place. Following are some ways teams can make observability actionable and scalable.

1 Understand the context and the topology

Understanding the context and topology of an IT environment involves instrumenting applications and infrastructure in a way that identifies relationships between every entity and interdependency among potentially billions of interconnected components. Rich context metadata enables real-time topology maps, providing an understanding of causal dependencies both vertically throughout the stack and horizontally across services, processes, and hosts.

2 Implement continuous automation

Automatic discovery, instrumentation, and baselining of every system component on a continuous basis shifts IT effort away from manual configuration work to value-add innovation projects that can prioritize understanding of the things that matter. Observability becomes “always-on” and scalable, so constrained teams can do more with less.

3 Establish true AIOps

Exhaustive AI-driven fault-tree analysis combined with code-level visibility enables teams to automatically pinpoint the root cause of anomalies without having to rely on time-consuming human trial and error. Additionally, causation-based AI can automatically detect any unusual change points to discover “unknown unknowns” that teams are not aware of or monitoring. These actionable insights drive the faster and more accurate responses that DevOps and SRE teams require.

4 Foster an open ecosystem

An open ecosystem extends observability to include external data sources, such as OpenTelemetry, which is an open-source project led by vendors such as Dynatrace, Google, and Microsoft. OpenTelemetry expands telemetry collection and ingestion for platforms that provide topology mapping, automated discovery and instrumentation, and actionable answers required for observability at scale.

5 Utilize AI

An AI-driven solution-based approach makes observability truly actionable by solving the challenges associated with cloud complexity. An observability solution makes it easier to interpret the vast stream of telemetry data arising from multiple sources at increasingly greater velocities. With a single source of truth, teams can quickly and accurately pinpoint root causes of issues before they result in degraded application performance or, in the event a failure has already occurred, accelerate their time to recovery.

Advanced observability also improves application availability through end-to-end distributed tracing across serverless platforms, Kubernetes environments, microservices, and open-source solutions. By gaining visibility into the complete journey of a request from start to finish, teams can proactively identify application performance issues and gain crucial insight into the end-user experience. This way, IT teams can quickly act on issues of concern, even as the organization scales its application infrastructure to support future growth.

Bring observability to everything

You can’t waste months or years trying to build your own tools or test out multiple vendors that only enable you to solve one piece of the observability puzzle. Instead, you need a solution that can help make all your systems and applications observable, give you actionable answers, and provide technical and business value as fast as possible.

Advanced observability from Dynatrace provides all these capabilities in a single platform, empowering your organization to tame modern cloud complexity and transform faster. Now, more than ever, it’s critical to make comprehensive observability part of every cloud migration. At Dynatrace, we call that approach cloud done right.

Read our free eBook, Upgrade to advanced observability for answers in cloud-native environments, to learn how advanced observability gives you actionable answers in cloud-native environments.

The post What is observability? Not just logs, metrics and traces appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/what-is-observability-2/feed/ 0
Turbocharge Dynatrace app development with the new Visual Studio Code extension `Dynatrace Apps` https://www.dynatrace.com/news/blog/turbocharge-dynatrace-app-development-with-the-new-visual-studio-code-extension-dynatrace-apps/ https://www.dynatrace.com/news/blog/turbocharge-dynatrace-app-development-with-the-new-visual-studio-code-extension-dynatrace-apps/#respond Thu, 06 Jun 2024 15:00:05 +0000 https://www.dynatrace.com/news/?p=64277 Dynatrace App Configuration

Dynatrace® Apps offer a powerful and intuitive way of extending Dynatrace in a secure, scalable, and enterprise-grade way. Build native applications directly on the Dynatrace platform and combine observability, security, and business data with the complete security and governance capabilities of Dynatrace.

The post Turbocharge Dynatrace app development with the new Visual Studio Code extension `Dynatrace Apps` appeared first on Dynatrace news.

]]>
Dynatrace App Configuration

As an app developer, you have many recurring tasks, such as starting the development server, creating app functions, querying data stored in Grail, managing app configurations, and building and deploying apps. Sound familiar? The Visual Studio Code extension Dynatrace Apps is here to streamline your development process and simplify app building.

App configuration and build management

Let’s start with the basics: Use the App Configuration form (select Configure app in the VS Code Project tree) to set up all necessary options from the app manifest, including name, icon, version, and environment URL, as well as app scopes and content security policies. The extension wraps all important functionality from the Dynatrace App Toolkit to build and deploy your app, manage app dependencies, and test the new version of your app without redeployment.

Working with data

Dynatrace Apps are often used to “bring logic to data,” addressing new use cases on top of data stored in Dynatrace. Now you can easily query live data directly within VS Code using the Dynatrace Query Language (DQL).

First, create a new file, such as logQuery.dql, and enter your new query, utilizing autocompletion similar to Notebooks. The Run query button allows you to select a timeframe, executes the query directly from within the IDE, and shows the result in a new editor window. Then, you can explore the result’s data structure and access the JSON object’s properties.DQL query screenshot in Dynatrace

(Re)Using queries within your app

Once you are happy with the result of your query, you can easily use it in your app code. Adding the annotation @name(“logsByLevel”) and saving the file auto-generates the corresponding TypeScript function runQueryLogsByLevel, ready for import and use within your app.

Define DQL query screenshot in Dynatrace

Dynamic queries

In real-world scenarios, most queries in the context of an app need some dynamic parameterization. This is achieved by adding another annotation. In the following example, we introduce a new parameter, logLevel, and define a default value ERROR by adding @param(“loglevel”, “ERROR”).

When saving the file, the generated function is automatically updated, and the new parameter can be used within the code.

DQL query in function screenshot in Dynatrace

Best practices when working with DQL

We recommend organizing multiple queries within a single file or across different DQL files to enhance the workspace structure. To seamlessly integrate query functions into React state management, utilize our react-hooks Dynatrace SDK package. This aids in effectively handling execution, loading, and error states.

The previously described process of generating a function for your query additionally produces a function named getQueryLogsByLevel, which returns the query along with the specified parameters. This feature proves particularly useful if you plan to implement an IntentButton. Such a button allows users to seamlessly open the query in another application, like Dynatrace Notebooks, for in-depth analysis and further exploration.

Working with app functions

App functions represent the backend, including all the business logic, of a Dynatrace app. They’re used for access to third-party APIs, heavy data processing and manipulation, or encapsulating functionality that needs elevated access rights.

You can create a new app function by selecting the Generate app function and entering a name. Once the function is generated, you can test the functionality by clicking the Play button next to the function name and reviewing the output in the console.

Now it’s your turn

Are you interested in learning more? Watch the latest video from Inside Dynatrace Apps, and see a live demo of the Dynatrace Apps VS Code extension.

If you want to start your own developer journey, head to the Microsoft Visual Studio Marketplace, download the extension, and benefit from increased developer efficiency thanks to simplified app configuration, effortless integration of DQL, and enhanced functionality through Dynatrace functions.

Are you new to the topic of developing apps for the Dynatrace platform? Head to our Developer Portal and learn how to get started by watching our new tutorial.

Happy coding!

Ready to try out the VS Code extension yourself? Install the Dynatrace Apps extension from Microsoft Visual  Studio Marketplace.

The post Turbocharge Dynatrace app development with the new Visual Studio Code extension `Dynatrace Apps` appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/turbocharge-dynatrace-app-development-with-the-new-visual-studio-code-extension-dynatrace-apps/feed/ 0
Dynatrace® Apps showcase: Akamas Kubernetes optimization https://www.dynatrace.com/news/blog/dynatrace-apps-showcase-akamas-kubernetes-optimization/ https://www.dynatrace.com/news/blog/dynatrace-apps-showcase-akamas-kubernetes-optimization/#respond Tue, 14 May 2024 07:58:16 +0000 https://www.dynatrace.com/news/?p=63997 Akamas optimization opportunities

Akamas is an application optimization technology company and a Dynatrace partner. The Akamas software platform was built by performance engineering experts to redefine what organizations can achieve with AI-driven optimization, enabling enterprises and online businesses to deliver unprecedented cost savings, service performance, and resilience for Kubernetes-based applications.

The post Dynatrace® Apps showcase: Akamas Kubernetes optimization appeared first on Dynatrace news.

]]>
Akamas optimization opportunities

An earlier blog post introduced how Akamas helps optimize Kubernetes clusters without “breaking the bank.” As one of the first Dynatrace partners, Akamas used its domain expertise and unique product capabilities to build a custom app on the Dynatrace platform. The app empowers platform engineering teams with insights into infrastructure health, FinOps, and security combined with alerting and automatic lifecycle management. In this way, the app reduces costs and improves the reliability and performance of all Kubernetes applications.

For example, the Akamas app optimized a Dynatrace-monitored Kubernetes environment, found 22 workloads with reliability issues and 53 workloads with sub-optimal performance, and identified potential savings of about $50,000 per month.

Insights into your Kubernetes environment

Leveraging Dynatrace observability data, Akamas analyzes all available Kubernetes workloads and identifies optimization opportunities such as cost reduction, reliability, and performance improvements.

As you can see in the screenshot below, you get a summary of all optimization opportunities identified in the environment. The table below lists all the individual Kubernetes workloads that can be optimized.

There is a cost reduction opportunity for the notification workload of $2,300 per month and other opportunities related to performance improvements and reliability issues due to misconfigured resource settings.

Overview of optimization opportunities and Kubernetes workloads
Figure 1: Overview of optimization opportunities and Kubernetes workloads

Optimize your workloads

Select Optimize next to the notification workload to open the Optimize workload page, which offers options for optimizing the workload. Akamas is a goal-driven optimization solution, which means you choose if you want to improve the application’s performance or lower the cost, which in Kubernetes translates to reducing the resource requirements of your containers for CPU, memory requests, and memory limits.

Goal-driven optimization of Kubernetes workloads.
Figure 2: Goal-driven optimization of Kubernetes workloads.

When you manually reduce your container’s resources to save costs, you risk impacting service performance and reliability. To avoid slowing down your apps or—even worse—harming SLOs, you can define constraints that need to be considered, such as response time or error rate. Once constraints are set, Akamas AI considers application-level performance signals so cost-reduction recommendations don’t impact your SLOs.

Before you can start optimizing, you need to define the scope. When you choose Container, Akamas tunes the Kubernetes CPU and memory limits. It also supports full stack optimization, which optimizes JVM parameters such as maximum heap size and garbage collection.

Stay in control: Monitor the optimization process

Returning to the overview, you can switch to the Optimizations tab, which summarizes all your running optimizations, including the optimization created in the previous step.

Overview of optimization opportunities
Figure 3: Overview of optimization opportunities

Select See details for any service to dive into more details about running optimization tasks. Besides showing a short summary, including optimization goals, constraints, and scope, you can dive into cost trends and the relevant SLOs.

Optimizing the Kubernetes workload cartservice
Figure 4: Optimizing the Kubernetes workload `cartservice`

In this example, despite the cost going down (from over $80 down to about $50), the application performance was not impacted and stayed well below the defined threshold of 340 milliseconds.

Apply configurations

Akamas can apply configurations automatically or suggest configurations that SRE teams can use for further manual improvements. In the example below, you can see the current CPU and memory limits and the new values suggested by Akamas: change the server.cpu_limit from 1,000 down to 982 millicores and reduce the server.memory_limit by 53 MB. Select the button in the top-right of the Pending Recommendation pane to reveal the kubectl command, which you can use to apply the suggested recommendations.

Additional recommendations for manual improvements
Figure 5: Additional recommendations for manual improvements

Summary

The Akamas app helps achieve three critical goals by optimizing Kubernetes application configurations, harnessing the power of artificial intelligence of the Akamas platform, and leveraging the capabilities of the Dynatrace platform:

  • Cost reduction: Identify opportunities to trim unnecessary expenses related to Kubernetes workloads. Imagine saving thousands of dollars each month by fine-tuning your resource allocations.
  • Reliability enhancement: Pinpoint reliability issues within your Kubernetes environment and ensure your applications run smoothly, minimizing downtime and improving overall system stability.
  • Performance optimization: Fine-tune your software stack configurations to maximize performance parameters, ensuring applications meet or exceed their service level objectives (SLOs).

This is a perfect example of how to easily build custom apps on top of observability data stored within Dynatrace and leverage its enterprise-grade platform. Utilize the power of the Dynatrace platform, with its easy-to-use building blocks, to address specific use cases based on your business requirements.

Interested in learning more?

Watch the recording of one of this year’s Perform breakout sessions, where Stefano Doni, CTO and Co-founder of Akamas, and Alois Mayr, Principal Product Manager at Dynatrace, gave a quick intro to how Dynatrace and Akamas work better together, introducing the new Dynatrace Kubernetes monitoring app and how Akamas helps optimize your Kubernetes environment.

Video thumbnail

Start building your own Dynatrace app to address the specific needs of your users:

Are you interested in learning more about Akamas and why Dynatrace and Akamas are better together? Have a look at the Dynatrace Hub, watch a video tour, or contact Akamas directly to book a demo and get more insights into how Akamas can optimize your Kubernetes environment.

Visit Dynatrace Developer to learn about the tools and technologies of Dynatrace AppEngine.

The post Dynatrace® Apps showcase: Akamas Kubernetes optimization appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-apps-showcase-akamas-kubernetes-optimization/feed/ 0
How platform engineering and IDP observability can accelerate developer velocity https://www.dynatrace.com/news/blog/how-platform-engineering-can-accelerate-developer-velocity/ https://www.dynatrace.com/news/blog/how-platform-engineering-can-accelerate-developer-velocity/#respond Wed, 06 Mar 2024 20:40:49 +0000 https://www.dynatrace.com/news/?p=62869 Data privacy by design, CrowdStrike

At Dynatrace Perform 2024, Dynatrace colleagues Andreas Grabner and Adam Gardner discussed how platform engineering accelerates developer velocity.

The post How platform engineering and IDP observability can accelerate developer velocity appeared first on Dynatrace news.

]]>
Data privacy by design, CrowdStrike

As organizations look to expand DevOps maturity, improve operational efficiency, and increase developer velocity, they are embracing platform engineering as a key driver. Indeed, recent research found that 54% of organizations are investing in platforms to enable easier integration of tools and collaboration between teams involved in automation projects.

Platform engineering creates and manages a shared infrastructure and set of tools, such as internal developer platforms (IDPs), to enable software developers to build, deploy, and operate applications more efficiently. The goal is to abstract away the underlying infrastructure’s complexities while providing a streamlined and standardized environment for development teams. As a result, teams can focus on writing code and building features, rather than dealing with infrastructure nuances.

During a breakout session at Dynatrace Perform 2024, Dynatrace DevSecOps activist Andreas Grabner and staff engineer Adam Gardner demonstrated how to use observability to monitor an IDP for key performance indicators (KPIs). The pair showed how to track factors, including developer velocity, platform adoption, DevOps research and assessment metrics, security, and operational costs.

Recent Dynatrace research has found that only 40% of a typical engineer’s time is spent on productive tasks, and 36% of developers resign because of a bad developer experience, Grabner noted. “If your developers are leaving the company, the IDP may have something to do with it,” he said.

Platform engineering: Build for self-service

Self-service deployment is a key attribute of platform engineering. It gives developers the means to create environments and toolsets unique to their projects.

“[An IDP] must be a product that developers want to use because it helps them get the job done,” Grabner said. “It makes them more productive . . . and reduces the complexity of things such as reading a new app or service. They shouldn’t worry about the platform; they should just start writing code.”

Because of their versatility, teams can use IDPs for all types of software engineering projects, not just those in cloud-native scenarios. IDPs can eliminate much of the administrative minutiae that stalls development projects. Grabner gave the example of one Dynatrace banking customer who built an IDP that enables developers to provision new Microsoft Azure machines or Chef policies without administrative help. “IDPs are not constrained to building microservices or a new serverless app,” Grabner noted.

Before putting an IDP in place, organizations must encourage their platform engineering teams to adopt a product mindset with feedback loops between developers and users. They should also establish milestones to ensure the built product solves a defined business problem.

Reference IDP with Dynatrace

The Dynatrace IDP encompasses platform services, delivery services, and access to observability and automation tools. The Dynatrace Operator automatically ingests all observability data from OpenTelemetry and Prometheus. Furthermore, OneAgent® software observes and gathers all remaining workload logs, metrics, traces, and events.

Automate deployment for faster developer velocity

Additionally, the IDP used during the session connects to the open source Backstage developer portal platform and a library of templates stored in a GitLab repository. The templates can deploy automatically into the development environment with just a few clicks.

Argo works in a GitOps fashion to automate the deployment of files stored in Git. “Argo has an eagle eye on the Git repository,” Gardner said. “Every time something changes, it’s synced to Kubernetes.”

Backstage holds many of an organization’s critical development resources that must be treated with the same respect as business-critical data. Observability is not only about measuring performance and speed but also about capturing granular business analytics to support data-driven decision-making. These metrics can include how many people are using the IDP, how quickly the tasks are running in the IDP, and more. “That means making it available, resilient, and secure,” Grabner said.

Intelligent monitoring is also crucial. “If you don’t monitor, you risk building a product that nobody needs,” Grabner continued.

Observability is a critical component of an IDP. It illuminates the activity of components such as Backstage, GitHub, Argo, and other tools. Service-level objectives (SLOs) are similarly important. SLOs help developers to accelerate their velocity and remain productive with an optimally functioning platform.

Test continuously

Synthetic testing simulates user behaviors within an application or service to pinpoint potential problems. This process is vital to an IDP’s effectiveness. An observability solution can monitor both synthetic and real-user tests to verify an application is on track.

GitLab, a source code repository and collaborative software development platform for DevOps and DevSecOps projects, is populated with a set of pre-filled templates. The combination gives developers a unique set of tools they can deploy on a self-service basis with full monitoring by Dynatrace.

“Every time [developers] pick a template in Backstage, they get their own version of the Git repository based on the template. Then, Argo deploys the app,” Grabner said. “It has worked kind of flawlessly.”

Observability at the core

How we built the IDP

Platform engineering is about being responsible for making sure platforms are available,” Gardner said. “Dynatrace can tell us whether Argo is up and whether it’s killing GitHub with too many syncs. It lets us see events such as starts and traces in a standardized manner.” This certainty can accelerate developer velocity and improve the developer experience, resulting in better software and happier, more productive developers.

Dynatrace has made the reference IDP architecture available on GitHub for anyone to use. It includes a notebook with configuration and deployment instructions.

“It explains every single step that was involved in building the IDP, creating the configuration, and setting up Argo,” Gardner said. “You can launch a code space that starts a container that shows you everything about how an app was built and deployed.”

Curious to learn more about observability to optimize KPI success? Check out the Perform 2024 session: Observability guide to platform engineering.

FAQs about platform engineering

What is the primary role of an internal developer platform (IDP)?

An IDP is a shared infrastructure and set of tools created and managed by platform engineering teams. By providing a self-service, standardized environment, IDPs enable software developers to more efficiently build, deploy, and operate applications. IDPs reduce the need for developers to manage underlying infrastructure nuances, thereby simplifying workflows and accelerating application delivery.

How does platform engineering improve software development?

Platform engineering accelerates developer velocity by providing a streamlined, standardized environment that abstracts away infrastructure complexities. This allows developers to focus on writing code and building features more efficiently.

It improves the developer experience by offering self-service tools and automating tasks, reducing the cognitive load and administrative minutiae that can otherwise hinder productivity and frustrate developers.

Why is observability crucial for successful platform engineering?

Observability is crucial for successful platform engineering because it provides deep insights into the performance, health, and activity of the internal developer platform and its components. It allows platform teams to monitor key performance indicators (KPIs) like platform adoption and operational costs, helping to support the platform’s availability, resiliency, and security. Continuous monitoring helps verify the effectiveness of the IDP and ensures developers maintain optimal productivity.

What are the key principles for building an effective internal developer platform?

Building an effective internal developer platform involves adopting a product mindset, treating developers as internal customers, and incorporating their feedback through continuous loops. Here are some principles to keep in mind as you build:

  • Provide clear self-service capabilities.
  • Automate repetitive tasks.
  • Focus on solving common developer pain points.

The platform should also emphasize scalability, security, and compliance, while fostering a culture of collaboration and knowledge sharing within the engineering organization.

Discover how unified observability unlocks platform engineering success in the free ebook: Driving DevOps and platform engineering for digital transformation.

The post How platform engineering and IDP observability can accelerate developer velocity appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-platform-engineering-can-accelerate-developer-velocity/feed/ 0
Container monitoring for VA Platform One helps VA achieve workload performance https://www.dynatrace.com/news/blog/container-monitoring-for-va-platform-one/ https://www.dynatrace.com/news/blog/container-monitoring-for-va-platform-one/#respond Wed, 28 Feb 2024 14:00:07 +0000 https://www.dynatrace.com/news/?p=62652 container monitoring, VA Platform One

At the U.S. Department of Veterans Affairs (VA), Dynatrace is helping VA monitor the development and testing of applications within VA Platform One (VAPO) containers to ensure teams are working better, safer, and faster.

The post Container monitoring for VA Platform One helps VA achieve workload performance appeared first on Dynatrace news.

]]>
container monitoring, VA Platform One

Through containers developed within VA Platform One (VAPO), the development team at the U.S. Department of Veterans Affairs (VA) is packaging application code along with its libraries and dependencies within an executable software unit. The containers can run anywhere, whether a private data center, the public cloud or a developer’s own computing devices. Dynatrace container monitoring supports customers as they collect metrics, traces, logs, and other observability-enabled data to improve the health and performance of containerized applications.

At Perform 2024, Matthew Fuqua, Technical Lead for VAPuO, sat down with Willie Hicks, Dynatrace Public Sector Chief Technologist, to discuss his team’s role in VA’s modernization journey—and how Dynatrace has significantly accelerated its progress while unleashing new capabilities.

What is VAPO?

VA Platform One (VAPO) is a comprehensive application development and delivery platform. The VAPO platform is used to develop, test, and deploy containerized application and middleware workloads that support 400,000 VA employees and 20 million veterans. VAPO relies on Dynatrace and its integration with Red Hat to monitor application development and testing within containers to ensure optimal performance and security.

“It’s an enterprise product that we use to help modernize the VA,” Fuqua said. “It’s one of our biggest modernization efforts, and it’s saving us money while providing better, quicker, and faster healthcare to our veterans.”

VAPO is available in both Microsoft Azure and AWS. It’s supported by the VA Enterprise Cloud (VAEC), a multi-vendor, FedRAMP High environment for hosting VA applications in the cloud.

VA Platform One enables rapid application development with “security first” oversight

VA Platform One provides developers with all the tools required to create, test, and push out applications swiftly. With VA’s heavy focus on security, the platform enables developers to incorporate security testing for applications during development. “In the development environment, you see exactly where in the pipeline security issues exist, and you can address them right there, so it speeds up development,” Fuqua said. “It’s easy to produce a container that we can rapidly test, then shift to pre-production and production. It’s helping us build applications more efficiently and faster and get them in front of veterans.”

Supporting application modernization with a focus on user experience (UX)

While VA is developing new applications, they also have monolithic applications to modernize. That modernization effort starts with Dynatrace. “We can look at user experiences and CPU memory usage over time,” Fuqua said. “Then, we can show our results once we containerize: what we’re saving, if the user experience is better or worse and, if not, we can improve that. We use Dynatrace as part of our migration pattern, so we can see how well we’re doing our jobs.”

With Dynatrace, Fuqua’s team has full observability of the applications themselves in a dashboard, so his team can make sure users are getting all they need. “We can log into an actual experience and it captures everything,” he said.

AI and automation simplify container monitoring and satisfy the “need for speed”

Container monitoring is inherently challenging because of containers’ highly dynamic nature. Dynatrace artificial intelligence (AI)-powered root cause analysis brings real-time insights and actionable answers to fix issues, automating operations so the VAPO team can focus on innovation. “We want to get to a point where we can identify something, then the platform automatically creates a ticket that goes to an approver,” Fuqua said. “If the approver says, ‘do it,’ then it schedules the action.”

With a consistent focus on UX, VA is leveraging synthetic monitoring to gain an accurate view of UX and then using automation to scale containers ahead of demand. “We’re using automation to kick off scaling events,” he said. “We want to be there in time.”

The veteran is the mission

As Hicks summarized, VA’s mission is focused on the veteran. “Your most important end user is the veteran,” he said.

The VAPO team appreciates how Dynatrace puts everything “all in one place” while making it so much easier to solve problems. “This is a continuous process,” Fuqua said. “You’re not going to wave a magic wand and have a container. With Dynatrace, we’re getting in there and doing things in phases and continuously improving.”

If you’d like to know more about how Dynatrace can help your government agency achieve this level of optimal performance quality, efficiency, and security, please contact us.

To hear the full story about how Dynatrace observability is empowering the VAPO team to overcome their challenges and achieve mission goals, watch the full session, Efficiency unleashed: VAPO’s dynamic approach to containerized workloads with Dynatrace.

The post Container monitoring for VA Platform One helps VA achieve workload performance appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/container-monitoring-for-va-platform-one/feed/ 0
Unlocking collaborative partner innovation: 2023 Dynatrace Partner Pro Club winners https://www.dynatrace.com/news/blog/unlocking-collaborative-partner-innovation-2024-dynatrace-partner-pro-club-winners/ https://www.dynatrace.com/news/blog/unlocking-collaborative-partner-innovation-2024-dynatrace-partner-pro-club-winners/#respond Mon, 26 Feb 2024 14:18:03 +0000 https://www.dynatrace.com/news/?p=62582 hybrid cloud network

In today’s complex digital landscape, organizations need to be able to scale and innovate in order to compete. The collaborative partner innovation showcased between Dynatrace and its strategic partnerships is a critical piece of enabling growth for our customers. While the curtains have fallen on Perform 2024, the Dynatrace annual user conference, the commitment and […]

The post Unlocking collaborative partner innovation: 2023 Dynatrace Partner Pro Club winners appeared first on Dynatrace news.

]]>
hybrid cloud network

In today’s complex digital landscape, organizations need to be able to scale and innovate in order to compete. The collaborative partner innovation showcased between Dynatrace and its strategic partnerships is a critical piece of enabling growth for our customers.

While the curtains have fallen on Perform 2024, the Dynatrace annual user conference, the commitment and collaborative nature of our partner relationships is consistently being demonstrated. During the conference, we unveiled the winners of the Dynatrace Partner Pro Club app competition. This challenge invited our partners to create impactful apps that solve real-world customer use cases using Dynatrace AppEngine. Below are the winners.

Elevating user experience and design with collaborative partner innovation

The top submissions stood out by solving complex business challenges and focusing on creating an exceptional user experience. The winning User Flow Analytics app by Andrea Caria of Spindox introduces a visual analysis of user navigation within web, mobile, or custom applications, presented through dynamic Sankey diagrams and funnels. This innovative approach enhances the understanding of user journeys through compelling visuals, setting the standard for future app developments.

The User Flow Analytics app was created to address real-life business challenges. For example, a multinational manufacturing company wanted to improve the performance and user experience of their mobile apps and websites. They asked Spindox to evaluate their existing product analytics software and suggest potential alternatives. After careful consideration, Spindox recommended the Dynatrace unified observability and security platform. Spindox also realized that if they could extend the Dynatrace Platform to cover user-funnel analytics and user-navigation Sankey diagrams, it would help the customer consolidate their tooling. Using Dynatrace AppEngine, Spindox was able to implement funnel analytics according to the customer’s specifications, providing them with all the information they needed within a single platform.

AppEngine is a stunning step forward in data monitoring because it allows you to merge the customization of coding with the high-quality data stored in Dynatrace, unlocking the limits of analysis beyond imagination. The more time you spend developing your application, the more you understand its potential by assembling the amazing number of ready-made components that Dynatrace provides.” — Andrea Caria, observability solution analyst at Spindox. 

Recognizing the runners up

While there was an overwhelming amount of quality submissions, we have to recognize two outstanding runners-up. Gil Givati’s (Matrix) HALO app impressed by allowing users to visually create holistic objectives for applications or products. Meanwhile, Michiel Otten’s (Eviden) Cross Platform Insights app garnered applause for providing a unified view of health across multiple Dynatrace environments.

“We were truly impressed by the apps showcased in the Dynatrace Pro Club app competition. Our partners combined their technical prowess and industry insight to engineer apps that effectively utilize observability, security, and business data to solve key customer challenges. I’m excited about the prospects of AppEngine empowering the creation of numerous groundbreaking custom apps.” — Ahmed El-Jafoufi, global partner enablement architect at Dynatrace.

A collective leap forward

AppEngine empowers partners to create custom solutions that are secure and enterprise-grade with the familiar Dynatrace out-of-the-box look and feel.

Our competition wasn’t just about winning; it was about pushing the boundaries of what’s possible in observability and security. The submissions we received were a testament to the power of innovative partner collaboration, and the shared commitment of Dynatrace and its partners to lead the way in shaping the future of digital experiences.

All of this year’s participants contributed to making this competition a success. The dedication of our partners to excellence elevates the Dynatrace ecosystem and continues to set the standard for what’s possible in the dynamic world of unified observability and security.

Dynatrace Partner Pro Club

Partner Pro Club is an initiative to recognize and reward Dynatrace partners who invest the time and commitment to achieving their Dynatrace Professional certification. Members gain access to exclusive events and training, fostering a community of champions.

The post Unlocking collaborative partner innovation: 2023 Dynatrace Partner Pro Club winners appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/unlocking-collaborative-partner-innovation-2024-dynatrace-partner-pro-club-winners/feed/ 0
Dynatrace to acquire Rookout to deliver code debugging in production environments https://www.dynatrace.com/news/blog/dynatrace-to-acquire-rookout-for-code-debugging/ https://www.dynatrace.com/news/blog/dynatrace-to-acquire-rookout-for-code-debugging/#respond Mon, 31 Jul 2023 20:30:15 +0000 https://www.dynatrace.com/news/?p=58928 Dynatrace + Rookout

Dynatrace recently announced it has signed a definitive agreement to acquire Rookout, a company that enables developers to quickly troubleshoot and debug actively running code in Kubernetes-hosted cloud-native applications. Adding Rookout to the Dynatrace platform will help developers accelerate innovation and deliver flawless and secure releases.

The post Dynatrace to acquire Rookout to deliver code debugging in production environments appeared first on Dynatrace news.

]]>
Dynatrace + Rookout

Developers are increasingly responsible for ensuring the quality and security of code throughout the software lifecycle. Traditional tools and approaches, however, only allow debugging in pre-production environments. Debugging in production often requires shutting down services. This can disrupt the users of the running application, slow down the application’s performance, or even crash it altogether.

Developer-first observability

Adding Rookout to the Dynatrace platform will provide developers with increased code-level observability of Kubernetes-hosted production environments. This will also add interactivity and control to troubleshooting and debugging in production, drastically reduce the need to replicate issues in pre-production environments and improve collaboration across development, IT, and security teams.

As Shahar Fogel, CEO at Rookout, said, “Our mission is to make debugging easy and fast for developers with state-of-the-art quality and a simple experience. We believe integrating Rookout into the Dynatrace platform and leveraging the artificial intelligence and automation capabilities Dynatrace is known for will accelerate this mission. This will also create a new standard for how engineers use developer-first, cloud-native observability to improve productivity by enabling them to spend less time on manual activities and more time delivering business value.”

Observability and security continue to converge

“Development teams are increasingly expected to incorporate observability and security capabilities into their solutions (shift-left) as well as perform testing, quality, and performance evaluation in production environments (shift right),” said Bernd Greifeneder, CTO at Dynatrace. “We believe acquiring Rookout will accelerate this process by providing our customers with developer-observability solutions that scale from a developer’s integrated development environment, or IDE, and are designed to enable their organizations to meet enterprise governance requirements. Our experience is that Rookout enables developers to troubleshoot and debug issues in production significantly faster than traditional tools and approaches, dramatically reducing the time they spend on maintenance activities.”

Closing of the proposed transaction is subject to customary closing conditions and is expected to occur later in the company’s second quarter, which ends on September 30, 2023. The proposed transaction will not have a material impact on Dynatrace’s fiscal year 2024 financials and will be funded from cash on hand.

The post Dynatrace to acquire Rookout to deliver code debugging in production environments appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-to-acquire-rookout-for-code-debugging/feed/ 0
Dynatrace to Acquire Rookout to Deliver Code Debugging in Production Environments https://www.dynatrace.com/news/press-release/dynatrace-to-acquire-rookout/ Mon, 31 Jul 2023 20:15:44 +0000 https://www.dynatrace.com/news/?post_type=press-release&p=58982 WALTHAM, Mass., July 31, 2023 – Dynatrace (NYSE: DT), the leader in unified observability and security, today announced it has signed a definitive agreement to acquire Rookout, a provider of enterprise-ready and privacy-aware solutions that enable developers to quickly troubleshoot and debug actively running code in Kubernetes-hosted cloud-native applications. The addition of Rookout to the […]

The post Dynatrace to Acquire Rookout to Deliver Code Debugging in Production Environments appeared first on Dynatrace news.

]]>
WALTHAM, Mass., July 31, 2023Dynatrace (NYSE: DT), the leader in unified observability and security, today announced it has signed a definitive agreement to acquire Rookout, a provider of enterprise-ready and privacy-aware solutions that enable developers to quickly troubleshoot and debug actively running code in Kubernetes-hosted cloud-native applications. The addition of Rookout to the Dynatrace® platform will help developers accelerate innovation and delivery of flawless and secure releases.

Developers are increasingly responsible for ensuring the quality and security of code throughout the lifecycle. Traditional tools and approaches, however, only allow debugging in pre-production environments. Debugging in production often requires shutting down services. This can disrupt the users of the running application, slow down the application’s performance, or even crash it altogether.

Adding Rookout to the Dynatrace platform will provide developers with increased code-level observability into production environments. This will also add interactivity and control to troubleshooting and debugging in production and drastically reduce the need to replicate issues in pre-production environments. The addition of Rookout to the Dynatrace platform will also improve collaboration across development, IT, and security teams by empowering them with a single platform for observability and security analytics and automation.

“Development teams are increasingly expected to incorporate observability and security capabilities into their solutions (shift-left) as well as perform testing, quality, and performance evaluation in production environments (shift-right),” said Bernd Greifeneder, CTO at Dynatrace. “We believe acquiring Rookout will accelerate this process by providing our customers with developer-observability solutions that scale from a developer’s integrated development environment, or IDE, and are designed to enable their organizations to meet enterprise governance requirements. Our experience is that Rookout enables developers to troubleshoot and debug issues in production significantly faster than traditional tools and approaches, dramatically reducing the time they spend on maintenance activities.”

“Our mission is to make debugging easy and fast for developers with state-of-the-art quality and a simple experience,” said Shahar Fogel, CEO at Rookout. “We believe integrating Rookout into the Dynatrace platform and leveraging the AI and automation capabilities Dynatrace is known for will accelerate this mission. This will also create a new standard for how engineers use developer-first, cloud-native observability to improve productivity by enabling them to spend less time on manual activities and more time delivering business value.”

Dynatrace plans to provide a seamless experience for customers by embedding Rookout into its unified observability and security platform.

Closing of the proposed transaction is subject to customary closing conditions and is expected to occur later in the company’s second quarter, which ends on September 30, 2023. The proposed transaction will not have a material impact on Dynatrace’s fiscal year 2024 financials and will be funded from cash on hand.

Cautionary Language Concerning Forward-Looking Statements

This press release includes certain “forward-looking statements” within the meaning of the Private Securities Litigation Reform Act of 1995, including statements regarding the anticipated benefits of the Rookout acquisition, the expected time for closing of the proposed acquisition, the impact of the proposed acquisition on Dynatrace’s fiscal year 2024 financials, and the source of funding of the transaction. These forward-looking statements include all statements that are not historical facts and statements identified by words such as “will,” “expects,” “anticipates,” “intends,” “plans,” “believes,” “seeks,” “estimates” and words of similar meaning. These forward-looking statements reflect our current views about our plans, intentions, expectations, strategies, and prospects, which are based on the information currently available to us and on assumptions we have made. Actual results may differ materially from those described in the forward-looking statements and will be affected by a variety of risks and factors that are beyond our control, including risks set forth under the caption “Risk Factors” in our Annual Report on Form 10-K for the fiscal year ended March 31, 2023 and our other SEC filings. We assume no obligation to update any forward-looking statements contained in this document as a result of new information, future events or otherwise.

The post Dynatrace to Acquire Rookout to Deliver Code Debugging in Production Environments appeared first on Dynatrace news.

]]>
Critical app observability in government including ArcGIS https://www.dynatrace.com/news/blog/critical-app-observability-in-government-including-arcgis/ https://www.dynatrace.com/news/blog/critical-app-observability-in-government-including-arcgis/#respond Tue, 18 Jul 2023 15:59:00 +0000 https://www.dynatrace.com/news/?p=58505 City map

For cities, counties, and states that use geographic information system (GIS) apps such as ArcGIS to drive mission critical services, application resilience is essential. With so much at risk during an emergency, ensuring performance apps don’t lag or crash when they’re most needed is vital. Advanced observability can eliminate blind spots surrounding application performance, health, […]

The post Critical app observability in government including ArcGIS appeared first on Dynatrace news.

]]>
City map

For cities, counties, and states that use geographic information system (GIS) apps such as ArcGIS to drive mission critical services, application resilience is essential. With so much at risk during an emergency, ensuring performance apps don’t lag or crash when they’re most needed is vital.

Advanced observability can eliminate blind spots surrounding application performance, health, and behavior for these critical applications and the infrastructure that supports them. Suppose ArcGIS has an unexpected outage. Usually, IT teams scramble with time-consuming and expensive fire drills to find and repair the root cause. Envision if IT teams could be proactively alerted before the field teams discover they can’t complete their gas line maintenance work or load the maps required to fix the freshwater infrastructure in the county. Alerts that deliver actionable data, such as ranking issues and pinpointing the root cause, enable teams to prioritize their efforts on the issues with the most significant user impact and accelerate the resolution.

When considering the financial costs of each hour of downtime, observability delivers a compelling return on investment (ROI). It empowers the team to recover faster as well as proactively prevent outages from happening in the first place.

Gaining a complete picture of app health in hybrid environments

As a starting point, teams should know if ArcGIS and other critical apps are available. While uptime is essential, availability doesn’t ensure optimal performance. This is why application performance monitoring (APM) is essential.

The challenge is that state and local governments operate in highly complex systems — legacy data centers or hybrid environments — crossing multiple clouds. That’s why it’s difficult to understand the root cause of performance anomalies. IT teams need full situational awareness of all assets and an understanding of how they’re performing, including third party vendor apps that were not built in-house.

Agency leaders have just as much to gain from adopting observability solutions as those on the front lines of DevSecOps efforts. Adopting APM can enable staff to improve citizen experience and by extension, agency reputation. Further, APM can improve employee satisfaction by eliminating the manual work of scrubbing every log to correlate events to discover what’s causing a problem with mission-critical software such as ArcGIS.

Understand user interactions with critical GIS software

While it’s important to have an application performance strategy, a top mission for state and local governments is to enhance the citizen experience. An advanced observability strategy includes end user experience monitoring to evaluate digital experiences—such as a website or mobile application—from the user’s perspective.

For example, Real User Monitoring (RUM) can provide an accurate picture of employee or citizen experience with an app. User Session Replay enables IT to resolve customer complaints by watching visual replays to determine exactly what went wrong during a session.

With these tools, agency leaders can empower IT teams to understand how citizens are using digital touchpoints, the obstacles they encounter as they’re using them, and how to eliminate those points of friction, whether in the app, database, or the infrastructure that hosts them.

Automatically discover how infrastructure and database changes impact critical app performance

Full-stack observability gives complete, real-time insight into applications’ behavior, health, the underlying infrastructure, and all other services on which the applications might depend.

Infrastructure monitoring automatically analyzes key health metrics and discovers performance problems caused by infrastructure bottlenecks or changes. As a result, IT teams can proactively identify and resolve potential infrastructure issues with ArcGIS and other critical applications before they impact performance and the citizen experience.

Similarly, observability provides detailed telemetry about the database performance that applications like ArcGIS depend on, including the health of the underlying infrastructure, the state of the database instance, and deep-dive analysis of each database statement and its performance.

Agency leaders can cut costs and improve productivity by eliminating the monitoring tool sprawl that pushes teams to work in silos. With a single source of truth for root cause analysis, IT and DevOps teams can quickly get on the same page about what needs to be done and who’s responsible for it.

Automate vulnerability detection for apps in the IT environment

With growing ransomware attacks, state and local governments must ensure applications don’t contain vulnerabilities that could allow illicit access to sensitive data, unauthorized code modification, or resource hijacking. Ensuring secure applications amid rising complexity is crucial to the digital transformation journey.

Further, applications aren’t as simple as they used to be, and ensuring they’re secure has become more challenging. Whether built in-house or purchased, any application can introduce vulnerabilities into the environment. However, it is more challenging to discover and remediate vulnerabilities when you don’t own the application code.

Full-stack runtime vulnerability analysis provides a holistic view and analysis across all layers of the application ecosystem, so teams have a single source of truth to provide the answers they need. Imagine the team discovers a vulnerability within a third-party app. They can immediately alert the vendor and hold them accountable for the fix to secure your environment. This approach helps agencies eliminate security blind spots and validate that vendors meet the agreed-upon service level agreements.

Ensure IT teams have actionable and scalable data to support the agency and its mission

The challenges agencies face in accelerating digital transformation are obvious, but meaningful answers are hard to come by. Purpose-built monitoring software often becomes shelfware because it’s too difficult to deploy, configure, or understand.

When building the technology stack, agencies must think in the context of the citizen and employee experience. ArcGIS software enables state and local agencies to enhance community service offerings, provide transparency, and engage citizens. Modern observability ensures that IT teams have the tools to appropriately manage and optimize critical application resilience, including ArcGIS performance.

A full stack observability solution enables staff to see it all, top to bottom, from end-user experience to infrastructure health. It also helps teams understand how everything is connected — including all the relationships and interdependencies among layers, components, or pieces of code. These capabilities can help your agency digitally transform faster and more easily — even as cloud complexity increases to ensure that the agency can provide the best possible services to citizens and the right tools to improve employee satisfaction.

The post Critical app observability in government including ArcGIS appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/critical-app-observability-in-government-including-arcgis/feed/ 0
OpenShift vs. Kubernetes: Understanding the differences https://www.dynatrace.com/news/blog/openshift-vs-kubernetes/ https://www.dynatrace.com/news/blog/openshift-vs-kubernetes/#respond Wed, 07 Jun 2023 19:30:18 +0000 https://www.dynatrace.com/news/?p=58129 OpenShift vs. Kubernetes

Many organizations consider OpenShift vs. Kubernetes for managing containerized apps at scale. What are the differences between OpenShift and Kubernetes?

The post OpenShift vs. Kubernetes: Understanding the differences appeared first on Dynatrace news.

]]>
OpenShift vs. Kubernetes

If you’re evaluating container orchestration software to manage containerized applications at scale, you may be wondering about the differences between OpenShift and Kubernetes. But as you contemplate OpenShift vs. Kubernetes, it’s important to understand what these container orchestration solutions are, how they relate, and their benefits and drawbacks.

A guide to container orchestration software

Container orchestration software automates the administration of containerized workloads and services, greatly reducing the time IT staff spend keeping an application environment running smoothly. Container orchestration allows an organization to digitally transform at a rapid clip without getting bogged down by slow, siloed development, difficult scaling, and high costs associated with optimizing application infrastructure.

As with Kubernetes vs. Docker, OpenShift vs. Kubernetes is a common debate, as they are two of the most widely used container orchestration tools. Although they share many features in common, there are also critical differences between OpenShift and Kubernetes. To put those differences into perspective, let’s look at what Kubernetes and OpenShift are and how they work.

What is Kubernetes?

Kubernetes is an open source container orchestration platform that enables organizations to automatically scale, manage, and deploy containerized applications in distributed environments. According to the Kubernetes in the Wild 2023 report, “Kubernetes is emerging as the operating system of the cloud.” In recent years, cloud service providers such as Amazon Web Services, Microsoft Azure, IBM, and Google began offering Kubernetes as part of their managed services. As a result of these services, organizations can further streamline the administrative overhead associated with application development.

Kubernetes containers are portable across environments, which enables developers to run them on nearly any type of infrastructure, whether in the cloud or locally. This flexibility helps organizations avoid vendor lock-in. Kubernetes also gives developers freedom of choice when selecting operating systems, container runtimes, storage engines, and other key elements for their Kubernetes environments. They can integrate their own applications in the Kubernetes API or use Kubernetes’ own tooling to roll out new features. One major Kubernetes advantage is its self-healing, continually making repairs and addressing failures that affect applications’ integrity.

Kubernetes architecture

That said, Kubernetes has some drawbacks. Its inherent complexity makes observability difficult — especially when used across highly distributed systems. IT teams can’t see into the internal state of Kubernetes containers, so they often collect a wide variety of telemetry data — such as logs, metrics, and distributed traces — to compensate for this lack of visibility.

While these data sources are helpful, they often can’t help IT understand the relationships and context necessary to quickly identify the root causes of application performance issues. As a result, organizations can have trouble transforming at scale, improving critical service-level agreements, and optimizing the user experience.

What is OpenShift?

Like Kubernetes, OpenShift is an open source Kubernetes-based container platform. OpenShift is developed by Red Hat and can run in a variety of environments — both cloud and on premises. In fact, it is a frequent choice for running Kubernetes on premises.

Because it’s based on Kubernetes, OpenShift provides containerization and orchestration of containerized workloads. Like Kubernetes, it allocates resources efficiently and ensures high availability and fault tolerance.

But OpenShift builds from there to provide integrated development tools, CI/CD (continuous integration/continuous deployment) pipelines, and built-in support for popular programming languages, frameworks, and databases. These tools enable OpenShift to support the entire application lifecycle, from development to production, including scaling, rolling updates, and version control. Without having to worry about underlying infrastructure concerns, such as storage, security, and lifecycle management, developers can focus on writing code.

Likewise, Red Hat OpenShift helps organizations administer Kubernetes more efficiently. For example, OpenShift simplifies Kubernetes management tools, giving developers everything they need to manage Kubernetes nodes, as well as the underlying control plane.

In addition, OpenShift provides numerous cloud services and self-managed deployment models to suit various applications and architectures, including the following:

  • OpenShift Container Platform (OCP). This self-managed offering can run on premises or in the cloud.
  • OpenShift Dedicated (OSD). The managed service runs on public clouds such as Amazon Web Services and Google Cloud.
  • Red Hat OpenShift Online (OSO). This fully managed service runs on Red Hat’s public cloud.

Despite its advantages, however, OpenShift also has its limitations. While Kubernetes supports all cloud and Linux distributions, making it widely accessible to organizations using various platforms, OpenShift supports only Red Hat distributions, such as Red Hat Enterprise Linux (RHEL), CentOS, and Fedora. As a commercial solution, OpenShift is also comparatively less flexible than open source Kubernetes, making it less customizable to an organization’s unique requirements.

OpenShift vs. Kubernetes: Weighing the key differences

While Kubernetes and OpenShift are both popular container orchestration platforms, they are used in slightly different ways and offer different features. Some of the key differences include the following:

  • Origin. Originally created by Google, Kubernetes is an open source project managed by the Cloud Native Computing Foundation (CNCF). OpenShift, on the other hand, is an open source Red Hat offering that is built on top of Kubernetes primarily on RHEL operating systems.
  • Ease of use. While Kubernetes offers increased flexibility and powerful features, it can be complex to set up and manage. In contrast, OpenShift provides a simplified, user-friendly interface, with built-in support for CI/CD pipelines.
  • Security. OpenShift has several built-in security features, while Kubernetes relies on the underlying infrastructure and additional tools for security. Additionally, OpenShift runs containers as a non-root user by default and provides additional security policies out of the box.
  • Networking. Kubernetes provides a basic networking model. However, it needs additional tools or plugins for more advanced networking features. OpenShift, on the other hand, includes a more advanced software-defined networking (SDN) solution, which supports network policies for finer control over container communication.
  • Updates and support. Kubernetes has frequent updates, which can sometimes lead to issues such as breaking changes. Red Hat OpenShift offers long-term support versions and commercial support.
  • Integration and extensions. Kubernetes is more of a bare-bones platform. Therefore, it relies on external tools and services for most integrations and extensions. Conversely, as a Red Hat offering, OpenShift provides built-in integration with other Red Hat products and offers a marketplace for third-party extensions.
  • Pricing. Unlike Kubernetes, which is a completely free and open source service, OpenShift has a pricing model for its enterprise version that includes additional features, support, and services.

A lesson in terminology: Kubernetes namespace vs. OpenShift project

It’s OpenShift vs. Kubernetes when it comes to terminology, too. In addition to the aforementioned feature differences, Kubernetes and OpenShift use different terminology, which can be confusing for organizations and practitioners alike.

For example, a namespace in Kubernetes is typically referred to as a project in OpenShift. Despite their similarities, there are a few differences between namespaces and projects, including the following:

  • Access control. In Kubernetes, users manage access control independently from namespaces. In contrast, OpenShift’s projects have a predefined set of permissions for project-level operations. This makes it easier for organizations to control who has access to what within a project.
  • Isolation. While teams use both namespaces and projects to isolate resources within a cluster, OpenShift’s projects provide additional features. These include the ability to limit the amount of resources that all containers can consume within a project.
  • User-friendly. Unlike Kubernetes namespaces, OpenShift projects are more user-friendly. When a user creates a project, for instance, they automatically become the project admin. With Kubernetes namespaces, this does not happen automatically.

Which container orchestration software is right for you?

If your organization needs a container orchestration solution with enterprise-level support and security, OpenShift is the clear choice. OpenShift offers a secure-by-default option to increase security, and its security policies are much stricter than Kubernetes. OpenShift is also a good choice if CI/CD is a priority for your organization.

Additionally, OpenShift is designed to meet the needs of industries with strong compliance and regulatory requirements, such as healthcare or finance. It addresses regulations such as the European Union’s General Data Protection Regulation and the U.S. Health Insurance Portability and Accountability Act.

On the other hand, Kubernetes is a strong option if you need more customization and flexibility, and you have in-house Kubernetes experts who can troubleshoot problems as they arise.

Kubernetes also works on the widest possible range of operating systems and platforms. If you’re a social media or gaming company that places an especially high priority on releasing updates at a rapid pace, Kubernetes may be a better fit.

Automatic and intelligent observability for OpenShift and Kubernetes

Whether you choose OpenShift vs. Kubernetes or vice versa, Dynatrace can make the most of your container orchestration solution. Dynatrace uses AIOps and cloud observability to combine metrics, logs, and traces with topology information, real user experience data, and meta information. With advanced observability of every Kubernetes cluster, pod, and node — and all connections and dependencies they touch — you can quickly pinpoint and solve performance problems as they arise.

Additionally, Dynatrace offers powerful monitoring capabilities for OpenShift, helping you manage costs, automate your operations, and release better software faster.

Whether using OpenShift or Kubernetes, the Dynatrace observability and security platform is the only Kubernetes monitoring system with continuous automation that identifies and prioritizes alerts from applications and infrastructure without changing code, container images, or deployments.

To learn more about how Dynatrace can help you achieve your container orchestration goals, check out our performance clinic, “Kubernetes platform observability with Dynatrace.”

The post OpenShift vs. Kubernetes: Understanding the differences appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/openshift-vs-kubernetes/feed/ 0
The top 7 Kubernetes challenges and how to solve them https://www.dynatrace.com/news/blog/the-top-seven-kubernetes-challenges/ https://www.dynatrace.com/news/blog/the-top-seven-kubernetes-challenges/#respond Tue, 06 Jun 2023 12:56:02 +0000 https://www.dynatrace.com/news/?p=58087 Weighing the top seven Kubernetes challenges

While Kubernetes offers many business benefits, it also has potential pitfalls. Discover the top seven Kubernetes challenges and how to gain control of your container environment.

The post The top 7 Kubernetes challenges and how to solve them appeared first on Dynatrace news.

]]>
Weighing the top seven Kubernetes challenges

Kubernetes has become the leading container orchestration platform for organizations adopting open source solutions to manage, scale, and automate application deployment. Adopting this powerful tool can provide strategic technological benefits to organizations — specifically DevOps teams. At the same time, it also introduces a large amount of complexity. This complexity has surfaced seven top Kubernetes challenges that strain engineering teams and ultimately slow the pace of innovation.

What is Kubernetes? And how does it benefit organizations?

Kubernetes is an open source container orchestration platform for managing, automating, and scaling containerized applications.

Containerized microservices have made it easier for organizations to create and deploy applications across multiple cloud environments without worrying about functional conflicts or software incompatibilities. This ease of deployment has led to mass adoption, with nearly 80% of organizations now using container technology for applications in production, according to the CNCF 2022 Annual Survey. However, as the use of containers has grown, so too has the need for more effective management of these highly distributed environments at scale.

To manage this complexity, teams have turned to container orchestration solutions such as Kubernetes. The platform aims to help DevOps teams optimize the allocation of compute resources across all containerized workloads in deployment.

The components of a Kubernetes cluster

While a handful of these solutions exist, Kubernetes has become the de facto industry standard, offering the following benefits:

  • Container deployment. Container orchestration platforms automate daily operations for processes such as the re-creation of failed containers and rolling deployments. This helps to avoid downtime for end users.
  • Automated scaling. Kubernetes enables efficient resource utilization by easily scaling applications and services based on demand.
  • Self-healing. The platform will automatically restart, replace, or kill failed containers, as well as reschedule unhealthy pods and manage node failures. This key feature helps in maintaining availability and reduces the need for manual intervention.
  • Extensibility and technology ecosystem. As an open source solution, Kubernetes has a large, active, and growing ecosystem of extensions, plug-ins, and technologies to enhance its capabilities. The ability to extend functionality across a wide range of use cases allows teams to tailor the platform to meet their organization’s specific requirements.

The top Kubernetes challenges and potential solutions

Despite its benefits, Kubernetes has some potential pitfalls that engineering leaders should consider when managing the complexity it introduces. The top seven Kubernetes challenges include the following:

1. Complexity. Kubernetes environments tend to be complex, multilayered, and dynamic. This creates limitations and blind spots when it comes to observability. Teams often need to know where to look to find issues and resolve them, but that can be time-consuming in large-scale deployments. Comprehensive platform solutions that provide full-stack monitoring with AI at their core can automatically detect anomalies and provide root-cause analysis to prevent future issues.

2. Networking. Large-scale, multicloud deployments can introduce challenges related to network visibility and interoperability. Traditional ways of operating networks using static IPs and ports simply don’t work in dynamic Kubernetes environments. Container Network Interface (CNI) provides a common way to seamlessly integrate various technologies with the underlying Kubernetes infrastructure. Additionally, service meshes — such as those offered by Istio, Linkerd, and Consul Connect — help to manage internetwork communication at the platform layer using purpose-built application programming interfaces.

3. Observability. While there are many observability and monitoring tools on the market today, most are specific in nature. A full-stack observability platform such as Dynatrace provides easy access to the three key types of monitoring signals — logs, traces, and metrics — in context. Automated data collection correlated with topology information, real-user experience, security events, and metadata makes it easier to identify issues and determine effective remediation paths. With the ability to monitor resource utilization metrics such as CPU and memory in real time, teams can optimize their operations, resulting in reduced cost and greater overall efficiency.

4. Cluster stability. Kubernetes containers are naturally short-lived and ephemeral, meaning they’re constantly being created, altered, and removed. This creates challenges in monitoring and debugging distributed applications at scale, often resulting in reliability issues. Effective monitoring, logging, and tracing mechanisms need to be in place to identify and resolve issues quickly. Additionally, ensuring cluster stability through monitoring critical components at the control plane is essential to preventing failures. To minimize negative effects, consider setting limits and alerts on CPU and memory resource requests.

5. Security. Kubernetes security incidents are primarily related to pod communications or misconfigurations that ultimately lead to delayed application deployment. Pods are not isolated, which leaves them vulnerable to malicious actors. Bad actors can use misconfigurations to gain access to sensitive data. However, Kubernetes configurations are complicated, which makes managing them at scale almost impossible. Network policies can restrict pod communications, and teams can use pod security policies to ensure pods are securely configured.

6. Logging. Logs provide critical visibility into the ongoing health of Kubernetes clusters. While Kubernetes makes it easier to generate logs from the various components and layers of a cluster, challenges remain in aggregation and analysis. Solutions that offer log management and analysis capabilities can help to streamline these efforts and provide actionable insights to teams maintaining the health of the workloads and underlying infrastructure.

7. Storage. Containers need to spin up and down easily. Therefore, they are built to be non-persistent by design. However, applications require persistent data to run successfully in production. Traditional storage solutions were not created to address these requirements, which are common among modern deployments. To ease the pain of managing storage at scale within a Kubernetes environment, Kubernetes has released features such as Container Storage Interface (CSI), StatefulSets, Persistent Volume (PV), and Persistent Volume Claim (PVC).

How a cloud-native observability and security platform can overcome Kubernetes challenges

To adequately address the complexity introduced by Kubernetes and other cloud-native technologies, organizations need more complete solutions that provide end-to-end, full-stack visibility, advanced performance and security analytics, and automated workflow capabilities all in a single, comprehensive platform.

Dynatrace integrates extensive Kubernetes observability with continuous runtime application security. This combination helps organizations more effectively meet business goals and minimize risk. The Dynatrace platform offers the following benefits:

  • Automated observability at scale. As organizations deploy Kubernetes clusters across on-premises, cloud, and edge environments, end-to-end observability becomes mandatory. Dynatrace OneAgent and open ingest deliver the deepest and broadest observability on the market, with hundreds of out-of-the-box integrations covering the complete Kubernetes ecosystem. Dynatrace Grail unifies observability, security, and business data at a limitless scale for any analysis at any time.
  • AI-powered analytics. Dynatrace provides precise and explainable answers in real time, identifying performance issues and anomalies across the full Kubernetes technology stack. Davis, the Dynatrace AI engine, helps organizations understand and optimize Kubernetes platform health and application performance, enabling IT teams to proactively pinpoint and rectify performance issues.
  • Platform and application security. Optimized for Kubernetes, Dynatrace Application Security automatically and continuously detects vulnerabilities and protects against injection attacks that exploit critical vulnerabilities, such as Log4Shell. These capabilities remove blind spots, ensuring development teams aren’t wasting time chasing false positives and providing business leaders with confidence in the security of their organizations’ applications.
  • Acceleration of innovation. Dynatrace enables IT teams to shift their effort from code maintenance and troubleshooting to innovation. Leveraging Dynatrace’s analytics and workflow automation enhances software quality, minimizes the mean time to resolve anomalies, and detects security vulnerabilities in near-real time.

For more information on how to solve common Kubernetes challenges, watch our performance clinic, “Kubernetes observability for SREs with Dynatrace

The post The top 7 Kubernetes challenges and how to solve them appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-top-seven-kubernetes-challenges/feed/ 0
A toggle to rule them all: Out-of-the-box alerting for Kubernetes https://www.dynatrace.com/news/blog/a-toggle-to-rule-them-all-out-of-the-box-alerting-for-kubernetes/ https://www.dynatrace.com/news/blog/a-toggle-to-rule-them-all-out-of-the-box-alerting-for-kubernetes/#respond Thu, 09 Feb 2023 18:07:38 +0000 https://www.dynatrace.com/news/?p=56066 Kubernetes alerting

The absolute foundation for managing a Kubernetes cluster is establishing basic alerting. With the introduction of out-of-the-box alerting for Kubernetes, Dynatrace offers a scalable and context-based solution for large Kubernetes environments that enables easy management for multiple teams with different needs.

The post A toggle to rule them all: Out-of-the-box alerting for Kubernetes appeared first on Dynatrace news.

]]>
Kubernetes alerting

In recent years, we have seen a drastic increase in the adoption of Kubernetes among our customer base. More and more companies are moving to containers and using Kubernetes in its various forms as their container platform of choice. While the list of benefits of using Kubernetes for managing containers is numerous, the list of details one must learn to use Kubernetes is even longer. The comparison to an iceberg is appropriate in this case. At the beginning of your Kubernetes journey, the task of managing your Kubernetes environment seems clear—then you begin discovering how many technical gotchas are “hidden beneath the surface.”

With a Kubernetes cluster often hosting multiple critical applications, at a minimum, you need basic alerting to ensure the uptime of your apps. For example, you need to be notified when nodes become unstable, or when a cluster is close to reaching its limit of CPU or memory requests. Otherwise, some pods might remain stuck in a pending state.

Prometheus alerting is great for small environments

Usually, teams turn to open source tools like Prometheus and Alertmanager to set up alerting. For small environments, this is a perfect solution. Alertmanager is free, and setup is easy. However, from our discussions with customers, we know that small environments can quickly grow into large environments of multiple Kubernetes clusters, each hosting multiple business-critical applications. This is often a consequence of companies eventually realizing the countless benefits of Kubernetes and then quickly going all in on this amazing technology.

While open source tools are great for small environments, they don’t scale easily, for example:

  • Training: Teaching multiple teams how to configure alerting using Prometheus, Alertmanager, and especially, PromQL, can be a huge project.
  • Access: Most of the time, each K8s cluster receives its own Prometheus instance so data is spread over multiple instances, with different access credentials, web portals, and, unfortunately, no context or correlation across clusters.
  • Data security: At the same time, restricting access to specific data within one Prometheus instance is not possible. So, you need to choose between allowing everyone to see all data or no one seeing any data.
  • Configuration: Usually, you want proper alerting for essential scenarios for each Kubernetes environment, which can be difficult to set up and keep in sync across multiple instances.

In summary, the Prometheus stack is amazing for Kubernetes alerting as long as you don’t need it for multiple teams.

Easily scale Kubernetes alerts for multiple teams

With Dynatrace, you can easily overcome the shortcomings of implementing a company-wide Kubernetes alerting solution based on Prometheus. To get started, open the global anomaly-detection settings for Kubernetes in the Dynatrace web UI and configure appropriate defaults for all Kubernetes clusters that are currently, or will in the future, be connected to Dynatrace.

You only need to flip a few settings toggles to activate critical Kubernetes alerts —there’s no need to research which essential metrics you need to alert on, learn a complex query language, figure out which metrics to use, or learn how to capture metrics. In other words, the training aspect of setting up alerts is reduced to near zero. When you use Dynatrace as a SaaS solution, you don’t need to take care of hosting—all your observability data is accessible in one place, while built-in access management allows you to easily define who is allowed to access which data.

Of course, each of your teams will want to adapt the default settings to their scope. This can be done easily by overwriting defaults on various scope levels. For example, if you have a Kubernetes development cluster, you probably don’t need the same alerts for nodes that you use in your production cluster. With Dynatrace, you can overwrite the global defaults for node alerts in the scope of every cluster. The same is true for common alerts on workloads, like alerting on frequent container restarts. You can easily define appropriate defaults for all workloads and overwrite them in the scope of a Kubernetes cluster or Kubernetes namespace. With this approach, you can provide each team with access to the settings related to their namespace. From their namespace page in Dynatrace, each team can directly navigate to the corresponding settings and adapt everything to their needs. There’s no need to learn a new query language or open a ticket to have someone else set and configure alerts.

Kubernetes namespace settings in Dynatrace

Now each team can set and customize whichever alerts they want to use, including detection sensitivity. Consequently, internal data protection across teams is ensured and your team doesn’t need to take care of all other teams’ adaptations.

When many teams are involved, you might think that all the overwrites will be difficult to track and understand. No worries—Dynatrace has you covered. In the global settings of the Dynatrace web UI, you can explore all the various overwrites that exist in your Dynatrace environment.

Kubernetes namespace anomaly detection settings in Dynatrace

It’s also easy for teams to understand if they are using custom alert configurations or the company-wide defaults.

Kubernetes namespace defined settings in Dynatrace

Of course, all alert configurations can also be automated using the Dynatrace API. Going forward, you will even be able to use our monitoring-as-code approach to configure alerts in a GitOps fashion.

How can I get this feature?

With Dynatrace 1.254, we’ve released a first set of alerts for common Kubernetes issues. Over the coming months, we plan to incrementally expand this set. You can find more information and provide your feedback in the Dynatrace Community.

For complete details on this feature, see Dynatrace Documentation.

Kubernetes in the wild report 2023

This Kubernetes survey shows how organizations actually use Kubernetes in production. The study analyzes factual Kubernetes production data from thousands of organizations worldwide that are using the Dynatrace Software Intelligence Platform to keep their Kubernetes clusters secure, healthy, and high performing.

Pie charts showing Kubernetes adoption of cloud-hosted clusters vs. on-premises clusters

The post A toggle to rule them all: Out-of-the-box alerting for Kubernetes appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/a-toggle-to-rule-them-all-out-of-the-box-alerting-for-kubernetes/feed/ 0
Dynatrace memory analysis helps Product Architects identify unknown unknowns https://www.dynatrace.com/news/blog/dynatrace-memory-analysis-helps-product-architects-identify-unknown-unknowns/ https://www.dynatrace.com/news/blog/dynatrace-memory-analysis-helps-product-architects-identify-unknown-unknowns/#respond Thu, 09 Feb 2023 17:34:58 +0000 https://www.dynatrace.com/news/?p=56060 Dashboard graphic

Excessive memory allocations or leaks can harm your organization’s clusters and lead to crashes or unresponsive services. To avoid this, it’s essential to monitor your KPIs for memory allocation and object churn as measures of the performance and health of a system.

The post Dynatrace memory analysis helps Product Architects identify unknown unknowns appeared first on Dynatrace news.

]]>
Dashboard graphic

Luckily, Dynatrace provides in-depth memory allocation monitoring, which allows fine-grained allocation analysis and can even point to the root cause of a problem.

While memory allocation analysis can show wasteful or inefficient code, it can also reveal different problems, one of which we’ll examine in this blog post. This real-world use case, caused by an issue in a customer environment, illustrates how Dynatrace memory analysis capabilities can contribute to root cause analysis within a Dynatrace Cluster.

The typical ratio is about 1.5X higher, but now it’s 3X higher—why?

At Dynatrace, we use dashboards to get a quick overview of the status of monitored services. One such dashboard is the Allocations dashboard which gives an overview of memory usage and allocations for an entire production environment, grouped by APIs.

We recently extended the pre-shipped code-level API definitions to group logical parts of our code so they’re consistently highlighted in all code-level views. For instance, everything related to our correlations engine is dark orange, and the different protocols are mustard colored. Another benefit of defining custom APIs is that the memory allocation and surviving object metrics are split by each custom API definition. So we can easily keep track of them on the Allocations dashboard.

One day while looking at a single cluster, we saw that the memory allocations were abnormally high. While the amount of bytes allocated for the Java API is typically 1.5X the average, in this case, the allocation for the Java API was more than 3X higher than the average, 41 TiB. What could be causing this?

Allocation Bytes dashboard in Dynatrace screenshot

We looked at one of the Dynatrace instances to investigate what was going on. Garbage collection suspension and CPU usage looked healthy. We know from experience that an average value of ~1% GC suspension is healthy, so it was still unclear what was causing the high number of allocations shown on the dashboard.

In Memory profiling view, we would normally expect to see allocations for protocols and database calls at the top of the list of allocation hotspots. In this case, all the top contributors are located in the cluster platform code (as shown by the package names).

Memory profiling All allocations in Dynatrace screenshot

Looking at the call stack of the top allocation, a familiar message handler can be identified, AgentClusterRuntimeInfoMsgHandler. This handler is responsible for sending configuration updates regarding usable communication endpoints (in other words, available ActiveGates) to connected OneAgents. Typically, the configuration does not change, and no responses are created for the OneAgents. In this case, the server appears to be continuously building responses, which is an expensive operation that indicates either we have a bug in the revision calculation of our message handler, or the list of ActiveGates is constantly changing, forcing frequent revision recalculation.

Selecting Called Methods next to the message handler opens the profiling view, which shows the full extent of the impact. The handler is responsible for ~3.5 TiB in allocations within 2 hours, allocating and removing about 75 billion objects during the process.

Profiling view of called methods in Dynatrace screenshot

Verification with Dynatrace custom metrics

As Dynatrace also exposes key metrics about our message handler via JMX, we can use those metrics to investigate further. In Further Details on the Host page, we instantly have the confirmation we’re looking for: We were constantly sending ~4.5MiB/s of ClusterRuntimeInfo responses, while on a healthy system the response size is typically 50KiB/s or less (depending on the number of connected agents).

Since other production systems are doing fine at the same time, a bug in the code might not be the problem. Instead, we investigate to see if we have many recalculations due to constantly changing ActiveGate connections.

Luckily, we have an audit log for ActiveGate connectivity on the Dynatrace Cluster, which can be seen in the log viewer.

Audit log for ActiveGate connectivity on the Dynatrace Cluster

Finding the root cause of the problem

In the audit log file, we can see that many ActiveGate registration and deregistration activities are taking place. By adding a filter for a single ActiveGate ID and increasing the timeframe, a pattern emerges: this ActiveGate is reconnecting once per hour.

ActiveGate registration and deregistration activities in audit log file

The other ActiveGates do the same at separate times, which explains the server behavior: every time an ActiveGate connects or disconnects, the endpoint list changes and so must be resent to the deployed OneAgents. The customer has more than 100 thousand OneAgents connected, which consumes many resources on the server and, more importantly, on the network. Following these insights, we contacted this customer to share our findings.

Conclusion

Memory allocation analysis can show wasteful or inefficient code, but it can also reveal unexpected problems, such as, in this case, numerous configuration updates sent out due to a problem on the customer side. Even though the server could easily handle the memory allocations (GC suspension was around 1%), the allocations showed up prominently, and they can be seen as an indicator of bugs in the system.

You can find out more about Dynatrace memory allocation analysis in our documentation:

New to Dynatrace?

Visit our trial page for a free 15-day Dynatrace trial.

The post Dynatrace memory analysis helps Product Architects identify unknown unknowns appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-memory-analysis-helps-product-architects-identify-unknown-unknowns/feed/ 0
Break down the barriers to end-to-end monitoring with Dynatrace https://www.dynatrace.com/news/blog/break-down-the-barriers-to-end-to-end-monitoring-with-dynatrace/ https://www.dynatrace.com/news/blog/break-down-the-barriers-to-end-to-end-monitoring-with-dynatrace/#respond Fri, 27 Jan 2023 17:18:28 +0000 https://www.dynatrace.com/news/?p=55885 End-to-end monitoring

Dynatrace is excited to release cross-environment tracing, which enables enterprises to follow individual distributed traces from the boundary of one monitoring environment into other environments while maintaining a complete end-to-end view.

The post Break down the barriers to end-to-end monitoring with Dynatrace appeared first on Dynatrace news.

]]>
End-to-end monitoring

In recent years, more and more large enterprises have embraced microservices-based architectures that run across clouds and geographies to deliver increased agility, improved performance, scale, and reliability. This approach has accelerated innovation and created new opportunities for distributed development and support models.

As companies have modernized, they have embraced new IT business models and faced tightening regulatory requirements. Specifically, the heightened awareness of data residency and sovereignty in many regions of the world has resulted in the requirement to maintain separate data centers with local data footprints across the globe. Furthermore, global enterprises, which typically have multiple business units and divisions, often require separating applications and services.

Easy end-to-end distributed tracing across multiple environments

Dynatrace is proudly committed to providing users with an integrated observability platform that provides true end-to-end monitoring and analysis. Simplifying complexity and delivering not just more data but answers has been our credo from the beginning.

Tracking a transaction from start to finish is critical for end-to-end visibility and is a relatively simple endeavor when applications are contained in a unified observability system. However, when applications span monitoring environments, end-to-end tracing becomes much more difficult because each environment only sees a piece of the complete transaction.

With millions of requests an hour processed, and some requests going into other environments, tracing a single transaction can be like finding a needle in a haystack.

Easy end-to-end distributed tracing through multiple environments

Seeing is believing, so let’s look at an example: a company that runs a travel platform is migrating its services to the cloud. They’re pursuing a hybrid cloud strategy where the front end runs on a hyperscaler cloud provider. The middleware and back end will continue to be managed internally to keep customer data in the local region. Multiple Dynatrace environments are deployed to ensure data residency.

Just before the start of the travel season, special attention is paid to the performance indicators, and an increased response time is noticed. The following distributed trace shows an end-to-end view through three Dynatrace environments from the front end (cloud hosted), to middleware, and the back end (hosted on-premises ). We can see an issue at the back end, but we can go further. A direct link to the Dynatrace back-end environment lets us jump right to the point where we can identify the problem.

Distributed trace cross-environment 1

Setup is fast and uncomplicated

Connecting multiple Dynatrace environments takes less than a minute. Just enter a unique environment name, the tenant URL of the Dynatrace environment, and a secure token. Be sure to enable the cross-environment feature in your OneAgent settings (Settings > Integration > Remote environments). That’s all there is to it.

Distributed trace cross-environment 2

How can I get this feature?

This new capability was released in Dynatrace SaaS version 1.248, Dynatrace Managed version 1.250, and OneAgent version 1.247. Please see Dynatrace Documentation to learn how to enable this capability and connect your environments.

The post Break down the barriers to end-to-end monitoring with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/break-down-the-barriers-to-end-to-end-monitoring-with-dynatrace/feed/ 0
Kubernetes in the Wild report 2023 https://www.dynatrace.com/news/blog/kubernetes-in-the-wild-2023/ https://www.dynatrace.com/news/blog/kubernetes-in-the-wild-2023/#respond Mon, 16 Jan 2023 13:00:02 +0000 https://www.dynatrace.com/news/?p=55562 Ships wheel logo for Kubernetes survey and Kubernetes adoption

Rapid Kubernetes adoption is driven—and challenged by—a growing ecosystem of advanced technologies. In this Kubernetes survey report, learn how top organizations use Kubernetes and related technologies in production, including observability, security, infrastructure models, and open source software.

The post Kubernetes in the Wild report 2023 appeared first on Dynatrace news.

]]>
Ships wheel logo for Kubernetes survey and Kubernetes adoption

Update: The new Kubernetes in the Wild report 2025 is now available.

Kubernetes adoption survey executive summary

Modern cloud-native computing is impossible to separate from containers and Kubernetes adoption. Although Kubernetes is still a relatively young technology, a majority of global enterprises use it to run business-critical applications in production. Its rapid adoption is driven—and challenged by—an ever-growing ecosystem of Kubernetes technologies that add advanced platform features, such as security, microservices communication, observability, scaling, resource utilization, and so on.

The Kubernetes in the Wild survey reveals how organizations actually use Kubernetes in production. The study analyzes factual Kubernetes production data from thousands of organizations worldwide that are using the Dynatrace unified observability and security platform to keep their Kubernetes clusters secure, healthy, and high-performing.

Findings provide insights into Kubernetes practitioners’ infrastructure preferences and how they use advanced Kubernetes platform technologies. The report also reveals the leading programming languages practitioners use for application workloads. As Kubernetes adoption increases and it continues to advance technologically, Kubernetes has emerged as the “operating system” of the cloud.

  1. Kubernetes moved to the cloud in 2022
  2. Kubernetes infrastructure models differ between cloud and on-premises
  3. Kubernetes is emerging as the “operating system” of the cloud
  4. The strongest Kubernetes growth areas are security, databases, and CI/CD technologies
  5. Open-source software drives a vibrant Kubernetes ecosystem
  6. Java, Go, and Node.js are the top 3 programming languages for Kubernetes application workloads

Insight 1

Kubernetes moved to the cloud in 2022

In 2022, Kubernetes became the key platform for moving workloads to the public cloud. At an annual growth rate of +127 percent, the number of Kubernetes clusters hosted in the cloud grew about five times as fast as clusters hosted on-premises. Likewise, the share of cloud-hosted clusters increased from 31% in 2021 to 45% in 2022. Cloud-hosted Kubernetes clusters are on par to overtake on-premises deployments in 2023.

Most Kubernetes clusters in the cloud (73%) are built on top of managed distributions from the hyperscalers like AWS Elastic Kubernetes Service (EKS), Azure Kubernetes Service (AKS), or Google Kubernetes Engine (GKE). Accordingly, the remaining 27% of clusters are self-managed by the customer on cloud virtual machines.

Kubernetes hosting decisions are guided by a set of parameters, including cost, ease of provisioning and scaling, data security, and regulatory compliance. As hyperscalers invest in all these areas and expand their presence into more geographic regions, they become more attractive to a broader set of organizations.

Pie charts showing cloud-hosted clusters vs. on-premises clusters
More clusters moved to the cloud from on-premises in 2022, and are thus on par to overtake on-premises deployments in 2023.

Insight 2

Kubernetes infrastructure models differ between cloud and on-premises

A typical cluster running in the public cloud consists of 5 relatively small nodes with just 16 to 32 GB of memory each. In comparison, on-premises clusters have more and larger nodes: on average, 9 nodes with 32 to 64 GB of memory.

The different infrastructure setup reflects economic and technical considerations. Hyperscalers offer a competitive price point for small to medium-sized hosts. Through effortless provisioning, a larger number of small hosts provide a cost-effective and scalable platform. On-premises data centers invest in higher capacity servers since they provide more flexibility in the long run, while the procurement price of hardware is only one of many cost factors.

Kubernetes survey results bar chart showing node memory sizes for nodes hosted on the cloud and on premises.

Kubernetes survey bar chart showing nodes and pods per cluster
Typical cloud-hosted clusters run on 5 relatively small nodes. Conversely, clusters hosted on-premises use 9 nodes with almost double the memory.

Insight 3

Kubernetes is emerging as the “operating system” of the cloud

As the ideal orchestration platform for running cloud-native microservice applications, Kubernetes comes with the benefit of built-in deployment, scaling, and resiliency capabilities. In 2021, in a typical Kubernetes cluster, application workloads accounted for most of the pods (59%). By contrast, all non-application workloads, such as system and auxiliary workloads, played a relatively smaller part.

But in 2022, this picture reverses. As Kubernetes adoption has grown, auxiliary workloads now outnumber application workloads (63% vs. 37%). This switch reflects that organizations are implementing more advanced Kubernetes platform technologies such as security controls, service meshes, messaging systems, and observability tools. At the same time, organizations are using Kubernetes for a broader range of use cases, including build pipelines and scheduled utility workloads, among others. Kubernetes becomes the platform for running almost anything. As such, Kubernetes is emerging as the “operating system” of the cloud.

Pie chart that shows application workloads vs. auxiliary workloads
In 2021, application workloads dominated, whereas in 2022, auxiliary workloads were predominant, showing a broader range of use cases.

“At Dynatrace, we use Kubernetes for any new software project, from build pipelines to SaaS offerings. We also see the same trend with our customers. Kubernetes effectively has emerged as the operating system for the cloud.”

Anita Schreiner, Dynatrace VP Delivery

Insight 4

Strongest Kubernetes growth areas are security, databases, and CI/CD technologies

In 2022, organizations identified Kubernetes security as a top priority. Starting from a low baseline, the percentage of organizations using Kubernetes security tools increased from 22% in 2021 to 34% in 2022. This corresponds to an annual growth rate of +55%. That trend will likely continue as Kubernetes security awareness further rises and a new class of security solutions becomes available.

Of the organizations in the Kubernetes survey, 71% run databases and caches in Kubernetes, representing a +48% year-over-year increase. Together with messaging systems (+36% growth), organizations are increasingly using databases and caches to persist application workload states.

Continuous integration and delivery (CI/CD) technologies grew by +43% year-over-year. This trend shows that organizations are dedicating significantly more Kubernetes clusters to running software build, test, and deployment pipelines.

Bar chart that shoes top Kubernetes adoption technologies

“The immense growth of Kubernetes presents new security challenges in runtime and increased complexity in hardening CI/CD pipelines in development. On the upside, new application security approaches address these challenges, reducing exposure to attacks and mitigating risks.”

Andreas Berger, Dynatrace Senior Principal Application Security

Insight 5

Open source software drives a vibrant Kubernetes ecosystem

Focusing on non-application workloads, organizations use an increasing variety of technologies. These results reflect the need to enhance Kubernetes with better observability, security, and service-to-service communications. Similarly, other technologies enable specific use cases like CI/CD tools or databases. Across all categories in the Kubernetes survey, open source projects rank among the most frequently used solutions.

Kubernetes survey Bar chart showing technologies used in Kubernetes environments

  • Open source observability: Prometheus is the clear leader in open source observability and is used by 65% of organizations. In general, metrics collectors and providers are most common, followed by log and tracing projects. Note: The survey excluded all commercial observability offerings, including Dynatrace.
  • Databases: Among databases, Redis is the most used at 60%. Redis is an in-memory key-value store and cache that simplifies processing, storage, and interaction with data in Kubernetes environments. Accordingly, for classic database use cases, organizations use a variety of relational databases and document stores.
  • Messaging: RabbitMQ and Kafka are the two main messaging and event streaming systems used. Specifically, they provide asynchronous communications within microservices architectures and high-throughput distributed systems.
  • Continuous integration and delivery: ArgoCD, Flux, GitLab, and Jenkins are the most widely adopted CI/CD tools. Organizations increasingly use the flexibility and elasticity of Kubernetes to run CI and CD jobs as well as their control planes.
  • Big data: To store, search, and analyze large datasets, 32% of organizations use Elasticsearch.
  • Security: For security, organizations mostly use policy checkers and enforcers, such as Gatekeeper. The need for runtime security observability is growing to automate vulnerability impact analysis.
  • Service meshes: Istio is the most used service mesh. Organizations are increasingly using service meshes in large Kubernetes clusters to automate secure service-to-service communication and expose telemetry data for better observability.

“Dynatrace believes in a strong open-source ecosystem and embraces the adoption of cloud-native technologies and practices. That’s why we actively contribute to and bootstrap projects and participate in the open source community in various roles. Dynatrace’s investment in open source technologies keeps growing.”

Alois Reitbauer, Dynatrace Chief Technology Strategist

Insight 6

Java, Go, and Node.js are the top Kubernetes programming languages

The Dynatrace OneAgent automatically detects the specific programming languages of every individual application workload running on Kubernetes. This provided unique insights into the Kubernetes programming languages organizations use.

Java Virtual Machine (JVM)-based languages are predominant. Accordingly, 65% of all application workloads run in a JVM, including related application servers like Tomcat or Spring. Most organizations, 72%, use Java to some degree.

Go ranks number 2 with a 58% adoption rate among organizations, with 14% of application workloads written in Go. Not counted are Kubernetes system workloads, sidecars, or any standard components of non-application workloads. In addition, Node.js ranks third in terms of workload count and organizational adoption.

Bar chart showing top programming languages used in Kubernetes adoption

“With Kubernetes, polyglot programming finally becomes a reality. As a result, Kubernetes empowers existing teams and makes onboarding of new ones easy, regardless of programming language and framework usage.”

Florian Ortner, Dynatrace Chief Product Officer

Kubernetes survey methodology

This report reflects Kubernetes adoption statistics based on the analysis of 4.1 billion Kubernetes pods from thousands of Dynatrace customers in all global regions. The data covers the period of January 2021 through September 2022. These customers are among the world’s largest 15,000 organizations from all major industries, including financial services, retail and e-commerce, technology, transportation, manufacturing, healthcare, and public-sector organizations.

The report only includes production data from Dynatrace customers and excludes all Kubernetes clusters Dynatrace uses internally or for hosting SaaS offerings.

The Kubernetes in the Wild report 2023 is also available as a printable PDF.

Update: The new Kubernetes in the Wild report 2025 is now available.

The post Kubernetes in the Wild report 2023 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/kubernetes-in-the-wild-2023/feed/ 0
Three smart log ingestion strategies in Dynatrace https://www.dynatrace.com/news/blog/three-smart-log-ingestion-strategies/ https://www.dynatrace.com/news/blog/three-smart-log-ingestion-strategies/#respond Thu, 15 Dec 2022 20:58:53 +0000 https://www.dynatrace.com/news/?p=55224 AppEngine: Create custom apps for data insights

Getting precise answers from log monitoring platforms gets challenging as cloud environments expand and grow more complex. Here are three log ingestion strategies to achieve scale in the Dynatrace platform—without OneAgent.

The post Three smart log ingestion strategies in Dynatrace appeared first on Dynatrace news.

]]>
AppEngine: Create custom apps for data insights

While many organizations have embraced cloud observability to better manage their cloud environments, they may still struggle with the volume of entities that observability platforms monitor. The key to getting answers from log monitoring at scale begins with relevant log ingestion at scale.

Engaging the automatic instrumentation of the Dynatrace OneAgent makes log ingestion automatic and scalable. However, our customers often have set up multiple other log ingestion methods. This flexibility enables logs from diverse environments and established configurations to complete the observability picture for automated troubleshooting and monitoring in Dynatrace.

In this blog, we share three log ingestion strategies from the field that demonstrate how building up efficient log collection can be environment-agnostic by using our generic log ingestion application programming interface (API).

As with all other log ingestion configurations, these examples work seamlessly with the new Log Management and Analytics powered by Grail that provides answers with any analysis at any time.

Log ingestion strategy no. 1: Welcome syslog, with the help of Fluentd

Syslog is a popular standard for transporting and ingesting log messages. Typically, these are streamed to a central syslog server. One option is to install OneAgent on that syslog server, which automatically discovers, instruments and sends the log data to the Dynatrace platform.

But there are cases where you might be limited in setting up a dedicated syslog server with OneAgent because of environment architecture or resources. Yet observability into syslog data on Dynatrace would help you monitor and troubleshoot infrastructure.

This is where it is prudent to configure syslog producers to send data to a log shipper like Fluentd.

What is Fluentd?

Fluentd logo for log ingestion and log monitoring

Fluentd is an open source data collector that decouples data sources from observability tools and platforms by providing a unified logging layer. Fluentd is known for its flexibility and is also highly scalable, which makes it a good choice for high-volume environments.

How does Fluentd work with Dynatrace?

Setting up the flow from syslog over Fluentd to Dynatrace takes three steps. First, point the syslog daemon to the Fluentd port by adding the following line to the syslog daemon configuration file:

*.* @@<fluentd host IP>:5140

*.* instructs the daemon to forward all messages to the specified Fluentd instance listening on port 5140 and <fluentd host IP> needs to point to the IP address of Fluentd.

As a second step, enable Fluentd to accept incoming syslog messages with the in_syslog plugin. Set up the configuration on the same port as specified for source data, in this example 5140.

Lastly, use the open source Dynatrace Fluentd plugin, which uses generic log ingestion. Just find the API token for log ingest API on your SaaS environment or your own Active Gate setup.

Now you should see log messages coming into the Dynatrace log viewer.

Log ingestion strategy No. 2: Point an existing log shipper to the generic Dynatrace ingest

Another common scenario is an environment where you have already invested a do-it-yourself or other log shipper solution. After spending time and budget on the tooling and configuration, it may be unwise to undo this custom work, despite the automatic instrumentation of the Dynatrace OneAgent. Although you preserve your custom work this way, it is a siloed approach for logs, which means you’ll miss out on the integrated observability and automated alerting of Dynatrace.

If that existing solution supports sending log data to an external HTTP endpoint, you can address log silos by integrating with Dynatrace generic ingest with minimal hassle.

To illustrate the solution, let’s look at how to configure log ingestion with the log shipper Cribl.

What is Cribl?

Cribl Stream logo for log ingestion and log monitoring

Cribl is a data operations platform that enables users to collect, route, transform, analyze, and act on data in real time. It provides a unified platform for handling every aspect of data operations, from collecting data to routing and transforming it. Cribl also allows users to orchestrate custom pipelines for their data to gain insights and take action on that information. As a data output, or what it calls a Cribl Stream destination, you can configure an HTTP endpoint.

How does Cribl work with Dynatrace?

The main part of the setup involves creating the configuration for the specific log shipper at hand—in this case, Cribl Stream.

In Cribl’s configuration, open “Data/Destinations” and find “Webhook.” Create a new webhook destination with a name of your choosing (for example, your Dynatrace environment ID, and provide the URL for the webhook). For a Dynatrace SaaS environment, this is the following:

https://{your-environment-id}.live.dynatrace.com/api/v2/logs/ingest

This points the data stream to your Dynatrace environment’s generic ingest.

But in Cribl’s case, you should provide two more settings under “Configure/Advanced Settings/Extra HTTP Headers.” Add two new headers with the following names and values:

  1. To authorize the request, add the header “Authorization” and provide the value Api-Token dt0c01.{your-token-here} where {your-token-here} is an API token with ingest logs scope.
  2. Then add a header “Content-Type” and provide the value “application/json; charset=utf-8
log ingestion, log management screenshot
Example configuration in Cribl of posting logs to Dynatrace API.

After committing and deploying the Cribl changes, you can select the newly created Dynatrace destination as the default destination for your logs. And just like that, all log data already collected by the existing shipper is being sent to Dynatrace for monitoring, analysis, alerting, and all other tasks.

Log ingestion strategy No. 3: Ingest AWS Fargate logs with Fluent Bit

Ingesting and working with Kubernetes logs in Dynatrace helps to provide a comprehensive view of application performance from the infrastructure layer to the application layer. The common approach for Kubernetes logging is to deploy OneAgent in the environment, where it auto-discovers log messages written to the containerized application’s stdout/stderr streams.

But not all environments, configurations, or privileges are created equal. One recurrent challenge is collecting Kubernetes logs if you’re limited in installing OneAgent because of technical or architectural restrictions.

In the case of AWS serverless container compute engine Fargate, for example, where OneAgent log collection is not supported, we recommend using Fluent Bit log forwarder.

Let’s take this example of AWS Fargate. AWS includes a log router called FireLens for Amazon ECS and AWS Fargate services, which gives you built-in access to FluentD and Fluent Bit. We covered FluentD support previously. Now let’s take a look at how to set up Fluent Bit.

What is Fluent Bit?

Fluent Bit logo - for log ingestion and log management

Fluent Bit is an open source and multiplatform log processor and forwarder that allows you to collect data/logs from different sources, unify and send them to multiple destinations and is fully compatible with Docker and Kubernetes environments.

When choosing between Fluentd or Fluent Bit shippers, the Fluent Bit is the preferred solution when resource consumption is critical because it is a lightweight component.

While Fluent Bit has configurable HTTP output, in this example, we look at the AWS Fargate context, where FireLens makes it easy to set up Fluent Bit more quickly.

Ingest AWS Fargate logs with Fluent Bit

When creating a new task definition using the AWS Management Console, the FireLens integration section makes it easy to add a log router container. Just pick the built-in Fluent Bit image.

Next, edit the container in which your app-generating logs are running. In the “Storage and Logging” section, select “awsfirelens” as the log driver.

The settings for the log driver should point to the log ingest API of your SaaS tenant. Note that you normally need to provide two headers for Fluent Bit: content type and authorization token. As FireLens supports only one header, you can pass the token as part of the URL. Your configuration for AWS FireLens should have the following:

  • Name: http
  • TLS: on
  • Format: json
  • Header: Content-Type application/json; charset=utf-8
  • Host: {your-environment-id}.live.dynatrace.com
  • Port: 443
  • URI: /api/v2/logs/ingest?api-token={your-API-token-here}
  • tls.verify: Off
  • Allow_Duplicated_Headers: false
  • match: *
  • json_date_format: iso8601
  • json_date_key: timestamp

To avoid publishing the token in plaintext, use AWS Secrets Manager to manage the token.

As your application starts publishing logs, you can view them in Dynatrace.

Read more about streaming logs to Dynatrace with Fluent Bit from our documentation.

More methods for log ingestion

These are just some of the ways you can ingest logs into the Dynatrace platform without using OneAgent. You’ll soon have even more methods for log ingestion into Dynatrace, for example:

  • Automated OpenTelemetry logs acquisition and processing
  • Syslog endpoint in your environment as a component on a private ActiveGate
  • Dynatrace Fluent Bit output plugin for out-of-the-box integration

Want to share your experiences with log ingestion? Head to the Dynatrace Community Feedback channel to share your thoughts with other users.

State of Log Management 2026

Download the report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

The post Three smart log ingestion strategies in Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/three-smart-log-ingestion-strategies/feed/ 0
Detecting network errors and their impact on services https://www.dynatrace.com/news/blog/detecting-network-errors-impact-on-services/ https://www.dynatrace.com/news/blog/detecting-network-errors-impact-on-services/#respond Wed, 14 Dec 2022 09:02:55 +0000 https://blog.ruxit.com/?p=12197 What is AIOps?

Detecting errors like dropped packets or retransmissions on the network level is relatively easy. Figuring out if those errors affect the performance and connectivity of your services is however another matter. Some network errors are mitigated and compensated for by network protocols and active networking components, like network interfaces. Meanwhile, other network errors lead to […]

The post Detecting network errors and their impact on services appeared first on Dynatrace news.

]]>
What is AIOps?

Detecting errors like dropped packets or retransmissions on the network level is relatively easy. Figuring out if those errors affect the performance and connectivity of your services is however another matter. Some network errors are mitigated and compensated for by network protocols and active networking components, like network interfaces. Meanwhile, other network errors lead to performance problems that negatively affect your services.

Following is an overview of common network errors and root causes, means and approaches of detecting such errors, and suggestions as to how monitoring tools can support you in staying on top of your services’ connectivity and performance.

TCP – Your protocol of choice since 1981

The TCP/IP protocol suite that we all know so well has been around for almost 40 years now. Although some alternatives have been developed over the years, TCP/IP still works well and it’s the foundation of almost all networking as we know it today. One of the reasons this protocol stack is still around is that it’s capable of compensating for many errors on its own. TCP, appropriate to the season, is the Santa Claus of protocols. It knows if your service is sleeping, it knows if it’s awake, it knows if the connections run bad or good, so [listen closely to what it says]. Your services need not worry about retransmissions or network congestion. TCP/IP does everything in its power to makes sure that your stateful connections are reliable and perform well. Nevertheless anybody running applications in production needs to understand TCP and its basics.

The top five common network errors

Network collisions

This is an oldie, but a goodie that’s now almost irrelevant because of full duplex switches and technology advances. Back in the days if two devices on the same Ethernet network (e.g., connected through a hub) tried to transmit data at the same time, the network would detect the collision and drop both packets. The CSMA/CD protocol, which made sure that nobody else was transmitting data before a device started transmitting its own data, was a step in the right direction. With full duplex switches, where communication end-points can talk to each other at the same time, this potential error is obsolete. Even in wireless networks, which still work basically like hubs, network collisions can be neglected because there are procedures in place to avoid collisions in the first place (e.g., CSMA/CA or RTS/CTS).

Checksum errors

When you download files from the Internet you often have the option of checking a file’s integrity with a MD5 or SHA-1 hash. With the help of checksums on the network level we are able to detect if a bit was toggled, missing, or duplicated by network data transmission. Checksums assure that received data is identical to the transmitted data.

checksums
Checksums are used in the Ethernet, IP, and TCP header

Packets with incorrect checksums aren’t processed by the receiving host. If the Ethernet checksum (CRC) is wrong the Ethernet frame is silently dropped by the network interface and is never seen by the operating system, not even with packet capturing tools. With the IP checksum and TCP checksum in the respective headers there are two additional supervisory bodies that can detect integrity errors. Be aware that despite the efforts of checksumming, there are some errors that can’t be detected.

Full queues

If the processing queue on a switch or router is overloaded, the incoming packets will be dropped. Also if the queue for incoming packets on the host you try to connect to is full, the packets will also be dropped. This behavior is actively exploited during DoS/DDoS attacks. So while it’s actually a good thing that a host only accepts the number of packets than it can process, this behavior can be used to take down your service.

Time to live exceeded

The Time to live (TTL) field in the IPv4 header has a misleading name. Every router that forwards an IP packet decreases the value of the field by one — it actually has nothing to do with time at all. In the IPv6 header this field is called “hop limit”. If the TTL value hits 0, an ICMP message “time to live exceeded” is sent to the dispatcher of the packet. Meanwhile, some network components drop packets with TTL equal to zero silently. This mechanism is useful for preventing packets from becoming caught up in an endless routing loop within your network. The observant reader and network veteran is familiar with this technique because traceroute uses it to identify all hops that a packet makes on its route to its destination.

Packet retransmissions

First off, retransmissions are essential for assuring reliable end-to-end communication in networks. Retransmissions are a sure sign that the self-healing powers of the TCP protocol are working — they are the symptom of a problem, not a problem in themselves. Common reasons for retransmissions include network congestion where packets are dropped (either a TCP segment is lost on its way to the destination, or the associated ACK is lost on the way back to the sender), tight router QoS rules that give preferential treatment to certain protocols, and TCP segments that arrive out of order at their destination, usually because the order of segments became mixed up on the way from sender to destination. The retransmission rate of traffic from and to the internet should not exceed 2%. If the rate is higher, the user experience of your service may be affected.

The three commands you need to know to gather information about network errors

Now we know about common errors – let’s take a look at network troubleshooting. The good news is that most of the problems are findable using standard tools that are usually part of your operating system.

ifconfig

The first place to go to find basic information about your network interfaces is good old ifconfig.

ifconfig output
ifconfig shows details about the specified network interface

Besides the MAC address and the IP address information for v4 and v6 you’ll find detailed statistics about received and transmitted packets. The line that starts with RX contains information about received packets. The TX lines contain information about transmitted packets.

RX information

rx details
Details about received packets
  • packets shows the number of successfully received packets.
  • errors can result from faulty network cables, faulty hardware (e.g., NICs, switch ports), CRC errors, or a speed or duplex mismatch between computer and switch, which would also manifest itself in a high number of collisions (CSMA/CD sends its regards). You can check the configuration on your computer using ethtool <device> to find out at which speed your network interface is operating and if the connection is full duplex or not.
  • dropped can indicate that your system can’t process incoming packets or send outgoing packets fast enough, you’re receiving or sending packets with bad VLAN tags, you’re using unknown protocols, or you’re receiving IPv6 packets and your computer doesn’t support IPv6. You can counter the first error by increasing the ring-buffer. This is the buffer that the NIC transfers frames to before raising an IRQ in the kernel, for RX of your network interface using ethtool.
  • overruns display the number of fifo overruns, which indicates that the kernel can’t keep up with the speed of the ring-buffer being emptied.
  • frame counts the number of received misaligned Ethernet frames.

TX information

tx details
Details about transmitted packets
  • packets shows the number of successfully transmitted packets.
  • errors shows the number of errors that occurred while transmitting packets due to carrier errors (duplex mismatch, faulty cable), fifo errors, heartbeat errors, and window errors.
  • dropped indicates network congestion, e.g., the queue on the switchport your computer is connected to is full and packets are dropped because it can’t transmit data fast enough.
  • overruns indicates that the ring-buffer of the network interface is full and the network interface doesn’t seem to get any kernel time to send out the frames stuck in the ring-buffer. Again, increasing the TX buffer using ethtool may help.
  • carrier shows the number of carrier errors, indicating a duplex mismatch or faulty hardware.
  • collisions shows the number of collisions that occurred while transmitting packets which, in modern networks, should be zero.
  • txqueuelen controls the length of the transmission buffer of the network interface. This parameter is relevant only for some queueing disciplines and can be overwritten using the tc command. For more information about queueing disciplines, take a look at this deep dive into Queueing in the Linux Network Stack and the tc-pfifo main page.

netstat

To see more detailed network statistics for the protocols TCP, UDP, IP, and ICMP you can use netstat -s. This returns a lot of information and the output format is in a human-readable format, like the number of retransmitted and dropped packets sorted by protocol. If you want to focus on TCP retransmissions you can filter out the relevant information.

netstat
netstat shows details about TCP retransmissions

netstat shows that there are 54 retransmitted segments. Meaning, for 54 TCP segments the corresponding ACK was not received within the timeout. Three TCP segments were “fast retransmitted” following the fast retransmission algorithm in RFC 2581. TCP SYN retransmission can happen if you want to connect to a remote host and the port on the remote host isn’t open (see example below).

tcp syn retransmission
Trying to connect to a closed port increases the TCP SYN retransmission counter

ethtool

This tool allows you to query and control the settings of the network interface and the network driver, as seen before. It shows you a detailed list of all errors that can occur on the network interface level, like CRC errors and carrier errors. If you have no retransmissions on the TCP layer but ifconfig still shows you a lot of erroneous packets, this is the place to look for the specifics. If a lot of errors show up in the ethtool output, it usually means that there is something wrong with the hardware (NIC, cable, switchport).

ethtool output
ethtool knows everything there is to know about your network interface, including errors

Some might still want to dig deeper to find out everything about those errors. The next step would be to read the Linux Device Drivers book, digest it, and then start reading through the kernel source code (e.g., linux/netdevice.h) and network driver code (e.g., Intel e1000 driver).

Three helpful tools for gathering information about network errors

tcpretrans

tcpretrans is part of the perf-tools package. It offers you a live ticker of retransmitted TCP segments, including source and destination address and port, and TCP state information. If you suspect that more than one application or service is responsible for TCP retransmissions, tcpretrans allows you to debug your network connections if you call your services in isolation from each other and watch the output of tcpretrans.

tcpretrans
netstat shows details about TCP retransmissions

tcpdump

tcpdump is a command-line network analyzer that shows the traffic specified by filters directly on the command line. With a command line  parameter you can write the output to a file for future analysis. tcpdump is available in almost every *nix distribution out of the box and is therefore the tool of  choice for a quick pragmatic network analysis.

Wireshark

Wireshark, formerly ethereal, is the Swiss Army knife of network and protocol analyzer tools for Windows and Unix when it comes to analyzing TCP sessions, identifying failed connections, and seeing all network traffic that travels to and from your computer. You can configure it to listen on a specific network interface, specify filters to, for example, concentrate on a certain protocol, host, or port, and you can dump captured traffic to a file for an future analysis. Also, Wireshark can read tcpdump files, so you can capture traffic on one host on the command line and open the file for analysis in Wireshark on your computer for an analysis. Another feature of wireshark is that it knows a lot of common application protocols (e.g., HTTP and FTP). Thus you can see what’s going on above layer 4 and get insight into the payloads that are sent using TCP.

wireshark
Wireshark shows MySQL payload

Following is an overview that shows which OSI layers the tools mentioned above cover and on which OSI layer the above-mentioned network errors occur.

<p>Network errors and analysis tools assigned to OSI layers</p>
Network errors and analysis tools assigned to OSI layers

Now you know how and where to find information about network errors. But what can you learn from this information? For starters, you can learn what type of errors you’re dealing with, which will guide your further investigation. Though do you really need to investigate anything at all? After all, investigating each retransmitted or dropped packet is pointless—the network protocol stack has self-healing powers and some of the alleged errors are simply part of the game.

What really counts

Usually, more than one computer, switches, and routers are involved in networking. When you have several hosts and detect a problem in your network it’s not efficient to ssh each computer and perform all these exercises to find out what’s going on. Ultimately, in more complex environments, you need tool support to stay on top of things.

You need a monitoring tool that monitors all the hosts that are part of your infrastructure — a tool that notifies you when something out of the ordinary occurs. The tool should automatically create performance baselines for all running services, as well as incoming/outgoing network traffic, average response time to service calls, and the availability of the service from the network’s point of view. You need to be notified if any of these measurements fall in comparison to the baseline.

Although network errors may be the root cause of why your services aren’t available or are performing poorly, as a service provider in the real world you shouldn’t need to focus on networking errors. Your main concern should be providing high-performance services that are easy to use and always available. In general, you don’t want to be notified about all errors that occur in the network layer (or anywhere else in the application for that matter).

There are a number of networking and service-related metrics you can measure and evaluate. The following three are a good starting point.

Network traffic

Measuring network traffic provides a good overview of the overall usage and performance of your service. It’s also a good indicator of whether or not you need to upscale your infrastructure (e.g., your one server may no longer be enough to handle all the load).

Responsiveness

Responsiveness measures the time from the last request packet that the service receives to the first response packet that the service sends. It measures the time a process needs to produce a response to a given request and should be watched in correlation with hardware resources.

Connectivity

Connectivity shows the percentage of properly established TCP connections compared to TCP connections that were refused or timed out. It shows when services were available to clients and when they were not, over time.

My tool of choice for network analysis in the datacenter is Dynatrace, but I’m obviously a bit biased. In analyzing the network health of one of my Tomcat servers (see the example below), I found out that my service had a responsiveness time of about 3 ms, not much traffic, and 100% availability over the last two hours.

Navigation from smartscape to process metrics
Navigating from smartscape to process metrics

Now the interesting part is how you can relate network errors to actual service response times. If response times or service availability deviate from the baseline you’ll see a summary of the resulting problems that shows how many users are affected and what the root cause of the issue is. The really neat thing about Dynatrace is how well it integrates all this information to help me assess and fix this problem.

Dynatrace problem view
Dynatrace shows relevant details to problems

If you take a close look at the problem view you’ll see that this problem affected real users, 688 user actions per minute to be specific. Furthermore, you can see that the JavaScript error rate increased and that the root cause of this problem is a crashed couchDB process (i.e., the TCP connectivity rate for the process decreased to 0%). If you click on the process name you’ll see the following screen, where you can clearly see that the TCP connections were refused and the connectivity dropped to 0% while the process was restarted. This is what a common network error looks like from a services’ point of view.

Dynatrace process connectivity
Dynatrace Network monitoring shows processes’ connectivity outage

Conclusion

Assessing the quality of your services on physical hosts with an underlying network consisting of physical switches and physical routers is a piece of cake with the right tools in place. However monitoring connectivity and performance in more complex infrastructures with network overlays and encapsulation, virtual switches that run as applications, and intra-VM traffic that you never see on any physical network interface add additional layers of complexity. But that’s another story. So stay tuned!

A word from my sponsor: Take your network monitoring to a new level with a Dynatrace!

If you’re curious about taking Dynatrace network monitoring for a test drive, you should definitely go for it. There is a free usage tier so you can walk through all the functionality described here and see for yourself how well it works in your own environment.

Discover how you can proactively identify connection issues with Dynatrace.

The post Detecting network errors and their impact on services appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/detecting-network-errors-impact-on-services/feed/ 0
Dynatrace Recognized as a Customers’ Choice in the 2022 Gartner® Peer Insights™ “Voice of the Customer”: Application Performance Monitoring and Observability Report https://www.dynatrace.com/news/press-release/dynatrace-recognized-in-2022-gartner-peer-insights-voice-of-the-customer-apm-and-observability-report/ Wed, 07 Dec 2022 13:00:08 +0000 https://www.dynatrace.com/news/?post_type=press-release&p=55095 WALTHAM, Mass., December 7, 2022 – Software intelligence company Dynatrace (NYSE: DT) today announced Gartner has named it an overall Customers’ Choice in the 2022 Gartner Peer Insights “Voice of the Customer”: Application Performance Monitoring (APM) and Observability report. In addition, Dynatrace was also recognized as a Customers’ Choice in three segment quadrants: Global Enterprise, […]

The post Dynatrace Recognized as a Customers’ Choice in the 2022 Gartner® Peer Insights™ “Voice of the Customer”: Application Performance Monitoring and Observability Report appeared first on Dynatrace news.

]]>
WALTHAM, Mass., December 7, 2022 – Software intelligence company Dynatrace (NYSE: DT) today announced Gartner has named it an overall Customers’ Choice in the 2022 Gartner Peer Insights “Voice of the Customer”: Application Performance Monitoring (APM) and Observability report. In addition, Dynatrace was also recognized as a Customers’ Choice in three segment quadrants: Global Enterprise, Large Enterprise, and North America.

Gartner based its analysis on feedback and ratings from end-user professionals with experience purchasing, implementing, and/or using Dynatrace. The Customers’ Choice distinction, or upper-right placement in the report, means Dynatrace meets or exceeds market average ratings from users in both the Overall Experience and User Interest and Adoption categories.

Customers rated the Dynatrace® platform 4.5 out of 5.0 stars, with 94% saying they would recommend it (as of September 2022.) A few of these customers’ reviews include:

“It is always a pleasure to receive recognition from Gartner Peer Insights, and it’s even better to be recognized by our customers,” said Steve Tack, SVP of Product Management at Dynatrace. “We are the only company to be named both a Leader in the 2022 Gartner Magic Quadrant for APM and Observability and a Customers’ Choice in the 2022 Voice of the Customer report for this market. I believe this reflects our focus on empowering the innovators in the world’s largest organizations to drive digital transformation at scale. Feedback from these customers and their willingness to recommend Dynatrace is the best indicator of the results and the positive impact we deliver.”

This Voice of the Customer: APM and Observability report adds to Gartner’s recent research including Dynatrace. Dynatrace is the only provider named a 2022 Customer’s Choice for APM and Observability and a Leader in the 2022 Gartner Magic Quadrant™ for APM and Observability. In addition, the 2022 Gartner Critical Capabilities for APM and Observability report evaluated 19 vendors. It recognized the Dynatrace platform with the highest overall scores in 4 of 6 use cases, including DevOps/AppDev, SRE/Platform Operations, IT Operations, and Digital Experience Monitoring.

Gartner Disclaimers

Gartner, Peer Insights ‘Voice of the Customer’: Application Performance Monitoring and Observability, 30 November 2022

Gartner, Magic Quadrant for Application Performance Monitoring and Observability, By Padraig Byrne, Gregg Siegfried, Mrudula Bangera, 7 June 2022.

Gartner, Critical Capabilities for Application Performance Monitoring and Observability, By Gregg Siegfried, Mrudula Bangera, Padraig Byrne, 8th June 2022.

Gartner® and Peer Insights™ are trademarks of Gartner, Inc. and/or its affiliates. All rights reserved. Gartner® Peer Insights™ content consists of the opinions of individual end users based on their own experiences, and should not be construed as statements of fact, nor do they represent the views of Gartner or its affiliates. Gartner does not endorse any vendor, product or service depicted in this content nor makes any warranties, expressed or implied, with respect to this content, about its accuracy or completeness, including any warranties of merchantability or fitness for a particular purpose.

Gartner does not endorse any vendor, product or service depicted in its research publications and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner research publications consist of the opinions of Gartner’s research organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this research, including any warranties of merchantability or fitness for a particular purpose.

The post Dynatrace Recognized as a Customers’ Choice in the 2022 Gartner® Peer Insights™ “Voice of the Customer”: Application Performance Monitoring and Observability Report appeared first on Dynatrace news.

]]>
Why averages suck and percentiles are great https://www.dynatrace.com/news/blog/why-averages-suck-and-percentiles-are-great/ https://www.dynatrace.com/news/blog/why-averages-suck-and-percentiles-are-great/#comments Mon, 14 Nov 2022 11:04:50 +0000 http://blog.dynatrace.com/?p=5543 Dynatrace employees

Anyone who ever monitored or analyzed an application uses or has used averages. They are simple to understand and calculate. We tend to ignore how wrong the picture is that averages paint of the world. To emphasize the point, let me give you a real-world example outside of the performance space I recently read in […]

The post Why averages suck and percentiles are great appeared first on Dynatrace news.

]]>
Dynatrace employees

Anyone who ever monitored or analyzed an application uses or has used averages. They are simple to understand and calculate. We tend to ignore how wrong the picture is that averages paint of the world. To emphasize the point, let me give you a real-world example outside of the performance space I recently read in a newspaper.

The article explained that the average salary in a certain region in Europe was 1,900 Euro (to be clear, this would be good in that region!). However, when looking closer, they found that most people, namely 9 out of 10 people, only earned around 1000 Euro and one would earn 10,000 (I oversimplified this of course, but you get the idea). If you do the math, you’ll see that the average of this is indeed 1,900 Euro, but we can all agree that this does not represent the “average” salary, as we would use the word in daily life. So now let’s apply this thinking to application performance.

The average response time

The average response time is by far the most commonly used metric in application performance management. We assume this represents a “normal” transaction, but this would only be true if the response time is always the same (all transactions run at equal speed) or the response time distribution is roughly bell-curved.

A Bell curve represents the "normal" distribution of response times in which the average and the median are the same. I rarely ever occurs in real applications
A Bell curve represents the “normal” distribution of response times in which the average and the median are the same. I rarely ever occurs in real applications

In a Bell Curve, the average (mean) and median are the same. In other words, observed performance would represent the majority (half or more than half) of the transactions.

In reality, most applications have few heavy outliers. A statistician would say the curve has a long tail. A long tail does not imply many slow transactions, but a few magnitudes slower than the norm.

This is a typical Response Time Distribution with few but heavy outliers - it has a long tail
This is a typical Response Time Distribution with few but heavy outliers – it has a long tail. The average here is dragged to the right by the long tail.

We recognize that the average no longer represents the bulk of the transactions, but can be much higher than the median.

You can now argue that this is not a problem, as long as the average doesn’t look better than the median. I would disagree, but let’s look at another real-world scenario experienced by many of our customers:

This is another typical Response Time Distribution. Here we have quite a few very fast transactions that drag the average to the left of the actual median
This is another typical Response Time Distribution. Here we have quite a few very fast transactions that drag the average to the left of the actual median

In this case, a considerable percentage of transactions are very, very fast (10-20 percent), while the bulk of transactions are several times slower. The median would still tell us the true story, but the average all of a sudden looks a lot faster than most of our transactions actually are. This is typical in search engines or when caches are involved. Some transactions are very fast, but the bulk are normal. Another reason for this scenario are failed transactions, more specifically transactions that failed fast. Many real-world applications have a failure rate of 1-10 percent (due to user errors or validation errors). These failed transactions are often magnitudes faster than the real ones, and consequently distorted an average.

Of course, performance analysts are not stupid and regularly try to compensate with higher frequency charts (compensating by looking at smaller aggregates visually) and by taking in minimum and maximum observed response times. However, we can often only do this if we know the application very well. Those unfamiliar with the application might easily misinterpret the charts. Because of the depth and type of knowledge required for this, it’s difficult to communicate your analysis to other people. Think how many arguments between IT teams have been caused by this. And that’s before we even think about communicating with business stakeholders!

A better metric by far are percentiles, because they allow us to understand the distribution. But before we look at percentiles, let’s take a look at a key feature in every production monitoring solution: Automatic baselining and alerting.

Automatic baselining and alerting

In real-world environments, performance gets attention when it is poor and negatively impacts the business and users. But how can we quickly identify performance issues to prevent negative effects? We cannot alert on every slow transaction since there are always some. In addition, most Operations teams have to maintain a large number of applications and are not familiar with all of them, so manually setting thresholds can be inaccurate, painful, and time-consuming.

The industry has come up with a solution called Automatic Baselining. Baselining calculates the “normal” performance and only alerts us when an application slows down or produces more errors than usual. Most approaches rely on averages and standard deviations.

Without going into statistical details, this approach again assumes the response times are distributed over a bell curve:

The Standard Deviation represents 33% of all transactions with the mean as the middle. 2xStandard Deviation represents 66% and thus the majority, everything outside could be considered an outlier.
The Standard Deviation represents 33% of all transactions with the mean as the middle. 2xStandard Deviation represents 66% and thus the majority; everything outside could be considered an outlier. However, most real-world scenarios are not bell-curved…

Typically, transactions that are outside 2 times standard deviation are treated as slow and captured for analysis. An alert is raised if the average moves significantly. In a bell curve, this would account for the slowest 16.5 percent (and you can of course adjust that), however, if the response time distribution does not represent a bell curve it becomes inaccurate. We either end up with a lot of false positives (transactions that are a lot slower than the average but when looking at the curve lie within the norm) or we miss a lot of problems (false negatives). In addition, if the curve is not a bell curve than the average can differ a lot from the median, applying a standard deviation to such an average can lead to quite a different result than you would expect! To work around this problem these algorithms have many tunable variables and a lot of “hacks” for specific use cases.

Percentile vs average

A percentile tells me at which part of the curve I am looking at and how many transactions are represented by that metric. To visualize this look at the following chart:

Average vs percentiles: This chart shows the median and 90th percentile along with the average of the same response time. It shows that the average is influenced far more heavily by the 90th, thus by outliers and not by the bulk of response times.
Average vs percentiles: This chart shows the median (50th percentile) and 90th percentile along with the average of the same response time. It shows that the average is influenced far more heavily by the 90th, thus by outliers and not by the bulk of response times.

As you can see in the above graph, the average is very volatile. The other two lines represent the median and 90th percentile. As we can see the median is rather stable but has a couple of jumps. These jumps represent real performance degradation for the majority (50%) of the transactions. The 90th percentile (this is the start of the “tail”) is a lot more volatile, which means that the outliers’ slowness depends on data or user behavior. What’s important here is that the average is heavily influenced (dragged) by the 90th percentile, the tail, rather than the bulk of the transactions.

If the 50th percentile (median) of a response time is 500ms, that means that 50% of my transactions are either as fast or faster than 500ms. If the 90th percentile of the same transaction is at 1000ms it means that 90% are as fast or faster and only 10% are slower. The average, in this case, could either be lower than 500ms (on a heavy front curve), a lot higher (long-tail), or somewhere in between. A percentile gives me a much better sense of my real-world performance because it shows me a slice of my response time curve.

For exactly, that reason percentiles are perfect for automatic baselining. If the 50th percentile moves from 500ms to 600ms I know that 50% of my transactions suffered a 20% performance degradation. You need to react to that.

In many cases, the 75th or 90th percentile does not change at all in such a scenario. This means the slow transactions didn’t get any slower; only the normal ones did. Depending on how long your tail is, the average might not have moved at all in such a scenario!

In other cases, we see the 98th percentile degrading from 1s to 1.5 seconds, while the 95th is stable at 900ms. This means your application is stable, but a few outliers got worse, nothing to worry about immediately. Percentile-based alerts do not suffer from false positives, are a lot less volatile and don’t miss any important performance degradations! Consequently, a baselining approach that uses percentiles does not require a lot of tuning variables to work effectively.

The following screenshot explains how percentile-based alerts do not suffer from false positives, are a lot less volatile, and don’t miss important performance degradation.

Percentile-based alerting
Percentile-based alerts do not suffer from false positives and are a lot less volatile.

How can we use percentiles for tuning?

Percentiles are also great for tuning and giving your optimizations a particular goal. Let’s say that something within my application is too slow in general and I need to make it faster. In this case, I want to focus on bringing down the 90th percentile. This would ensure that the overall response time of the application goes down. In other cases, I have unacceptably long outliers I want to focus on bringing down response time for transactions beyond the 98th or 99th percentile (only outliers). We see many applications that have perfectly acceptable performance for the 90th percentile, with the 98th percentile being magnitudes worse.

In throughput-oriented applications on the other hand I would want to make the majority of my transactions very fast while accepting that optimization makes a few outliers slower. I might therefore make sure that the 75th percentile goes down while trying to keep the 90th percentile stable or not getting a lot worse.

I could not make the same kind of observations with averages, minimum and maximum, but with percentiles, they are very easy indeed.

Conclusion

Averages are ineffective because they are too simplistic and one-dimensional. Percentiles are a really great and easy way of understanding the real performance characteristics of your application. They also provide a great basis for automatic baselining, behavioral learning, and optimizing your application with a proper focus. In short, percentiles are great!

Start a free trial!

Dynatrace is free to use for 15 days! Want to see intelligent autobaselining in action? Just enter your email address, and get up and running in under 5 minutes.

The post Why averages suck and percentiles are great appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/why-averages-suck-and-percentiles-are-great/feed/ 23
Prevent potential problems quickly and efficiently with Davis exploratory analysis https://www.dynatrace.com/news/blog/davis-ai-exploratory-analysis/ https://www.dynatrace.com/news/blog/davis-ai-exploratory-analysis/#respond Tue, 25 Oct 2022 12:00:01 +0000 https://www.dynatrace.com/news/?p=53962 How to implement an AIOps strategy at scale

Dynatrace saves time for site reliability engineers (SREs) by bringing the full power of Davis® AI to exploratory and proactive analyses. When a negative trend arises for an important service level indicator (SLI)—for example, a slowly depleted error budget for availability, request latency, or throughput—SREs can now trigger Davis to perform exploratory analysis. Davis surfaces all signal anomalies connected to the slowly depleted error budget and provides an explanation. This allows SREs to understand the origins of unexpected changes. Davis considers all domain and topology knowledge to identify corresponding observability signals.

The post Prevent potential problems quickly and efficiently with Davis exploratory analysis appeared first on Dynatrace news.

]]>
How to implement an AIOps strategy at scale

SREs typically start their days with meetings to ensure that various stakeholders—such as clients, IT team leaders, Kubernetes platform operations team leaders, and business leaders—that their projects are developing smoothly. Each group of SREs has different needs and goals that must align at the end of the day. To ensure continuous availability, it‘s essential to proactively analyze potential problems and optimize the environment in advance to minimize the negative impact on users and improve user experience.

With Davis, Dynatrace enables rapid MTTR for SRE and DevOps teams by identifying the path to the root causes of detected problems.

With the increasing complexity of cloud-native environments, the number of observed signals grows, as does the effort required for humans to find and analyze these signals. This increased complexity makes it impossible to analyze all relevant situations. The proper focus and best optimization level must be chosen wisely to get the most out of the available time.

Just one click to your preventive analysis

Dynatrace now goes a step further and makes it possible for SREs and DevOps to perform proactive exploratory analysis of observability signals with intelligent answers. This is done by extending our Davis AI engine with a new capability that considers domain and topology knowledge. This significantly reduces the time needed to assess and analyze potential problems and helps prevent production outages.

If one or more anomalies occur, all relevant observability data in the domain context can be displayed with just one click. This is done without the need to create custom dashboards and is complemented by efficient analysis capabilities that automatically guide SREs to potential root causes of anomalies, enabling more efficient work and freeing up time for essential workflows.

“The work of SREs and platform owners is becoming increasingly complex due to the explosion of information that we need to deal with to operate a production environment properly. With Davis exploratory analysis we can now automatically analyze thousands of signals before incidences even arise. This saves valuable time for engineers and architects for innovation.” Henrik Rexed, Open-Source Advocate

Let’s look at an example related to Kubernetes.

Example: Unintended side effects of introducing service mesh technology

With the distribution of Kubernetes, there is growing interest in using service mesh technology to add secure service-to-service communication and fine-grained management of ingress/egress traffic rules while keeping platform operations teams in the driver’s seat.

In this example, a K8s platform operator installs a service mesh that uses a sidecar container as a proxy managing the communication with the various services of the cluster. This introduces an unwanted pattern when used with CronJobs. The sidecar container remains active even when the job is already completed. Since a CronJob pod is only deleted after all containers have been stopped, such pods continue to run and block the requested resources indefinitely. These “zombie” pods are never removed without human intervention, resulting in an accumulation of blocked but unused resources. In this case, a CronJob that creates a new pod every time it runs is executed periodically, for example, once an hour. This leads to systematic growth of used resources and could lead to an unhealthy cluster where zombie pods slowly consume all the resources of the nodes and negatively impact the cluster-health SLO. The SLO provides a ratio of the number of healthy nodes to unhealthy nodes. The SRE begins investigating this issue when less than 90% of the nodes are healthy.

Avoid the zombie-pod apocalypse with Davis exploratory analysis

Now let’s look at how to diagnose and prevent such a zombie-pod apocalypse, which can occur when introducing service meshes or other tools deployed via sidecar containers. You’ll see how to prevent zombie pods and how Kubernetes best practices could have reduced the impact of this problem.

First, the SRE notices a deviation in an SLO representing the health of the K8s cluster.

Node Availability of the cluster Dynatrace screenshot

The SRE needs to identify the namespace that’s consuming this cluster CPU. Looking at the “Kubernetes cluster overview,” it’s clear that the Otel-demo namespace is the namespace allocating most of the CPU requests of the cluster.

Kubernetes namespaces by CPU requests

With this knowledge, the SRE analyzes the otel-demo namespace and notices a time frame during which there were multiple CPU resource spikes. The SRE selects this time frame for Davis exploratory analysis, which takes domain knowledge and topological context into account, and analyzes the problematic workload.

This workload is a CronJob, which creates new pods that run concurrently, rather than sequentially. This CronJob is responsible for the increased CPU usage due to the accumulation of running pods.



Video thumbnail

Looking at the loadgeneratorservice workload, the SRE can further analyze the reasons behind the increased number of pods. The available resources are not able to serve the increased load. At this point, the SRE contacts the K8s operator, notifying them that they need to immediately scale their nodes horizontally or vertically to remediate the issue.

Following this initial bandaid solution of adding resources, the K8s operator can dig deeper into the problem. Using Dynatrace to analyze the details of the pod, the K8s operator understands that, while the job container has already exited successfully, the sidecar proxy injected by the service mesh is still running, which prevents the pod from stopping and being deleted.

Following additional research on this issue, the operator finds that the service mesh provides a specific endpoint for this scenario, which must be called by the CronJob before it exists. This information can be passed along to the responsible teams who can resolve this problem.

How to prevent this with K8s best practices

There are three best practices that could have drastically reduced the impact of this problem. First, the application teams responsible for CronJobs could have leveraged the activeDeadlineSeconds spec of a job, ensuring the termination of all running pods created by a job after a fixed time. In this case, the jobs would have exited with a failed status due to the sidecar still running after the timeout period.

The other recommended CronJob setting is to set concurrencyPolicy to Replace. This setting will avoid having multiple pods running and will only leave one pod for this job. With this setting, there would have been an indicator that something is not quite right and at least there wouldn’t be an accumulation of running pods consuming additional cluster resources. Adding this requirement for every job could be enforced with policy agents like Open Policy Agent or Kyverno. Another counter-measure is namespace quotas, which would have at least reduced the scope of exhausted resources to a single namespace instead of the entire cluster.

Having now explained two Kubernetes best practices for reducing the impact of this problem, it’s time to share our recommended solution, which completely solves this issue. When running CronJobs in combination with sidecar proxies, the solution is to delete the sidecar proxy running alongside the same pod at the end of your job.

But how can a proxy container be deleted? This is possible with the help of a feature implemented by most service mesh solutions in the market: An HTTP endpoint offered by the proxy container allows you to stop the container gracefully.

You can do this by sending the following:
– For istio:  HTTP post http://localhost:15020/quitquitquit
– For linkerd: HTTP POST localhost:4191/shutdown

How to get started with Davis exploratory analysis

The new Davis exploratory analysis feature will be released at the beginning of November with Dynatrace SaaS version 1.254.

The new Kubernetes web UI pages shared in this blog post will be available in January 2023.

There are many more use cases

Besides avoiding a Kubernetes zombie-pod apocalypse, various other use cases for exploratory Davis Analysis exist. In fact, this new analysis feature is available not only for Kubernetes pages but also for the host overview page, services pages, queues pages, container pages, and domain-specific unified analysis pages, like F5, or SNMP.

We’ll cover all these scenarios in future blog posts, so please stay tuned for more details.

Kubernetes in the wild report

Uncover global Kubernetes adoption trends, cost-optimization strategies, and key tools driving innovation for thousands of organizations worldwide. This report highlights global trends in the technology’s adoption and usage in production environments from thousands of organizations across diverse industries.

Kubernetes clusters hosted in the Cloud

The post Prevent potential problems quickly and efficiently with Davis exploratory analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/davis-ai-exploratory-analysis/feed/ 0
Business KPI tracking for mobile applications with Dynatrace: The value of an end-to-end platform for mobile app owners https://www.dynatrace.com/news/blog/business-kpi-tracking-for-mobile-applications-with-dynatrace/ https://www.dynatrace.com/news/blog/business-kpi-tracking-for-mobile-applications-with-dynatrace/#respond Thu, 15 Sep 2022 17:55:36 +0000 https://www.dynatrace.com/news/?p=53253 incident response and business analytics

If you’re responsible for the overall performance and success of a mobile application, you likely have many KPIs to measure and monitor, along with business KPI tracking. For example, is your application meeting business goals? Are customers satisfied with the app and leaving positive reviews? What issues are causing users to abandon or uninstall the app? […]

The post Business KPI tracking for mobile applications with Dynatrace: The value of an end-to-end platform for mobile app owners appeared first on Dynatrace news.

]]>
incident response and business analytics

If you’re responsible for the overall performance and success of a mobile application, you likely have many KPIs to measure and monitor, along with business KPI tracking. For example, is your application meeting business goals? Are customers satisfied with the app and leaving positive reviews? What issues are causing users to abandon or uninstall the app? How can you optimize performance and user experience?

To answer these questions for the business as well as work with your mobile developers to prioritize efforts and implement changes, it’s critical to have a single source of truth that provides the operational and business answers you need. An all-in-one software intelligence platform with comprehensive native mobile app monitoring integrated with end-to-end observability meets this need.

This article is part two of the mobile application monitoring series covering the benefits the Dynatrace platform provides for mobile application owners. In the first part, we cover how easy it is for mobile developers to instrument a native mobile application with Dynatrace. Once instrumented, the app immediately starts sending telemetry data back to the platform that application owners can use.

In this second part, we cover how Dynatrace gives mobile app owners visibility into everything they need to know about an app. This visibility includes how users are interacting with it, the user experience, and whether the app is overall healthy, performing well, crash-free, error-free, and providing users across different regions with the same experience. Watch the video below or read on to learn more about the benefits of an end-to-end platform for mobile app owners.

Video thumbnail

When it comes to mobile app development, it’s vital that owners get the full picture. For example, teams can further segment the telemetry data captured from a mobile app based on operating system, device, region, app version, and other custom metrics, to provide more granular insights on users and their behavior. With this information, app owners along with mobile developers and UX leads can make critical decisions, such as where and how to enhance user journeys to continuously improve their overall offering on their mobile channels.

Use cases from the banking industry illustrate the importance of mobile app monitoring

Mobile apps have increased in importance across most industries. Take banking for example. Now, and especially with the increase in digital demand activity due to the pandemic, many people are reluctant to go to bank branches for frequent and common banking transactions. In addition, many use mobile apps almost exclusively instead of their web counterparts. So, it is imperative that app owners have complete visibility and automatic, contextual analysis to know exactly what their users encounter when they navigate through their primary channel for managing their finances.

Using Dynatrace as an end-to-end platform for mobile monitoring, one of Peru’s largest banks was able to increase revenue by almost 40% within a quarter, with the single biggest contributor being increased transactions on their mobile channel by 30%. With an improved and better-performing app, usage went up by 250% over a span of just 6 months. They accomplished these improvements by increasing visibility into business goals, building up customer satisfaction, and raising adoption levels of Dynatrace capabilities within the organization, which also fostered collaboration across different teams.

Tracking every mobile user interaction

Out-of-the-box, Dynatrace provides experience metrics using Apdex ratings to illustrate how users are experiencing a particular application. Application owners can observe over time to understand how many users are interacting with the application. They can also see when new users adopt newer versions of the application as they roll out on the Apple App Store or the Google Play Store. Since not all users update their mobile applications at the same time, app owners can zero in on how users are interacting with particular versions of their app.

Mobile application Dynatrace screenshot

On the same screen, a mobile app owner can see issues or problems in real time for the selected application. They can also expand the user actions captured within the context of a user session when they are interacting with the app. This view captures what that user is doing from the time they start up the application, including logging in, filling out forms, scrolling through widgets, and more. Dynatrace automatically captures these user actions and then tracks them over time. Using this information, Dynatrace automatically baselines these actions so app owners can understand specific functions or journeys they might want to encourage or discourage. If a particular action is slow or more error prone, Dynatrace automatically highlights it and provides the root cause.

For applications that cater to users across multiple cities, regions, or countries, practitioners can drill down to granular levels of detail. For example, this data can determine which vicinity a set of users came from and compare their experience of a particular issue with another set of users from a different location. This can help isolate, identify and resolve issues that are more nebulous at first glance.

Dynatrace also provides more technically detailed information and their impact on the application, such as web requests from 3rd party services, as depicted below on the left, so you can isolate which of these have higher than acceptable error rates likely contributing to a sub-par user experience. Furthermore, unlike web applications, different versions of a mobile app could be installed across devices. Isolating and identifying errors by versions, as depicted below on the right, helps you take the first crucial step towards identifying which new features resulted in an app becoming more error prone. Unlike point solutions, where you would have to switch between at least two different tools, on Dynatrace, with one click you can scroll to find the details you are looking for.

Mobile user sessions Dynatrace screenshot

Dynatrace’s rich privacy controls enable application owners to comply with data privacy laws. For example, the masking feature allows practitioners to effectively limit the capture of sensitive data. You can choose predefined levels of protection or have adequate flexibility to fine-tune the masking options to your organization’s specific needs.

Users data privacy Dynatrace screenshot

Capturing entire user journeys

After analyzing an application’s performance, an application owner can see that Dynatrace captures everything based on user actions out-of-the-box and can analyze them in a variety of ways. These can include timescale, frequency, and other ways to help with business KPI tracking as depicted in the screenshot below. To develop an understanding from a business perspective as to what users are doing, Dynatrace can create visualizations or dashboards tailored specifically to selected mobile application owners.

Capturing an entire user journeys Dynatrace screenshot

An example of this is the mobile user journey dashboard depicted below. A feature of Dynatrace called user session query language (USQL), allows practitioners to pick out very specific user functions. The dashboard represents a funnel to help practitioners understand how many users visited the application, the number of abandons based on entry and exit points, and how much time users spent on particular screens, all within the context of that session. This dashboard below captures the entire user journey from entry, to search, to logging in, and ultimately up to the purchase. It shows the complete end-to-end flow from a business perspective, identifying abandonments and the causes behind them, such as latencies or crashes at specific parts of the user journey that hinder conversions.

Further down, you can see that Dynatrace’s Davis AI isolates the satisfied abandons from those that are not. As an app owner, you can then dive into the frustrated abandons and examine the individual sessions and even play them back using Session Replay to pinpoint the exact reasons behind them.

Mobile User Journey Dynatrace screenshot

Mobile Abandon Analysis Dynatrace screenshot

Tracking mobile application KPIs made easy with Dynatrace

With Dynatrace, mobile application owners can easily capture usage data and view a wide assortment of visualizations from the metrics Dynatrace collects. Whether KPIs are specifically tailored toward business needs or application health and performance metrics, Dynatrace tracks them from the back end to the front.

Watch a video overview of mobile app monitoring with Dynatrace and read part one and part three of the mobile application monitoring series

Start your free trial today to start monitoring and optimizing your mobile apps.

The post Business KPI tracking for mobile applications with Dynatrace: The value of an end-to-end platform for mobile app owners appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/business-kpi-tracking-for-mobile-applications-with-dynatrace/feed/ 0