Kayla Bondy | Dynatrace news https://www.dynatrace.com/news/blog/author/kayla-bondy/ The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Fri, 06 Mar 2026 08:11:44 +0000 en hourly 1 Stop treating all IT problems the same: Prioritize what matters most with business process entities now in Dynatrace Smartscape https://www.dynatrace.com/news/blog/stop-treating-all-it-problems-the-same-prioritize-what-matters-most-with-business-process-entities-now-in-dynatrace-smartscape/ https://www.dynatrace.com/news/blog/stop-treating-all-it-problems-the-same-prioritize-what-matters-most-with-business-process-entities-now-in-dynatrace-smartscape/#respond Thu, 26 Feb 2026 16:43:06 +0000 https://www.dynatrace.com/news/?p=73183 Business Process analytics

Dynatrace introduced a new entity in Smartscape® that connects your IT topology with your critical business processes. Monitoring a business process as an entity via Business Flow establishes a direct link between your IT issues and business operations. This connection helps you prioritize problem resolution by showing which issues impact your business. Group IT services […]

The post Stop treating all IT problems the same: Prioritize what matters most with business process entities now in Dynatrace Smartscape appeared first on Dynatrace news.

]]>
Business Process analytics

Dynatrace introduced a new entity in Smartscape® that connects your IT topology with your critical business processes. Monitoring a business process as an entity via Business Flow establishes a direct link between your IT issues and business operations. This connection helps you prioritize problem resolution by showing which issues impact your business.

Group IT services by business purpose

Your business runs on processes that matter: purchase orders, loan approvals, employee onboarding, and more. But these business processes often span multiple IT systems, making it hard to see the full picture. When something breaks, it becomes nearly impossible to quickly understand which IT issues threaten critical business outcomes.

With Dynatrace Real User Monitoring, organizations can understand the impact of IT problems by seeing how many users are actively using an application, which users are affected, and how failures degrade user experience. Dynatrace also allows telemetry to be scoped and analyzed by business dimensions such as region, team, or customer-facing application, using segments.

This business context provides strong value, but it stops short of showing how individual IT services come together to support an end‑to‑end business process. One way to bridge that gap is tagging, but in practice, that approach is manual and difficult to keep up to date. It rarely reflects the current, real-time reality of dynamic environments, and many legacy systems can’t be tagged at all. Instead of relying on static tags, the integration of business flows within the new Smartscape captures these relationships directly, showing how services collectively support a business process as it actually runs. By showing which IT services underpin a critical business process and how disruptions propagate across them, teams can respond immediately and decisively when issues arise.

The new Smartscape: making business outcomes a part of your topology

Dynatrace models business flow configurations as entities in the new Smartscape, making business processes a first-class part of the topology. Representing critical business processes as entities creates a single source of truth across the enterprise, aligning IT and business priorities. This allows teams to:

  • Understand business purpose across IT and OT, eliminating guesswork about what matters most.
  • Prioritize security risks by business impact by linking vulnerabilities to revenue-critical systems.
  • See business impact during problem resolution by connecting IT problems to the processes that drive growth in the Problems app.
  • Align cost and sustainability with business outcomes by mapping spend or carbon emissions to specific business processes.

Connect your Business Flow configurations to Smartscape

Business events are the cornerstone of any business process monitoring initiative. Each event is automatically enriched with important topology and application context.

Business process milestones (or steps) are defined in Business Flow, where specific types of business events are grouped. By analyzing business-event groups, Dynatrace identifies the topological entities involved in each milestone.

Example

If a business event is collected by a specific process (software component) and a host, and that event corresponds to a step such as “Payment approval”, the business process entity will be linked to that host and process in Business Flow. These links are refreshed every two hours to reflect changes in the environment.

Once connected, business flow entities and their dependencies can instantly be visually explored in the Smartscape app. They can be viewed alongside applications, services, and infrastructure, or isolated to show all IT relationships supporting a specific business process.

A Business Flow entity shown in the Smartscape on Grail view.
Figure 1. A Business Flow entity shown in the Smartscape on Grail view.

Turning business flows into real-time business metrics

When your business flow configuration becomes an entity in Smartscape, it doesn’t just give you visibility; it becomes a continuous engine of business insight. Dynatrace automatically transforms each monitored business process into a powerful stream of business metrics, generating and storing new business events that are enriched with the KPIs that matter most to your organization. This means you’re not just observing how your IT impacts your business; you’re measuring how your business performs in real time.

With this new capability, you’re in complete control:

  • Choose how often metrics are generated, from high-frequency snapshots for fast-moving operations to broader intervals for strategic monitoring of longer business processes.
  • Define the timeframe of analysis, enabling a moving window of business performance that updates continuously.
  • Track every business KPI with precision, from conversion rates to fulfillment times, total revenue, value of approved loans, business exceptions, and more.

These automatically-generated business metrics give you a dynamic, always-on view of how your business processes evolve over time. The result is a smarter, data-driven way to detect trends, anticipate issues, and optimize performance long before it impacts your customers or your bottom line.

Business impact of an IT problem

Modeling business flows as entities allows Dynatrace to show when IT problems impact critical business processes directly in Problems.

Here’s how the analysis works:

  • Root-cause analysis identifies all IT entities impacted by the problem.
  • When exploring a specific problem, the Business Impact Analyzer checks whether affected entities are connected to any business flow entities.
  • Problems displays this connection in a new single-value tile, showing whether one or more business flows are impacted.

From this tile, you can navigate to Business Flow for detailed exploration.

An IT problem affecting a critical business process
Figure 2. An IT problem affecting a critical business process

Problem investigation mode in Business Flow

From Problems, you can drill down into the Business Flow app for affected business flows. This opens in problem investigation mode, which highlights:

  • The impacted business flows from the list of entities.
  • A timeframe aligned with the IT problem.
  • A detailed view of all the business KPIs over time.

Exploring the different parts of the business flow tree or checking the business KPIs over time allows you to discover the direct impact of an IT problem on a business process.

Currently, the analysis focuses on business processes that are connected to affected IT entities. Soon, this will extend to step-level impact analysis—identifying fulfillment drops and enabling navigation to the exact step affected.

Problem investigation mode in Business Flow
Figure 3. Problem investigation mode in Business Flow

Take the next step into business observability

Dynatrace continues to enhance IT problem contextualization with the Business Impact Analyzer, connecting technical issues to critical business processes mapped in Business Flow.

Discover how to create your first business flow configuration and activate it as an entity in Smartscape.

Ready to bring together your business operations and observability teams?

The post Stop treating all IT problems the same: Prioritize what matters most with business process entities now in Dynatrace Smartscape appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/stop-treating-all-it-problems-the-same-prioritize-what-matters-most-with-business-process-entities-now-in-dynatrace-smartscape/feed/ 0
Logs and traces: Why context is everything for seamless investigations https://www.dynatrace.com/news/blog/correlating-logs-and-traces-with-observability/ https://www.dynatrace.com/news/blog/correlating-logs-and-traces-with-observability/#respond Fri, 04 Jul 2025 10:05:27 +0000 https://www.dynatrace.com/news/?p=69744 logs and traces

It’s 3:00 AM. Alerts are firing. Something’s broken, latency is spiking, there’s too much noise, and you’re under pressure to find the root cause fast. You need to be able to understand how your system interacts to solve the problem as soon as possible. But systems just keep getting more complex as your organization adds […]

The post Logs and traces: Why context is everything for seamless investigations appeared first on Dynatrace news.

]]>
logs and traces

It’s 3:00 AM. Alerts are firing. Something’s broken, latency is spiking, there’s too much noise, and you’re under pressure to find the root cause fast. You need to be able to understand how your system interacts to solve the problem as soon as possible. But systems just keep getting more complex as your organization adds new technologies, AI models, and container-based microservices. That’s why, as complexity scales, so does the need for connected insights. In the world of observability, logs and traces serve distinct but complementary purposes. When used together, they unlock a powerful view into system health, performance, and behavior.

The secret lives of logs and traces

In theory, correlating logs and traces should be straightforward. However, in practice, teams often find themselves context-switching to follow the path of a trace and all the logs involved. To understand why, let’s take a closer look at the roles and responsibilities of logs and traces.

Traces: The big picture view

Traces follow the journey of a request as it moves through various services in a distributed system. They provide end-to-end observability of how different components interact, making them ideal for understanding latency, bottlenecks, and service dependencies. Distributed traces connect events into a cohesive timeline, helping engineers see how one service’s performance affects others.

Traces shine when you’re trying to answer questions like, “Where did this request slow down?” or “Which service caused the failure?”

Logs: The detailed detective work

Logs are detailed, timestamped records of events generated by applications and infrastructure. They’re rich in context, often containing error messages, debug information, and custom outputs that developers write into the code. While traces show the flow, logs show the details. Logs can exist independently of traces and are often the first place developers look when something goes wrong.

Logs shine when you’re trying to answer: “What exactly happened here?”

Don’t forget metrics and other telemetry signals

Although we’re focusing here on logs and traces, metrics and other telemetry data are also essential for observability and deeper context. For more about why it’s important to unify the full spectrum of observability signals, see What is observability and Unified observability: Why storing OpenTelemetry signals in one place matters.

The power of correlating logs and traces from a single, full-context platform

Isolated telemetry signals can lead to blind spots and wasted time searching for answers. Some of the main ways to use logs and traces are to simplify troubleshooting, enhance performance, improve security posture, and meet compliance standards.

When you can correlate logs and traces from a single source of observability data, you eliminate the constant context switching that slows down investigations. Instead of toggling between tracing tools and log viewers, you get a unified view that connects the dots fast.

Correlating logs and traces from a single platform transforms troubleshooting from a fragmented hunt into streamlined analysis, where you spend time solving problems instead of searching for information. Core technologies like Grail®, OneAgent®, and Davis® AI provide the scalable foundation while embracing open-source frameworks like OpenTelemetry for flexibility.

Cracking the case of the failed checkout: Investigating logs and traces

Not every investigation starts the same way. Sometimes a trace gives you the high-level view you need to spot an issue and dive deeper. Other times, a log entry is the first clue that something is off. In the next section, we’ll walk through two examples, one that starts with traces and the other with logs, to show how you can get the answers you need.

Scenario 1: Investigating from traces to logs

Investigating from traces to logs in Dynatrace video

While doing some routine monitoring in the Distributed Tracing app, we notice a series of failed requests in our Kubernetes prod namespace. So we filter for unsuccessful transactions to examine them more closely.

One request stands out: “/cart/checkout”. It’s a critical transaction path, and we’re seeing failures.

We dive into the trace waterfall. Just below it, we find the logs tied to each span, giving us deeper insight. That’s where we find the message:

error: failure to complete the order

Distributed Tracing requests in Dynatrace screenshot

Digging further, another log reveals the root cause: only Visa and Mastercard are accepted, which is in line with our policy, but potentially limiting our business. This raises a new question: how often is this happening?

With a single click, we pivot to the logs app, where we can search for this specific message and quantify how many transactions may have been impacted.

This approach turns scattered signals into a cohesive story, helping us move from surface-level symptoms to actionable insights with speed and precision.

Scenario 2: Investigating from logs to traces

Logs to Traces video thumbnail

No matter how you start your day, whether you are coming from PagerDuty, Slack or start directly in Dynatrace through one of the many apps like Kubernetes or the Clouds app, you can always see logs in context of your investigation.

In this scenario, we’re investigating this case from another angle, starting with the logs app using the prefiltered segment for the Kubernetes prod namespace. The view is tailored to the services we own. A quick scan reveals something suspicious: numerous errors in some of the log files.

screenshot of logs affected by errors in logs and traces investigation
Figure 1. A quick scan reveals numerous errors in some log files.

To dig deeper, we navigate in the logs app and use the content filter for “payment” and “error”, and we find several logs with the following message:

Could not charge card for user id = xxxxxxxxxxxxx

But what is causing the failure? We click Show surrounding logs, which reveals all logs associated with the trace ID. Now we can view log messages sequentially as they happened.

Investigating some of the surrounding logs, we see that the user is using a credit card other than Visa or Mastercard, which our organization doesn’t support. Now that we understand why things are failing, let’s investigate further to see if we can optimize this experience.

To understand the full impact, we pivot seamlessly to the trace view. Here, we see the full waterfall breakdown of the request: service calls, timing, and span-level metadata.

One detail stands out: it took 5 seconds for the user to receive the failure message. That’s a long time to wait just to be told their card isn’t supported.

With this insight, we can now make targeted improvements so that the user does not have to wait a long time to understand that their payment method is not supported and deliver a better user experience.

While these examples highlight how seamless navigation between logs and traces accelerates troubleshooting, they’re just one part of the story. With Dynatrace Grail and Notebooks, you can take things a step further by running advanced queries, automating repetitive tasks, and building collaborative, data-rich workflows. These tools empower teams to go beyond reactive troubleshooting and into proactive, scalable observability.

Why seamless navigation between logs and traces matters

Seamless navigation between logs and traces isn’t just a convenience; it’s a game-changer. Whether you start with a trace or a log, the ability to pivot instantly between signals means you spend less time hunting for answers and more time solving problems. It accelerates root cause analysis, improves team collaboration, and gives you the full context needed to act with confidence. This is just one example of how Dynatrace helps you move from fragmented troubleshooting to unified intelligent observability.

Ready to start investigating?

Explore Distributed Tracing and Log Management and Analytics, complete with prepopulated data in the Dynatrace Playground.

Want to get started with your own data instead?

The post Logs and traces: Why context is everything for seamless investigations appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/correlating-logs-and-traces-with-observability/feed/ 0
Distributed tracing best practices for the software development lifecycle https://www.dynatrace.com/news/blog/distributed-tracing-best-practices/ https://www.dynatrace.com/news/blog/distributed-tracing-best-practices/#respond Thu, 01 May 2025 15:01:40 +0000 https://www.dynatrace.com/news/?p=68986 Distributed tracing best practices

Distributed tracing helps analysts throughout the organization understand the relationships among services to troubleshoot problems. But distributed tracing is a crucial capability at every stage of the software development lifecycle. Discover distributed tracing best practices to help troubleshoot, gain critical input for design, feedback for implementation, and precision answers for automating. These best practices help developers deliver seamless services for customers and essential business goals for the organization.

The post Distributed tracing best practices for the software development lifecycle appeared first on Dynatrace news.

]]>
Distributed tracing best practices

For developers, the ultimate win is creating impactful features, shipping them quickly, and confidently implementing fixes when needed. Distributed tracing plays a critical role in achieving these goals. As a core component of observability, it provides the insights developers need to write better code, collaborate efficiently, and increase application reliability.

Developers know how challenging it can be to follow traces from beginning to end across multiple entities and data providers. Harder still is following those traces in context with metrics, logging data, security details, and real user experience data. Certain tools may provide visibility into small segments or isolated details. But developers often have to piece these clues together manually, which makes it difficult to extract meaningful and reliable results.

Dynatrace delivers distributed tracing with capabilities designed to streamline workflows. From understanding the impact of code changes to quickly resolving issues flagged in production, here’s how distributed tracing can help you increase developer velocity and deliver meaningful innovation throughout the software development lifecycle (SDLC).

An figure 8 infinity loop showing the stages of software development where you can apply distributed tracing best practices
Figure 1. Stages of the software development lifecycle where you can apply distributed tracing best practices to deliver better software and automation.

Distributed tracing best practices key takeaways:

  1. Visualize and understand your system. Distributed tracing plays a key role in uncovering the topology, interactions, and dependencies of your services.
  2. Analyze and fix problems fast. Combine distributed tracing with AI analysis, context-driven querying, and collaboration tools.
  3. Build better software with telemetry data in context. Unify telemetry data
  4. Automate to ship better code faster. Use distributed tracing to automate and streamline workflows and set up intelligent alerts.

Visualize and understand your system

Planning

Requirements
gathering

Design

Before implementing changes or designing new features, it’s essential to have a clear understanding of your system’s current state. Without a clear picture of your current system, any change risks overlooking critical dependencies and misalignment with your architecture, whether in the cloud or on-premises. Distributed tracing plays a key role in uncovering the topology, interactions, and dependencies of your services, giving developers the visibility they need to make informed decisions.

Dynatrace Smartscape® technology automatically maps horizontal service relationships, offering real-time visualization of service connections and dependencies. This end-to-end view helps developers quickly understand how components interact within the overall system.

For a more granular understanding, inspecting individual traces allows you to focus on specific workflows or problematic behaviors. This depth of insight can help you plan better and identify challenges early, reducing surprises during later stages of development.

Trace example: Verifying service for “platinum” loyalty users

Say you’re a developer on an app called “AuthenticationService.” You can select a trace that lets you visualize its path through your system and examine each span’s properties to ensure “Platinum” loyalty users are being served effectively. While understanding the state of your system is important at each stage of the SDLC, it’s often one of the first steps in planning, requirements gathering, and designing changes.

In Dynatrace, you would open the Distributed Tracing app and enter a filter for the request attribute that indicates loyalty status. This filter returns a list of traces you can inspect individually to understand any bottlenecks or issues that could be impacting this specific subset of customers.

Distributed Tracing app screehshot showing a slow response time associated with a database call, indicating a service that needs improvement.
Figure 2. A filter in the Distributed Tracing app shows a slow response time associated with a database call, indicating a service that needs improvement.

Analyze and fix problems fast using distributed tracing for developers

Testing

Deployment

Maintenance

Identifying root causes in complex microservice-based architectures can be tough, but Dynatrace Davis® AI analyzes dependencies in real time, helping you quickly and confidently resolve issues.

But what if you want to visualize data outside of an individual problem occurrence? Some issues occur repeatedly in specific areas of a system, while others arise from application changes. Problems often span multiple teams, requiring collaboration to resolve effectively.

For broader insights, the Dynatrace Distributed Tracing app provides flexible access to raw trace data. You can filter by namespace, release version, or database calls, then Dynatrace automatically generates a DQL (Dynatrace Query Language) query so you can view results in notebooks or pin visuals to dashboards, which enable you to collaborate seamlessly with other teams. This approach is particularly useful during testing, deployment, and maintenance, where visibility into system behavior is critical.

This example shows a plot generated by a DQL query that indicates the number of database calls an application makes and how they perform over time.

A line chart showing database calls reveals the executeQuery-select-oracle-DB1 as the service that needs attention.
Figure 3. This plot of database calls reveals the executeQuery-select-oracle-DB1 as the service that needs attention.

Let’s assume you’re also interested in how your database queries are performing. To get an instant performance summary, you can easily add a table to your dashboard that summarizes the count, average, and percentiles of the durations of your database spans.

A DQL query that shows an output that tabulates database query performance.
Figure 4. Easily create a table that shows the performance of your database queries.

Build better software with telemetry data in context

Testing

Deployment

Maintenance

Developers juggle many responsibilities, from testing and deploying code to maintaining stability. However, fragmented workflows and the endless search for the right information can often slow progress and create frustration, making it harder to meet deadlines and maintain a focus on innovation.

As a developer, how can you get the information you need without wasting time chasing down data and trying to understand all its upstream and downstream effects? The answer: distributed traces and telemetry data in context.

Flexible access to unified telemetry data is a game-changer for developers. With Dynatrace, you can seamlessly query multiple telemetry data sets and visualize the results in ways that align with your specific needs. This flexibility empowers you to connect the dots across complex workflows, analyze incidents within their broader context, and make data-driven decisions that support automation and enhance software quality.

The example below brings together spans and logs. Dynatrace automatically enriches every log with trace and span IDs, making it simple to track correlations among related data. With Dynatrace Query Language (DQL), you can dynamically combine spans and related log content on-the-fly.

A DQL query that joins logs and their associated spans
Figure 5. A simple DQL query joins logs and spans to find those affected by the ‘Card verification failed‘ log message

To understand the impact of an error, you can chart which services and endpoints are emitting this log error.

A DQL query and line chart that shows card verification failed errors.
Figure 6: You can plot which services the card verification failed error affects.

As a developer, you can unlock deeper insights, linking operations to specific errors and enabling you to assess the scope and impact of incidents accurately. This gives you and your teams the ability to move beyond reactive problem-solving into proactive, intelligent development.

Use distributed tracing for developers to automate and ship better code faster

Testing

Deployment

Maintenance

Automation is key to accelerating development while maintaining quality. By streamlining workflows, automating repetitive tasks, and setting up intelligent alerts, teams can proactively address issues and minimize manual effort.

For example, as a developer, you can use trace data to establish quality gates or set alerts for specific exceptions, ensuring you can quickly flag and resolve any recurring problems. This approach boosts efficiency and ensures that teams can consistently deliver high-performing, reliable code to production.

Assume you’ve been working on a bug that appeared as an exception in your traces. You want to check if the bug has been fixed and set up an alert to get a reminder if this particular bug comes back.

Let’s start by creating a metric based on this particular exception in the Dynatrace stream-processing data ingestion technology, OpenPipeline™. By creating a metric at ingestion using OpenPipeline’s centralized data handling and advanced data processing capabilities, you can continually monitor for evidence of the exception recurring.

Span Pipeline screen that shows creating a metric to track a data signature
Figure 7. Creating a metric at ingestion with the data signature of the exception helps you ensure the exception is fixed and doesn’t recur.

You can use this extracted metric to visualize its frequency over time in a dashboard.

Screenshot showing a line chart showing exception spans over time
Figure 8. Track your metric over time in a dashboard.

Finally, once you’ve set up a metric, you can use it to create an alert and send a notification to your system of choice (such as Slack, Teams, ServiceNow, and so on).

Increase reliability and automate more with distributed tracing best practices

As development responsibilities continue to expand in today’s complex environments, distributed tracing has emerged as a critical tool for developers throughout the entire software development lifecycle.

Applyin these distributed tracing best practices helps developers visualize systems, resolve issues quickly, and streamline workflows. By providing deep insights and context of upstream and downstream events, traces and spans give teams the data they need to automate tasks and deliver more reliable, higher-quality software faster and with greater confidence. The result allows you to focus on what matters most: creating impactful features.

Start integrating distributed tracing into one part of your workflow today and see how quickly it becomes indispensable to your development practice.

Try out the new Distributed Tracing experience with prepopulated data on the Dynatrace Playground.

If you’re new to Dynatrace and want to explore distributed tracing best practices with your own data, check out our free trial.

The post Distributed tracing best practices for the software development lifecycle appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/distributed-tracing-best-practices/feed/ 0
Distributed tracing with Dynatrace just got even better https://www.dynatrace.com/news/blog/distributed-tracing-with-dynatrace-just-got-even-better/ https://www.dynatrace.com/news/blog/distributed-tracing-with-dynatrace-just-got-even-better/#respond Tue, 11 Mar 2025 14:58:10 +0000 https://www.dynatrace.com/news/?p=68257 Distributed tracing wth Dynatrace and OpenTelemetry

Get ready to experience a whole new world of limitless tracing power. With our latest enhancements, we’re transforming the way you work with trace data. The Dynatrace® platform now enables comprehensive data exploration and interactive analytics across data sets (trace, logs, events, and metrics)—empowering you to solve complex use cases, handle any observability scenario, and […]

The post Distributed tracing with Dynatrace just got even better appeared first on Dynatrace news.

]]>
Distributed tracing wth Dynatrace and OpenTelemetry

Get ready to experience a whole new world of limitless tracing power. With our latest enhancements, we’re transforming the way you work with trace data. The Dynatrace® platform now enables comprehensive data exploration and interactive analytics across data sets (trace, logs, events, and metrics)—empowering you to solve complex use cases, handle any observability scenario, and gain unprecedented visibility into your systems. Whether you’re using OpenTelemetry or OneAgent, operating in the cloud or on-premises—we’ve got you covered.

Introducing a new era of distributed tracing with advanced analytics

In today’s complex systems landscape, understanding the root causes of issues can be daunting, especially as applications scale and OpenTelemetry adds complexity.

Davis® AI automatically pinpoints root causes, offering immediate answers. For deeper exploration, our Distributed Tracing app empowers you to analyze raw trace data and uncover insights, whether troubleshooting errors, optimizing performance, or discovering the “unknown unknowns.”

But why stop there? Building on this solid foundation, we’re thrilled to announce two powerful platform enhancements. Say hello to advanced trace analytics and new data storage and capture options. These game-changing features elevate your data interactions, opening up vast possibilities for advanced queries and efficient data management tailored to your needs.

Figure 1. Explore every detail of your traces with in-depth exception analysis, providing easy access to exception details with full trace context.
Figure 1. Explore every detail of your traces with in-depth exception analysis, providing easy access to exception details with full trace context.

Site reliability engineers, performance architects, and developers can now leverage dynamic analysis tools like dashboards and workflows to explore trends, automate processes, and maintain control at an unprecedented level. Additionally, these queries serve as excellent starting points for more complex data explorations with Notebooks.

Get ready to maximize the full potential of your trace data—unlock deeper insights and automate like never before, all within a single platform.

Level up your analytics game: Enhanced team collaboration and advanced data insights

With traces now stored in Dynatrace Grail™, our scalable data lakehouse, you can unlock powerful new analytics capabilities, handle massive volumes of data, and run complex queries seamlessly. Combining traces with logs, metrics, Kubernetes events, and telemetry attributes gives you a complete, contextual view of your environment for unmatched end-to-end observability.

Unlock deeper insights

Using Dynatrace Query Language (DQL), you can extract game-changing insights from raw span data with precision. Use these queries to start more complex data exploration with Notebooks. This enables you to uncover hidden patterns, discover unknown unknowns, and make confident, data-driven decisions. These powerful insights can easily be transformed into interactive dashboards.

Example: Exception analysis

Understanding patterns, especially regarding exceptions, is no easy feat. However, you can begin unlocking additional insights using the Distributed Tracing app. For example, you can filter to understand endpoint performance where exception messages contain the string, access denied.

Once filtered, you can easily open a notebook with a pre-populated DQL query. You can then add additional details or modify the query as needed. Combine multiple findings in a notebook or dashboard to share your analysis with your team, allowing them to see these focused updates live in real time.

Figure 2. Open a notebook with a pre-populated DQL query, modify it, and share real-time updates with your team.
Figure 2. Open a notebook with a pre-populated DQL query, modify it, and share real-time updates with your team.

Achieve superior analytics

Transform trace data into intelligent insights by combining trace data with logs, metrics, and events. This combination allows you to enrich all data and get details in context. Use this intelligent data to see what matters to you most in real time by creating interactive dashboards and driving better decision-making. This gives you the power to break down silos, spark collaboration, and extract actionable insights with ease. It democratizes access to critical data, ensuring all teams can leverage the same reliable insights to drive impactful outcomes.

Example: Combine trace data with logs

A common scenario is understanding which frontend API requests have log messages on the backend indicative of a specific problem. Let’s look at an example where the log message contains timeout, and we want to understand the response time of traces in the context of these messages.

By linking trace data with logs, you can query across spans and related log messages. You can also summarize with DQL to understand how often a specific pattern occurs. To visualize span duration, use p99 to see the slowest percentile of span response times and then navigate directly to the traces.

Figure 3. Combine trace data with logs to identify frontend API requests with backend "timeout" log messages, and analyze response times in context.
Figure 3. Combine trace data with logs to identify frontend API requests with backend “timeout” log messages and analyze response times in context.

Automate with Dynatrace OpenPipeline

With Dynatrace, you can create custom metrics from trace data using Dynatrace OpenPipeline™, unlocking powerful new automation capabilities. It’s now possible to create metrics on OpenTelemetry and OneAgent spans with any available attribute, giving you the power to define operational, request, and method-level metrics.

OpenPipeline provides the flexibility to build metrics tailored to your specific needs, enabling you to integrate metrics seamlessly with advanced Dynatrace automation features, such as AutomationEngine and SRE Guardian. You can streamline workflows, intelligently automate repetitive tasks, proactively resolve issues, and spend more time innovating with automation, ensuring that only reliable, high-performing code reaches production.

Figure 4. OpenPipeline ingests, processes, and manages observability, security, and business data at any scale.
Figure 4. OpenPipeline ingests, processes, and manages observability, security, and business data at any scale.

We’ve only scratched the surface of scenarios where advanced analytics can make an impact—the possibilities are virtually endless. By combining advanced trace analysis, intuitive query capabilities, and seamless automation, your team can enable sharper analysis, streamline workflows, and foster innovation. This powerful approach ensures you can focus on delivering better outcomes with greater efficiency, empowering your organization to tackle complex issues with precision and agility, ultimately bringing unprecedented value.

Extended trace retention: Retain data longer when it matters

Need to analyze trends over the long term or adhere to compliance requirements? With extended trace retention, you can store trace data for up to 10 years. This feature ensures your organization is well-equipped for trend analysis and detailed post-mortem reviews—all while meeting regulatory requirements.

Retention policies are fully configurable in OpenPipeline. You can share data in buckets with varying retention times depending on the use case, allowing you to target specific applications or error-prone services for longer storage. This precision reduces storage costs while ensuring you retain the data that matters most.

Extended trace ingest

You can now customize trace ingestion rates to meet your specific needs. While most trace data is already ingested at high coverage rates, this option gives you more granular control over your trace volume. Ingest as much data as you want, above and beyond what is already included in your Dynatrace license.

Maximize the value of your OpenTelemetry data

At Dynatrace, we love OpenTelemetry. We champion open source innovation and recognize OpenTelemetry’s influence in setting observability standards (we’re also a top contributor). That’s why our advanced capabilities were designed from the ground up with OpenTelemetry at its core. We built our entire new tracing experience on OpenTelemetry semantic conventions and expanded from there. Now, OpenTelemetry users can troubleshoot and analyze while leveraging  OTel standards. This commitment empowers you to simplify complexity and innovate faster by extracting maximum value from your data, regardless of origin.

Experience the future of Distributed Tracing

At Dynatrace, we believe that observability should be effortless and completely on your terms. With our latest advancements, we’re helping you manage complexity, innovate faster, and push boundaries. Our solution adapts seamlessly to your ecosystem, whether you use OpenTelemetry or run cloud-native or on-premises workloads. Built to handle enterprise scale, the Dynatrace platform processes massive volumes of data in real time while unifying insights across all teams in your organization.

We’re rolling out this functionality to existing Dynatrace Platform Subscription (DPS) customers. Elevate your observability journey with these new possibilities—tailored to your needs, your way.

If you’re not a DPS customer, you can try out the new Distributed Tracing experience with prepopulated data on the Dynatrace Playground.

If you’re new to Dynatrace and want to try out the new Distributed Tracing experience with your own data, check out our free trial

Take the leap today and discover how Dynatrace can revolutionize your approach to observability.

The post Distributed tracing with Dynatrace just got even better appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/distributed-tracing-with-dynatrace-just-got-even-better/feed/ 0
New Distributed Tracing app provides effortless trace insights https://www.dynatrace.com/news/blog/new-distributed-tracing-app-provides-effortless-trace-insights/ https://www.dynatrace.com/news/blog/new-distributed-tracing-app-provides-effortless-trace-insights/#respond Wed, 23 Oct 2024 19:37:45 +0000 https://www.dynatrace.com/news/?p=66311 Distributed Tracing app graphic

In today’s digital era, the shift to the cloud and complex microservices architectures makes understanding the flow of requests across various services essential for maintaining system reliability and performance. Distributed tracing provides visibility, allowing you to track and analyze requests across system components. Such insights are invaluable for discovering unknown unknowns, diagnosing issues, and ensuring efficient operations. However, as systems grow more complex, the volume of data can be overwhelming, making it difficult to identify trends, pinpoint issues, and isolate root causes.

The post New Distributed Tracing app provides effortless trace insights appeared first on Dynatrace news.

]]>
Distributed Tracing app graphic

We’re excited to announce the first version of our new Distributed Tracing app, a part of the new Dynatrace user experience that leverages the full power of the Dynatrace platform. With the Distributed Tracing app, you can flexibly slice and dice raw trace data to understand what went wrong and why. Find what you’re looking for faster with:

  • Enhanced charting and data visualization: Easily filter, group, search, and visualize trace data to gain deeper insights into your system’s behavior.
  • Automatic data capture and display: More data, including span attributes, is available for out-of-the-box analysis, with no additional configuration necessary.
  • Seamless OpenTelemetry integration: Make the most of your trace data with native support of OpenTelemetry traces. For more details, see our recent blog post explaining how new Dynatrace capabilities help modern app teams analyze OpenTelemetry traces and log data at scale.

Whether you’re troubleshooting a specific issue or looking to improve overall system performance, Distributed tracing equips you with the tools you need to make informed decisions and maintain a high standard of application performance.

To understand the benefits of the Distributed Tracing app, let’s take a look at a typical scenario.

Use Distributed Tracing to improve application performance and troubleshoot faster

In this scenario, an e-commerce business uses Dynatrace to monitor the performance of its online store. They use Kubernetes to power their marketplace. The team decides to dig into the “prod” namespace to perform exploratory analysis of their critical production workloads.

By opening the time series view filtered by the “prod” cluster, the team immediately notices spikes in the 90% decile of request response times. These performance outliers in production are impacting customer experience, so the team needs to investigate further.

Distributed Tracing Explorer chart namespaces in Dynatrace screenshot

By analyzing the response time distribution in the histogram, the team notices that the outliers occur when the response time is around 5 seconds.

Next, the team leverages the interactive chart, hovering over the outlier requests to see real-time details. In the image below, they select the range of slow requests (3.7 s – 7.24 s) to investigate further.

Distributed Tracing Explorer chart requests in Dynatrace screenshot

Now filtered, the image below shows only requests in the time bucket selected (3.7 s -7.24 s).

Distributed Tracing Explorer chart in Dynatrace screenshot

The filter bar displays all the filters applied during the analysis. To better understand where slow response times are occurring, the e-commerce team decides to group requests by service and endpoint.

To focus on an essential endpoint for the e-commerce website, they use a wildcard (*) to filter on endpoints that start with "/cart".

Distributed Tracing Explorer filter in Dynatrace screenshot

This investigation reveals that requests in the “/cart/checkout” endpoint are failing. The team filters further by the “/cart/checkout endpoint” attribute value.

Distributed Tracing Explorer cart checkout in Dynatrace screenshot

To pinpoint the exact requests that are failing, the e-commerce team filters by excluding successful HTTP 200 status codes. This refinement reveals that only a few requests are failing. The team can now dive deeper to find out why.

Distributed Tracing Explorer chart in Dynatrace screenshot

To understand what happened in detail, the team clicks on an impacted trace and opens the waterfall view of the full trace.

In the waterfall view, the team can quickly and easily switch among multiple traces and their attributes when analyzing issues. A span can have many different attributes, and the search function helps the team quickly find interesting insights. They search for “failure” and explore the exception tab. Here we’ve found some exceptions that happen rarely.

In this view, the root cause of the issue is clear: an exception is generated when a user tries to purchase an item with a card that is not a Visa or Mastercard. Instead of returning a 500 error, the application should provide the user with a failed payment message letting them know their card is the reason why the transaction was not completed. With these details, the issue is found, and they have the information needed to escalate the situation to their development teams and prioritize the creation and deployment of the needed fix.

With logs now integrated directly into the waterfall view, the team can see detailed log messages tied to each span without switching tools. This seamless access to both trace flow and detailed log context accelerates the investigation from symptom identification to root cause analysis

Explore service telemetry data

With the Services app you can view traces in an aggregated format, easily sift through problems, and ensure general services are functioning properly. A comprehensive view of trace data organized by services provides additional context.

When a problem arises, the Services app is an excellent tool for analysis. Davis® AI automatic root cause analysis highlights abnormal behaviors, such as increased failure rates at the /cart/checkout endpoint, in real time to accelerate the analysis process

Distributed Tracing Explorer chart in Dynatrace screenshot

Get started with the Distributed Tracing and Services apps

If you’re new to Dynatrace and want to try out the Distributed Tracing app, check out our free trial.

We’re rolling out this new functionality for our existing Dynatrace Platform Subscription (DPS) customers. As soon as the new Distributed Tracing Experience is available for your environment, you’ll see a teaser banner in your classic Distributed Traces app.

If you’re not yet a DPS customer, you can use the Dynatrace playground instead. You can even walk through the same example above. The new Services app is already available to all DPS and non-DPS customers.

This is just the beginning. stay tuned for more enhancements and features.

Make your voice heard after you’ve tried out this new experience. Provide feedback for Distributed Tracing in the Distributed Tracing feedback channel (Dynatrace Community). To share your feedback regarding the Services app, go to the Services feedback channel (Dynatrace Community).

The post New Distributed Tracing app provides effortless trace insights appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/new-distributed-tracing-app-provides-effortless-trace-insights/feed/ 0
CrowdStrike update: How Dynatrace helped customers recover in hours https://www.dynatrace.com/news/blog/crowdstrike-update-crisis-dynatrace-customers-recovery/ https://www.dynatrace.com/news/blog/crowdstrike-update-crisis-dynatrace-customers-recovery/#respond Wed, 31 Jul 2024 13:29:47 +0000 https://www.dynatrace.com/news/?p=65026 CrowdStrike update

The ability to recover quickly in case of a sudden IT outage is crucial for business resilience. This blog is part of a series that explores how organizations can maintain business resilience by having the right capabilities to recover quickly from an IT outage.

The post CrowdStrike update: How Dynatrace helped customers recover in hours appeared first on Dynatrace news.

]]>
CrowdStrike update

On July 19th, 2024, countless organizations had their operations disrupted by a routine software update from CrowdStrike, a popular cybersecurity software. The resulting outages wreaked havoc on customer experiences and left IT professionals scrambling to quickly find and repair affected systems.

A wide variety of companies and industries have suffered the effects of this incident, from delayed flights to disruptions in healthcare, insurance, and the financial industry. The ripple effects on the global supply chain have been equally significant. The crisis has emphasized the importance of having a strategy for maintaining stability and performance.

Time is of the essence in any crisis—so is having the right tools and capabilities. Although Dynatrace can’t help with the manual remediation process itself, end-to-end observability, AI-driven analytics, and key Dynatrace features proved crucial for many of our customers’ remediation efforts.

Following are some of the critical capabilities many of our customers relied on to recover from the CrowdStrike update crisis in just hours.

1. Real-time monitoring with out-of-the-box features

Real-time data and monitoring are crucial for maintaining situational awareness of IT environment stability and performance, especially during a crisis. Knowing what’s offline and which dependencies connect to mission-critical services is key to determining the impact of an incident and determining where to start remediation. Dynatrace offers various out-of-the-box features and applications to provide a high-density overview of system health for all hosts and related metrics in a single view.

Smartscape topology mapping

Dynatrace Smartscape® provides a dynamic, real-time visualization of the entire application topology so teams can quickly identify and address issues. Understanding application dependencies helps teams prioritize what to address first. For example, a good course of action is knowing which impacted servers run mission-critical services and remediating those first.

Problems application

The Problems application automatically identifies issues, collects the context behind them, and presents their root cause and impacts in a single view. Powered by Davis® AI, the app helps teams immediately see the problem’s duration, root cause, and business impact.

Dashboards and visualizations

Standard dashboards and visualizations also provide situational awareness out of the box. The following honeycomb visualization shows a healthy environment before an incident and after the incident starts. All the problems, offline hosts, databases, and failing services appear in red.

Honeycomb visualization: Before CrowdStrike outage
Systems before the outage
Honeycomb visualization: CrowdStrike outage in progress
Systems during the outage

Dynatrace OneAgent full-stack and Foundation and Discovery modes

Dynatrace OneAgent provides automatic discovery and monitoring. In addition to using OneAgent for full stack monitoring of the most critical applications, Dynatrace offers OneAgent Foundation and Discovery mode. This lightweight alternative provides full coverage of an environment in scenarios where teams need cost-effective yet comprehensive monitoring. Foundation and Discovery provide essential metrics and topology discovery, making it useful to quickly identify and recover affected hosts.

Together, these technologies enable organizations to maintain real-time visibility and control, swiftly mitigating the impact of incidents and efficiently restoring critical services. They also enable companies to measure the effectiveness of their remediation activities to ensure that recoveries proceed as expected.

How out-of-the-box monitoring features helped one company recover from the CrowdStrike outage within hours

A US-based pharmaceutical company recovered its most critical systems within hours of the CrowdStrike incident using out-of-the-box real-time monitoring features. Dynatrace automatically found the hosts that were unavailable or having problems. The key information displayed on the standard Dynatrace Problems app and the Infrastructure and Operations App became the basis of their team’s remediation plan. The company was back to normal business operations as other companies continued to struggle with recovering days after the initial software push.

2. Synthetic monitoring

Synthetic monitoring is a critical tool for ensuring application reliability and performance, especially during a crisis. By simulating user interactions and running tests from various locations worldwide, synthetic monitoring provides a comprehensive view of application performance and availability. This proactive approach allows organizations to detect and resolve issues early, optimize performance, and maintain a high-quality user experience.

Organizations can use synthetic monitoring to continuously monitor API endpoints and ensure that critical user journeys perform well and meet service level agreements (SLAs). This awareness can help reduce downtime and minimize disruption, enabling swift action if an incident occurs.

Many businesses rely on third-party services, such as payment processors, content delivery networks (CDNs), and ticketing systems to get through their day-to-day operations. Even if the business isn’t directly affected by a crisis, their third-party suppliers may be, which can still disrupt their operations.

How synthetic monitoring helped one company get an early warning of the CrowdStrike impact

During the recent CrowdStrike crisis, a US life insurance provider was impacted indirectly through its third-party ticketing system. The Dynatrace synthetic monitors they used to monitor the performance of this critical third-party application immediately detected the outage when the CrowdStrike software affected the vendor’s servers.

Dynatrace created a problem notification and Davis AI determined the root cause on the vendor’s side, hours before the company publicly announced the CrowdStrike incident affected it. This advanced warning allowed the life insurance provider to execute a contingency plan and course of action much earlier than if they had to wait for the third-party provider to notify them of the problem.

Synthetic monitoring view of the CrowdStrike outage

Synthetic monitoring view of the CrowdStrike outage
Synthetic monitoring views of the CrowdStrike outage

3. Dynatrace Query Language (DQL)

Dynatrace Query Language (DQL) is a structured syntax for exploring, querying, and processing observability data in Dynatrace. It allows users to chain commands together to filter, manipulate, and analyze data efficiently.

As a data query tool, DQL provides flexibility and customization, allowing organizations to tailor their investigations to meet specific needs, addressing unique challenges and optimizing performance.

Using DQL, investigators can find specific answers so they can quickly identify and respond to issues, which is crucial during crises like the CrowdStrike incident.

Dynatrace Notebooks is an interactive capability that enables users across the organization to collaborate using code, text, and rich media to build, evaluate, and share insights for exploratory analytics. This ability to track and collaborate on issue details is a crucial capability in a crisis.

How DQL and Notebooks helped companies pinpoint affected systems and prioritize remediation

During the CrowdStrike crisis, a North American telecommunications provider used DQL and a notebook to create custom charts filtered by application. This helped the company prioritize remediating its most critical servers first, restoring essential services promptly.

DQL and Notebooks investigation showing systems affected by the CrowdStrike outage
DQL and Notebooks investigation showing systems affected by the CrowdStrike outage

When a major US airline began experiencing the CrowdStrike outage, its IT team also used DQL and Dynatrace Notebooks to identify which systems were no longer forwarding logs back to Dynatrace. This newfound visibility enabled the kiosk management team to focus their recovery efforts effectively.

To see an example of how to use DQL to find when BSOD issues are being written to Windows system logs, see the blog Crowdstrike BSOD: Quickly find machines impacted by the CrowdStrike issue by Dynatrace Principal Solutions Engineer Josh Wood, Ph.D.

4. Real user monitoring to understand business impact

Real User Monitoring (RUM) offers comprehensive insights into user experiences across web, mobile, and custom applications. By capturing user sessions, RUM provides a detailed view of user journeys, helping businesses understand critical actions for conversions.

When an incident occurs, Dynatrace automatically generates a problem notification. Davis AI analyzes details from the front end to the backend to identify the root cause, severity, and impact of the issue.

Dynatrace RUM also tracks application downtime, enabling organizations to calculate the cost of business interruptions and account for lost revenue. RUM offers flexible options for tracking and reporting key information, such as conversion goals. Examples include successful checkouts, newsletter signups, or demo requests. By monitoring conversion rates, businesses can estimate expected revenue to better understand the financial impact of an incident and help organizations account for lost revenue.

How Dynatrace RUM helped one company identify the business impacts of the CrowdStrike outage

For example, during the CrowdStrike crisis, a North American mortgage provider received an alert for unexpected low traffic. The problem card helped them identify the affected application and actions, as well as the expected traffic during that period. This allowed them to prioritize remediation efforts on their most critical services.

Dynatrace RUM shows the user impact of the CrowdStrike outage
Dynatrace RUM shows the user impact of the CrowdStrike outage

5. Using SLOs to verify recovery

Service level objectives (SLOs) are essential for maintaining and enhancing the performance of applications and services, especially during and after a crisis. Dynatrace makes it easy to create, capture, and visualize SLOs in real time. Establishing and monitoring SLOs can play an instrumental role before, during, and after a crisis.

  • Before a crisis. Setting up SLOs for mission-critical services helps establish and maintain standards for availability and performance. Dynatrace AI continuously monitors these benchmarks, allowing teams to identify and address potential issues proactively.
  • During a crisis. SLOs provide real-time monitoring and immediate feedback on service performance. Dynatrace AI can quickly pinpoint the root cause of issues, enabling swift resolution and minimizing user impact.
  • After a crisis. SLOs ensure that application performance returns to the same standard of performance as before the incident. They play a crucial role in post-incident analysis, helping teams understand the business impact of the incident and implement improvements to prevent future occurrences.

By implementing Dynatrace SLOs, organizations can ensure robust performance management before, during, and after a crisis, leading to more resilient and reliable services.

Prepare for any crisis with observability and the right capabilities

The recent CrowdStrike crisis has highlighted the critical need for robust monitoring and observability tools. For organizations navigating disruptions, observability is crucial for rapid detection and remediation. Dynatrace offers comprehensive solutions with real-time data, synthetic monitoring, DQL querying capabilities, real user monitoring, and SLOs. These tools empower organizations to maintain stability, swiftly identify and resolve issues, and ensure the continuity of essential services.

The real-world examples of our customers demonstrate how Dynatrace monitoring solutions have enabled organizations to recover quickly and maintain operational stability, ultimately safeguarding their customer experiences and bottom lines. As we continue to lean heavily on technology for day-to-day operations, observability will be essential for navigating the complexities of an increasingly digital world.

Using Dynatrace, organizations can not only react fast to mitigate the immediate impacts of crises but also be proactive and build resilient IT infrastructure that’s prepared for future challenges.

Contact us to learn how you can gain the same situational awareness and responsiveness that enabled these customers to recover so quickly from the CrowdStrike outage.

To learn more about the recent CrowdStrike update outage and explore more resources to help you maintain business resilience, check out the resource center, Business Resilience through CrowdStrike and Beyond.

The post CrowdStrike update: How Dynatrace helped customers recover in hours appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/crowdstrike-update-crisis-dynatrace-customers-recovery/feed/ 0
10 tips for migrating from monolith to microservices https://www.dynatrace.com/news/blog/10-tips-for-migrating-from-monolith-to-microservices/ https://www.dynatrace.com/news/blog/10-tips-for-migrating-from-monolith-to-microservices/#respond Tue, 03 Oct 2023 00:09:17 +0000 https://www.dynatrace.com/news/?p=59857 Dynatrace named to Constellation Research's annual ShortList for top vendors

Transforming an application from monolith to microservices-based architecture can be daunting, and knowing where to start can be difficult. Because monolithic applications combine database, client-side interfaces, and server-side application elements in a single executable, they’re difficult to understand, even for their own administrators. Today’s customer expectations can’t tolerate tightly coupled dependencies, difficulties deploying changes, or […]

The post 10 tips for migrating from monolith to microservices appeared first on Dynatrace news.

]]>
Dynatrace named to Constellation Research's annual ShortList for top vendors

Transforming an application from monolith to microservices-based architecture can be daunting, and knowing where to start can be difficult. Because monolithic applications combine database, client-side interfaces, and server-side application elements in a single executable, they’re difficult to understand, even for their own administrators.

Today’s customer expectations can’t tolerate tightly coupled dependencies, difficulties deploying changes, or long release cycles. Unsurprisingly, organizations are breaking away from monolithic architectures and moving toward event-driven microservices. However, the move to microservices comes with its own challenges and complexities.

Limits of a lift-and-shift approach

A traditional lift-and-shift approach, where teams migrate a monolithic application directly onto hardware hosted in the cloud, may seem like the logical first step toward application transformation. However, it’s common that teams need to refactor an application to help with performance after the lift-and-shift operation. It is better to consider refactoring as part of the application transformation process before migrating, if possible.

forklift representing the limits of a lift-and-shift approach to migrating from monolith to microservices

Although lifting and shifting to the cloud may provide some cost advantages at the outset, applications refactored for a microservices architecture can take full advantage of a cloud-first approach and provide cost savings in the long run. Microservices applications comprise independent services that teams develop, deploy, and maintain separately. Because they’re separate, they allow for faster release cycles, greater scalability, and the flexibility to test new methodologies and technologies.

However, the distributed system of a microservices architecture comes with its own cost: increased application complexity and convoluted testing. Migration is time-consuming and involved. Likewise, refactoring and rewriting code takes a lot of time and effort. Therefore, it’s important to do it right.

In the past, we’ve covered various topics on how to break the monolith, define monolithic architecture, and identify the advantages and disadvantages of microservices. Since the ultimate goal is to migrate to the cloud and refactor applications to a microservices architecture, here are 10 tips for getting started on the path to a successful migration.

10 tips for migrating from monolith to microservices

1. Understand the monolith

Monolithic applications take work to understand. In fact, it can be difficult to make code changes that won’t disrupt the entire system. This is also true of splitting the application into microservices. It’s important not to disrupt a running application and user experience. Start by evaluating the monolithic application components to understand all the dependencies and their business functions. Teams can gain this understanding through topology mapping, with telemetry data from request traces, and understanding how the frontend ties to backend functions. While this sounds simple, teams often get stuck trying to map out dependencies accurately when they try to do it manually. Automatic discovery and intelligent observability are the keys to overcoming this hurdle.

2. Find the right candidates for refactoring

There will inevitably be parts of a monolith application that teams can’t fully understand. When it comes to refactoring, teams should start with what they can understand. Components that are already loosely coupled and have few dependencies will be easier to migrate, with less chance of impacting application performance.

Use domain-driven design when creating new microservices by separating microservices via their underlying business functions. This allows individual teams to own components and makes it easier to pinpoint problems with mission-critical functions. Start with components that have high business value, such as a webpage checkout action.

3. Incrementally refactor

The saying “Rome wasn’t built in a day” rings true when it comes to refactoring microservices. It is important to remember that refactoring is essentially re-architecting and rebuilding your application. This is something that will take time. The best approach is incremental, using the Strangler Fig pattern: Gradually replacing parts of the monolithic application until only the microservices architecture remains.

strangler fig model

The approach takes place in three stages: 1. create a microservice; 2. use both the microservice and monolith for the same functionality; 3. remove the dependency on the monolith after all testing is successful.

4. Ensure the microservices architecture is loosely coupled

Monolithic applications are traditionally tightly coupled, meaning dependencies among services within the app are intertwined. Because teams often can’t know or understand all dependencies, it can be difficult for them to make changes. When creating new microservices, it is better to keep them loosely coupled with minimal dependencies to allow for flexible changes and ease of deployment. One way to minimize dependencies is to use asynchronous messaging and message queues wherever possible.

5. Choose the right technology for each service

One advantage of using a microservices architecture is being able to choose different technologies for the application function at hand. Microservices can operate independently of each other. As a result, teams can leverage a polyglot architecture that uses the language and technology best suited for the job. However, there can be drawbacks to using too many different languages and technologies. It may make sense to use fewer languages depending on team bandwidth, capabilities, and size.

6. Instrument for end-to-end observability

One challenge that comes with migrating to microservices from a monolithic architecture is an increase in application complexity. Many microservices and different supporting technologies run independently of each other to support the application. With so many different components, it is easy to lose track of what is happening in which component when potential problems arise. End-to-end observability starts with tracking logs, metrics, and traces of all the components, providing a better understanding of service relationships and application dependencies.

end-to-end observability

This visibility is essential to understanding if the application is performing well and identifying where problems arise. Visibility is also the key to remediating problems and implementing automation.

There are many ways to instrument microservices for observability, including automatic instrumentation using a unified observability and security platform. Many organizations also find it useful to use an open source observability tool, such as OpenTelemetry. An observability platform approach makes it easy to combine both methods. Such an approach also considers additional important details, such as business impact and security implications.

7. Factor in security at every stage

With the increasing number and sophistication of security breaches, such as Log4Shell, integrating security into all application changes is more important than ever. This includes when teams refactor applications for microservices architecture. In fact, security risks differ among microservices. Security should be an integral part of each stage of the software delivery lifecycle, from development to monitoring in real time. Real-time security monitoring—evaluating continuously updated security data about systems, processes, and events—and runtime security monitoring—analyzing security information from a running system—can be particularly useful during migration since there will be an immediate alert if there is a security risk with this new microservices architecture.

8. Monitor the application before, during, and after migration

Migrating and changing code can be a tricky business. To ensure that the migration doesn’t affect user experience, teams should monitor application performance before, during, and after migration. Use SLAs, SLOs, and SLIs as performance benchmarks for newly migrated microservices. Repeat this process throughout the different environments before development, staging, release, and production. As each service migration is complete, continuously validate the existing code base as the team releases new code. Intelligent dependency mapping and automated baselining can help easily identify performance degradation and problems caused by new releases.

9. Automate wherever possible

Creating a system and a flow that teams can replicate and automate is a desirable objective. An automated CI/CD pipeline allows for an extremely smooth release process, accomplishing one of the goals of migrating to microservices. Every step of the way, define checks based on SLAs, SLOs, SLIs, and security scans, and automate the transitions from continuous integration, delivery, and deployment. teams can even build auto-remediation into a CI/CD pipeline, so if a problem arises, the system can trigger a fix or roll back to a previous version.

10. Optimize performance and user experience using observability data

Once the application has successfully migrated, it’s time to take advantage of all the newfound benefits a microservices architecture offers. Deploy changes fast to optimize performance and eliminate bottlenecks. Continuously update, improve, and easily add new features to provide an exceptional user experience. Utilize observability data to monitor and improve digital experiences and analyze data that can affect the business.

A unified observability and security platform approach to migrating from monolith to microservices

Migrating from monolith to microservices can be difficult, complex, and time consuming. However, making this architecture change is often the best way to take advantage of the agility of the cloud. For a simplified and smooth migration, teams need a way to assess a monolithic application’s starting state. They then need to track progress during migration to identify any drops in application performance.

As an AI-driven, unified observability and security platform, Dynatrace uses topology and dependency mapping and artificial intelligence to automatically identify all entities and their dependencies. This comprehensive view helps teams gain an initial understanding of a monolithic application so they can develop a migration strategy.

Once teams start introducing microservices, unified observability enables teams to visualize changes and the real-time performance of all the services in context. This AI platform approach automatically discovers all entities and identifies any issues developing among them. The observability extends to on-premises environments, Kubernetes infrastructure, multicloud platforms, and the multitude of proprietary and open source tools they depend on.

Keeping track of the migration stages, phases, and environments is not always easy. With real-time observability, teams can easily plan their migration and fine-tune performance as they migrate microservices. Since core observability with Dynatrace includes logs, traces, metrics, security, and user experience, teams can make decisions using these details in context.

After migration, observability continues to help teams monitor and optimize their microservices, resulting in better user experiences and, as an extension, better business outcomes.

To learn more about how migrating from monolith to microservices with real-time observability works in practice, join us for the on-demand webinar, 10 things you didn’t know about cloud migration and adoption.

The post 10 tips for migrating from monolith to microservices appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/10-tips-for-migrating-from-monolith-to-microservices/feed/ 0