log management and analytics | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Tue, 07 Jul 2026 12:05:27 +0000 en hourly 1 Log management for AI workloads: How to bring your logs and telemetry plan into the AI-first century https://www.dynatrace.com/news/blog/2026-log-management-for-ai-workloads-action-plan/ https://www.dynatrace.com/news/blog/2026-log-management-for-ai-workloads-action-plan/#respond Wed, 24 Jun 2026 12:19:44 +0000 https://www.dynatrace.com/news/?p=74644 2026 Logs Report action plan blog

AI is stretching the boundaries of traditional log management. More data without context slows insight, increases risk, and stalls AI progress. Teams need to rethink how they capture, process, and use telemetry from ingest to analysis. This action plan outlines how to unify telemetry, optimize pipelines, and turn data into real‑time, trusted intelligence teams can […]

The post Log management for AI workloads: How to bring your logs and telemetry plan into the AI-first century appeared first on Dynatrace news.

]]>
2026 Logs Report action plan blog

AI is stretching the boundaries of traditional log management. More data without context slows insight, increases risk, and stalls AI progress. Teams need to rethink how they capture, process, and use telemetry from ingest to analysis. This action plan outlines how to unify telemetry, optimize pipelines, and turn data into real‑time, trusted intelligence teams can use to scale AI operations with confidence.

AI workloads aren’t just increasing telemetry—they’re exposing the limits of how teams capture, store, and use it. Traditional log management was built for predictable systems and finite telemetry. AI systems break those assumptions. The result: more telemetry, less context, and increasing cost pressures to stay operational. The shift is how to manage logs better—and changing how teams capture, process, and make logs telemetry available across the entire lifecycle.

Here’s how that shift looks in practice:

Traditional log management AI-ready log management
Indexing and re-indexing, schema-first, archiving and rehydrating  Schema-on-read, always hydrated, always queryable 
Tool-specific context  Unified telemetry context 
Reactive troubleshooting  Preventive operations 

The State of Log Management 2026 research report—based on a global survey of 450 senior IT leaders—examines how AI is reshaping log economics, instrumentation, and observability strategies, and what shifts technical leaders must make to telemetry capture, storage, and management to support and scale agentic AI projects.

5 actions to kickstart your new log management plan

  • Create a single source of truth for AI systems by centralizing all telemetry in a unified, continuously queryable context layer and platform.
  • Establish causation across AI systems by automatically unifying logs with metrics, traces, and lifecycle context—not relying on logs alone.
  • Control costs without losing visibility by optimizing telemetry before ingest and eliminating indexing, archiving, and rehydration dependencies.
  • Standardize and govern telemetry at ingest to ensure data quality, compliance, and real-time usability at AI scale.
  • Enable preventive AI operations by turning contextual telemetry into real-time insight and automated remediation.

Why should teams unify telemetry on a single observability platform for AI workloads?

Unifying telemetry reduces manual correlation, preserves context, and keeps logs, metrics, and traces continuously queryable as telemetry scale increases.

AI workloads are exacerbating an existing problem by fragmenting even more telemetry across tools just as systems require more context. Teams now use an average of seven log tools, forcing manual correlation that doesn’t scale.

  • Unify all telemetry—logs, metrics, traces, security signals, user behavior, business events—into a single, continuously queryable context layer where telemetry is correlated automatically.
  • Enrich telemetry at ingest starting at the edge with shared technical and business context to explain system behavior as dependencies multiply.
  • Democratize access using intuitive querying so more teams can validate AI behavior and act faster with confidence.

How do logs and traces work together for reliable and explainable autonomous operations?

Logs don’t explain AI behavior independently. Understanding comes from unifying logs with traces and other telemetry signals optimized throughout the telemetry lifecycle.

Autonomous systems demand deterministic signals that explain what happened, why it happened, and how to respond—something logs alone can’t fully provide.

  • Instrument logs to capture AI‑specific details at every inference layer to preserve the exact sequence of events.
  • Correlate logs with traces automatically to establish causation and pinpoint root causes.
  • Automate remediation using continuously enriched telemetry to enable reliable, explainable autonomous operations at scale.

How can teams optimize log management costs without sacrificing insight?

Teams can manage costs by retaining high‑value telemetry without rigid schemas, indexing overhead, or rehydration delays that limit analysis.

Managing log costs involves data strategy, not just storage. Logs consume nearly half of observability budgets, yet even after reducing volume by filtering, masking, and aggregating, 50% of organizations don’t collect or discard an average of 86% of logs specifically to manage costs, and 74% say indexing and rehydration costs are barriers to value.

  • Ingest and retain telemetry without rigid schemas or indexes, eliminating the need to predict questions in advance.
  • Store exabytes of data in one queryable layer, avoiding cold archives and rehydration costs and delays.
  • Analyze telemetry in full context to reduce waste and maximize business value from AI‑generated data.

What changes to instrumentation and ingest should teams make to support AI workloads?

Teams must optimize telemetry before ingest—standardizing instrumentation and automating parsing and configurations—so data remains high-quality, contextual, and continuously queryable at AI scale.

Fragmented instrumentation and brittle ingest pipelines slow insight and delay AI projects from reaching production. 85% of organizations struggle to ingest logs at AI scale, and 80% say turning telemetry into insight delays AI initiatives.

  • Standardize instrumentation across logs, traces, metrics, and other telemetry signals to maintain context and reduce downstream correlation.
  • Streamline ingestion in real time by automating parsing, configurations, and enrichment to retain only high‑value, compliant data.
  • Sustain telemetry at scale with an always‑queryable data layer that supports real‑time analytics and automation.

Why are preventive operations critical to AI-native environments?

As AI workloads increase telemetry volume and autonomous operations, teams need detailed intelligence about what’s happening in AI output to predictively detect early signals of unexpected results.

Because reactive troubleshooting can’t keep up with autonomous systems, teams need real‑time, contextual telemetry to detect drift and prevent failures early. 84% say customer trust in AI depends on their ability to use log analytics to predict and prevent problems.

  • Correlate logs with end‑to‑end traces automatically to create a reliable understanding of AI behavior before failures escalate.
  • Analyze AI telemetry in real time and full context to detect early signs of drift or degradation.
  • Automate response to reduce risk and scale AI‑driven operations safely.

Upleveling log management to advance trustworthy agentic AI

Expanding AI workloads demand more from log management—an approach built on unified observability, open, optimized ingest at massive scale, and real‑time analytics without rigid schemas, indexing overhead, or rehydration delays. Logs remain the accountability anchor, but trust emerges only when all telemetry signals come together in context.

Get the State of Log Management 2026 report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI operations.

FAQ: Log management action plan for AI workloads

Why do AI workloads require a different log management approach?

AI workloads are variable and generate significantly more telemetry, which demands explainability, reliability, and cost control capabilities that traditional log architectures weren’t designed to support.

What is the first step teams should take to modernize log management for AI?

Unify logs, metrics, and traces on a single observability platform so telemetry is always available in context and doesn’t require manual correlation.

How can organizations reduce log management costs without losing insight?

By starting observability at the edge and retaining high‑value telemetry without rigid schemas, indexes, or rehydration delays, teams avoid discarding data while controlling cost.

Why aren’t logs alone enough to support autonomous operations?

Logs show what happened and why, but traces show how and what’s affected; together they explain AI behavior and enable reliable, automated remediation.

What enables preventive operations in AI‑native environments?

Real‑time analysis of contextual telemetry that detects early signs of drift or degradation and triggers automated guardrails before failures escalate.

The post Log management for AI workloads: How to bring your logs and telemetry plan into the AI-first century appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/2026-log-management-for-ai-workloads-action-plan/feed/ 0
How AI workloads are changing what logs must deliver, forcing a new strategy https://www.dynatrace.com/news/blog/log-management-research-key-findings/ https://www.dynatrace.com/news/blog/log-management-research-key-findings/#respond Wed, 17 Jun 2026 12:34:03 +0000 https://www.dynatrace.com/news/?p=73713 Future of log management: 2026 research shows AI demands of logs

AI workloads are redefining what logs must deliver—and exposing where traditional approaches fall short. New research reveals how rising scale, cost pressures, and missing context are reshaping log management. The path forward is clear: unify logs with traces and other telemetry in context to turn fragmented signals into reliable insight and build the foundation for trusted, scalable AI operations.

The post How AI workloads are changing what logs must deliver, forcing a new strategy appeared first on Dynatrace news.

]]>
Future of log management: 2026 research shows AI demands of logs

New Dynatrace research reveals how logs are becoming the accountability anchor for AI systems and why cost-driven trade-offs (sampling, cold storage, and discarding data) increase operational risk. In fact, with legacy log management tools now consuming 45% of observability budgets, 67% say the costs of these tools outweigh their value, indicating a need to modernize quickly (source: Dynatrace State of Log Management 2026 research).

3 takeaways for engineering and business leaders

  • Scale is accelerating. Log and telemetry volume increased 93% on average with 1 in 5 organizations seeing growth above 150%.
  • Costs are breaking traditional models. Existing log management tools consume 45% of observability budgets with respondents estimating an average annual spend near $2.5M.
  • Context with traces is the path to AI trust. 73% say logs reveal only part of what’s happening with AI workloads, and 70% rank traces as a top source for evaluating AI performance and behavior. But 65% rely heavily on logs because they don’t have easy visibility into other telemetry signals.

Logs capture the precise details of events from cloud, AI, and infrastructure. As AI workloads and agentic systems (AI systems that act autonomously) make up an increasing proportion of technology stacks, logs serve as a crucial shared language for humans and agents to reason and troubleshoot system state and autonomously build and optimize infrastructure and applications.

Accountability depends on placing high-fidelity log telemetry in context: connecting it with traces, metrics, security events, user behavior, and business signals to understand AI behavior and guide remediation at scale.

How is AI breaking the economics of traditional log management?

AI workloads break the financials of traditional logging approaches by driving massive telemetry growth that forces many teams to not collect or discard data to avoid runaway costs.

Over the past year, AI workloads triggered a 93% average increase in log and telemetry volume, with one in five experiencing expansions above 150%. At the same time, teams rely on an average of seven different log and telemetry tools, forcing manual correlation that doesn’t scale.

The financial impact is just as stark. According to the report, existing log management tools now consume 45% of observability budgets, with average annual spend among survey respondents nearing an estimated $2.5M per organization. To contain costs, many teams limit ingestion, sample telemetry, or push logs into cold storage, losing crucial context and increasing security, compliance, and operational risk. In fact, 67% say the cost of existing log management tools now outweighs their value. And even after filtering, masking, and aggregating to reduce volume, the limitations of traditional logging tools force teams into imprecise compromises, resulting in half of organizations not collecting or discarding 86% of logs specifically to manage costs.

Traditional vs. AI-native log management

What changes

Traditional log management
(cost-first)

AI-native log management
(unified observability)

How teams find answers  Manual stitching slows analysis; teams spend 58% of analysis time correlating telemetry  Automated ingestion, enrichment, and correlation reduces manual work and accelerates time to answers 
AI trust and validation  Logs alone are incomplete; 73% say logs reveal only part of what’s happening in AI workloads  Context builds trust: logs + traces (top-ranked by 70%) + metrics + events show behavior and causality 
Readiness for what’s next  79% worry current ingest/storage won’t meet future needs; instrumentation lags AI requirements  Updated instrumentation starting at the edge and open, automated processing at scale (supported by 81%) for accelerated AI innovation 
Data strategy  Many teams control spend by reducing ingest using sampling, cold storage, or discarding data  Retain high-fidelity telemetry at scale without rehydration/indexing friction 
Tooling approach  Fragmented tooling; teams use an average of seven tools, forcing manual correlation  One real-time observability context layer that unifies logs with metrics, traces, security events, and business signals 
Operational impact  Blind spots and risks grow as half of teams don’t collect or discard 86% of logs  Answers, not guesses; more context preserved for faster diagnosis and safer automation 

What must change in instrumentation and ingestion for autonomous systems?

To build trust in AI workloads and advance the business value of autonomous decision-making, optimizing and streamlining telemetry must start before ingest and be open and automated at a massive scale.

Traditional logging architectures weren’t built for AI‑driven scale or autonomy.

  • As AI workloads proliferate, 79% of technical leaders worry their current ingest and storage approaches won’t meet future needs.
  • 80% say they must update instrumentation to support new AI‑specific metrics.

The operational toll is significant. Teams spend 58% of their analysis time stitching together logs, metrics, and traces before extracting insight, which slows decisions and delays AI projects from moving into production.

To break this bottleneck, 81% of organizations say log ingestion and processing must be open and automated at massive scale, enabling real‑time analysis without rigid schemas, indexing overhead, or rehydration delays.

Why do logs need an observability ecosystem to build AI trust?

While logs are a crucial component of AI observability, they need context to tell the whole story.

  • 73% of organizations say logs reveal only part of what’s happening in AI workloads
  • Teams spend 58% of analysis time stitching telemetry together
  • 72% say standalone log management tools are obsolete—AI workloads demand a platform approach that combines all types of telemetry in one place to accelerate time to answers

As a result, most organizations rely on logs alongside other telemetry signals—especially traces, which 70% rank as the top source for evaluating AI performance and behavior.

Trust in AI develops when analysis is based on high-fidelity telemetry in context:

  • Logs provide the fact basis
  • Traces expose flow and causality
  • Metrics quantify performance
  • Security events surface risk
  • Business signals connect system behavior to outcomes

Unified observability turns this telemetry into a coherent narrative—enabling teams to validate AI behavior, assess impact radius, guide remediation, and move from reactive troubleshooting to preventive operations.

What does AI-native log management look like with unified observability?

Unified observability transforms log management for the AI era by optimizing telemetry instrumentation and ingest, unifying telemetry with context, and eliminating crippling cost-cutting measures.

With exponential increase in telemetry volumes due to AI workloads, the path forward can’t just be shrinking log ingest to manage costs. AI innovation depends on leaning into telemetry volume with the right strategy and capabilities. The report points to key actions leaders can take to adopt AI-native log management practices.

  • Centralize all telemetry (logs, traces, metrics, events) in a unified, continuously queryable context layer to eliminate silos and scale AI visibility.
  • Automatically correlate logs with traces and lifecycle context to establish causation, understand AI behavior end to end, and enable reliable autonomous operations.
  • Control log costs before ingest while retaining full-fidelity telemetry in exabyte-capacity storage with no rigid schemas, indexes, cold archives, or rehydration.
  • Standardize instrumentation and optimize ingestion to capture high-value, governed telemetry that tracks agent actions, reasoning traces, and lifecycle events.
  • Enable preventive operations by detecting early signals and automating remediation to reduce risk, strengthen reliability, and safely scale AI projects.

Fragmented vs Unified log management

The goal is reliable operations and AI accountability at scale. Together, logs, traces, metrics, and events in context can enable preventive operations, reliable autonomy, and confident decision‑making as AI systems move from pilots into production. Organizations that build this unified foundation will be best positioned to scale AI without sacrificing trust.

Download the State of Log Management 2026 report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

FAQ: State of Log management 2026

Why is log management getting harder in the AI era?

Because AI workloads are driving rapid telemetry growth—organizations saw a 93% average increase in log and telemetry volume in the past year.

Why do log management costs feel out of control?

Existing log management tools consume 45% of observability budgets, and average annual spend is nearing $2.5M per organization.

Do teams still see value in log management at today’s price?

Not consistently. 67% of respondents say costs of existing log management tools now outweigh their value.

Why do teams discard so many logs?

Cost pressure. Using existing tools, 50% of organizations don’t collect or discard 86% of their logs on average, often using sampling or limiting ingestion specifically to reduce spend.

What’s the biggest operational bottleneck with today’s tooling?

Manual correlation. Teams spend 58% of analysis time stitching together logs, metrics, and traces before they can extract insight.

Why aren’t standalone log tools enough for AI workloads?

Because logs alone rarely tell the whole story. 72% say standalone log management tools are obsolete, and 73% say logs reveal only part of what’s happening in AI workloads.

What telemetry signal do teams rely on most to evaluate AI behavior?

Traces—70% rank them as the top source for evaluating AI performance and behavior beyond logs alone.

The post How AI workloads are changing what logs must deliver, forcing a new strategy appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/log-management-research-key-findings/feed/ 0
How to mask PII like email addresses appearing in logs with Dynatrace: An advanced use case https://www.dynatrace.com/news/blog/how-to-mask-pii-like-email-addresses-appearing-in-logs-with-dynatrace-an-advanced-use-case/ https://www.dynatrace.com/news/blog/how-to-mask-pii-like-email-addresses-appearing-in-logs-with-dynatrace-an-advanced-use-case/#respond Wed, 21 Jan 2026 22:50:02 +0000 https://www.dynatrace.com/news/?p=72599 OpenTelemetry logs

Protecting Personally Identifiable Information (PII) data in logs is crucial for supporting privacy and compliance. While Dynatrace OneAgent can mask data at capture/source, sensitive information like social security numbers (SSNs), payment details, IP addresses, or email addresses might still get through — especially when logs are ingested via other mechanisms such as directly using an […]

The post How to mask PII like email addresses appearing in logs with Dynatrace: An advanced use case appeared first on Dynatrace news.

]]>
OpenTelemetry logs

Protecting Personally Identifiable Information (PII) data in logs is crucial for supporting privacy and compliance. While Dynatrace OneAgent can mask data at capture/source, sensitive information like social security numbers (SSNs), payment details, IP addresses, or email addresses might still get through — especially when logs are ingested via other mechanisms such as directly using an API, via Fluent Bit, Cribl, or using the OpenTelemetry collector.

In this blog, practitioners will walk away with an in-depth understanding of how to create your own parsing rules. We’ll focus on masking/obfuscating email addresses in log events during the ingest process, leveraging Dynatrace OpenPipeline before logs are retained. Outside of this blog’s scope is the attribute-based access capability and mask on read.

Best Practice: Before we start masking data at the time of ingest, we’ll validate our configuration within a Dynatrace Notebook to prevent unwanted data loss caused by inaccurate patterns. We’ll walk through:

  • Identifying patterns
  • Creating a masking pipeline in OpenPipeline
  • Validating the results
If you’ve already registered for free access to our Playground tenant, you can follow the steps and demonstrations there. Alternatively, you can also analyze your own logs within your own tenant. At the bottom of this blog in the addendum section, you can find more details and demo data enabling you to follow every individual step, including the OpenPipeline configuration.

Step 1: Simulate logs with sensitive data

To demonstrate, we’ll work with logs containing email addresses and IP addresses. These logs use example domains and IPs (e.g., example.com and 203.0.113.x) for documentation purposes.

Here are the eight sample logs sent from various demo applications to the Dynatrace Playground:

  • 2025-04-22 11:46:38 [ERROR|[203.0.113.13]
    |K9p8Q3xJwZrStUv4XyZ7AbC2n5M6h8T1v0L2r4 
    |com.example.exmplstore.security.examplePersistentTokenRepository] 
    Can't find credentials for series marie_curie@example.com
  • 2025-04-22 13:50:43 [INFO 
    |[203.0.113.112] |A1b2C3d4EfGhIjK5LmN6oPq7RsT8uVw9XyZ0 
    |class com.example.exmplfacades.order.exampleCheckoutLogger Checkout ABC] 
    Receive API Request:method=placePayOrder, cartCode=302132310, 
    customerID=freddie@example.com
    |com.example.exmplstore.checkout.request.PayPlaceOrderRequestDto@3d3a53d5|]
  • 2025-04-22 13:58:24 [INFO |||com.example.exmpl.BusinessProcessLoggingAspect] 
    Finish Action:[ PerformSubscriptionAction ], BusinessProcessCode: 
    [ customerRegistrationProcess-martin_luther@example.com-1745294290737],
    OrderCode:[ n/a ]
  • 2025-04-29 10:16:13 [INFO ||
    |com.example.exmplmarketing.action.exampleSendCustomerNotificationAction] 
    Successfully sent email forgottenPassword message for process 
    forgottenPasswordProcess-jane_austen@example.com-1745885767984
  • 2025-04-29 10:24:26 [INFO |[203.0.113.14] 
    |M3n4P5q6RsTuVwXyZ7aBc8DeF9gHi0JkL2|com.example.exmpl.UserDeleteInterceptor] 
    Deleting userId: cleopatra@example.com, actioned by userId: anonymous
  • 2025-04-29 10:24:36 [INFO |[203.0.113.19] 
    |T2u3V4w5XyZaBc6DeF7gHi8JkL9mNo0PqR1|com.example.exmpl.AdvantageApiClient] 
    Loyalty: Check loyalty account exist for email: leonardo_davinci@example.com , 
    wodCorrelationId 2fee137d-915e-46fb-b390-c69aae3f150
  • 2025-04-29 10:22:47 [INFO ||
    |com.example.exmplbusproc.aop.BusinessProcessLoggingAspect] 
    Begin Action: [ exampleAdvantageLinkEmailAction ], BusinessProcessCode: 
    [ advantageLinkEmailProcess-albert_einstein@example.com-1745886166709], 
    OrderCode:[ n/a ]
  • 10.96.152.11 - - [29/Apr/2025:10:33:43 +1000] 
    "GET /reminder/shakespeare@example.com/1 HTTP/1.1" 200 90 
    "https://www.example.com/my-account/order-details/306399901" 
    "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 
    (KHTML, like Gecko) Chrome/135.0.0.0 Safari/537.36" 203.0.113.145 - 
    T2u3V4w5XyZaBc6DeF7gHi8JkL9mNo0PqR1 100 

As you can see, all of these logs contain email addresses, which must be masked before storing in Dynatrace Grail™.

Step 2: Identify patterns in logs

We can use the Notebooks app to query stored logs and search for log lines that contain email addresses. To do so, we create a new notebook, add a DQL section, and execute the following command:

Adding a DQL section in Dynatrace Notebooks app
Figure 1- Adding a DQL section by selecting “+ New section” and “DQL” or using the shortcut “Shift + D”
fetch logs
| filter matchesPattern(content, "LD '@' [A-Za-z0-9'_']* '.'
[A-Za-z]* (LD | EOS)") 

The second part of this command executes a pattern matching command, which makes use of the Dynatrace Pattern Language (DPL) to define the pattern and reduces the result to show only those entries that contain an email address.

Resulting set of log entries identified containing email addresses
Figure 2 – Resulting set of log entries identified containing email addresses.

Let’s have a closer look at the DPL pattern itself:

LD '@' [A-Za-z0-9'-']* '.' [A-Za-z]* LD EOS 

and break down the elements:

  • LD: Matches any characters (letters, digits, or others) until the next non-optional matcher (in this case, ‘@’) within the scope of a line.
  • '@': Matches the actual @ symbol.
  • [A-Za-z0-9'-']*: Matches zero or more letters, digits, or hyphens, forming the domain part of the email address.
  • '.': Matches the literal “.” (e.g. in .com).
  • [A-Za-z]*: Matches zero or more letters, representing the top-level domain. Note: If you wish to match double-dotted TLD’s like .co.uk use this pattern instead: [A-Za-z’.’]*
  • (LD | EOS): The pattern matches either a sequence of characters (letters, digits, or others) via LD to the end of the line after the email address. However, if the last word in the log line is the email address, then optionally the EOS will match the end of string.

Step 3: Analyze logs for email address patterns

After confirming email addresses exist, we examine the log lines to identify patterns preceding them. Here are the literals just before the email addresses, found in our returned sample logs:

  • order confirmation email sent to
  • find credentials for series
  • customerID=
  • customerRegistrationProcess-
  • forgottenPasswordProcess-
  • Deleting userId: 
  • Check loyalty account exist for email: 
  • advantageLinkEmailProcess-
  • GET /reminder/ 

These literals will help us target the email addresses for masking.

Step 4: Check the pattern to test the process of masking the email addresses

We will now parse the email addresses and mask the local part (before the @ symbol) using Dynatrace Query Language (DQL). This is parsing at query time to test whether the pattern is valid or not. It does not change or update the unmasked data that is already stored in Grail. This step ensures that the pattern we’ll be using in the following steps with OpenPipeline are valid. To achieve this, we must open the DPL architect by selecting the 3 dots on the right hand of a row-element in the content column and select “Extract fields.”

Extract fields option in Dynatrace Notebooks app
Figure 3 – Extract fields option is selected from the row item menu.

Parse the email field

We will use the following parsing pattern with DQL to extract email addresses based on the literals identified:

Pattern:

 
LD  
( 
'order confirmation email sent to' |  
'find credentials for series' | 
'customerID=' |  
'customerRegistrationProcess-' |  
'forgottenPasswordProcess-' |  
'Deleting userId: ' |  
'Check loyalty account exist for email: ' | 'advantageLinkEmailProcess-' |  
'GET /reminder/') LD:email 
 

Breakdown: 

  • LD 
    Matches any non-whitespace characters before the literal.
  • ('Can't find credentials for series' | ...) 
    Matches any of the specified literals.
  • LD:email 
    Captures the email address into a field named email.

Validate the parsing rule using DPL architect in Dynatrace Notebooks by pasting the parse command (starting with LD):

DPL Architect with pattern detection sample in Dynatrace Notebooks app
Figure 4 – DPL Architect with pattern detection sample.

Make sure you select the ‘Add to preview’ action button to ensure there are no “Unmatched records” and all the matched records based on the previous filter can be truly passed:

Unmatched records view in Dynatrace Notebooks app
Figure 5 – Unmatched records view is empty, ensuring that no log items are missed by the parsing rule.
Matched records in Dynatrace Notebooks app
Figure 6 – Matched records view validates all records are selected.
Email addresses capture enabled in Dynatrace Notebooks app
Figure 7 – validate that you are able capture the email addresses.

Test the process of masking the email

In the screenshot below, you can notice that the logs do contain unmasked email addresses:

Unmasked email addresses detected in Dynatrace Notebooks app
Figure 8 – Unmasked email addresses detected.

Test replacing and obfuscating the email address in the log content with a masked value (e.g., xyz):
Add below your existing query the fieldsAdd operation:

| fieldsAdd content = replacePattern(content,  
 
" 
<< 
( 
 'order confirmation email sent to' | 
'find credentials for series ' |  
 'customerID=' |  
 'customerRegistrationProcess-' |  
 'forgottenPasswordProcess-' |  
 'Deleting userId: ' |  
 'Check loyalty account exist for email: ' |  
 'advantageLinkEmailProcess-' |  
 'GET /reminder/' 
) LD:email  
>> 
'@' 
",  
 
“xyz”) 
 

This replaces the email address (stored in the email field) with xyz in the content field.

Masked email addresses in Dynatrace Notebooks app
Figure 9 – Email addresses are masked at query time with xyz value for testing and masking validation.

What it does:

This DQL command modifies the content field in a Dynatrace log record by replacing specific patterns defined – email addresses in this case – with the defined string “xyz”.
Let’s have a closer look at the technical details:

  1. replacePattern Function:
    The replacePattern function searches the content field for matches of a specified pattern and replaces them with a given string (here, “xyz”).
  2. DPL ModifierLookaround:
    Positive Look Behind Modifier <<
    Pattern:
<<
(
'for series' |
'customerID=' | 
'order confirmation email sent to' |
'customerRegistrationProcess-' |
' process forgottenPasswordProcess-' |
'userId: ' |
'Check OnePass account exist for email: ' |
'MobileApp customerRef=' |
'advantageLinkEmailProcess-' |
'GET /reminder/'
)
    • Explanation: The << modifier looks up to 64 bytes before the current position in the log to check for any of the specified phrases. If one is found, the pattern continues. Our pattern includes <<('customerID=' 'GET /reminder/'). If the log line contains customerID=marie_currie@example.com, the modifier confirms that customerID= appears within 64 bytes before the email address, so the match succeeds. If that phrase isn’t found in the preceding 64 bytes, the match fails. Read more about DPL modifiers in our product documentation
  1. Positive Look Ahead >>
    Pattern:
    LD:email >> ‘@' Explanation:  The >> modifier performs a look-ahead check. It ensures the pattern only matches if a specific condition appears after the current position. In this case, after capturing the local part of the email with LD:email, the pattern verifies that an @ symbol follows. This confirms the captured text is indeed part of an email address. For example, take the logline “Can't find credentials for series marie_currie@example.com” The pattern first uses LD:email to capture marie_currie as the local part. The >> '@' check then looks ahead and confirms that @ immediately follows. Because the condition is met, the match succeeds. But if the @ symbol were missing (e.g., Can't find credentials for series marie_curieexample.com), the match would fail.
  2. Replacement:
    The matched pattern (e.g., marie_currie) is replaced with “xyz”. This masks the local part of the email address while leaving the rest of the log entry intact. For instance:

    • Before: customerID=marie_currie@example.com
    • After: xyz@example.com
  3. fieldsAdd content = …: 
    The modified content (with the email local part masked) is stored back into the content field, overwriting the original value.

Additional guidance:

Follow these steps if you wish to mask the entire email address, including the domain:
If you wish to mask or extract the entire email address, you could use a pattern like below:

<<
( 
'order confirmation email sent to ' | 
'find credentials for series ' |  
'customerID=' |  
'customerRegistrationProcess-' |  
'forgottenPasswordProcess-' |  
'Deleting userId: ' |  
'Check loyalty account exist for email: ' |  
'advantageLinkEmailProcess-' |  
'GET /reminder/' 
) 
(LD '@'[A-Za-z0-9'-'']* '.' [A-Za-z]*):email 

In addition to the modifiers and pattern matching literals that were explained earlier, below is the explanation of the last line on how the breakdown of the DPL pattern language:

(LD '@'[A-Za-z0-9'-'']* '.' [A-Za-z]*):email 

  1. LD: Matches line data.
  2. '@‘: Matches the ‘@’ symbol.
  3. [A-Za-z0-9'-'']*: Matches any sequence of alphanumeric characters, hyphens, and single quotes, occurring zero or more times.  Read more here.
  4. '.': Matches the ‘.’ (dot) character.
  5. [A-Za-z]*: Matches any sequence of alphabetic characters, occurring zero or more times.
  6. :email: Captures the matched data and labels it as email. This is not needed if you don’t wish to capture the email address.

Step 5: Filter logs to avoid unintended masking

To ensure we only mask logs from specific sources and pattern containers, apply additional filters to isolate those logs only.

As a best practice, you should pre-filter your logs, by applying this filter directly underneath your fetch logs statement.

Filter by Container Names:

| filter k8s.container.name == "checkout"
OR k8s.container.name == "astroshop"
OR k8s.container.name == "aks-playground"

Filter by Content Patterns:

filter  
( 
matchesPhrase(content, "find credentials for series ") OR  
matchesPhrase(content, "order confirmation email sent to ") | 
matchesPhrase(content, "customerID=") OR  
matchesPhrase(content, "customerRegistrationProcess-") OR  
matchesPhrase(content, "forgottenPasswordProcess-") OR  
matchesPhrase(content, "Deleting userId: ") OR  
matchesPhrase(content, "Check loyalty account exist for email: ") OR 
matchesPhrase(content, "advantageLinkEmailProcess-") OR 
matchesPhrase(content, "GET /reminder/") 
) 

Complete DQL Query

Here’s the full DQL query combining all steps for your reference and to copy/paste:

fetch logs 
| filter  
( 
k8s.container.name == "checkout" OR  
k8s.container.name == "astroshop" OR  
k8s.container.name == "aks-playground" 
) 
AND 
// Match only logs that have specific phrases  
( 
matchesPhrase(content, "find credentials for series ") OR  
matchesPhrase(content, "order confirmation email sent to ") OR  
matchesPhrase(content, "customerID=") OR  
matchesPhrase(content, "customerRegistrationProcess-") OR  
matchesPhrase(content, "forgottenPasswordProcess-") OR  
matchesPhrase(content, "Deleting userId: ") OR  
matchesPhrase(content, "Check loyalty account exist for email: ") OR 
matchesPhrase(content, "advantageLinkEmailProcess-") OR 
matchesPhrase(content, "GET /reminder/") 
) 
| // extract the contents just after these literals that contains the localpart/username within the email address and replace the pattern that comes after these strings. 
fieldsAdd content = replacePattern(content,  
 
" 
<< 
( 
'order confirmation email sent to' | 
'find credentials for series ' |  
'customerID=' |  
'customerRegistrationProcess-' |  
'forgottenPasswordProcess-' |  
'Deleting userId: ' |  
'Check loyalty account exist for email: ' |    
'advantageLinkEmailProcess-' |  
 'GET /reminder/' 
) LD:email  
>> 
'@' 
",  
 
“xyz”) 

Store this query in a Dynatrace Notebook for future reference, and feel free to select the Run button to validate that the output matches your expectations.

Result validation in Notebook app
Figure 10 – Result validation in Notebook.

All our earlier steps did not actually mask the incoming logs, but applied our masking at read/query, and secondly, just for us as the users of the Notebook.

The Notebook and its DQL snippet simply masked the data we visualized. Although this is useful for preventing users from displaying unmasked data when automated, it is even better if logs with PII are masked during ingest time.

Now that we have validated that our pattern works, we can use OpenPipeline to ensure that any log matching the DQL processor rule will undergo a series of processes that transform the log events during the ingest process.

Step 6: Set up OpenPipeline for real-time masking

Let’s configure OpenPipeline to mask email addresses at log ingestion.

  1. Launch OpenPipeline in Settings app: Navigate to OpenPipeline using search or the CTRL+K shortcut. Then select Logs from the OpenPipeline menu.
  2. Create a New Pipeline:
    • Go to Pipelines and create a new pipeline.
    • Provide a meaningful name (e.g., “Mask Email Addresses”).
    • In the Processing tab, add a DQL processor.
  3. Configure the DQL Processor:
    • Name the processor (e.g., “mask email address in astroshop OR checkout OR aks-playground”).
    • Replace the default matching condition true with our patterns defined:
    • ( 
      k8s.container.name == "astroshop" OR  
      k8s.container.name == "checkout" OR  
      k8s.container.name == "aks-playground" 
      ) AND 
      // Match only logs that have specific phrases 
      ( 
      matchesPhrase(content, "order confirmation email sent to ") OR matchesPhrase(content, "find credentials for series ") OR  
      matchesPhrase(content, "customerID=") OR  
      matchesPhrase(content, "customerRegistrationProcess-") OR  
      matchesPhrase(content, "forgottenPasswordProcess-") OR  
      matchesPhrase(content, "Deleting userId: ") OR  
      matchesPhrase(content, "Check loyalty account exist for email: ") OR 
      matchesPhrase(content, "advantageLinkEmailProcess-") OR 
      matchesPhrase(content, "GET /reminder/") 
      )
    • Define what action should be applied when the condition matches, by defining the DQL processor:
      fieldsAdd content = replacePattern(content,  
       
      " 
      << 
      ( 
       'find credentials for series ' |  
       'order confirmation email sent to ' |  
       'customerID=' |  
       'customerRegistrationProcess-' |  
       'forgottenPasswordProcess-' |  
       'Deleting userId: ' |  
       'Check loyalty account exist for email: ' |  
       'advantageLinkEmailProcess-' |  
       'GET /reminder/' 
      ) LD:email  
      >> 
      '@' 
      ",  
       
      “xyz”)

Email ID masking in Dynatrace Notebooks app

While it is suggested to test your new masking with sample data, you can alternatively input ‘example’ as text and select to save the processor and pipeline, as we have validated them earlier in our Notebook.

4. Set up and define Dynamic Routing:

  • Select the Dynamic routing tab
  • Create a new Dynamic route named “Email ID masking for specific container names.”
  • Use the same condition as the matching condition before.
  • Connect the route to the pipeline you’ve just created.
Matching conditions in Dynatrace OpenPipeline with Dynamic Routing
Figure 12 – Matching conditions in Dynatrace OpenPipeline with Dynamic Routing
Once configured and saved, OpenPipeline will automatically mask email addresses during ingestion.

Step 7: Validate the masking

To confirm the email addresses are masked:

  1. Re-run the DQL query in your Dynatrace Notebook, but this time without the parsing.
  2. Check the content field in the logs to ensure email addresses are replaced with xyz.

For example, marie_currie@example.com should now appear as xyz@example.com.
We can verify this by looking for these logs, and we can observe that the email addresses have been fully masked without any parsing at query time. Additionally, you can notice that the logs have the dt.openpipeline.pipelines attribute attached with the pipeline that was used during ingest time. This also confirms that the log was processed in that pipeline.

Processed logs in Dynatrace Notebooks app

Conclusion

In this blog, we used email masking at ingest to demonstrate how to effectively obfuscate various types of sensitive data in logs using OpenPipeline, protecting sensitive data in real-time during the ingest process. By leveraging OpenPipeline’s powerful DQL-based processing and dynamic routing capabilities, you can precisely target specific Kubernetes containers or just any log source with patterns to enhance data privacy and security.

Take the next step

Explore OpenPipeline’s advanced features to:

Streamline Compliance

For comprehensive data-subject rights management, consider the Sensitive Data Center. app. Efficiently manage end-user personal data requests in Grail, with support for regulations like GDPR and CCPA.

Addendum: Exploring this approach in the Dynatrace Playground tenant:

A sample dataset, preloaded with tailored pattern matching and masking for the following steps and explanations, is available in the Dynatrace Playground tenant within this Dynatrace Notebook for hands-on exploration.

The individual steps can also be observed in the playground tenant with the logs ingested from the astroshop demo, while unfortunately, you lack the permissions to configure new pipelines and processors in Playground tenant.

If you look to test this within your personal tenant or a free trial tenant without actual logs containing PII, or running Astroshop yourself, you can download a copy of this Notebook or test with the pattern below, using the data command to generate sample data in a Dynatrace Notebook.

data  
record(content = "2025-04-22 11:46:38 [ERROR|[203.0.113.13] |K9p8Q3xJwZrStUv4XyZ7AbC2n5M6h8T1v0L2r4 |com.example.exmplstore.security.examplePersistentTokenRepository] Can't find credentials for series marie_curie@example.com", log.source = "bash-script"), 
record(content = "2025-04-22 13:50:43 [INFO |[203.0.113.112] |A1b2C3d4EfGhIjK5LmN6oPq7RsT8uVw9XyZ0 |class com.example.exmplfacades.order.exampleCheckoutLogger Checkout ABC] Receive API Request:method=placePayOrder, cartCode=302132310, customerID=freddie@example.com|com.example.exmplstore.checkout.request.PayPlaceOrderRequestDto@3d3a53d5|", log.source = "bash-script"), 
record(content = "2025-04-22 13:58:24 [INFO |||com.example.exmpl.BusinessProcessLoggingAspect] Finish Action: [ PerformSubscriptionAction ], BusinessProcessCode: [ customerRegistrationProcess-martin_luther@example.com-1745294290737], OrderCode:[ n/a ]", log.source = "bash-script"), 
record(content = "2025-04-29 10:16:13 [INFO |||com.example.exmplmarketing.action.exampleSendCustomerNotificationAction] Successfully sent email forgottenPassword message for process forgottenPasswordProcess-jane_austen@example.com-1745885767984", log.source = "bash-script"), 
record(content = "2025-04-29 10:24:26 [INFO |[203.0.113.14] |M3n4P5q6RsTuVwXyZ7aBc8DeF9gHi0JkL2|com.example.exmpl.UserDeleteInterceptor] Deleting userId: cleopatra@example.com, actioned by userId: anonymous", log.source = "bash-script"), 
record(content = "2025-04-29 10:24:36 [INFO |[203.0.113.19] |T2u3V4w5XyZaBc6DeF7gHi8JkL9mNo0PqR1|com.example.exmpl.AdvantageApiClient] Loyalty: Check loyalty account exist for email: leonardo_davinci@example.com , wodCorrelationId 2fee137d-915e-46fb-b390-c69aae3f150", log.source = "bash-script"), 
record(content = "2025-04-29 10:22:47 [INFO |||com.example.exmplbusproc.aop.BusinessProcessLoggingAspect] Begin Action: [ exampleAdvantageLinkEmailAction ], BusinessProcessCode: [ advantageLinkEmailProcess-albert_einstein@example.com-1745886166709], OrderCode:[ n/a ]", log.source = "bash-script"), 
record(content = "10.96.152.11 - - [29/Apr/2025:10:33:43 +1000] \"GET /reminder/shakespeare@example.com/1 HTTP/1.1\" 200 90 \"https://www.example.com/my-account/order-details/306399901\" \"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/135.0.0.0 Safari/537.36\" 203.0.113.145 - T2u3V4w5XyZaBc6DeF7gHi8JkL9mNo0PqR1 100", log.source = "bash-script") 
 
| filter matchesPattern(content, "LD '@' [A-Za-z0-9'_']* '.' [A-Za-z]* (LD | EOS)")
Generated demo sample data showcasing PII data found in results in Dynatrace Notebooks app
Figure 14 – Generated demo sample data showcasing PII data found in results.

The post How to mask PII like email addresses appearing in logs with Dynatrace: An advanced use case appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-to-mask-pii-like-email-addresses-appearing-in-logs-with-dynatrace-an-advanced-use-case/feed/ 0
Next-level batch job monitoring and alerting part 2: Using AI to automatically identify issues and workflows to remediate them https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-part-2-using-ai-to-automatically-identify-issues-and-workflows-to-remediate-them/ https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-part-2-using-ai-to-automatically-identify-issues-and-workflows-to-remediate-them/#respond Wed, 21 Jan 2026 14:38:37 +0000 https://www.dynatrace.com/news/?p=72554 Next-level batch job monitoring and alerting

Let’s say you have a nightly account update batch job that processes user transactions, recalculates balances, and synchronizes data across services in a financial application. If it fails or gets stuck and consumes so many resources that your service degrades, you’ll suffer significant business impact. Every second of such an incident can lead to customer […]

The post Next-level batch job monitoring and alerting part 2: Using AI to automatically identify issues and workflows to remediate them appeared first on Dynatrace news.

]]>
Next-level batch job monitoring and alerting

Let’s say you have a nightly account update batch job that processes user transactions, recalculates balances, and synchronizes data across services in a financial application. If it fails or gets stuck and consumes so many resources that your service degrades, you’ll suffer significant business impact. Every second of such an incident can lead to customer frustration, SLA breaches, and lost revenue.

In our previous blog, we demonstrated how you can use Dynatrace to observe batch jobs and how and why to treat them as first-class citizens in your monitoring strategy. We demonstrated how to uncover performance trends, pinpoint failures, and reveal bottlenecks. But detection isn’t enough.

This blog post takes you to the next level. We’re no longer just talking about detecting problems—we’re talking about fixing them automatically with the help of Dynatrace Davis® AI, Workflows, EdgeConnect, and ServiceNow integration.

We’ll walk you through the automated process of determining root cause, creating an incident in ServiceNow, and automatically remediating the problem.

Architecture overview: A modern, cloud-native setup

Our example environment is a Kubernetes-based architecture that mirrors many modern cloud-native enterprise deployments.

Here’s what the setup looks like:

  • A Kubernetes cluster orchestrating all workloads
  • A public-facing NGINX proxy handling incoming traffic
  • A React front-end delivering the user interface
  • A Java-based broker service responsible for backend processing and database interaction
  • A persistent database service for storing and retrieving application data
Architecture and flow of automatic detection and remediation
Figure 1 – Architecture and flow of automatic detection and remediation

This is automatically instrumented using the Dynatrace Kubernetes Operator, which injects monitoring capabilities directly into the cluster without manual configuration. This provides end-to-end visibility across services, infrastructure, and dependencies with no code changes or changes to existing containers.

Enhancing AI context: Tracking batch job events in Dynatrace

For custom use cases and unique business scenarios, Davis® AI needs to be made aware of specific events through custom approaches. While traditional monitoring captures standard metrics, enabling intelligent reasoning around specific workflows—like batch jobs—means providing targeted event tracking. This ensures Davis AI has full visibility into these activities, allowing it to understand and respond to them effectively.

Contextual events, such as the start of the batch job, could be sent to Dynatrace. This functionality allows you to maintain a clear timeline of all activities, including initiating and completing batch jobs, enhancing visibility and control over critical processes.

In our case, we use an HTTP call that follows this pattern, which utilizes the Dynatrace event types documented here:

An example Event API Payload for attaching batch job information to services:

{ 
  "entitySelector": "type(\"SERVICE\"),tag(service:broker-service 
  "eventType": "CUSTOM_INFO", 
  "properties": { 
    "batch-job-name": "update-account", 
    "process-id": "%s", 
    "workload-name": "account-updater", 
    "namespace": "easytrade" 
  }, 
  "title": "%s update-account batch-job" 
}

This payload attaches batch job information to services matching the entitySelector criteria. The event will appear in the timeline of services tagged with broker-service linking batch job execution data to the relevant services.

The Service with all events recorded about the start and completion of the batch job
Figure 2 – The Service with all events recorded about the start and completion of the batch job.

Davis AI in action: Instant root cause analysis

In this real-world example, the application team noticed latency spikes and increased error rates reported by end users. But before anyone raised a ticket or sent an alert, Dynatrace Davis AI had already detected the anomaly.

Here’s what Davis AI surfaced immediately:

  • Problem severity and duration
  • Number of users affected
  • Which services were degraded, and which business transactions were impacted
  • And most importantly: the root cause
AI-powered problem detection with affected users, events, SLOs, and the root cause of the problem
Figure 3 – AI-powered problem detection with affected users, events, SLOs, and the root cause of the problem.

In this case, the root cause was traced to a batch job that had started consuming excessive resources—CPU, memory, and I/O—all of which starved the Java broker service and caused downstream failures.

Without AI, isolating this root cause across layers could have taken hours. With Davis, it was done in minutes.

Identifying service failures due to batch job failures

The screenshot below displays distributed traces filtered to show batch job-related activity and reported failures. It reveals issues in other services, such as BrokerService, where failures from stuck batch jobs are impacting downstream services.

Screenshot showing BrokerService failures from stuck batch jobs
Figure 4 – Screenshot showing BrokerService failures from stuck batch jobs.

Tapping ServiceNow for instant remediation while Davis gave us the diagnosis, we still need a solution. This is where our integrations with ServiceNow come in. Dynatrace Workflows identifies the issue and checks ServiceNow for an existing ticket. If none exists, it leverages Dynatrace EdgeConnect to execute a Kubernetes API, suspending the affected batch job for remediation. If a ticket is already open, the Workflow adds comments to ServiceNow and follows the same remediation process.

Orchestrating recovery: Dynatrace Workflows

We created a custom Dynatrace Workflow specifically for this batch job scenario. It’s designed to react to Davis AI-detected problems that match a particular pattern—resource-intensive jobs degrading core services.

Here’s how the workflow functions:

  1. Trigger. Once the problem is detected and Davis confirms the root cause, the workflow is automatically triggered.
  2. ServiceNow integration. Without manual input, it instantly opens a ServiceNow incident with all contexts—affected services, severity, root cause.
  3. Live updates. As the situation evolves, the ticket is updated in real time, ensuring that both engineering and service management teams are in sync.
  4. Conditional logic. The workflow checks for remediation eligibility—e.g., is it safe to pause or reschedule the batch job?
  5. Remediation. The workflow now initiates the remediation by EdgeConnect by executing a Kubernetes API to suspend or delete the job
  6. Validation and closure. The causal AI validates successful remediation by confirming the batch job suspension, then automatically closes the problem ticket. Optionally, EdgeConnect can perform additional verification to ensure the root cause has been eliminated and system stability is restored.

All of this happens without anyone needing to log into a dashboard.

Closing the loop: Secure remediation with Dynatrace EdgeConnect

Now, let’s talk about the actual fix.

With the root cause identified and the ServiceNow ticket tracking the issue, the final step is to initiate remediation safely and securely inside the Kubernetes cluster.

This is achieved using Dynatrace EdgeConnect, a lightweight, secure connector that allows Dynatrace to command and interact with your private infrastructure; in this case, Kubernetes.

In our case, EdgeConnect was configured to:

  • Connect securely to the Kubernetes API
  • Execute a custom remediation action, such as scaling down the batch job, adjusting its priority, or rescheduling it to an off-peak window
  • Verify the effect of the remediation and confirm resolution back in Dynatrace and ServiceNow
Workflow that detects such custom problems and initiates remediation and notification to IT Service Management
Figure 5 – Workflow that detects such custom problems and initiates remediation and notification to IT Service Management.

All of this occurred without breaking security boundaries, using EdgeConnect’s outbound-only architecture.

Instant notifications with Slack integration

SysAdmins, DevOps, and SRE teams need to stay informed in real time to ensure system reliability. The Dynatrace integration with Slack delivers instant notifications for critical events, such as batch job failures or system errors, directly to your team’s Slack channels.

In this case, they are simply notified as all the action has already been taken care of by intelligent automation.

Notification seen in Slack, which updates the Problem in ServiceNow
Figure 6 – Notification seen in Slack, which updates the Problem in ServiceNow.

In this example, we conducted remediation by automatically suspending the batch job, as confirmed by running a relevant command to check the job status. Note that remediation approaches and use cases may vary, and we recommend tailoring solutions to your specific environment and requirements.

Demonstration of the automatic suspension of the cronjob through automation
Figure 7 – Demonstration of the automatic suspension of the cronjob through automation.

From reactive to proactive with AI-driven automation

Batch jobs have long been a blind spot in observability—often managed outside core APM tooling or considered too niche for automated handling.

But today, Dynatrace brings batch job management into the mainstream of modern observability and automation.

  • With Davis AI, you get root cause detection in real time.
  • With Workflows, you turn those insights into action.
  • With EdgeConnect, you push changes securely to production environments.
  • And with ServiceNow integration, you keep ITSM workflows up to date without lifting a finger.

This is a new era of autonomous cloud operations—where custom, complex issues like batch job failures are no longer exceptional cases, but standard parts of your AI-driven automation strategy.

How to get started

It’s considered best practice to send custom events to Dynatrace to enhance monitoring capabilities. These events help keep Dynatrace informed about key activities or changes in your environment.

If any of these events correlate with issues in the landscape, Davis AI automatically analyzes them and identifies the root cause. This insight can then be used to trigger Workflows that help remediate potential problems proactively.

Start optimizing your observability today!

The post Next-level batch job monitoring and alerting part 2: Using AI to automatically identify issues and workflows to remediate them appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-part-2-using-ai-to-automatically-identify-issues-and-workflows-to-remediate-them/feed/ 0
Significantly improve your Mainframe availability by connecting logs with traces https://www.dynatrace.com/news/blog/significantly-improve-your-mainframe-availability-by-connecting-logs-with-traces/ https://www.dynatrace.com/news/blog/significantly-improve-your-mainframe-availability-by-connecting-logs-with-traces/#respond Fri, 31 Oct 2025 08:00:35 +0000 https://www.dynatrace.com/news/?p=71629 Connecting logs and traces related content

Speed up resolution of mainframe application issues and switch to a proactive and preventive mode of operations, through z/OS logs and traces in context with both your Dynatrace SaaS and Managed deployments.

The post Significantly improve your Mainframe availability by connecting logs with traces appeared first on Dynatrace news.

]]>
Connecting logs and traces related content

Logs are essential observability telemetry

Organizations are increasingly challenged to deliver seamless digital services, an essential component for achieving business-critical objectives. These challenges are amplified by complex hybrid cloud environments, where managing diverse technologies across cloud providers and platforms like IBM Z Mainframe becomes particularly demanding.

To address this complexity, it’s vital to unify observability telemetry and its signals across all layers of the application delivery chain, including the mainframe platform, within a single AI-powered observability solution.

Dynatrace enrichment capabilities enhance ingested log records by adding contextual metadata such as trace IDs, span IDs, and process group instance IDs automatically and without any manual tagging efforts. This allows seamless correlation between logs, application performance metrics, and traces. This is essential for accelerating AI-driven root cause analysis and significantly reducing time spent on shortening RCA and MTTx, as many of our customers can confirm.

Simplify and automate z/OS log collection

Dynatrace Log Management & Analytics extends beyond distributed technology stacks to include support for the IBM Z Mainframe platform. It automatically captures and ingests logs from monitored IBM CICS and IBM IMS regions and offers you advanced ingest rules. All collected logs are enriched automatically with topological metadata, allowing seamless mapping to Dynatrace’s topology and entity model for z/OS Hosts (LPARs) and z/OS Processes (regions).

Dynatrace automatically maps log lines to z/OS entities, in this case, to the process name and job ID of a CICS region
Figure 1. Dynatrace automatically maps log lines to z/OS entities, in this case, to the process name and job ID of a CICS region

Additionally, logs can be enriched with Trace IDs and Span IDs to precisely correlate each log line with the corresponding CICS or IMS trace or span that generated it.

Dynatrace can map log lines to a specific z/OS trace, which allows you to directly navigate to the trace that created the specific log line (via “View trace”).
Figure 2. Dynatrace can map log lines to a specific z/OS trace, which allows you to directly navigate to the trace that created the specific log line (via “View trace”).

This dramatically simplifies navigation for any user in your organization. By selecting the log line containing the Trace ID and Span ID, you can directly navigate to the related trace while understanding the load times and delay.

This example shows an end-to-end trace and the log line that was written by a CICS COBOL program
Figure 3. This example shows an end-to-end trace and the log line that was written by a CICS COBOL program

Let’s summarize what we’ve seen so far. Log enrichment significantly enhances and accelerates:

  • Correlation with distributed traces, enabling end-to-end visibility across systems.
  • Troubleshooting, by linking logs to specific spans or transactions—accelerating troubleshooting.
  • Observability for both structured and unstructured log data, ensuring comprehensive insights regardless of log format.

Beyond the agent: Stream mainframe logs to Dynatrace with OpenTelemetry

Dynatrace OneAgent® already supports ingestion of logs out of the box. However, when it comes to mainframe environments, the story is more nuanced.

Why all logs are not created equal

Some logs—especially those on mainframes—are proprietary, customer-specific, and deeply embedded in legacy workflows. And while Dynatrace is constantly adding additional and automated coverage for additional log types, some might never be supported natively by OneAgent, simply because their structure and relevance are unique to each customer.

But that doesn’t mean they’re out of reach.

OpenTelemetry to the rescue

For logs that fall outside OneAgent’s native scope, OpenTelemetry offers a powerful alternative. By deploying an OpenTelemetry Collector, customers can stream log data from their LPARs to a distributed host—preferably to Linux, Windows, or zLinux to save MSU consumption. But even z/OS itself is an option for hosting an OpenTelemetry Collector.

The Collector supports:

  • Filelog receiver for arbitrary text files
  • Syslog receiver for structured system logs
  • Filter processor for preprocessing and enrichment
  • Concurrent export to multiple backends, including Dynatrace via OTLP

Getting logs off the mainframe

There are several ways to move logs from LPARs to distributed systems:

  • SFTP: A blunt but reliable method
  • z/OSMF: Offers REST API access to SMF records
  • z/OS Data Gatherer: SMF REST Services
  • Custom scripts or processes: Tailored to specific datasets
  • Streaming frameworks: Kafka, MQ, or even FTP-to-Collector bridges

The bottom line: Just transfer log data from the mainframe to the host where the OpenTelemetry Collector is located, and it will handle everything for you from there.

There’s no strict requirement to provision a dedicated host to run your OpenTelemetry Collector. If Dynatrace OneAgent is already monitoring one of your LPARs, the ActiveGate hosting the Dynatrace zRemote component is a perfectly suitable environment for the Collector.

The ActiveGate hosting the zRemote mediates the ingestion of both out-of-the-box logs and OpenTelemetry logs.
Figure 4. The ActiveGate hosting the zRemote mediates the ingestion of both out-of-the-box logs and OpenTelemetry logs.

Preprocessing and enrichment

Both the Dynatrace Distro and the Contrib Distro of the OpenTelemetry Collector support advanced preprocessing:

  • Timestamp normalization (for example, converting z/OS timestamps to Unix time)
  • Resource attribute extraction (for example, job name, LPAR ID, subsystem)
  • Sensitive data masking and filtering

This ensures that even complex logs are transformed into structured telemetry before reaching the backend.

The cherry on top: Dynatrace OpenPipeline

While OpenTelemetry handles ingestion and transformation, Dynatrace OpenPipeline® adds another layer of intelligence during ingestion and processing:

  • Further enrich logs with business context
  • Extract metrics, events, and business observability events
  • Apply AI-driven baselining and anomaly detection
  • Transform, mask, or drop data
  • Some of these capabilities overlap with the Collector, giving users flexibility to choose where to apply logic based on performance, cost, and control.

Operlog: A prime candidate

Let’s take Operlog as an example—a system log that many customers are keen to stream. While it’s not a simple text dataset, creative solutions can bridge the gap. If you can extract Operlog records and store them as text on a Linux host, the Collector can ingest them immediately. From there, Dynatrace visualizations and alerting kick in.

Here’s a snapshot of how Operlog looks once streamed and processed in Dynatrace:

Log entries captured from Operlog visualized in Dynatrace
Figure 5. Log entries captured from Operlog visualized in Dynatrace

Not satisfied with only logs?

You’re not limited to ingesting only proprietary log files via OpenTelemetry.

Do you have access to metrics that are relevant for tracking the health of the subsystems on your mainframe? Would you rather feed in certain data as events instead of logs?

Just as the OpenTelemetry Collector is highly customizable, Dynatrace offers multiple ways of ingesting all these signals.

Possible sources for OpenTelemetry signals and how Dynatrace ingests them
Figure 6. Possible sources for OpenTelemetry signals and how Dynatrace ingests them

Conclusion

Mainframe logs may be complex, but they’re not unreachable. With OpenTelemetry and Dynatrace working in tandem, even the most proprietary datasets can be brought into the fold. Whether you’re using OneAgent, OpenTelemetry, or a hybrid approach, the key is creativity—and the right tooling.

What’s next

Dynatrace currently supports CICS MSGUSR and IMS Master Terminal Logs via its z/OS Agents, and remains committed to enhancing these capabilities. This includes exploring support for additional z/OS log types and expanding the scope of information captured through the z/OS Agents.

Log monitoring is also available for Linux on IBM Z and LinuxONE. For more details, see the blog post, Enable full observability for Linux on IBM Z mainframe now with logs.

Get started with Dynatrace log observability

If you’re looking to elevate your end-to-end observability and explore tailored possibilities within your specific z/OS environment, we’d be happy to connect. Reach out to us to request a demo and dive deeper into what Dynatrace can offer.

The post Significantly improve your Mainframe availability by connecting logs with traces appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/significantly-improve-your-mainframe-availability-by-connecting-logs-with-traces/feed/ 0
Predictable costs for Log Management & Analytics with new simplified licensing plan https://www.dynatrace.com/news/blog/predictable-costs-for-log-management-analytics-simplified-licensing-plan/ https://www.dynatrace.com/news/blog/predictable-costs-for-log-management-analytics-simplified-licensing-plan/#respond Thu, 19 Dec 2024 18:21:04 +0000 https://www.dynatrace.com/news/?p=67091 Dynatrace Log Management & Analytics graphic

As cloud complexity increases and security concerns mount, organizations need log analytics to discover and investigate issues and gain critical business intelligence. But exploring the breadth of log analytics scenarios with most log vendors often results in unexpectedly high monthly log bills and aggressive year-over-year costs. To give organizations the freedom to explore log analytics […]

The post Predictable costs for Log Management & Analytics with new simplified licensing plan appeared first on Dynatrace news.

]]>
Dynatrace Log Management & Analytics graphic

As cloud complexity increases and security concerns mount, organizations need log analytics to discover and investigate issues and gain critical business intelligence. But exploring the breadth of log analytics scenarios with most log vendors often results in unexpectedly high monthly log bills and aggressive year-over-year costs. To give organizations the freedom to explore log analytics without barriers due to cost concerns, Dynatrace is proud to announce a new Dynatrace Platform Subscription (DPS) pricing model option called Retain with Included Queries.

With this new DPS pricing model option, customers can retain data at a fixed low cost with no additional cost to query for up to 35 days. This model provides a predictable way for customers to manage and analyze logs, drive log management tool consolidation, and reduce costs while gaining maximum value from their log data.

Based on customer feedback, we’re offering the Retain with Included Queries pricing model as an alternative to our existing usage-based plan. Both plans offer the same low ingest price. However, the new all-access plan combines retention and queries into one low price to simplify scoping and budgeting.

Retain with Included-Query pricing simplifies forecasting and annual usage calculation costs. Customers who choose this pricing option get:

  • Retention cost: $0.02 per GiB per day
  • No cost to query for up to 35 days
  • Ingest cost: $0.20 per GiB ingested (no change)

With this approach, the whole team can leverage the power of Grail queries and dashboards without worrying about limiting query usage. Customers can configure the Retain with Included Queries option with retention periods ranging from 10 to 35 days. Customers requiring longer retention periods should opt for our existing usage-based pricing, which supports retention for up to 10 years.

Queries are included

  • Predictable pricing: If you know the number of logs you ingest daily, then you’ll know roughly your total annual cost upfront, providing peace of mind and less managerial overhead.
  • Simple scoping: Remove the complexity associated with predicting query search volumes. Realize cost savings immediately for high-query usage scenarios. Get started quickly!
  • No cost management required: Once your configured retention period ends (a maximum of 35 days), logs are automatically deleted. No oversight is needed over query usage.
Dynatrace Log Management & Analytics pricing
Figure 1. Dynatrace Log Management & Analytics pricing

Usage-based pricing is still an option

Over time, our existing usage-based pricing is the more cost-optimized option, as you only pay for the queries your users execute, and you benefit from the competitive $0.0007 per GiB per day to retain logs for up to 10 years. Queries are charged at $0.0035 per GiB scanned. Usage-based pricing is ideal for organizations with longer retention requirements and known query patterns. This pricing flexibility allows customers to optimize their log analysis expenses by paying only for what they use.

Cost-efficient:

  • Lowest upfront cost
  • Charges are strictly based on query execution

Scalability:

  • Ideal for businesses with varying query demands
  • Adapt dynamically to usage patterns

Retention:

  • Supports log storage from 1 day to 10 years
  • Optimal for longer-term log analytics needs

Guidance on using both plans

The Retain with Included Queries pricing option is a great way to get started while you learn about your query usage. Dynatrace includes a ready-made cost dashboard that provides insights into query usage and DQL best practices. Once you develop best practices and are confident with your consumption patterns, you can switch to usage-based pricing to maximize the value of your DPS investment.

Innovations on the horizon*

We’re very excited about our new Retain with Included Queries pricing, but we expect to deliver more updates. This pricing model is part of our plan to introduce new features that help customers align the right pricing strategies to their use cases. With these features, customers can easily see, manage, and choose how to align the Retain with Included Queries pricing with the usage-based pricing model.

Customers will soon be able to mix and match log pricing options on a per-bucket basis and provide users with access to both models simultaneously. This flexibility will allow customers to optimize the pricing selection based on the anticipated use case associated with each bucket, yielding even greater savings and value.

Retain with Included Queries: Start here

With the Retain with Included Queries pricing model, Dynatrace now offers a more cost-effective way to get started with log analytics. Drive efficiency and get more value out your logs with this predictable pricing model while you’re building your log analytics practices.

State of Log Management 2026

Download the report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

* Disclaimer: This publication may include references to the planned testing, release, and/or availability of Dynatrace products and services. The information provided in this publication is for informational purposes only; its contents are subject to change without notice, and it should not be relied on in making a purchasing decision. The information is not a commitment, promise, or legal obligation to deliver any material, code, or functionality. The development, release, and timing of any features or functionality described for products remains at the sole discretion of Dynatrace

The post Predictable costs for Log Management & Analytics with new simplified licensing plan appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/predictable-costs-for-log-management-analytics-simplified-licensing-plan/feed/ 0
Easily troubleshoot z/OS application issues through logs with Dynatrace https://www.dynatrace.com/news/blog/troubleshoot-z-os-application-issues-logs-with-dynatrace/ https://www.dynatrace.com/news/blog/troubleshoot-z-os-application-issues-logs-with-dynatrace/#respond Wed, 09 Oct 2024 19:42:19 +0000 https://www.dynatrace.com/news/?p=66117 Explore logs graphics

New support for the IBM z/OS operating system automates log discovery and enables collection at scale. Get better hybrid cloud observability with the Dynatrace platform by including automatically enriched log data and faster issue troubleshooting.

The post Easily troubleshoot z/OS application issues through logs with Dynatrace appeared first on Dynatrace news.

]]>
Explore logs graphics

Logs become an integrated part of observability

Organizations face daily challenges in delivering integrated digital services essential to meeting their business goals. This is particularly true in hybrid cloud architectures, where system complexity, security, and performance are difficult to manage across cloud providers and the IBM Z mainframe platform.

A prerequisite for a successful hybrid cloud is maintaining observability, including all telemetry signals: logs, metrics, and traces. Without combining these signals in a unified AI-powered observability platform, the effectiveness of AIOps workflows in remediating problems is diminished, leading to wasted investment.

The Dynatrace® software intelligence platform can help you manage the complexity of digital services and hybrid clouds by providing holistic end-to-end visibility from the frontend, where your customers interact with your application, to the backend, where business transactions are processed.

Extend root cause analysis to logs on IBM z/OS

Dynatrace provides a platform for observing hybrid clouds and introduces support for log collection of the IBM z/OS operating system. This includes IBM CICS regions and IBM IMS subsystems.

You can now extend root cause analysis for any issue identified by Davis® AI with logs that are automatically linked to z/OS applications, transactions, or other identifiers specific to the environment hosting the resources.

Dynatrace offers log management and collection in a single place for public and private clouds and mainframe platforms such as IBM Z or LinuxONE. This makes log collection policies much more effective and transparent.

For example, you can apply a filter change to ignore certain logs from your central Dynatrace environment on all your monitored platforms without making any manual adjustments.

Centralized data masking rules allow administrators to easily configure the masking of sensitive data. This allows customers to address local or industry-specific regulatory or privacy requirements. The configured rulesets are directly deployed to the OneAgents, where the logs are collected.

Speed up your troubleshooting processes

Log analysis is typically one of the first steps in troubleshooting frontend problems. When a critical issue arises, it’s essential to have the right logs available to quickly and easily understand the full scope of what’s happening within your applications on the backend.

Dynatrace automatically discovers and collects logs from monitored IBM CICS regions and IBM IMS subsystems. All collected logs are enriched with metadata to map them to the entity model of z/OS hosts (logical partitions) and z/OS processes (regions and subsystems).

Dynatrace dramatically shortens Mean Time To Identify (MTTI) and Remediate (MTTR) times for incidents, as relevant log lines are provided in the context of detected problems.

Incidents are often not tied to a single component, as surrounding components and services can cause disruptions. With a single click, Dynatrace provides a surrounding logs view, showing the log lines of related components.

The newly released Dynatrace Logs app offers broad insights for manual investigation. Easy click-to-filter elements and a newly introduced DQL editor capable of translating these selected filters into Dynatrace Query Language (DQL) improve the experience for novice users.

Thanks to the enriched log data, log lines are connected to the respective z/OS Host pages.

Monitoring logs with Dynatrace facilitates novel ways to analyze telemetry data, significantly expanding the observability use cases for IBM Z mainframes. For example, with DQL queries, operators can quickly access all abends or drill down into specific job statistics.

Dynatrace Notebooks allows deep analysis by leveraging the Dynatrace Query Language within notebooks along with the newly introduced Davis CoPilot™ integration, a natural text-to-DQL builder.

Get started with logs on Dynatrace

Get started with Log Management and Analytics powered by Dynatrace Grail™ or Log Monitoring Classic (for Dynatrace Managed deployments). To start, deploy Dynatrace on your IBM z/OS operating system and set up log collection. You can easily set up log collection by turning on the provided z/OS log ingest rules globally or for a specific host.

Go to Settings > Log Monitoring > Log ingest rules and turn on z/OS CICS message user and z/OS IMS master terminal to start collecting logs.

  • Log monitoring on IBM z/OS is available with the release of Dynatrace OneAgent version 1.297 and ActiveGate version 1.297 (with the zRemote module) for Dynatrace SaaS and Dynatrace Managed.

The post Easily troubleshoot z/OS application issues through logs with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/troubleshoot-z-os-application-issues-logs-with-dynatrace/feed/ 0
Next-level batch job monitoring and alerting: Elevate performance and reliability https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-elevate-performance-and-reliability/ https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-elevate-performance-and-reliability/#respond Fri, 27 Sep 2024 15:45:45 +0000 https://www.dynatrace.com/news/?p=65730 Dynatrace Security

Batch jobs are the backbone of automated, scheduled processes that execute tasks in bulk, such as data processing, system maintenance, or report generation. These jobs, which typically run in the background without user interaction, are critical and indispensable for handling large-scale operations efficiently.

The post Next-level batch job monitoring and alerting: Elevate performance and reliability appeared first on Dynatrace news.

]]>
Dynatrace Security

As batch jobs run without user interactions, failure or delays in processing them can result in disruptions to critical operations, missed deadlines, and an accumulation of unprocessed tasks, significantly impacting overall system efficiency and business outcomes. The urgency of monitoring these batch jobs can’t be overstated.

Monitor batch jobs

Monitoring is critical for batch jobs because it ensures that essential tasks, such as data processing and system maintenance, are completed on time and without errors. Failures, delays, or resource issues can lead to operational disruptions, financial losses, or compliance risks. Continuous monitoring enables early detection of problems, allowing quick remediation and maintaining business continuity.

Most jobs provide detailed information about job execution, including status, errors, and processing times in logs. The first step in monitoring batch jobs is to ingest these logs into Dynatrace. This is achieved by identifying the log files generated by the batch job program.

Apply basic filtering to ensure the availability of batch job-related logs. In this case, filter the logs based on relevant phrases or keywords.
Figure 1. Apply basic filtering to ensure the availability of batch job-related logs. In this case, filter the logs based on relevant phrases or keywords.

In this case, batch job statuses are constantly written from the deployment name get-cc-status-*. Thus we can create a rule in Dynatrace to ingest these logs via OneAgent without making any changes to the container, cluster, or host. Logs can also be ingested from various sources, including OpenTelemetry and Fluentbit.

A great reference is our blog post, Leverage edge IoT data with OpenTelemetry and Dynatrace, in which we documented the required steps to parse and ingest a single JSON log file into Dynatrace via OpenTelemetry.

Once logs are ingested, parsing the key messages is crucial. Below is a sample query that demonstrates how batch jobs can be parsed to extract important fields:

fetch logs
| filter matchesPhrase(content, "JOBS") AND matchesPhrase(content, "RunID")
| filter matchesValue(dt.entity.host, "HOST-HOSTID12345678")
| parse content , "
LD 'JOBS.' WORD:Job
LD 'RunID ' STRING:RunId
LD:status"
| fields timestamp, Job, status, content, RunId
| filterOut status == "."
| fieldsAdd start_time=if(contains(content,"started."),timestamp)
| fieldsAdd end_time=if(contains(content,"ended normally."),timestamp)
| fields timestamp, content, Job, status, RunId,start_time,end_time
Parsing the log lines that have critical data related to batch job status
Figure 2. Parsing the log lines that have critical data related to batch job status

Now that we can parse critical information, we can make informed decisions. However, it’s important to know if a job that started has ended within the expected timeframe. When a batch job exceeds its allotted time, the issue must be quickly identified and remediated.

Capture the time difference between two log entities

We use JavaScript within Dynatrace Dashboards to determine whether a previously started job was successfully completed. This three-level approach helps track how long a job took to complete and identifies any stuck jobs.

  1. Identify the unique property of each job and initialize its structure.
    const batch = {};
    
    /* Reiterate through each record and populate the data-structure*/
    for (const record of recordSet) {
      const runId = record['RunId'];
      if (!batch[runId]) {
        batch[runId] = {
          Job: record["Job"],
          run_id: runId,
          Status: "",
          JobStarted: null,
          JobEnded: null,
          Duration: "NA"
        };
      }
    }
  2. Process each job’s start time, end time, and status from the DQL parsed output.
    if (record["start_time"]) batch[runId].JobStarted = utcToLocal(record["start_time"]);
    if (record["end_time"]) batch[runId].JobEnded = utcToLocal(record["end_time"]);
    
    if (!statusLocked[runId]) {
      let status = record["status"]?.trim() || "";
    
      if (status.toLowerCase().includes("ended with return code")) {
        batch[runId].Status = "Failed";
        statusLocked[runId] = true;
      } else if (status == "started.") {
        batch[runId].Status = "Running";
      } else if (status == "ended normally.") {
        batch[runId].Status = "Completed without errors";
        statusLocked[runId] = true;
      } else {
        batch[runId].Status = status;
      }
    }
  3. Update the job status based on specific conditions (running, failed, completed).
    /* Leverage pre-populated data to identify duration for the completed jobs*/
    for (const runId in batch) {
      const job = batch[runId];
    
      if (job.JobStarted && job.JobEnded) {
        const startTime = new Date(job.JobStarted);
        const endTime = new Date(job.JobEnded);
        const duration = endTime - startTime;
    
        job.Duration = `${duration / 1000} seconds`;
      }
    }

Resources for the dashboard and workflow mentioned above can be found in this GitHub repository.

Individual batch job status with processing times and status
Figure 3. Individual batch job status with processing times and status
Advanced statistics for further analysis of batch jobs (median duration and job by status)
Figure 4. Advanced statistics for further analysis of batch jobs (median duration and job by status)

Correlate the impact of batch jobs with the application

Batch jobs should not impact applications because they run in the background. While they consume resources, they shouldn’t impact resource usage or client-facing applications. We can use Dynatrace Grail™ data lakehouse for unified observability data.

Correlate batch job runs with Application and Service resource utilization
Figure 5. Correlate batch job runs with Application and Service resource utilization

Adjust log parsing to account for varying log patterns

No two batch jobs are the same, and the log patterns you encounter might differ from what you see here. You can achieve the same results by parsing the logs. Parsing logs, as shown above, can be done using DPL Architect.

DPL Architect is a handy tool, accessible through the Notebooks app, that helps you quickly extract fields from records. It helps create patterns, provides instant feedback, and allows you to save and reuse DPL patterns, for faster access to data analytics use cases. This blog post offers further details about DPL architect.

Alerting for long-running or failed batch jobs

Constantly monitoring a dashboard isn’t practical, so you need automated alerting. Dynatrace workflows can check the status of batch jobs every 15 minutes and send alerts for failures or long-running jobs. These alerts can trigger actions or notifications sent via Slack, Teams, or as a ticket in your IT service management tool. In this example, the notifications are sent via email.

Automate batch job alerting and reporting
Figure 6. Automate batch job alerting and reporting

Conclusion

Monitoring batch jobs is essential to ensure they run smoothly and within expected timeframes. We can effectively identify issues such as long-running or failed jobs by ingesting logs into Dynatrace, parsing critical job information, and using custom logic to track job completion times. Implementing automated alerts triggering actions and notifications ensures proactive management, allowing teams to quickly resolve problems and maintain operational efficiency. With these tools in place, organizations can improve the reliability and performance of their batch-processing systems.

Use the approach detailed in this blog post to implement advanced batch job monitoring in your environment. Download the dashboards and Notebooks from this GitHub repository and start your automation journey today.

What’s next

In a future blog post, we’ll show how batch job management can be efficiently orchestrated using workflows and predictive analysis to schedule and run jobs optimally. With Davis® AI identifying root causes, workflows can be used to stop erroneous batch executions.

Additionally, Davis® AI prediction analysis, in conjunction with workflows, can reschedule or pause jobs to ensure optimal resource utilization, preventing any negative impact on the application landscape.

Download the Dashboards and Notebooks used in this blog post from our GitHub repository.

The post Next-level batch job monitoring and alerting: Elevate performance and reliability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-elevate-performance-and-reliability/feed/ 0
Unlock log analytics: Seamless insights without writing queries https://www.dynatrace.com/news/blog/log-analytics-seamless-insights-without-writing-queries/ https://www.dynatrace.com/news/blog/log-analytics-seamless-insights-without-writing-queries/#respond Tue, 28 May 2024 14:48:22 +0000 https://www.dynatrace.com/news/?p=64183

Logs are an integral part of the daily workflow for your DevOps and SRE teams to understand what’s happening in your tech stack. No matter the industry you operate in or the scale of your business, getting value from log data is often slowed down by challenges: making sure the right logs are monitored, finding the relevant logs when you need answers, and making sense of logs in the context of other data like traces, events, and metrics.

The post Unlock log analytics: Seamless insights without writing queries appeared first on Dynatrace news.

]]>

Logs provide answers, but monitoring is a challenge

Manual tagging is error-prone

Making sure your required logs are monitored is a task distributed between the data owner and the monitoring administrator. Often, it comes down to provisioning YAML configuration files and listing the files or log sources required for monitoring. This manual, error-prone approach can lead to monitoring gaps, which become critical when a host or service has an outage or incident.

Finding the right logs is cumbersome

Even if your logs are monitored, you need to make sense of the vast data volume. As the scale and complexity of your tech stack grows, you might need to navigate the maze of hosts or Kubernetes clusters, apps, and microservices and understand the relevance and risks associated with logs originating from these entities. Challenges compound: Manual tagging of log sources has long been difficult regarding monitoring coverage. And you can’t assume the tagging is 100% correct to pinpoint the correct logs.

In the past, more work was needed to understand the context of log data. What about correlated trace data, host metrics, real-time vulnerability scanning results, or log messages captured just before an incident occurs? This context is vital to understanding issues.

Dynatrace automatically puts logs into context

Dynatrace Log Management and Analytics directly addresses these challenges. First, OneAgent takes care of log autodiscovery. Once logs are selected for monitoring, OneAgent enriches log data with the topological context you need. For example, OneAgent helps you monitor the logs from a Kubernetes environment with automatic enrichment that identifies the right cluster, namespace, container, and pod ID.

Once logs are stored in Dynatrace Grail™, our purpose-built data lakehouse for observability data, the logs are automatically shown in the right context. Finding answers begins with opening the right app for your use case.

Kubernetes logs in context in Dynatrace screenshot

You can easily pivot between a hot Kubernetes cluster and the log file related to the issue in 2-3 clicks in these Dynatrace® Apps: Infrastructure & Observability (I&O), Databases, Clouds, and Kubernetes.

Open a host, cluster, cloud service, or database view in one of these apps, and you immediately see logs alongside other relevant metrics, processes, SLOs, events, vulnerabilities, and data offered by the app.

By eliminating slow and manual correlation, lack of context, and getting visibility into the surrounding data, you reduce the risk of prolonged outages, mean time to repair, and tool sprawl.

Log data in Dynatrace

Get quicker answers

Let’s look at how logs in context can make your teams more effective.

Video thumbnail

Log histograms: Insight into log volumes and patterns

Open one of these Dynatrace Apps and select Logs for any listed entity (host, Kubernetes workload, cloud service, or database instance):

  • Infrastructure & Operations
  • Kubernetes
  • Databases
  • Clouds

You’ll see a histogram chart of log data with various severity levels (such as Error, Info, or Warning) relevant to the selected Dynatrace entity, giving you a clear understanding of log patterns and volumes over time. Is there a sudden spike in errors? A sudden drop in received log data? Depending on which app is in use, one glance at a histogram provides invaluable insight into managing clouds, databases, Kubernetes environments, and infrastructure.

hosts logs in context

Log analytics simplified: Deeper insights, no DQL required

Your team will immediately notice the streamlined log analysis capabilities below the histogram. Jump directly into log insights by selecting a recommended query, for example, to see the errors related to a problem detected by Davis® AI during the selected timeframe. Furthermore, your team can easily access all error logs within the specified timeframe displayed on the histogram or view all logs within that timeframe, all without writing any queries from scratch.

Surrounding logs display: Effortlessly navigate log context

You can see the result after opening a recommended query without leaving an app’s context. Upon expanding a single log entry, all relevant context provided by OneAgent during the ingestion process is displayed, making it easy to expand your analysis to the infrastructure or entity related to the error logs. For a single log record found, you can easily see the surrounding logs.

Look at this example of an online store payment service generating errors. The application owner found error logs related to unsupported credit cards. Select Surrounding logs to view the log messages for the whole transaction, based on the trace ID, that ended up with an error and a failed order.

Surrounding logs

In Infrastructure & Operations, surrounding logs can also be displayed based on other criteria, like the host file or log source from which logs are collected. This allows quick and easy troubleshooting without writing or editing queries.

Logs in context across Dynatrace Apps

  • Infrastructure & Operations leverages advanced AI capabilities that automatically discover and map all components within your infrastructure, including hosts, virtual machines, containers, and cloud instances.
  • Databases offers comprehensive database monitoring capabilities, providing organizations with real-time visibility into the performance and health of their database environments.
  • Clouds is a central hub for monitoring and managing multicloud environments, providing organizations with a unified view of their cloud infrastructure and services.
  • Kubernetes delivers comprehensive monitoring and management capabilities for Kubernetes environments, enabling organizations to ensure the performance, availability, and scalability of their containerized workloads.

Stay tuned for even wider support of log data embedded seamlessly into the context of Dynatrace Apps, and better ways to get answers from logs without writing queries.

See for yourself

Already have a Dynatrace account? See logs in context for yourself in the Dynatrace Playground.

The post Unlock log analytics: Seamless insights without writing queries appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/log-analytics-seamless-insights-without-writing-queries/feed/ 0
Stream logs to Dynatrace with Amazon Data Firehose to boost your cloud-native journey https://www.dynatrace.com/news/blog/stream-logs-to-dynatrace-with-amazon-data-firehose/ https://www.dynatrace.com/news/blog/stream-logs-to-dynatrace-with-amazon-data-firehose/#respond Fri, 03 May 2024 15:25:51 +0000 https://www.dynatrace.com/news/?p=63910 Dynatrace and Amazon Data Firehose

Your cloud logs can provide the root cause of high-impact issues or reveal the details of security incidents. Now, you can integrate an Amazon Data Firehose high-frequency data stream directly with the high-performant Dynatrace Grail™ analytics engine and use the Dynatrace AI-powered observability platform to mitigate issues with minimal impact to your business.

The post Stream logs to Dynatrace with Amazon Data Firehose to boost your cloud-native journey appeared first on Dynatrace news.

]]>
Dynatrace and Amazon Data Firehose

Real-time streaming needs real-time analytics

As enterprises move their workloads to cloud service providers like Amazon Web Services, the complexity of observing their workloads increases. Log data—the most verbose form of observability data, complementing other standardized signals like metrics and traces—is especially critical. As cloud complexity grows, it brings more volume, velocity, and variety of log data.

Managing this change is difficult. Without the ability to see the logs that are relevant to your service, infrastructure, or cloud function—at exactly the right time and in exactly the right format—your cloud or DevOps engineers lose the ability to find the root causes of the issues they troubleshoot. Even the AIOps approach doesn’t cut it if you don’t have proper logs in your observability platform.

Amazon CloudWatch is the most common method of collecting logs across your AWS footprint. As a native tool used by many enterprises, CloudWatch supports a wide range of AWS resources, applications, and services.

Amazon Data Firehose helps stream logs to the right destination

But your SREs and DevOps engineers know CloudWatch is not the terminal destination for data but rather an intermediate station. Their job is to find out the root cause of any SLO violations, ensure visibility into the application landscape to fix problems efficiently and minimize production costs by reducing errors. SREs and DevOps engineers need cloud logs in an integrated observability platform to monitor the whole software development lifecycle.

When trying to address this challenge, your cloud architects will likely choose Amazon Data Firehose. This fully managed native service is indispensable for streaming high-frequency logs collected by CloudWatch.

In some deployment scenarios, you might skip CloudWatch altogether. Take the example of Amazon Virtual Private Cloud (VPC) flow logs, which provide insights into the IP traffic of your network interfaces. VPC flow logs can be used as the source for troubleshooting connectivity issues, implementing security incident investigations, detecting intrusions, or managing access control issues. VPC flow logs can be massive in volume as your cloud deployment footprint grows, and directly streaming these logs with Amazon Data Firehose can be the most cost-effective method.

After configuring Amazon Data Firehose, your teams discover they have completed only the first part of the observability jigsaw puzzle. They also need a high-performance, real-time analytics platform to make that data actionable.

Dynatrace delivers the missing piece for AWS cloud observability with native Firehose integration. This complements our existing AWS logging integrations like S3 log forwarder, Lambda layer log forwarding, or direct log ingest API. These already provide a common integration with AWS log sources. The new Firehose integration removes intermediary components that previously required additional maintenance and provides a direct link from AWS to Grail data lakehouse.

This means high-frequency streamed logs from Firehose can be captured in your Dynatrace environment, automatically processed, stored in Grail for the retention period of your choice, and included in the full observability automation suite of the Dynatrace® platform, apps, and Davis® AI problem detection.

With this out-of-the-box support for scalable data ingest, log data is immediately available to your teams for troubleshooting and observability, investigating security issues, or auditing. As logs are first-class citizens alongside traces, metrics, business events, and other data types, you have an observability platform ready to scale with you in your cloud-native journey.

Easy setup takes just a few steps

Setting up a direct ingest of Firehose log data is quick and easy.

First, you need to generate an API key to ingest logs. In the Dynatrace web UI, go to Access tokens and select Generate new token. Select ingest logs as the scope of the token. Then, generate the token.

Next, go to the AWS console to configure the forwarding of data streams defined in your log groups. Data Firehose stream requires a trusted relationship with CloudWatch through an IAM role. Follow the instructions available in Dynatrace documentation to allow proper access and configure Firehose settings.

Now, you can set up your Firehose stream. The preferred way is to use a CloudFormation template that streamlines and automates the process. See CloudFormation template documentation for details.

Alternatively, you can configure the stream in the AWS web console. Choose Dynatrace as the Destination in the AWS console and complete the other fields with the correct parameters.

Choose Dynatrace as the destination in AWS console.
Figure 1. Choose Dynatrace as the destination in AWS console.

Now, you can view your cloud logs in Dynatrace!

For example, open the Clouds app with integrated logs in the context of your Lambda functions observability for one-click access to error logs.

See logs in context in the Dynatrace Platform, including the relevant logs for this AWS Lambda function.
Figure 2. See logs in context in the Dynatrace Platform, including the relevant logs for this AWS Lambda function.

When Dynatrace Davis AI detects a problem in your environment, you can also see relevant logs streamed via AWS Firehose that are related to the problem. When analyzing a problem, look at the related service, which displays related log data. This lets you jump right to the error that provides details of the problem.

When doing proactive health checks or analysis, you can inspect log data in Notebooks. For example, pick a template to explore data or write your own DQL query and chart incoming error rates from logs streamed via AWS Firehose.

Easily visualize Lambda error log distribution over time with Notebooks.
Figure 3. Easily visualize Lambda error log distribution over time with Notebooks.

Try it out today

Share your experience

We’d love to hear from you. Share your use cases for Amazon Data Firehose integration with the Dynatrace Community.

Stream AWS service logs collected in CloudWatch or directly via Firehose.

The post Stream logs to Dynatrace with Amazon Data Firehose to boost your cloud-native journey appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/stream-logs-to-dynatrace-with-amazon-data-firehose/feed/ 0
Enable full observability for Linux on IBM Z mainframe now with logs https://www.dynatrace.com/news/blog/enable-full-observability-for-linux-on-ibm-z-mainframe-now-with-logs/ https://www.dynatrace.com/news/blog/enable-full-observability-for-linux-on-ibm-z-mainframe-now-with-logs/#respond Wed, 17 Apr 2024 17:25:45 +0000 https://www.dynatrace.com/news/?p=63665 Hosts fetch logs

Complement your hybrid cloud journey with a resilient observability setup with log monitoring on Linux on IBM Z and LinuxONE. Include the familiar mainframe OS into one integrated observability platform and thus eliminate the need for platform-specific component upkeep and management, reduce the risk of prolonged outages, and get full details from logs for troubleshooting.

The post Enable full observability for Linux on IBM Z mainframe now with logs appeared first on Dynatrace news.

]]>
Hosts fetch logs

Mainframe is a strong choice for hybrid cloud, but it brings observability challenges

IBM Z is a mainframe computing platform chosen by many organizations with a hybrid cloud strategy because of its security, resiliency, performance, scalability, and sustainability. With the availability of Linux on IBM Z and LinuxONE, the IBM Z platform brings a familiar host operating system and sustainability that could yield up to 75% energy reduction compared to x86 servers.

That’s why a hybrid cloud scenario, where workloads are shared between public clouds and highly performant mainframe platforms like IBM Z, is a robust and effective strategy.

The challenge for hybrid cloud deployments is maintaining critical observability, which must include the full set of monitoring signals: logs, metrics, and traces. Without combining these signals in a unified AI-powered observability platform, monitoring apps, infrastructure, and troubleshooting issues are nothing more than a patchwork of manual correlation.

Deploying your critical applications on additional host operating systems increases the dependencies for observability. It means maintaining platform-specific observability components or tools, managing security updates, and deploying changes, all of which lead to configuration spread.

This creates a risk that can impact your time to problem resolution in troubleshooting, the effectiveness of AIOps workflows to remediate issues before they affect your end-users, and ultimately your business metrics.

Logs become an integrated part of observability

Dynatrace provides a unified and integrated platform to observe such hybrid cloud deployments, now with added support to monitor logs effortlessly on Linux on IBM Z and LinuxONE.

OneAgent® is a core component of the Dynatrace platform; it enables observability with minimal setup effort while offering extensive and flexible central configuration options. By including logs in your hybrid cloud observability, you have everything you need in one place to make smarter, faster decisions when troubleshooting and measuring the health of your application environments.

You can now seamlessly expand your analysis of the root cause of any problem identified by Davis® AI with logs automatically available in the correct context of hosts, applications, or other identifiers specific to your environment.

Because Dynatrace provides a unified and central place to configure your observability, there is a single place where you manage your log collection for public and private clouds and mainframe components like Linux on IBM Z or LinuxONE.

This makes your log collection policies much more effective and transparent. You can push a filtering change to filter out all unwanted logs from your central Dynatrace environment and apply the change automatically to all your monitored platforms.

It’s also easy to minimize the risk of violating data access policies or regulations by masking sensitive data in logs. By centrally configuring masking rules for sensitive data, you can push the rules out to all of your deployed OneAgents wherever they are deployed to make sure you stay compliant.

Configure log collection across all your hosts

Start by deploying OneAgent for Linux on IBM Z by going to Deploy OneAgent (for earlier versions of Dynatrace and Dynatrace Managed deployments, go to Settings > Deploy Dynatrace), choose Linux as the underlying platform, and s390 as the installer type.

You can now install OneAgent on Linux with s390 architecture.
Figure 1. You can now install OneAgent on Linux with s390 architecture.

Next, set up log ingest. As log monitoring is now available with OneAgent for Linux on IBM Z, a single log ingest rule can cover all your Linux operating systems no matter what architecture is utilized under the hood. This means OneAgents deployed on Linux with s390, ARM, AIX, or x86 are covered.

Go to Settings > Log Monitoring > Log ingest rules and turn on Ingest all logs to start log collection.

Enabling the log ingestion setting applies to deployed OneAgents on Linux on IBM Z and LinuxONE.
Figure 2. Enabling the log ingestion setting applies to deployed OneAgents on Linux on IBM Z and LinuxONE.

Next you can start using logs in your troubleshooting and analysis tasks. For example, on the Dynatrace platform, open the new Infrastructure & Operations app and navigate to any monitored host running on Linux on IBM Z (s390 architecture). You can see the Logs tab for the host, which displays insights about automatically contextualized logs from that host.

Infrastructure & Operations app shows a monitored host with s390 architecture, and the Logs tab shows log data for that host.
Figure 3. The infrastructure & Operations app shows a monitored host with s390 architecture, and the Logs tab shows log data for that host.

You can take your Dynatrace Grail™ analysis of log data further in Notebooks. Start with a query builder to get error logs for your Linux on IBM Z hosts, and continue your exploration with Dynatrace Query Language.

Error logs in Notebooks with distribution chart
Figure 4. Error logs in Notebooks with distribution chart

Start monitoring logs on Linux on IBM Z

What’s next

Stay tuned for an upcoming blog post about log collection in OpenShift for Linux on IBM Z and LinuxONE.

Are you running containerized applications on IBM Z?

The post Enable full observability for Linux on IBM Z mainframe now with logs appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enable-full-observability-for-linux-on-ibm-z-mainframe-now-with-logs/feed/ 0
Bring syslog into Dynatrace using OpenTelemetry to get open source value with enterprise support https://www.dynatrace.com/news/blog/bring-syslog-into-dynatrace-over-opentelemetry-to-get-open-source-value-with-enterprise-support/ https://www.dynatrace.com/news/blog/bring-syslog-into-dynatrace-over-opentelemetry-to-get-open-source-value-with-enterprise-support/#respond Fri, 15 Mar 2024 16:45:45 +0000 https://www.dynatrace.com/news/?p=63075 Fetch logs

Syslog is a standard protocol for system and network device monitoring. Integrating syslog into enterprise observability solutions is tricky due to its strict support and security patching requirements.

The new Dynatrace OTel Collector distribution unlocks the power of syslog and open source community contributions with the power of Dynatrace support and the value of Dynatrace Grail™ to analyze log data from devices at scale.

The post Bring syslog into Dynatrace using OpenTelemetry to get open source value with enterprise support appeared first on Dynatrace news.

]]>
Fetch logs

Getting insights into the health and disruptions of your networking or infrastructure is fundamental to enterprise observability. Syslog is the go-to protocol that delivers infrastructure administrators, network engineers, and security team logs that tell them all they need to know about their systems’ delivery, performance, availability, and security.

Without syslog, you’re blind to what happens on your infrastructure

While syslog is a common way to gain insights into enterprise infrastructure operations, integrating it with other signals into an observability overview is often a painful experience.

Syslog is a protocol with clear specifications that require a dedicated syslog server. This is needed to collect messages across your systems because many different types of devices and applications can produce logs in the syslog format.

However, enterprise adoption at scale typically has much higher requirements for components than for supported features—components must have proper vendor support. Without vendor support, you’re betting your business on goodwill. Even for a supported component, delivering logs from applications and infrastructure to DevSecBizOps workflows requires significant manual configuration.

For example, a supported syslog component must support the masking of sensitive data at capture to avoid transmitting personally identifiable information or other confidential data over the network. Log batching, enrichment, transformation, log source distinction, and application offloading are also regular requirements.

As enterprise environments scale enormously, filtering and dropping data “at the edge” before transmission to a central collection point must be a supported option.

Compliance, retention, archiving, or data governance regulations often require multicasting logs from the original source to multiple destinations, like an observability platform with long-term log storage.

In the end, site reliability engineering (SRE) and security teams need to have data delivered via syslog to their observability platform, in the context of other data types.

Syslog will remain a proven log solution because without understanding why connections are dropped, server starts or stops, or which requests your firewall blocked, your organization runs like a ship where the captain on the bridge has no understanding of what’s going on at the lower levels of the ship. However the challenges in maintaining syslog in a cloud-native era create a maze of requirements that SRE teams and infrastructure administrators must navigate, often finding themselves maintaining multiple tools and components.

This increases the risk of multiple points of failure, adds overhead, and ultimately fractures observability overview with prolonged time needed to recover from potential outages.

Start monitoring syslog using OpenTelemetry under the Dynatrace umbrella of support

OpenTelemetry has been a rising star in the observability landscape and is often a preferred way to achieve end-to-end visibility with a vendor-agnostic footprint. Dynatrace has been a part of the OpenTelemetry journey for years and has contributed to its rise.

With the new Dynatrace OTel Collector distribution, we provide a streamlined and supported way to collect logs using the syslog protocol. This fills all the requirements enterprises have and makes it hassle-free to stream syslog to Grail data lakehouse integrating logs with other observability data.

The Dynatrace OTel Collector for syslog has numerous benefits. Our approach is to understand what components our customers need and value. We then integrate them with our observability platform and offer support, so you don’t have to worry about unsupported bugs or lack of ownership.

We also provide security updates and patches to critical vulnerabilities that may arise in the components. This alleviates the risk of open source components with unpatched vulnerabilities remaining open to exploitation long after they have been revealed.

Ultimately this combination of Dynatrace support and the OpenTelemetry standard gives you the best of both worlds—enterprise-grade software support with open source community contributions.

Dynatrace OTel Collector fits with your existing setup

The new Dynatrace OTel Collector fits nicely into your existing Dynatrace setup to bring in syslog data. Our existing log ingest API already supports your logs using the OpenTelemetry protocol, so you just need to deploy the collector and point your syslog producers to it.

To start using the Dynatrace OTel Collector, take the following steps:

  1. Generate an API token for the OTLP endpoint in your environment.
  2. Find our newly released Dynatrace OTel Collector, deploy it, and configure the exporter with your API key and environment ID.
  3. Configure receivers to enable different log sources for your syslog producers.
  4. Point your syslog sources to the collector and you’re done!

This diagram explains how the components communicate with each other.

This diagram explains how the components of the Dynatrace OTel Collector communicate with each other.

Take a look at this example for configuration. After generating an API token and deploying the collector, configure your instance. You need to configure each component (receiver, optional processor, and exporter) individually in a YAML file and enable them via pipelines. Follow the examples below or refer to Collector configuration documentation.

To point the exporter to your environment’s OTLP endpoint, add the following configuration:

exporters:
  logging:
    verbosity: detailed

  otlphttp/tenant_1:
    endpoint: "https://{your-tenant}.live.dynatrace.com/api/v2/otlp"
    headers:
      Authorization: "Api-Token {your-api-token}"

Next, you can add receivers to your collectors, for example, F5 BIG-IP systems to log to a remote syslog server (version 11.x-17.x). Refer to F5 BIG-IP documentation for detailed and up-to-date instructions regarding remote Syslog configuration. Take a look at Syslog (Dynatrace OTel Collector) in Dynatrace Hub for an example configuration file for the receiver, so you can enable two separate syslog endpoints for F5 and host syslogs. This allows you to differentiate log sources (attribute.device.type) for analysis in Dynatrace.

You can also make the Dynatrace OTel Collector multicast incoming syslog messages to multiple destinations. For example, you can set up exporters for your Dynatrace production environment and sandbox environment:

service:
  pipelines:
    logs:
      receivers: [syslog/f5, syslog/host]
      processors: [batch]
      exporters: [logging, otlphttp/tenant_1, otlphttp/tenant_2]

As a result, you should see logs in Dynatrace with corresponding log.source and device.type attributes:

Logs in Dynatrace with corresponding log.source and device.type attributes

Deploy Dynatrace OTel Collector for syslog now

What’s next

  • Stay tuned for direct syslog ingestion into Dynatrace, which brings syslog endpoints to an Environment ActiveGate, fully configurable from the cluster. This will enable you to use Dynatrace ActiveGate to ingest syslog data.

Go to Syslog (Dynatrace OTel Collector) in Dynatrace Hub to see examples and continue to the installation.

The post Bring syslog into Dynatrace using OpenTelemetry to get open source value with enterprise support appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/bring-syslog-into-dynatrace-over-opentelemetry-to-get-open-source-value-with-enterprise-support/feed/ 0
Enhance data collection with Dynatrace OTel Collector distribution https://www.dynatrace.com/news/blog/enhance-data-collection-with-dynatrace-otel-collector-distribution/ https://www.dynatrace.com/news/blog/enhance-data-collection-with-dynatrace-otel-collector-distribution/#respond Fri, 15 Mar 2024 16:38:36 +0000 https://www.dynatrace.com/news/?p=63069 OpenTelemetry demo app Astronomy Shop

As organizations strive for observability and data democratization, OpenTelemetry emerges as a key technology to create and transfer observability data. OpenTelemetry is gaining popularity because it’s considered a standard, and that’s why it’s a common choice for creating future-proof solutions for years to come. To answer the growing demand for OpenTelemetry, Dynatrace is proud to […]

The post Enhance data collection with Dynatrace OTel Collector distribution appeared first on Dynatrace news.

]]>
OpenTelemetry demo app Astronomy Shop

As organizations strive for observability and data democratization, OpenTelemetry emerges as a key technology to create and transfer observability data. OpenTelemetry is gaining popularity because it’s considered a standard, and that’s why it’s a common choice for creating future-proof solutions for years to come.

To answer the growing demand for OpenTelemetry, Dynatrace is proud to announce the release of the Dynatrace OTel Collector distribution. This collector, fully supported and maintained by Dynatrace, is entirely open source. Before we get into the specifics, let’s first recap the benefits OpenTelemetry offers and why using collectors is a best practice.

Understanding OpenTelemetry

OpenTelemetry is an open, vendor-neutral standard for creating, collecting, and transferring telemetry data, like traces, metrics, and logs. Developers and operators can gain insights into their applications and infrastructure without fear of vendor lock-in because OpenTelemetry is fully open source and owned by CNCF. The OpenTelemetry project is supported and maintained by representatives from Microsoft, Google, Amazon, and many others, including Dynatrace.

Why do I need an OpenTelemetry collector?

As the name suggests, an OpenTelemetry collector gathers data from multiple sources and sends it to observability backends, like Dynatrace, for analysis. A collector helps developers control their telemetry data streams for each signal. Different data streams can be directed to different backends or even multicast to multiple backends simultaneously. The configuration is highly flexible in solving various user needs.

A collector is also a powerful component for data processing. It removes the burden of managing retries, batching, and sampling from monitored applications, which can reduce the CPU and memory requirements of applications. A collector can also transform and enrich the data with additional context. For example, in a Kubernetes environment, a collector can automatically attach metadata about pods and namespaces to all observability data. This ensures that application telemetry is contextualized with the infrastructure, enabling the observability backend to link the application and infrastructure for enhanced insights and root cause analysis.

From a user perspective, a collector can also serve as an open source platform that can be extended with custom components. You can create internal collector components of your own, for example, to receive telemetry data in a special format or to process it in a certain way. Using a collector as a telemetry processing platform can be much easier than creating an entirely new application.

Why the Dynatrace OTel Collector

The OpenTelemetry community releases different distributions of the collector, many of which our customers use to send OpenTelemetry data to Dynatrace. So, why should you consider using the Dynatrace Otel Collector? Quite simply, support and stability.

We have seen many customers identify collectors as a potential solution for their needs. Still, they haven’t been able to deploy collectors in production due to the lack of support. Deploying an open source component without external support and the needed expertise is undoubtedly a risk. That’s why we provide Dynatrace customers with a Dynatrace-supported solution.

The Dynatrace Otel Collector comes with collector components that have been verified by Dynatrace for seamless operation. This removes the burden of manually validating each component and use case. To further help you with your collector journey, we publish configuration examples of typical Dynatrace use cases and best practices to provide a good starting point.

The Dynatrace Otel Collector includes components that we know run stably in production, which means we can offer full Dynatrace support. At the same time, innovations from the OpenTelemetry community can be added to the Dynatrace Otel Collector only after they are mature enough and have proven their stability.

Deployment and the typical use cases

The Dynatrace Otel Collector can be deployed on Kubernetes or Docker using a provided container image or directly on a host with the published binary. For more details, refer to our Dynatrace OTel Collector deployment guide.

In the initial release, the Dynatrace Otel Collector comes with components for:

Additionally, the Dynatrace Otel Collector includes a rich toolset for data processing to enrich, filter, transform, sample, and batch. The complete list of the components is available in our GitHub repository.

Dynatrace OTel Collector diagram with telemetry sources

What’s next

After the initial release, we’ll continue enhancing the Dynatrace OTel Collector with new features and capabilities to make it even easier to integrate with Dynatrace. We also intend to offer more automated methods to deploy the collector with pre-configurations. So, stay tuned.

As a significant contributor to the OpenTelemetry project, Dynatrace remains committed to working with the community and other vendors to enrich its capabilities and make it user-friendly for everyone.

Deploying the Dynatrace Otel Collector takes only minutes and uses the tooling you already know: standalone binary, Docker image, Kubernetes Operator, Helm chart, or a standard manifest file. The configuration maps one-to-one with the collector distributions from the OpenTelemetry community.

The post Enhance data collection with Dynatrace OTel Collector distribution appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enhance-data-collection-with-dynatrace-otel-collector-distribution/feed/ 0
Dynatrace log collection for ARM unlocks power-efficient architecture for your enterprise https://www.dynatrace.com/news/blog/dynatrace-log-collection-for-arm-unlocks-power-efficient-architecture-for-your-enterprise/ https://www.dynatrace.com/news/blog/dynatrace-log-collection-for-arm-unlocks-power-efficient-architecture-for-your-enterprise/#respond Tue, 05 Dec 2023 16:33:49 +0000 https://www.dynatrace.com/news/?p=60963 log collection for ARM

ARM (Advanced RISC Machine) architecture is finally mature enough to offer enterprise customers its promised energy efficiency and powerful performance boost. Implementing the Dynatrace® observability platform with metrics, traces, and logs gives you full visibility into your ARM architecture environment. It unlocks AI-powered problem detection with root cause identification and log-based analytics. Leverage the Dynatrace “secret sauce" by installing OneAgent® on ARM-based hosts to enable automatic log discovery and ingest logs to Grail™ data lakehouse at a massive scale.

The post Dynatrace log collection for ARM unlocks power-efficient architecture for your enterprise appeared first on Dynatrace news.

]]>
log collection for ARM

Without observability, the benefits of ARM are lost

Over the last decade and a half, a new wave of computer architecture has overtaken the world. ARM architecture, based on a processor type optimized for cloud and hyperscale computing, has become the most prevalent on the planet, with billions of ARM devices currently in use. This growth was spurred by mobile ecosystems with Android and iOS operating systems, where ARM has a unique advantage in energy efficiency while offering high performance.

While ARM processors have been humming in consumers’ pockets for over a decade, enterprise IT adoption has been slower. Legacy data center infrastructure and software support have kept all the benefits of ARM at, well… arm’s length. Still, we at Dynatrace speak to customers who recognize this value and want to implement it.

Energy efficiency and carbon footprint outshine x86 architectures

The first clear benefit of ARM in the enterprise IT landscape is energy efficiency. This is a crucial factor for all data centers, cloud or managed, where power consumption and cooling costs are reflected in the total cost of ownership. Traditional x86 architecture needs more power, so switching to ARM can offer a clear advantage.

Energy efficiency is coupled with the total enterprise carbon footprint. As organizations look to take ownership of their total ecological footprint and help mitigate climate change, it’s critically important for organizations to measure, monitor, and reduce their IT carbon footprints. Initiatives like the Carbon Impact app can be used to measure the footprint of monitored ARM-based hosts compared to x86 hosts.

Huge performance leaps in recent years

The top priority is often performance, where ARM resources have improved significantly. You would have been unimpressed if you evaluated public cloud vendor ARM-based offerings five or more years ago. But take a look at the latest iterations of, for example, AWS Graviton2, which delivered a 40% price/performance boost, and Graviton3, which had an additional 27% price/performance improvement over Graviton2.

These Improvements and other optimizations show that mature service offerings can rely on ARM architecture. You’re no longer required to use a single offering or choose from a few instance families; Graviton includes general-purpose and accelerated-computing offerings, plus compute-, memory-, and storage-optimized instances.

No observability, no gains

All these factors have made ARM an attractive computing architecture for innovative companies. However, the lack of end-to-end integrated observability with full support for log collection is a clear blocker for many organizations looking to adopt ARM-based services.

Without an observability platform to collect and process signals, and provide AI-powered answers for problem detection, root cause identification, and impact analysis, the migration to ARM remains only a roadmap item for many AIOps teams.

Even if some part of your codebase can be instrumented to collect observability data, having all three signal types (metrics, traces, and logs) is crucial. Without collecting logs from the observed platform in a scalable AI-powered data lakehouse like Grail, it’s more of a challenge to identify the root cause of problems and provide details for troubleshooting or security incidents.

Having no access to logs or relying on a siloed approach to logging poses a risk of blind spots in your IT landscape, prolonged outages, and an increase in Mean Time To Identify (MTTI) for incidents, all of which can nullify the benefits of ARM.

Hassle-free ARM deployment with automated log collection

With automated log observability, the Dynatrace platform enables your AIOps team to reap the full benefits of ARM architecture. Dynatrace OneAgent can automatically detect common logs for ARM-based hosts and ingest them on a massive scale out of the box.

As soon as you install just a single OneAgent on a host, for example, an ARM64 (AArch64) based Linux host, the OneAgent scans and autodiscovers logs every 60 seconds.

After installation, you can control OneAgent behavior using a powerful central configuration mechanism in your Dynatrace environment. You can tune the granularity of OneAgent filtering for each host to meet your requirements for discovery and ingestion, share configuration details with other hosts in the same host group, or apply global settings for your whole environment.

As OneAgent collects logs from your ARM-based host, it automatically ties the discovered data to your environment topology. In this way, log data is always associated with the host, service, or other entity that generated it.

AI-powered problem detection can then connect the log data, which often contains the source of truth, to detected issues in your environment, thereby speeding up troubleshooting and remediation by orders of magnitude compared to manual correlation.

In addition, the integration with auto-baselined metrics and automatic connection between distributed traces and logs puts you in complete control of the software you run on ARM-based architectures.

Immediately see the root cause in ARM host logs

You can enable the full value of Dynatrace with log monitoring on ARM in just a few steps.

  1. Install OneAgent version 1.269+ on an ARM host (Go to Apps In the Dynatrace web UI and search for Deploy OneAgent to access the installer). For full details, see OneAgent installation on Linux.
  2. Following OneAgent installation, you can verify on the host page that the ARM 64-bit instruction set is used.
    OneAgent installation
    Next, go to Settings > Log monitoring > Set up log ingest and review the log ingestion rules in the Hosts section. For the simplest quickstart, select Show rules.
    Set up log ingest
  3. Turn on the rule [Built-in] Ingest all logs to enable all rules and gain complete visibility into the host.
    Ingest all logs

Now you have the additional value of log data informing your Dynatrace root cause analysis should a problem arise on this host.

In the following example, a synthetic monitor is set up for webpage uptime monitoring. Dynatrace Davis® AI has automatically discovered a problem on the host, and the root cause analysis points to a PHP web service as the source of the problem.

Problem page

The problem page gives you a direct link to the service causing the disruption, and you can quickly access all related logs.

Problem page

Error logs for the process reveal a PHP syntax error with code-level information that caused the outage.

Problem page

As you can see, with the Dynatrace AI-powered observability platform at your disposal, you can benefit from resource-efficient ARM architectures in your organization without the risk of blindspots or guesswork when you encounter incidents or outages.

Start monitoring ARM and see what’s next

Coming next

We’re working on making Dynatrace log monitoring even easier to use at scale with a focus on areas highlighted by you and other customers:

  • Are you ingesting logs across multiple sources with a wildcard (*) ingest rule? Soon, you can enhance log records with custom file names or log path attributes for later analysis in Dynatrace.
  • Get automatic log ingest recommendation rules relevant to your environment based on the log sources discovered by OneAgent.
  • Tools to troubleshoot the status of your log sources and instructions for resolving issues in cases where a log file is inaccessible.
  • See the complete list of log sources on each host page, even if the sources are tied to another process group.

The post Dynatrace log collection for ARM unlocks power-efficient architecture for your enterprise appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-log-collection-for-arm-unlocks-power-efficient-architecture-for-your-enterprise/feed/ 0
Complete Kubernetes observability with logs in topology context https://www.dynatrace.com/news/blog/complete-kubernetes-observability-with-logs-in-topology-context/ https://www.dynatrace.com/news/blog/complete-kubernetes-observability-with-logs-in-topology-context/#respond Fri, 25 Aug 2023 15:28:19 +0000 https://www.dynatrace.com/news/?p=59373 complete Kubernetes observability with logs in context

Kubernetes is the open source container orchestration system that many companies use to run containerized workloads across hybrid cloud and on-premises environments. Dynatrace Log Management and Analytics can now maintain the complete context of your Kubernetes (K8s) architecture and platform logs.

The post Complete Kubernetes observability with logs in topology context appeared first on Dynatrace news.

]]>
complete Kubernetes observability with logs in context

Kubernetes workload management is easier with a centralized observability platform

When deploying applications with Kubernetes, the configuration is flexible and declarative, allowing for scalability. However, due to the distributed nature of Kubernetes, it can be difficult to understand overall deployment health and the status of Kubernetes clusters.

To properly monitor Kubernetes clusters and containers, it’s necessary to have access to relevant logs. Although K8s doesn’t provide a specific logging mechanism, the underlying container runtime collects the container logs in K8s. It makes them available for a log analytics platform to gain automated, contextual, and actionable insights into the services and underlying platforms. The complexity of hybrid environments with multiple virtual machines and cloud solutions like AWS EKS, Azure AKS, or GCP GKE with hundreds of containers and their constantly changing lifecycles creates challenges for app owners, developers, SREs, and infrastructure owners. The Kubernetes community aims to prioritize the understanding of workload health in the context of topology and provide the ability to troubleshoot deployed workloads quickly after deployment.

Managing Kubernetes application logs with Dynatrace is a seamless process

By following a few simple steps to deploy the OneAgent daemon set using Dynatrace Operator, It’s easy to collect logs and gain complete observability within the context of your topology.

It’s recommended that Kubernetes applications write logs to standard Linux streams, namely standard output (stdout) and standard error (stderr). These logs are gathered by Kubernetes and saved as files on the node. Dynatrace OneAgent® can automatically identify logs on the Kubernetes node and link them to their respective pods using the pod UID from the file path.

Furthermore, Dynatrace OneAgent enhances these logs by including Kubernetes metadata such as namespace, node, workload, pod, container, and more. More detailed information about the included metadata can be found in Dynatrace Documentation.

This additional metadata helps link logs with the entity models of Kubernetes clusters, namespaces, workloads, and pods. Logs can then be located in the Kubernetes entity model on the Dynatrace platform, enhancing the observability of your infrastructure and workloads.

Dynatrace OneAgent has an internal mechanism that ingests only necessary log data. Within your Dynatrace tenant, various options are available for filtering logs at the source, masking sensitive data, and selecting relevant logs by matching Kubernetes-specific values such as K8s container name, K8s deployment name, and K8s namespace name. You can filter logs based on their content, source, or process technology.

With these built-in mechanisms, you can control the ingested log sources and the associated monitoring costs.

Kubernetes

If a log shipping solution is already in place, such as Fluent Bit, Fluentd, or Logstash, it’s possible to utilize the Logs Generic Ingest API endpoint to easily send and analyze logs in Dynatrace. This feature is readily available on both the Dynatrace tenant and Environment ActiveGate. It’s also a great option for situations where an application writes logs inside pods or if serverless k8s deployments, such as AWS Fargate, are utilized.

Gain log visibility within Kubernetes workloads: a practical example

Dynatrace provides out-of-the-box alerting for K8s clusters and workloads by leveraging Dynatrace Davis® AI. In the example below, Dynatrace detected unexpected out-of-memory kills for pods of a workload called cartservice in the namespace online-boutique.

Problem card showing an out-of-memory error
Figure 2. Problem card showing an out-of-memory error

As Dynatrace provides all observability signals in context, troubleshooting can easily be conducted in the context of the respective Kubernetes workload, which drastically reduces the time required to pinpoint the root cause of issues.

Video thumbnail

In this example, the root cause can easily be determined by further analyzing the Kubernetes events and logs for the cartservice workload. Davis AI automatically detected out-of-memory kills, and the specification of this workload was changed within the most recent deployment release.

Error logs provide further valuable insights

As the Redis port was modified, the cartservice is no longer able to connect to the Redis instance, which led to out-of-memory kills for this workload. To fix this problem and minimize MTTR, we can immediately reach out to the team in charge of this workload by leveraging ownership information from this workload. This process can easily be automated for different types of problems.

Service ownership details
Figure 3. Service ownership details

In addition, Dynatrace offers powerful log analytics in the Dynatrace Log Viewer. You can drill down by K8s cluster, namespace, workload, pod, or filter for a certain severity level or log pattern, just to name a few examples. Every log line is enriched with Kubernetes metadata and, if you use Dynatrace Application Observability and/or Real User Monitoring, each log line is automatically linked to traces and/or user sessions to provide end-to-end observability in context.

Log Viewer
Figure 4. Log Viewer

Begin your logging journey effortlessly at scale

Analyzing log data is crucial to avoiding potential release issues, allowing you to take proactive measures to prevent any negative impact on users. To get started, just dive into the documentation!

Please share your feedback and comments in the Dynatrace Community. Kubernetes support requirements are reflected in our roadmap with items like Kubernetes labels support, which will help you manage your environment in Dynatrace the same way you manage your Kubernetes cluster, and our planned investments related to short-living pods and K8s serverless cloud deployments.

The post Complete Kubernetes observability with logs in topology context appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/complete-kubernetes-observability-with-logs-in-topology-context/feed/ 0
Accelerate your cloud journey with Dynatrace observability for AWS S3 logs https://www.dynatrace.com/news/blog/accelerate-your-cloud-journey-with-dynatrace-observability-for-aws-s3-logs/ https://www.dynatrace.com/news/blog/accelerate-your-cloud-journey-with-dynatrace-observability-for-aws-s3-logs/#respond Tue, 27 Jun 2023 14:54:35 +0000 https://www.dynatrace.com/news/?p=58359 observability for AWS S3 logs

Dynatrace accelerates enterprise observability, troubleshooting, security, and automation use cases that rely on log data from Amazon Web Services S3. AWS S3 is a popular storage location for AWS services and third party solutions, which can now be more easily integrated with Dynatrace Grail™ data lakehouse. This allows you to answer any question at any time using log data while avoiding the complexity of storage management.

The post Accelerate your cloud journey with Dynatrace observability for AWS S3 logs appeared first on Dynatrace news.

]]>
observability for AWS S3 logs

Logs complement metrics and enable automation

Cloud practitioners agree that observability, security, and automation go hand in hand. The increasing complexity of cloud service architectures requires a rock-solid understanding of the activity, health status, and security of cloud services. Logs complement out-of-the-box metrics and enable automated actions for responding to availability, security, and other service events.

Many AWS services and third party solutions use AWS S3 for log storage. We hear from our customers how important it is to have a centralized, quick, and powerful access point to analyze these logs; hence we’re making it easier to ingest AWS S3 logs and leverage Dynatrace Log Management and Analytics powered by Grail.

Centralized log management for scalable ingestion into Grail

As AWS S3 proves to be the preferred way of storing cloud logs, enterprise customers face mounting challenges in putting S3 log data to use. Dynatrace Log Management and Analytics powered by Grail enables you to get answers from logs with any query at any time. However, as a first step, logs from S3 need to be ingested into Dynatrace Grail.

To date, some customers have used open source or community-backed components to forward logs from S3 to Dynatrace. A Dynatrace S3 log forwarder has been available for some time to early adopters, with community support only. While these are great examples of innovation and the power of the community, enterprises often require the type of continuous support and maintenance that only comes with official software.

Another painful blocker is the need for more support for multiple S3 accounts and AWS regions. Some enterprise customers use over a thousand accounts for cloud services, which dramatically increases complexity and overhead within production environments. If an organization operates in multiple geographies and AWS regions, they can centralize logs in a regional S3 bucket, as AWS services send log data to S3 buckets in the same region where they run.

Because data context is missing for logs, it’s slow or even impossible to build causal relationships in your observability data. Most importantly, it’s impossible to establish relationships between infrastructure and application events, business impact, and real user events.

Without the ability to connect logs in S3 with Dynatrace, more expensive or cumbersome alternatives are often used, which slow down troubleshooting and have a lasting business impact.

Ingest logs from AWS S3 with one forwarder

Dynatrace now has a solution for forwarding logs from AWS S3 to its industry-leading log analysis platform, providing enterprise-level support. This makes S3 logs a robust and reliable way of collecting logs, forwarding them to Dynatrace, analyzing log data via DQL or in apps, and thus putting log data to use in solving your observability and security use cases.

The AWS S3 log forwarder can ingest any text- or JSON-formatted logs, which unlocks not only log data from AWS services but also from common third-party logs. Service providers such as Akamai, Fastly, Netlify, open source tools like Apache Airflow, and many others have built-in log delivery integration with S3. The Dynatrace AWS S3 log forwarder is designed to be extensible, so you can customize ingestion settings to fit your specific use cases as well as add custom attributes to your logs.

We designed the Dynatrace AWS S3 log forwarder to scale up to meet the demands of our largest customers, some of whom operate thousands of AWS accounts. The model of deploying one forwarder per AWS region and AWS account doesn’t scale well, so the Dynatrace AWS S3 log forwarder has built-in support for log ingestion from multiple AWS Accounts and AWS regions with a single log forwarder deployment.

The S3 log forwarder also keeps the metadata of log messages, so the originating AWS account and region as well as service-specific attributes like resource ID are preserved. This makes it possible to tie log messages back to the apps, infrastructure, and cloud services where they originated, and enables the unified observability of the Dynatrace platform.

Use Notebooks on the Dynatrace platform to analyze logs from AWS Application Load Balancer.
Use Notebooks on the Dynatrace platform to analyze logs from AWS Application Load Balancer.

As the cloud footprint of a company grows, so does its log data volume. This is why Dynatrace AWS S3 log forwarder throughput is aligned with the ingest volume of Grail data lakehouse to support your growing data needs.

After you ingest logs into Grail, you can put the data to use with exploratory analytics in Notebooks or transform complex data into easy-to-understand visualizations using Dashboards.

Build a custom dashboard on the Dynatrace platform to instantly visualize AWS Application Load Balancer logs.
Build a custom dashboard on the Dynatrace platform to instantly visualize AWS Application Load Balancer logs.

Set up log forwarding on AWS S3

Dynatrace Amazon S3 log forwarder is an AWS Lambda function that supports out-of-the-box parsing and forwarding of logs for the following AWS Services:

Additional context for these and other AWS services can be covered thanks to a built-in parsing mechanism. This is achieved either by Dynatrace AWS S3 forwarder or log processing mechanisms in Dynatrace.

The log forwarder sends the data to the generic log ingest API in your Dynatrace SaaS tenant for Grail analysis. This is another use case where S3 log ingestion can be used to address a wide range of use cases.

Get started today

All the information you need to get started is listed below:

Do you work extensively with AWS S3 logs?

If so, stay tuned for more news about direct AWS Kinesis Data Firehose configuration in AWS console. Or explore the recently introduced support for AWS Lambda logs.

The post Accelerate your cloud journey with Dynatrace observability for AWS S3 logs appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/accelerate-your-cloud-journey-with-dynatrace-observability-for-aws-s3-logs/feed/ 0
Stay in control of your data retention with Dynatrace Grail—from 10 days to 10 years https://www.dynatrace.com/news/blog/stay-in-control-of-your-data-retention/ https://www.dynatrace.com/news/blog/stay-in-control-of-your-data-retention/#respond Fri, 28 Apr 2023 08:00:16 +0000 https://www.dynatrace.com/news/?p=57284 Database graphic

Managing observability and business-data storage is essential to getting data-driven answers and setting up automation workflows. Traditionally, these efforts have led to compromises in cost, business requirements, and compliance with applicable regulations. And relying on an archive-and-retrieve solution isn’t an option because it’s slow and expensive to get value from your data. Thankfully, the new custom buckets in the Dynatrace Grail™ data lakehouse keep you in control of your data, make your data available at all times, and abolish data management overhead.

The post Stay in control of your data retention with Dynatrace Grail—from 10 days to 10 years appeared first on Dynatrace news.

]]>
Database graphic

Optimize cost and availability while staying compliant

Observability data like logs and metrics provide automated answers, root cause detection, and security issues.

Customer decisions about data retention are often determined by important security, privacy, and legal issues. Customers must comply with internal and external policies and regulations that might demand them to keep specific data stored for a minimum period of time (for example, audit logs). However, the opposite is also true—in some cases data must be deleted after a certain period of time. This is the case when a company no longer has legal grounds to retain its customer data, as outlined in privacy protection regulations.

Often customers make business decisions about data retention based on the value they get from keeping historical data and the associated data retention costs. This means compromising between keeping data available as long as possible for analysis while juggling the costs and overhead of storage, archiving, and retrieval. For example, suppose data has to be retained for a longer period because of legal or business reasons. In such a case, the data is archived in cold storage where it can only be accessed for analysis following a delay, re-ingestion into a log analysis tool, and reindexing to prepare the data for analysis.

Ultimately this leads to a lose-lose situation for customers—they have to pay for and maintain data storage but they can’t get answers from their data quickly and effortlessly when needed.

Grail gives you control and the answers from data

With Grail, Dynatrace provides control over data retention and access policies for granular portions of data called “buckets.” This allows you to design data management and retention policies based on individual requirements, starting from days of retention up to a decade.

By introducing control over data retention, Dynatrace doesn’t impose any additional complexities. Even with the flexibility of buckets, there is no additional overhead of data storage management, no archiving, no retrieval from archives, and no performance degradations when using retained data for answers.

The price of data retention is always transparent and uniform, based on the number of days the data is retained, with no hidden fees for managing data. The same applies to querying data with transparent pricing based on read-data volume, with no extra costs for querying older data.

Use buckets for any use case in a secure way

When using Log Management and Analytics or Business Observability with Grail, you can create custom buckets with specified data-retention periods. For example, you can route incoming log data to a specific bucket so selected team members can access it.

App developers might need to read logs from their environment for debugging purposes, but only for a specific timeframe. With Grail, it’s easy to create a bucket with ten days of retention time and provide all developers access to the data.

Infrastructure teams may need to work with host logs from recent months or quarters. To do this, infrastructure logs can be routed to a bucket with a retention period of three months to a year.

Local regulation often requires that security or audit logs be retained for 7 to 10 years. Such logs can be collected in a bucket with the required retention period, with only the security operations team having access to the logs.

A bucket can be wiped if, at any point in time, there is a need to delete the data stored in it. The reasons for this can vary from a changing business justification to data-privacy regulations. It’s also possible to extend or shorten a bucket’s retention period, which impacts how long existing data in a bucket is stored.

To support configuration-as-code for enterprise environments, creating, updating, and deleting data buckets in Grail is available through an API endpoint. This allows you to create new buckets, change the retention period of existing buckets, or delete buckets via an API call.

Bucket management follows a strict permission policy approach, where only users with corresponding permissions can create, update, or delete buckets. Every API call is saved in audit logs to document the complete picture of activities in your environments.

Get value from your data with the Dynatrace Grail today

The post Stay in control of your data retention with Dynatrace Grail—from 10 days to 10 years appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/stay-in-control-of-your-data-retention/feed/ 0
Three smart log ingestion strategies in Dynatrace https://www.dynatrace.com/news/blog/three-smart-log-ingestion-strategies/ https://www.dynatrace.com/news/blog/three-smart-log-ingestion-strategies/#respond Thu, 15 Dec 2022 20:58:53 +0000 https://www.dynatrace.com/news/?p=55224 AppEngine: Create custom apps for data insights

Getting precise answers from log monitoring platforms gets challenging as cloud environments expand and grow more complex. Here are three log ingestion strategies to achieve scale in the Dynatrace platform—without OneAgent.

The post Three smart log ingestion strategies in Dynatrace appeared first on Dynatrace news.

]]>
AppEngine: Create custom apps for data insights

While many organizations have embraced cloud observability to better manage their cloud environments, they may still struggle with the volume of entities that observability platforms monitor. The key to getting answers from log monitoring at scale begins with relevant log ingestion at scale.

Engaging the automatic instrumentation of the Dynatrace OneAgent makes log ingestion automatic and scalable. However, our customers often have set up multiple other log ingestion methods. This flexibility enables logs from diverse environments and established configurations to complete the observability picture for automated troubleshooting and monitoring in Dynatrace.

In this blog, we share three log ingestion strategies from the field that demonstrate how building up efficient log collection can be environment-agnostic by using our generic log ingestion application programming interface (API).

As with all other log ingestion configurations, these examples work seamlessly with the new Log Management and Analytics powered by Grail that provides answers with any analysis at any time.

Log ingestion strategy no. 1: Welcome syslog, with the help of Fluentd

Syslog is a popular standard for transporting and ingesting log messages. Typically, these are streamed to a central syslog server. One option is to install OneAgent on that syslog server, which automatically discovers, instruments and sends the log data to the Dynatrace platform.

But there are cases where you might be limited in setting up a dedicated syslog server with OneAgent because of environment architecture or resources. Yet observability into syslog data on Dynatrace would help you monitor and troubleshoot infrastructure.

This is where it is prudent to configure syslog producers to send data to a log shipper like Fluentd.

What is Fluentd?

Fluentd logo for log ingestion and log monitoring

Fluentd is an open source data collector that decouples data sources from observability tools and platforms by providing a unified logging layer. Fluentd is known for its flexibility and is also highly scalable, which makes it a good choice for high-volume environments.

How does Fluentd work with Dynatrace?

Setting up the flow from syslog over Fluentd to Dynatrace takes three steps. First, point the syslog daemon to the Fluentd port by adding the following line to the syslog daemon configuration file:

*.* @@<fluentd host IP>:5140

*.* instructs the daemon to forward all messages to the specified Fluentd instance listening on port 5140 and <fluentd host IP> needs to point to the IP address of Fluentd.

As a second step, enable Fluentd to accept incoming syslog messages with the in_syslog plugin. Set up the configuration on the same port as specified for source data, in this example 5140.

Lastly, use the open source Dynatrace Fluentd plugin, which uses generic log ingestion. Just find the API token for log ingest API on your SaaS environment or your own Active Gate setup.

Now you should see log messages coming into the Dynatrace log viewer.

Log ingestion strategy No. 2: Point an existing log shipper to the generic Dynatrace ingest

Another common scenario is an environment where you have already invested a do-it-yourself or other log shipper solution. After spending time and budget on the tooling and configuration, it may be unwise to undo this custom work, despite the automatic instrumentation of the Dynatrace OneAgent. Although you preserve your custom work this way, it is a siloed approach for logs, which means you’ll miss out on the integrated observability and automated alerting of Dynatrace.

If that existing solution supports sending log data to an external HTTP endpoint, you can address log silos by integrating with Dynatrace generic ingest with minimal hassle.

To illustrate the solution, let’s look at how to configure log ingestion with the log shipper Cribl.

What is Cribl?

Cribl Stream logo for log ingestion and log monitoring

Cribl is a data operations platform that enables users to collect, route, transform, analyze, and act on data in real time. It provides a unified platform for handling every aspect of data operations, from collecting data to routing and transforming it. Cribl also allows users to orchestrate custom pipelines for their data to gain insights and take action on that information. As a data output, or what it calls a Cribl Stream destination, you can configure an HTTP endpoint.

How does Cribl work with Dynatrace?

The main part of the setup involves creating the configuration for the specific log shipper at hand—in this case, Cribl Stream.

In Cribl’s configuration, open “Data/Destinations” and find “Webhook.” Create a new webhook destination with a name of your choosing (for example, your Dynatrace environment ID, and provide the URL for the webhook). For a Dynatrace SaaS environment, this is the following:

https://{your-environment-id}.live.dynatrace.com/api/v2/logs/ingest

This points the data stream to your Dynatrace environment’s generic ingest.

But in Cribl’s case, you should provide two more settings under “Configure/Advanced Settings/Extra HTTP Headers.” Add two new headers with the following names and values:

  1. To authorize the request, add the header “Authorization” and provide the value Api-Token dt0c01.{your-token-here} where {your-token-here} is an API token with ingest logs scope.
  2. Then add a header “Content-Type” and provide the value “application/json; charset=utf-8
log ingestion, log management screenshot
Example configuration in Cribl of posting logs to Dynatrace API.

After committing and deploying the Cribl changes, you can select the newly created Dynatrace destination as the default destination for your logs. And just like that, all log data already collected by the existing shipper is being sent to Dynatrace for monitoring, analysis, alerting, and all other tasks.

Log ingestion strategy No. 3: Ingest AWS Fargate logs with Fluent Bit

Ingesting and working with Kubernetes logs in Dynatrace helps to provide a comprehensive view of application performance from the infrastructure layer to the application layer. The common approach for Kubernetes logging is to deploy OneAgent in the environment, where it auto-discovers log messages written to the containerized application’s stdout/stderr streams.

But not all environments, configurations, or privileges are created equal. One recurrent challenge is collecting Kubernetes logs if you’re limited in installing OneAgent because of technical or architectural restrictions.

In the case of AWS serverless container compute engine Fargate, for example, where OneAgent log collection is not supported, we recommend using Fluent Bit log forwarder.

Let’s take this example of AWS Fargate. AWS includes a log router called FireLens for Amazon ECS and AWS Fargate services, which gives you built-in access to FluentD and Fluent Bit. We covered FluentD support previously. Now let’s take a look at how to set up Fluent Bit.

What is Fluent Bit?

Fluent Bit logo - for log ingestion and log management

Fluent Bit is an open source and multiplatform log processor and forwarder that allows you to collect data/logs from different sources, unify and send them to multiple destinations and is fully compatible with Docker and Kubernetes environments.

When choosing between Fluentd or Fluent Bit shippers, the Fluent Bit is the preferred solution when resource consumption is critical because it is a lightweight component.

While Fluent Bit has configurable HTTP output, in this example, we look at the AWS Fargate context, where FireLens makes it easy to set up Fluent Bit more quickly.

Ingest AWS Fargate logs with Fluent Bit

When creating a new task definition using the AWS Management Console, the FireLens integration section makes it easy to add a log router container. Just pick the built-in Fluent Bit image.

Next, edit the container in which your app-generating logs are running. In the “Storage and Logging” section, select “awsfirelens” as the log driver.

The settings for the log driver should point to the log ingest API of your SaaS tenant. Note that you normally need to provide two headers for Fluent Bit: content type and authorization token. As FireLens supports only one header, you can pass the token as part of the URL. Your configuration for AWS FireLens should have the following:

  • Name: http
  • TLS: on
  • Format: json
  • Header: Content-Type application/json; charset=utf-8
  • Host: {your-environment-id}.live.dynatrace.com
  • Port: 443
  • URI: /api/v2/logs/ingest?api-token={your-API-token-here}
  • tls.verify: Off
  • Allow_Duplicated_Headers: false
  • match: *
  • json_date_format: iso8601
  • json_date_key: timestamp

To avoid publishing the token in plaintext, use AWS Secrets Manager to manage the token.

As your application starts publishing logs, you can view them in Dynatrace.

Read more about streaming logs to Dynatrace with Fluent Bit from our documentation.

More methods for log ingestion

These are just some of the ways you can ingest logs into the Dynatrace platform without using OneAgent. You’ll soon have even more methods for log ingestion into Dynatrace, for example:

  • Automated OpenTelemetry logs acquisition and processing
  • Syslog endpoint in your environment as a component on a private ActiveGate
  • Dynatrace Fluent Bit output plugin for out-of-the-box integration

Want to share your experiences with log ingestion? Head to the Dynatrace Community Feedback channel to share your thoughts with other users.

State of Log Management 2026

Download the report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

The post Three smart log ingestion strategies in Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/three-smart-log-ingestion-strategies/feed/ 0
Getting answers from data starts with automated log acquisition, at any scale https://www.dynatrace.com/news/blog/getting-answers-from-data-starts-with-automated-log-acquisition-at-any-scale/ https://www.dynatrace.com/news/blog/getting-answers-from-data-starts-with-automated-log-acquisition-at-any-scale/#respond Fri, 14 Oct 2022 16:28:49 +0000 https://www.dynatrace.com/news/?p=53819 What is FinOps?

Dynatrace enables configuration-as-code for logs, making data acquisition, filtering, masking, and anonymizing scalable in enterprise environments. Orchestrating a fleet of OneAgents from a central configuration point removes error-prone manual work for DevOps and SRE teams across dynamic IT landscapes.

The post Getting answers from data starts with automated log acquisition, at any scale appeared first on Dynatrace news.

]]>
What is FinOps?

Unlock the full value of log data with Dynatrace

Log data provides a unique source of truth for debugging applications, optimizing infrastructure, and investigating security incidents. To further enrich log data for automated observability, it’s necessary to dynamically tie logs to distributed traces on the code level, user sessions in the app front-end, and the topology of your IT landscape. This contextualization of log data enables AI-powered problem detection and root cause analysis at scale.

The key to 360-degree observability is acquiring the right data at the right time from the right places. Achieving this from a typical enterprise’s various apps, systems, and configurations is the beginning of the observability journey and, therefore, critical to get right.

Dynamic landscape and data handling requirements result in manual work

All this is easier said than done because:

  • Kubernetes-based dynamic architecture is becoming the norm. Many Dynatrace customers complain that typical sources for logs are ephemeral and short-lived pods or microservices, which brings new challenges in scaling up observability efforts. Such ever-changing log-source landscapes bring pain and hinder teams in scaling their observability efforts.
  • A typical enterprise environment includes multiple teams with varying requirements that can be conflicting. So, setting the same configuration rules for all teams is hardly a satisfying solution as it slows down more agile teams and reduces competitiveness. A customer recently told us, “We monitor hosts in our enterprise from more than 50 countries. Some monitoring requirements are the same for all teams; some country teams need to create a subset of rules.”
  • Data sovereignty and governance establish compliance standards that regulate or prohibit the collection of certain data in logs. To remain compliant, such logs are typically subject to pseudonymization or anonymization procedures (masking), which makes sensitive data inaccessible while preserving the contextual information of the event, service, or host.
  • As data volumes explode, teams have different prioritization for the data they want to acquire for observability. Collecting logs that aren’t relevant to their business case creates noise, overloads congested networks, and slows down teams.
  • Every manual step in growing enterprise environments becomes a hurdle. As the number of services or hosts reaches tens or hundreds of thousands, automation becomes the only way to tame the complexity.
  • Basic automation and out-of-the-box solutions might not cover all the edge cases within enterprise-scale installations. A customer recently shared their pain with manual configuration files, stating, “We must set up custom logging rules for some hosts, which requires logging into that host to create a configuration file. But in our case, the number of hosts with custom rules goes to four figures.”

The new configuration empowers log acquisition at scale

Dynatrace now provides a log acquisition toolkit that guarantees maximum value from your log data on the Dynatrace observability platform. By enabling configuration-as-code for central orchestration of your deployed OneAgents, automation can observe tens of thousands of hosts, services, or short-lived Kubernetes pods. With complete identity and access management support, scaling log management across your enterprise is more manageable.

Relevant log sources are automatically discovered as OneAgents are deployed across your IT landscape. Dynatrace offers an easy and powerful configuration flow for determining which auto-discovered logs are relevant to which teams and need to be ingested into Dynatrace. The central management of this configuration offers default setups and blanket rules, with granular rules supporting each team’s needs.

To control local network data volume and potential congestion, Dynatrace also allows filtering of log data on-source—by specific host, service, or even log content—before data is sent to the cloud.

Such filtering allows pseudonymizing or anonymizing (masking) sensitive data at the OneAgent level. This way, it‘s possible to flexibly select what confidential or sensitive information (for example, PII) is hashed or completely removed before it leaves the enterprise premises. This helps you stay compliant while working with sanitized logs without losing the event context, which provides valuable insights into DevOps, SRE, or business teams’ observability goals.

Even if a team has unique requirements for a specific edge case, the new custom log source capability eliminates the need for time-consuming manual work at the host level. By configuring custom log sources in Dynatrace, it’s possible to direct OneAgents to find and ingest log data relevant to a specific use case.

Flexible rules cover default and edge-case business needs

The new log acquisition configuration works by setting up rules on three levels:

  • Host
  • Host group
  • Tenant

OneAgent processes rules precisely in this order, with host scope rules processed before host group and tenant scope rules. If a more granular rule is present on the host level, that rule will precede any blanket rule on, for example, the tenant level.

Dynatrace OneAgent processes order rule

Setting up rules on each scope can require creating matches for logs across specific paths or process groups, instructions for masking sensitive data, or picking up custom log sources.

For example, consider that you need to monitor all NGINX logs across all hosts and Kubernetes nodes for pod logs. This involves setting up a rule on the tenant scope: telling OneAgents to upload all logs discovered in the /var/log/nginx path.

If the Kubernetes nodes belong to a EKS host group, then a second rule can be created for the host group. All new hosts appearing in that group would fall under that rule, and logs outside the EKS host group will not be stored.

This allows you to create flexible and powerful log storage configurations on any level by utilizing the unique autodiscovery capabilities of Dynatrace OneAgent or a custom setup. Sensitive data masking enables you to comply with applicable standards and data protection laws in a simple way, as the selected data never leaves your enterprise premises. By creating or modifying the log storage configuration through an API, teams can enable software intelligence as code, automate manual tasks, and achieve higher efficiency with a lower error rate.

Try it out yourself

The new power of log acquisition, masking and custom log sources in Dynatrace is supported for both SaaS and Managed deployments. It’s delivered in three parts:

The post Getting answers from data starts with automated log acquisition, at any scale appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/getting-answers-from-data-starts-with-automated-log-acquisition-at-any-scale/feed/ 0
How unified data and analytics offers a new approach to software intelligence https://www.dynatrace.com/news/blog/new-approach-to-software-intelligence/ https://www.dynatrace.com/news/blog/new-approach-to-software-intelligence/#respond Tue, 04 Oct 2022 09:00:09 +0000 https://www.dynatrace.com/news/?p=53551 Causal AI use cases for modern observability; exploratory data analytics

Today's organizations need a new approach to software intelligence. See how a data and analytics-powered approach unifies observability and security data while generating real-time insights.

The post How unified data and analytics offers a new approach to software intelligence appeared first on Dynatrace news.

]]>
Causal AI use cases for modern observability; exploratory data analytics

Software and data are a company’s competitive advantage. That’s because every company is now a software company. As a result, organizations need software to work perfectly to create customer experiences, deliver innovation, and generate operational efficiency. But for software to work perfectly, organizations need to use data to optimize every phase of the software lifecycle. That’s exactly what our platform does.

Much of the software developed today is cloud native. However, cloud infrastructure has become increasingly complex. This requires hundreds of interdependent services to work perfectly. Organizations must update services and apps dozens of times a day. Further, the delivery infrastructure that makes this happen has also become complex. The only way to address these challenges is through observability data — logs, metrics, and traces. But it doesn’t stop there.

Teams interact with myriad data types. For example, users generate user data, ecommerce sites generate business data, and service portals generate service desk tickets and call volume data. But how is this data connected? That’s where context comes into play.

Organizations need to unify all this observability, business, and security data based on context and generate real-time insights to inform actions taken by automation systems, as well as business, development, operations, and security teams.

Traditionally, though, to gain true business insight, organizations had to make tradeoffs between accessing quality, real-time data and factors such as data storage costs. IT pros want a data and analytics solution that doesn’t require tradeoffs between speed, scale, and cost.

With a data and analytics approach that focuses on performance without sacrificing cost, IT pros can gain access to answers that indicate precisely which service just went down and the root cause.

The next frontier: Data and analytics-centric software intelligence

Modern software intelligence needs a new approach. It should be open by design to accelerate innovation, enable powerful integration with other tools, and purposefully unify data and analytics. Enter Grail-powered data and analytics.

Grail is a purpose-built data lakehouse for observability, security, and AIOps. Grail makes it possible to converge real-time analytics, historical analytics, and predictive analytics on a single platform.

This purpose-built data lakehouse approach for observability, security, and AIOps offers schema-less ingestion of various data types, including logs, metrics, traces, user data, business data, and topology data. Additionally, it provides index-free storage and direct analytics access to source data without requiring data rehydration.

Modern software intelligence needs a new approach. Enter the Grail-powered data and analytics platform.

Ultimately, this helps address data scale and access performance constraints that prevent organizations from unlocking data’s full potential.

Here are six steps to creating a modern data stack and AI strategy for observability, AIOps, and application security.

1. Identify the business outcomes

It’s important to understand what business outcomes you want to achieve. Your key business objectives will drive your strategy and metrics. An example is improving customer experience. Customers today expect a very high level of experience when they engage with an organization.

Consider the data needed and its source. Collecting logs, metrics, events, and trace data is great. But for full-stack observability, you also need to bring together the topology data model, code-level details, and user experience data. If you’re using an approach that employs disparate tools to monitor individual components of the stack separately, then you’re leaving value on the table by failing to take a platform approach. Individual tools continue to promote and generate data silos and prevent organizations from using data effectively.

Modern observability platforms make it possible to centralize observability data from even the most complex stacks and get answers that help you achieve your desired business outcomes.

2. Ingest all the data you need from anywhere

A combination of proprietary and open source technology can speed data ingestion from common and long-tail sources. You need data from myriad sources centralized to get the right context to power precise answers. This is where a data lakehouse with software intelligence comes into play. A purpose-built data lakehouse can ingest a variety of data sources without requiring tradeoffs between data storage and performance.

Don’t reinvent the wheel. Between Dynatrace OneAgent and open source observability frameworks such as OpenTelemetry, you are well covered. These two technologies enable you to ingest all types of observability data — logs, metrics, traces, user experience, and business data.

Logs. Systems automatically generate logs, which record events that took place. Log entries usually contain the following:

  • Date and time of event;
  • System or resource name;
  • App name; and
  • Event severity.

Logs come in different formats depending on the source system — including key-value pairs, JSON, CSV, and more.

Metrics. This data is aggregated over a period of time. For example, this includes CPU utilization, memory percentage in use, and average load times.

Traces. A trace is the path a transaction took in an application for completion — for example, querying a database or executing a customer transaction. A trace is usually shown from the beginning of the transaction to the end.

User experience data. Web and mobile apps record data about every user interaction. This data includes information about crashes, lags, rage clicks, user interface hangs, and time spent in apps.

Business data. This data is anything generated from business operations. For example, this includes conversion rate, average order value, cart abandonment rate, service desk ticket, and call volume data.

3. Unify and centralize observability, security, and business data

Centralizing all organizational data is unrealistic. However, observability, security, and business data are different. Organizations need to instantly process, enrich, contextualize, and analyze all the data that supports mission-critical operations. All the infrastructure metrics, application performance data, and user experience data contain records of not only performance degradation events, but also security threats, fraudulent activities, and customer behavior.

Advances in data storage technology and architectures make it possible to store huge volumes and a variety of observability data in an efficient way that scales as requirements evolve. Centralizing observability data makes it easy to curate high-quality data. As a result, organizations accelerate the process of identifying relationships between entities, connecting the dots between disparate data sets, and gaining ROI from aggregated data with actionable insights.

If data sets are in siloed tools and systems, it slows down the process of delivering precise answers. In contrast, if you preserve and manage all the relationships between all data, it makes it possible to do more deterministic, causational AI on this data, resulting in precise answers.

4. Use AI to generate answers and insights from data

Once you unify and centralize the relevant data, you can apply real-time data processing to identify the precise root cause of issues and generate actionable insights. Successfully doing so at scale requires AI to identify the relationships and context between data types.

Having a real-time topology map that tracks all entities is useful. It helps derive context between different data slices. Doing so manually is beyond human capability because of data volumes, ingestion speeds, formats, and the dynamic and complex nature of the application environment and infrastructure. This is where AI excels in data and analytics-powered software intelligence. Continuously processing data from every layer of the stack opens the door to numerous possibilities.

Real-time anomaly detection. Real-time monitoring detects issues within infrastructure and applications before they become costly, customer-facing problems. Anomaly detection helps identify issues that deviate from the norm so teams can proactively resolve them. Getting relevant information to the right teams at the right moment is critical. In part, that involves reducing alerts so IT teams aren’t overwhelmed and can identify high-priority issues. Providing context and a prioritized list of issues helps them focus on the most important tasks.

Automated root-cause analysis (RCA). Automated RCA breaks down an issue into its components and identifies the precise root cause. But without full-stack observability, RCA is impossible. This is because, in many cases, an application issue is tied to a microservice on the back end. Having a real-time topology with deterministic AI helps immediately find the root cause. It saves engineers a lot of time by showing exactly what went wrong and how it happened.

Runtime application security. With continuous software intelligence from full-stack observability data, DevSecOps teams are notified if vulnerable code is called in production applications. By enriching data with vulnerability databases, operations engineers can create a risk-weighted priority list of security issues. Additionally, gaining a complete understanding of vulnerability severity and frequency becomes useful for developers.

5. Use exploratory analytics for lightning-fast answers

Sometimes, teams need to drill deeper into an answer, or a question pops into your mind. Exploratory analytics on a data lakehouse architecture with software intelligence makes it possible to write any query and get an instant answer, thanks to distributed query execution.

6. Automate actions and optimizations powered by AIOps

Automated root-cause analysis eliminates guesswork and human effort. Therefore, IT pros know exactly which code in a software release was problematic, or which server had an issue. Rolling back code deployment or restarting a server makes sense, and you can build rollback into an automated remediation workflow. A shift-left approach helps in designing the remediation mechanism during the early stages of software development. Notifying the right teams of remediation activity closes the loops and removes the burden of manual action.

Data and AI are key to making software work perfectly

Assembling, cleaning, combining, and enriching observability data from various systems is key to getting correct answers. Contrary to popular belief, AI systems easily fall prey to “garbage in, garbage out” principles — that is, the systems are only as good as the quality of the data. Controlling the quality of data is key to getting the right data and the right answers to achieve business goals. A data and analytics platform that includes a data lakehouse design and software intelligence facilitates unifying not only data, but also various analytics workloads for what ultimately matters — time to response and action.

The Dynatrace difference, powered by Grail

Dynatrace offers a unified platform that supports your mission to accelerate cloud transformation, eliminate inefficient silos, and streamline processes. By managing observability data in Grail — the Dynatrace data lakehouse with massively parallel processing — all your data is automatically stored with causational context, with no rehydration, indexes, or schemas to maintain.

With Grail, Dynatrace provides unparalleled precision in its ability to cut through the noise and empower you to focus on what is most critical. Thanks to the platform’s automation and AI, Dynatrace helps organizations tame cloud complexity, create operational efficiencies, and deliver better business outcomes.

Discover how software intelligence as code enables tailored observability, AIOps, and application security at scale.

The post How unified data and analytics offers a new approach to software intelligence appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/new-approach-to-software-intelligence/feed/ 0