log analytics | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Fri, 03 Jul 2026 09:42:53 +0000 en hourly 1 How Dynatrace supercharged log observability in 2025 https://www.dynatrace.com/news/blog/how-dynatrace-supercharged-log-observability-in-2025/ https://www.dynatrace.com/news/blog/how-dynatrace-supercharged-log-observability-in-2025/#respond Thu, 15 Jan 2026 17:18:49 +0000 https://www.dynatrace.com/news/?p=72456 Dynatrace Logs icon

Large enterprises such as Western Union, Vodafone, and United Airlines are ditching legacy log solutions in favor of a single, unified observability platform that delivers real-time insights and scalability at the petabyte level, as you’ll hear firsthand from them at Perform 2026. In this blog post, we’ll look back at the log-focused Dynatrace product releases […]

The post How Dynatrace supercharged log observability in 2025 appeared first on Dynatrace news.

]]>
Dynatrace Logs icon

Large enterprises such as Western Union, Vodafone, and United Airlines are ditching legacy log solutions in favor of a single, unified observability platform that delivers real-time insights and scalability at the petabyte level, as you’ll hear firsthand from them at Perform 2026.

In this blog post, we’ll look back at the log-focused Dynatrace product releases of 2025, while keeping in mind the three benefits that customers love most about Dynatrace:

  1. Fast log onboarding with unified ingestion from any source
  2. It’s easy to get started, yet powerful for your daily work
  3. Productivity boosts with Davis AI

Boost productivity with Davis AI

The promise of “Logs in Context” is simple:
Find the right log line at the right time, automatically and powered by AI.

The magic of Dynatrace is not a single feature or hyped AI. It’s the sum of many Dynatrace capabilities that comprise the foundation of the Dynatrace platform: Grail®, Smartscape®, Davis® AI, OpenPipeline®, and many others, that come at no extra cost, providing the automation and assistance you need.

Easily identify root causes and create tickets using the Dynatrace Problems app and logs.
Figure 1. Easily identify root causes and create tickets using the Dynatrace Problems app and logs.

If you aren’t yet using Dynatrace for your logs, stop stitching together clues across tools and say goodbye to manual swivel chair ops:

  • Logs in context: The right log lines appear automatically within the workflow or Dynatrace app you’re using. Whether that’s troubleshooting a service, reviewing Kubernetes node health, or investigating performance incidents of Infrastructure or cloud native apps.
  • Free of charge: Every in-context query, including surrounding logs, is now zero-rated (non-billable) when you view logs inside these Dynatrace core apps: Clouds, Infrastructure & Operations, Services, and Distributed Traces. While these apps don’t generate query consumption, ingestion and retention consumption are billed individually. We’re delivering the logs you need to take action – instantly, efficiently, and automatically correlated.
  • Leverage the power of Dynatrace Davis AI: With Dynatrace, features like “Explain logs” dramatically shorten time to action. Our customers report that their teams can more easily understand the possible causes and impacts of incidents without having to manually search for error codes in logs on Google.
  • By leveraging Davis AI, Workflow Automation, and integrations such as our ServiceNow partnership, customers can dramatically reduce the number of incidents; one of our customers reported reducing MTTI by 90%.

AI summaries are available across the Dynatrace platform and MCP server.

Explore logs, expand log messages, and comprehend them faster using the “explain log” AI feature.
Figure 2. Explore logs, expand log messages, and comprehend them faster using the “explain log” AI feature.

With Dynatrace, observability is not limited to cloud native apps. These features work seamlessly across cloud native, on-premises, hybrid, and traditional IT stacks. So, whether you’re on Kubernetes, a Mainframe, or an AWS Lambda function, the experience is the same.

Effortless for everyone, powerful for experts

Once your logs are ingested, you need to be able to understand them. This is where our Logs app shines for both new and expert Dynatrace users.

Pre-defined and admin-curated views boost productivity

Earlier this year, we improved the simplicity of applying complex and advanced queries with new data segmentation and advanced filters.

Using segments, admins and power-users can provide reusable and pre-scoped filters. When paired with dynamic variables, users can easily modify filter conditions.

Simultaneously, we continued enhancing the Logs app to provide advanced click-to-filter capabilities in various areas, like pinning frequent queries and filters:

  • Filter field: Suggest attributes, operators, and entities
  • Facets: Gain a quick understanding of patterns and groups, or build queries
  • Advanced filtering: Intuitive click-to-filter side pane, including JSON-structure log support with nesting
Combine segments and facets to create a pre-filtered view
Figure 3. Combine segments and facets to create a pre-filtered view

JSON‑structured log handling

Log messages aren’t always clean. A field might be hidden inside a nested message attribute or buried three levels deep in nested JSON.

Dynatrace log handling:

  • Detects and normalizes JSON.
  • Exposes nested fields in the UI without manual mapping.
  • Provides human-readable log messages in the results across all apps that use logs.

This way, you and your users can focus on analysis, not plumbing and normalizing logs.

Free text search surfaces the content you're looking for instantly, with human-readable results, even for JSON-structured log records
Figure 4. Free text search surfaces the content you’re looking for instantly, with human-readable results, even for JSON-structured log records

Correlation at scale

With Traces on Grail, your traces are automatically correlated in context with surfaced logs within the Distributed Tracing app, including associated exceptions.

The value you and your teams gain

If you’re accustomed to working with traces, you can continue using your troubleshooting routine and easily navigate from traces to logs and error exception messages. If you prefer to start your work by focusing on logs, you can achieve the same outcome.

The Dynatrace Distributed Tracing app automatically links logs with traces or spans.
Figure 5. The Dynatrace Distributed Tracing app automatically links logs to traces or spans.

Remember, Logs in context are free with the Distributed Tracing app!

Fast log onboarding with unified ingestion from any source

You want all your logs, and you want them fast. You don’t want to wrestle with YAML files, forward scripts, or configure custom collectors.

Centralized configuration, self-service management, and enabling teams with granular permissions to collect and ingest logs—these are what customers asked for:

  • OneAgent + Journald – Enhanced capabilities for automatically capturing logs on Linux machines with a single, centralized, configured agent: Dynatrace OneAgent®. Just deploy and watch the logs magically appear in your tenant.

Kubernetes logging made easy – The Dynatrace Kubernetes Logs Module gives you complete visibility without requiring OneAgent to operate in Full-Stack mode or to configure OTel manually.

Onboarding your Kubernetes cluster and logs using the Log Onboarding Wizard.
Figure 6. Onboarding your Kubernetes cluster and logs using the Log Onboarding Wizard.
  • Log Onboarding Wizard – To further simplify the onboarding experience, we’ve introduced a new wizard across several apps. When logs are missing, or you manually launch the wizard, it provides guided steps to onboard your logs, including creating an API key.
If you already have a standardized intake process in place for your teams, simply don’t provide one or all of the required permissions. Then your users won't be able to see the wizard or onboarding recommendations.
Figure 7. If you already have a standardized intake process in place for your teams, simply don’t provide one or all of the required permissions. Then your users won’t be able to see the wizard or onboarding recommendations.

Scale that never breaks

You can ingest up to 1 PB of logs per day per tenant, which should eliminate most sizing or scaling headaches. This bandwidth is part of the Dynatrace SaaS magic: Dynatrace Grail stores and processes everything in an indexless manner and using schema on-read. At the same time, OpenPipeline® routes the telemetry according to your rules and requirements defined in the pipelines.

  • Thousands of pipelines and self-service: Every pipeline has fine-grained permissions, so teams can create isolated pipelines and self-service onboarding, processing, and routing to buckets for retention.
  • 120+ parsing processors – From JSON normalization to custom field extraction, you can assign a processor to a pipeline and let OneAgent do the matching magic for you. Are you using OpenTelemetry or Cribl? Matching conditions or technology attribution offers you the same experience, regardless of the log source.
    10 MB log records – In our Go big with Dynatrace blog post, we discussed why large log records aren’t an anomaly and how customers benefit from out-of-the-box support for large log records.
    Figure 8. 10 MB log records – In our Go big with Dynatrace blog post, we discussed why large log records aren’t an anomaly and how customers benefit from out-of-the-box support for large log records.

With Dynatrace, your log ingestion stays ahead of your growth curve, no matter how many new sources you add.

Keep your costs predictable

Large enterprises often need to charge back to internal business units. We’ve introduced increased flexibility for existing features related to chargebacks:

  • Retain with Included Queries: Configurable on the individual bucket level, and seamlessly combinable with the established usage-based IRQ model.
  • Cost Allocation: Attribute your logs, metrics, and traces with business‑unit and product labels. This supports your FinOps efforts, as recently discussed in our blog post, Cost Allocation for Logs.

Best Practices: Not everything we delivered in 2025 was a product enhancement. We’ve also delivered a new best practices section in our product documentation, based on field feedback from pre-sales, post-sales, and support teams.

If you prefer to watch a webinar recording instead of reading, we recorded a video that walks you through all the best practices detailed in this blog post.

Dynatrace YouTube Series  | Optimize your logs: Save money and boost performance

Ready to get started?

Let’s make observability effortless, not overwhelming.

Your team can spend less time chasing logs and more time delivering value. Dive in today and experience the power of an observability platform that was built for the future of IT.

If you’re using Dynatrace SaaS with a DPS contract, all the features mentioned in this blog post are available to you. If you’re not, why not start a free trial today and experience the value yourself?

Resources

Dynatrace University – Free training that covers everything from basic log ingestion to advanced analytics.

Dynatrace Playground – Our free sandbox tenant with sample log files, ready to explore log

State of Log Management 2026 – Download the report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

The post How Dynatrace supercharged log observability in 2025 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-dynatrace-supercharged-log-observability-in-2025/feed/ 0
Unified privacy and sensitive data management for logs with the Sensitive Data Center https://www.dynatrace.com/news/blog/unified-privacy-and-sensitive-data-management-for-logs-with-the-sensitive-data-center/ https://www.dynatrace.com/news/blog/unified-privacy-and-sensitive-data-management-for-logs-with-the-sensitive-data-center/#respond Fri, 05 Dec 2025 18:02:28 +0000 https://www.dynatrace.com/news/?p=72151 Unified privacy and sensitive data management

Managing sensitive data in log files is getting even easier with Dynatrace. Our new Sensitive Data Center unifies privacy request workflow, data cleanup, and the new Sensitive Data Scanner to help you streamline working with sensitive data in Dynatrace. It allows teams to respond to data subject rights with confidence, remove records safely when required, […]

The post Unified privacy and sensitive data management for logs with the Sensitive Data Center appeared first on Dynatrace news.

]]>
Unified privacy and sensitive data management

Managing sensitive data in log files is getting even easier with Dynatrace. Our new Sensitive Data Center unifies privacy request workflow, data cleanup, and the new Sensitive Data Scanner to help you streamline working with sensitive data in Dynatrace. It allows teams to respond to data subject rights with confidence, remove records safely when required, and proactively discover and govern sensitive data ingested into Grail®. Sensitive Data Scanner will be available in a limited preview release by the end of 2025.

Handle sensitive data throughout your growing data ecosystem

Organizations face increasing pressure to demonstrate responsible practices for managing sensitive data while maintaining efficient operations and minimizing downtime. In addition to meeting end users’ data subject rights requests, such as the export or deletion of personal data, organizations must take proactive steps to prevent unnecessary exposure and storage of sensitive information. Achieving these objectives is challenging, especially with fragmented tools, siloed teams, and manual processes. As telemetry volumes increase, these inefficiencies lead to slower response times, higher operational overhead, and increased compliance risks.

Streamline privacy operations in Dynatrace with the Sensitive Data Center

The Sensitive Data Center brings privacy operations and sensitive data management together in a single app on the Dynatrace platform and complements existing privacy controls with an additional layer of control. Aligning scanning, cleanup, and data subject rights workflows with where your data resides helps teams reduce manual work and improve accuracy, all in a transparent process where every scan, cleanup, and request is logged and auditable, supporting your regulatory obligations with clarity and control.

Continuously scan for unintentionally ingested sensitive data in Logs on Grail

Imagine a service administrator who suspects that sensitive data might have been unintentionally ingested in their observability data. They need a quick way to confirm whether it happened and, if so, where the sensitive data resides and what type of sensitive data is affected—ideally without having to build custom scripts or pull engineers off priority work. The Sensitive Data Scanner is a new module in the Sensitive Data Center that helps you discover sensitive data at the time of ingestion, allowing you to govern it more effectively in three steps:

  1. Configure the scanner
  2. Review the scan results
  3. Mitigate any potential findings

Configure scans in the Sensitive Data Scanner

The setup is straightforward. You can choose to monitor specific buckets or the entire environment. Select the sensitive data type or types from built-in rules such as email, credit card, or IP address. You can set up several fine-grained scans with different scopes to accommodate different scan areas. A scan runs at a defined cadence every 6, 12, or 24 hours, depending on your compliance needs, and alerts you when data matching the selected criteria is found.

Sensitive data scans set up

Review scan results

A dashboard provides a clear overview of scan statuses and highlights when sensitive data is found. From there, you can drill down into a specific scan to review detailed findings and understand exactly what was detected.

Sensitive data scan dashboard

You can review the results and examine the data flow from ingestion to the storage location.

Sensitive data scan dashboard

Mitigate potential findings

With these results, you can immediately take action. You can configure or adjust masking rules to prevent similar data from being ingested in the future, change access to stored data, update retention periods, or utilize the cleanup functionality to delete the data as needed.

Sensitive Data Scanner preview

The Sensitive Data Scanner will be available in a preview release by the end of 2025. As we gather feedback, we will continue to refine the experience and expand coverage, allowing teams to move confidently from identification to action within the same app.

Act decisively with precise, auditable cleanup

Data cleanup is available directly within the Sensitive Data Center, allowing you to take action when your organization’s regulatory obligations or policies require the removal of data.

The “Cleanup data” workflow in Sensitive Data Center allows you to easily locate, review, and delete an entire time frame of data, as well as any selected individual records that contain sensitive data defined in your DQL search query. To improve accuracy and minimize the risk of accidentally deleting data, you can also select a reviewer who will review and approve deletion requests before the data is deleted. Learn more about deleting data in Grail.

Efficiently locate, export, and delete end users’ personal data

Let’s walk through another common scenario. The services administrator, ensuring Dynatrace is operating smoothly, receives an urgent request from the privacy legal team: “Please locate, compile, and delete all personal data associated with this email address in Dynatrace as part of our end-user’s right to be forgotten.” Privacy requests in the Sensitive Data Center offer an end-to-end experience for managing data subject rights requests within the Dynatrace platform, allowing for quick, compliant, and efficient handling of sensitive data.

Dynatrace empowers you with a ready-made solution for submitting, tracking, and verifying the status of requests. With a user interface designed for compliance needs, you can efficiently manage data export and deletion requests. A dashboard summarizes key details, including the request reference, status, and due date alignment.

Sensitive data scan dashboard

Together, scanning and cleanup ensure that only the sensitive data you intend to process is stored in Logs on Grail, while privacy requests provide the workflows and approvals necessary to comply with user privacy rights for lawfully processed personal data.

Try Sensitive Data Center in the Dynatrace Playground.

This blog may contain forward-looking statements about our product plans, upcoming features, and anticipated improvements.  These statements are for informational purposes only and are not promises or guarantees.  The development, release, and timing of any features or functionality described remain at the sole discretion of Dynatrace LLC and may be modified, delayed, or canceled without notice.  We encourage readers to make decisions based on the product’s current capabilities and features.

The post Unified privacy and sensitive data management for logs with the Sensitive Data Center appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/unified-privacy-and-sensitive-data-management-for-logs-with-the-sensitive-data-center/feed/ 0
Significantly improve your Mainframe availability by connecting logs with traces https://www.dynatrace.com/news/blog/significantly-improve-your-mainframe-availability-by-connecting-logs-with-traces/ https://www.dynatrace.com/news/blog/significantly-improve-your-mainframe-availability-by-connecting-logs-with-traces/#respond Fri, 31 Oct 2025 08:00:35 +0000 https://www.dynatrace.com/news/?p=71629 Connecting logs and traces related content

Speed up resolution of mainframe application issues and switch to a proactive and preventive mode of operations, through z/OS logs and traces in context with both your Dynatrace SaaS and Managed deployments.

The post Significantly improve your Mainframe availability by connecting logs with traces appeared first on Dynatrace news.

]]>
Connecting logs and traces related content

Logs are essential observability telemetry

Organizations are increasingly challenged to deliver seamless digital services, an essential component for achieving business-critical objectives. These challenges are amplified by complex hybrid cloud environments, where managing diverse technologies across cloud providers and platforms like IBM Z Mainframe becomes particularly demanding.

To address this complexity, it’s vital to unify observability telemetry and its signals across all layers of the application delivery chain, including the mainframe platform, within a single AI-powered observability solution.

Dynatrace enrichment capabilities enhance ingested log records by adding contextual metadata such as trace IDs, span IDs, and process group instance IDs automatically and without any manual tagging efforts. This allows seamless correlation between logs, application performance metrics, and traces. This is essential for accelerating AI-driven root cause analysis and significantly reducing time spent on shortening RCA and MTTx, as many of our customers can confirm.

Simplify and automate z/OS log collection

Dynatrace Log Management & Analytics extends beyond distributed technology stacks to include support for the IBM Z Mainframe platform. It automatically captures and ingests logs from monitored IBM CICS and IBM IMS regions and offers you advanced ingest rules. All collected logs are enriched automatically with topological metadata, allowing seamless mapping to Dynatrace’s topology and entity model for z/OS Hosts (LPARs) and z/OS Processes (regions).

Dynatrace automatically maps log lines to z/OS entities, in this case, to the process name and job ID of a CICS region
Figure 1. Dynatrace automatically maps log lines to z/OS entities, in this case, to the process name and job ID of a CICS region

Additionally, logs can be enriched with Trace IDs and Span IDs to precisely correlate each log line with the corresponding CICS or IMS trace or span that generated it.

Dynatrace can map log lines to a specific z/OS trace, which allows you to directly navigate to the trace that created the specific log line (via “View trace”).
Figure 2. Dynatrace can map log lines to a specific z/OS trace, which allows you to directly navigate to the trace that created the specific log line (via “View trace”).

This dramatically simplifies navigation for any user in your organization. By selecting the log line containing the Trace ID and Span ID, you can directly navigate to the related trace while understanding the load times and delay.

This example shows an end-to-end trace and the log line that was written by a CICS COBOL program
Figure 3. This example shows an end-to-end trace and the log line that was written by a CICS COBOL program

Let’s summarize what we’ve seen so far. Log enrichment significantly enhances and accelerates:

  • Correlation with distributed traces, enabling end-to-end visibility across systems.
  • Troubleshooting, by linking logs to specific spans or transactions—accelerating troubleshooting.
  • Observability for both structured and unstructured log data, ensuring comprehensive insights regardless of log format.

Beyond the agent: Stream mainframe logs to Dynatrace with OpenTelemetry

Dynatrace OneAgent® already supports ingestion of logs out of the box. However, when it comes to mainframe environments, the story is more nuanced.

Why all logs are not created equal

Some logs—especially those on mainframes—are proprietary, customer-specific, and deeply embedded in legacy workflows. And while Dynatrace is constantly adding additional and automated coverage for additional log types, some might never be supported natively by OneAgent, simply because their structure and relevance are unique to each customer.

But that doesn’t mean they’re out of reach.

OpenTelemetry to the rescue

For logs that fall outside OneAgent’s native scope, OpenTelemetry offers a powerful alternative. By deploying an OpenTelemetry Collector, customers can stream log data from their LPARs to a distributed host—preferably to Linux, Windows, or zLinux to save MSU consumption. But even z/OS itself is an option for hosting an OpenTelemetry Collector.

The Collector supports:

  • Filelog receiver for arbitrary text files
  • Syslog receiver for structured system logs
  • Filter processor for preprocessing and enrichment
  • Concurrent export to multiple backends, including Dynatrace via OTLP

Getting logs off the mainframe

There are several ways to move logs from LPARs to distributed systems:

  • SFTP: A blunt but reliable method
  • z/OSMF: Offers REST API access to SMF records
  • z/OS Data Gatherer: SMF REST Services
  • Custom scripts or processes: Tailored to specific datasets
  • Streaming frameworks: Kafka, MQ, or even FTP-to-Collector bridges

The bottom line: Just transfer log data from the mainframe to the host where the OpenTelemetry Collector is located, and it will handle everything for you from there.

There’s no strict requirement to provision a dedicated host to run your OpenTelemetry Collector. If Dynatrace OneAgent is already monitoring one of your LPARs, the ActiveGate hosting the Dynatrace zRemote component is a perfectly suitable environment for the Collector.

The ActiveGate hosting the zRemote mediates the ingestion of both out-of-the-box logs and OpenTelemetry logs.
Figure 4. The ActiveGate hosting the zRemote mediates the ingestion of both out-of-the-box logs and OpenTelemetry logs.

Preprocessing and enrichment

Both the Dynatrace Distro and the Contrib Distro of the OpenTelemetry Collector support advanced preprocessing:

  • Timestamp normalization (for example, converting z/OS timestamps to Unix time)
  • Resource attribute extraction (for example, job name, LPAR ID, subsystem)
  • Sensitive data masking and filtering

This ensures that even complex logs are transformed into structured telemetry before reaching the backend.

The cherry on top: Dynatrace OpenPipeline

While OpenTelemetry handles ingestion and transformation, Dynatrace OpenPipeline® adds another layer of intelligence during ingestion and processing:

  • Further enrich logs with business context
  • Extract metrics, events, and business observability events
  • Apply AI-driven baselining and anomaly detection
  • Transform, mask, or drop data
  • Some of these capabilities overlap with the Collector, giving users flexibility to choose where to apply logic based on performance, cost, and control.

Operlog: A prime candidate

Let’s take Operlog as an example—a system log that many customers are keen to stream. While it’s not a simple text dataset, creative solutions can bridge the gap. If you can extract Operlog records and store them as text on a Linux host, the Collector can ingest them immediately. From there, Dynatrace visualizations and alerting kick in.

Here’s a snapshot of how Operlog looks once streamed and processed in Dynatrace:

Log entries captured from Operlog visualized in Dynatrace
Figure 5. Log entries captured from Operlog visualized in Dynatrace

Not satisfied with only logs?

You’re not limited to ingesting only proprietary log files via OpenTelemetry.

Do you have access to metrics that are relevant for tracking the health of the subsystems on your mainframe? Would you rather feed in certain data as events instead of logs?

Just as the OpenTelemetry Collector is highly customizable, Dynatrace offers multiple ways of ingesting all these signals.

Possible sources for OpenTelemetry signals and how Dynatrace ingests them
Figure 6. Possible sources for OpenTelemetry signals and how Dynatrace ingests them

Conclusion

Mainframe logs may be complex, but they’re not unreachable. With OpenTelemetry and Dynatrace working in tandem, even the most proprietary datasets can be brought into the fold. Whether you’re using OneAgent, OpenTelemetry, or a hybrid approach, the key is creativity—and the right tooling.

What’s next

Dynatrace currently supports CICS MSGUSR and IMS Master Terminal Logs via its z/OS Agents, and remains committed to enhancing these capabilities. This includes exploring support for additional z/OS log types and expanding the scope of information captured through the z/OS Agents.

Log monitoring is also available for Linux on IBM Z and LinuxONE. For more details, see the blog post, Enable full observability for Linux on IBM Z mainframe now with logs.

Get started with Dynatrace log observability

If you’re looking to elevate your end-to-end observability and explore tailored possibilities within your specific z/OS environment, we’d be happy to connect. Reach out to us to request a demo and dive deeper into what Dynatrace can offer.

The post Significantly improve your Mainframe availability by connecting logs with traces appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/significantly-improve-your-mainframe-availability-by-connecting-logs-with-traces/feed/ 0
Leverage logs for an end-to-end view of your business processes via Dynatrace OpenPipeline https://www.dynatrace.com/news/blog/logs-for-end-to-end-view-of-business-processes-via-dynatrace-openpipeline/ https://www.dynatrace.com/news/blog/logs-for-end-to-end-view-of-business-processes-via-dynatrace-openpipeline/#respond Fri, 27 Sep 2024 17:45:49 +0000 https://www.dynatrace.com/news/?p=65788 Logs for business process

Organizations in today’s data-driven world often struggle with fragmented data sources that hinder comprehensive business insights. With Dynatrace OpenPipeline, you can ingest logs from any system and extract relevant business data to get a cohesive end-to-end view of your business processes.

The post Leverage logs for an end-to-end view of your business processes via Dynatrace OpenPipeline appeared first on Dynatrace news.

]]>
Logs for business process

Unrealized optimization potential of business processes due to monitoring gaps

Imagine a retail company facing gaps in its business process monitoring due to disparate data sources. Due to separated systems that handle different parts of the process, the view of the process is fragmented. On top of that, the data sources are inconsistent. While some data comes from modern systems with APIs, other data stems from older systems that generate log files, and some data originates from external vendors. This inconsistency complicates uniform gathering and data analysis, resulting in incomplete or inaccurate insights. This scenario is shared among large organizations that rely on multiple internal and external systems for data collection.

Figure 1. Incomplete view of the ordering process due to older systems
Figure 1. Incomplete view of the ordering process due to older systems

Get business process observability across data silos with Dynatrace OpenPipeline

Traditional observability platforms often focus on technical metrics and logs, which are essential for troubleshooting, but they don’t take business value into consideration. Dynatrace OpenPipeline goes beyond this by offering the possibility of extracting business-relevant data from logs and using them to monitor end-to-end business observability.

In our retail company example, older systems are involved in shipping the order. They only produce logs that have not yet been monitored from a business perspective, although they contain valuable business information.

By leveraging Dynatrace OpenPipeline, the retail company can integrate data from all sources across the whole process, including Dynatrace OneAgent®, logs, and external business tools. This approach ensures that the company can easily track the entire journey from online order to its successful delivery.

Figure 2. Extracting business events from logs enables an end-to-end view of the ordering process
Figure 2. Extracting business events from logs enables an end-to-end view of the ordering process

Benefits of capturing business events

Logs often contain valuable insights into your business; however, this information can be difficult to process, particularly as you probably only need data from some specific log lines.

Dynatrace OpenPipeline extracts this business information from logs and stores it as a separate data type, so-called business events. This has several advantages:

  • Separation of concerns: Logs are often used for technical troubleshooting. Extracting business events allows for a clear separation between technical and business-relevant data.
  • Access control: Different teams can access different types of data. For example, support teams might access logs for troubleshooting, while business teams access business events for analytics.
  • Ease of access: Having business events in a uniform format simplifies querying and visualization, making it easier to analyze and derive insights.

How to find valuable business information in logs

In our example of a retail company, we need to extract business information from shipping logs to know if our order has already been shipped. See a typical log file with shipping details below.

Figure 3. Log file with business information
Figure 3. Log file with business information

From this log file, we can identify and capture the following business information:

1. Correlation ID – The Correlation ID is part of the log payload and identifies this event in the context of a business process.

2. Message – The message is also part of the log payload and contains information on which part of the business process it represents.

3. Context – The context contains information on which IT system handled this part of the business process.

4. Status—The status information tells us if something was successful or if we hit an issue.

Use OpenPipeline to identify relevant log events

To turn this information into a business event for analytics, we utilize the capabilities of Dynatrace OpenPipeline to identify relevant log events, parse the data, and create a business event out of it. The approach is the following:

  1. Go to the OpenPipeline app (available by default in your Dynatrace tenant)
  1. Create a route that specifies which logs should be processed. For example, you can only process log events coming from a specific ingest source or containing a specific phrase in its log message.
  2. Define a pipeline that will process these logs.
  3. Configure the data extraction rules within the pipeline. This involves:
    • Extracting correlation IDs: Identify and extract correlation IDs from the log content. Rename these IDs to make them uniform with other events (for example, rename “order ID”).
    • Extracting the message: The relevant message will be extracted from the log content and used as part of the business event.
    • Parse out contextual information such as the context (which is the related host) or status information in our case.
    • Defining event types: Specify the event type using the extracted message and provider information. For example, you might name the provider retail.logs for consistency.

The extracted information is re-ingested into the pipeline as a new business event. The event is then stored in Grail™ datalake house and can be used for further analysis or process visualization.

Do you want to dig deeper into extracting the data? Watch the dedicated Dynatrace Lab episode with Andreas Grabner and Alistair Emslie for a step-by-step guide on how this is done:

Turn your business data into tangible value

The extracted business data enables you to optimize your processes and gain valuable new insights. You can easily create a dashboard like the one below: it provides the example retail company with real-time data and KPIs, such as the success rate of shipping orders. Unlike traditional dashboards focusing on specific applications or technical aspects, this dashboard tracks the performance of the complete business journey with business KPIs and metrics, utilizing the information extracted from the log files.

Of course, you can also leverage the full power of the Dynatrace platform, like AI-powered monitoring of your business KPIs through Davis® AI. You can also visualize the end-to-end process flow in our Business Flow app, which enables you to identify delays, errors, and exceptions, providing insights into IT and business-related issues.

Figure 4. A dashboard provides a holistic view of business process performance
Figure 4. A dashboard provides a holistic view of business process performance

Learn more about the capabilities of Dynatrace OpenPipeline

If you’re struggling with gaps in your business process monitoring caused by disparate data sources and outdated systems, consider how Dynatrace can transform your approach.

To learn more about Dynatrace OpenPipeline capabilities, see our documentation.

The post Leverage logs for an end-to-end view of your business processes via Dynatrace OpenPipeline appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/logs-for-end-to-end-view-of-business-processes-via-dynatrace-openpipeline/feed/ 0
Privacy spotlight: Ensure compliance by hard deleting individual records in Grail https://www.dynatrace.com/news/blog/hard-deleting-individual-records-in-grail/ https://www.dynatrace.com/news/blog/hard-deleting-individual-records-in-grail/#respond Thu, 11 Jul 2024 16:52:33 +0000 https://www.dynatrace.com/news/?p=64686

Dynatrace introduces record-level hard deletion to more effectively comply with end-user deletion requests in line with privacy laws. By ensuring only relevant data is efficiently and permanently removed, deletion requests help drive innovation and don’t negatively impact your data quality.

The post Privacy spotlight: Ensure compliance by hard deleting individual records in Grail appeared first on Dynatrace news.

]]>

Data deletion might seem simple, but it’s critical in data management and privacy laws. It’s not just about hitting the Delete button; it’s about securely meeting compliance requirements while ensuring data integrity. Dynatrace Grail™ is a data lakehouse optimized for high performance, automated data collection and processing, and queries of petabytes of data in real time. A data lakehouse is schema-on-read and indexless, making deletion operations complex. Adding to the technical challenges, effective deletion involves a combination of policies, procedures, and technologies to ensure data is appropriately managed throughout its lifecycle.

Strategically handle end-to-end data deletion

Two key elements form the backbone of an effective deletion strategy in Dynatrace SaaS data management: retention-based and on-demand deletion. Retention-based deletion is governed by a policy outlining the duration for which data is stored in the database before it’s deleted automatically. The retention period can vary based on the nature of the data, regulatory requirements, and the business’s specific needs. With Dynatrace Grail™, you can customize data retention periods with day-level granularity for up to 10 years.

On-demand deletions are initiated in response to specific events or requests. For instance, if data is mistakenly ingested into the database, it may need to be deleted to prevent inaccuracies or sensitive data from being stored. Another consideration is compliance with end-user privacy rights to delete personal data processed about them in line with data protection laws like GDPR and CCPA.

Hard deletion on the record level is the gold standard

On-demand data deletion in a Dynatrace SaaS environment relies on two core characteristics: granularity and hard deletion. Granularity, specifically at the record level, allows for precise control over what data is deleted. This means that individual records can be targeted for deletion without affecting the rest of the dataset. Hard deletion refers to the permanent and secure erasure of data. These characteristics are crucial in ensuring data deletion is effective, secure, and compliant with data privacy regulations.

Many other SaaS vendors mistakenly assume that a soft delete—merely hiding or marking data as deleted while retaining it elsewhere—is sufficient. However, this approach leaves room for accidental exposure, unauthorized access, or data breaches, jeopardizing compliance and customer trust. Industry standards refer to hard deletion by overwriting data as the appropriate data sanitization method.[1]

In addition, as with some other solutions, data deletion is often limited to a timeframe or requires removing all data in an index. This approach lacks flexibility and results in losing large volumes of important business data.

[1] NIST Special Publication 800-88 Revision 1, Guidelines for Media Sanitization

Keep valuable data in Grail with targeted deletion

With record-level hard deletion, Dynatrace gives you control and transparency over your data. It enables you to effectively identify and promptly delete targeted data in a thorough and irreversible manner. So, while the relevant data is completely erased, all other valuable data remains intact. This ensures that your business operations continue smoothly, and you can focus on extracting value from Dynatrace.

In addition to the possibility of deleting data in buckets in Grail, record deletion in Grail now enables you to identify and select the records to be removed by leveraging DQL. You can use the Grail Storage Record Deletion API to trigger a deletion request.

Here are some tips to consider when deleting records in Grail:

  • Leverage Notebooks: Use Notebooks to create a DQL query to filter all the records you intend to delete. With Notebooks, you can easily query data from Grail and visualize the results. This step also helps you determine the timeframe within which the relevant records are stored. To delete the records, use the Storage Record Deletion API. You might need to invoke the API multiple times if your data spans multiple days. Be aware that only one deletion process can be executed at a time.
  • Verify your permissions: Before proceeding with any data deletion tasks, ensure that you have the necessary permissions, specifically storage:records:delete.
  • Pause data ingestion: To avoid mistakenly deleting any recent data, stop ingesting new data that contains information you plan to delete. Only data older than 4 hours can be deleted.
  • Keep track of deletion: Use the status command to gain insights into the status of the current deletion process. If necessary, use the cancel command to cancel a running process.

An audit trail is maintained to provide a clear and transparent record of what data was deleted, when, and by whom. By following these best practices, you can ensure efficient and safe data management, allowing you to focus on extracting value from Dynatrace while maintaining smooth and compliant business operations.

Cleanup data in Grail with Privacy Rights

You can also hard delete on the record level in Grail using the out-of-the-box deletion functionalities in Privacy Rights. With an interface designed for your compliance needs, you can efficiently manage data deletion requests and cleanup tasks. A dashboard provides you with a summary of key details, including the request reference, status, and due date alignment. This streamlined process—from request creation to record deletion—is logged and auditable within the app, ensuring transparency and accountability.

The cleanup workflow in Privacy Rights allows you to refine your query to locate all relevant data points, guaranteeing a comprehensive and precise deletion process. In addition, a multi-user approval process with an assignee and a reviewer minimizes deletion risks and ensures accuracy. With Privacy Rights, what once seemed like a daunting compliance task has transformed into a simple and transparent workflow.

Get started

What’s next

As we continue to innovate and enhance our privacy features, we remain committed to providing you with the industry’s most robust and transparent data protection solutions. We’re working on making data deletion even easier by onboarding log deletion in Grail to the Privacy Rights app. Check our Privacy Rights documentation to stay tuned to our continuous improvements.

See documentation for Record deletion in Grail via API.

The post Privacy spotlight: Ensure compliance by hard deleting individual records in Grail appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/hard-deleting-individual-records-in-grail/feed/ 0
Stream logs to Dynatrace with Amazon Data Firehose to boost your cloud-native journey https://www.dynatrace.com/news/blog/stream-logs-to-dynatrace-with-amazon-data-firehose/ https://www.dynatrace.com/news/blog/stream-logs-to-dynatrace-with-amazon-data-firehose/#respond Fri, 03 May 2024 15:25:51 +0000 https://www.dynatrace.com/news/?p=63910 Dynatrace and Amazon Data Firehose

Your cloud logs can provide the root cause of high-impact issues or reveal the details of security incidents. Now, you can integrate an Amazon Data Firehose high-frequency data stream directly with the high-performant Dynatrace Grail™ analytics engine and use the Dynatrace AI-powered observability platform to mitigate issues with minimal impact to your business.

The post Stream logs to Dynatrace with Amazon Data Firehose to boost your cloud-native journey appeared first on Dynatrace news.

]]>
Dynatrace and Amazon Data Firehose

Real-time streaming needs real-time analytics

As enterprises move their workloads to cloud service providers like Amazon Web Services, the complexity of observing their workloads increases. Log data—the most verbose form of observability data, complementing other standardized signals like metrics and traces—is especially critical. As cloud complexity grows, it brings more volume, velocity, and variety of log data.

Managing this change is difficult. Without the ability to see the logs that are relevant to your service, infrastructure, or cloud function—at exactly the right time and in exactly the right format—your cloud or DevOps engineers lose the ability to find the root causes of the issues they troubleshoot. Even the AIOps approach doesn’t cut it if you don’t have proper logs in your observability platform.

Amazon CloudWatch is the most common method of collecting logs across your AWS footprint. As a native tool used by many enterprises, CloudWatch supports a wide range of AWS resources, applications, and services.

Amazon Data Firehose helps stream logs to the right destination

But your SREs and DevOps engineers know CloudWatch is not the terminal destination for data but rather an intermediate station. Their job is to find out the root cause of any SLO violations, ensure visibility into the application landscape to fix problems efficiently and minimize production costs by reducing errors. SREs and DevOps engineers need cloud logs in an integrated observability platform to monitor the whole software development lifecycle.

When trying to address this challenge, your cloud architects will likely choose Amazon Data Firehose. This fully managed native service is indispensable for streaming high-frequency logs collected by CloudWatch.

In some deployment scenarios, you might skip CloudWatch altogether. Take the example of Amazon Virtual Private Cloud (VPC) flow logs, which provide insights into the IP traffic of your network interfaces. VPC flow logs can be used as the source for troubleshooting connectivity issues, implementing security incident investigations, detecting intrusions, or managing access control issues. VPC flow logs can be massive in volume as your cloud deployment footprint grows, and directly streaming these logs with Amazon Data Firehose can be the most cost-effective method.

After configuring Amazon Data Firehose, your teams discover they have completed only the first part of the observability jigsaw puzzle. They also need a high-performance, real-time analytics platform to make that data actionable.

Dynatrace delivers the missing piece for AWS cloud observability with native Firehose integration. This complements our existing AWS logging integrations like S3 log forwarder, Lambda layer log forwarding, or direct log ingest API. These already provide a common integration with AWS log sources. The new Firehose integration removes intermediary components that previously required additional maintenance and provides a direct link from AWS to Grail data lakehouse.

This means high-frequency streamed logs from Firehose can be captured in your Dynatrace environment, automatically processed, stored in Grail for the retention period of your choice, and included in the full observability automation suite of the Dynatrace® platform, apps, and Davis® AI problem detection.

With this out-of-the-box support for scalable data ingest, log data is immediately available to your teams for troubleshooting and observability, investigating security issues, or auditing. As logs are first-class citizens alongside traces, metrics, business events, and other data types, you have an observability platform ready to scale with you in your cloud-native journey.

Easy setup takes just a few steps

Setting up a direct ingest of Firehose log data is quick and easy.

First, you need to generate an API key to ingest logs. In the Dynatrace web UI, go to Access tokens and select Generate new token. Select ingest logs as the scope of the token. Then, generate the token.

Next, go to the AWS console to configure the forwarding of data streams defined in your log groups. Data Firehose stream requires a trusted relationship with CloudWatch through an IAM role. Follow the instructions available in Dynatrace documentation to allow proper access and configure Firehose settings.

Now, you can set up your Firehose stream. The preferred way is to use a CloudFormation template that streamlines and automates the process. See CloudFormation template documentation for details.

Alternatively, you can configure the stream in the AWS web console. Choose Dynatrace as the Destination in the AWS console and complete the other fields with the correct parameters.

Choose Dynatrace as the destination in AWS console.
Figure 1. Choose Dynatrace as the destination in AWS console.

Now, you can view your cloud logs in Dynatrace!

For example, open the Clouds app with integrated logs in the context of your Lambda functions observability for one-click access to error logs.

See logs in context in the Dynatrace Platform, including the relevant logs for this AWS Lambda function.
Figure 2. See logs in context in the Dynatrace Platform, including the relevant logs for this AWS Lambda function.

When Dynatrace Davis AI detects a problem in your environment, you can also see relevant logs streamed via AWS Firehose that are related to the problem. When analyzing a problem, look at the related service, which displays related log data. This lets you jump right to the error that provides details of the problem.

When doing proactive health checks or analysis, you can inspect log data in Notebooks. For example, pick a template to explore data or write your own DQL query and chart incoming error rates from logs streamed via AWS Firehose.

Easily visualize Lambda error log distribution over time with Notebooks.
Figure 3. Easily visualize Lambda error log distribution over time with Notebooks.

Try it out today

Share your experience

We’d love to hear from you. Share your use cases for Amazon Data Firehose integration with the Dynatrace Community.

Stream AWS service logs collected in CloudWatch or directly via Firehose.

The post Stream logs to Dynatrace with Amazon Data Firehose to boost your cloud-native journey appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/stream-logs-to-dynatrace-with-amazon-data-firehose/feed/ 0
Privacy Spotlight: Easily comply with data subject rights in Dynatrace https://www.dynatrace.com/news/blog/privacy-spotlight-data-subject-rights-in-dynatrace/ https://www.dynatrace.com/news/blog/privacy-spotlight-data-subject-rights-in-dynatrace/#respond Thu, 02 May 2024 15:15:19 +0000 https://www.dynatrace.com/news/?p=63849 business resiliency

Introducing Privacy Rights, a Dynatrace® app built to streamline data-subject rights management. With Privacy Rights, you’re in control of your compliance tasks. You can efficiently follow up on your end-user's personal data requests in Grail™ in line with privacy laws like GDPR and CCPA.

The post Privacy Spotlight: Easily comply with data subject rights in Dynatrace appeared first on Dynatrace news.

]]>
business resiliency

Across the globe, privacy laws grant individuals data subject rights, such as the right to access and delete personal data processed about them. Rising consumer expectations for transparency and control over their data, combined with increasing data volumes, contribute to the importance of swift and efficient management of privacy rights requests.

“By 2026, fines due to mismanagement of subject rights will have increased tenfold from 2022 to total over $1 billion.” [1]
–Gartner®

These drivers and the growing complexity of data privacy regulations make manual handling of these requests unsustainable, necessitating automated and scalable solutions.

Handling privacy rights throughout your complex data ecosystem

Your obligation to uphold privacy rights extends beyond the boundaries of your own organization to the vendors in your supply chain. Because of that, SaaS providers, such as Dynatrace, play a crucial role in facilitating compliance with privacy rights. When a privacy request is initiated, it cascades through every layer of your supply chain. Successful compliance with privacy rights requests involves tracking and verifying requests across the entire data ecosystem, including third-party services.

Lack of control over the management of privacy rights can lead to increased administrative burdens, higher risk of errors, and overall frustration for organizations who fear a loss of customer trust in case of non-compliance with data protection regulations like GDPR and CCPA.

“While the need for scalable subject rights delivery and fulfillment will not go away, the demand for more automation will lead to a faster move toward a zero-touch model.” [2]
— Nader Henein, VP Analyst, Gartner

The Privacy Rights app is designed to streamline this process in Dynatrace.

Efficiently locate and export your end users’ personal data

Let’s walk through a fictitious example to explore privacy rights handling in Dynatrace. A diligent services administrator is busy making sure that Dynatrace is deployed and working correctly. But today, something new awaits in their inbox: an urgent data-subject rights request from the privacy legal team. It states: “Please locate, compile, and delete all personal data associated with this email address in Dynatrace as part of our end-user’s right to access their data.” How can this services administrator meet this request in a quick, compliant, and efficient way?

Unlike many other SaaS vendors, Dynatrace empowers you by providing a tool for submitting, tracking, and verifying the status of requests. Dynatrace overcomes the challenge of manual, error-prone processes with workflows that enhance transparency and accountability in privacy rights handling. With an out-of-the-box solution, Privacy Rights enables your team to focus on deriving value from Dynatrace rather than being stuck with time-consuming compliance tasks.

Privacy Rights in action

With an interface tailored to your compliance needs, Dynatrace users can track privacy rights requests. A dashboard summarizes relevant details, such as reference to the request and its status and alignment with any due date. This approach effortlessly keeps you and your privacy team informed about end-user requests without requiring additional follow-up through support for status updates. The entire process, from the creation of the request to the export or deletion of the relevant data, is logged and auditable within the app.

When creating an export or deletion request, an authorized user specifies the end user’s details. The details are used to search Grail for any personal data related to that user. This step lets you fine-tune your query to identify all matching data points, ensuring a thorough and accurate retrieval process.

Create export request in Dynatrace screenshot

Once the data in Grail that matches the run query is returned, a second authorized user reviews the results in terms of volume (the number of log records, volume, data residency, and the number of systems). Based on this overview, the export of the matching logs is either approved or rejected. Similarly, in the case of personal data deletion, the reviewer has the opportunity to accept or reject the deletion request. This multi-user approval process reduces export and deletion risks and serves as an accuracy check.

Export request in Dynatrace screenshot

With just a few clicks, what might have initially seemed like a cumbersome compliance task becomes a seamless and transparent workflow. Privacy Rights helps organizations confidently navigate the complexities of data subject rights handling in Dynatrace.

Get started with Privacy Rights today

What’s next

We’re working on making Dynatrace privacy management even easier. Check out Privacy Rights documentation to stay informed about our continuous improvements.

To install the Privacy Rights app, search for “Privacy Rights” in Dynatrace Hub and select Install.

[1] Gartner Press Release, “Gartner Predicts Fines Related to Mismanagement of Data Subject Rights Will Exceed $1 Billion by 2026,” 24 August 2023, https://www.gartner.com/en/newsroom/press-releases/2023-08-24-gartner-predicts-fines-related-to-mismanagement-of-data-subject-rights-will-exceed-1-billion-dollars-by-2026. GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.

[2] Ibid.

The post Privacy Spotlight: Easily comply with data subject rights in Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/privacy-spotlight-data-subject-rights-in-dynatrace/feed/ 0
Enable full observability for Linux on IBM Z mainframe now with logs https://www.dynatrace.com/news/blog/enable-full-observability-for-linux-on-ibm-z-mainframe-now-with-logs/ https://www.dynatrace.com/news/blog/enable-full-observability-for-linux-on-ibm-z-mainframe-now-with-logs/#respond Wed, 17 Apr 2024 17:25:45 +0000 https://www.dynatrace.com/news/?p=63665 Hosts fetch logs

Complement your hybrid cloud journey with a resilient observability setup with log monitoring on Linux on IBM Z and LinuxONE. Include the familiar mainframe OS into one integrated observability platform and thus eliminate the need for platform-specific component upkeep and management, reduce the risk of prolonged outages, and get full details from logs for troubleshooting.

The post Enable full observability for Linux on IBM Z mainframe now with logs appeared first on Dynatrace news.

]]>
Hosts fetch logs

Mainframe is a strong choice for hybrid cloud, but it brings observability challenges

IBM Z is a mainframe computing platform chosen by many organizations with a hybrid cloud strategy because of its security, resiliency, performance, scalability, and sustainability. With the availability of Linux on IBM Z and LinuxONE, the IBM Z platform brings a familiar host operating system and sustainability that could yield up to 75% energy reduction compared to x86 servers.

That’s why a hybrid cloud scenario, where workloads are shared between public clouds and highly performant mainframe platforms like IBM Z, is a robust and effective strategy.

The challenge for hybrid cloud deployments is maintaining critical observability, which must include the full set of monitoring signals: logs, metrics, and traces. Without combining these signals in a unified AI-powered observability platform, monitoring apps, infrastructure, and troubleshooting issues are nothing more than a patchwork of manual correlation.

Deploying your critical applications on additional host operating systems increases the dependencies for observability. It means maintaining platform-specific observability components or tools, managing security updates, and deploying changes, all of which lead to configuration spread.

This creates a risk that can impact your time to problem resolution in troubleshooting, the effectiveness of AIOps workflows to remediate issues before they affect your end-users, and ultimately your business metrics.

Logs become an integrated part of observability

Dynatrace provides a unified and integrated platform to observe such hybrid cloud deployments, now with added support to monitor logs effortlessly on Linux on IBM Z and LinuxONE.

OneAgent® is a core component of the Dynatrace platform; it enables observability with minimal setup effort while offering extensive and flexible central configuration options. By including logs in your hybrid cloud observability, you have everything you need in one place to make smarter, faster decisions when troubleshooting and measuring the health of your application environments.

You can now seamlessly expand your analysis of the root cause of any problem identified by Davis® AI with logs automatically available in the correct context of hosts, applications, or other identifiers specific to your environment.

Because Dynatrace provides a unified and central place to configure your observability, there is a single place where you manage your log collection for public and private clouds and mainframe components like Linux on IBM Z or LinuxONE.

This makes your log collection policies much more effective and transparent. You can push a filtering change to filter out all unwanted logs from your central Dynatrace environment and apply the change automatically to all your monitored platforms.

It’s also easy to minimize the risk of violating data access policies or regulations by masking sensitive data in logs. By centrally configuring masking rules for sensitive data, you can push the rules out to all of your deployed OneAgents wherever they are deployed to make sure you stay compliant.

Configure log collection across all your hosts

Start by deploying OneAgent for Linux on IBM Z by going to Deploy OneAgent (for earlier versions of Dynatrace and Dynatrace Managed deployments, go to Settings > Deploy Dynatrace), choose Linux as the underlying platform, and s390 as the installer type.

You can now install OneAgent on Linux with s390 architecture.
Figure 1. You can now install OneAgent on Linux with s390 architecture.

Next, set up log ingest. As log monitoring is now available with OneAgent for Linux on IBM Z, a single log ingest rule can cover all your Linux operating systems no matter what architecture is utilized under the hood. This means OneAgents deployed on Linux with s390, ARM, AIX, or x86 are covered.

Go to Settings > Log Monitoring > Log ingest rules and turn on Ingest all logs to start log collection.

Enabling the log ingestion setting applies to deployed OneAgents on Linux on IBM Z and LinuxONE.
Figure 2. Enabling the log ingestion setting applies to deployed OneAgents on Linux on IBM Z and LinuxONE.

Next you can start using logs in your troubleshooting and analysis tasks. For example, on the Dynatrace platform, open the new Infrastructure & Operations app and navigate to any monitored host running on Linux on IBM Z (s390 architecture). You can see the Logs tab for the host, which displays insights about automatically contextualized logs from that host.

Infrastructure & Operations app shows a monitored host with s390 architecture, and the Logs tab shows log data for that host.
Figure 3. The infrastructure & Operations app shows a monitored host with s390 architecture, and the Logs tab shows log data for that host.

You can take your Dynatrace Grail™ analysis of log data further in Notebooks. Start with a query builder to get error logs for your Linux on IBM Z hosts, and continue your exploration with Dynatrace Query Language.

Error logs in Notebooks with distribution chart
Figure 4. Error logs in Notebooks with distribution chart

Start monitoring logs on Linux on IBM Z

What’s next

Stay tuned for an upcoming blog post about log collection in OpenShift for Linux on IBM Z and LinuxONE.

Are you running containerized applications on IBM Z?

The post Enable full observability for Linux on IBM Z mainframe now with logs appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enable-full-observability-for-linux-on-ibm-z-mainframe-now-with-logs/feed/ 0
Bring syslog into Dynatrace using OpenTelemetry to get open source value with enterprise support https://www.dynatrace.com/news/blog/bring-syslog-into-dynatrace-over-opentelemetry-to-get-open-source-value-with-enterprise-support/ https://www.dynatrace.com/news/blog/bring-syslog-into-dynatrace-over-opentelemetry-to-get-open-source-value-with-enterprise-support/#respond Fri, 15 Mar 2024 16:45:45 +0000 https://www.dynatrace.com/news/?p=63075 Fetch logs

Syslog is a standard protocol for system and network device monitoring. Integrating syslog into enterprise observability solutions is tricky due to its strict support and security patching requirements.

The new Dynatrace OTel Collector distribution unlocks the power of syslog and open source community contributions with the power of Dynatrace support and the value of Dynatrace Grail™ to analyze log data from devices at scale.

The post Bring syslog into Dynatrace using OpenTelemetry to get open source value with enterprise support appeared first on Dynatrace news.

]]>
Fetch logs

Getting insights into the health and disruptions of your networking or infrastructure is fundamental to enterprise observability. Syslog is the go-to protocol that delivers infrastructure administrators, network engineers, and security team logs that tell them all they need to know about their systems’ delivery, performance, availability, and security.

Without syslog, you’re blind to what happens on your infrastructure

While syslog is a common way to gain insights into enterprise infrastructure operations, integrating it with other signals into an observability overview is often a painful experience.

Syslog is a protocol with clear specifications that require a dedicated syslog server. This is needed to collect messages across your systems because many different types of devices and applications can produce logs in the syslog format.

However, enterprise adoption at scale typically has much higher requirements for components than for supported features—components must have proper vendor support. Without vendor support, you’re betting your business on goodwill. Even for a supported component, delivering logs from applications and infrastructure to DevSecBizOps workflows requires significant manual configuration.

For example, a supported syslog component must support the masking of sensitive data at capture to avoid transmitting personally identifiable information or other confidential data over the network. Log batching, enrichment, transformation, log source distinction, and application offloading are also regular requirements.

As enterprise environments scale enormously, filtering and dropping data “at the edge” before transmission to a central collection point must be a supported option.

Compliance, retention, archiving, or data governance regulations often require multicasting logs from the original source to multiple destinations, like an observability platform with long-term log storage.

In the end, site reliability engineering (SRE) and security teams need to have data delivered via syslog to their observability platform, in the context of other data types.

Syslog will remain a proven log solution because without understanding why connections are dropped, server starts or stops, or which requests your firewall blocked, your organization runs like a ship where the captain on the bridge has no understanding of what’s going on at the lower levels of the ship. However the challenges in maintaining syslog in a cloud-native era create a maze of requirements that SRE teams and infrastructure administrators must navigate, often finding themselves maintaining multiple tools and components.

This increases the risk of multiple points of failure, adds overhead, and ultimately fractures observability overview with prolonged time needed to recover from potential outages.

Start monitoring syslog using OpenTelemetry under the Dynatrace umbrella of support

OpenTelemetry has been a rising star in the observability landscape and is often a preferred way to achieve end-to-end visibility with a vendor-agnostic footprint. Dynatrace has been a part of the OpenTelemetry journey for years and has contributed to its rise.

With the new Dynatrace OTel Collector distribution, we provide a streamlined and supported way to collect logs using the syslog protocol. This fills all the requirements enterprises have and makes it hassle-free to stream syslog to Grail data lakehouse integrating logs with other observability data.

The Dynatrace OTel Collector for syslog has numerous benefits. Our approach is to understand what components our customers need and value. We then integrate them with our observability platform and offer support, so you don’t have to worry about unsupported bugs or lack of ownership.

We also provide security updates and patches to critical vulnerabilities that may arise in the components. This alleviates the risk of open source components with unpatched vulnerabilities remaining open to exploitation long after they have been revealed.

Ultimately this combination of Dynatrace support and the OpenTelemetry standard gives you the best of both worlds—enterprise-grade software support with open source community contributions.

Dynatrace OTel Collector fits with your existing setup

The new Dynatrace OTel Collector fits nicely into your existing Dynatrace setup to bring in syslog data. Our existing log ingest API already supports your logs using the OpenTelemetry protocol, so you just need to deploy the collector and point your syslog producers to it.

To start using the Dynatrace OTel Collector, take the following steps:

  1. Generate an API token for the OTLP endpoint in your environment.
  2. Find our newly released Dynatrace OTel Collector, deploy it, and configure the exporter with your API key and environment ID.
  3. Configure receivers to enable different log sources for your syslog producers.
  4. Point your syslog sources to the collector and you’re done!

This diagram explains how the components communicate with each other.

This diagram explains how the components of the Dynatrace OTel Collector communicate with each other.

Take a look at this example for configuration. After generating an API token and deploying the collector, configure your instance. You need to configure each component (receiver, optional processor, and exporter) individually in a YAML file and enable them via pipelines. Follow the examples below or refer to Collector configuration documentation.

To point the exporter to your environment’s OTLP endpoint, add the following configuration:

exporters:
  logging:
    verbosity: detailed

  otlphttp/tenant_1:
    endpoint: "https://{your-tenant}.live.dynatrace.com/api/v2/otlp"
    headers:
      Authorization: "Api-Token {your-api-token}"

Next, you can add receivers to your collectors, for example, F5 BIG-IP systems to log to a remote syslog server (version 11.x-17.x). Refer to F5 BIG-IP documentation for detailed and up-to-date instructions regarding remote Syslog configuration. Take a look at Syslog (Dynatrace OTel Collector) in Dynatrace Hub for an example configuration file for the receiver, so you can enable two separate syslog endpoints for F5 and host syslogs. This allows you to differentiate log sources (attribute.device.type) for analysis in Dynatrace.

You can also make the Dynatrace OTel Collector multicast incoming syslog messages to multiple destinations. For example, you can set up exporters for your Dynatrace production environment and sandbox environment:

service:
  pipelines:
    logs:
      receivers: [syslog/f5, syslog/host]
      processors: [batch]
      exporters: [logging, otlphttp/tenant_1, otlphttp/tenant_2]

As a result, you should see logs in Dynatrace with corresponding log.source and device.type attributes:

Logs in Dynatrace with corresponding log.source and device.type attributes

Deploy Dynatrace OTel Collector for syslog now

What’s next

  • Stay tuned for direct syslog ingestion into Dynatrace, which brings syslog endpoints to an Environment ActiveGate, fully configurable from the cluster. This will enable you to use Dynatrace ActiveGate to ingest syslog data.

Go to Syslog (Dynatrace OTel Collector) in Dynatrace Hub to see examples and continue to the installation.

The post Bring syslog into Dynatrace using OpenTelemetry to get open source value with enterprise support appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/bring-syslog-into-dynatrace-over-opentelemetry-to-get-open-source-value-with-enterprise-support/feed/ 0
Enhance data collection with Dynatrace OTel Collector distribution https://www.dynatrace.com/news/blog/enhance-data-collection-with-dynatrace-otel-collector-distribution/ https://www.dynatrace.com/news/blog/enhance-data-collection-with-dynatrace-otel-collector-distribution/#respond Fri, 15 Mar 2024 16:38:36 +0000 https://www.dynatrace.com/news/?p=63069 OpenTelemetry demo app Astronomy Shop

As organizations strive for observability and data democratization, OpenTelemetry emerges as a key technology to create and transfer observability data. OpenTelemetry is gaining popularity because it’s considered a standard, and that’s why it’s a common choice for creating future-proof solutions for years to come. To answer the growing demand for OpenTelemetry, Dynatrace is proud to […]

The post Enhance data collection with Dynatrace OTel Collector distribution appeared first on Dynatrace news.

]]>
OpenTelemetry demo app Astronomy Shop

As organizations strive for observability and data democratization, OpenTelemetry emerges as a key technology to create and transfer observability data. OpenTelemetry is gaining popularity because it’s considered a standard, and that’s why it’s a common choice for creating future-proof solutions for years to come.

To answer the growing demand for OpenTelemetry, Dynatrace is proud to announce the release of the Dynatrace OTel Collector distribution. This collector, fully supported and maintained by Dynatrace, is entirely open source. Before we get into the specifics, let’s first recap the benefits OpenTelemetry offers and why using collectors is a best practice.

Understanding OpenTelemetry

OpenTelemetry is an open, vendor-neutral standard for creating, collecting, and transferring telemetry data, like traces, metrics, and logs. Developers and operators can gain insights into their applications and infrastructure without fear of vendor lock-in because OpenTelemetry is fully open source and owned by CNCF. The OpenTelemetry project is supported and maintained by representatives from Microsoft, Google, Amazon, and many others, including Dynatrace.

Why do I need an OpenTelemetry collector?

As the name suggests, an OpenTelemetry collector gathers data from multiple sources and sends it to observability backends, like Dynatrace, for analysis. A collector helps developers control their telemetry data streams for each signal. Different data streams can be directed to different backends or even multicast to multiple backends simultaneously. The configuration is highly flexible in solving various user needs.

A collector is also a powerful component for data processing. It removes the burden of managing retries, batching, and sampling from monitored applications, which can reduce the CPU and memory requirements of applications. A collector can also transform and enrich the data with additional context. For example, in a Kubernetes environment, a collector can automatically attach metadata about pods and namespaces to all observability data. This ensures that application telemetry is contextualized with the infrastructure, enabling the observability backend to link the application and infrastructure for enhanced insights and root cause analysis.

From a user perspective, a collector can also serve as an open source platform that can be extended with custom components. You can create internal collector components of your own, for example, to receive telemetry data in a special format or to process it in a certain way. Using a collector as a telemetry processing platform can be much easier than creating an entirely new application.

Why the Dynatrace OTel Collector

The OpenTelemetry community releases different distributions of the collector, many of which our customers use to send OpenTelemetry data to Dynatrace. So, why should you consider using the Dynatrace Otel Collector? Quite simply, support and stability.

We have seen many customers identify collectors as a potential solution for their needs. Still, they haven’t been able to deploy collectors in production due to the lack of support. Deploying an open source component without external support and the needed expertise is undoubtedly a risk. That’s why we provide Dynatrace customers with a Dynatrace-supported solution.

The Dynatrace Otel Collector comes with collector components that have been verified by Dynatrace for seamless operation. This removes the burden of manually validating each component and use case. To further help you with your collector journey, we publish configuration examples of typical Dynatrace use cases and best practices to provide a good starting point.

The Dynatrace Otel Collector includes components that we know run stably in production, which means we can offer full Dynatrace support. At the same time, innovations from the OpenTelemetry community can be added to the Dynatrace Otel Collector only after they are mature enough and have proven their stability.

Deployment and the typical use cases

The Dynatrace Otel Collector can be deployed on Kubernetes or Docker using a provided container image or directly on a host with the published binary. For more details, refer to our Dynatrace OTel Collector deployment guide.

In the initial release, the Dynatrace Otel Collector comes with components for:

Additionally, the Dynatrace Otel Collector includes a rich toolset for data processing to enrich, filter, transform, sample, and batch. The complete list of the components is available in our GitHub repository.

Dynatrace OTel Collector diagram with telemetry sources

What’s next

After the initial release, we’ll continue enhancing the Dynatrace OTel Collector with new features and capabilities to make it even easier to integrate with Dynatrace. We also intend to offer more automated methods to deploy the collector with pre-configurations. So, stay tuned.

As a significant contributor to the OpenTelemetry project, Dynatrace remains committed to working with the community and other vendors to enrich its capabilities and make it user-friendly for everyone.

Deploying the Dynatrace Otel Collector takes only minutes and uses the tooling you already know: standalone binary, Docker image, Kubernetes Operator, Helm chart, or a standard manifest file. The configuration maps one-to-one with the collector distributions from the OpenTelemetry community.

The post Enhance data collection with Dynatrace OTel Collector distribution appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enhance-data-collection-with-dynatrace-otel-collector-distribution/feed/ 0
The state of observability in 2024: AI, analytics, and automation https://www.dynatrace.com/news/blog/the-state-of-observability-in-2024/ https://www.dynatrace.com/news/blog/the-state-of-observability-in-2024/#respond Wed, 06 Mar 2024 16:52:47 +0000 https://www.dynatrace.com/news/?p=62911 Dynatrace AI-powered observability is now on Google Cloud

In 2024, digital transformation continues to dominate as a key priority, where organizations seek to build and maintain a competitive edge. To unlock the agility to drive this innovation, organizations are embracing multicloud environments and Agile delivery practices. But these environments are, by nature, difficult to manage. They create an explosion of data that is […]

The post The state of observability in 2024: AI, analytics, and automation appeared first on Dynatrace news.

]]>
Dynatrace AI-powered observability is now on Google Cloud

In 2024, digital transformation continues to dominate as a key priority, where organizations seek to build and maintain a competitive edge.

To unlock the agility to drive this innovation, organizations are embracing multicloud environments and Agile delivery practices.

But these environments are, by nature, difficult to manage. They create an explosion of data that is extremely challenging to manually capture, analyze, and act on. As a result, it will become increasingly difficult for IT and security teams to optimize user experiences and improve the resilience of their applications.

The latest Dynatrace report, “The state of observability 2024: Overcoming complexity through AI-driven analytics and automation,” explores these challenges and highlights how IT, business, and security teams can overcome them with a mature AI, analytics, and automation strategy.

Fragmented monitoring and analytics can’t keep up

The continued reliance on fragmented monitoring tools and manual analytics strategies is a particular pain point for IT and security teams.

According to the report, the average multicloud environment spans 12 different platforms and services. On average, organizations use 10 different observability or monitoring tools to manage applications, infrastructure, and user experience across these environments.

This fragmented approach leaves teams struggling to access the answers they need to accelerate innovation and optimize digital services effectively. Some 85% of technology leaders say the number of tools, platforms, dashboards, and applications they rely on adds to the complexity of managing a multicloud environment.

Manual monitoring processes are also too time-consuming, which distracts teams from tasks that create new value for customers and the business. In fact, 81% of technology leaders say the effort their teams invest in maintaining monitoring tools and preparing data for analysis steals time from innovation.

Kubernetes adds to the complexity of technology stacks

Alongside the challenges of managing multicloud environments, IT and security teams struggle to maintain visibility into cloud-native architectures as Kubernetes continues to become the platform of choice for modern applications.

Kubernetes architectures make it easier to quickly scale services to new users and drive efficiency gains through dynamic resource provisioning. However, due to the dynamic nature of Kubernetes, 76% of technology leaders say it’s more difficult to maintain visibility into this architecture compared with traditional technology stacks.

More specifically, organizations have highlighted attack detection and blocking (48%), observability (45%), and cloud cost management (38%) as key challenges as they continue to shift toward Kubernetes.

Organizations are drowning in data

“The state of observability 2024” report also found that the rise of dynamic cloud-native technology stacks has unleashed a flood of data that ITOps and security teams are struggling to contain.

The data has clear value in providing the insights that modern digital enterprises need to drive smarter decisions and better business outcomes. However, cloud environments generate data at such volume and velocity that 86% of technology leaders consider it impossible for teams to cost-effectively capture and analyze it using outdated practices and fragmented monitoring tools.

The report shows that log analytics are a particular challenge, as the long-term storage cost of all this data has begun to overshadow the value organizations can unlock from querying it. As they continue in this way, teams are forced to decide which logs to retain for real-time analytics and which to discard or archive in a lower-cost, less-accessible storage.

To address this, 79% of organizations are currently using or planning to adopt a unified platform for observability and security data within the next 12 months.

A mature AI, analytics, and automation strategy is essential

Organizations are increasingly turning to advanced AI, analytics, and automation capabilities to reduce the complexity of managing their multicloud environments.

Nearly three-quarters (72%) of organizations have adopted AIOps (or AI for IT operations) solutions to tackle the complexity of their multicloud environments. However, 97% of technology leaders say probabilistic machine learning approaches limit the value that AIOps tools deliver given the manual effort required to gain reliable insights.

To unlock all the benefits of AI-driven approaches, organizations need a unified observability platform that combines multiple AI techniques–including causal, predictive, and generative AI. These capabilities enable security and IT teams to power custom applications and analytics use cases to drive smarter decision making and more efficient ways of working.

For more information on what technology leaders across the globe are saying, read the report, “The state of observability 2024.”
For more information on how observability is becoming the control plane for AI transformation, read the report, “The state of observability 2025.”

The post The state of observability in 2024: AI, analytics, and automation appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-state-of-observability-in-2024/feed/ 0
Dynatrace log collection for ARM unlocks power-efficient architecture for your enterprise https://www.dynatrace.com/news/blog/dynatrace-log-collection-for-arm-unlocks-power-efficient-architecture-for-your-enterprise/ https://www.dynatrace.com/news/blog/dynatrace-log-collection-for-arm-unlocks-power-efficient-architecture-for-your-enterprise/#respond Tue, 05 Dec 2023 16:33:49 +0000 https://www.dynatrace.com/news/?p=60963 log collection for ARM

ARM (Advanced RISC Machine) architecture is finally mature enough to offer enterprise customers its promised energy efficiency and powerful performance boost. Implementing the Dynatrace® observability platform with metrics, traces, and logs gives you full visibility into your ARM architecture environment. It unlocks AI-powered problem detection with root cause identification and log-based analytics. Leverage the Dynatrace “secret sauce" by installing OneAgent® on ARM-based hosts to enable automatic log discovery and ingest logs to Grail™ data lakehouse at a massive scale.

The post Dynatrace log collection for ARM unlocks power-efficient architecture for your enterprise appeared first on Dynatrace news.

]]>
log collection for ARM

Without observability, the benefits of ARM are lost

Over the last decade and a half, a new wave of computer architecture has overtaken the world. ARM architecture, based on a processor type optimized for cloud and hyperscale computing, has become the most prevalent on the planet, with billions of ARM devices currently in use. This growth was spurred by mobile ecosystems with Android and iOS operating systems, where ARM has a unique advantage in energy efficiency while offering high performance.

While ARM processors have been humming in consumers’ pockets for over a decade, enterprise IT adoption has been slower. Legacy data center infrastructure and software support have kept all the benefits of ARM at, well… arm’s length. Still, we at Dynatrace speak to customers who recognize this value and want to implement it.

Energy efficiency and carbon footprint outshine x86 architectures

The first clear benefit of ARM in the enterprise IT landscape is energy efficiency. This is a crucial factor for all data centers, cloud or managed, where power consumption and cooling costs are reflected in the total cost of ownership. Traditional x86 architecture needs more power, so switching to ARM can offer a clear advantage.

Energy efficiency is coupled with the total enterprise carbon footprint. As organizations look to take ownership of their total ecological footprint and help mitigate climate change, it’s critically important for organizations to measure, monitor, and reduce their IT carbon footprints. Initiatives like the Carbon Impact app can be used to measure the footprint of monitored ARM-based hosts compared to x86 hosts.

Huge performance leaps in recent years

The top priority is often performance, where ARM resources have improved significantly. You would have been unimpressed if you evaluated public cloud vendor ARM-based offerings five or more years ago. But take a look at the latest iterations of, for example, AWS Graviton2, which delivered a 40% price/performance boost, and Graviton3, which had an additional 27% price/performance improvement over Graviton2.

These Improvements and other optimizations show that mature service offerings can rely on ARM architecture. You’re no longer required to use a single offering or choose from a few instance families; Graviton includes general-purpose and accelerated-computing offerings, plus compute-, memory-, and storage-optimized instances.

No observability, no gains

All these factors have made ARM an attractive computing architecture for innovative companies. However, the lack of end-to-end integrated observability with full support for log collection is a clear blocker for many organizations looking to adopt ARM-based services.

Without an observability platform to collect and process signals, and provide AI-powered answers for problem detection, root cause identification, and impact analysis, the migration to ARM remains only a roadmap item for many AIOps teams.

Even if some part of your codebase can be instrumented to collect observability data, having all three signal types (metrics, traces, and logs) is crucial. Without collecting logs from the observed platform in a scalable AI-powered data lakehouse like Grail, it’s more of a challenge to identify the root cause of problems and provide details for troubleshooting or security incidents.

Having no access to logs or relying on a siloed approach to logging poses a risk of blind spots in your IT landscape, prolonged outages, and an increase in Mean Time To Identify (MTTI) for incidents, all of which can nullify the benefits of ARM.

Hassle-free ARM deployment with automated log collection

With automated log observability, the Dynatrace platform enables your AIOps team to reap the full benefits of ARM architecture. Dynatrace OneAgent can automatically detect common logs for ARM-based hosts and ingest them on a massive scale out of the box.

As soon as you install just a single OneAgent on a host, for example, an ARM64 (AArch64) based Linux host, the OneAgent scans and autodiscovers logs every 60 seconds.

After installation, you can control OneAgent behavior using a powerful central configuration mechanism in your Dynatrace environment. You can tune the granularity of OneAgent filtering for each host to meet your requirements for discovery and ingestion, share configuration details with other hosts in the same host group, or apply global settings for your whole environment.

As OneAgent collects logs from your ARM-based host, it automatically ties the discovered data to your environment topology. In this way, log data is always associated with the host, service, or other entity that generated it.

AI-powered problem detection can then connect the log data, which often contains the source of truth, to detected issues in your environment, thereby speeding up troubleshooting and remediation by orders of magnitude compared to manual correlation.

In addition, the integration with auto-baselined metrics and automatic connection between distributed traces and logs puts you in complete control of the software you run on ARM-based architectures.

Immediately see the root cause in ARM host logs

You can enable the full value of Dynatrace with log monitoring on ARM in just a few steps.

  1. Install OneAgent version 1.269+ on an ARM host (Go to Apps In the Dynatrace web UI and search for Deploy OneAgent to access the installer). For full details, see OneAgent installation on Linux.
  2. Following OneAgent installation, you can verify on the host page that the ARM 64-bit instruction set is used.
    OneAgent installation
    Next, go to Settings > Log monitoring > Set up log ingest and review the log ingestion rules in the Hosts section. For the simplest quickstart, select Show rules.
    Set up log ingest
  3. Turn on the rule [Built-in] Ingest all logs to enable all rules and gain complete visibility into the host.
    Ingest all logs

Now you have the additional value of log data informing your Dynatrace root cause analysis should a problem arise on this host.

In the following example, a synthetic monitor is set up for webpage uptime monitoring. Dynatrace Davis® AI has automatically discovered a problem on the host, and the root cause analysis points to a PHP web service as the source of the problem.

Problem page

The problem page gives you a direct link to the service causing the disruption, and you can quickly access all related logs.

Problem page

Error logs for the process reveal a PHP syntax error with code-level information that caused the outage.

Problem page

As you can see, with the Dynatrace AI-powered observability platform at your disposal, you can benefit from resource-efficient ARM architectures in your organization without the risk of blindspots or guesswork when you encounter incidents or outages.

Start monitoring ARM and see what’s next

Coming next

We’re working on making Dynatrace log monitoring even easier to use at scale with a focus on areas highlighted by you and other customers:

  • Are you ingesting logs across multiple sources with a wildcard (*) ingest rule? Soon, you can enhance log records with custom file names or log path attributes for later analysis in Dynatrace.
  • Get automatic log ingest recommendation rules relevant to your environment based on the log sources discovered by OneAgent.
  • Tools to troubleshoot the status of your log sources and instructions for resolving issues in cases where a log file is inaccessible.
  • See the complete list of log sources on each host page, even if the sources are tied to another process group.

The post Dynatrace log collection for ARM unlocks power-efficient architecture for your enterprise appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-log-collection-for-arm-unlocks-power-efficient-architecture-for-your-enterprise/feed/ 0
Privacy spotlight: Retain data in Grail with 1-day precision to better meet your compliance requirements https://www.dynatrace.com/news/blog/privacy-spotlight-retain-data-in-grail-with-1-day-precision-for-up-to-10-years/ https://www.dynatrace.com/news/blog/privacy-spotlight-retain-data-in-grail-with-1-day-precision-for-up-to-10-years/#respond Tue, 05 Dec 2023 16:15:02 +0000 https://www.dynatrace.com/news/?p=60976 Retain data in Grail with 1-day precision

The constantly evolving challenges of privacy-conscious data management require more than traditional data retention settings, which are limited to weeks or months. With Dynatrace Grail™ you can configure data retention periods with day-level granularity for up to 10 years, ensuring your ability to meet legal requirements while maximizing the value you get from Dynatrace.

The post Privacy spotlight: Retain data in Grail with 1-day precision to better meet your compliance requirements appeared first on Dynatrace news.

]]>
Retain data in Grail with 1-day precision

Streamline privacy requirements with flexible retention periods

Data retention is a critical aspect of data handling, and it’s not just about privacy compliance—it’s about having the flexibility to optimize data storage times in Grail for your Dynatrace use cases. Most observability solutions rely on fixed, default retention periods, or provide only limited configurability with pre-defined retention periods to choose from (number of days, weeks, or months), or require you to set up rules to archive and extract logs.

Holding onto personal data beyond the minimum required timeframe can introduce privacy risks, inefficiency, and excessive storage costs. Implementing clear data deletion policies is also closely linked with the legal requirement of personal data minimization. The core privacy requirement to consider when defining retention periods is to keep data no longer than required for a specific purpose. While this might seem straightforward, the challenge lies in determining precisely what this means in diverse scenarios. Deciding on the correct duration for retaining data is context-specific; it depends on the nature of the data, its intended use, legal requirements, and your evolving business needs.

For example, log data, which can include personal data, should have a shorter retention period when used to troubleshoot and debug performance issues in your testing environment (compared to log data used for audit logs) because the data might not be needed after a few days. Fixed or predetermined timeframes lack the adaptability needed for diverse data management scenarios. Archiving options fail to provide a dynamic, user-friendly solution for tailoring retention periods.

Customize retention periods with buckets in Grail

You can define retention periods in custom buckets for all the data you store in Grail. Grail buckets function like folders in a file system. Each bucket should contain only those records that should be handled together as a set. You can learn more about custom bucket retention periods in our recent blog post, which explains how to enhance data management and includes best practices for setting up buckets with security context in Grail. Besides improving query speed, reducing related costs, and enabling organizations to target a specific business unit or production stage, custom buckets enable you to segment your data for specific use cases.

Segmentation of use cases establishes the perfect landscape for tailoring retention periods to the purposes for which data is collected. You might have specific data retention needs based on your industry and use cases. For example, industries such as finance and healthcare have specific regulations that dictate audit log retention periods ranging from months to several years.

With Dynatrace, you have full control over Grail retention periods—you choose how long to store each portion of data in line with your requirements. You can select the desired retention period for your data in bucket configuration settings, with the available retention periods ranging from 1 day to 10 years, plus an additional week (3,657 days). A hard deletion process is initiated when there’s no legal or business reason to continue retaining data in alignment with your configuration settings. The system reviews the timestamp of all data segments to see if any have reached the end of their specified retention period. Data that has exceeded its retention period is flagged for deletion, and the data is permanently removed from Grail, making it unrecoverable.

Storage management overview in Dynatrace screenshot

The Storage Management app makes it simple to set data retention periods

While the Grail bucket management API is a powerful approach you can apply to your data management policies in an enterprise setting at scale, a simpler method is also available. The Storage Management app on the Dynatrace® platform lets you manage all custom buckets for Grail in a visual user interface. This makes the data retention rules easy to grasp and implement even for non-technical users and auditors.

The Storage Management app provides an intuitive overview of Grail buckets for different types of data you store with Dynatrace (for example, logs and business events. Other data types will be available soon). For each type you can inspect existing buckets along with their defined retention periods and current status. The app also provides a quick way to extend or shorten a retention period for existing buckets, rename buckets, add new buckets, or delete buckets to completely wipe the data they contain.

Grail data retention

What’s next?

  • Sign up for a Dynatrace free trial account to start using Dynatrace today.
  • Start managing your Grail buckets with the Storage Management app, to be released with Dynatrace version 1.281.
  • Explore our documentation, which explains data retention periods and how to manage custom Grail buckets.
  • Learn more about our commitment to providing you with control and transparency over your customers’ personal data in the Dynatrace Trust Center.

The post Privacy spotlight: Retain data in Grail with 1-day precision to better meet your compliance requirements appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/privacy-spotlight-retain-data-in-grail-with-1-day-precision-for-up-to-10-years/feed/ 0
Enhance data management with Grail: Ultimate guide to custom buckets and security policies https://www.dynatrace.com/news/blog/enhance-data-management-with-grail-ultimate-guide-to-custom-buckets-and-security-policies/ https://www.dynatrace.com/news/blog/enhance-data-management-with-grail-ultimate-guide-to-custom-buckets-and-security-policies/#respond Fri, 06 Oct 2023 14:26:12 +0000 https://www.dynatrace.com/news/?p=59912 Application Security graphic

Logs now complete the observability picture alongside traces and metrics in the Dynatrace Grail™ data lakehouse. By following best practices for setting up buckets with security context in Grail, the appropriate log data is now always at the fingertips of the right teams.

The post Enhance data management with Grail: Ultimate guide to custom buckets and security policies appeared first on Dynatrace news.

]]>
Application Security graphic

Grail: Enterprise-ready data lakehouse

Grail, the Dynatrace causational data lakehouse, was explicitly designed for observability and security data, with artificial intelligence integrated into its foundation. We’ve further enhanced its capabilities to meet the high standards of large enterprises by incorporating record-level permission policies.

To fully utilize Grail features, it’s recommended that you incorporate its unique buckets and security policies at the beginning of your observability journey. Utilizing built-in mechanisms and customizing organization-specific policies maximizes the benefits of Grail capabilities.

Custom data buckets for faster queries, increased control, and custom retention periods

>> Scroll down to the bottom of this blog post to view a ~7-minute video demonstration of custom data buckets.

The first layer in the Grail data model consists of buckets and tables (and views for entities, which is outside this blog post’s scope).

Tables are a physical data model, essentially the type of observability data that you can store. Buckets are similar to folders, a physical storage location.

There is a default bucket for each table. Here is the list of tables and corresponding default buckets in Grail.

Table name Default bucket
logs default_logs
events default_events
metrics default_metrics
bizevents default_bizevents
dt.system.events dt_system_events
spans default_spans

The default buckets let you ingest data immediately, but you can also create additional custom buckets to make the most of Grail.

Address specific use cases with custom buckets

It’s logical to segregate high-volume data into its own bucket. This allows the data to be frequently queried and used separately from other scenarios.

For example, a separate bucket could be used for detailed logs from Dynatrace Synthetic nodes. Debug-level logs, which also generate high volumes and have a shorter lifespan or value period than other logs, could similarly benefit from dedicated storage. Keeping these logs separate decreases the data volume for other troubleshooting logs. This improves query speeds and reduces related costs for all other teams and apps.

Custom data buckets with Dynatrace Grail

Address organizational structure with custom buckets

Depending on your organization’s structure, you may find it beneficial to keep logs used by specific business units or departments in separate buckets. This approach makes queries faster for individual units, as they only query relevant logs, and it ensures distinct access separation.

Suppose a single Grail environment is central storage for pre-production and production systems. In that case, the use of separate buckets makes it easier to distinguish different stages, keeping production data separate from development or staging data.

Adopting this level of data segmentation helps to maximize Grail’s performance potential. If your typical queries only target a specific use case, business unit, or production stage, ensuring they don’t include unrelated buckets helps maintain efficiency and relevance.

Custom buckets unlock different retention periods.

Segmenting your data into multiple buckets also puts you in control of the data retention period. The simplicity of storing data in Grail is reflected in its retention policies; you choose how long to store each portion of your data, and you never have to think about managing archives or retrieving archived data. Coupled with the transparent pricing of GiB/day, you can set up buckets to exactly match your business needs.

While the built-in default_logs bucket has a retention period of 35 days, you have more options to choose from. Use Grail’s public bucket management API to create new buckets for which you can select data retention periods of 1 day to 10 years , or use the new Storage Management app to that with just a few clicks.

In conjunction with the previous example of keeping high-volume and short-lived logs separate, you might also need to keep your application data longer. For example, transaction data and user-profile logs might need to be retained for 12 or 18 months.

Use buckets to query only the log data you need

Whether looking for a “needle in a haystack” or reporting on data stored in Grail, you can start an advanced query with DQL by fetching data from one of the tables. At this point, you should familiarize yourself with the blog post Tailored access management, Part 2: Onboard users to Grail and AppEngine, which covers access to Grail tables and buckets.

Now, let’s take a look at a query example that puts this all into use.

fetch logs

| filter loglevel=="ERROR"

In this example, we query a certain table (logs) and filter the results by a field (loglevel) with a certain value (ERROR). Note that with such a query, you fetch logs from all the buckets the end-user can access.

Custom data buckets with Dynatrace Grail

Although this initially only includes the default bucket, you might also include other buckets (if these are available to the user). As you bring in more data and users to Grail, relying just on the default buckets is not the optimal setup.

This is where filtering on custom buckets comes in. This allows you to query data from a specific bucket.

fetch logs

| filter dt.system.bucket=="prod_infra_logs" and loglevel=="ERROR"

This example now includes an additional filter that restricts data retrieval from a certain bucket (prod_infra_logs).

Custom data buckets with Dynatrace Grail

Using buckets to query only the data you need significantly speeds up queries and reduces query costs.

Buckets for data with high-security requirements

Buckets can also be used for managing high-level access control of data.

You can store access control logs or payment provider events in separate buckets to grant only limited SecOps or business administrators access to these logs.

Custom data buckets with Dynatrace Grail

Grail and AppEngine’s new policy-based access management provides a way to do this with buckets. For example, by adding a WHERE clause to the policy statement, you can define a specific bucket (=) or a range of buckets (STARTSWITH). In this example, the policy grants access to all buckets that have names starting with prod_infra_.

ALLOW storage:buckets:read WHERE storage:bucket-name STARTSWITH "prod_infra_";

However, creating access policies solely on the bucket and table level is not scalable in a enterprise landscape, as one Dynatrace tenant can have a limited number of custom buckets. Instead, access control based on specific attributes like host groups or team assignments can be achieved using different policies that are based on supported attributes or security context.

Record-level permissions and security context

As covered in the previously linked blog post about access management, Dynatrace Grail brings a new architecture to permissions management. The new approach that uses security policies provides you with new dynamic controls for user authorization.

As using custom buckets opened up a basic approach to access where users could get access to a whole bucket, record-level permissions allow you to take a fine-grained approach.

This means that whenever you run a DQL query to fetch data from Grail, your policy-based access rights are evaluated, and records without defined access are filtered out.

This means your teams’ permissions are not constrained by data management decisions on a bucket level. Let’s say it makes sense to consolidate all short-living app debug logs to a bucket that has a short retention period. With record-level permissions, you can now ensure multiple app owners can see only their data in that bucket.

Another example would be a business unit admin who needs to have access to departmental data across buckets.

Custom data buckets with Dynatrace Grail

Permissions based on DQL fields and security context

To implement this in the Log Management and Analytics context, you can create policies with additional clauses that provide access.

ALLOW storage:logs:read
WHERE storage:k8s.namespace.name="abc"

 AND storage:dt.host_group.id STARTSWITH "org1-";

In this example, the policy allows access to logs that have a certain Kubernetes namespace (abc) and originate from a host group whose name must start with a specific string (org1-).

The list of standard table fields used in security policies provides flexibility for defining individual policies.

DQL fields Mainly used with
event.kind events, bizevents
event.type events, bizevents
event.provider events, bizevents
k8s.namespace.name events, bizevents, logs, metrics, spans
k8s.cluster.name events, bizevents, logs, metrics, spans
host.name events, bizevents, logs, metrics, spans
dt.host_group.id events, bizevents, logs, metrics, spans
metric.key metrics
service.name events, bizevents, logs, metrics, spans
log.source logs
dt.security_context events, bizevents, system, logs, metrics, spans, entities
gcp.project.id events, bizevents, logs, metrics
aws.account.id events, bizevents, logs, metrics
azure.subscription events, bizevents, logs, metrics
azure.resource.group events, bizevents, logs, metrics

When you look at the list of DQL fields, you’ll notice one reserved field, dt.security_context.

You can assign a value to dt.security_context during data ingest for use in a security policy, which is not covered by the list of previous DQL fields. You can explicitly set a value for dt.security_context for some logs, or take the value of an existing field.

Log monitoring security context in Dynatrace settings

In this example, logs from a particular source (dsfm) are enriched with a literal value for dt.security_context (sec-lvl-7) during log ingest. A policy with a specific clause can provide access to only logs with this security context. Security context rule management is available via settings In the Dynatrace web UI and API.

This approach to granular record-level permissions opens up the flexibility needed in enterprise environments. You can craft policies based on existing fields like Kubernetes cluster or namespace, a host or a host group, a service name, or a log source. Or you can use security context for any other use cases, like granting access based on an AWS account, a GCP project, an Azure subscription, a username, or a team name.

Take the first step now

Many organizations have found immediate value in working with logs in Grail. Starting from optimizing their business and opening revenue streams based on data in logs to unlocking real-time insights from observability data and eliminating hours of manual work per process, as the Bank of Montreal did recently. Utilizing the core aspects of Grail, like buckets and permissions, sets organizations on the path to success.

Next steps

  • Start a Dynatrace free trial and explore Log Management and Analytics powered by Grail
  • Read our documentation explaining buckets, permissions in Grail, security context for logs, and IAM
  • Record-level permissions for Grail are generally available with Dynatrace version 1.277. Custom buckets, in addition to bucket and table permissions, are available with Dynatrace version 1.265.
  • Use Storage Management app to create and manage custom Grail buckets and unlock custom retention times for your data since Dynatrace version 1.281.

Special thanks to Dominik Punz and Christian Kiesewetter for contributing to this blog post.

Video demo

Watch this 7-minute video to see how you can use custom data buckets to separate use cases, data retention, and access permissions in Dynatrace.

Video thumbnail
Use buckets to separate use cases, data retention, and access permissions (7-minute video)

The post Enhance data management with Grail: Ultimate guide to custom buckets and security policies appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enhance-data-management-with-grail-ultimate-guide-to-custom-buckets-and-security-policies/feed/ 0
TTP-based threat hunting with Dynatrace Security Analytics and Falco Alerts solves alert noise https://www.dynatrace.com/news/blog/ttp-based-threat-hunting-solves-alert-noise/ https://www.dynatrace.com/news/blog/ttp-based-threat-hunting-solves-alert-noise/#respond Wed, 09 Aug 2023 11:58:55 +0000 https://www.dynatrace.com/news/?p=59078 TTP-based threat hunting with Dynatrace Grail and Falco for Security Analytics

Today’s security analysts have no easy job. Not only are cyberattacks increasing, but they’re also becoming more sophisticated, with tools such as WormGPT putting generative AI technology in the hands of attackers. As a result, analysts are turning to AI and TTP-based threat-hunting techniques to uncover how attackers are trying to exploit their environments. While AIOps with generative AI […]

The post TTP-based threat hunting with Dynatrace Security Analytics and Falco Alerts solves alert noise appeared first on Dynatrace news.

]]>
TTP-based threat hunting with Dynatrace Grail and Falco for Security Analytics

Today’s security analysts have no easy job. Not only are cyberattacks increasing, but they’re also becoming more sophisticated, with tools such as WormGPT putting generative AI technology in the hands of attackers. As a result, analysts are turning to AI and TTP-based threat-hunting techniques to uncover how attackers are trying to exploit their environments.

While AIOps with generative AI will certainly empower security teams to mitigate threats faster and with greater precision, attackers will just as certainly utilize the same technology to create novel malware, more convincing phishing campaigns, and uncover high-risk zero-day vulnerabilities quicker.

Not only that, teams struggle to correlate events and alerts from a wide range of security tools, need to put them into context, and infer their risk for the business. But the industry as a whole is still hampered by a ubiquitous tool sprawl to achieve that critical mission under a barrage of alert noise.

In this blog post, we’ll use Dynatrace Security Analytics to go threat hunting, bringing together logs, traces, metrics, and, crucially, threat alerts. We use the power of DQL on Grail to derive high-level attacker tactics, techniques, and procedures (TTPs), which are much easier to interpret and act upon.

TTP-based threat hunting: Tactics, techniques, procedures

At Dynatrace, we don’t want to bombard you with alert noise and uncorrelated warnings. Instead, we want to focus on detecting and stopping attacks before they happen: In your applications, in context, at the exact line of code that is vulnerable and in use. But even when an attack happens, Dynatrace detects and blocks them in real time while providing you with rich technical details on the concrete attack procedure. Procedures describe the specific technical details that an adversary used to carry out an attack, for example, what script they ran to exploit a weakness.

TTP-based threat hunting with Dynatrace: tactic, technique, procedure

When investigating advanced cyberattacks, it’s helpful to map attack procedures to attack techniques. Techniques describe the tactical goal an adversary is pursuing by executing a specific procedure. One of the most critical attack techniques within the MITRE ATT&CK® knowledge base of adversary tactics and techniques — and one example of what Dynatrace can prevent, detect, and block in real-time — is attack technique T1190, “Exploit Public-Facing Application”.

Public-facing applications can be an initial access vector an attacker could exploit to gain entry into a system. Attack tactics describe why an attacker performs an action, for example, to get that first foothold into your network.

Thinking in terms of tactics, techniques, and procedures (TTPs) brings many benefits. For example, security analysts can more easily stitch together advanced cyberattacks on an abstract level. Likewise, operation specialists can prioritize their efforts on monitoring the highest-risk tactics, and executives can better communicate the business risk.

Threat hunting and analyzing threat alerts with Dynatrace Security Analytics and Grail

Dynatrace offers Runtime Application Protection to detect a wide range of injection attacks in your applications. However, our customers often want to augment the data Dynatrace provides with data from third-party tools. Customers also want to carry out their own analysis tailored to specific use cases and forensic needs.

Dynatrace Grail is a data lakehouse that provides context-rich analytics capabilities for observability, security, and business data. You may also ingest additional data into our unified intelligence platform: One popular choice to gather fine-grained security data is Falco. Falco is an open-source, cloud-native security tool that utilizes the Linux kernel technology eBPF, to generate fine-grained networking, security, and observability events.

In the following sections, we demo the following:

  1. Introduce Unguard, our insecure cloud-native microservices demo application.
  2. Install Falco in AWS EKS to gather security-relevant events from all the happenings in Unguard.
  3. Ingest those Falco events into Dynatrace Grail using falcosidekick.
  4. Query Falco events in Dynatrace Grail, map them to TTPs, and conduct structural multi-step attack detection.

In other words, we find attacks that are composed of multiple steps by using TTPs and Dyntrace Smartscape for DQL in a way that eliminates alert noise.

First, Dynatrace OneAgent will automatically monitor and trace our infrastructure and communicate with Dynatrace. Second, we will enrich our data in the Grail data lakehouse by also ingesting Falco events using falcosidekick.

threat hunting architecture with Dynatrace and Falco

Setting up our TTP-based threat-hunting demo environment

Before we start threat hunting, we’ll first walk through how to set up the demo environment.

Introducing Unguard, our insecure cloud-native demo app

As our playground, we introduce Unguard, a microblogging demo application that embodies the challenges of modern cloud-native environments. It consists of eight services, written in at least four different languages, with countless vulnerabilities and misconfigurations. To keep it real, we have a load generator that creates benign traffic. It also generates OpenTelemetry traces.

TTP-based threat hunting: Unguard demo application

Unguard was first introduced at DEFCON 31 by our colleagues Simon Ammer and Christoph Wedenig.

For the demonstration in this blog post, we want to deploy Unguard in AWS EKS and hunt for attacks within that environment. You can easily play around with Unguard by installing its Helm chart:

helm install unguard \ 
  oci://ghcr.io/dynatrace-oss/unguard/chart/unguard \ 
  --wait --namespace unguard --create-namespace

(Please read the Unguard README for detailed and up-to-date instructions)

This demo assumes your Kubernetes cluster is already monitored by Dynatrace. For instructions, see Set up Dynatrace on Kubernetes.

Deploy Falco and falcosidekick in AWS EKS

You can install Falco in various ways. For this demo, we installed it with the Helm chart in our AWS EKS cluster:

helm repo add falcosecurity https://falcosecurity.github.io/charts 
helm repo update 
helm install falco falcosecurity/falco --namespace falco --create-namespace

(See the Falco README for detailed and up-to-date instructions)

Next, we set up falcosidekick, which is a daemon that forwards Falco events to many possible outputs. We’re proud to announce that, with Falco version 2.29, currently in pre-release, you can now also use Dynatrace as an output.

You can use this minimal values.yaml configuration file for the Helm chart:

# values.yaml 
 
falcosidekick: 
  enabled: true 
  image: 
    tag: 2.29.0-rc.1 
  config: 
    # as of 2023-08-02, this feature is still a pre-release so we 
    # have to manually override the environment variables for now 
    extraEnv: 
      - name: DYNATRACE_APITOKEN 
        value: dt0c01.EXAMPLE_TOKEN_REPLACE_THE_ENTIRE_STRING 
      - name: DYNATRACE_APIURL 
        value: https://ENVIRONMENTID.live.dynatrace.com/api

(Please read the Helm chart README for detailed and up-to-date instructions)

We insert the apitoken we generated within Dynatrace and grant the token the scope logs.ingest. See the topic Dynatrace API – Tokens and authentication to learn more about creating tokens. As the apiurl, use the following:

Dynatrace SaaS:

https://ENVIRONMENTID.live.dynatrace.com/api

Dynatrace Managed:

https://YOURDOMAIN/e/ENVIRONMENTID/api

See the topic Environment ID to learn more about environment IDs.

Finally, we update the Falco Helm chart with this new configuration:

helm upgrade falco falcosecurity/falco -f values.yaml

If everything worked out well (check the pod logs otherwise), we are now able to successfully query Falco events with DQL. To verify, we open a new Notebook and see how Dynatrace automatically infers the fields from our events already:

fetch logs, from:now() - 5m 
| filter (event.provider == "Falco")

TTP-based threat hunting: Dynatrace automatically infers the fields from our events

Observing TTPs using Dynatrace Security Analytics

For the sake of this demonstration, our internal red team unleashed a novel attack on our Unguard application. The attack lit up our Falco deployment with more than 100,000 events in 24 hours, more than 3,000 of them critical.

As security analysts, we know we can’t find sophisticated attacks by manually scrolling through thousands of audit logs and events. We need automation, full contextual knowledge of our infrastructure, and very often, domain-specific expertise from security analysts.

To get an initial overview, we can use DQL on Grail to visualize what MITRE techniques Falco observed in our infrastructure over the past 72 hours. We can summarize events using mitre.tactic or mitre.technique. These fields exist on many Falco alerts and are automatically ingested by the Dynatrace output of falcosidekick. We can explore the distribution of techniques with this query:

fetch logs, from:now() - 72h 
| filter event.provider == "Falco" and isNotNull(mitre.technique) 
| filterOut in(mitre.technique, {"T1548.001", "T1083", "T1565", "T1055.008"}) 
| summarize event_count = count(), by:{mitre.technique}

Threat hunting technique chart

In this example, we also observe that we can attribute most events to the following MITRE techniques:

After manually investigating these alerts, however, we conclude they’re noisy false positives. Some of our applications were treating environment variables in an insecure way or communicating with the Kubernetes API server with improperly configured service accounts. Therefore, we filtered them out with DQL.

Observability and context: Attributing reconnaissance activity to TTPs using distributed traces

So far in our TTP-based threat hunting, we’ve utilized Dyntrace Security Analytics to visualize ingested alerts from third-party tools.

But truly magical things arise when we combine this with the rich and high-quality observability data that our customers have valued since the beginning of Dynatrace. Using observability data, we can close an important security-relevant gap. Attackers often probe systems using automated scanning tools. Their many access attempts leave behind a lot of traces. Dynatrace PurePath is one of the core platform technologies that captures and analyzes those distributed traces across an entire infrastructure.

Attack sub-technique T1595.003 “Active Scanning: Wordlist Scanning” describes how attackers use scanners to learn about the many endpoints an application might expose to the internet. Typically, they use large lists with well-known path names, where many of them could be potentially vulnerable. Such wordlists often contain common path names, such as wp-admin, .git, or .htaccess. The following query looks for five indicative files and expresses how many of them match with the new recon.confidence field we set up to track wordlist-based scanners. The more matches, the more confident we can be that these requests came from a wordlist-based scanner:

fetch spans, from:now() - 72h 
| filter in(http.target, {"/wp-admin", "/.git", "/.htaccess", "/.ssh", "/cgi"}) 
| summarize { 
    recon.confidence = countDistinct(http.target) / 5, 
    recon.first_seen = min(timestamp), 
    recon.last_seen = max(timestamp) 
  }, by:{host.name, k8s.container.name, k8s.namespace.name, k8s.pod.name} 
| fieldsAdd mitre.technique = "T1595.003", mitre.tactic = "mitre_reconnaissance" 
| filter recon.confidence > 0.5

fetch spans results

Indeed, it did find some reconnaissance attempts! This query scanned 2.5 million spans in less than 50 ms and reduced them to three comprehensible TTP records. With a conventional database, a query that scans millions of records would take many seconds to complete and require that we structure our queries up front. But Grail completes the search across millions of records in milliseconds, and automatically parses and infers the structure for us so we can just start writing queries directly.

We now know that the Envoy proxy in our environment most likely got scanned by an attacker. Let us bring all the bits and pieces together in the next sequence.

Structural multi-step attack detection with Dynatrace Security Analytics

Attackers typically perform many small steps to achieve their mission. Security experts like to think in terms of so-called kill chains, which describe the many stages of an attack. When you work with TTPs, the attack tactics represent those stages. While the MITRE ATT&CK® knowledge base describes as many as 14 tactics, we can distill this into three broad categories:

  1. Land – First, hackers investigate their target, looking for an initial way to gain access and establish a foothold in your system.
  2. Expand – Then, hackers typically try to escalate their privileges and move laterally within your system, compromising neighboring hosts.
  3. Execute – Finally, they find their target and execute their mission.

Next in our demonstration of TTP-based threat hunting with Dynatrace Security Analytics, we’re going to show you a simple but effective strategy that can uncover such advanced attacks: Structural multi-step attack detection. This strategy is structural since it utilizes Smartscape for DQL to take the topological relationship of events into account when hunting for attacks that are composed of multiple steps.

The following DQL query looks for the filtered Falco alerts, and for each Kubernetes pod, it records how many distinct tactics and techniques we just observed. This way, we’re not just looking at whatever pod was the noisiest, but instead, which pod generated alerts from the most tactics and techniques. The more tactics and techniques, the higher the chances that an attacker carried out a full kill chain on that pod.

fetch logs, from:now() - 72h 
| filter (event.provider == "Falco") 
| filterOut in(mitre.technique, {"T1548.001", "T1083", "T1565", "T1055.008"}) 
| summarize { 
    num_tactics = countDistinct(mitre.tactic), 
    num_techniques = countDistinct(mitre.technique), 
    tactics = collectDistinct(mitre.tactic), 
    techniques = collectDistinct(mitre.technique) 
  }, by:{k8s.pod.name} 
| filter num_tactics > 1 and num_techniques > 1 
| sort num_tactics desc, num_techniques desc

threat hunting: attack detection query

This result is highly interesting and confirms our previous suspicion. There is one instance of the Envoy proxy that captured alerts for the following techniques:

Further, the Dynatrace spans we looked at in the previous sequence that explored wordlist scanning indicated TA0043 “Reconnaissance”. This query just scanned through more than 23 million records in 300 ms, providing us with an abstract description of a full kill chain.

But we don’t stop here. We can drill down and observe the individual steps our attacker has taken:

fetch logs, from:now() - 72h 
| filter (event.provider == "Falco") 
| filterOut in(mitre.technique, {"T1548.001", "T1083", "T1565", "T1055.008"}) 
| filter k8s.pod.name == "unguard-envoy-proxy-666464f76d-5p26f" 
| fields timestamp, event.name, mitre.technique, content.output_fields.proc.cmdline 
| sort timestamp asc

TTP threat hunting attack chain query

In the result, we see the records that explain the attack procedure in detail. The above screenshot shows only an excerpt of all 47 records. Here’s what we learned about our attacker’s steps:

  • Scanned our Envoy proxy with a well-known wordlist, as we learned by mining the traces for wordlist entries.
  • Launched a Perl-based reverse shell on Envoy, indicated by the perl command opening a socket, giving them full code execution access.
  • Downloaded a couple of binaries like nmap and nc, indicated by the curl command that pulled them from the internet.
  • Scanned our internal network with nmap.
  • Exfiltrated large volumes of data from our Redis database, indicated by the queries to Redis that request all keys with the KEYS * command
  • Tried to cover their tracks by deleting the shell history, indicated by the rm /home/envoy/.bash_history command

Isn’t this a truly elegant way to hunt for attacks?

TTP-based threat hunting with context-rich observability and security analytics

This demonstration shows how modern attack detection strategies become a reality with context-rich security analytics on a unified observability and security platform. With this approach, you can do the following:

  • Utilize observability data to capture security-relevant reconnaissance alerts and map them to TTPs.
  • Enrich the Dynatrace platform with more data of your own, such as ingesting Falco alerts into Grail.
  • Use DQL and Grail to find the needle in the haystack to scan tens of millions of records in milliseconds to identify the chain of only a handful of events that exposed an attacker and the exact methods they used.

For another great demonstration, we recommend reading the blog post Log forensics: Finding malicious activity in multicloud environments with Dynatrace Grail by Liisa Tallinn.

If this blog post made you eager to try out Dyntrace and learn more about Grail, join us for the on-demand webinar, Get to know Dynatrace: Grail edition.

The post TTP-based threat hunting with Dynatrace Security Analytics and Falco Alerts solves alert noise appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ttp-based-threat-hunting-solves-alert-noise/feed/ 0
OpenTelemetry logs in Grail unlock full observability https://www.dynatrace.com/news/blog/opentelemetry-logs-in-grail-unlock-full-observability/ https://www.dynatrace.com/news/blog/opentelemetry-logs-in-grail-unlock-full-observability/#respond Tue, 11 Jul 2023 20:07:24 +0000 https://www.dynatrace.com/news/?p=58569 OpenTelemetry logs

Dynatrace now offers native support for OpenTelemetry logs, which opens up the ability to collect all your observability data in a single platform and benefit from unified observability with other OpenTelemetry signals. This complements existing Dynatrace support for collecting traces and metrics via OpenTelemetry Protocol (OTLP) and allows you to get actionable answers from log data with the powerful combination of Grail and Dynatrace Query Language.

The post OpenTelemetry logs in Grail unlock full observability appeared first on Dynatrace news.

]]>
OpenTelemetry logs

Without native log support, overhead and complexity grow

OpenTelemetry, the Cloud Native Computing Foundation (CNCF) incubating project, introduced standards that enable companies to instrument, generate, and export telemetry data. When combined with out-of-the-box correlation, such telemetry data provides context-rich observability. Dynatrace has supported the OpenTelemetry project for years as a key contributor and contributed to its rise to a popular open source observability framework for cloud-native software. Many global enterprises have instrumented their code to emit traces, metrics, and logs in a standardized and vendor-neutral way using OpenTelemetry.

While ingestion of OpenTelemetry traces and metrics into Dynatrace is supported, companies often prefer to collect logs in the OpenTelemetry format. Earlier this year, OpenTelemetry announced that their logs API/SDK specification is stable, making it ripe for broader adoption. This enables unified observability because logs are indispensable for troubleshooting apps, monitoring infrastructure, auditing or investigating security incidents, tracking business events, and many other use cases.

Without such a holistic view of system behavior and performance, organizations typically need to invest in separate instrumentation and integration efforts for each telemetry type, which brings additional overhead, costs, and complexity.

Unify OpenTelemetry logs, traces, and metrics in Dynatrace

Dynatrace now includes full support for OpenTelemetry logs, which provides unified observability for organizations with vendor-neutral and open-source tech stacks. Our commitment to this open standard allows you to cover all three pillars of observability with minimal configuration effort because OpenTelemetry traces, metrics, and logs can be exported to Dynatrace using the same OTLP exporter.

By ingesting OTLP logs into Dynatrace, you can utilize the Grail™ data lakehouse and its massively-parallel processing analytics engine. This allows you to eliminate log forwarding and collection solutions, which not only add maintenance overhead and complexity but can also become bottlenecks in performance and log volume throughput.

With the added support of logs to OpenTelemetry traces and metrics, Dynatrace now gives you a unified and holistic overview of observability signals, with integrated linking of traces and logs. This enables you to connect the traces in your stack directly with root-cause information in logs. One customer recently shared why working with all telemetry signals together really makes sense for them.

Native support for OpenTelemetry (OTLP) logs also supports enterprises that have highly diverse technical architectures. While Dynatrace OneAgent® is often the preferred way of discovering and ingesting logs from traditional hosts or Kubernetes environments, there are certain environments for which OneAgent is not a viable option. As an alternative, OpenTelemetry lets you extend Dynatrace technology coverage with log data.

Generic ingest of log data now works with OTLP

OTLP log ingest API

Ingesting OTLP logs is now supported via the OpenTelemetry Logs Ingest API. In SaaS deployments, you can use this approach to ingest log data into Grail and analyze it via Log Management & Analytics in your environment. The same API is also available for Log Monitoring Classic for Dynatrace Managed deployments (via an Environment ActiveGate).

All you need to do is configure your OpenTelemetry collector or any other OTLP log source to send logs to the OpenTelemetry Logs Ingest API endpoint.

Start using the full OpenTelemetry set today

The post OpenTelemetry logs in Grail unlock full observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/opentelemetry-logs-in-grail-unlock-full-observability/feed/ 0
Log forensics: Finding malicious activity in multicloud environments with Dynatrace Grail https://www.dynatrace.com/news/blog/log-forensics-with-dynatrace-grail/ https://www.dynatrace.com/news/blog/log-forensics-with-dynatrace-grail/#respond Mon, 22 May 2023 06:00:17 +0000 https://www.dynatrace.com/news/?p=57709 Logs forensics graphic

Log forensics—investigating security incidents based on log data—has become more challenging as organizations adopt cloud-native technologies. Organizations are increasingly turning to these cloud environments to stay competitive, remain agile, and grow. But as organizations rely more on cloud environments, data and complexity have proliferated. Teams struggle to maintain control of and gain visibility into all […]

The post Log forensics: Finding malicious activity in multicloud environments with Dynatrace Grail appeared first on Dynatrace news.

]]>
Logs forensics graphic

Log forensics—investigating security incidents based on log data—has become more challenging as organizations adopt cloud-native technologies. Organizations are increasingly turning to these cloud environments to stay competitive, remain agile, and grow.

But as organizations rely more on cloud environments, data and complexity have proliferated. Teams struggle to maintain control of and gain visibility into all the applications, microservices and data dependencies these environments generate. Without visibility, application performance and security are easily compromised.

As a result, teams are turning to technologies such as observability to understand events in their cloud environments. Moreover, they have come to recognize that they need to understand data in context. But most observability technologies today provide information in silos. Without unifying these silos, teams miss critical context that can lead to blind spots or application problems—problems that compound when there’s a need to investigate security events.

Modern observability enables log forensics

Dynatrace is a software intelligence platform that provides deep visibility into and understanding of applications and infrastructure. It started as an observability platform; over time, it has expanded to provide real user monitoring, business analytics, and security insights. The recent innovation around log storage, processing, and analysis—Grail—makes Dynatrace a great solution for security use cases such as threat hunting and investigating the who-what-when-where-why-how of an incident.

Grail is a data lakehouse that retains data context without requiring upfront categorization of that data. Unifying data in Grail brings critical security capabilities to bear as teams seek to understand malicious events.

Grail enables organizations to find and analyze security events in the context of their broader cloud environments. Moreover, with capabilities such as log forensics—the analysis of log data to identify when a security-related event occurred—organizations can explore historical application data in its full context.

Grail makes it easy to query historical data without data rehydration or indexing, re-indexing, and up-front schema management. This accessibility gives users quick and precise results about when malicious activity occurred, when reconnaissance was first seen in the systems, what was attempted, and if the attackers were successful.

Demo: “Ludo Clinic” uses log forensics to discover and investigate attacks using Grail

Imagine working as a security analyst for a respectable medical institution called Ludo Clinic. As Ludo Clinic started using Dynatrace, the platform’s Runtime Application Protection feature detected a SQL-injection (SQLI) attack. Thanks to the details provided by the code-level vulnerability functionality, the developers knew where in the code the exploited vulnerability was and were able to address and patch it quickly.

The task now is to investigate whether the system experienced any suspicious activity before the attack so we can determine if any other systems are affected. The good news is there are metrics available a few days before your team detected the attack, and you also ingested three months of application and access logs into Grail. Because these logs are ready for querying with no rehydration, the investigation can start immediately. The bad news is we don’t know exactly what to look for. “Find suspicious activity” can mean anything. So, we’ll start by exploring the data using the hints and context information we already detected with Dynatrace.

Hint one: Blocked SQL injection report details

Here’s the report from Dynatrace on the blocked SQL injection details on 10 February from the IP 104.132.226.34. This report shows details of the attack, such as the entry point, the vulnerability that was exploited, the IP address of the attacker, and so on.

screenshot of Dynatrace blocked SQL injection report showing attack details

Hint two: Failed logins spike

A quick look at the metrics dashboard dating back to 8 February shows a spike in failed logins metrics before your team detected the attack. Indeed, there’s a spike on 8 February.

screenshot of failed logins spike

It would make sense to see if there’s any activity from that IP before we set up monitoring. Has the attacker been doing reconnaissance from that same IP in our systems even before this? If yes, how? Did they try something else during those three months? Were they successful?

Notebooks, DQL, and DPL: Tools of the Grail log forensics trade

Now that we have some clues about where to look for suspicious activity, we’ll dig into the logs using Notebooks, DQL, and DPL.

Dynatrace Notebooks

Dynatrace Notebooks is a collaborative data exploration feature that operates on data stored in Grail for ad-hoc exploratory analytics. Notebooks enable cross-functional teams, such as IT, development, security, and business analysts, to build, evaluate, and share insights for exploratory analytics using code, text, and rich media. The ability to build insights from the same data using the expertise of different roles helps organizations truly understand everything their data has to say.

Dynatrace Query Language (DQL)

Dynatrace Query Language (DQL) is a piped SQL-like query language, similar to Linux commands executed in sequence. You can look at the queries like a series of building blocks applied in an order you happen to need at this moment. Select fields, summarize, and count a value, apply more filters, select additional fields, extract data from a particular field, and so on. DQL is great for exploring and experimenting with data, which makes it a great ally in log forensics and security analytics.

The first query of our investigation uses DQL in Notebooks to fetch logs from Grail, filter the access log, and limit the result to 1000 records for initial exploration.

log forensics using Notebooks to start the investigation

Dynatrace Pattern Language (DPL)

In our investigation, we’ll also use DPL. DPL stands for Dynatrace Pattern Language, a parsing language that also consists of intuitive building blocks that help to extract meaningful fields from data on read. That means there is no need to manage indexes and rehydrate archived data; simply specify an ad hoc schema using DPL as part of the query.

What’s more, with DPL, the parsed results return typed fields, so you can be sure that a timestamp is a timestamp and an IP address is an IP address, not some random octets separated by a dot like 320.255.255.586. Working with typed data means excellent quality and precision for investigation results because you can run type-specific queries like calendar operations, calculations on numeric data, working with JSON objects, and so on. Working with typed data means excellent quality and precision for investigation results, as you can run type-specific queries like calendar operations, calculations on numeric data, working with JSON objects, and so on.

Log forensics: Querying the access log

Remember: our task is to investigate whether any other systems are affected. The first step is to query whether the IP address where the SQLI attack came from has been used before. Can we see it in the web application access log months prior to the attack?

screenshot of log forensics query of the access log using Dynatrace Grail

The query result shows there is no activity from that IP earlier than records on 10 February, the day Dynatrace detected the SQL injection attack. This means that the attack appeared “out of the blue,” and it is likely the attackers were using other IP addresses to do reconnaissance on our systems.

Because we can’t find the attacker by the IP address, let’s look at abnormalities in login behavior because there is a chance they’ll be related to reconnaissance. This means we’ll investigate the spike in failed logins we saw earlier in the metrics graph. Are there any other failed login spikes three months prior to the attack? Where do the failed logins originate from?

We can see that a failed login attempt takes users to a specific URL:

/ludo-clinic/login?authenticationFailure=true

So, let’s see if and how often this URL appears in the logs by adding the following filter to the query.

| filter contains(content, "/ludo-clinic/login?authenticationFailure=true") 
| limit 10000

Indeed, the query gives us 6531 records containing a failed login URL:

screenshot of log forensics query result showing 6531 records

Making sense of the access log

For the next stage of our investigation, let’s make more sense of these ~6,300 records and find out how many unique IP addresses were the origin of failed login attempts. The hypothesis is: some of the IP addresses stand out when it comes to the number of login failures. This means we first need to extract the IP address to run this aggregation.

We can utilize the schema-on-read functionality, that is, extract only the fields we need for a specific query. Taking a closer look at the content field of the access log, we can see a traditional HTTP access log: clientIP, timestamp, requestURL, HTTP response code, and so on.

screenshot of query results showing extracted fields clientIP, requestURL, HTTP response code, and so on

Notice that the timestamp field (the ingest timestamp) is similar for all log records (15/05/2023 14:09:48). This is because Ludo ingested the historical log records in bulk. To analyze the event time, we need to extract the timestamp from the content field as event time. To count IP addresses, we also need the IP address.

Quick ad-hoc parsing to aggregate login failures

To parse out data (timestamp and IP) from the content field, we’ll select the content field and select Extract fields to open the DPL Architect. To retrieve the timestamp and clientIP, we’ll replace the default DPL pattern with the following:

IPADDR:client_ip LD HTTPDATE:event_time

screenshot showing ad-hoc parsing timestamps and clientIP using the DPL architect

This matches and extracts the timestamp and the IP address from the content field and gives them a name (event_data and client_ip) and leaves the rest of the pattern unmatched, as we don’t need it for the following query. Clicking Insert pattern brings us back to the query view, adding a parse command to the newly created pattern.

screenshot showing results of parsing fields in DPL architect

Now with the extracted IP address, we can proceed with queries and use the summarize command to count the number of failed logins per unique IP address to see if there were any failed logins originating from a specific IP. Sorting the result set based on the number of failed logins in descending order gives us the largest outliers.

fetch logs 
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, "IPADDR:client_ip LD HTTPDATE:event_time" 
| fields event_time, client_ip, content 
| summarize total=count(), failed=countIf(contains(content, "/ludo-clinic/login?authenticationFailure=true")), by:client_ip 
| sort failed desc

This pays off! The results reveal that a significant portion of logins (181,774) and failed logins (6161) originate from the IP address: 172.31.24.11. This seems interesting and is worth taking a closer look.

screenshot showing the count of failed logins from the originating IP address

Timing of login failures

Next in our log forensics journey, let’s see when these failed logins from that particular IP address occurred to get more information on the potential reconnaissance activity. Did the requests all occur within a short period or regularly across a longer period?

Because we’re interested in the behavior of a specific IP and investigating the reasons behind failed logins, let’s also extract the session ID from the log line. As the session ID is the only field that occurs both in the access and application weblogs, it will be also useful later when we need to join the two for investigating affected users.

We already extracted the IP and included the timestamp (HTTPDATE). We will now extend the pattern and skip the part of the record we don’t need by not naming the three double-quoted strings (DQS). Finally, we’re extracting the last field that contains the session ID.

IPADDR:client_ip LD HTTPDATE:event_time LD DQS LD DQS LD DQS SPACE LD:session_id

screenshot showing a query that extracts session IDs involving the target IP address

When we select Insert pattern, we again get a parse command populated with the DPL pattern we just created.

Focusing on the suspicious IP

Next, let’s select only the fields we’re interested in and then aggregate fields. These actions reveal more about the extended activity that involves the suspicious IP address responsible for many of the failed logins.

| fields time, client_ip, session_id, content

screenshot showing the results of a query that extracts session IDs involving the target IP address

Filtering the attacker IP and sorting the fields based on the timestamp we just parsed out, it appears this IP address was first seen on 24 December 2022. We now know the start of the suspicious activity. For malicious actors, it is quite common to act during the holiday period.

fetch logs  
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, " IPADDR:client_ip LD HTTPDATE:event_time LD DQS LD DQS LD DQS SPACE LD:session_id" 
| fields event_time, client_ip, session_id, content 
| filter contains(content,"172.31.24.11") 
| sort event_time asc 
| limit 300000

screenshot showing a query that extracts the event times involving the target IP address

Find the suspicious activity pattern across time

To see the activity pattern of this suspicious IP across time, let’s count the number of failed logins in one-hour time intervals.

fetch logs 
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, "IPADDR:ip LD HTTPDATE:time LD DQS LD DQS LD DQS SPACE LD:sessionID EOS" 
| fields time, ip, sessionID, content 
| filter contains(content,"172.31.24.11") 
| summarize failed=countIf(contains(content, "/ludo-clinic/login?authenticationFailure=true")), by:bin(time, 1h)

screenshot showing a query that counts the number of failed logins involving the target IP address

It appears as though failed logins from this IP appear to follow a very regular pattern: 24 failed attempts every hour. Looks like this activity is automated and most probably refers to a dictionary attack: regular (automated) attempts from the attacker to try out different usernames and passwords, mostly with failed results.

But to escape the clinic’s countermeasures (failed login attempts velocity check), the attacker also conducts a successful login every now and then. If we count all activity from that IP address (not just the failed logins but successful attempts as well), the results are again very symmetrical:

fetch logs 
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, "IPADDR:ip LD HTTPDATE:time LD DQS LD DQS LD DQS SPACE LD:sessionID EOS" 
| fields time, ip, sessionID, content 
| filter toString(ip) == "172.31.24.11" 
| summarize count=count(), by:bin(time, 1h)

screenshot showing a query that returns all logins from the target IP address.

Identify targeted users

Next, it would be useful to know which users the attacker has targeted and whether any attempts have been successful. Let’s aggregate the activity from this IP using sessionIDs:

fetch logs 
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, "IPADDR:ip LD HTTPDATE:event_time LD DQS LD DQS LD DQS SPACE LD:sessionID EOS" 
| fields timestamp, event_time, ip, sessionID, content 
| filter contains(content, "172.31.24.11") 
| summarize count=count(), 
            by:{sessionID 
               } 
| sort count desc 
| limit 10000

The result is again peculiar, suggesting automated activity: 59 log lines per session.

screenshot showing a query that identifies logins by session ID that suggests automated activity.

Next log forensics dataset: The web application log

Next, let’s see what was happening based on the web application log, using data from what was going on during those sessions that originated from the suspicious IP address we discovered from the access log dataset.

First, to familiarize ourselves with the content of the webapp log, let’s run a basic query to see what the content field of the web application log looks like:

screenshot showing a log forensics query that shows content of the web application log.

We can see a timestamp, log severity, traces and spans, a session ID, result, and username. There are plenty of interesting fields to play with, so the next step is to parse the content into fields that are ready for querying. We can extract the fields using the DPL Architect. The following DPL pattern extracts the event time session ID, result, and username from the webapp log.

'[' TIMESTAMP('dd/MMM/yyyy:HH:mm:ss.S'):event_time LD ' - ' LD:sessionIdApp ' ' LD:result_text ': ' LD:username (';' | EOS)

screenshot showing a query that parses out the fields of interest for the log forensics

Inserting the pattern, this is what the query looks like when parsing out session IDs and usernames from the web application log.

fetch logs, from:-300d   
| filter contains(log.source, "ludo-clinic-webapp.log") 
| parse content, "'[' TIMESTAMP('dd/MMM/yyyy:HH:mm:ss.S'):event_time LD ' - ' LD:sessionIdApp ' ' LD:result_text ': ' LD:username (';' | EOS) " 
| limit 10000 
| fields event_time, sessionIdApp, result_text, username

screenshot showing the results of parsing the fields of interest in the web application log

Filter out records from the authentication provider

To see which users were targeted and how successful the attacker was, we will continue working with only those web application records that contain authentication responses. First, we filter out the records that originate from the authentication provider, then we skip the responses we’re not interested in:

| filter contains(content, "CustomAuthenticationProvider")  
  AND NOT contains(content, "Starting findUsersByUsernameAndPassword") // we want to see only auth response log records 
  AND NOT contains(content, "retrieved matching list")

The full query now looks like this and returns the following results:

fetch logs  
| filter contains(log.source, "ludo-clinic-webapp.log") 
| filter contains(content, "CustomAuthenticationProvider")  
  AND NOT contains(content, "Starting findUsersByUsernameAndPassword") // we want to see only auth response log records 
  AND NOT contains(content, "retrieved matching list") 
| parse content, "'[' TIMESTAMP('dd/MMM/yyyy:HH:mm:ss.S'):event_time LD ' - ' LD:sessionIdApp ' ' LD:result_text ': ' LD:username (';' | EOS) " 
| limit 10000 
| fields event_time, sessionIdApp, result_text, username

screenshot showing the full query with the fields of interest from the web application log.

Review the user sessions that originate from the attacker

Next, to see which user sessions in the webapp log originated from the attacker’s activity, we use a lookup query to join aggregated sessions from the attacker IP address we discovered in the access log with sessions in the webapp log. In short, we will see what was happening during the suspicious sessions according to the webapp log.

fetch logs  
| filter contains(log.source, "ludo-clinic-webapp.log") 
| fields content 
| filter contains(content, "CustomAuthenticationProvider")  
  AND NOT contains(content, "Starting findUsersByUsernameAndPassword") // we want to see only auth response log records 
  AND NOT contains(content, "retrieved matching list") 

| limit 10000 
| parse content, "'[' TIMESTAMP('dd/MMM/yyyy:HH:mm:ss.S'):event_time LD ' - ' LD:sessionIdApp ' ' LD:result_text ': ' LD:username (';' | EOS) " 
| lookup [fetch logs, from:-300d 
                | filter contains(log.source, "ludo-clinic-access.log") 
                | filter contains(content,"172.31.24.11") 
                | parse content, "IPADDR:ip LD HTTPDATE:time LD DQS LD DQS LD DQS SPACE LD:sessionID EOS" 
                | filter isNotNull(sessionID) 
                | limit 100000 
                | summarize accesscount=count(), by:{sessionID} 
                | fields sessionID], sourceField:sessionIdApp, lookupField:sessionID 

| fieldsRemove content

Screenshot showing the results of attempted authentications.

The result shows the attacker has achieved both successful authentications as well as failed authentications. Finally, we see which usernames the attack targeted the most by looking for the response “No users found requested username.” The system returns this value when it receives a non-existent user or a wrong password. By aggregating the result based on unique usernames, we get a list of the most (unsuccessfully) targeted users.

| filter result_text == "No users found requested username" 
| summarize count(), by:{username} 
| sort `count()`desc

screenshot showing the no users found query that reveals the targeted user accounts

These results are fascinating – we can see five usernames that the attacker continuously entered and received failed authentication results. We can also see SQL commands instead of regular usernames.

Determine successfully targeted users

Next question: Did they achieve anything besides ‘No users found’ when targeting these users? Let’s have a look by excluding the “No users found requested username” response and concentrating on those five users from the last query result, and adding the following line to the query:

|  filter not matchesPhrase (result_text, "No users found requested username") and in (username, "arnie", "herman", "krzysztofs", "bernice", "sherry")

screenshot showing drilldown to identify affected usernames.

Indeed, we see a lot of “successfully authenticated” responses in the result text field. This confirms the attackers were successfully conducting a dictionary attack: trying out several usernames and passwords to authenticate as real users of Ludo Clinic. When counting the number of successful authentications per these five users, the results are quite similar:

|  filter not matchesPhrase (result_text, "No users found requested username") and in (username, "arnie", "herman", "krzysztofs", "bernice", "sherry") 
| summarize count(), by:{username} 
| sort `count()`desc

screenshot showing top targeted users with successful authentication.

Investigation results from log forensics and metrics with Dynatrace

As a result of Dynatrace detecting a SQL vulnerability, anomalies in metrics, and subsequently running forensic queries on three months of logs prior to the attack, we’ve been able to construct the following timeline:

  • As Ludo Clinic started using Dynatrace, they were able to observe a spike in metrics capturing failed logins on 8 February
  • The system detected and blocked a SQL injection attack on 10 Feb (Fri)
  • There was no other activity from that IP in the access logs (the logs reach back three months)
  • When aggregating failed login activity, we discovered the following details: the IP address 172.31.24.11 stands out from the rest, counting to 6161 failed logins based on the access log during the past three months
  • This IP was first seen in the logs on 24 December 2022 (the earliest timestamp for this set of logs is 20 November 2022)
  • The sessions contain identical activities during identical timeframes, which suggests the attacker was using automated tools
  • Joining sessions from the access log to the application log reveal almost ten thousand records originating from the suspicious IP 172.31.24.11
  • It looks like the attackers were attempting a dictionary attack because it targeted several users at regular intervals, resulting in failed as well as successful authentications.
  • Users stafford, ray, orrel, doug and joby were targeted to discover their passwords. The attacker successfully authenticated 35-37 times per user.
  • The attackers also entered SQL statements instead of usernames attempting SQL injection attacks.

The DQL and DPL advantage

For such investigations, DQL and DPL make it convenient to quickly investigate and query logs for security analytics use cases that require drawing broad conclusions from the data one minute and then zooming into the activities of a specific session the next. An interesting find inspires the analyst to parse out yet another field and run aggregations on this data. As historical data is always ready for querying, all hypotheses can be quickly verified or dismissed. A curious mind and the right log forensics tools (DQL and DPL) make a great combination for fighting evil.

To see more of Grail in action for log forensics and exploratory analytics, join us for the Observability Clinic: The Practitioner’s Guide to Analytics without Boundaries with Dynatrace.

The post Log forensics: Finding malicious activity in multicloud environments with Dynatrace Grail appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/log-forensics-with-dynatrace-grail/feed/ 0
Log auditing and log forensics benefit from converging observability and security data https://www.dynatrace.com/news/blog/log-auditing-and-log-forensics/ https://www.dynatrace.com/news/blog/log-auditing-and-log-forensics/#respond Thu, 13 Apr 2023 14:26:36 +0000 https://www.dynatrace.com/news/?p=56968 log analytics log output for log management best practices

Log auditing and log forensics are essential practices for securing apps and infrastructure. But the complexity of cloud-native environments requires a new approach to keep investigations real-time and relevant. Converging observability and security data gives security teams end-to-end visibility into application security issues for real-time answers at scale.

The post Log auditing and log forensics benefit from converging observability and security data appeared first on Dynatrace news.

]]>
log analytics log output for log management best practices

The growing complexity of multicloud environments and ever-increasing number of application vulnerabilities have made it harder than ever to protect against attacks. Log auditing—and its investigative partner, log forensics—are becoming essential practices for securing cloud-native applications and infrastructure.

Many organizations don’t know for months (or years) after a security attack when, why, or how it happened. This represents a significant risk, with the same attack vector repeatedly exploited because the vulnerability wasn’t detected on time. The massive volumes of log data associated with a breach have made cybersecurity forensics a complicated, costly problem to solve.

As organizations adopt more cloud-native technologies, observability data—telemetry from applications and infrastructure, including logs, metrics, and traces—and security data are converging. This alignment provides security teams with the opportunity to track application security issues through the ever-increasing volume and variety of log data. Together, this data makes teams more effective in identifying and responding to critical security incidents as quickly as possible. Overall, this results in a better security posture. Let’s explore how a log auditing and log forensics program can benefit from the convergence of observability and security data.

What is log auditing?

Log auditing is a cybersecurity practice that involves examining logs generated by various applications, computer systems, and network devices to identify and analyze security-related events. Logs can include information about user activities, system events, network traffic, and other various activities that can help to detect and respond to critical security incidents.

Log auditing is a crucial part of building a comprehensive security program. Log auditing helps ensure that teams are following security policies and procedures and that they are identifying and addressing any anomalies or suspicious activities in a timely manner.

What is log forensics?

Log forensics is a practice that involves collecting, analyzing, and preserving log data to identify the time a security incident was initiated, who initiated the incident, the sequence of actions they took, and the impact it had on an organization. It also helps to identify the data that has been affected by an attack and to identify the attack pattern.

Traditionally, log forensics has been based primarily on logs that can help teams to identify the source of a cyberattack, the surface area of damage, and any other relevant details surrounding the nature of the attack. Forensics is crucial for incident response and post-incident analysis. Forensics allows organizations to learn from critical security incidents and take proactive steps to prevent similar recurrences in the future.

Together, log auditing and log forensics are critically important components of security best practices, as they help organizations detect, respond, and recover from security incidents. However, cloud-native technologies have introduced a level of complexity that make a logs-only approach to auditing and forensics limiting. t’s also now critical for organizations to have detailed observability data to improve the quality and context of security investigations.

Cloud complexity introduces new challenges to security audit and forensics

Log auditing and forensics of cloud infrastructure requires a different approach and capabilities compared with traditional on-premises environments. It requires an understanding of cloud architecture and distributed systems, with the goal of automating processes.

Organizations face many challenges when it comes to log audit and forensics, including the following:

  • The large volume of data. Cloud infrastructure and applications scale dynamically, which generates a large volume of logs at high This can make it challenging to process, store, and analyze them in real time.
  • Distributed and complex topologies. In the cloud, infrastructure components are often distributed across multiple regions, availability zones, and even multiple cloud providers. This can make it difficult to understand the relationships between different entities, identify root cause, and determine resolution.
  • Incomplete. Siloed data, incomplete traces, lack of context, and insufficient instrumentation and metrics are all factors that lead to organizations needing more trustworthy, automatable answers in critical security investigations.
  • Skills and expertise. Efficient and effective log audit and forensics practices can require specialized understanding of cloud environments, applications, and log formats. This expertise may exist in teams that may not have the bandwidth to provide them for security incident response.
  • Time and resources. Typically, the process of analyzing log data includes data rehydration, reloading, reindexing, or re-ingesting, which takes time and resources. Through this lengthy process and due to the pressure of an ongoing security investigation, teams are often expected to provide prompt answers to these questions, which conflicts with the additional time needed to conduct a precise analysis of all security incidents.

Organizations need to be aware of these challenges and take steps to address them to ensure their log audit and forensics programs are effective in detecting and responding to critical security incidents. An observability approach, one that covers logs comprehensively along with metrics, end-to-end traces, and real-time context, can enable teams to keep pace with the complexity.

Need security answers now? Observability context can help provide them quickly

Let’s consider an organization that is conducting an investigation of a current or suspected security incident. The company has applications that produce a high volume of logs per day, and a new wave of attacks has targeted this company’s sector. In the aftermath of a critical zero-day vulnerability, such as Log4Shell, it’s vital for teams to determine whether the vulnerability affects them and to identify signs of compromise. Since there is a possibility of the vulnerability existing for weeks or even months prior to discovery, an organization’s investigation should be thorough and span all logs, including a historical span.

Teams can quickly answer questions such as the following by querying not only logs but also observability data (including traces and metrics) and topology context.

  • Were there attack attempts? (for example, query web server logs from the past year for specific attack strings containing ).
  • Are there any indicators of compromise? (for example, query topology to cross reference entity information to narrow down attacked services to those that are running Java)
  • To what extent could we be compromised? (for example, collate which and how many Java applications were attacked)
  • Did we lose any critical data? (for example, query logs for all attacked java applications to find if there were there any suspicious outgoing connections)
  • What data did we lose? (for example, what was the payload on these outgoing connections?)
  • How can we protect ourselves against future attacks? (for example, query application traces to understand how an attack progressed through the application)

How to boost log auditing and log forensics with observability data from Dynatrace

Log auditing and log forensics can be intimidating and complex. But with a platform approach to log analytics based on observability at a cloud-native scale, organizations can accomplish much more.

Dynatrace Grail alleviates the burden of identifying security risks in multicloud and hybrid cloud infrastructures. Grail magnifies Dynatrace Application Security capabilities by enabling teams to make boundless queries of all observability data types. Dynatrace Query Language (DQL) offers a superior approach to query data and includes a unique high-performance Dynatrace Pattern Language (DPL) for easier and faster parsing and data matching. With Grail, Dynatrace customers can now leverage a unified point of governance for their DevSecOps strategies by doing the following:

  • Automatically contextualizing security data with Dynatrace observability insights.
  • Performing collaborative analysis through Dynatrace Notebooks, a method of context-based data sharing.
  • Automating workflows for repetitive tasks and utilizing customized Dynatrace data analysis apps with the new AppEngine.
  • Operationalize data findings in one unified platform, connecting data from development and runtime environments.
  • Eliminate data silos between DevSecOps teams and increase data sharing and analysis capabilities.

With these capabilities, Grail enables customers to extend the existing OneAgent-focused Application Security solution with an end-to-end data analytics-driven approach, and unique convergence of observability and security data in context. Dynatrace Grail can help organizations overcome cloud complexity through instant, cost-efficient, AI-powered analytics for observability, security, and business data at any scale.

Overcome cloud complexity through instant, cost-efficient, AI-powered analytics for observability, security, and business data at any scale with Grail.

The post Log auditing and log forensics benefit from converging observability and security data appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/log-auditing-and-log-forensics/feed/ 0
Observability vs. monitoring: What’s the difference? https://www.dynatrace.com/news/blog/observability-vs-monitoring/ https://www.dynatrace.com/news/blog/observability-vs-monitoring/#respond Thu, 23 Feb 2023 23:45:46 +0000 https://www.dynatrace.com/news/?p=46949 Observability pillars include logs, metrics, and traces.

Organizations are depending more on distributed architectures to provide application services. This trend is prompting advances in both observability and monitoring. But exactly what are the differences between observability vs. monitoring? Understanding when something goes wrong along the application delivery chain is essential so you can identify the root cause and correct it before it […]

The post Observability vs. monitoring: What’s the difference? appeared first on Dynatrace news.

]]>
Observability pillars include logs, metrics, and traces.

Organizations are depending more on distributed architectures to provide application services. This trend is prompting advances in both observability and monitoring. But exactly what are the differences between observability vs. monitoring?

Understanding when something goes wrong along the application delivery chain is essential so you can identify the root cause and correct it before it impacts your business. Monitoring and observability provide a two-pronged approach. Monitoring supplies situational awareness, and observability helps pinpoint what’s happening and what to do about it.

To better understand of observability vs. monitoring, we’ll explore the differences between the two. Then we’ll look at how you can best utilize both to improve business outcomes.

Monitoring vs. observability

First, let’s define what we mean by observability and monitoring.

What is meant by monitoring?

By textbook definition, monitoring is the process of collecting, analyzing, and using information to track a program’s progress toward reaching its objectives and to guide management decisions. Monitoring focuses on watching specific metrics. Logging provides additional data but is typically viewed in isolation of a broader system context.

What is meant by observability?

Observability is the ability to understand a system’s internal state by analyzing the data it generates, such as logs, metrics, and traces. Observability helps teams analyze what’s happening in context across multicloud environments so you can detect and resolve the underlying causes of issues.

What is the difference between observability and monitoring?

Monitoring is capturing and displaying data, whereas observability can discern system health by analyzing its inputs and outputs. For example, we can actively watch a single metric for changes that indicate a problem — this is monitoring. A system is observable if it emits useful data about its internal state, which is crucial for determining the root cause.

What are the similarities between observability and monitoring?

Observability and monitoring are closely related concepts in systems and software engineering. Both aim to provide insights into the health, performance, and behavior of a system. They utilize data collection, analysis, and visualization techniques to enable proactive detection and troubleshooting of issues. Ultimately, they empower engineers to ensure system reliability, performance optimization, and efficient resource utilization.

Between observability and monitoring, which is better?

So how do you know which model is best for your environments?

Monitoring typically provides a limited view of system data focused on individual metrics. This approach is sufficient when systems failure modes are well understood. Because monitoring tends to focus on key indicators such as utilization rates and throughput, monitoring indicates overall system performance. For example, when monitoring a database, you’ll want to know about any latency when writing data to a disk or average query response time. Experienced database administrators learn to spot patterns that can lead to common problems. Examples include a spike in memory utilization, a decrease in cache hit ratio, or an increase in CPU utilization. These issues may indicate a poorly written query that needs to be terminated and investigated.

Conventional database performance analysis is simple compared to diagnosing microservice architectures with multiple components and an array of dependencies. Monitoring is helpful when we understand how systems fail, but as applications become more complex, so do their failure modes. It is often not possible to predict how distributed applications will fail. By making a system observable, you can understand the internal state of the system and from that, you can determine what is not working correctly and why.

However, correlations between a few metrics often do not diagnose incidents in modern applications. Instead, these modern, complex applications require more visibility into the state of systems, and you can accomplish this using a combination of observability and more powerful monitoring tools.

The “three pillars” of observability and beyond

As mentioned earlier, traditionally, observability is understanding what’s happening inside a system from its logs, metrics, and traces. Modern observability includes these three original pillars along with user experience and security. Systems are observable when they generate and readily expose the type of data that enables you to evaluate the state of the system. Here’s a closer look at logs, metrics, distributed traces, user experience, and security.

The pillars of observability
The pillars of observability
  • Logs include application- and system-specific data that details the operations and flow of control within a system. Log entries describe events, such as starting a process, handling an error, or simply completing some part of a workload. Logging complements metrics by providing context for the state of an application when metrics are captured. For example, log messages might indicate a large percentage of errors in a particular API function. At the same time, metrics on a dashboard are showing resource exhaustion issues, such as a lack of available memory. Metrics may be the first sign of a problem, but logs can provide details about what is contributing to the problem and how it impacts operations.
  • Metrics in this context are sets of measurements taken over time, and there are a few types:
    • Gauge metrics measure a value at a specific point in time, such as the CPU utilization rate at the time of measurement.
    • Delta metrics capture differences between previous and current measurements, such as a change in throughput since the last measurement.
    • Cumulative metrics capture changes over time — for example, the number of errors returned by an API function call in the last hour.
  • Distributed tracing is the third pillar of observability and provides insights into the performance of operations across microservices. An application may depend on multiple services, each with its own set of metrics and logs. Distributed tracing is observing requests as they move through distributed cloud environments. In these complex systems, traces highlight any problems that can happen with the relationships among services.
  • User experience considers how users interact with the front end; understanding where time is spent and which actions are critical helps prioritize and identify users’ needs. This is essential when the goal is to deliver an exceptional customer experience. This important piece of the puzzle takes into consideration things like revenue, conversions, and customer engagement. All of these are important inputs to get a full understanding of the application landscape.
  • Security is an essential component in understanding the internal state of a system. Organizations are shifting away from siloed security teams and taking a DevSecOps approach. This includes security at each stage of the SDLC. So should it be in observability, where security is one element that affects the health, performance, and customer experience of an application.

True observability, however, relies on more data than just key indicators.

Why monitoring and observability need a next-gen approach

When trying to effectively monitor, manage, and improve complex microservices-based applications, observability and monitoring are both vital. Monitoring and observability represent a continuum from basic telemetry of single servers to profound insights about complete applications and dependencies.

Many organizations start with monitoring and realize these tools lack contextual insights. Context is critical to understanding why problems exist and how they impact the business. Organizations look to observability to provide the data they need for contextual analysis. Understanding the problem means they can understand the root cause and its effects.

DevOps practitioners need help to maintain highly available and scalable applications. That’s because these complex, interdependent systems behave in unpredictable ways and issues originate from sources that are often not apparent. Practices and tools that worked when we built monolithic applications simply can’t handle the level of data distributed environments generate. They don’t ingest enough data or provide enough insight into the state of applications to understand how to correct problems quickly. Luckily, some tools and practices address these challenges.

An automatic and intelligent approach to monitoring and observability

An advanced software intelligence solution like Dynatrace automatically collects and analyzes highly scalable data to make sense of these sprawling multicloud environments. Dynatrace’s causal AI engine, Davis, sifts through massive volumes of disparate, high-velocity data streams, and analyzes them through a unified interface. This single source of truth tears down information silos that traditionally separate teams that perform different functions on many application components. This centralized, automatic approach eliminates the need for manual diagnostics. It also provides paths to remediation to keep the technology users rely on functioning smoothly.

Learn more about observability vs. monitoring

Check out this Dynatrace eBook!

Explore Dynatrace OpenTelemetry observability

Register now for the on-demand power demo!

Incorporate OpenTelemetry into your observability strategy

Learn more now!

The Developer’s Guide to Observability

Read the full eBook now!

The post Observability vs. monitoring: What’s the difference? appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/observability-vs-monitoring/feed/ 0
Dynatrace memory analysis helps Product Architects identify unknown unknowns https://www.dynatrace.com/news/blog/dynatrace-memory-analysis-helps-product-architects-identify-unknown-unknowns/ https://www.dynatrace.com/news/blog/dynatrace-memory-analysis-helps-product-architects-identify-unknown-unknowns/#respond Thu, 09 Feb 2023 17:34:58 +0000 https://www.dynatrace.com/news/?p=56060 Dashboard graphic

Excessive memory allocations or leaks can harm your organization’s clusters and lead to crashes or unresponsive services. To avoid this, it’s essential to monitor your KPIs for memory allocation and object churn as measures of the performance and health of a system.

The post Dynatrace memory analysis helps Product Architects identify unknown unknowns appeared first on Dynatrace news.

]]>
Dashboard graphic

Luckily, Dynatrace provides in-depth memory allocation monitoring, which allows fine-grained allocation analysis and can even point to the root cause of a problem.

While memory allocation analysis can show wasteful or inefficient code, it can also reveal different problems, one of which we’ll examine in this blog post. This real-world use case, caused by an issue in a customer environment, illustrates how Dynatrace memory analysis capabilities can contribute to root cause analysis within a Dynatrace Cluster.

The typical ratio is about 1.5X higher, but now it’s 3X higher—why?

At Dynatrace, we use dashboards to get a quick overview of the status of monitored services. One such dashboard is the Allocations dashboard which gives an overview of memory usage and allocations for an entire production environment, grouped by APIs.

We recently extended the pre-shipped code-level API definitions to group logical parts of our code so they’re consistently highlighted in all code-level views. For instance, everything related to our correlations engine is dark orange, and the different protocols are mustard colored. Another benefit of defining custom APIs is that the memory allocation and surviving object metrics are split by each custom API definition. So we can easily keep track of them on the Allocations dashboard.

One day while looking at a single cluster, we saw that the memory allocations were abnormally high. While the amount of bytes allocated for the Java API is typically 1.5X the average, in this case, the allocation for the Java API was more than 3X higher than the average, 41 TiB. What could be causing this?

Allocation Bytes dashboard in Dynatrace screenshot

We looked at one of the Dynatrace instances to investigate what was going on. Garbage collection suspension and CPU usage looked healthy. We know from experience that an average value of ~1% GC suspension is healthy, so it was still unclear what was causing the high number of allocations shown on the dashboard.

In Memory profiling view, we would normally expect to see allocations for protocols and database calls at the top of the list of allocation hotspots. In this case, all the top contributors are located in the cluster platform code (as shown by the package names).

Memory profiling All allocations in Dynatrace screenshot

Looking at the call stack of the top allocation, a familiar message handler can be identified, AgentClusterRuntimeInfoMsgHandler. This handler is responsible for sending configuration updates regarding usable communication endpoints (in other words, available ActiveGates) to connected OneAgents. Typically, the configuration does not change, and no responses are created for the OneAgents. In this case, the server appears to be continuously building responses, which is an expensive operation that indicates either we have a bug in the revision calculation of our message handler, or the list of ActiveGates is constantly changing, forcing frequent revision recalculation.

Selecting Called Methods next to the message handler opens the profiling view, which shows the full extent of the impact. The handler is responsible for ~3.5 TiB in allocations within 2 hours, allocating and removing about 75 billion objects during the process.

Profiling view of called methods in Dynatrace screenshot

Verification with Dynatrace custom metrics

As Dynatrace also exposes key metrics about our message handler via JMX, we can use those metrics to investigate further. In Further Details on the Host page, we instantly have the confirmation we’re looking for: We were constantly sending ~4.5MiB/s of ClusterRuntimeInfo responses, while on a healthy system the response size is typically 50KiB/s or less (depending on the number of connected agents).

Since other production systems are doing fine at the same time, a bug in the code might not be the problem. Instead, we investigate to see if we have many recalculations due to constantly changing ActiveGate connections.

Luckily, we have an audit log for ActiveGate connectivity on the Dynatrace Cluster, which can be seen in the log viewer.

Audit log for ActiveGate connectivity on the Dynatrace Cluster

Finding the root cause of the problem

In the audit log file, we can see that many ActiveGate registration and deregistration activities are taking place. By adding a filter for a single ActiveGate ID and increasing the timeframe, a pattern emerges: this ActiveGate is reconnecting once per hour.

ActiveGate registration and deregistration activities in audit log file

The other ActiveGates do the same at separate times, which explains the server behavior: every time an ActiveGate connects or disconnects, the endpoint list changes and so must be resent to the deployed OneAgents. The customer has more than 100 thousand OneAgents connected, which consumes many resources on the server and, more importantly, on the network. Following these insights, we contacted this customer to share our findings.

Conclusion

Memory allocation analysis can show wasteful or inefficient code, but it can also reveal unexpected problems, such as, in this case, numerous configuration updates sent out due to a problem on the customer side. Even though the server could easily handle the memory allocations (GC suspension was around 1%), the allocations showed up prominently, and they can be seen as an indicator of bugs in the system.

You can find out more about Dynatrace memory allocation analysis in our documentation:

New to Dynatrace?

Visit our trial page for a free 15-day Dynatrace trial.

The post Dynatrace memory analysis helps Product Architects identify unknown unknowns appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-memory-analysis-helps-product-architects-identify-unknown-unknowns/feed/ 0
Three smart log ingestion strategies in Dynatrace https://www.dynatrace.com/news/blog/three-smart-log-ingestion-strategies/ https://www.dynatrace.com/news/blog/three-smart-log-ingestion-strategies/#respond Thu, 15 Dec 2022 20:58:53 +0000 https://www.dynatrace.com/news/?p=55224 AppEngine: Create custom apps for data insights

Getting precise answers from log monitoring platforms gets challenging as cloud environments expand and grow more complex. Here are three log ingestion strategies to achieve scale in the Dynatrace platform—without OneAgent.

The post Three smart log ingestion strategies in Dynatrace appeared first on Dynatrace news.

]]>
AppEngine: Create custom apps for data insights

While many organizations have embraced cloud observability to better manage their cloud environments, they may still struggle with the volume of entities that observability platforms monitor. The key to getting answers from log monitoring at scale begins with relevant log ingestion at scale.

Engaging the automatic instrumentation of the Dynatrace OneAgent makes log ingestion automatic and scalable. However, our customers often have set up multiple other log ingestion methods. This flexibility enables logs from diverse environments and established configurations to complete the observability picture for automated troubleshooting and monitoring in Dynatrace.

In this blog, we share three log ingestion strategies from the field that demonstrate how building up efficient log collection can be environment-agnostic by using our generic log ingestion application programming interface (API).

As with all other log ingestion configurations, these examples work seamlessly with the new Log Management and Analytics powered by Grail that provides answers with any analysis at any time.

Log ingestion strategy no. 1: Welcome syslog, with the help of Fluentd

Syslog is a popular standard for transporting and ingesting log messages. Typically, these are streamed to a central syslog server. One option is to install OneAgent on that syslog server, which automatically discovers, instruments and sends the log data to the Dynatrace platform.

But there are cases where you might be limited in setting up a dedicated syslog server with OneAgent because of environment architecture or resources. Yet observability into syslog data on Dynatrace would help you monitor and troubleshoot infrastructure.

This is where it is prudent to configure syslog producers to send data to a log shipper like Fluentd.

What is Fluentd?

Fluentd logo for log ingestion and log monitoring

Fluentd is an open source data collector that decouples data sources from observability tools and platforms by providing a unified logging layer. Fluentd is known for its flexibility and is also highly scalable, which makes it a good choice for high-volume environments.

How does Fluentd work with Dynatrace?

Setting up the flow from syslog over Fluentd to Dynatrace takes three steps. First, point the syslog daemon to the Fluentd port by adding the following line to the syslog daemon configuration file:

*.* @@<fluentd host IP>:5140

*.* instructs the daemon to forward all messages to the specified Fluentd instance listening on port 5140 and <fluentd host IP> needs to point to the IP address of Fluentd.

As a second step, enable Fluentd to accept incoming syslog messages with the in_syslog plugin. Set up the configuration on the same port as specified for source data, in this example 5140.

Lastly, use the open source Dynatrace Fluentd plugin, which uses generic log ingestion. Just find the API token for log ingest API on your SaaS environment or your own Active Gate setup.

Now you should see log messages coming into the Dynatrace log viewer.

Log ingestion strategy No. 2: Point an existing log shipper to the generic Dynatrace ingest

Another common scenario is an environment where you have already invested a do-it-yourself or other log shipper solution. After spending time and budget on the tooling and configuration, it may be unwise to undo this custom work, despite the automatic instrumentation of the Dynatrace OneAgent. Although you preserve your custom work this way, it is a siloed approach for logs, which means you’ll miss out on the integrated observability and automated alerting of Dynatrace.

If that existing solution supports sending log data to an external HTTP endpoint, you can address log silos by integrating with Dynatrace generic ingest with minimal hassle.

To illustrate the solution, let’s look at how to configure log ingestion with the log shipper Cribl.

What is Cribl?

Cribl Stream logo for log ingestion and log monitoring

Cribl is a data operations platform that enables users to collect, route, transform, analyze, and act on data in real time. It provides a unified platform for handling every aspect of data operations, from collecting data to routing and transforming it. Cribl also allows users to orchestrate custom pipelines for their data to gain insights and take action on that information. As a data output, or what it calls a Cribl Stream destination, you can configure an HTTP endpoint.

How does Cribl work with Dynatrace?

The main part of the setup involves creating the configuration for the specific log shipper at hand—in this case, Cribl Stream.

In Cribl’s configuration, open “Data/Destinations” and find “Webhook.” Create a new webhook destination with a name of your choosing (for example, your Dynatrace environment ID, and provide the URL for the webhook). For a Dynatrace SaaS environment, this is the following:

https://{your-environment-id}.live.dynatrace.com/api/v2/logs/ingest

This points the data stream to your Dynatrace environment’s generic ingest.

But in Cribl’s case, you should provide two more settings under “Configure/Advanced Settings/Extra HTTP Headers.” Add two new headers with the following names and values:

  1. To authorize the request, add the header “Authorization” and provide the value Api-Token dt0c01.{your-token-here} where {your-token-here} is an API token with ingest logs scope.
  2. Then add a header “Content-Type” and provide the value “application/json; charset=utf-8
log ingestion, log management screenshot
Example configuration in Cribl of posting logs to Dynatrace API.

After committing and deploying the Cribl changes, you can select the newly created Dynatrace destination as the default destination for your logs. And just like that, all log data already collected by the existing shipper is being sent to Dynatrace for monitoring, analysis, alerting, and all other tasks.

Log ingestion strategy No. 3: Ingest AWS Fargate logs with Fluent Bit

Ingesting and working with Kubernetes logs in Dynatrace helps to provide a comprehensive view of application performance from the infrastructure layer to the application layer. The common approach for Kubernetes logging is to deploy OneAgent in the environment, where it auto-discovers log messages written to the containerized application’s stdout/stderr streams.

But not all environments, configurations, or privileges are created equal. One recurrent challenge is collecting Kubernetes logs if you’re limited in installing OneAgent because of technical or architectural restrictions.

In the case of AWS serverless container compute engine Fargate, for example, where OneAgent log collection is not supported, we recommend using Fluent Bit log forwarder.

Let’s take this example of AWS Fargate. AWS includes a log router called FireLens for Amazon ECS and AWS Fargate services, which gives you built-in access to FluentD and Fluent Bit. We covered FluentD support previously. Now let’s take a look at how to set up Fluent Bit.

What is Fluent Bit?

Fluent Bit logo - for log ingestion and log management

Fluent Bit is an open source and multiplatform log processor and forwarder that allows you to collect data/logs from different sources, unify and send them to multiple destinations and is fully compatible with Docker and Kubernetes environments.

When choosing between Fluentd or Fluent Bit shippers, the Fluent Bit is the preferred solution when resource consumption is critical because it is a lightweight component.

While Fluent Bit has configurable HTTP output, in this example, we look at the AWS Fargate context, where FireLens makes it easy to set up Fluent Bit more quickly.

Ingest AWS Fargate logs with Fluent Bit

When creating a new task definition using the AWS Management Console, the FireLens integration section makes it easy to add a log router container. Just pick the built-in Fluent Bit image.

Next, edit the container in which your app-generating logs are running. In the “Storage and Logging” section, select “awsfirelens” as the log driver.

The settings for the log driver should point to the log ingest API of your SaaS tenant. Note that you normally need to provide two headers for Fluent Bit: content type and authorization token. As FireLens supports only one header, you can pass the token as part of the URL. Your configuration for AWS FireLens should have the following:

  • Name: http
  • TLS: on
  • Format: json
  • Header: Content-Type application/json; charset=utf-8
  • Host: {your-environment-id}.live.dynatrace.com
  • Port: 443
  • URI: /api/v2/logs/ingest?api-token={your-API-token-here}
  • tls.verify: Off
  • Allow_Duplicated_Headers: false
  • match: *
  • json_date_format: iso8601
  • json_date_key: timestamp

To avoid publishing the token in plaintext, use AWS Secrets Manager to manage the token.

As your application starts publishing logs, you can view them in Dynatrace.

Read more about streaming logs to Dynatrace with Fluent Bit from our documentation.

More methods for log ingestion

These are just some of the ways you can ingest logs into the Dynatrace platform without using OneAgent. You’ll soon have even more methods for log ingestion into Dynatrace, for example:

  • Automated OpenTelemetry logs acquisition and processing
  • Syslog endpoint in your environment as a component on a private ActiveGate
  • Dynatrace Fluent Bit output plugin for out-of-the-box integration

Want to share your experiences with log ingestion? Head to the Dynatrace Community Feedback channel to share your thoughts with other users.

State of Log Management 2026

Download the report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

The post Three smart log ingestion strategies in Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/three-smart-log-ingestion-strategies/feed/ 0
What is log monitoring? See why it matters in a hyperscale world https://www.dynatrace.com/news/blog/why-log-monitoring-and-log-analytics-matter-in-a-hyperscale-world/ https://www.dynatrace.com/news/blog/why-log-monitoring-and-log-analytics-matter-in-a-hyperscale-world/#respond Fri, 15 Jul 2022 14:35:46 +0000 https://www.dynatrace.com/news/?p=47194 log analytics log output for log management best practices

Log monitoring and management are now crucial as organizations adopt more cloud-native technologies, containers, and microservices-based architectures. In fact, the global log management market is expected to grow from $1.9 billion in 2020 to $4.1 billion by 2026, according to numerous market research reports. The increasing adoption of hyperscale cloud providers — such as Amazon […]

The post What is log monitoring? See why it matters in a hyperscale world appeared first on Dynatrace news.

]]>
log analytics log output for log management best practices

Log monitoring and management are now crucial as organizations adopt more cloud-native technologies, containers, and microservices-based architectures.

In fact, the global log management market is expected to grow from $1.9 billion in 2020 to $4.1 billion by 2026, according to numerous market research reports. The increasing adoption of hyperscale cloud providers — such as Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP) — as well as containerized microservices are driving this growth. But the flexibility of these environments also makes them more complex. As such, this complexity brings an exponential increase in the volume, velocity, and variety of logs.

To identify what’s happening in these increasingly complex environments — and, more importantly, to harness their operational and business value — teams need a smarter way to monitor and analyze logs.

Therefore, it’s critical to explore logs and log monitoring to understand why they’re so critical for healthy cloud architectures.

What are logs?

A log is a timestamped record of an event generated by an operating system, application, server, or network device. Logs can include data about user inputs, system processes, and hardware states.

Log files contain much of the data that makes a system observable — for example, records of all events that occur throughout the operating system, network devices, or pieces of software. Logs even record communication between users and application systems. Logging is the practice of generating and storing logs for later analysis.

Watch the Perform 2024 session “Turning logs to learnings faster with Dynatrace”

What is log monitoring?

Log monitoring is a process by which developers and administrators continuously observe logs as they’re recorded. With log monitoring software, teams can collect information and trigger alerts if something affects system performance and health.

DevOps teams (or development and operations teams) often use a log monitoring solution to ingest application, service, and system logs so they can detect issues throughout the software delivery lifecycle (SDLC). Whether a situation arises during development, testing, deployment, or in production, a log monitoring solution detects conditions in real time to help teams troubleshoot issues before they slow down development or affect customers.

But to determine root causes, teams must be able to analyze logs.

How log monitoring facilitates log analytics

Log monitoring and log analytics are related — but different — concepts that work in conjunction. Together, they ensure the health and optimal operation of applications and core services.

Whereas log monitoring is the process of tracking logs, log analytics evaluates logs in context to understand their significance. This includes troubleshooting issues with software, services, applications, and any infrastructure with which they interact. Such infrastructure includes multicloud platforms, container environments, and data repositories.

Log monitoring and analytics work together to ensure applications are performing optimally and to determine how systems can improve.

Log analytics also helps identify ways to make infrastructure environments more predictable, efficient, and resilient. Together, they provide continuous value to businesses by providing a window into issues and how to run systems optimally.

Reap the benefits of log monitoring

Log monitoring helps teams to maintain situational awareness in cloud-native environments. This practice provides myriad benefits, including the following:

  • Faster incident response and resolution. Log monitoring helps teams respond to incidents faster and discover issues before they affect end users.
  • More IT automation. With clear insight into crucial system metrics, teams can automate more processes and responses with greater precision.
  • Optimized system performance. Log monitoring can reveal potential bottlenecks and inefficient configurations so teams can fine-tune system performance.
  • Increased collaboration. A single log monitoring solution benefits cloud architects and operators so they can create more resilient multicloud environments.

Log monitoring use cases

Anything connected to a network that generates a log of activity is a candidate for log monitoring. As solutions evolve to use artificial intelligence, the variety of use cases has extended beyond break-fix scenarios to address a wide range of technology and business concerns.

These include the following:

  • Infrastructure monitoring automatically tracks modern cloud infrastructure, including the following:
    • Hosts and virtual machines;
    • Platform as a service, such as AWS, Azure, and GCP;
    • Container platforms, such as Kubernetes, OpenShift, and Cloud Foundry;
    • Network devices, process detection, resource utilization, and network usage and performance;
    • Third-party data and event integration; and
    • Open source software.
  • Application performance monitoring and monitoring of microservices discovers dynamic microservices workloads running inside containers and detects and pinpoints issues before they affect real users.
  • Digital experience monitoring, including real-user monitoring, synthetic monitoring, and mobile app monitoring, ensures that every application is available, responsive, fast, and efficient across every channel.
  • Application Security automatically detects vulnerabilities across cloud and Kubernetes environments.
  • Business Observability provide real-time visibility into business key performance indicators to improve IT and business collaboration.
  • Cloud automation and orchestration for DevOps and site reliability engineering teams build better-quality software faster by bringing observability, automation, and intelligence to DevOps pipelines.

Overcoming log monitoring challenges

In modern environments, turning the crush of incoming logs and data into meaningful use cases can quickly become overwhelming. While log monitoring is essential to IT operations, practicing it effectively in cloud-native environments has some challenges.

One major challenge for organizations is a lack of end-to-end observability, which enables users to measure an individual system’s current state based on the data it generates. As environments use thousands of interdependent microservices across multiple clouds, observability becomes increasingly difficult.

Organizations also struggle with inadequate context. Logs are often collected in data silos, with no relationships between them, and aggregated in meaningless ways. Without meaningful connections, you’re often looking for a few traces among billions to know whether two alerts are related or how they affect users.

Too often, logging tools leave you clicking through data and poring through logs to deduce root causes based on simple correlations. Lack of causation makes it difficult to quantify the effect on users. It also makes it hard to determine which optimization efforts are delivering performance improvements.

Additionally, log monitoring’s high cost and blind spots often plague enterprises. To avoid the high data-ingest costs of traditional log monitoring solutions, many organizations exclude large portions of their logs. This results in minimal sampling. Although cold storage and rehydration can mitigate high costs, it is inefficient and creates blind spots.

With the complexity of modern multicloud environments, traditional aggregation and correlation approaches are inadequate. Teams need to quickly discover faulty code, anomalies, and vulnerabilities. Too often, organizations implement multiple disparate tools to address different problems at different phases, which only compounds the complexity.

How Dynatrace unlocks the value of log monitoring

Logs are an essential part of the three fundamental pillars of observability: metrics, logs, and traces. End-to-end observability is crucial for gaining situational awareness into cloud-native architectures. But logs alone aren’t enough. To attain true observability, organizations need the ability to determine the context of an issue, both upstream and downstream. Equally important is leveraging user experience data to understand what’s affected, the root cause, and how it affects the business.

To overcome these challenges and to get the best intelligence for log monitoring, organizations need to work with a solution that takes analytics to the next level.

Using a real-time map of the software topology and deterministic AI, Dynatrace helps DevOps teams automatically monitor and analyze all logs in context of their upstream and downstream dependencies. This broad yet granular visibility enables analysts to understand the business context of an issue. Using AI, teams can automatically pinpoint an issue’s precise root cause down to a particular line of code.

Unlike tools that rely solely on correlation and aggregation, the Dynatrace AIOps platform enables teams to speed up and automate processes. As a single source of truth, Dynatrace combines precise answers and automation to free teams to optimize apps and processes. With a sole source of analytics, teams can improve system reliability and drive better business outcomes.

State of Log Management 2026

Download the report to explore benchmark data on how AI workloads are exploding log volume and costs, and why unified observability is now essential for reliable, trustworthy AI.

The post What is log monitoring? See why it matters in a hyperscale world appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/why-log-monitoring-and-log-analytics-matter-in-a-hyperscale-world/feed/ 0