Notebooks | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Thu, 09 Jul 2026 14:32:50 +0000 en hourly 1 Dynatrace Release Radar 04.26 https://www.dynatrace.com/news/blog/dynatrace-release-radar-04-26-whats-new-and-why-it-matters/ https://www.dynatrace.com/news/blog/dynatrace-release-radar-04-26-whats-new-and-why-it-matters/#respond Wed, 27 May 2026 17:07:44 +0000 https://www.dynatrace.com/news/?p=74179 Release Radar

This series covers recent Dynatrace releases and updates, focusing on what’s new, what’s changed, and how these recent enhancements can benefit you and your organization. Each post covers newly available capabilities and points you toward where to explore them.

The post Dynatrace Release Radar 04.26 appeared first on Dynatrace news.

]]>
Release Radar

The April 2026 Dynatrace SaaS releases bring six updates aimed at a familiar problem: too much manual work between a signal and an answer. The updates focus on native cloud visibility, deeper Kubernetes insights, consistent severity handling, faster investigations, and smoother analytics work.

Explore all updates hands-on in the Release Radar launchpad.

A native Azure experience in Clouds

What does Dynatrace add for Azure practitioners?

Dynatrace now extends the enhanced Clouds experience to Microsoft Azure, putting Azure subscriptions on the same footing as AWS. Metrics, logs, metadata, and topology now sit in one managed view, so Azure teams can move from inventory to investigation without stitching the picture together by hand.

What you get out of the box

  • Opinionated insights and ready-made dashboards built from enriched Azure telemetry, so investigations start with answers instead of a blank canvas.
  • Pre-configured health alerts for Azure, created and managed directly in Clouds, with drill-down, search, and filtering directly in the Clouds app.
  • Broad metric coverage for any Azure Monitor native platform metric across Azure services.
  • A rich Azure topology inventory that periodically scans Azure environments and enriches resources with native metadata such as tags and subscription IDs, all queryable with Dynatrace Query Language (DQL).
  • Simple onboarding and lifecycle management that turns Azure subscriptions into native Dynatrace connections and manages them centrally.

Dynatrace also adds drilldowns from cloud entities into the relevant Dynatrace experiences, so teams can keep moving instead of bouncing between cloud and platform views.

The new Clouds experience for Azure lets you optimize cloud operations at scale.
The new Clouds experience for Azure lets you optimize cloud operations at scale.

Kubernetes visibility for autoscaling and custom resources

What’s new in Kubernetes observability?

Dynatrace extends Kubernetes visibility to two additional object types that SREs and platform teams rely on daily: Horizontal Pod Autoscalers (HPA) and Custom Resources (CRs).

Horizontal Pod Autoscaler as a first-class object

HPA is now a first-class object in enhanced Kubernetes visibility. You can see when scaling kicked in, what triggered it, and how desired and actual replica counts lined up next to the workloads involved.

Custom Resource insights

You can monitor up to five Custom Resources per cluster, surfaced the same way as built-in Kubernetes objects. This brings CRD-heavy ecosystems such as Argo, Istio, Cert-Manager, Kyverno, and operator-managed databases into the same investigation scope as the rest of your cluster.

For clusters connected through cloud integrations, the Kubernetes cluster details page now exposes the underlying cloud configuration (EKS, AKS, or GKE) in YAML or JSON, making cloud-side and cluster-side state accessible in one place.

HorizontalPodAutoscaler visibility in the Kubernetes app experience.
HorizontalPodAutoscaler visibility in the Kubernetes app experience.

A unified severity model for alerts and problems

What is event.severity in Dynatrace?

Dynatrace introduces a standardized event.severity field for alerts and problems, aligned with the ITIL Incident Management framework. Severity is stored in Grail as an integer from 1 (Critical) to 5 (Informational) and is shown as a human-readable label across the platform.

Severity levels at a glance

Value Label Description
1 Critical Major business disruption; service outage
2 Major Significant impact; workaround may exist
3 Minor Limited or non-critical impact
4 Warning Low impact; no business disruption
5 Informational No business impact

Severity automatically propagates from correlated alerts to the parent problem, with the highest severity always taking precedence. This gives teams one severity model to filter on, route with, and escalate from.

You can now:

  • Filter the problem feed by severity
  • Display a severity column with visual icons in problem lists
  • Use severity as a condition in Workflows for alert routing and notifications
Event severity in the Problems app experience.
Event severity in the Problems app experience.

Faster Investigations with Smartscape navigation

What changed in Smartscape?

Smartscape now offers all six ready-made views, such as vertical topology, horizontal topology, and visual resolution path, just a click away in a persistent side panel. You no longer need to return to the landing page in the middle of an investigation.

The new Recent views section shows your latest investigations, making it easy to reopen them, compare them, and keep working as you test a root-cause hypothesis.

The result is less backtracking in the middle of an incident.

The new sidebar navigation in Smartscape
The new sidebar navigation in Smartscape

Dashboards and notebooks: productivity improvements

What’s new for dashboard authors and analysts?

The latest release adds several practical upgrades for team members who build dashboards and work in notebooks every day.

  • Treemap visualization for identifying dominant categories in hierarchical data, such as requests per service by Kubernetes namespace.
    Treemap visualization example
    Treemap visualization example
  • Dashboard variables for dynamic coloring and thresholds, so visual conditions stay in sync with environment or team selectors.
    Use dashboard variables for dynamic coloring and threshold conditions
    Use dashboard variables for dynamic coloring and threshold conditions
  • Centralized tile indicator controls, allowing you to show or hide warnings, descriptions, and custom timeframes at the dashboard level.
    Select or clear tile indicators on a dashboard
    Select or clear tile indicators on a dashboard
  • Direct image upload in Markdown using a built-in image library shared across Dashboards, Notebooks, and the Launcher.
    Upload image directly in Markdown
    Upload image directly in Markdown
  • Row marker coloring for tables, making it easier to visually group related rows without sacrificing readability.
    Highlight table rows with color markers in Dashboards and Notebooks
    Highlight table rows with color markers in Dashboards and Notebooks

User experience improvements

Why does the platform feel faster?

This release smooths out the path from the first symptom to root cause analysis. Tracing and services workflows now handle high-span traces more reliably, show timing more clearly, and surface useful sample traces earlier.

Table-first workflows also benefit from richer entity-detail tables, better filtering, and clearer structure, helping teams answer more questions without switching views. Navigation patterns, overlays, and error messaging are now more consistent across the platform, reducing mental overhead when time is tight.

Explorer new table experience with entity details, alerts, and schema links
New Explorer table experience with entity details, alerts, and schema links

Why these updates matter

Taken together, these updates eliminate inefficiencies in the work that teams do every day. Cloud operations teams, Kubernetes SREs, on-call engineers, and analytics authors get richer context, faster paths to answers, and simpler ways to share what they find. This is where Dynatrace earns its keep under pressure.

Explore the updates live in the Release Radar launchpad.

The post Dynatrace Release Radar 04.26 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-release-radar-04-26-whats-new-and-why-it-matters/feed/ 0
How lookup tables turn observability data into business insight https://www.dynatrace.com/news/blog/how-lookup-tables-turn-observability-data-into-business-insight/ https://www.dynatrace.com/news/blog/how-lookup-tables-turn-observability-data-into-business-insight/#respond Wed, 25 Mar 2026 23:02:27 +0000 https://www.dynatrace.com/news/?p=73533 Blog thumbnail

Lookup tables are a simple feature with serious leverage. They give your team a lightweight, maintainable way to layer business context onto your observability data at query time so your raw telemetry stays clean and enrichment remains flexible.

The post How lookup tables turn observability data into business insight appeared first on Dynatrace news.

]]>
Blog thumbnail

Observability data can tell you a lot, but often not in a language that your business understands. Traces fire, business events flow, and dashboards light up, yet P1 incidents still break out. The first question from the war room is typically, “Who’s impacted?” Too often, the honest answer is, “Let me check the database and get back to you.” This gap between raw telemetry and human-readable business context is exactly what lookup tables in Dynatrace are designed to close.

In our introductory blog post on Dynatrace lookup tables, we explained the fundamentals: what lookup tables are, how to ingest them, and initial use cases spanning security to business context enrichment.

In this blog, we’ll go deeper, with real-world examples of lookup tables, practical patterns, and how teams are already putting them to work.

The problem: Dashboards display cryptic identifiers, not insights

Daniel Adams, Observability Engineer at FreedomPay, put it plainly in a recent Dynatrace Tips & Tricks session:

“A lot of our apps are not coded to show customer names or plain text human-readable stuff. It’s mostly IDs being passed back and forth.”

The result is that dashboards are filled with cryptic identifiers that operations teams and business stakeholders can’t mentally decode. And splitting data by customer, country, or store requires trips back to the SQL database for further contextual configuration.

Many organizations that run distributed systems at scale face the same challenge: telemetry is rich, but the business context is missing.

The solution: Use lookup tables to make your reference data available in Grail

Lookup tables let you upload reference data (CSV, JSON, or XML files) directly into Dynatrace Grail®, where it lives alongside your logs, traces, metrics, and business events. Once uploaded, two DQL commands unlock the enrichment from your lookup tables:

  • Load fetches the contents of a lookup table, working just like querying any other data type in Grail.
  • lookup joins records from your observability data with matching rows in the lookup table, using a defined key field.

In FreedomPay’s case, the lookup table holds customer names, customer codes, store IDs, and countries. The join key is the store ID, the same cryptic identifier that was previously displayed on their dashboards. One DQL lookup command later, those IDs become readable names, and the data can be split by country, grouped by customer, and surfaced in a dashboard that anyone on the team can use.

Using lookup tables in DQL to enrich business events with customer and response context at query time.
Figure 1: Using lookup tables in DQL to enrich business events with customer and response context at query time.

Lookup tables reduce user cognitive load

The impact of lookup tables is not just cosmetic. Daniel described the before-and-after in P1 situations:

“Previously, everyone was asking: who’s impacted, which customer? Now it’s just very apparent and real-time. People can share screenshots directly from Dynatrace.”

The shift from “let me go check” to “it’s right here” compresses triage time and reduces the cognitive load on everyone in the room.

Beyond incident response, lookup tables also power dashboard variables. FreedomPay uses dashboard variables to populate dropdown filters (for example, clients and countries) directly from the lookup table itself, keeping the UI dynamic without manual maintenance. The load command behaves like any other fetch command but operates on reference data rather than telemetry, making it a natural fit for driving interactive dashboards.

And the enrichment isn’t limited to business events. The load and lookup commands work across all data types in Grail.

A dashboard powered by lookup tables, grouping throughput, errors, and success rates by customer and country.
Figure 2: A dashboard powered by lookup tables, grouping throughput, errors, and success rates by customer and country.

In practice: cost allocation with lookup tables

Cost allocation is a good example of how lookup tables bridge the gap between operational data and business reality. Many organizations already maintain mappings that associate users, teams, or services with products or cost centers, but that context rarely lives inside observability data. Lookup tables allow you to introduce such context.

Consider DPS (Dynatrace Platform Subscription) consumption driven by user activity such as running queries, triggering automation, or executing functions. On its own, such usage is hard to attribute to a specific part of the organization. By associating users with products or cost centers in a lookup table, you can attribute usage to organizational ownership rather than just technical activity, align cost views with existing financial or product structures, and keep that attribution logic centralized and reusable rather than embedding it in individual analyses.

What makes this pattern scale is that business mappings change over time, while usage data continues to flow. Lookup tables allow you to evolve the mapping without rewriting your analyses. The same mechanism used for cost allocation works equally well for mapping users to teams, services to portfolios, or environments to internal programs. Cost allocation simply highlights how powerful this becomes when operational data is interpreted through a business lens.

Importantly, this is just one way to approach cost allocation, not the only way. Some teams rely on host-based attribution, others on pipeline metadata or external financial systems. Lookup tables complement these models by making it easy to incorporate existing reference data wherever it adds clarity.

Keeping lookup data fresh

For lookup tables to deliver lasting value, their data must remain current. FreedomPay currently uses Postman to push updated CSVs via the upload API, but they’re building automation to pull this data from their SQL backend daily and sync changes using the overwrite: true parameter. The lookup table path stays the same; the data underneath gets refreshed. Downstream dashboards and queries are updated automatically.

More broadly, there are several approaches to keeping lookup data up to date:

  • Dynatrace Workflows automate extraction and upload from SQL databases or APIs on a schedule.
  • Periodic refresh updates your lookup data programmatically as part of CI/CD or data sync processes.
  • Manual refresh uses Postman, cURL, or other API-based uploads with the overwrite flag when data changes infrequently.

Dynatrace lookup tables are now generally available

With the release of Dynatrace SaaS version 1.334, lookup tables are generally available for all customers running Dynatrace SaaS with an active Dynatrace Platform Subscription (DPS). GA brings production-readiness, improved query performance for the lookup command, and full integration with Grail’s scalability and access controls.

To get started

  1. Identify a data enrichment use case: cost allocation, customer context, error code mapping, or security allow lists.
  2. Prepare your reference data as a CSV, JSON, or JSONL (JSON Lines) file.
  3. Upload the file to Grail using the Dynatrace API, Workflows, or a custom app.
  4. Use the load and lookup commands in your DQL queries, notebooks, and dashboards.

For detailed instructions, visit our lookup data documentation. And to see how FreedomPay built its implementation from end to end, watch the full demo on YouTube.

The post How lookup tables turn observability data into business insight appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-lookup-tables-turn-observability-data-into-business-insight/feed/ 0
Explore without friction: Deeper insights with Dynatrace expanded analytics app portfolio https://www.dynatrace.com/news/blog/deeper-insights-with-dynatrace-expanded-analytics-app-portfolio/ https://www.dynatrace.com/news/blog/deeper-insights-with-dynatrace-expanded-analytics-app-portfolio/#respond Wed, 28 Jan 2026 16:55:12 +0000 https://www.dynatrace.com/news/?p=72769 Dynatrace analytics app portfolio

Modern IT systems operate under constant pressure to deliver efficiency, resilience, security, agility, and business alignment simultaneously. That balance isn’t achieved overnight; it’s built through continuous improvement driven by learning and insight. This is where exploratory analytics becomes essential: analytics help you probe deeper, ask sharper questions, uncover patterns, anticipate issues, and optimize resources. With the Grail® unified data lakehouse, all live production data is unified in context and ready to explore. Discover the enhanced Dynatrace analytics app portfolio and see how embedding exploration into your processes requires little effort yet delivers transformative results.

The post Explore without friction: Deeper insights with Dynatrace expanded analytics app portfolio appeared first on Dynatrace news.

]]>
Dynatrace analytics app portfolio

Spark continuous improvement by making data exploration a habit

In practice, exploration often stalls. Logs live in one tool, metrics in another, traces somewhere else, and the broader business context lives with different teams. Answering a single question, such as “Did this pattern exist before the last release?” can mean switching contexts, running separate queries, and manually correlating results. The friction builds, and a deeper investigation is deferred until the next incident forces it.

There is more than one kind of exploration. Sometimes exploration is structured: embedded within workflows as part of postmortems, release validations, or SLO reviews, tracing signals across logs, metrics, traces, events, and business data to link impact to outcomes and prevent repeat failures. Other times, exploration is curiosity-driven: a quieter moment where you notice a memory pattern tied to a batch job, a timeout spike with specific clients, or a cloud spend anomaly you’d never have caught in a scheduled report.

Both of these exploration modes matter. Together, they foster learning and a culture of continuous improvement, both essential to modern enterprises.

Dynatrace Grail, a unified data lakehouse, makes exploration easy and rewarding: a single place for all your live production data across logs, metrics, traces, events, business, and security data, so you can follow the evidence wherever it leads without rigid queries or manual joins. With Dynatrace Query Language (DQL), every field and relationship is at your fingertips, enabling broad searches, contextual pivots, precise slicing, and rapid iteration to test and discard weak hypotheses.

Dynatrace Intelligence® makes exploration accessible to everyone. Use Assist to query and explore your data in natural language, or get support in interpreting and understanding your findings. Leverage AI-powered analysis to detect patterns and anomalies at scale, or forecast trends to predict future behavior.

To support different exploration needs, Dynatrace offers a portfolio of use-case optimized apps:

  • Dashboards for persistent, shared visibility and ongoing monitoring.
  • Smartscape for visual analytics of real-time topology and dependency context.
  • Notebooks for collaborative, ad hoc exploration and rapid hypothesis testing.
  • Investigations for sequential, forensic depth in complex scenarios.

Move seamlessly between apps without losing context. Start in Dashboards, drill into a data point, and continue to explore your data in Notebooks. From there, you might run a deep, focused analysis in Investigations, then pivot to Smartscape for a dependency graph. Finally, bring your findings back into a dashboard for continuous monitoring, making your entire exploration journey seamless.

Let’s look at the apps in more detail, starting with Dashboards, often a natural entry point for your exploration journey.

The exploratory apps portfolio, each app optimized for different use cases.
Figure 1. The exploratory apps portfolio, each app optimized for different use cases.

Dashboards: from real-time visualizations to taking action

Dashboards provide a powerful way to transform complex data visualizations into actionable insights, serving as the cornerstone of the Dynatrace exploratory analytics portfolio where exploration meets operational excellence. By offering real-time visibility into key metrics, dashboards help teams monitor performance, identify trends, and make informed decisions. With ready-made dashboards for common use cases, such as Kubernetes, infrastructure, and digital experience monitoring, teams gain immediate access to critical insights, allowing for faster and more proactive responses to their daily challenges.

Dashboards are designed to foster operational clarity with intuitive, interactive visualizations that allow you to drill down into metrics, apply filters, and segment data to uncover meaningful patterns. While Notebooks and Investigations are ideal for deep dives and custom analyses, Dashboards deliver concise, shareable, real-time views that keep teams aligned and informed.

Deeply integrated with Dynatrace’s AI-powered analytics, dashboards enhance visualizations with contextual explanations, anomaly detection, and forecasting. These capabilities allow teams not only to monitor what’s happening but also to understand why it’s happening and predict what might happen next. By making insights accessible to both technical and non-technical stakeholders, dashboards foster collaboration, break down silos, and empower teams to stay aligned and proactive.

Use Dashboards to:

  • Monitor KPIs and SLOs in real time
  • Identify anomalies and emerging trends early
  • Align teams with shared, role-based views and a single source of truth
  • Trigger deeper analysis via drill-downs into charts and entities
  • Track progress against goals and initiatives over time
  • Surface business and technical context side by side for informed decisions
Get instant insights into infrastructure health with ready-made dashboards.
Figure 2. Get instant insights into infrastructure health with ready-made dashboards.

Smartscape: visualize the topology and dependencies of your complete digital systems

Smartscape® is the latest addition to the Dynatrace exploratory analytics app portfolio, and it’s a game-changer for exploring highly dynamic IT systems. Purpose-built for real-time visual analytics, Smartscape gives you a dynamic, interactive view of your entire IT ecosystem—spanning all layers, including services, cloud, Kubernetes, and on-premises infrastructure. Unlike static diagrams or manual dependency maps, Smartscape updates continuously, so you can understand changes as they happen.

Smartscape’s visual analytics capabilities go far beyond simple mapping. It provides multidimensional, domain-specific views that allow teams to see how services, processes, and infrastructure interact in real time. This real-time visualization helps uncover hidden dependencies, assess the blast radius of outages, spot drift or misconfigurations, and validate architecture after deployments. Apply your business context by using Segments, and pivot from other apps like Problems, Kubernetes, or Clouds into Smartscape without losing context.

Visualize and explore dependencies across your IT systems at scale with the new Smartscape app.
Figure 3. Visualize and explore dependencies across your IT systems at scale with the new Smartscape app.

Use Smartscape to:

  • Visualize real-time dependencies and communication paths across services and infrastructure
  • Assess blast radius and map out highly connected and interdependent entities during incidents
  • Validate architecture and changes after deployment
  • Identify and understand hotspots, bottlenecks, and hidden dependencies
  • Navigate readymade domain views for clouds, Kubernetes, services, and infrastructure with zero setup
  • Align engineering, ops, and business teams with a shared, always-current understanding

Notebooks: collaborate, explore, and solve problems in real-time

Notebooks bridge the gap between the two modes of exploration and play an important role in both standardized processes and curiosity-driven exploration.

As a workspace for free exploration, Notebooks give you a playground to experiment with data, quickly visualize insights with a large set of chart types from a curated library, and iterate quickly. You can slice massive datasets in real time, pivot on context, and uncover patterns without constraints.

At the same time, Notebooks shine in collaborative workflows. Teams can work together to document and share findings during incident resolution or postmortems, create troubleshooting guides, and also generate automated reports from queries, all within the same space. Notebooks documenting incidents are automatically surfaced in the Problems app via vector search when similar issues occur, and snapshots of investigations can be preserved as long as needed outside of retention period settings, ensuring insights remain accessible.

Whether you’re just performing free-form discovery or creating documents within processes, Notebooks make it effortless to turn exploration into reusable assets.

Use Notebooks to:

  • Collaborate on incident investigations
  • Document postmortems for future reference
  • Report insights ad hoc or on a schedule
  • Analyze your data using generative AI
  • Prototype and validate DQL for alerts, workflows, and automation
  • Tell data stories with rich visuals and narrative
  • Build a reusable knowledge base to reduce MTTR
  • Extract data on demand
  • Transform and shape data on read
Notebooks are the perfect place for ad-hoc data exploration, collaboration, and data storytelling.
Figure 4. Notebooks are the perfect place for ad-hoc data exploration, collaboration, and data storytelling.

Investigations: dive deeper with sequential analysis and forensics

When exploration moves from curiosity to critical analysis, Investigations is your go-to tool. Built for structured, forensic deep dives, it’s the perfect complement to ad hoc exploration in Notebooks.

Investigations works with DQL across all data in Grail, including logs, events, metrics, traces, business data, and security signals. Teams can pivot from initial findings to comprehensive analysis without friction, comparing scenarios and following evidence trails wherever they lead. The query tree tracks your analytical path, letting you branch into parallel hypotheses and return to previous queries and results at any point.

Compare different scenarios and follow evidence trails with the query tree in Investigations.
Figure 5. Compare different scenarios and follow evidence trails with the query tree in Investigations.

Imagine this flow: a security team is alerted to unusual login attempts or wants to follow up on an anomaly spotted in Dashboards or Notebooks. With a single click, they transition to Investigations to trace lateral movement, simulate attack scenarios, and preserve evidence for future reference. Investigations support sequential workflows, allowing you to pivot queries based on metadata, visualize intricate patterns, and even enrich analysis with lookup tables and external data joins.

Pivot queries based on metadata, visualize patterns across multiple dimensions, and enrich your analysis with lookup tables and external data joins. Because all Grail data is accessible, you can seamlessly connect the dots, linking a suspicious error to its pod’s resource consumption, or tracing a payment failure back to the infrastructure event that caused it. Custom pivots let you select multiple findings and branch into separate queries automatically.

When you’ve found what you’re looking for, save the investigation, or just the relevant branches, as a Notebook to share with your team.

This isn’t just incident response, but everyday analytics for complex environments. Investigations help teams validate hypotheses, document findings, and strengthen resilience across IT systems.

Use Investigations to:

  • Conduct structured, multi-step analyses across services and systems
  • Correlate signals from applications, infrastructure, and user activity
  • Follow evidence trails to confirm or refute hypotheses
  • Explore parallel scenarios with branching query paths
  • Diagnose complex integration and dependency issues across environments
  • Collect and preserve evidence for auditability and knowledge reuse
  • Enrich analyses with context (for example, lookups, reputation data, metadata)

Ready to make more of your data with Dynatrace?

Chances are your IT organization is currently focused on increasing efficiency, strengthening resilience and security, increasing agility, or aligning more closely with business priorities. Within your Dynatrace data, there are likely far more insights waiting to be uncovered, helping you accelerate these goals. Start by formulating the right questions: Where are inefficiencies hiding? Which patterns precede incidents and impact uptime? How exposed are critical services?

Give yourself time and room to explore: Visualize dependencies in Smartscape. Use Assist to turn your questions into DQL queries in Notebooks; experiment and try things out. When findings require structured follow-up, Investigations will help you get a complete understanding. And finally, for anything interesting you see on a chart, dashboards let you drill down into the details.

Start exploring now and make exploration a cornerstone of continuous improvement.

For inspiration and an overview of available Exploratory Analytics resources, take a look at our Perform 2026 – Exploratory Analytics Launchpad.

The post Explore without friction: Deeper insights with Dynatrace expanded analytics app portfolio appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/deeper-insights-with-dynatrace-expanded-analytics-app-portfolio/feed/ 0
Understand and validate DQL queries using Dynatrace Davis CoPilot https://www.dynatrace.com/news/blog/understand-and-validate-dql-queries-using-dynatrace-davis-copilot/ https://www.dynatrace.com/news/blog/understand-and-validate-dql-queries-using-dynatrace-davis-copilot/#respond Mon, 03 Nov 2025 16:54:00 +0000 https://www.dynatrace.com/news/?p=71670 Dynatrace Davis CoPilot

Dynatrace Query Language (DQL) delivers unlimited contextual analytics, but if you’re not writing queries every day, the learning curve can feel steep. Davis CoPilot® makes things easier by generating even complex DQL queries from natural language alone. With its latest enhancement, Davis CoPilot can also summarize and explain existing queries in context, showing what a […]

The post Understand and validate DQL queries using Dynatrace Davis CoPilot appeared first on Dynatrace news.

]]>
Dynatrace Davis CoPilot

Update: We’ve launched Dynatrace Assist, our next-generation AI chat that goes far beyond answering questions.
Dynatrace Assist is the evolution of Davis CoPilot®.

Dynatrace Query Language (DQL) delivers unlimited contextual analytics, but if you’re not writing queries every day, the learning curve can feel steep. Davis CoPilot® makes things easier by generating even complex DQL queries from natural language alone. With its latest enhancement, Davis CoPilot can also summarize and explain existing queries in context, showing what a query does, why it’s structured the way it is, and how the results relate to the underlying data. This helps teams validate intent, spot gaps, and confidently build on each other’s work without requiring deep query expertise.

Suppose you’ve opened a dashboard or notebook and found a complex query you didn’t write. You know the struggle: queries can be overwhelming to look at, key details may be nested or referenced elsewhere, and you might not be familiar with the specific data syntax or the user’s original intent. Even revisiting your own work after a few weeks can mean trying to remember what the dashboard was designed to show and how the query fits together. Reverse-engineering shouldn’t be a prerequisite for collaboration.

With the Summarize and explain queries Davis CoPilot skill, you get a clear, contextual explanation of any query, helping you quickly understand what it does.

Get an explanation of any DQL query

Figure 1. Get an explanation of any DQL query. (video)

From creating queries to explaining them

Last year, we introduced natural language querying, allowing anyone to explore their data without learning DQL syntax. Now, Davis CoPilot can also interpret existing queries using the Dynatrace data model, explaining what the query does, how it filters and calculates results, and which data sources it uses. This reduces the effort required to work with complex syntax, facilitating the onboarding of new users while enabling experts to validate intent and iterate more efficiently.

The Explain and summarize queries skill is available in Notebooks and Dashboards. Review queries from teammates, tailor ready-made dashboards to your specific needs, and accelerate knowledge sharing.

Try it out on the Dynatrace Playground

Dynatrace offers a wide range of ready-made dashboards to help you get started instantly; however, sometimes you need to tailor dashboards to your unique use cases. With the Explain and summarize queries skill, you can instantly understand how the underlying queries were built by Dynatrace experts, giving you guidance and inspiration for your own customizations. See the examples below:

Log ingest overview dashboard: The table below highlights your noisiest log sources, helping you quickly pinpoint where excessive volume might be driving up ingest costs or masking real issues. With Davis CoPilot, you can see exactly how the underlying query is constructed, making it easy to extend the logic or use it as a template for your own ranking and cost-optimization dashboards.

Davis CoPilot explains a query of the top 20 log producers in your system Davis CoPilot explains a query of the top 20 log producers in your system query explanation

Figure 2. Davis CoPilot explains a query of the top 20 log producers in your system.

Databases overview dashboard: The following chart identifies slow or inefficient SQL statements that degrade application responsiveness, allowing you to focus your tuning efforts where they matter most. Use Davis CoPilot to break down the logic behind the analysis so you can adapt its scope, filter for critical services, or enrich results with additional business context.

Davis CoPilot explains a query that identifies the 20 most resource-intensive statements from Oracle databases Davis CoPilot explains a query that identifies the 20 most resource-intensive statements from Oracle databases queries explanaion

Figure 3. Davis CoPilot explains a query that identifies the 20 most resource-intensive statements from Oracle databases.

Kubernetes cluster dashboard: Optimizing Kubernetes requires clear visibility into how workloads consume cluster resources. This query returns CPU usage broken down by namespace, helping you detect saturation early and maintain efficiency. With Davis CoPilot, you can see exactly how the query works and then adapt it to create your own capacity-planning tool.

Davis CoPilot explains a query returning CPU quotas per Kubernetes namespace Davis CoPilot explains a query returning CPU quotas per Kubernetes namespace query explanation

Figure 4. Davis CoPilot explains a query returning CPU quotas per Kubernetes namespace.

Ready to try it out yourself?

This feature is already available in all environments; you just need to ensure that Davis CoPilot is turned on and that you have the necessary permissions for this skill. Learn more about Davis CoPilot summarization and explanation of DQL queries in Dynatrace Documentation.

The post Understand and validate DQL queries using Dynatrace Davis CoPilot appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/understand-and-validate-dql-queries-using-dynatrace-davis-copilot/feed/ 0
Enrich your Dynatrace data with the newly introduced lookup tables https://www.dynatrace.com/news/blog/enrich-your-dynatrace-data-with-the-newly-introduced-lookup-tables/ https://www.dynatrace.com/news/blog/enrich-your-dynatrace-data-with-the-newly-introduced-lookup-tables/#respond Thu, 07 Aug 2025 15:15:05 +0000 https://www.dynatrace.com/news/?p=70214 Observability data

With the introduction of a new file storage system in Dynatrace Grail®, you can now easily enrich your observability and security data by storing and querying lookup data, with no additional data ingest or manipulation required.

The post Enrich your Dynatrace data with the newly introduced lookup tables appeared first on Dynatrace news.

]]>
Observability data

Enriching observability data with additional context means improved data quality, which leads to better decision-making and faster troubleshooting. Instead of switching to an external data source and searching for a specific identifier across multiple documents, you now gain immediate insights at query time, effectively streamlining your work.

In this blog post, you’ll learn how to ingest lookup data and use it to effortlessly enrich your observability data. Practical use cases outline scenarios in which lookup data improves user workflows and makes root cause analysis and troubleshooting more efficient.

How to ingest lookup data

Lookup data files can be uploaded in formats such as CSV, JSON, or XML. You can upload data files using Workflows, via API, or by creating your own custom app. Once ingested, you can query lookup data just like any other Grail data, using Dynatrace Query Language (DQL) commands like lookup and join, or built-in Dynatrace® Apps like Dashboards, Notebooks, and Security Investigator for exploratory analytics.

Grail architecture: Streaming observability data (logs, metrics, traces, and events) is stored in buckets and structured into tables. Static files (such as lookup data) provide contextual enrichment via Dynatrace Query Language.
Figure 1. Grail architecture: Streaming observability data (logs, metrics, traces, and events) is stored in buckets and structured into tables. Static files (such as lookup data) provide contextual enrichment via Dynatrace Query Language.

In addition to an uploaded data file, you also need to provide a parse pattern written in Dynatrace Pattern Language (DPL) that defines the structure of the lookup data.

Once uploaded, you can access lookup tables via the load command. Be aware that files are organized in a directory-like structure in Grail. To make it easier to find stored files, we’ve introduced autocomplete functionality. Just start typing and jump directly to the respective file.

With autocomplete, you can type a filename, instantly surface matching entries, and jump directly to the respective file.
Figure 2. With autocomplete, you can type a filename, instantly surface matching entries, and jump directly to the respective file.

To learn more about supported file types, available attributes for data ingest, or the structure of parse patterns, please have a look at our documentation.

Practical use cases

Populating lookup data is a fantastic choice for enriching data with additional context in several scenarios:

  • Mapping error codes in your logs to readable text for streamlined troubleshooting,
  • Enriching IP addresses or IDs with respective account names to convert meaningless identifiers into meaningful qualifiers that speed up triage and root cause analysis.
  • Accelerating security investigations with allow lists for security data.

Enrich your data with business context

Imagine that your system’s business-relevant events logged in Grail contain product IDs, and you’d like to enrich the IDs with the vendor’s name and some additional information from an external source.

This can easily be done with lookup tables by ingesting data containing the product and vendor information. In the example below, we use the product ID as the lookup field and enrich the business events with the mapped vendor values from the lookup table.

fetch bizevents
| lookup [ load "/lookups/vendorlist" ],
    sourceField: product.id,
    lookupField: product.id
Enriching observability with context: With the addition of custom lookup data – such as vendor metadata – we get deeper correlation of the data as well as faster insights.
Figure 3. Enriching observability with context: With the addition of custom lookup data, such as vendor metadata, we get deeper correlation of the data as well as faster insights.

Improved insights when working with security data

In another use case, imagine a security analyst is tasked with finding suspicious login attempts to your company’s network outside of business hours. Let’s assume corporate policy allows IT engineers to work from home any day, but that is not the case for accountants. The security analyst wants to understand which usernames belong to which role. Doing this manually would mean spending considerable time cross-referencing employees with their respective roles and manually creating filters based on usernames.

Creating a lookup table containing employees’ usernames and roles could significantly streamline this work, allowing the analyst to use external data to filter and summarize more accurate results.

Filtering for malicious IP addresses

Suppose your security analyst has obtained a list of fraudulent IP addresses from a threat intelligence feed that tracks malicious IP activity. These IP addresses are associated with spam, malware, botnets, or other malicious activities that expose your applications to potential threats.

The security analyst can now store this suspicious IP list as lookup data in Grail, update it whenever necessary, use it to detect and flag requests from any listed IP addresses, and leverage the data for further analysis in Security Investigator.

Flag TOR exit nodes

Going a step further, your security analyst can identify, flag, and track requests from TOR networks. TOR is an anonymizer that hides your tracks on the internet. By rerouting your internet activity via at least three other nodes before reaching your website, the TOR network obscures where requests originate, allowing bad actors to hide their identity and explore the internet with malicious intent.

The analyst creates a lookup table and populates it regularly with the latest list of TOR exit nodes. This data is used for further analysis, for example, in Security Investigator to detect login attempts that originate from TOR.

To utilize the Dynatrace® platform’s full power and set the TOR data in context, the analyst automates the fetching, writing, and uploading of the list of IP addresses using Workflows and visualizes the data with Dashboards.

Lookup data in Security Investigator: Use a DQL query to filter log entries and cross-reference IP addresses against a lookup table containing known malicious IP addresses.
Figure 4. Lookup data in Security Investigator: Use a DQL query to filter log entries and cross-reference IP addresses against a lookup table containing known malicious IP addresses.

What’s next?

Lookup tables provide a method to efficiently add context to any type of data stored in Grail. They can be used to integrate operational and transactional data, supporting your users in their day-to-day lives.

Stay tuned for further updates, such as improving our existing Snowflake Workflow Connector by adding capabilities to create and manage lookup tables, and using Security Investigator to create new lookup tables or view and filter existing tables.

Are you interested in trying out lookup tables in your own environment? This new capability is available as a public preview for all customers running the latest version of Dynatrace SaaS with an active Dynatrace Platform Subscription (DPS). It is super simple to activate; head over to our documentation to learn how.

Start enriching your observability data with lookup data to understand your business like never before!

The post Enrich your Dynatrace data with the newly introduced lookup tables appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/enrich-your-dynatrace-data-with-the-newly-introduced-lookup-tables/feed/ 0
Tell data-driven stories with new world map, gauge, and heatmap data visualizations https://www.dynatrace.com/news/blog/tell-data-driven-stories-with-new-world-map-gauge-and-heatmap-visualizations/ https://www.dynatrace.com/news/blog/tell-data-driven-stories-with-new-world-map-gauge-and-heatmap-visualizations/#respond Mon, 21 Jul 2025 15:33:09 +0000 https://www.dynatrace.com/news/?p=70095 data visualizations from Dynatrace

We’ve just expanded our visualization catalog with six powerful new additions: four new world map visualizations, the much-requested gauge chart, and an even more advanced heatmap. Available in both Dashboards and Notebooks, these new visualizations unlock a new level of visual storytelling and data analysis across any vertical or use case. Just like all our visualizations, these new additions combine smart defaults, exceptional flexibility, and a consistent, familiar configuration experience, thus delivering immediate value while still giving you the power to tailor every detail to your needs.

The post Tell data-driven stories with new world map, gauge, and heatmap data visualizations appeared first on Dynatrace news.

]]>
data visualizations from Dynatrace

Harness geographical insights from your data

With the new world map data visualizations, you can easily bring datasets into a geographical context by leveraging geospatial attributes. From Apdex scores by country, sales by branch location, and user errors by region, to global traceroutes, delivery vehicle tracking, and LLM usage by location, we’ve got you covered. All world map visualizations are world-view aware, adapting to different geopolitical perspectives and regional boundaries, with no additional configuration required.

Figure 1. World map (Choropleth) visualizing user errors by country.
Figure 1. World map (Choropleth) visualizing user errors by country.

To support a wide range of geospatial use cases, we introduced four distinct map types, each focusing on specific “stories” that can be told:

  • Dot distribution: Marks precise locations (for example, real users, store locations, delivery vehicles, cruise ships).
  • Bubble: Visualizes volume or intensity at specific points (for example, traffic, error counts).
  • Connection: Shows relationships or flows between locations (for example, service-to-service traffic, traceroutes).
  • Choropleth: Highlights aggregated metrics by country or region (for example, showing total revenue by country to quickly spot which regions are driving the most sales).

How to use world-map data visualizations effectively

World maps are among the most powerful and unique visualizations for exploring data from a geographic perspective. They make spatial patterns and regional insights immediately clear. What might be difficult to spot in a table often becomes instantly visible in a sequentially colored choropleth map, helping you uncover trends tied to real-world locations.

For the best possible results:

  • Choose the right map type for your data.
    • Use a choropleth map for aggregated data by country or region.
    • Use dot distribution or bubble maps for precise location-based data points.
    • Use connection maps to visualize routes, flows, or relationships between locations.
  • Tailor color schemes and granularity to your audience. For example, executive dashboards may benefit from simplified, high-level views, while operational dashboards might require more detailed, granular data.
  • Leverage dashboard variables (for example, region, service group, or environment) to dynamically filter and update the map in real time, enabling more interactive and context-aware exploration.
Data visualization. A choropleth map visualizing Apdex scores per European country, and a dot distribution map visualization of air traffic over the US.
Figure 2. A choropleth map visualizing Apdex scores per European country, and a dot distribution map visualization of air traffic over the US.

Track key thresholds with gauge visualizations

Gauge charts are a powerful way to visualize real-time metrics against thresholds, providing immediate visual feedback on whether you’re operating within acceptable limits. Gauge charts are perfect for tracking key performance indicators (KPIs) like response times, error rates, or throughput. They’re also well-suited for visualizing SLOs, making it easy to see how close you are to breaching targets, and for monitoring LLM performance, visualizing token usage, model load, or latency per prompt in AI Observability use cases. In business analytics, gauge charts can be used to represent conversion rates, revenue goals, or customer satisfaction scores.

Data visualization: Gauge charts used to visualize SLO statuses.
Figure 3. Gauge charts are used to visualize SLO statuses.

How to use gauge charts effectively

Select a gauge chart when you need to display a single metric and compare its current value against one or more thresholds.

When configuring a gauge chart, don’t forget to:

  • Define meaningful minimum and maximum values to ensure the gauge data visualization accurately reflects the expected range of the metric.
  • Apply custom color coding to define threshold ranges, for example, use green for “healthy,” yellow for “warning,” and red for “critical” states. This enhances visual clarity and helps users quickly interpret a metric’s status.
  • Pair gauge charts with time series visualizations to provide both real-time status and historical context, helping users understand not just the current state, but how it’s evolved over time.
Figure 4: Current status and trend over time for a latency SLO.
Figure 4. Current status and trend over time for a latency SLO.

Unlock hidden patterns with enhanced heatmaps

Traditionally, heatmaps were limited to showing how a metric changes over time. With our new enhanced heatmap, that constraint is gone. You can now combine any axis type, time series, numerical, or categorical, unlocking a much broader range of use cases and analytical depth.

This flexibility allows you to visualize complex relationships across dimensions, such as service vs. error code, user segment vs. response time, or model version vs. token usage. Whether you’re analyzing RUM data, Application Security alerts, or LLM performance, the enhanced heatmap data visualization helps uncover patterns and hotspots that were previously hidden.

Figure 5: Heatmap visualizing LLM token consumption per model
Figure 5. Heatmap visualizing LLM token consumption per model

How to use heatmap visualizations effectively

Use a heatmap to explore relationships across multiple dimensions, especially when those dimensions include a mix of value types.

Heatmaps are ideal for uncovering patterns, correlations, and anomalies in complex datasets. They’re particularly useful for comparing metrics, enabling analysis that goes far beyond traditional time-based heatmaps. In the example below, the heatmap visualizes how request durations are distributed over time: darker areas reveal when and where high volumes of similar-duration requests occur, helping you spot performance trends and anomalies instantly.

Data visualization: Heatmap showing request duration over time.
Figure 6. Heatmap showing request duration over time.
  • Choose meaningful dimensions: Combine time, numerical, and categorical axes in ways that reveal useful patterns, for example, service vs. error code, region vs. response time, or model version vs. token usage.
  • Use consistent binning and grouping: Aggregate your data in a way that supports meaningful comparisons. For example, group timestamps into hourly or daily intervals, or bucket numerical values into ranges.
  • Apply intuitive color palettes: Use color gradients that clearly communicate intensity or severity, such as diverging color palettes to highlight differences around a midpoint or sequential color palettes to represent gradual changes in intensity, ensuring the colors align with the data’s structure and purpose.
  • Avoid overcrowding by using dashboard variables and filters: Heatmaps can become hard to interpret when too many values are shown at once. Use dashboard variables (for example, service, region, environment) and filters to dynamically narrow the scope, focusing on what’s relevant without overwhelming the user.

Try the new data visualizations

Explore the new visualizations directly in Dashboards or Notebooks.

  • See applied LLM Observability use cases for these data visualizations in our latest blog post covering all new features in the Dashboards app.
  • Use the Dynatrace Playground to test out visualizations without touching your production environment.
  • If you’re interested in previous additions to our data visualization catalog, such as honeycomb or histogram, have a look at our previous blog post.

The post Tell data-driven stories with new world map, gauge, and heatmap data visualizations appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/tell-data-driven-stories-with-new-world-map-gauge-and-heatmap-visualizations/feed/ 0
Distributed tracing with Dynatrace just got even better https://www.dynatrace.com/news/blog/distributed-tracing-with-dynatrace-just-got-even-better/ https://www.dynatrace.com/news/blog/distributed-tracing-with-dynatrace-just-got-even-better/#respond Tue, 11 Mar 2025 14:58:10 +0000 https://www.dynatrace.com/news/?p=68257 Distributed tracing wth Dynatrace and OpenTelemetry

Get ready to experience a whole new world of limitless tracing power. With our latest enhancements, we’re transforming the way you work with trace data. The Dynatrace® platform now enables comprehensive data exploration and interactive analytics across data sets (trace, logs, events, and metrics)—empowering you to solve complex use cases, handle any observability scenario, and […]

The post Distributed tracing with Dynatrace just got even better appeared first on Dynatrace news.

]]>
Distributed tracing wth Dynatrace and OpenTelemetry

Get ready to experience a whole new world of limitless tracing power. With our latest enhancements, we’re transforming the way you work with trace data. The Dynatrace® platform now enables comprehensive data exploration and interactive analytics across data sets (trace, logs, events, and metrics)—empowering you to solve complex use cases, handle any observability scenario, and gain unprecedented visibility into your systems. Whether you’re using OpenTelemetry or OneAgent, operating in the cloud or on-premises—we’ve got you covered.

Introducing a new era of distributed tracing with advanced analytics

In today’s complex systems landscape, understanding the root causes of issues can be daunting, especially as applications scale and OpenTelemetry adds complexity.

Davis® AI automatically pinpoints root causes, offering immediate answers. For deeper exploration, our Distributed Tracing app empowers you to analyze raw trace data and uncover insights, whether troubleshooting errors, optimizing performance, or discovering the “unknown unknowns.”

But why stop there? Building on this solid foundation, we’re thrilled to announce two powerful platform enhancements. Say hello to advanced trace analytics and new data storage and capture options. These game-changing features elevate your data interactions, opening up vast possibilities for advanced queries and efficient data management tailored to your needs.

Figure 1. Explore every detail of your traces with in-depth exception analysis, providing easy access to exception details with full trace context.
Figure 1. Explore every detail of your traces with in-depth exception analysis, providing easy access to exception details with full trace context.

Site reliability engineers, performance architects, and developers can now leverage dynamic analysis tools like dashboards and workflows to explore trends, automate processes, and maintain control at an unprecedented level. Additionally, these queries serve as excellent starting points for more complex data explorations with Notebooks.

Get ready to maximize the full potential of your trace data—unlock deeper insights and automate like never before, all within a single platform.

Level up your analytics game: Enhanced team collaboration and advanced data insights

With traces now stored in Dynatrace Grail™, our scalable data lakehouse, you can unlock powerful new analytics capabilities, handle massive volumes of data, and run complex queries seamlessly. Combining traces with logs, metrics, Kubernetes events, and telemetry attributes gives you a complete, contextual view of your environment for unmatched end-to-end observability.

Unlock deeper insights

Using Dynatrace Query Language (DQL), you can extract game-changing insights from raw span data with precision. Use these queries to start more complex data exploration with Notebooks. This enables you to uncover hidden patterns, discover unknown unknowns, and make confident, data-driven decisions. These powerful insights can easily be transformed into interactive dashboards.

Example: Exception analysis

Understanding patterns, especially regarding exceptions, is no easy feat. However, you can begin unlocking additional insights using the Distributed Tracing app. For example, you can filter to understand endpoint performance where exception messages contain the string, access denied.

Once filtered, you can easily open a notebook with a pre-populated DQL query. You can then add additional details or modify the query as needed. Combine multiple findings in a notebook or dashboard to share your analysis with your team, allowing them to see these focused updates live in real time.

Figure 2. Open a notebook with a pre-populated DQL query, modify it, and share real-time updates with your team.
Figure 2. Open a notebook with a pre-populated DQL query, modify it, and share real-time updates with your team.

Achieve superior analytics

Transform trace data into intelligent insights by combining trace data with logs, metrics, and events. This combination allows you to enrich all data and get details in context. Use this intelligent data to see what matters to you most in real time by creating interactive dashboards and driving better decision-making. This gives you the power to break down silos, spark collaboration, and extract actionable insights with ease. It democratizes access to critical data, ensuring all teams can leverage the same reliable insights to drive impactful outcomes.

Example: Combine trace data with logs

A common scenario is understanding which frontend API requests have log messages on the backend indicative of a specific problem. Let’s look at an example where the log message contains timeout, and we want to understand the response time of traces in the context of these messages.

By linking trace data with logs, you can query across spans and related log messages. You can also summarize with DQL to understand how often a specific pattern occurs. To visualize span duration, use p99 to see the slowest percentile of span response times and then navigate directly to the traces.

Figure 3. Combine trace data with logs to identify frontend API requests with backend "timeout" log messages, and analyze response times in context.
Figure 3. Combine trace data with logs to identify frontend API requests with backend “timeout” log messages and analyze response times in context.

Automate with Dynatrace OpenPipeline

With Dynatrace, you can create custom metrics from trace data using Dynatrace OpenPipeline™, unlocking powerful new automation capabilities. It’s now possible to create metrics on OpenTelemetry and OneAgent spans with any available attribute, giving you the power to define operational, request, and method-level metrics.

OpenPipeline provides the flexibility to build metrics tailored to your specific needs, enabling you to integrate metrics seamlessly with advanced Dynatrace automation features, such as AutomationEngine and SRE Guardian. You can streamline workflows, intelligently automate repetitive tasks, proactively resolve issues, and spend more time innovating with automation, ensuring that only reliable, high-performing code reaches production.

Figure 4. OpenPipeline ingests, processes, and manages observability, security, and business data at any scale.
Figure 4. OpenPipeline ingests, processes, and manages observability, security, and business data at any scale.

We’ve only scratched the surface of scenarios where advanced analytics can make an impact—the possibilities are virtually endless. By combining advanced trace analysis, intuitive query capabilities, and seamless automation, your team can enable sharper analysis, streamline workflows, and foster innovation. This powerful approach ensures you can focus on delivering better outcomes with greater efficiency, empowering your organization to tackle complex issues with precision and agility, ultimately bringing unprecedented value.

Extended trace retention: Retain data longer when it matters

Need to analyze trends over the long term or adhere to compliance requirements? With extended trace retention, you can store trace data for up to 10 years. This feature ensures your organization is well-equipped for trend analysis and detailed post-mortem reviews—all while meeting regulatory requirements.

Retention policies are fully configurable in OpenPipeline. You can share data in buckets with varying retention times depending on the use case, allowing you to target specific applications or error-prone services for longer storage. This precision reduces storage costs while ensuring you retain the data that matters most.

Extended trace ingest

You can now customize trace ingestion rates to meet your specific needs. While most trace data is already ingested at high coverage rates, this option gives you more granular control over your trace volume. Ingest as much data as you want, above and beyond what is already included in your Dynatrace license.

Maximize the value of your OpenTelemetry data

At Dynatrace, we love OpenTelemetry. We champion open source innovation and recognize OpenTelemetry’s influence in setting observability standards (we’re also a top contributor). That’s why our advanced capabilities were designed from the ground up with OpenTelemetry at its core. We built our entire new tracing experience on OpenTelemetry semantic conventions and expanded from there. Now, OpenTelemetry users can troubleshoot and analyze while leveraging  OTel standards. This commitment empowers you to simplify complexity and innovate faster by extracting maximum value from your data, regardless of origin.

Experience the future of Distributed Tracing

At Dynatrace, we believe that observability should be effortless and completely on your terms. With our latest advancements, we’re helping you manage complexity, innovate faster, and push boundaries. Our solution adapts seamlessly to your ecosystem, whether you use OpenTelemetry or run cloud-native or on-premises workloads. Built to handle enterprise scale, the Dynatrace platform processes massive volumes of data in real time while unifying insights across all teams in your organization.

We’re rolling out this functionality to existing Dynatrace Platform Subscription (DPS) customers. Elevate your observability journey with these new possibilities—tailored to your needs, your way.

If you’re not a DPS customer, you can try out the new Distributed Tracing experience with prepopulated data on the Dynatrace Playground.

If you’re new to Dynatrace and want to try out the new Distributed Tracing experience with your own data, check out our free trial

Take the leap today and discover how Dynatrace can revolutionize your approach to observability.

The post Distributed tracing with Dynatrace just got even better appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/distributed-tracing-with-dynatrace-just-got-even-better/feed/ 0
Transform data into insights with Dynatrace Dashboards and Notebooks https://www.dynatrace.com/news/blog/transform-data-into-insights-with-dynatrace-dashboards-and-notebooks/ https://www.dynatrace.com/news/blog/transform-data-into-insights-with-dynatrace-dashboards-and-notebooks/#respond Wed, 16 Oct 2024 18:45:55 +0000 https://www.dynatrace.com/news/?p=66227 Explore Kubernetes metrics graphic

When we launched the new Dynatrace experience, we introduced major updates to the platform, including Grail™, our innovative data lakehouse unifying observability, security, and business data, and Dynatrace Query Language (DQL) for accessing and exploring unified data. While Grail and DQL opened up nearly limitless possibilities for data exploration, mastering DQL was necessary to fully […]

The post Transform data into insights with Dynatrace Dashboards and Notebooks appeared first on Dynatrace news.

]]>
Explore Kubernetes metrics graphic


When we launched the new Dynatrace experience, we introduced major updates to the platform, including Grail™, our innovative data lakehouse unifying observability, security, and business data, and Dynatrace Query Language (DQL) for accessing and exploring unified data. While Grail and DQL opened up nearly limitless possibilities for data exploration, mastering DQL was necessary to fully leverage the power of Grail. Our latest enhancements to the Dynatrace Dashboards and Notebooks apps make learning DQL optional in your day-to-day work, speeding up your troubleshooting and optimization tasks.

In this blog post, we look at these enhancements, exploring methods for monitoring your Kubernetes environment and showcasing how modern dashboards can transform your data. Furthermore, we illustrate how these methods work seamlessly with Dashboards and Notebooks to enhance their effectiveness.

Get real-time insights by transforming complex data into dynamic, interactive dashboards

The many paths to building a dashboard or notebook

Getting started with the new Dashboards is now easier than ever, offering unprecedented ease and capabilities for exploring your data. We’ve not only improved how you interact with data in dashboards and notebooks, we also enhanced the way that underlying data can be shared across apps. These updates expand your options for exploration and creation, helping you to build your dashboards and notebooks quicker and more intuitively.

You can now:

Let’s look at each of these paths through an end-to-end use case focused on Kubernetes monitoring.

Kickstart your creation journey using ready-made dashboards and notebooks

Creating dashboards and notebooks from scratch can take time, particularly when figuring out available data and how to best use it. Ready-made dashboards and notebooks address this concern by offering pre-configured data visualizations and filters designed for common scenarios like troubleshooting and optimization.

These ready-made dashboards offer your platform engineers, who oversee Kubernetes environments, immediate and comprehensive data visibility. This allows platform engineers to focus on high-value tasks like resolving issues and optimizing performance rather than spending time on data discovery and exploration.

Kickstarting the dashboard creation process is, however, just one advantage of ready-made dashboards. Let’s assume you’re already using the new Kubernetes app, which offers a comprehensive overview of your Kubernetes environments and their telemetry. There are cases where more flexible data presentation is needed. Our new ready-made dashboards for Kubernetes not only provide instant insights into your clusters, nodes, workloads, or pods but also enable you to extend and customize the data shown in the Kubernetes app, leveraging the context-rich data from Dynatrace Grail. So, for example, if you need to seamlessly integrate metrics with logs for your workloads, you can create a customized view based on the pre-configured dashboard that consolidates all critical signals in one place, which is particularly essential for troubleshooting.

Finding ready-made dashboards is straightforward. Navigate to the list of dashboards and set the filter at the top left to Ready-made. Select the title of any dashboard that interests you, or use the search bar to narrow down the results.

Visualization: Leverage ready-made dashboards to create yours video thumbnail

Accelerate data exploration with seamless integration between apps

In developing the new Dynatrace experience, our goal was to integrate apps seamlessly by sharing the context when navigating between them (known as “intent”), much like sharing a photo from your smartphone to social media. This approach acknowledges that in any organization, software doesn’t work in isolation; boundaries and responsibilities are often blurred. This is even more true for critical scenarios like troubleshooting, which often requires more than the capabilities of a single person or app.

Let’s make this more tangible by using the Kubernetes cluster dashboard and demonstrating how this concept helps you to:

  • Seamlessly navigate between apps while maintaining context.
  • Effortlessly explore data in Dynatrace and create dashboards from it.

When working with the Kubernetes cluster dashboard, you have two options for digging deeper into further analysis, both using the Kubernetes app. You can use dynamic markdown links, which include the values of the actual dashboard variables, or you can utilize the open-with feature (the “intent” concept), which uses the actual context of the dashboard tile you’re viewing. With this latter approach, you even have the choice of passing a single value (Open field with) or all underlying data (Open record with) for the respective element (row, series, cells, etc.) when navigating to another app.

Visualization: Accelerate data exploration with seamless integration between apps video thumbnail

Next, let’s use the Kubernetes app to investigate more metrics. The intent concept and the open with feature can also be applied in reverse to include data or specific visualizations from an app on a particular dashboard. An example of this is shown in the video above, where we incorporated network-related metrics into the Kubernetes cluster dashboard.

Start from scratch with the new Explore interface for metrics in Dashboards and Notebooks

Once you’ve learned how to monitor your Kubernetes cluster using a ready-made dashboard and extending it with context from other apps, the next step is understanding how to create and extend such dashboards using the Dashboards or Notebooks app.

Exploring and adding metrics from scratch

Let’s revisit our example from the last chapter and add the same Kubernetes network metrics, this time by using the new Explore metric interface that allows you to:

  • Browse and add multiple metrics to a single tile
  • Apply basic commands such as aggregation, filter, and split
  • Use expressions to do calculations based on previously added metrics

Visualization: Exploring and adding metrics from scratch video thumbnail

The revised Explore interface, as shown in the clip above, now includes logs, metrics, events and business events, offering an improved filtering experience that enables you to:

  • Type ahead to add, edit, or remove available filters
  • Control how filters are applied via a rich set of operators (=, !=, in, not in, >=, <=, >, <)
  • Place wildcards before and after your filter values to automatically generate the best matching DQL when using startsWith, endsWith, or contains.
  • Control how filters are combined with logical operators, such as AND or OR
  • Easily filter entities by ID, name, or tags in the web UI
  • Get suggestions for metric values and entities (IDs, names, tags) for all data types

Build your dashboard effortlessly with only a few clicks

Blend metrics with data from Explore logs for a more comprehensive view to start log analysis

With the enhanced Explore Logs interface, retrieving and viewing logs from your Kubernetes workloads is straightforward. By incorporating a new tile, you can integrate these logs into your dashboard along with key metrics, such as the new Kubernetes network metrics we added earlier.

Leverage dashboards to monitor your environment in real time through log data. Once you identify an anomaly that requires your attention, you can start troubleshooting by delving further into the issue using the Open with option and the intent mechanism in the new Logs app. This app provides advanced analytics, such as highlighting related surrounding traces and pinpointing the root cause, as illustrated in the example below.

Visualization: Enhanced Explore Logs interface video thumbnail

Leveraging the capabilities of Grail and Smartscape® topology, Dynatrace seamlessly integrates logs, metrics, and traces to offer enhanced context for troubleshooting and in-depth analysis. This integration facilitates a comprehensive understanding of individual transactions through the Distributed Traces app and aids in pinpointing the root cause of issues when using the Problems app.

Intuitive data access with Davis CoPilot AI assistant

There’s also a brand new and completely different option for analyzing data using natural language; using the power of generative AI, Davis CoPilot™ converts your conversational prompts into accurate DQL commands, allowing both non-technical users as well as experienced data analysts to make data-driven decisions faster than ever before.

Davis CoPilot in Dynatrace screenshot

To learn more about how Davis CoPilot empowers you and your teams, see our blog post, Announcing General Availability of Davis CoPilot: Your new AI assistant.

Search metrics from anywhere

A speedy way to begin your data exploration journey, particularly if you already know which data you need, is to utilize our global search feature to effortlessly find and explore metrics from anywhere on the Dynatrace platform. This feature lets you explore any available metric and add it to Notebooks or Dashboards.

Imagine a colleague mentioning a newly released metric for Kubernetes during a coffee break. Rather than manually exploring the Kubernetes app you can simply open the Dynatrace global search and enter “Kubernetes network.” The relevant metrics are then immediately displayed alongside further details.

This efficient method allows you to easily browse and identify the appropriate metrics; adding them to your notebooks and dashboards requires just a single click.

Browse and identify the appropriate metrics in Dynatrace screenshot

Get started discovering and exploring your data

It has never been easier to analyze data within Dynatrace. Kickstart your data exploration journey and familiarize yourself with ready-made dashboards and the new Explore data interface. By the way, we also added new data visualization capabilities by adding new chart types and chart interactions.

Curious about our latest releases and upcoming features for Dashboards and Notebooks? Check out our community roadmap thread to stay updated!

Are you ready to try out the new Explore Data features?

The post Transform data into insights with Dynatrace Dashboards and Notebooks appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/transform-data-into-insights-with-dynatrace-dashboards-and-notebooks/feed/ 0
Next-level interaction and customization of data visualizations in Dynatrace Dashboards and Notebooks https://www.dynatrace.com/news/blog/custom-data-visualizations-in-dashboards-and-notebooks/ https://www.dynatrace.com/news/blog/custom-data-visualizations-in-dashboards-and-notebooks/#respond Thu, 10 Oct 2024 15:19:43 +0000 https://www.dynatrace.com/news/?p=66089 Abstract image depicting unlocking business potential with Dynatrace using power dashboarding

Enhanced data visualization options change how you present, analyze, and interact with your data in the Dynatrace Dashboards and Notebooks apps. We added honeycomb and histogram visualizations, made the visualizations more interactive, and introduced numerous custom settings, giving you all the tools you need to extract maximum value from your unified data.

The post Next-level interaction and customization of data visualizations in Dynatrace Dashboards and Notebooks appeared first on Dynatrace news.

]]>
Abstract image depicting unlocking business potential with Dynatrace using power dashboarding

Take your monitoring, data exploration, and storytelling to the next level with outstanding data visualization

All your applications and underlying infrastructure produce vast volumes of data that you need to monitor or analyze for insights. Visualizations help to curate data into a form that is more accessible to understand, highlighting trends and outliers, gaps, clusters, or patterns. At a single glance, visualizations can raise questions that stimulate further exploration or indicate problems.

Each type of visualization tells a different story and is best suited for a particular use case. If you want your data to speak to its audience, you need a comprehensive toolkit of visualizations and customization options. Good visualizations are not just static, unintelligent data presentations; they enable interaction and ideally serve as a starting point for subsequent analysis.

Dynatrace unified analytics capabilities for observability are top-of-the-class (Gartner Magic Quadrant 2024), enabling you to query and analyze all your observability data across your enterprise. The Dynatrace Notebooks and Dashboards apps are the perfect starting point for visualizing and understanding your data for monitoring or in-depth analysis.

Over the last year, we introduced many new functionalities and updated our visual presentation to provide you with an all-new experience:

  • Optimized for exploring large data sets: We’ve optimized our entire interface, components, and visualizations to present large volumes of data while providing all the required flexibility.
  • Visualize with a single click: Even inexperienced users can visualize data sets and create graphs in seconds.
  • Broad range of visualizations: Our curated catalog provides numerous visualization types and hundreds of customization options, including the newly added honeycomb and histogram visualizations.
  • Interact with your data: We integrated additional data interaction methods to provide more information immediately.
  • Quick analysis with Davis CoPilot™: Explore your data through natural language by translating your conversational prompts directly into DQL queries.

Now, let’s introduce you to our two newest entries to our visualization catalog and tell you about the great things you can do with them.

New: identify hotspots with the honeycomb visualization

Honeycombs are great for visualizing health in complex and distributed systems, enabling you to visualize countless entities effectively and at scale. They have become a quasi-standard in the industry, especially for infrastructure monitoring visualizations.

The new honeycomb visualization in Dynatrace enhances your health dashboards and offers many customization options to tailor it to your needs. For example, it supports string and numerical values, enabling a multitude of different use cases. There are many practical applications of honeycombs; here is a small sample of just a few of them:

Ready-made dashboard for problem reporting
Figure 1. Ready-made dashboard for problem reporting
  • Problem visualization: The new, ready-made dashboard for the Problems app features two honeycomb visualizations. At one glance, you see which entities are particularly affected by problems, and you can also identify the blast radius of a single problem to understand how many entities have been affected. Have a look at them on our Dynatrace Playground.
  • Infrastructure health: A honeycomb chart is often used to visualize infrastructure health. You can use it to visualize CPU utilization across your hosts, disk space used, server-side response time, web request/service failure rates, or any other area where you need to spot outliers immediately.
  • Service Level Objectives (SLO) tracking: Honeycomb charts can visualize SLOs, helping you monitor whether your services meet performance and reliability targets. Based on the color, you immediately see if any SLOs are off track. This can guide you in prioritizing issues that impact user experience.
Honeycomb visualization highlighting outliers
Figure 2. Honeycomb visualization highlighting outliers

How you get the best results with honeycombs

Honeycombs highlight hotspots that require attention. To achieve the best visual outcome, we recommend experimenting with the available customization options.

  • Try different cell shapes. The honeycomb visualization also supports circles and square variants, allowing you to differentiate between use cases clearly.
  • Use color coding to tell a story. Use different diverging and sequential color palettes to highlight patterns and insights in your data. The chart also supports conditional coloring for both string and numeric values.
  • Min and max limits. Set an applicable, expected value range in the honeycomb visualization to ensure effective coloring and highlighting of hotspots and outliers. For example, set the value range for CPU consumption from 0% to 100%. In other use cases, you want to ensure a consistent zero-baseline (that is min = 0), but with an automatic scaling max value (max = auto).

Go to our documentation to learn more about implementing honeycomb visualizations on your dashboards or notebooks.

New: explore your data with histograms to identify patterns

The histogram chart is a crucial visualization for understanding the distribution patterns of values within a given dataset. It shows where the peaks of the distribution are, whether the distribution is skewed or symmetric, and whether there are any outliers.

While histograms look much like time-series bar charts, they’re different in that each bar represents a count (often termed frequency) of metric values. These bars are called bins or buckets; their width represents a value range. The height of the bar reflects the frequency or count of data points within a bucket.

Histogram showing the distribution of failed payments, split by credit card provider
Figure 3. Histogram showing the distribution of failed payments, split by credit card provider

The use cases and underlying metrics analyzed via histograms are extremely broad:

  • Latency distribution: Histograms can show the distribution of request latencies, helping you understand how many requests fall into different latency buckets. This is useful for identifying performance bottlenecks and understanding the overall user experience.
  • Resource utilization: Use histograms to visualize the distribution of CPU or memory usage across different instances or containers. This helps identify outliers and understand the overall resource consumption patterns.
  • Distributed tracing: Histograms can analyze the response times of different endpoints or services, allowing you to pinpoint which parts of your system are slower and need optimization.
  • User behavior tracking: Track user interactions, such as login or page load times, to understand how users are experiencing your application. This can help identify areas for improvement in user experience.

How you get the best results with histogram charts

Experiment with bucket sizes. Use DQL’s built-in range function, together with the summarize command, to bucketize your data. The choice of bin size has an inverse relationship with the number of bins. The larger the bin sizes, the fewer bins will be needed to cover the whole range of data. With a smaller bin size, you’ll get more bins.

It is worth taking some time to test out different bin sizes to see how the distribution looks in each one, then choose the best plot that represents the data. Try out different range sizes or bin ranges (“widths”) with your data set – doing so can help you identify distinct patterns in the data.

If you have too many bins, the data distribution will look rough, making it difficult to discern the signal from the noise. On the other hand, with too few bins, the histogram will lack the details needed to discern any helpful patterns from the data. Identify common distribution patterns such as bimodal, comb, edge peak, normal, skewed, and uniform.

Add split by parameter. This can help compare sub-distributions; however, it is ideally limited to only two sub-divisions, such as A vs. B, North vs. South, Open vs. Closed, Male vs. Female, etc.

If you want to learn more about how to best use histograms for OTel observability, check out Mikko Viitanen’s OpenTelemetry histograms blog post. In this post, you’ll learn how to define and monitor service-level objectives with histograms that can be used to set up alerts.

After introducing the new visualizations, we will now look at our new custom settings.

Optimize your visualizations with many new configuration options

We elevated the handling of visualization settings across Dashboards and Notebooks in three powerful ways:

  1. We extended the customization options for all data visualizations by offering an additional 35+ configuration options, such as custom column types in the table, so you can better tailor your visualizations to your specific requirements.
  2. We improved the usability of all visualization settings by introducing new unified UI controls across both apps. Now, it’s even easier for users to customize visualizations as they see fit for any use case.
  3. We introduced a visualization settings search so that you can instantly search all settings and find the exact configuration or customization option you’re looking for.
New configuration options
Figure 4. New configuration options

Get more details immediately

In our recent release, we added more functionality to enable you to interact with your data. You can now zoom into your data or pan left and right in all time-series visualizations. This allows you to zoom in on a particular data set, see underlying details, or change the displayed data altogether. This allows you to dive deeper into your data anytime without needing to modify/rewrite your original query.

Taking this interaction a step further, in Dashboards, if you zoom in on one visualization, it will automatically apply to all other visualizations on the same dashboard.

The functionality is automatically available in all time-series charts in the Dashboards and Notebook apps.

Zoom-in

Click and drag any time series chart in your Dashboard or Notebook to zoom in on your data. The interaction will re-fetch data for the selected timeframe (in most cases, with a more fine-grained interval).

Alternatively, you can zoom in and out with the integrated chart toolbar, located in the top right corner of the chart, keyboard shortcuts, or touch gestures.

In Dashboards, the zoom interaction adapts the timeframe for the current chart and updates the entire dashboard, automatically synchronizing all other tiles and visualizations.

Dynatrace dashboard zoom interaction video thumbnail

Figure 5. Zoom-in

Panning

You can click and drag the chart’s x-axis to pan it to the left or right while maintaining the current zoom level. Alternatively, you can pan left and right using the middle mouse button, activating the chart’s pan mode (via the integrated chart toolbar), keyboard shortcuts, or touch gestures.

Keyboard accessibility

If you prefer interacting with your visualizations via the keyboard, we also have you covered: all interactions are also available via the keyboard. Switch between modes with “e” to explore your data or “p” to pan the chart left and right. Holding down the “CMD/CTRL” key while pressing “arrow up” or “arrow down” will let you zoom in or out of the chart. Pressing “r” will quickly let you reset all the zoom and pan changes.

What’s next

World map and heatmap visualization

We constantly exchange with our community and add further visualizations to our curated catalog where it makes sense. We’re currently working on introducing multiple world map visualizations and a heatmap, which you can expect to be released in the upcoming quarters.

World Map preview
Figure 6. World Map preview

Synchronized crosshair

In addition to the interactions described above, we will soon support automatically synchronizing all crosshairs across all time-series charts. That way, you can compare multiple charts more easily, regardless of the metric or time span.

Try our new visualizations now

If you’re an existing Dynatrace customer, visit your Dashboards app and check out our ready-made dashboards. Our new Getting Started document explains how to create and customize honeycomb and histogram visualizations. Also, have a look at our new Problems Dashboard, which you can access and download directly from the Dynatrace Playground.

To learn more about how data visualizations can enhance your app insights, check out our latest episode of Inside Dynatrace Apps with Mikele Hasson and Penny Scully

From data to decisions video thumbnail

The post Next-level interaction and customization of data visualizations in Dynatrace Dashboards and Notebooks appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/custom-data-visualizations-in-dashboards-and-notebooks/feed/ 0
Learn how to create a Davis AI anomaly detector on Grail https://www.dynatrace.com/news/blog/create-a-davis-ai-anomaly-detector-on-grail/ https://www.dynatrace.com/news/blog/create-a-davis-ai-anomaly-detector-on-grail/#respond Tue, 11 Jun 2024 19:22:55 +0000 https://www.dynatrace.com/news/?p=64325 Davis AI Fetch logs

From working with Dynatrace Notebooks, you know that exploratory analytics are crucial for uncovering the narratives within your organization’s data. By leveraging visual data analytics and collaboration input from development, security, and business teams, such insights become transparent, enabling immediate understanding and action on the implications for your business. Further, it’s essential to take automated actions to proactively use anomaly detection to determine if your business is at risk. Such anomaly detection should be implemented in straightforward steps, as described in this blog post.

The post Learn how to create a Davis AI anomaly detector on Grail appeared first on Dynatrace news.

]]>
Davis AI Fetch logs
Update: We’ve enhanced anomaly detection on Grail with Dynatrace Intelligence, enabling smarter, AI-powered insights and automated actions across the Dynatrace platform.
Dynatrace Intelligence builds on Davis AI®, advancing how teams detect, analyze, and respond to anomalies in their data.

Dynatrace Grail™ data lakehouse provides contextual analytics across unified observability, security, and business data. It allows you to query and combine data anytime using the Dynatrace Query Language (DQL). This enables exploratory data analysis and the ability to collaborate visually on the results with your colleagues.

Anomaly detection in Notebooks

You likely encounter “why” questions in your daily work. Why did we have an outage? Why did the system behave differently? Why did I receive an alert? These questions can be effectively investigated in Dynatrace Notebooks, where you can easily compile the necessary data and break it down into a time series. However, in the time series example below, we must determine whether the number of access attempts to our example Travel Mobile app is normal or abnormal.

Figure 1: Generated time series based on access logs in Notebooks
Figure 1: Generated time series based on access logs in Notebooks

In many cases, it’s evident, based on your past experiences looking at time series data, whether or not something is an anomaly. But how can you automate your expertise? Such automation could ensure that you and your colleagues don’t have to manually monitor time series to identify whether or not they include anomalies.

Davis® AI provides such automated anomaly detection out of the box. Still, your business requires the flexibility of Davis AI to detect anomalies based on your specific requirements, for example, to automatically generate a Davis problem based on a detected anomaly. For this purpose, we provide the Davis AI Analyzer, which allows you to select a specific analyzer. Three anomaly detection analyzers are available, each equipped with unique mechanisms to detect anomalies in your data that significantly deviate from the norm.

One unique feature of the Davis AI Analyzer is that it works on any time series, regardless of its origin—whether generated with makeTimeseries from events, business events, logs, or other sources or the joining of different time series. As you can see in the screenshot below, Davis AI Analyzer gains the full power of DQL, making Davis anomaly detection even more flexible and stronger than ever. This power can be easily experienced by selecting the desired Davis anomaly detection analyzer in Notebooks or Dashboards.

Figure 2: Using the seasonal baseline anomaly detection analyzer in Notebooks.
Figure 2: Using the seasonal baseline anomaly detection analyzer in Notebooks.

By selecting the seasonal baseline analyzer, Davis AI recognizes that the number of attempted accesses to the app in this example doesn’t deviate from the norm based on the past data during the same period. The time series falls within the seasonal green confidence band. A potential alert would be visually simulated if the time series fell outside this band.

This anomaly detector observes the number of attempted accesses per minute and triggers an event when anomalies are detected. You can create a similar Davis anomaly detector in a few simple steps.

Automate your experience with Davis Anomaly Detection

In Notebooks, select open with and choose Davis Anomaly Detection; all settings required for creating an anomaly detector will be carried over.

Create a new anomaly detector in Davis Anomaly Detection.
Figure 3: Create a new anomaly detector in Davis Anomaly Detection.

The new anomaly detector is created in four steps; the first two steps are carried over automatically from Notebooks. Let’s start with the most straightforward step, Get started, where you define a title for your anomaly detector and a description for the configuration.

The next two steps, as mentioned, have already been prefilled from Notebooks. In the Configure your query step, you’ll find the DQL query you predefined, and in the Customize parameters step, you’ll find your selected anomaly detection analyzer. The last significant step, the Create an event template step, remains. Here, you can define the template for your event and describe all essential information for the subsequent process.

Define the description and properties in the event template.
Figure 4: Define the description and properties in the event template.

What makes this template exceptional is that you can use {placeholder} hints to add additional context to the text about the event. For example, the value of the violation or the source entity where the anomaly was detected. This ensures that all essential information about the event is immediately visible to the Site Reliability Engineer (SRE). After completing all four steps, we can create the Davis anomaly detector by selecting Create. The anomaly detector will automatically monitor your defined time series every minute and trigger your specified event upon detection of an anomaly.

The new anomaly detector is now listed in Davis Anomaly Detection. Here, you’ll find all anomaly detector configurations, and you can filter them according to your specific criteria. Additionally, you can expand this table with extra information about the configurations, such as when the anomaly detectors were last modified.

Overview of anomaly detectors available within Davis Anomaly Detection.
Figure 5: Overview of anomaly detectors available within Davis Anomaly Detection.

Of course, you always have the option to reopen an anomaly detector directly in Notebooks, where all configuration settings are carried over. You also have the option to display a preview of your anomaly detector directly in Davis Anomaly Detection.

Figure 6: Visualize your custom anomaly detectors in Notebooks without leaving Davis Anomaly Detection.
Figure 6: Visualize your custom anomaly detectors in Notebooks without leaving Davis Anomaly Detection.

The exciting challenge is finding answers to your everyday “why” questions using Grail and DQL analytics capabilities. If the answer is successfully identified in a time series and you want to automate the result with anomaly detection, this can be done in just a few steps. We recommend you explore the new Davis Anomaly Detection analyzer in Notebooks; we’re confident you’ll quickly discover its many uses.

Try out Davis Anomaly Detection

Want to know more? Check out the following video, in which Andreas Grabner and I collaborated on a new episode of the Dynatrace Observability Clinic. Here, we share a live introduction to Anomaly Detection based on DQL.

We also recommend watching the exciting use case for Anomaly Detection and the 5 Pillars of Data Observability.

What’s next

Davis Anomaly Detection is automatically enabled for all Dynatrace SaaS environments with the release of Dynatrace version 1.291. No effort is needed from your side. We’re, of course, highly interested in your feedback. So, please head to the Dynatrace Community and share your suggestions and product ideas to help us continuously improve Dynatrace Anomaly Detection.

Are you interested in learning more? In Dynatrace Documentation, you can learn more about Davis Anomaly Detection and how to use anomaly detection within Notebooks, or look at our Playground, where you can explore practical examples of how to utilize Davis AI Analyzer in your anomaly detection.

See examples of using Davis AI to detect anomalies. Visit Dynatrace Playground.

The post Learn how to create a Davis AI anomaly detector on Grail appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/create-a-davis-ai-anomaly-detector-on-grail/feed/ 0
Discover the power of the Dynatrace platform on Azure with Azure Native Dynatrace Service https://www.dynatrace.com/news/blog/azure-with-azure-native-dynatrace-service/ https://www.dynatrace.com/news/blog/azure-with-azure-native-dynatrace-service/#respond Thu, 23 May 2024 14:55:49 +0000 https://www.dynatrace.com/news/?p=64121 Azure Native Dynatrace Service

In March 2024, Dynatrace made its AI-powered platform generally available on Microsoft Azure. You can now leverage the unique features of the Dynatrace® platform and Grail with the Azure Native Dynatrace Service (ANDS) to get instant insights into Azure resources within your Azure subscriptions. This offers a best-in-class product experience with deep integration into Microsoft Azure, such as instant deployment of Dynatrace OneAgent® from the Microsoft Azure Portal and data flowing into the Dynatrace platform with no-touch configuration. Furthermore, leveraging the unique features of the Dynatrace platform allows you to get deep insights into any workload, whether in the cloud or on-premises.

The post Discover the power of the Dynatrace platform on Azure with Azure Native Dynatrace Service appeared first on Dynatrace news.

]]>
Azure Native Dynatrace Service

Dynatrace platform’s unique features are now available on Azure

As of today, Azure customers can leverage the latest Dynatrace core innovations on Microsoft Azure, including:

  • Dynatrace Grail™ data lakehouse unifies the massive volume and variety of observability, security, and business data from cloud-native, hybrid, and multicloud environments while retaining data context to deliver instant, cost-efficient, and precise analytics.
  • Dynatrace® AutomationEngine features a no- and low-code toolset and leverages Davis® AI to empower teams to create and extend customized, intelligent, and secure workflow automation across cloud ecosystems.
  • Dynatrace® AppEngine features a no- and low-code toolset and leverages Davis AI to empower teams to easily create and share custom, intelligent, and secure apps that leverage insights from data generated by their clouds.
  • The new Dynatrace user experience, including powerful dashboarding capabilities and interactive Dynatrace Notebooks, drives tighter cross-team collaboration and enables more people within the organization to make data-backed decisions.

Azure Native Dynatrace Service allows easy access to new Dynatrace platform innovations

Dynatrace has long offered deep integration into Azure and Azure Marketplace with its Azure Native Dynatrace Service, developed in collaboration with Microsoft. With the AI-powered Dynatrace platform now generally available on Azure, Azure Native Dynatrace Service customers can now leverage the full AI power of the Dynatrace platform directly from Azure. With Dynatrace directly connected to your Azure subscriptions, Dynatrace OneAgent® can be deployed directly to VMs and App Services from the Microsoft Azure Portal. One-click activation of log collection and Azure Monitor metric collection in the Microsoft Azure Portal allows instant ingest of Azure Monitor logs and metrics into the Dynatrace platform. For more details, see the blog post, Set up AI-powered observability for your Microsoft Azure cloud resources in just one click.

The following figure shows the benefits of Azure Native Dynatrace Service. For more details, please see the blog post Dynatrace and Microsoft Azure integrate to help accelerate your cloud transformation.

Benefits of Dynatrace for Azure native integration

Using the Dynatrace platform on Azure allows Azure Native Dynatrace Service customers to get instant insights into their workloads with a comprehensive set of new innovative features with the best product experience from the Microsoft Azure Portal:

  • The Dynatrace platform is automatically provisioned as part of the Azure Marketplace purchase process, allowing you to set up observability for workloads in minutes.
  • Configuring the Dynatrace platform and deploying Dynatrace OneAgent from the Microsoft Azure Portal ensures that observability data from Azure resources arrives in Dynatrace within seconds.
  • New Dynatrace platform features—with the Grail data lakehouse at their core—allow you to easily query and get insights into your workloads with the new Dynatrace user experience using DQL, Notebooks, and Dashboards, driven by a new Automation Engine, and AppEngine.

Leverage Dynatrace to get observability into all resources in your Azure tenant

This blog post shows you how the Dynatrace platform, together with the Azure Native Dynatrace Service—and unique new capabilities such as Grail, Dashboards, Notebooks, and the Cloud app—can give instant insights into all the Azure resources deployed in your Azure tenant.

Set up complete monitoring for your Azure subscription with Azure Monitor integration

After activating the Azure Native Dynatrace Service (see Dynatrace Documentation for details), the Azure Monitor integration is enabled easily via the Microsoft Azure Portal, as shown in the following screenshot. There’s no need for configuration or setup of any infrastructure.

Microsoft Azure Metrics and logs

Enabling the Azure Monitor integration in the Microsoft Azure Portal triggers the linked Dynatrace environment to start polling metrics and resource metadata for the Dynatrace Smartscape® topology model. Furthermore, as shown in the above screenshot, the collection of resource logs and activity logs is easy to turn on.

Integrate multiple Azure subscriptions under a single Dynatrace environment

If you have multiple Azure subscriptions in your Azure tenant, it’s best practice to integrate the Azure subscriptions with a single Dynatrace environment. Establishing a single source of truth for all observability data reveals the full power of the Dynatrace platform by querying all data with comprehensive new capabilities from the one Dynatrace Grail data lake house for data storage.

To monitor multiple subscriptions within your Azure tenant

  1. Create a new Dynatrace environment within the first subscription by creating the first Dynatrace resource (Option Create a new Dynatrace environment as shown in the following screenshot).
  2. Then, create a Dynatrace resource in all other Azure subscriptions and link the Dynatrace environment created in Step 1 (the option Link Azure subscription to an existing Dynatrace environment is shown in the following screenshot).

Create a Dynatrace resource in Azure

More information can be found at Dynatrace link to existing. Also, see how to automate this process with bicep and how to automate with Azure CLI.

Azure use cases with Dynatrace Apps

The following new Dynatrace® Apps can be leveraged to get instant insights into all Azure resources of Azure subscriptions connected via the Azure Native Dynatrace Service:

  • Clouds enables cross-subscription and cross-region observability in one place.
  • Dashboards leverages the power of DQL for Azure monitoring in one place.
  • Notebooks offers advanced Azure observability analytics with DQL.

Clouds

Clouds allows easy investigation and troubleshooting of all Azure resources from the subscriptions previously connected to Dynatrace.

The following screenshot shows that you can easily search for resources by name, resource type, region, and tags across all connected Azure subscriptions.

Azure OpenAI in Dynatrace screenshot

Clouds provides resource properties, metrics, problems, and events in a single view, as shown below.

Azure OpenAI in Dynatrace screenshot

Azure OpenAI in Dynatrace screenshot

Davis AI automatically analyzes ongoing problems with your resources. Clouds makes it easy to search for all resources affected by detected problems and navigate directly to the problem details.

Azure OpenAI in Dynatrace screenshot

Azure OpenAI in Dynatrace screenshot

Clouds also supports getting the necessary insights for cloud governance. For example, Dynatrace uses automatic tagging to mark the user who created an Azure resource. Searching for all resources with a specific tag (for example, created-by) across your connected Azure subscriptions is easy with Clouds.

Azure OpenAI in Dynatrace screenshot

Dashboards

Dashboards provides new dashboarding capabilities powered by the new Dynatrace platform query language (DQL), which fully leverages the new Dynatrace experience.

With DQL and Dashboards, combining all observability data in one dashboard is easy. For example, it’s possible to combine resource metadata, resource health from activity logs, Azure Monitor metrics, Davis AI problems, OneAgent process details, Synthetic monitoring results, and much more in a single view, as shown below.

Azure Dashboard in Dynatrace screenshot Azure Dashboard in Dynatrace screenshot

For each section, it’s possible to drill down for detailed context and learn the underlying DQL query.

Azure Dashboard in Dynatrace screenshot

The following shows a simple DQL summarizing all Azure Virtual Machine cores in the connected Azure subscriptions. By querying the Dynatrace entity model, cores are summed up from entity metadata.

fetch dt.entity.azure_vm
| parse azureVmSize, "KVP{LD[a-zA-Z]+:key'='(LONG:valueLong | STRING:valueStr)','?}:vmsize"
| fieldsAdd cores =   vmsize[`numCores`], memoryMb =   vmsize[`memoryMb`]
| summarize Cores = sum(cores)

Notebooks

Dynatrace Platform offers Notebooks, which allows advanced ad-hoc analysis of Azure environments based on metrics, logs, spans, events, and entities.

The following DQL from Notebooks queries data for all recent Azure health events and aggregates all resource health events from all Azure subscriptions connected via the Azure Native Dynatrace service.

fetch logs //, scanLimitGBytes: 500, samplingRatio: 1000
| filter matchesValue(cloud.provider, "Azure")
| filter azure.resource.type == "MICROSOFT.COMPUTE/VIRTUALMACHINES" or azure.resource.type == "MICROSOFT.COMPUTE/VIRTUALMACHINESCALESETS/VIRTUALMACHINES"
| parse content, "JSON:json"
| fieldsAdd operationName=json[operationName],
category=json[category],
resultType=json[resultType],
callerIpAddress=json[callerIpAddress],
action=json[identity][authorization][action] ,
principalId=json[identity][authorization][evidence][principalId] ,
principalType=json[identity][authorization][evidence][principalType] ,
name=coalesce(json[identity][claims][name],json[identity][claims][xms_mirid]),
time= json[time],
title=json[properties][title],
details=json[properties][details],
current=json[properties][currentHealthStatus],
previous=json[properties][previoustHealthStatus],
type=json[properties][type],
cause=json[properties][cause]
| filter  category== "ResourceHealth"
| fieldsKeep  time, title, details, type, cause, azure.resource.name
| sort  time desc

Notebooks cloud observability sandbox

Utilizing the Dynatrace platform to observe your Azure subscriptions allows you to address compliance requirements beyond the standard; for example, the retention period for Azure activity logs is 90 days. With Dynatrace, you can configure up to a 10-year retention period in Dynatrace Grail (for full details, see data retention periods).

Here are further examples that show the power of DQL and Grail for Azure analysis:

Count the number of VMs per Azure Subscription

fetch dt.entity.azure_vm
| fieldsAdd  azure_subscription = accessible_by[dt.entity.azure_subscription][0], azure_region = belongs_to[dt.entity.azure_region]
| lookup [fetch dt.entity.azure_subscription | fieldsAdd name = entity.name ,  uuid = azureSubscriptionUuid], sourceField:azure_subscription, lookupField:id, prefix:"azure.subscription."
| lookup [fetch dt.entity.azure_region | fieldsAdd name = entity.name], sourceField:azure_region, lookupField:id, prefix:"azure.region."
| summarize  count=count(), by: azure.subscription.name

Count the number of VMs per Azure region

fetch dt.entity.azure_vm
| fieldsAdd  azure_subscription = accessible_by[dt.entity.azure_subscription][0], azure_region = belongs_to[dt.entity.azure_region]
| lookup [fetch dt.entity.azure_subscription | fieldsAdd name = entity.name ,  uuid = azureSubscriptionUuid], sourceField:azure_subscription, lookupField:id, prefix:"azure.subscription."
| lookup [fetch dt.entity.azure_region | fieldsAdd name = entity.name], sourceField:azure_region, lookupField:id, prefix:"azure.region."
| summarize  count=count(), by: azure.region.name

Identify the VMs with top CPU usage

timeseries max= max(dt.cloud.azure.vm.cpu_usage),      
 by:  {dt.entity.azure_vm}
| fieldsAdd maxCpu = arrayMax(max)
| sort maxCpu desc
| lookup [fetch dt.entity.azure_vm | fieldsAdd name = entity.name, azure_subscription=accessible_by[dt.entity.azure_subscription][0],  azure_region=belongs_to[dt.entity.azure_region] ], sourceField:`dt.entity.azure_vm`, lookupField:id, prefix:"azure.vm."
| lookup [fetch dt.entity.azure_subscription | fieldsAdd name = entity.name ,  uuid = azureSubscriptionUuid], sourceField:azure.vm.azure_subscription, lookupField:id, prefix:"azure.subscription."
| lookup [fetch dt.entity.azure_region | fieldsAdd name = entity.name], sourceField:azure.vm.azure_region, lookupField:id, prefix:"azure.region."
| fieldsKeep   maxCpu,azure.vm.name, azure.subscription.name, azure.region.name
| limit 20

Get the latest 10 log lines from Azure

fetch logs | filter cloud.provider == "Azure" or cloud.provider == "azure" | sort  timestamp desc  | limit 10

Upgrade to Dynatrace for Azure

If you’re an existing Dynatrace customer, please contact us to learn how to upgrade to Dynatrace on Azure. An overview of how to upgrade to Dynatrace is available in our guide, Upgrade to Dynatrace SaaS.

Get started with the AI-powered Dynatrace platform and Azure Native Dynatrace Service

Visit the Azure marketplace to start a trial of the Dynatrace platform on Azure and the Azure Native Dynatrace Service.

Please see Data Security Controls in Dynatrace documentation for an overview of the Azure regions currently supported by Dynatrace.

Check out our website to learn more about the Dynatrace AI-powered observability platform.

The post Discover the power of the Dynatrace platform on Azure with Azure Native Dynatrace Service appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/azure-with-azure-native-dynatrace-service/feed/ 0
Introducing Dynatrace built-in data observability on Davis AI and Grail https://www.dynatrace.com/news/blog/introducing-dynatrace-built-in-data-observability-on-davis-ai-and-grail/ https://www.dynatrace.com/news/blog/introducing-dynatrace-built-in-data-observability-on-davis-ai-and-grail/#respond Wed, 31 Jan 2024 17:00:22 +0000 https://www.dynatrace.com/news/?p=61558 Database observability graphic

“Great! I have ingested important custom data into Dynatrace, critical to running my applications and making accurate business decisions… but can I trust the accuracy and reliability?” Welcome to the world of data observability. The Dynatrace open platform is well-positioned to take advantage of the exponential increase in data generation. However, coupled with the increase […]

The post Introducing Dynatrace built-in data observability on Davis AI and Grail appeared first on Dynatrace news.

]]>
Database observability graphic

“Great! I have ingested important custom data into Dynatrace, critical to running my applications and making accurate business decisions… but can I trust the accuracy and reliability?”

Welcome to the world of data observability.

The Dynatrace open platform is well-positioned to take advantage of the exponential increase in data generation. However, coupled with the increase of external data sources that can now be ingested, there are new challenges in data management that need to be addressed.

 “Every year, poor data quality costs organizations an average $12.9 million”
– Gartner

Data observability is a practice that helps organizations understand the full lifecycle of data, from ingestion to storage and usage, to ensure data health and reliability. Data observability involves monitoring and managing the internal state of data systems to gain insight into the data pipeline, understand how data evolves, and identify any issues that could compromise data integrity or reliability. At its core, data observability is about ensuring the availability, reliability, and quality of data.

Data observability is crucial to analytics and automation, as business decisions and actions depend on data quality. In the age of AI, data observability has become foundational and complementary to AI observability, data quality being essential for training and testing AI models.

Dynatrace now addresses many of the issues customers experience around the health, quality, freshness, and general usefulness of data that is externally sourced into Dynatrace Grail™, allowing them to make better-informed decisions and optimize their efforts for digital transformation and data-driven operations.

The rise of data observability in DevOps

Data forms the foundation of decision-making processes in companies across the globe. Data is the foundation upon which strategies are built, directions are chosen, and innovations are pursued. Consequently, the importance of continuously observing data quality, and ensuring its reliability, is paramount. Surveys from our recent Automation Pulse Report underscore this sentiment: 57% of C-level executives say the absence of data observability and data flow analysis makes it difficult to drive automation in a compliant way. This not only underscores the universal significance of data, it also hints at its pivotal role within DevOps. For DevOps teams that inform deployment strategies, optimize processes, and drive continuous improvement, the integrity and timeliness of data are of significant importance.

As organizations scale and accelerate their digital transformation journeys, a major hurdle to proper DevOps adoption is the trustworthiness of the massive volume of data coming from various sources, much of which goes into data silos such as log management tools, SIEM solutions, and others.

The rise of data observability needs is where Dynatrace capabilities around Grail, analytics, and Davis® AI are in an outstanding and unmatched position to deliver the currently missing value to the market: a leading and single solution for all data observability analytics needs. This reduces the demand for further data flow analysis tools and clears any hurdles to making data useable for DevOps automation use cases.

Davis AI, Grail, and data observability

By grouping common data observability issues into industry-standard pillars, we can provide tangible examples and showcase current capabilities. The five pillars we focus on are freshness, volume, distribution, schema, and lineage.

Freshness: Timeliness of data

In an ideal ecosystem, actionable data should be as recent as possible, supported by learnings from accurate, historical data. Observing the freshness of data helps to ensure that decisions are based on the most recent and relevant information.

Scenario: Due to an undetected configuration issue, a flight status system from a popular airline had been buffering data for the last two hours before sending it on in one batch. Downstream dashboards and system automations were using outdated data, leading to incorrect statuses of flights in reports.

Solution: After setting up data ingestion into Grail, Dynatrace Query Language (DQL) is used to add a freshness field (Figure 1) which is calculated from the delta between when the signal was written and when it was ingested. This freshness measurement can then be used by out-of-the-box Dynatrace anomaly detection to actively alert on abnormal changes within the data ingest latency to ensure the expected freshness of all the data records. Furthermore, the new Alert on missing data feature in the Anomaly Detector panel can be used to trigger notifications when data is not coming in as expected after being baselined.

Value: The possibility of alerting on data freshness issues, based on a learned baseline through Davis AI, allows for faster time-to-detect where there are seemingly no infrastructure issues. Normally this would have left an issue undetected for much longer, providing a false sense of security, eventually leading to a much bigger customer and monetary impact for the organization.

Use of Dynatrace Notebook to track when a flight status table was last updated.
Figure 1. Use of Dynatrace Notebook to track when a flight status table was last updated.

Volume: Quantity of data generated or processed within a given timeframe

Unexpected increases or drops in the volume of data are often a good indication of an undetected issue.

Scenario: For many B2B SaaS companies, the number of reported customers is an important metric. It heavily influences downstream reports, and dashboards, shaping decisions from daily operations to strategic monthly reviews. In this scenario, a manually triggered run of a production pipeline had the unintended consequence of duplicating the reported customer metric. If left unchecked, this misrepresentation of a single KPI could lead to misguided decision-making processes through multiple layers of the organization.

Solution: Like the freshness example, Dynatrace can monitor the record count over time. Once a DQL query has been set up, it can be used in an automation workflow (Figure 2) where scheduling, prediction, comparison to actual value, and, finally, alerting are all taken care of to enable a fully flexible way to detect anomalies in data volume.

Value: KPIs and metrics such as the number of reported customers are central to an organization’s business and strategic processes. Any issues here will result in a loss of trust in the data, and, if left undetected, they will eventually lead to monetary impact, including loss of reputation for an organization.

Using Dynatrace Workflows to alert on data volume anomalies
Figure 2 Using Dynatrace Workflows to alert on data volume anomalies

Distribution: The statistical spread or ranges of data

The distribution of data is essential in identifying patterns, outliers, or anomalies in the data. Deviation from the expected distribution can signal an issue in data collection or processing.

Scenario: A financial institution processes millions of transactions daily, ranging from credit card purchases and mortgage payments to interbank transfers and ATM withdrawals. An erroneous change in the database system leads to a subset of the data being categorized incorrectly. After several days, the fraud detection system starts triggering on a frequent basis, and liquidity management dashboards begin showing questionable values.

Solution: Baselining and raising alerts on anomalies are core capabilities of Davis AI. After setting up ingestion for the data that you want to monitor, it’s simple to use Dynatrace full AI capabilities to observe and alert on any anomalies in the data. In the example above, ingesting the number of transactions as business events, anomaly detection could be based on this to proactively alert and trigger mitigation activities.

Value: While variations are expected in financial trends, anomalies should be auto-detected, and manual detection should not be relied on. Earlier detection of these issues will keep the fallout as low as possible.

Schema: Structure and relationships of data between entities

Observing the schema can help identify and flag unanticipated changes, such as the addition of new fields or deletion of existing fields.

Scenario: An externally connected database system made an update that inadvertently dropped the account_id column in the customers table. The automated data pipeline propagated these changes, leading to downstream reports, dashboards, and applications breaking as the previous field reference is now missing.

Solution: Using the DQL FieldsSummary command, we can keep track of the number of distinct field keys within a given family of data records. Once confirmed in a notebook, the number of field keys can be used in an automated workflow to continuously monitor the count and write it back to a new metric (Figure 3). Once the new metric is established, out-of-the-box Dynatrace anomaly detection can be used to alert on either a static threshold or a learned baseline.

Value: Observing incoming data Schemas, and thus placing expectations on what the external data should look like and must contain, allows for pro-active alerting and mitigation of issues long before they can lead to widespread business impact such as broken reports, dashboards, or further analytics on top of the data.

Keeping track of the field count in a new metric (data.observability.fields) using Workflows and Typescript.
Figure 3. Keeping track of the field count in a new metric (data.observability.fields) using Workflows and Typescript.

Lineage: Journey of data through a system

Data lineage provides insights into where the data came from (upstream) and what is impacted (downstream). It plays a crucial role in root cause analysis as well as informing impacted systems about an issue as quickly as possible.

Scenario: The hourly_consumption table was deprecated and removed by an overzealous database administrator as there were no known downstream consumers of this data, breaking a monthly integration check used for consumption reporting for shareholders.

Solution: In the future, Dynatrace Smartscape® could be used, which already builds a dependency graph, to enable a data lineage view. This would enable faster root cause analysis of any data-related problems, as well as allow for easy notification of downstream consumers who would be impacted.

Value: A proper understanding of the source of the data, as well as where it is used, helps drive down time-to-alert and time-to-repair. Time-to-alert is achieved by quickly and automatically alerting those who are impacted by a data issue by quickly understanding downstream consumers of the data, while time-to-repair informs on the source of where the data originated from, to quickly drill down into those systems.

Data availability: A prerequisite

You could implement the most contemporary, accurate, and useful data observability solution possible, but what good will it be if all the data simply does not arrive as expected? Broken pipelines or missing data sources would mean that there is simply no data to observe and that data may never arrive, forever lost.

A truly valuable data observability solution should be able to alert on data issues as early in the process as possible. This requires monitoring of the upstream infrastructure, applications, or platform supporting those data streams. This is where the power of Dynatrace end-to-end observability comes into play. Dynatrace can leverage existing Infrastructure Monitoring and Application Observability solutions to surface problems that can affect later data observability workstreams—long before a traditional data observability solution would pick up the issue.

Leverage the power of Dynatrace and Davis AI—now and into the Future

Anomaly Detection

Anomaly detection is grounded in the idea of baselining typical patterns of ingested data, designed to alert where a change or deviation from the norm is observed. These patterns typically go beyond simple flat or trend lines, often exhibiting complex seasonal behaviors, such as business hours or weekly patterns related to the industry. Dynatrace is particularly strong in this area: Davis predictive AI has been enriched over the years with a set of advanced machine learning (ML) algorithms optimized for time-series observability datasets to cope with these challenges. Davis AI anomaly detection, leveraging these ML algorithms, can already be used on the results of DQL queries. (Embedding ML algorithms into DQL as functions is on the Dynatrace platform roadmap.)

Considering the examples and solutions provided above, anomaly detection plays a pivotal role in numerous data observability use cases and can be harnessed to effectively address these challenges.

Triage and resolution of a data incident

Triaging requires an ability to identify the root cause of a data incident, which is particularly challenging as an organization scales up the volume and speed of data ingest typical of an enterprise environment. It’s easy to see how Davis causal AI problem detection could be extended in the future to identify root-cause data observability issues.

Depending on the incident, there might be different paths to resolution. One acceptable path could be full auto-remediation, whereby Dynatrace AutomationEngine could be triggered, scripts executed, permissions granted, security checked, and data corrected. A second path might require Jira tickets to be created and human intervention through an approval process. A data problem alert could be used as the event allowing for multiple methods to notify the correct data owners, stewards, governors, or data teams.

An incident requires not only resolution but also understanding and alerting upstream data providers and downstream data subscribers to the potential impact. Dynatrace is strong on the observability of data pipelines ingesting data into Grail and consumers of Grail data, although this is an area that will be enhanced and improved in the future product roadmap.

Prevention of future incidents

Not all data quality incidents can be prevented, especially because ELT/ETL data pipelines typically tend to grow over time and span many different heterogeneous collectors that have different ownerships. There are, however, mitigation techniques you can use, for example:

  • Health tracking of key datasets or streams over time—alerting on anomalies
  • Monitoring standard query results and changes over time
  • Well-designed, data-focused dashboards for monitoring
  • Auto remediation where appropriate with built-in audit logging
  • Forensic abilities for ad-hoc data analysis

Summary

Dynatrace is uniquely positioned to provide even more value by extending our world-class observability platform into the data observability realm. To achieve this, we leverage Infrastructure Monitoring and Application Observability for early warnings on data pipeline issues and use DQL, Workflows, and Grail for data observability—all enabled by our best-in-class Davis AI engine.

Ensuring the quality and reliability of underlying data is more crucial than ever now that many organizations are deploying Generative AI models. Data observability is becoming a mandatory part of business analytics, automation, and AI. Davis AI and data observability together uniquely ensure the quality and reliability of data at the level of hypermodal AI—predictive, causal, and generative.

You can now monitor sources and incoming data pipelines for freshness, volume, distribution, lineage, and availability issues early on without added noise and in a central location, the Dynatrace platform. This gives your teams additional confidence over data quality, saves time, prevents inaccurate analyses and automation outcomes, leads to more trustworthy AI models, and supports efforts to consolidate or reduce the number of IT tools they rely on.

Ready to get started with Dynatrace data observability? For complete details, best practices, and detailed use cases, see Dynatrace data observability documentation.

The post Introducing Dynatrace built-in data observability on Davis AI and Grail appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/introducing-dynatrace-built-in-data-observability-on-davis-ai-and-grail/feed/ 0
Speed up your security investigations with DPL Architect https://www.dynatrace.com/news/blog/speed-up-your-security-investigations-with-dpl-architect/ https://www.dynatrace.com/news/blog/speed-up-your-security-investigations-with-dpl-architect/#respond Thu, 12 Oct 2023 18:33:12 +0000 https://www.dynatrace.com/news/?p=60004 DPL Architect

Grail™, the Dynatrace causational data lakehouse, offers you instant access to any kind of data, enabling anyone to get answers within seconds. As all historical data is immediately available in Grail, no data rehydration is needed, even for analysis of suspicious timeframes that are lengthy or ancient.

The post Speed up your security investigations with DPL Architect appeared first on Dynatrace news.

]]>
DPL Architect

To help you raise the quality of your investigation results, Dynatrace offers an easy way of structuring data using DPL Architect. This tool lets you quickly extract typed fields from unstructured text (such as log entries) using the Dynatrace Pattern Language (DPL), enabling you to extract timestamps, determine status codes, identify IP addresses, or work with real JSON objects. This allows you to answer even the most complex questions with ultimate precision. The best thing: the whole process is performed on read when the query is executed, which means you have full flexibility and don’t need to define a structure when ingesting data.

>> Scroll down to see Dynatrace DPL Architect in action (24-second video)

Investigating log data with the help of DQL

Let’s look at a practical example. The simplest log analysis use cases can be solved in seconds by applying a simple filter to the query results. If a CISO asks, “Have we seen the IP address 40.30.20.1 in our logs within the last year?” a simple DQL query seems to suffice:

fetch logs, from: -365d
| filter contains(content, “40.30.20.1”)

This query returns all the records that contain the string value 40.30.20.1.

In the next step, adding DQL aggregation functions enables you to answer more complex questions like “How many times in an hour has this IP address visited our website within the last two weeks?” Again, a simple DQL query helps you out:

fetch logs, from: -14d
| filter contains(content, “40.30.20.1”)
| summarize by: bin(timestamp, 1h), count()

However, this kind of simple filtering using a sub-string search is prone to errors and is not precise enough. The mentioned filters return all records that contain 40.30.20.1, but are not necessarily from the source IP portion of the log format. The string 40.30.20.1 might also appear elsewhere, for example, in query parameters, and a simple string filter will return all records containing this search term:

19.31.99.1 - - [28/Aug/2023:10:27:10 -0300] "GET /index.php?ip=40.30.20.1 HTTP/1.1" 200 3395
40.30.20.109 - - [28/Aug/2023:10:22:04 -0300] "GET / HTTP/1.1" 200 2216

The first log record is matched because of its query parameter—something we don’t care about in our case. The second log record came up because the source IP contains the IP address—this is what we’re really looking for. In a nutshell, if you’re querying terabytes of logs and wish to drill down to specific records, filtering strings from plaintext log content is not enough.

The issue is that questions from CISOs aren’t usually so trivial when it comes to security use cases. Or, the log format where the answers should be looked for is more complex (for example, AWS CloudTrail logs). As a consequence, we need to search structured data to avoid mistakes. To minimize false positives, increase the precision of queries, and get the maximum out of DQL, fields can be extracted from record content.

Extracting patterns using DPL

Dynatrace Pattern Language (DPL) is a pattern language that allows you to describe a schema using matchers, where a matcher is a mini-pattern that matches a certain type of data. These matchers can be used to extract new fields from the schemaless data in Grail to enable more precise log filtering. Consider the same Apache access log example from above:

19.31.99.1 - - [28/Aug/2023:10:27:10 -0300] "GET /index.php?ip=40.30.20.1 HTTP/1.1" 200 3395

To extract the client IP addresses from the beginning of the log, a simple DPL pattern can be used:

IPADDR:client_ip

The IPADDR matches both versions of IP addresses (IPv4 addresses in dot-decimal notation and IPv6 addresses in hextet notation), leaving all other IP-like strings unmatched. For example, if the log record contains a dot-decimal number that is NOT an IPv4 address (999.999.999.999), it won’t be matched.

DPL patterns can be applied in DQL using the parse command. The simplest way to get started with DPL is to use Dynatrace DPL Architect.

DPL Architect to the rescue

DPL Architect is a handy tool, accessible through the Notebooks app, which supports you in quickly extracting fields from records. It helps create patterns, provides instant feedback, and allows you to save and reuse DPL patterns, for faster access to data analytics use cases.

To open the DQL Architect, you have to execute a DQL query, select the content, and choose Extract fields.

Video thumbnail

Figure 1: Extract fields in Notebooks using DPL Architect (24-second video)

Starting with preset patterns

The simplest way to extract data is using one of the ready-to-use preset patterns available for the most popular technologies, such as AWS, Microsoft, or GCP. Start exploring AWS VPC Flow logs or analyzing Kubernetes audit logs by choosing the patterns from the panel on the left side. You can also customize the list by adding your own individual patterns.

Figure 2: Choose from a list of available patterns
Figure 2: Choose from a list of available patterns

Developing a new pattern using DPL Architect

DPL Architect helps you by creating your own patterns and provides instant feedback. Start typing a DPL expression in the pattern field, and review the matching data (highlighted below) in the preview editor.

Figure 3: Match preview editor
Figure 3: Match preview editor

For accessing the extracted fields in the resultset, it’s necessary to add a name. In the example below, we’re only interested in the client_ip and the related response_code. The pattern matches the whole record, but only two fields are being extracted since they have defined extract names. Any other data between those two fields won’t be visible in the results (however, it is very easy to add them later if necessary).

Figure 4: Match preview editor
Figure 4: Match preview editor
Figure 5: Review extracted fields in Results tab
Figure 5: Review extracted fields in the Results tab

Remember, we started our journey with the DPL Architect by selecting query results in a notebook. The colored bar at the top of the DPL Architect shows how many records from your query result match the current DPL pattern. This enables you to modify the DPL pattern to ensure it matches all records in the resultset.

Figure 6: Add records to the Match preview that don't match the current pattern.
Figure 6: Add records to the Match preview that don’t match the current pattern.

Records that don’t match the current pattern can easily be added to the Match preview panel by selecting Add to preview.

Figure 7: Add unmatching records to Match preview dataset
Figure 7: Add unmatching records to the Match preview dataset

After changing the DPL pattern, any previously unmatched records will be highlighted and the progress bar on the top informs you that 100% of the records in the Notebooks resultset are matched. It’s now time to insert our pattern into the DQL query by selecting Insert Pattern.

Figure 8: Review the result and insert the created pattern
Figure 8: Review the result and insert the created pattern

Precise extraction of fields from complex data

DPL also provides matchers for more complex data structures, like key-value pairs, structures, and JSON objects, that enable you to parse individual sub-elements from the whole object. Consider the following log record:

1693230219 230.4.130.168 C "GET / HTTP/1.1" "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:105.0) Gecko/20100101 Firefox/105.0" 186 {"username":"james","result":0 }

Imagine your CISO comes with a request: “Give me the list of all IP addresses where the user ‘james’ has successfully logged in from.” Extracting these three fields and building a DQL query is as easy as pie:

LD IPADDR:src_ip LD JSON{STRING+:username, INT:result}(flat=true)

The DPL will extract three fields: an IP address from the record, a string from the JSON object element username, and an integer from the JSON element result.

Figure 9: Extracted fields in the Results tab.
Figure 9: Extracted fields in the Results tab.

Summarizing the results based on these fields can be done with the following DQL:

fetch logs
| parse content, "LD IPADDR:src_ip LD JSON{STRING+:username, INT:result}(flat=true)"
| filter username == "james" and result == 0
| summarize by: src_ip, count()

Imagine that you don’t have DPL available and need to filter all log records based on only searching for string values: the result would contain a lot of false positives (which only add “noise” to time-critical investigations). This is especially true with more complex log records containing nested JSON objects. The above example was oversimplified intentionally, but considering complex log records like AWS CloudTrail log records, having precise access to data in a specific node of an object makes a huge difference.

Summary

When performing security investigations or threat-hunting activities, it’s important to have precision in place to get reliable results. Historical data needs to be available, and access to object details is required for precise answers. DPL Architect enables you to quickly create DPL patterns, speeding up the investigation flow and delivering faster results. With the possibility of using Dynatrace-provided patterns for selected technology stacks, investigators can deliver answers even faster!

Check out the following video, where Andreas Grabner and I teamed up for a new episode of Dynatrace Observability Clinic. In this video, we dig deeper into the topic of extracting data via DPL, including a live demonstration of DPL Architect.

The post Speed up your security investigations with DPL Architect appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/speed-up-your-security-investigations-with-dpl-architect/feed/ 0
TTP-based threat hunting with Dynatrace Security Analytics and Falco Alerts solves alert noise https://www.dynatrace.com/news/blog/ttp-based-threat-hunting-solves-alert-noise/ https://www.dynatrace.com/news/blog/ttp-based-threat-hunting-solves-alert-noise/#respond Wed, 09 Aug 2023 11:58:55 +0000 https://www.dynatrace.com/news/?p=59078 TTP-based threat hunting with Dynatrace Grail and Falco for Security Analytics

Today’s security analysts have no easy job. Not only are cyberattacks increasing, but they’re also becoming more sophisticated, with tools such as WormGPT putting generative AI technology in the hands of attackers. As a result, analysts are turning to AI and TTP-based threat-hunting techniques to uncover how attackers are trying to exploit their environments. While AIOps with generative AI […]

The post TTP-based threat hunting with Dynatrace Security Analytics and Falco Alerts solves alert noise appeared first on Dynatrace news.

]]>
TTP-based threat hunting with Dynatrace Grail and Falco for Security Analytics

Today’s security analysts have no easy job. Not only are cyberattacks increasing, but they’re also becoming more sophisticated, with tools such as WormGPT putting generative AI technology in the hands of attackers. As a result, analysts are turning to AI and TTP-based threat-hunting techniques to uncover how attackers are trying to exploit their environments.

While AIOps with generative AI will certainly empower security teams to mitigate threats faster and with greater precision, attackers will just as certainly utilize the same technology to create novel malware, more convincing phishing campaigns, and uncover high-risk zero-day vulnerabilities quicker.

Not only that, teams struggle to correlate events and alerts from a wide range of security tools, need to put them into context, and infer their risk for the business. But the industry as a whole is still hampered by a ubiquitous tool sprawl to achieve that critical mission under a barrage of alert noise.

In this blog post, we’ll use Dynatrace Security Analytics to go threat hunting, bringing together logs, traces, metrics, and, crucially, threat alerts. We use the power of DQL on Grail to derive high-level attacker tactics, techniques, and procedures (TTPs), which are much easier to interpret and act upon.

TTP-based threat hunting: Tactics, techniques, procedures

At Dynatrace, we don’t want to bombard you with alert noise and uncorrelated warnings. Instead, we want to focus on detecting and stopping attacks before they happen: In your applications, in context, at the exact line of code that is vulnerable and in use. But even when an attack happens, Dynatrace detects and blocks them in real time while providing you with rich technical details on the concrete attack procedure. Procedures describe the specific technical details that an adversary used to carry out an attack, for example, what script they ran to exploit a weakness.

TTP-based threat hunting with Dynatrace: tactic, technique, procedure

When investigating advanced cyberattacks, it’s helpful to map attack procedures to attack techniques. Techniques describe the tactical goal an adversary is pursuing by executing a specific procedure. One of the most critical attack techniques within the MITRE ATT&CK® knowledge base of adversary tactics and techniques — and one example of what Dynatrace can prevent, detect, and block in real-time — is attack technique T1190, “Exploit Public-Facing Application”.

Public-facing applications can be an initial access vector an attacker could exploit to gain entry into a system. Attack tactics describe why an attacker performs an action, for example, to get that first foothold into your network.

Thinking in terms of tactics, techniques, and procedures (TTPs) brings many benefits. For example, security analysts can more easily stitch together advanced cyberattacks on an abstract level. Likewise, operation specialists can prioritize their efforts on monitoring the highest-risk tactics, and executives can better communicate the business risk.

Threat hunting and analyzing threat alerts with Dynatrace Security Analytics and Grail

Dynatrace offers Runtime Application Protection to detect a wide range of injection attacks in your applications. However, our customers often want to augment the data Dynatrace provides with data from third-party tools. Customers also want to carry out their own analysis tailored to specific use cases and forensic needs.

Dynatrace Grail is a data lakehouse that provides context-rich analytics capabilities for observability, security, and business data. You may also ingest additional data into our unified intelligence platform: One popular choice to gather fine-grained security data is Falco. Falco is an open-source, cloud-native security tool that utilizes the Linux kernel technology eBPF, to generate fine-grained networking, security, and observability events.

In the following sections, we demo the following:

  1. Introduce Unguard, our insecure cloud-native microservices demo application.
  2. Install Falco in AWS EKS to gather security-relevant events from all the happenings in Unguard.
  3. Ingest those Falco events into Dynatrace Grail using falcosidekick.
  4. Query Falco events in Dynatrace Grail, map them to TTPs, and conduct structural multi-step attack detection.

In other words, we find attacks that are composed of multiple steps by using TTPs and Dyntrace Smartscape for DQL in a way that eliminates alert noise.

First, Dynatrace OneAgent will automatically monitor and trace our infrastructure and communicate with Dynatrace. Second, we will enrich our data in the Grail data lakehouse by also ingesting Falco events using falcosidekick.

threat hunting architecture with Dynatrace and Falco

Setting up our TTP-based threat-hunting demo environment

Before we start threat hunting, we’ll first walk through how to set up the demo environment.

Introducing Unguard, our insecure cloud-native demo app

As our playground, we introduce Unguard, a microblogging demo application that embodies the challenges of modern cloud-native environments. It consists of eight services, written in at least four different languages, with countless vulnerabilities and misconfigurations. To keep it real, we have a load generator that creates benign traffic. It also generates OpenTelemetry traces.

TTP-based threat hunting: Unguard demo application

Unguard was first introduced at DEFCON 31 by our colleagues Simon Ammer and Christoph Wedenig.

For the demonstration in this blog post, we want to deploy Unguard in AWS EKS and hunt for attacks within that environment. You can easily play around with Unguard by installing its Helm chart:

helm install unguard \ 
  oci://ghcr.io/dynatrace-oss/unguard/chart/unguard \ 
  --wait --namespace unguard --create-namespace

(Please read the Unguard README for detailed and up-to-date instructions)

This demo assumes your Kubernetes cluster is already monitored by Dynatrace. For instructions, see Set up Dynatrace on Kubernetes.

Deploy Falco and falcosidekick in AWS EKS

You can install Falco in various ways. For this demo, we installed it with the Helm chart in our AWS EKS cluster:

helm repo add falcosecurity https://falcosecurity.github.io/charts 
helm repo update 
helm install falco falcosecurity/falco --namespace falco --create-namespace

(See the Falco README for detailed and up-to-date instructions)

Next, we set up falcosidekick, which is a daemon that forwards Falco events to many possible outputs. We’re proud to announce that, with Falco version 2.29, currently in pre-release, you can now also use Dynatrace as an output.

You can use this minimal values.yaml configuration file for the Helm chart:

# values.yaml 
 
falcosidekick: 
  enabled: true 
  image: 
    tag: 2.29.0-rc.1 
  config: 
    # as of 2023-08-02, this feature is still a pre-release so we 
    # have to manually override the environment variables for now 
    extraEnv: 
      - name: DYNATRACE_APITOKEN 
        value: dt0c01.EXAMPLE_TOKEN_REPLACE_THE_ENTIRE_STRING 
      - name: DYNATRACE_APIURL 
        value: https://ENVIRONMENTID.live.dynatrace.com/api

(Please read the Helm chart README for detailed and up-to-date instructions)

We insert the apitoken we generated within Dynatrace and grant the token the scope logs.ingest. See the topic Dynatrace API – Tokens and authentication to learn more about creating tokens. As the apiurl, use the following:

Dynatrace SaaS:

https://ENVIRONMENTID.live.dynatrace.com/api

Dynatrace Managed:

https://YOURDOMAIN/e/ENVIRONMENTID/api

See the topic Environment ID to learn more about environment IDs.

Finally, we update the Falco Helm chart with this new configuration:

helm upgrade falco falcosecurity/falco -f values.yaml

If everything worked out well (check the pod logs otherwise), we are now able to successfully query Falco events with DQL. To verify, we open a new Notebook and see how Dynatrace automatically infers the fields from our events already:

fetch logs, from:now() - 5m 
| filter (event.provider == "Falco")

TTP-based threat hunting: Dynatrace automatically infers the fields from our events

Observing TTPs using Dynatrace Security Analytics

For the sake of this demonstration, our internal red team unleashed a novel attack on our Unguard application. The attack lit up our Falco deployment with more than 100,000 events in 24 hours, more than 3,000 of them critical.

As security analysts, we know we can’t find sophisticated attacks by manually scrolling through thousands of audit logs and events. We need automation, full contextual knowledge of our infrastructure, and very often, domain-specific expertise from security analysts.

To get an initial overview, we can use DQL on Grail to visualize what MITRE techniques Falco observed in our infrastructure over the past 72 hours. We can summarize events using mitre.tactic or mitre.technique. These fields exist on many Falco alerts and are automatically ingested by the Dynatrace output of falcosidekick. We can explore the distribution of techniques with this query:

fetch logs, from:now() - 72h 
| filter event.provider == "Falco" and isNotNull(mitre.technique) 
| filterOut in(mitre.technique, {"T1548.001", "T1083", "T1565", "T1055.008"}) 
| summarize event_count = count(), by:{mitre.technique}

Threat hunting technique chart

In this example, we also observe that we can attribute most events to the following MITRE techniques:

After manually investigating these alerts, however, we conclude they’re noisy false positives. Some of our applications were treating environment variables in an insecure way or communicating with the Kubernetes API server with improperly configured service accounts. Therefore, we filtered them out with DQL.

Observability and context: Attributing reconnaissance activity to TTPs using distributed traces

So far in our TTP-based threat hunting, we’ve utilized Dyntrace Security Analytics to visualize ingested alerts from third-party tools.

But truly magical things arise when we combine this with the rich and high-quality observability data that our customers have valued since the beginning of Dynatrace. Using observability data, we can close an important security-relevant gap. Attackers often probe systems using automated scanning tools. Their many access attempts leave behind a lot of traces. Dynatrace PurePath is one of the core platform technologies that captures and analyzes those distributed traces across an entire infrastructure.

Attack sub-technique T1595.003 “Active Scanning: Wordlist Scanning” describes how attackers use scanners to learn about the many endpoints an application might expose to the internet. Typically, they use large lists with well-known path names, where many of them could be potentially vulnerable. Such wordlists often contain common path names, such as wp-admin, .git, or .htaccess. The following query looks for five indicative files and expresses how many of them match with the new recon.confidence field we set up to track wordlist-based scanners. The more matches, the more confident we can be that these requests came from a wordlist-based scanner:

fetch spans, from:now() - 72h 
| filter in(http.target, {"/wp-admin", "/.git", "/.htaccess", "/.ssh", "/cgi"}) 
| summarize { 
    recon.confidence = countDistinct(http.target) / 5, 
    recon.first_seen = min(timestamp), 
    recon.last_seen = max(timestamp) 
  }, by:{host.name, k8s.container.name, k8s.namespace.name, k8s.pod.name} 
| fieldsAdd mitre.technique = "T1595.003", mitre.tactic = "mitre_reconnaissance" 
| filter recon.confidence > 0.5

fetch spans results

Indeed, it did find some reconnaissance attempts! This query scanned 2.5 million spans in less than 50 ms and reduced them to three comprehensible TTP records. With a conventional database, a query that scans millions of records would take many seconds to complete and require that we structure our queries up front. But Grail completes the search across millions of records in milliseconds, and automatically parses and infers the structure for us so we can just start writing queries directly.

We now know that the Envoy proxy in our environment most likely got scanned by an attacker. Let us bring all the bits and pieces together in the next sequence.

Structural multi-step attack detection with Dynatrace Security Analytics

Attackers typically perform many small steps to achieve their mission. Security experts like to think in terms of so-called kill chains, which describe the many stages of an attack. When you work with TTPs, the attack tactics represent those stages. While the MITRE ATT&CK® knowledge base describes as many as 14 tactics, we can distill this into three broad categories:

  1. Land – First, hackers investigate their target, looking for an initial way to gain access and establish a foothold in your system.
  2. Expand – Then, hackers typically try to escalate their privileges and move laterally within your system, compromising neighboring hosts.
  3. Execute – Finally, they find their target and execute their mission.

Next in our demonstration of TTP-based threat hunting with Dynatrace Security Analytics, we’re going to show you a simple but effective strategy that can uncover such advanced attacks: Structural multi-step attack detection. This strategy is structural since it utilizes Smartscape for DQL to take the topological relationship of events into account when hunting for attacks that are composed of multiple steps.

The following DQL query looks for the filtered Falco alerts, and for each Kubernetes pod, it records how many distinct tactics and techniques we just observed. This way, we’re not just looking at whatever pod was the noisiest, but instead, which pod generated alerts from the most tactics and techniques. The more tactics and techniques, the higher the chances that an attacker carried out a full kill chain on that pod.

fetch logs, from:now() - 72h 
| filter (event.provider == "Falco") 
| filterOut in(mitre.technique, {"T1548.001", "T1083", "T1565", "T1055.008"}) 
| summarize { 
    num_tactics = countDistinct(mitre.tactic), 
    num_techniques = countDistinct(mitre.technique), 
    tactics = collectDistinct(mitre.tactic), 
    techniques = collectDistinct(mitre.technique) 
  }, by:{k8s.pod.name} 
| filter num_tactics > 1 and num_techniques > 1 
| sort num_tactics desc, num_techniques desc

threat hunting: attack detection query

This result is highly interesting and confirms our previous suspicion. There is one instance of the Envoy proxy that captured alerts for the following techniques:

Further, the Dynatrace spans we looked at in the previous sequence that explored wordlist scanning indicated TA0043 “Reconnaissance”. This query just scanned through more than 23 million records in 300 ms, providing us with an abstract description of a full kill chain.

But we don’t stop here. We can drill down and observe the individual steps our attacker has taken:

fetch logs, from:now() - 72h 
| filter (event.provider == "Falco") 
| filterOut in(mitre.technique, {"T1548.001", "T1083", "T1565", "T1055.008"}) 
| filter k8s.pod.name == "unguard-envoy-proxy-666464f76d-5p26f" 
| fields timestamp, event.name, mitre.technique, content.output_fields.proc.cmdline 
| sort timestamp asc

TTP threat hunting attack chain query

In the result, we see the records that explain the attack procedure in detail. The above screenshot shows only an excerpt of all 47 records. Here’s what we learned about our attacker’s steps:

  • Scanned our Envoy proxy with a well-known wordlist, as we learned by mining the traces for wordlist entries.
  • Launched a Perl-based reverse shell on Envoy, indicated by the perl command opening a socket, giving them full code execution access.
  • Downloaded a couple of binaries like nmap and nc, indicated by the curl command that pulled them from the internet.
  • Scanned our internal network with nmap.
  • Exfiltrated large volumes of data from our Redis database, indicated by the queries to Redis that request all keys with the KEYS * command
  • Tried to cover their tracks by deleting the shell history, indicated by the rm /home/envoy/.bash_history command

Isn’t this a truly elegant way to hunt for attacks?

TTP-based threat hunting with context-rich observability and security analytics

This demonstration shows how modern attack detection strategies become a reality with context-rich security analytics on a unified observability and security platform. With this approach, you can do the following:

  • Utilize observability data to capture security-relevant reconnaissance alerts and map them to TTPs.
  • Enrich the Dynatrace platform with more data of your own, such as ingesting Falco alerts into Grail.
  • Use DQL and Grail to find the needle in the haystack to scan tens of millions of records in milliseconds to identify the chain of only a handful of events that exposed an attacker and the exact methods they used.

For another great demonstration, we recommend reading the blog post Log forensics: Finding malicious activity in multicloud environments with Dynatrace Grail by Liisa Tallinn.

If this blog post made you eager to try out Dyntrace and learn more about Grail, join us for the on-demand webinar, Get to know Dynatrace: Grail edition.

The post TTP-based threat hunting with Dynatrace Security Analytics and Falco Alerts solves alert noise appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ttp-based-threat-hunting-solves-alert-noise/feed/ 0
Automate predictive capacity management with Davis AI for Workflows https://www.dynatrace.com/news/blog/automate-predictive-capacity-management-with-davis-ai-for-workflows/ https://www.dynatrace.com/news/blog/automate-predictive-capacity-management-with-davis-ai-for-workflows/#respond Tue, 11 Jul 2023 20:19:17 +0000 https://www.dynatrace.com/news/?p=58574 predictive capacity management

>> Scroll down to see predictive capacity management in action (14-second video) Our recent blog post, Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI, showed how Dynatrace Notebooks are used to predict the future behavior of time series data stored in Grail™. This follow-up post introduces Davis® AI for […]

The post Automate predictive capacity management with Davis AI for Workflows appeared first on Dynatrace news.

]]>
predictive capacity management
>> Scroll down to see predictive capacity management in action (14-second video)

Our recent blog post, Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI, showed how Dynatrace Notebooks are used to predict the future behavior of time series data stored in Grail™. This follow-up post introduces Davis® AI for Workflows, showing you how to fully automate prediction and remediation of your future capacity demands. The anticipation of future capacity demands makes it possible to completely avoid critical outages by notifying you days in advance, well before incidents arise.

Predictive capacity management starts within a Dynatrace Notebook, where the operations team explores important capacity indicators, such as the percentage of free disks, as shown below.

Figure 1. Example forecast of remaining disk capacity with upper/lower bounds and an anticipated value.
Figure 1. Example forecast of remaining disk capacity with upper/lower bounds and an anticipated value.

After exploring and selecting the most important capacity indicators for your environment, a workflow triggers forecast reporting at regular intervals. The example workflow below is triggered every Monday at 8:00 AM to provide a capacity report for all the disks that will likely run out of space within the next week.

Figure 2. Over of the predict disk capacity workflow
Figure 2. Predict disk capacity workflow

Define the forecast

The workflow uses the Davis for Workflows action to automatically trigger a forecast for a selected set of disks. The forecast operation is selected within the Davis action, and a DQL query is used to specify the set of disks and the capacity indicator metric that should be predicted. Note that you can use any time series data you can fetch from Grail using DQL within the forecast action.

While this example uses the metric dt.host.disk.free, you can choose any kind of capacity metric, such as host CPU, memory, or network load—you can even extract a metric value from a given log line.

The forecast is trained on a relative timeframe (for example, the last seven days) which is specified in the configured DQL query. The DQL query example below trains forecasting on a relative timeframe of the last seven days:

timeseries avg(dt.host.disk.free), by:{dt.entity.host, dt.entity.disk}, bins: 120, from:now()-7d, to:now()

The configuration below shows that a forecast horizon of 100 data points is requested, which means that 100 additional predicted points will expand the initially fetched 120 data bins of the source DQL query. This predicts one week into the future.

Figure 3. Detail of the forecasting workflow step
Figure 3. Detail of the forecasting workflow step

The prediction action returns all its forecasted time series lines, which can include hundreds or even thousands of individual disk predictions.

Evaluate the forecast results

Within the following TypeScript action, each disk prediction is tested against a threshold to determine if the disk will run out of space in the next week. The TypeScript code snippet below is responsible for checking for threshold violations and for preparing all the violations in a result object for subsequent actions to follow up on:

Figure 4. Evaluating the results with a custom TypeScript action
Figure 4. Evaluating the results with a custom TypeScript action

The TypeScript action returns a custom object that uses a Boolean flag (violation) to tell the follow-up actions about violations and an array of all the violation details (violations).

const predictionSummary = { violation: false, violations: new Array<Record<string, string>>() };

Tip: Download the TypeScript template from our documentation.

Trigger remediation actions

A collection of remediation actions can be used to follow up on predicted capacity shortages. In this example, two parallel actions are defined. One action sends out an email notification; the other raises a Davis problem for each violating disk. All remediation actions use the Boolean violation flag of the previous workflow action to avoid invocations when there are no violations.

Here you can see the invocation condition used in the follow-up actions that control the invocation.

Figure 5. Conditional execution
Figure 5. Conditional execution

Raise events in case of disk capacity shortage!

A TypeScript remediation action is used to iterate through all the predicted disk shortages and to raise individual alarm events. Each alarm event has custom event properties that can be used to deliver further details about the situation and to further identify the disk or host.

Figure 6. Create an alarm event for predicted shortages.
Figure 6. Create an alarm event for predicted shortages.

Tip: Download the TypeScript template from our documentation.

Review all Davis-predicted capacity problems

Navigating to the Davis problems feed, the operations team can review all the predicted disk capacity shortages. Remember, raising events and problems is an optional remediation step that can be skipped entirely by directly sending emails or Slack messages to the responsible teams.

The creation of alerting events within this workflow example highlights the flexibility and power of the Dynatrace AutomationEngine combined with the analytical capabilities of Davis AI and Grail.

Figure 7. List of events created by the workflow.
Figure 7. List of events created by the workflow.

Summary

The combination of Davis AI forecasts with Dynatrace AutomationEngine and Grail opens the door for many valuable use cases—anticipative management of capacity being the most prominent of these. Predicting future capacity shortages for thousands of disks or hosts allows operations teams to anticipate critical situations weeks before incidents occur. The flexibility and power of the Dynatrace AutomationEngine allow operations teams to react to detected shortages flexibly and to customize and implement their remediation flows.

You can install Davis® for Workflows via the Dynatrace Hub. As a starting point for implementing your own anticipative capacity management workflow, you can download all the TypeScript code used in this example from our documentation:

For full details, see Davis AI analysis in workflows documentation.

Predictive capacity management in action (14-second video)

The post Automate predictive capacity management with Davis AI for Workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automate-predictive-capacity-management-with-davis-ai-for-workflows/feed/ 0
Log forensics: Finding malicious activity in multicloud environments with Dynatrace Grail https://www.dynatrace.com/news/blog/log-forensics-with-dynatrace-grail/ https://www.dynatrace.com/news/blog/log-forensics-with-dynatrace-grail/#respond Mon, 22 May 2023 06:00:17 +0000 https://www.dynatrace.com/news/?p=57709 Logs forensics graphic

Log forensics—investigating security incidents based on log data—has become more challenging as organizations adopt cloud-native technologies. Organizations are increasingly turning to these cloud environments to stay competitive, remain agile, and grow. But as organizations rely more on cloud environments, data and complexity have proliferated. Teams struggle to maintain control of and gain visibility into all […]

The post Log forensics: Finding malicious activity in multicloud environments with Dynatrace Grail appeared first on Dynatrace news.

]]>
Logs forensics graphic

Log forensics—investigating security incidents based on log data—has become more challenging as organizations adopt cloud-native technologies. Organizations are increasingly turning to these cloud environments to stay competitive, remain agile, and grow.

But as organizations rely more on cloud environments, data and complexity have proliferated. Teams struggle to maintain control of and gain visibility into all the applications, microservices and data dependencies these environments generate. Without visibility, application performance and security are easily compromised.

As a result, teams are turning to technologies such as observability to understand events in their cloud environments. Moreover, they have come to recognize that they need to understand data in context. But most observability technologies today provide information in silos. Without unifying these silos, teams miss critical context that can lead to blind spots or application problems—problems that compound when there’s a need to investigate security events.

Modern observability enables log forensics

Dynatrace is a software intelligence platform that provides deep visibility into and understanding of applications and infrastructure. It started as an observability platform; over time, it has expanded to provide real user monitoring, business analytics, and security insights. The recent innovation around log storage, processing, and analysis—Grail—makes Dynatrace a great solution for security use cases such as threat hunting and investigating the who-what-when-where-why-how of an incident.

Grail is a data lakehouse that retains data context without requiring upfront categorization of that data. Unifying data in Grail brings critical security capabilities to bear as teams seek to understand malicious events.

Grail enables organizations to find and analyze security events in the context of their broader cloud environments. Moreover, with capabilities such as log forensics—the analysis of log data to identify when a security-related event occurred—organizations can explore historical application data in its full context.

Grail makes it easy to query historical data without data rehydration or indexing, re-indexing, and up-front schema management. This accessibility gives users quick and precise results about when malicious activity occurred, when reconnaissance was first seen in the systems, what was attempted, and if the attackers were successful.

Demo: “Ludo Clinic” uses log forensics to discover and investigate attacks using Grail

Imagine working as a security analyst for a respectable medical institution called Ludo Clinic. As Ludo Clinic started using Dynatrace, the platform’s Runtime Application Protection feature detected a SQL-injection (SQLI) attack. Thanks to the details provided by the code-level vulnerability functionality, the developers knew where in the code the exploited vulnerability was and were able to address and patch it quickly.

The task now is to investigate whether the system experienced any suspicious activity before the attack so we can determine if any other systems are affected. The good news is there are metrics available a few days before your team detected the attack, and you also ingested three months of application and access logs into Grail. Because these logs are ready for querying with no rehydration, the investigation can start immediately. The bad news is we don’t know exactly what to look for. “Find suspicious activity” can mean anything. So, we’ll start by exploring the data using the hints and context information we already detected with Dynatrace.

Hint one: Blocked SQL injection report details

Here’s the report from Dynatrace on the blocked SQL injection details on 10 February from the IP 104.132.226.34. This report shows details of the attack, such as the entry point, the vulnerability that was exploited, the IP address of the attacker, and so on.

screenshot of Dynatrace blocked SQL injection report showing attack details

Hint two: Failed logins spike

A quick look at the metrics dashboard dating back to 8 February shows a spike in failed logins metrics before your team detected the attack. Indeed, there’s a spike on 8 February.

screenshot of failed logins spike

It would make sense to see if there’s any activity from that IP before we set up monitoring. Has the attacker been doing reconnaissance from that same IP in our systems even before this? If yes, how? Did they try something else during those three months? Were they successful?

Notebooks, DQL, and DPL: Tools of the Grail log forensics trade

Now that we have some clues about where to look for suspicious activity, we’ll dig into the logs using Notebooks, DQL, and DPL.

Dynatrace Notebooks

Dynatrace Notebooks is a collaborative data exploration feature that operates on data stored in Grail for ad-hoc exploratory analytics. Notebooks enable cross-functional teams, such as IT, development, security, and business analysts, to build, evaluate, and share insights for exploratory analytics using code, text, and rich media. The ability to build insights from the same data using the expertise of different roles helps organizations truly understand everything their data has to say.

Dynatrace Query Language (DQL)

Dynatrace Query Language (DQL) is a piped SQL-like query language, similar to Linux commands executed in sequence. You can look at the queries like a series of building blocks applied in an order you happen to need at this moment. Select fields, summarize, and count a value, apply more filters, select additional fields, extract data from a particular field, and so on. DQL is great for exploring and experimenting with data, which makes it a great ally in log forensics and security analytics.

The first query of our investigation uses DQL in Notebooks to fetch logs from Grail, filter the access log, and limit the result to 1000 records for initial exploration.

log forensics using Notebooks to start the investigation

Dynatrace Pattern Language (DPL)

In our investigation, we’ll also use DPL. DPL stands for Dynatrace Pattern Language, a parsing language that also consists of intuitive building blocks that help to extract meaningful fields from data on read. That means there is no need to manage indexes and rehydrate archived data; simply specify an ad hoc schema using DPL as part of the query.

What’s more, with DPL, the parsed results return typed fields, so you can be sure that a timestamp is a timestamp and an IP address is an IP address, not some random octets separated by a dot like 320.255.255.586. Working with typed data means excellent quality and precision for investigation results because you can run type-specific queries like calendar operations, calculations on numeric data, working with JSON objects, and so on. Working with typed data means excellent quality and precision for investigation results, as you can run type-specific queries like calendar operations, calculations on numeric data, working with JSON objects, and so on.

Log forensics: Querying the access log

Remember: our task is to investigate whether any other systems are affected. The first step is to query whether the IP address where the SQLI attack came from has been used before. Can we see it in the web application access log months prior to the attack?

screenshot of log forensics query of the access log using Dynatrace Grail

The query result shows there is no activity from that IP earlier than records on 10 February, the day Dynatrace detected the SQL injection attack. This means that the attack appeared “out of the blue,” and it is likely the attackers were using other IP addresses to do reconnaissance on our systems.

Because we can’t find the attacker by the IP address, let’s look at abnormalities in login behavior because there is a chance they’ll be related to reconnaissance. This means we’ll investigate the spike in failed logins we saw earlier in the metrics graph. Are there any other failed login spikes three months prior to the attack? Where do the failed logins originate from?

We can see that a failed login attempt takes users to a specific URL:

/ludo-clinic/login?authenticationFailure=true

So, let’s see if and how often this URL appears in the logs by adding the following filter to the query.

| filter contains(content, "/ludo-clinic/login?authenticationFailure=true") 
| limit 10000

Indeed, the query gives us 6531 records containing a failed login URL:

screenshot of log forensics query result showing 6531 records

Making sense of the access log

For the next stage of our investigation, let’s make more sense of these ~6,300 records and find out how many unique IP addresses were the origin of failed login attempts. The hypothesis is: some of the IP addresses stand out when it comes to the number of login failures. This means we first need to extract the IP address to run this aggregation.

We can utilize the schema-on-read functionality, that is, extract only the fields we need for a specific query. Taking a closer look at the content field of the access log, we can see a traditional HTTP access log: clientIP, timestamp, requestURL, HTTP response code, and so on.

screenshot of query results showing extracted fields clientIP, requestURL, HTTP response code, and so on

Notice that the timestamp field (the ingest timestamp) is similar for all log records (15/05/2023 14:09:48). This is because Ludo ingested the historical log records in bulk. To analyze the event time, we need to extract the timestamp from the content field as event time. To count IP addresses, we also need the IP address.

Quick ad-hoc parsing to aggregate login failures

To parse out data (timestamp and IP) from the content field, we’ll select the content field and select Extract fields to open the DPL Architect. To retrieve the timestamp and clientIP, we’ll replace the default DPL pattern with the following:

IPADDR:client_ip LD HTTPDATE:event_time

screenshot showing ad-hoc parsing timestamps and clientIP using the DPL architect

This matches and extracts the timestamp and the IP address from the content field and gives them a name (event_data and client_ip) and leaves the rest of the pattern unmatched, as we don’t need it for the following query. Clicking Insert pattern brings us back to the query view, adding a parse command to the newly created pattern.

screenshot showing results of parsing fields in DPL architect

Now with the extracted IP address, we can proceed with queries and use the summarize command to count the number of failed logins per unique IP address to see if there were any failed logins originating from a specific IP. Sorting the result set based on the number of failed logins in descending order gives us the largest outliers.

fetch logs 
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, "IPADDR:client_ip LD HTTPDATE:event_time" 
| fields event_time, client_ip, content 
| summarize total=count(), failed=countIf(contains(content, "/ludo-clinic/login?authenticationFailure=true")), by:client_ip 
| sort failed desc

This pays off! The results reveal that a significant portion of logins (181,774) and failed logins (6161) originate from the IP address: 172.31.24.11. This seems interesting and is worth taking a closer look.

screenshot showing the count of failed logins from the originating IP address

Timing of login failures

Next in our log forensics journey, let’s see when these failed logins from that particular IP address occurred to get more information on the potential reconnaissance activity. Did the requests all occur within a short period or regularly across a longer period?

Because we’re interested in the behavior of a specific IP and investigating the reasons behind failed logins, let’s also extract the session ID from the log line. As the session ID is the only field that occurs both in the access and application weblogs, it will be also useful later when we need to join the two for investigating affected users.

We already extracted the IP and included the timestamp (HTTPDATE). We will now extend the pattern and skip the part of the record we don’t need by not naming the three double-quoted strings (DQS). Finally, we’re extracting the last field that contains the session ID.

IPADDR:client_ip LD HTTPDATE:event_time LD DQS LD DQS LD DQS SPACE LD:session_id

screenshot showing a query that extracts session IDs involving the target IP address

When we select Insert pattern, we again get a parse command populated with the DPL pattern we just created.

Focusing on the suspicious IP

Next, let’s select only the fields we’re interested in and then aggregate fields. These actions reveal more about the extended activity that involves the suspicious IP address responsible for many of the failed logins.

| fields time, client_ip, session_id, content

screenshot showing the results of a query that extracts session IDs involving the target IP address

Filtering the attacker IP and sorting the fields based on the timestamp we just parsed out, it appears this IP address was first seen on 24 December 2022. We now know the start of the suspicious activity. For malicious actors, it is quite common to act during the holiday period.

fetch logs  
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, " IPADDR:client_ip LD HTTPDATE:event_time LD DQS LD DQS LD DQS SPACE LD:session_id" 
| fields event_time, client_ip, session_id, content 
| filter contains(content,"172.31.24.11") 
| sort event_time asc 
| limit 300000

screenshot showing a query that extracts the event times involving the target IP address

Find the suspicious activity pattern across time

To see the activity pattern of this suspicious IP across time, let’s count the number of failed logins in one-hour time intervals.

fetch logs 
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, "IPADDR:ip LD HTTPDATE:time LD DQS LD DQS LD DQS SPACE LD:sessionID EOS" 
| fields time, ip, sessionID, content 
| filter contains(content,"172.31.24.11") 
| summarize failed=countIf(contains(content, "/ludo-clinic/login?authenticationFailure=true")), by:bin(time, 1h)

screenshot showing a query that counts the number of failed logins involving the target IP address

It appears as though failed logins from this IP appear to follow a very regular pattern: 24 failed attempts every hour. Looks like this activity is automated and most probably refers to a dictionary attack: regular (automated) attempts from the attacker to try out different usernames and passwords, mostly with failed results.

But to escape the clinic’s countermeasures (failed login attempts velocity check), the attacker also conducts a successful login every now and then. If we count all activity from that IP address (not just the failed logins but successful attempts as well), the results are again very symmetrical:

fetch logs 
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, "IPADDR:ip LD HTTPDATE:time LD DQS LD DQS LD DQS SPACE LD:sessionID EOS" 
| fields time, ip, sessionID, content 
| filter toString(ip) == "172.31.24.11" 
| summarize count=count(), by:bin(time, 1h)

screenshot showing a query that returns all logins from the target IP address.

Identify targeted users

Next, it would be useful to know which users the attacker has targeted and whether any attempts have been successful. Let’s aggregate the activity from this IP using sessionIDs:

fetch logs 
| filter contains(log.source, "ludo-clinic-access.log") 
| parse content, "IPADDR:ip LD HTTPDATE:event_time LD DQS LD DQS LD DQS SPACE LD:sessionID EOS" 
| fields timestamp, event_time, ip, sessionID, content 
| filter contains(content, "172.31.24.11") 
| summarize count=count(), 
            by:{sessionID 
               } 
| sort count desc 
| limit 10000

The result is again peculiar, suggesting automated activity: 59 log lines per session.

screenshot showing a query that identifies logins by session ID that suggests automated activity.

Next log forensics dataset: The web application log

Next, let’s see what was happening based on the web application log, using data from what was going on during those sessions that originated from the suspicious IP address we discovered from the access log dataset.

First, to familiarize ourselves with the content of the webapp log, let’s run a basic query to see what the content field of the web application log looks like:

screenshot showing a log forensics query that shows content of the web application log.

We can see a timestamp, log severity, traces and spans, a session ID, result, and username. There are plenty of interesting fields to play with, so the next step is to parse the content into fields that are ready for querying. We can extract the fields using the DPL Architect. The following DPL pattern extracts the event time session ID, result, and username from the webapp log.

'[' TIMESTAMP('dd/MMM/yyyy:HH:mm:ss.S'):event_time LD ' - ' LD:sessionIdApp ' ' LD:result_text ': ' LD:username (';' | EOS)

screenshot showing a query that parses out the fields of interest for the log forensics

Inserting the pattern, this is what the query looks like when parsing out session IDs and usernames from the web application log.

fetch logs, from:-300d   
| filter contains(log.source, "ludo-clinic-webapp.log") 
| parse content, "'[' TIMESTAMP('dd/MMM/yyyy:HH:mm:ss.S'):event_time LD ' - ' LD:sessionIdApp ' ' LD:result_text ': ' LD:username (';' | EOS) " 
| limit 10000 
| fields event_time, sessionIdApp, result_text, username

screenshot showing the results of parsing the fields of interest in the web application log

Filter out records from the authentication provider

To see which users were targeted and how successful the attacker was, we will continue working with only those web application records that contain authentication responses. First, we filter out the records that originate from the authentication provider, then we skip the responses we’re not interested in:

| filter contains(content, "CustomAuthenticationProvider")  
  AND NOT contains(content, "Starting findUsersByUsernameAndPassword") // we want to see only auth response log records 
  AND NOT contains(content, "retrieved matching list")

The full query now looks like this and returns the following results:

fetch logs  
| filter contains(log.source, "ludo-clinic-webapp.log") 
| filter contains(content, "CustomAuthenticationProvider")  
  AND NOT contains(content, "Starting findUsersByUsernameAndPassword") // we want to see only auth response log records 
  AND NOT contains(content, "retrieved matching list") 
| parse content, "'[' TIMESTAMP('dd/MMM/yyyy:HH:mm:ss.S'):event_time LD ' - ' LD:sessionIdApp ' ' LD:result_text ': ' LD:username (';' | EOS) " 
| limit 10000 
| fields event_time, sessionIdApp, result_text, username

screenshot showing the full query with the fields of interest from the web application log.

Review the user sessions that originate from the attacker

Next, to see which user sessions in the webapp log originated from the attacker’s activity, we use a lookup query to join aggregated sessions from the attacker IP address we discovered in the access log with sessions in the webapp log. In short, we will see what was happening during the suspicious sessions according to the webapp log.

fetch logs  
| filter contains(log.source, "ludo-clinic-webapp.log") 
| fields content 
| filter contains(content, "CustomAuthenticationProvider")  
  AND NOT contains(content, "Starting findUsersByUsernameAndPassword") // we want to see only auth response log records 
  AND NOT contains(content, "retrieved matching list") 

| limit 10000 
| parse content, "'[' TIMESTAMP('dd/MMM/yyyy:HH:mm:ss.S'):event_time LD ' - ' LD:sessionIdApp ' ' LD:result_text ': ' LD:username (';' | EOS) " 
| lookup [fetch logs, from:-300d 
                | filter contains(log.source, "ludo-clinic-access.log") 
                | filter contains(content,"172.31.24.11") 
                | parse content, "IPADDR:ip LD HTTPDATE:time LD DQS LD DQS LD DQS SPACE LD:sessionID EOS" 
                | filter isNotNull(sessionID) 
                | limit 100000 
                | summarize accesscount=count(), by:{sessionID} 
                | fields sessionID], sourceField:sessionIdApp, lookupField:sessionID 

| fieldsRemove content

Screenshot showing the results of attempted authentications.

The result shows the attacker has achieved both successful authentications as well as failed authentications. Finally, we see which usernames the attack targeted the most by looking for the response “No users found requested username.” The system returns this value when it receives a non-existent user or a wrong password. By aggregating the result based on unique usernames, we get a list of the most (unsuccessfully) targeted users.

| filter result_text == "No users found requested username" 
| summarize count(), by:{username} 
| sort `count()`desc

screenshot showing the no users found query that reveals the targeted user accounts

These results are fascinating – we can see five usernames that the attacker continuously entered and received failed authentication results. We can also see SQL commands instead of regular usernames.

Determine successfully targeted users

Next question: Did they achieve anything besides ‘No users found’ when targeting these users? Let’s have a look by excluding the “No users found requested username” response and concentrating on those five users from the last query result, and adding the following line to the query:

|  filter not matchesPhrase (result_text, "No users found requested username") and in (username, "arnie", "herman", "krzysztofs", "bernice", "sherry")

screenshot showing drilldown to identify affected usernames.

Indeed, we see a lot of “successfully authenticated” responses in the result text field. This confirms the attackers were successfully conducting a dictionary attack: trying out several usernames and passwords to authenticate as real users of Ludo Clinic. When counting the number of successful authentications per these five users, the results are quite similar:

|  filter not matchesPhrase (result_text, "No users found requested username") and in (username, "arnie", "herman", "krzysztofs", "bernice", "sherry") 
| summarize count(), by:{username} 
| sort `count()`desc

screenshot showing top targeted users with successful authentication.

Investigation results from log forensics and metrics with Dynatrace

As a result of Dynatrace detecting a SQL vulnerability, anomalies in metrics, and subsequently running forensic queries on three months of logs prior to the attack, we’ve been able to construct the following timeline:

  • As Ludo Clinic started using Dynatrace, they were able to observe a spike in metrics capturing failed logins on 8 February
  • The system detected and blocked a SQL injection attack on 10 Feb (Fri)
  • There was no other activity from that IP in the access logs (the logs reach back three months)
  • When aggregating failed login activity, we discovered the following details: the IP address 172.31.24.11 stands out from the rest, counting to 6161 failed logins based on the access log during the past three months
  • This IP was first seen in the logs on 24 December 2022 (the earliest timestamp for this set of logs is 20 November 2022)
  • The sessions contain identical activities during identical timeframes, which suggests the attacker was using automated tools
  • Joining sessions from the access log to the application log reveal almost ten thousand records originating from the suspicious IP 172.31.24.11
  • It looks like the attackers were attempting a dictionary attack because it targeted several users at regular intervals, resulting in failed as well as successful authentications.
  • Users stafford, ray, orrel, doug and joby were targeted to discover their passwords. The attacker successfully authenticated 35-37 times per user.
  • The attackers also entered SQL statements instead of usernames attempting SQL injection attacks.

The DQL and DPL advantage

For such investigations, DQL and DPL make it convenient to quickly investigate and query logs for security analytics use cases that require drawing broad conclusions from the data one minute and then zooming into the activities of a specific session the next. An interesting find inspires the analyst to parse out yet another field and run aggregations on this data. As historical data is always ready for querying, all hypotheses can be quickly verified or dismissed. A curious mind and the right log forensics tools (DQL and DPL) make a great combination for fighting evil.

To see more of Grail in action for log forensics and exploratory analytics, join us for the Observability Clinic: The Practitioner’s Guide to Analytics without Boundaries with Dynatrace.

The post Log forensics: Finding malicious activity in multicloud environments with Dynatrace Grail appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/log-forensics-with-dynatrace-grail/feed/ 0
Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI https://www.dynatrace.com/news/blog/stay-ahead-of-the-game-forecast-it-capacity-with-dynatrace-grail-and-davis-ai/ https://www.dynatrace.com/news/blog/stay-ahead-of-the-game-forecast-it-capacity-with-dynatrace-grail-and-davis-ai/#respond Wed, 26 Apr 2023 12:03:45 +0000 https://www.dynatrace.com/news/?p=57280 Cloud observability graphic

Anticipatory management of cloud resources within highly dynamic IT systems is a critical success factor for modern companies. Operators need to closely observe business-critical resources such as storage, CPU, and memory to avoid outages that are driven by resource shortages.

The post Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI appeared first on Dynatrace news.

]]>
Cloud observability graphic
>> Scroll down to see Davis Forecasting in action

Traditionally, cloud-resource management is done by collecting telemetry data for critical-capacity resources and configuring multi-level reactive alerting (warnings, errors, and critical errors) for those resources.

While the traditional approach to cloud-resource management might have been acceptable in the past, it doesn’t scale up to address the requirements of modern cloud environments. Highly dynamic services are deployed to the cloud globally, where resources are requested and deployed on demand. The end result of this global scale is that—without the right tools—operators are completely lost in alert storms.

Some of our customers run tens of thousands of storage disks in parallel, all needing continuous resizing. This can lead to hundreds of warnings and errors every week. The most annoying aspect of this, according to the operations teams, is that alerts are often sent after business hours, including on weekends.

Disk alert storm
Figure 1. In the past, disk alerting was triggered using static capacity thresholds.

One effective capacity-management strategy is to switch from a reactive approach to an anticipative approach: all necessary capacity resources are measured, and those measurements are used to train a prediction model that forecasts future demand.

The example below shows how the reactive approach can be transformed into a scheduled, predictive capacity management model. With this approach, when available capacity of a critical resource is forecast to soon fall below acceptable levels, the operations team is notified with a single report, sent during business hours, well in advance of the resource actually experiencing a resource shortage.

Disk forecast report
Figure 2. Capacity planning with Dynatrace Davis® AI forecasts actual usage rather than waiting for thresholds to be exceeded before alerts are sent.

Use Grail and Davis to predict capacity demands

The Dynatrace Query Language (DQL) allows you to analyze all data that’s stored for your environment within the Dynatrace Grail™ data lakehouse. With the newly introduced Notebooks, you can use DQL for exploratory analysis of any capacity-related telemetry.

The screenshot below shows a section of an example notebook that plots the average percentage of free disk space over the last 7 days.

Disk capacity notebook
Figure 3. Example notebook showing average percentage of free disk space during past 7 days.

Select any line in the chart to display the available actions for that line. For example, you can request that Davis AI forecast the future of any given time series.

Disk capacity notebook forecast
Figure 4. Select Filter and forecast for any chart line to start Davis Forecast.

Davis AI analyzes the selected time series, automatically chooses the best prediction model based on the characteristics of the time series, and then trains a prediction model. After the training is finished, the notebook chart shows a probabilistic forecast of the given time series with upper and lower bounds as well as the predicted value.

The example result of the trained prediction model on this disk capacity measurement shows that the lower bound of the prediction (worst case scenario) will fall below 6% free disk capacity in early April.

Given this information, the operations team can anticipate that they need to resize this disk before early April (during business hours, of course).

Disk capacity notebook forecast result
Figure 5. Example forecast of remaining disk capacity with upper/lower bounds and an anticipated value.

While this notebook focuses on only one disk, Davis Forecast can learn and predict the future capacity needs of thousands of individual disks in parallel. For example, in our own cloud infrastructure at Dynatrace, we track over 8,000 disks that require periodic resizing. By running a scheduled, weekly forecast, our cloud automation teams avoid reactive alerts sent outside of business hours.

Reactive alerts as last line of defense

Of course, unexpected things still happen and a weekly forecast can’t, for example, anticipate a customer onboarding 5,000 OneAgents on a Sunday morning. Therefore, reactive alert conditions remain in place, to ensure that alerts are still sent if needed during unexpected events. Scheduled forecasts do not replace these reactive alerts, rather they serve as a last line of defense for unforeseen situations.

AutoML detects seasonality and chooses the best prediction model

Dynatrace invested significantly into simplifying prediction and forecasting for you. You can use any DQL query that yields a time series to train a prediction model. This AutoML approach analyzes the statistical characteristics of any time series (variance, seasonality, trend, and noise) to determine the best prediction model.

The AutoML approach also helps when you need to automate your forecasts and the underlying metric characteristics change over time.  Please see Davis Forecast analysis documentation to learn more about our AutoML approach and which algorithms are used within the Davis Forecast service.

So far, we’ve only discussed resource consumption measurements, which by nature show more linear changes than seasonal characteristics. Now let’s see how the AutoML approach selects the best suitable method for a time series that has seasonal behavior. The example below shows that Davis Forecast automatically detects any given seasonality, independent of the noise level, and correctly returns a probabilistic forecast.

Probabilistic seasonal forecasting with AutoML
Figure 6. Probabilistic seasonal forecasting with AutoML

Automate your Davis capacity forecasts

While the operations team could regularly check this notebook, see what Davis Forecast anticipates as the upcoming capacity need, and then proactively resize all disks that will run out of space the following week, the better option is to use the newly introduced AutomationEngine to schedule an automated weekly Davis prediction workflow. With this approach, an automated workflow can automatically run a forecast for all disks, check against a critical capacity limit, and notify the operations team with a list of the disks that need their attention. Below is an example of such a workflow that includes a Davis Forecast action and a notification email action. The new forecasting capabilities together with Dynatrace AutomationEngine and the Workflows app, allow you to automate any predictive analytics in a few simple steps.

Davis Forecast as part of a workflow automation
Figure 7. Use Davis Forecast as part of a workflow automation.

This use case can be spun even further: so far, we’ve automated the forecast and introduced reporting for the disks that need to be resized. Once the operations team becomes familiar with the anticipatory approach, full automation can easily be configured by adding additional action steps within the existing workflow, for example, automated provisioning of new disk space.

Summary

Davis Forecast provides a powerful mechanism on top of the Grail data lakehouse that enables organizations to switch from reactive strategies to more proactive anticipative strategies. Such predictive approaches help avoid outages and reactive alert storms outside business hours.

By offering a standard forecast mechanism on top of the powerful DQL query language, Dynatrace opens predictive analytics for any kind of anticipative use cases, including the business-critical topic of predictive capacity management. These new analytics capabilities can be used as part of your exploratory analytics in Notebooks, as a step within workflows, or as part of your custom app–addressing your specific business needs with Dynatrace AppEngine.

See Davis Forecasting in action

Check out the “Forecasting with Dynatrace” Observability Clinic, where Linda Gratzer, Andreas Grabner, and Bernhard Kepplinger dig deeper into the topic of forecasting and the data science behind it, and also share a live demonstration.

For further details, have a look at our Davis AI Forecast Analytics documentation, or watch the recording of my Perform breakout session, Easy forecasting and predictive analytics with Davis AI.

We are of course highly interested in your feedback! We encourage you to try out Davis Forecasting and then head over to the Dynatrace Community and share your suggestions and product ideas, to help us continuously improve the Dynatrace platform.

The post Stay ahead of the game: Forecast IT capacity with Dynatrace Grail and Davis AI appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/stay-ahead-of-the-game-forecast-it-capacity-with-dynatrace-grail-and-davis-ai/feed/ 0
Expanded Grail data lakehouse and new Dynatrace user experience unlock boundless analytics https://www.dynatrace.com/news/blog/boundless-exploratory-observability-and-security-analytics/ https://www.dynatrace.com/news/blog/boundless-exploratory-observability-and-security-analytics/#respond Wed, 15 Feb 2023 18:00:12 +0000 https://www.dynatrace.com/news/?p=56090 three pillars of observability converge on the Grail data lakehouse

Last October, we introduced Dynatrace Grail™, our causational data lakehouse. From day one, Grail disrupted the log management and analytics market by unifying observability, security, and business data and providing instant answers thanks to its massively parallel processing (MPP) capabilities.

Further extending our platform's analytics capabilities, we're increasing Grail's capabilities by adding new data types and unlocking support for graph analytics. These capabilities enable Davis®, the Dynatrace causal AI engine, to gather even more insights. They also enable an entirely new way of interacting with data and performing any analysis without boundaries.

The post Expanded Grail data lakehouse and new Dynatrace user experience unlock boundless analytics appeared first on Dynatrace news.

]]>
three pillars of observability converge on the Grail data lakehouse

Grail – the foundation of exploratory analytics

Grail can already store and process log and business events. Now we’re adding Smartscape to DQL and two new data sources to Grail: Metrics on Grail and Traces on Grail.

Grail infographic
Grail is addressing a lot of shortcomings of common databases.

Introducing Metrics on Grail

Despite their many advantages, modern cloud-native architectures can result in scalability and fragmentation challenges. Ensuring observability across these environments requires access to data at a massive scale. The proliferation of metrics can quickly result in a high cardinality challenge, with each service, host, or Kubernetes pod adding its own unique values to the data set.

Grail solves this scalability issue! Metrics on Grail is architected to manage billions of metrics to cope with cardinalities and unique value combinations of 1 trillion potential permutations for timeframes beyond a year. This is only possible because of our no-index approach and massive parallel processing capabilities, which enable Dynatrace to offer extra-long data retention (15+ months) at full granularity that is cost-efficient and fast.

You no longer need to split, distribute, or pre-aggregate your data. Let Grail do the work, and benefit from instant visualization, precise analytics in context, and spot-on predictive analytics.

Get instant visualization, precise analytics in context, and spot-on predictive analytics from Grail

Introducing Traces on Grail

A distributed trace follows a transaction on its journey through every service, cloud platform, and host in your environment. Having access to traces that span the full hybrid and multicloud stack enables developers to debug their applications in production and understand dependencies in live environments. For more complex cloud-native architectures, adding more services and applications leads to a massive increase in the volume of collected traces.

With Grail, we address these customer challenges by offering the most powerful and future-proof trace analytics solution on the market, which:

  • Handles data volumes of hundreds of terabytes a day
  • Retains large data volumes for up to 15 months in a highly cost-efficient way
  • Ensures that data retains its context by assembling trace spans into PurePath® distributed traces (including additional code and thread profiling data)
  • Returns instant query results in real-time using indexless queries

Traces in Grail

Smartscape for DQL: Context is king

Bill Gates wrote an essay in 1996 entitled “Content is King” in which he described the future of the internet as a marketplace for content. In DevSecOps, content includes applications and services—in addition to information about the environments where they run and the users who use them. Observability and application security use cases rely on data. However, data on its own, without context, doesn’t reveal all its insights. Whereas Bill Gates’ observation is still valid, for the DevSecOps industry today, a more accurate description is “context is king.”

In a traditional monitoring environment, metrics are aggregated data points that lose their context and granularity when data sets are trimmed to make them more manageable. With Dynatrace and Smartscape for DQL, metrics are a completely different game. Whether it’s metrics, logs, events, traces, or any other data type, Dynatrace not only retains the data context but also enables you to analyze data in its semantic context without boundaries.

These capabilities are powered by Smartscape for DQL, a directional graph representing the real-time topology and dependencies of a data architecture. Smartscape unifies the different data types ingested into Dynatrace and retains the full context of this data to enable holistic and precise data analytics.

With the Dynatrace Query Language (DQL), teams can perform these analyses by asking questions that weren’t possible in the past. There are now boundless possibilities, such as identifying users affected by a service outage in a red-alert scenario or doing forensic research on a recent data breach. With DQL, you can easily combine different data types into a single query.

Sample DQL query combining multiple data types
Sample DQL query combining multiple data types. Thanks to Smartscape for DQL, this query filters on causal-dependent information.

The power of Smartscape is, of course, not limited to manual queries. The same data model fuels Davis, the causational AI engine at the core of the Dynatrace platform. Dynatrace has used Davis for many years and is leveraging its power for root cause analysis, identifying security risks, and many other use cases. Davis doesn’t rely on machine learning or statistical correlations—the models that power most available AIs and try to correlate data points by timestamp analysis, searching for similarities, or processing manual instrumentations. Alternatively, Davis is causal AI that reflects continuously updated topology and dependencies (powered by Dynatrace Smartscape) and understands the precise relationships and dependencies between isolated signals.

Whereas other AIs must guess, Davis knows and eliminates false positives. With the addition of Dynatrace Grail, which ingests, retains, and maintains data in context, we’re revolutionizing the observability industry and extending Dynatrace further to provide answer-driven analytics and automation for unlimited observability and security use cases.

Exploratory analytics – empowering people and data

While data is considered by some to be the new gold, it’s people that still make the difference. Gaining insights from data stored within Dynatrace has traditionally been limited to people within an organization who have specific expertise and training. This is no longer the case.

With the new Dynatrace user experience, we’re introducing new concepts and changing how people across organizations work and interact with data.

New user experience

How many user interfaces have you used that are defined by the data and data types they show rather than the use cases they support? How often have you spent time decluttering or trying to make sense of the information presented on a dashboard? How often have you wished you could interact with data in the same intuitive way you interact with information on your smartphone, quickly switching between visualizations, easily understanding the context behind a spike in a chart or diagram, or digging deeper to perform ad-hoc analysis?

When we started working on Strato—the new Dynatrace design language that powers our new user experience—we developed a few core principles to address the design challenges stated above:

  • Designing software for DevSecOps use cases means handling data—large volumes of data that need to be accessible for in-depth analysis in an easy-to-digest interface. We therefore completely rethought the user experience: the interface is user-centric rather than data-centric. We designed Dynatrace in a way that places the user in the middle, offering a flexible UI—tailored to individual needs and deriving rich insights from different perspectives.
  • The interface is simple—whether you’re a first-time user, an occasional user, or an SRE using Dynatrace as your single source of truth, the experience is simple and easy to learn.
  • The interface offers infinite possibilities. Users need to be able to work efficiently regardless of how large their environments are. Sharing and collaborating with teams is now easier than ever before.

Video thumbnail

These principles all align with a single, overarching goal: making data and insights derived from analytics available to a wider audience. To achieve this, we designed the new Dynatrace user experience (UX) to facilitate collaboration with teams across organizations—IT, development, security, and business—and solve everyday problems. We focused on democratizing the user interface, making it less trivial and more accessible, and empowering teams to better understand data signals and make data-backed decisions.

“When you allow data access to any tier of your company, it empowers individuals at all levels of ownership and responsibility to use the data in their decision making.”

—  @BernardMarr

User in context

We already mentioned above that putting data in context is vital for Dynatrace. This is also true from a user experience perspective. With Strato, we add user context to the Dynatrace UI.

Charts are now interactive—data points are clickable. Think of a chart that shows a spike in response time because of a deployment two hours earlier—any user can now hover over this data point and begin interacting with the underlying data, whether it’s a drill-down or just the context of the spike.

Introducing Every component and view in the Dynatrace web UI is interlinked based on user context or “intents”. Similar to what you know from your smartphone when sharing an image on your favorite social media channel, when opening a page within Dynatrace, you can easily pass and share the context of your analyses to any other app on the Dynatrace platform. The selection of available apps that are presented is completely context-sensitive and can even be expanded based on your needs.

Simplifying data analytics with Dashboards and Notebooks

In addition to new concepts revolutionizing the overall Dynatrace user experience, we’re introducing two new apps to the Dynatrace platform: Dynatrace® Dashboards, a complete overhaul of the dashboarding experience, and Dynatrace® Notebooks, for on-demand data exploration. These apps make it easier for more team members to explore, visualize, and collaborate on analytics projects. The following principles guided the development of these new capabilities:

  • Offer drastically faster and simpler flows, guided by Strato, the new Dynatrace design system
  • Fetch all Dynatrace data from one place and even combine it in a single query with Grail
  • Integrate external data easily with Dynatrace functions
  • Add flexibility and versatile filtering with variables
  • Take context with you as you seamlessly navigate the Dynatrace UI using intents

Dashboards and Notebooks have individual strengths that make them the best choice for solving specific use cases.

Observe data with Dynatrace Dashboards

Dynatrace® Dashboards transforms complex data into easy-to-understand visualizations. Dashboards is your go-to app for quick and clear data overviews, whether you need a status-quo view that can be observed over time or you need to share aggregated views with management or business teams.

Dynatrace Dashboards

Dashboards offers an interactive experience with full support for all capabilities and datatypes offered by Grail, allowing you to query not only metrics, but also logs, events, and even external data. It serves as your starting point for further deep-dive analysis, offering more detailed drill-downs via Notebooks (see below).

The all-new Dashboards app is your answer to live data visualization and observation. From data to insights in seconds, and we’re just at the beginning.

Video thumbnail

Explore data with Dynatrace Notebooks

Dynatrace Notebooks is your on-demand window into data exploration. It addresses the challenge of finding the right data, cleaning, filtering and transforming the data and finally connecting with other data to understand underlying dependencies. You no longer need the help of a data scientist for such tasks—Notebooks enables every Dynatrace user to perform any type of analysis on data stored in Dynatrace.

Dynatrace Notebooks

Start creating data-driven documents and perform custom analytics. Depending on your use case, Notebooks can persist a status quo and create a snapshot whenever necessary or be “self-updating” using current data to always reflect the actual status. You can easily interact with any query result by “slicing and dicing” the data stored in Grail: advanced filters, refinements, aggregations, and sort orders, are just a click away. It’s even possible to harness the power of Davis by adding predictive forecasts to identify future trends with a simple click.

Suppose you need the limitless power of Grail. In that case, you can easily create and edit DQL queries to filter, join, and transform data any way you need it, or even extend Notebooks with custom logic and external data by adding ad-hoc functions powered by Dynatrace® AppEngine.

Whether you’re analyzing new opportunities, the latest vulnerability post-mortem, or your executive production report, Notebooks does it all and ensures that both your query and results are persisted and ready to be shared with your team members. Empowered by Grail, Notebooks is the Swiss Army knife of the Dynatrace platform—built for collaborative data exploration and analysis.

Video thumbnail

Summary

With these newly added capabilities, Dynatrace users can now perform any custom query, leveraging Dynatrace Grail AI-fueled graph analytics power. This delivers instant and precise answers for an unlimited array of use cases:

  • Business impact: quickly identify impacted users by mapping observability and security findings to your business context and improve customer satisfaction by querying for e-commerce customers who are unable to finalize their check-outs due to a service outage.
  • Automation: understand the potential impact of remediation actions on dependent components.
  • Security: protect customers and brands by conducting application security forensics to identify, mitigate, and prevent data breaches.
  • Business process health: show and analyze the health status of complex processes even if dependencies are non-transactional.
  • Optimize: enable more efficient multicloud operations by predicting cloud performance and utilization over time to optimize resource allocation based on user needs.

The Dynatrace analytics platform converges security and observability data, enabling cost-effective end-to-end analytics at a large scale with long retention times, in context with your business, thus multiplying your value from data: every imaginable analysis of data in Dynatrace is now possible!

What’s next?

The new Dynatrace user experience, including the newly designed Dashboards, Notebooks, and Dynatrace Grail support for metrics and Dynatrace® Smartscape for DQL, will be available in Q2 2023. Grail support for Dynatrace PurePath® distributed traces will be open for preview in Q2 2023.

In the meantime, watch out for upcoming Observability Clinics and “Ask me anything” sessions covering the main topics of this blog post. You can either view a list of the next webinars on our website or follow us on LinkedIn to stay up to date with upcoming announcements and activities.

The post Expanded Grail data lakehouse and new Dynatrace user experience unlock boundless analytics appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/boundless-exploratory-observability-and-security-analytics/feed/ 0