Smartscape | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Fri, 10 Jul 2026 15:05:30 +0000 en hourly 1 Dynatrace Release Radar 06.26 https://www.dynatrace.com/news/blog/dynatrace-release-radar-06-26/ https://www.dynatrace.com/news/blog/dynatrace-release-radar-06-26/#respond Thu, 09 Jul 2026 16:52:43 +0000 https://www.dynatrace.com/news/?p=74748 Release Radar

This series covers recent Dynatrace releases and updates, focusing on what’s new, what’s changed, and how these recent enhancements can benefit you and your organization. Each post covers newly available capabilities and where to explore them.

The post Dynatrace Release Radar 06.26 appeared first on Dynatrace news.

]]>
Release Radar

If you want to see them in action, head over to our Release Radar launchpad on the Dynatrace Playground.

Smartscape gets a unified topology view and ad-hoc filters

In a significant Smartscape update, a new All topology view shows every relationship for a given node in a single graph: the infrastructure stack, communication flows, and relationships such as monitoring, load balancing, routing, and API dependencies. Where the existing Vertical and Horizontal views each focus on a subset of relationships, the All view provides a more comprehensive view of relationships, from any node in any app across the platform.

Two changes make these views faster and more focused:

  • The AWS and Kubernetes views now use flat layouts instead of nested ones, bringing the same relevance-based edge fetching and priority-driven node loading used elsewhere in Smartscape for more consistent visibility across your cloud landscape.
  • New ad-hoc node and edge filters let you narrow any view by node type, cloud and infrastructure labels, team ownership, environment, and other properties.

Filters work alongside segments and are saved in the URL, so you can bookmark and share a focused view with segment, timeframe, and filters all preserved.

Ad-hoc filters narrow a Smartscape view by team ownership and environment, highlighting matching nodes and preserving the filter state in the URL.
Ad-hoc filters narrow a Smartscape view by team ownership and environment, highlighting matching nodes and preserving the filter state in the URL.

Press Ctrl+F (Cmd+F on Mac) in any Smartscape view to find nodes by name or ID. Matching nodes are highlighted in the graph, and the legend is narrowed to matching entity groups.

For broader context on Smartscape, see The new Dynatrace Smartscape improves operational efficiency across clouds, Kubernetes, infrastructure, and more.

AI Observability gains LLM evaluation, OpenInference, and Python instrumentation

AI applications can fail without obvious indicators — returning responses at normal speed with no errors, while delivering answers that are inaccurate, unsafe, or inconsistent. Traditional performance monitoring misses this entirely.

dt-evals is a new open source CLI that closes that gap. It pulls live gen_ai.* spans directly from your Dynatrace environment. Built-in evaluators use an LLM judge to score real production interactions for faithfulness, hallucination, relevance, toxicity, bias, PII leakage, prompt injection, and drift. The judge writes structured results back to Dynatrace as business events. Evaluation scores sit alongside latency and error metrics in the same dashboards. These scores can trigger alert workflows and gate CI/CD releases based on quality thresholds the same way that performance metrics do. For the thinking behind this approach, see Evaluate LLM and agent quality in Dynatrace AI Observability and LLM evaluations as a foundation for trustworthy agentic AI systems.

Evaluation quality scores, pass rates, and drift trends from dt-evals running alongside model latency and token usage — turning AI quality into the same kind of operational signal as performance.
Evaluation quality scores, pass rates, and drift trends from dt-evals running alongside model latency and token usage — turning AI quality into the same kind of operational signal as performance.

Dynatrace OneAgent now automatically instruments Python applications that use AWS Bedrock, OpenAI, Azure OpenAI, and LangChain. Dynatrace captures distributed traces, logs, and AI-related telemetry for supported model interactions — provider, operation, model, duration, token usage, and prompt and completion metadata where available. To capture prompt and completion content, go to OneAgent features and turn on Python OpenAI prompt capture.

The same visibility extends to teams using OpenInference with OpenTelemetry (OTel). Dynatrace ingests OpenInference traces and normalizes them to the same gen_ai.* attribute schema — covering model usage, token consumption, prompts, completions, agents, tools, embeddings, and guardrails — so OTel-instrumented applications get consistent telemetry without switching instrumentation frameworks.

As AI adoption grows, evaluation and instrumentation together turn AI services into observable, governable assets rather than black boxes.

Logs gains pattern analysis, Kubernetes insights, and in-context traces

Log analysis gets three meaningful upgrades.

Log pattern analysis (Preview) lets you aggregate query results in Logs into patterns that cluster similar logs together. You can focus quickly on recurring errors, reduce thousands of similar logs to a handful of patterns, recognize the changing parts of a pattern (and their datatypes), and reuse the generated Dynatrace Pattern Language (DPL) for other queries or in OpenPipeline.

Log pattern analysis grouping thousands of similar entries into a handful of patterns, with dynamic segments highlighted and DPL ready to reuse.
Log pattern analysis grouping thousands of similar entries into a handful of patterns, with dynamic segments highlighted and DPL ready to reuse.

In-context trace details mean that when you investigate a log entry with trace context, you can open the associated trace directly inside Logs. A waterfall icon signals that you stay in context rather than navigating away to Distributed Tracing.

Log insights in ready-made Kubernetes dashboards provide built-in log analytics for clusters, namespace workloads, namespace pods, and node pods. Error log counts appear alongside health metrics, with log level distribution and severity trends below. Direct links to the Logs app ensure that a deeper investigation is only one click away.

Faster service investigation with the Services Explorer Preview

The Services app now includes a visual service map that overlays performance and health indicators on service-to-service relationships and messaging flows. It’s the fastest way to understand blast radius during an incident, providing a single view of topology context, performance signals, and bottlenecks without switching views.

The Services Explorer service map overlaying performance indicators on service-to-service relationships to pinpoint blast radius during an incident.
The Services Explorer service map overlays performance indicators on service-to-service relationships to pinpoint the blast radius during an incident.

You can also filter services directly by primary Grail fields such as k8s.cluster.name, k8s.namespace.name, aws.region, and azure.location — the same attributes that power segments across Dynatrace. Both capabilities are available in the Explorer Preview view and open for feedback before general availability; see the Community post for details.

New security integrations and a Kubernetes security tab

Threat Observability expands its ingestion options with new integrations. Dynatrace now integrates with Checkmarx for software composition analysis and container security findings, and adds CrowdStrike and Kyverno integrations — pulling detection findings and Kubernetes policy compliance data into Dynatrace as security events. For Kyverno, see Ingest Kyverno compliance findings.

Kubernetes monitoring also gets a dedicated security tab (Kubernetes app version 1.42.0+) that replaces the Vulnerability tab in the Explorer, bringing security context into the same place teams already investigate cluster health.

The Security tab surfacing vulnerability, detection, and misconfiguration findings alongside Kubernetes cluster health — without leaving the monitoring context.
The Security tab surfacing vulnerability, detection, and misconfiguration findings alongside Kubernetes cluster health — without leaving the monitoring context.

Runtime Vulnerability Analytics now has a native interface, replacing the legacy management-zone-based monitoring rules with a single consolidated workflow.

One change worth flagging for security teams: ingested security.events must now carry a timestamp within −1h/+10min, tightened from the previous −24h/+10min window. Events with older timestamps are dropped, so please review any pipelines that backfill security events.

Performance, drilldowns, and navigation improvements

Improved discovery of ready-made dashboards. Ready-made dashboards deliver instant insights without requiring complex queries. Finding, installing, configuring, and customizing them is now more straightforward — so new users get value faster and experienced users can build confidently on best-practice templates.

The Hub discovery workflow guides you from platform search to installable ready-made dashboards.
The Hub discovery workflow guides you from platform search to installable, ready-made dashboards.

Contents tab added to all extension apps. All extension apps in Dynatrace Hub now include a Contents tab that surfaces the extension’s ready-made dashboards, so you can quickly go from installation to insights.

Session Replay has two improvements:

  • Full-screen mode is now available, removing viewport constraints during playback.
  • Navigating to a session through Error Inspector now opens Session Replay directly in context, keeping the investigation continuous.

Cleaner Smartscape topology. Inactive Synthetic Locations no longer appear in Smartscape, keeping topology views focused on what’s live.

Smartscape navigation for database tables and indexes. Direct navigation intents let you jump from a database node to its table or index detail view in one click.

Filters stay with you. Automated filtering suggestions scope correctly to OR and AND conditions across all apps. Filter state, search terms, and highlights survive page reloads. HTTP Status Filter selections persist through navigation steps in Distributed Tracing.

DQL durations support decimals. Duration literals (h, m, s, ms, us, ns) now accept decimal numbers — for example, 0.5h or .2m. Note, however, that this doesn’t apply to calendar durations.

More headroom in Distributed Tracing. The log viewer no longer caps at 1,000 entries, with full deduplication across trace and span IDs. Span scan limits are configurable from settings (default 5,000, up to 10,000). Field naming is also cleaned up — Smartscape fields drop the redundant prefix, and classic ME fields are clearly labeled.

More allowlist entries for external requests. You can now add up to 100 allowlist entries, double the previous limit of 50, with existing entries preserved across all environments.

Affected entity names enriched in problem records. A new affected_entity_names array is now populated alongside the existing affected_entity_ids and affected_entity_types arrays, index-aligned across all three.

The Problems feed displaying affected entity names alongside IDs, enabling notification workflows and integrations to reference entities without a separate lookup.
The Problems feed displays affected entity names alongside IDs, enabling notification workflows and integrations to reference entities without a separate lookup.

This brings the 3rd-gen platform to parity with classic problem notifications and enables notification workflows and external integrations to reference entity names without additional lookup. The Problems app v1.27 reached General Availability on June 29.

Proactive Cost Intelligence across your entire stack

Dynatrace now makes it easier to understand costs, act before they spike, and optimize with less effort. New Optimize documentation walks Dynatrace Platform Subscription (DPS) customers through the full journey from understanding to optimizing costs, aligned with the FinOps Foundation framework.

Dynatrace Assist surfaces the root cause of a cost spike directly from billing usage events, without requiring specialist knowledge.
Dynatrace Assist surfaces the root cause of a cost spike directly from billing usage events, without requiring specialist knowledge.

The bigger shift is that Dynatrace Assist can now do the cost analysis work for you, designed to reduce the need for specialist expertise. You can ask it to:

  • Understand spikes — “I received a notification that costs have increased. Can you find anything notable?” returns the root cause along with a full drilldown into your billing_usage
  • Predict costs — “Based on my log ingest usage over the last 90 days, can you predict my usage for the next 30 days?” returns a capability-level forecast based on actual consumption, useful when onboarding new teams.
  • Optimize usage — “Are there any log queries duplicated by multiple users?” surfaces overlapping queries with concrete suggestions to improve them.

For more on building cost discipline into your observability practice, see Driving your FinOps strategy with observability best practices.

Why these changes matter

Taken together, the June releases make everyday investigation work feel less fragmented. You get more context in the places where teams already troubleshoot: a fuller Smartscape view, AI quality signals alongside performance data, log patterns that identify root causes faster, service maps for incident response, and security and cost insights that are easier to act on without switching tools or relying on specialists.

These are the kinds of changes that add up across a week of real work.

Check out all these updates in action on our Release Radar launchpad.

The post Dynatrace Release Radar 06.26 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-release-radar-06-26/feed/ 0
Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals https://www.dynatrace.com/news/blog/evaluate-llm-and-agent-quality-in-dynatrace-ai-observability/ https://www.dynatrace.com/news/blog/evaluate-llm-and-agent-quality-in-dynatrace-ai-observability/#respond Thu, 11 Jun 2026 19:44:47 +0000 https://www.dynatrace.com/news/?p=74476

AI applications fail in ways that differ from traditional software. They can return responses quickly, with no errors, and still deliver answers that are inaccurate, ungrounded, unsafe, or unusable. That's why AI quality can't be treated as a side project.

The post Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals appeared first on Dynatrace news.

]]>


For AI systems, reliability is defined by response quality, factual grounding, data security, and usability — and those signals need to live alongside the same observability data teams already trust to monitor performance and availability.

When evaluation scores are isolated in notebooks, spreadsheets, standalone tools, or CI logs, they’re hard to operationalize. By bringing AI quality metrics into Dynatrace AI Observability—next to latency, cost, errors, traces, and user behavior—teams can connect poor responses and hallucinations directly to the prompts, models, retrieval contexts, tool calls, services, and traces that produced them.

What is dt-evals?

dt-evals is an open source CLI for evaluating LLM and agent quality from real GenAI traces, agentic interactions. Teams can run online evaluations against live or recent interactions, score outputs with an LLM judge, and send structured results back to Dynatrace AI Observability so quality becomes visible, queryable, trendable, and actionable.

A minor prompt edit, model change, or retrieval update to an AI application can improve one behavior while quietly breaking another. The challenge to tracking down where and why these systems break is that evaluation results are often maintained outside the operational workflow, making it difficult to connect a low score to the exact trace, prompt, model version, retrieval context, tool call, or service that produced the unwanted behavior.

Dynatrace AI Observability closes this loop. With dt-evals and the AI Observability Evaluation Preview teams can pull recent gen_ai.*  spans, score real interactions with an LLM judge, and write structured evaluation results back as business events. These scores can be viewed with the originating trace, queried for custom analysis, trended in dashboards, and used to trigger alerts or workflow-driven remediation.

A failing faithfulness score is no longer just a number in a report. It’s now an operational signal.

What are LLM evaluations?

An LLM evaluation system scores an AI response against a range of quality and safety dimensions. Common examples include whether the answer is relevant to the question, faithful to the provided context, free of hallucinations, safe for users, complete enough to be useful, and resistant to prompt-injection attempts.

LLM evaluations are typically applied in two modes:

Offline evaluations run before release against a fixed test set or curated trace dataset. These are used to compare a proposed prompt, model, retriever, or agent-tool change against a known baseline before shipping. For example, replay 500 representative support questions in CI and block the release if faithfulness drops below the configured threshold.

Online evaluations run after deployment against sampled production or user traffic. Use online evaluations to detect regressions caused by live inputs, changing retrieval results, tool behavior, traffic mix, or model drift. For example, evaluate 10% of support-agent traces from the last hour and alert the team if hallucination failures exceed the configured window.

With dt-evals, you can run evaluations from the command line, use them in CI/CD, or schedule them to detect quality regressions autonomously after deployment as a post-processing quality gate for your AI agents and LLM output.

AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability
Figure 1. AI Evaluation & Agentic App Performance dashboard showing dt-evals results in Dynatrace AI Observability

Run evaluations from the command line

dt-evals is an open source evaluation toolkit for teams that want to bring their own data, judge provider, and evaluation logic while keeping traces, scores, dashboards, and alerts connected.

Install the CLI:

npm install -g @dynatrace-oss/dt-evals

Or run it directly with npx:

npx @dynatrace-oss/dt-evals <command>

A typical first run has three steps:

  1. Configure your environment and judge provider (Bring Your Own AI API key):

dt-evals configure

  1. Verify your local setup and connection:

dt-evals doctor

  1. Run evaluations on recent GenAI traces:

dt-evals run --since 1h --sample 10

This command evaluates traces from the last hour and samples 10% of them. In other words, dt-evals evaluates roughly one out of every ten matching traces, including the prompt and completion messages associated with each selected trace.

During configuration, you provide the connection to your Dynatrace environment and the credentials for the LLM judge provider you want to use. dt-evals does not require teams to send evaluations through a fixed provider. You bring your own judge credentials and control where evaluation execution happens.

For CI/CD use cases, run in CI mode:

dt-evals run --since 6h –ci

In CI mode, dt-evals emits machine-readable output and can fail the pipeline when a configured threshold is breached. This makes quality checks part of the same delivery process used for prompt changes, model upgrades, retrieval updates, and agent releases.

dt-evals in action
Video 1. dt-evals in action

Bring your own LLM judge provider

The “LLM-as-judge” evaluation approach involves using an AI model to score another model or agent response. The LLM judge needs to come from an AI provider your team trusts and has approved for the type of data being evaluated.

dt-evals supports common LLM judge AI models and inference providers, including OpenAI, Anthropic, Google/Vertex/Gemini, AWS Bedrock, and Azure OpenAI. Depending on the package and configuration path you use. Teams provide their own credentials, choose the judge model, and can tune execution settings such as thresholds and concurrency.

This matters for both governance and cost control. Teams can decide which LLM provider is assigned to evaluate which traffic, how many judge calls run in parallel, and where evaluation results are stored.

Which quality and safety dimensions are evaluated by dt-evals?

dt-evals supports built-in LLM-as-Judge evaluators for a range of quality and safety dimensions, including:

  • Relevance: Does the response answer the user’s question?
  • Faithfulness: Is the response supported by the provided context?
  • Hallucination: Does the response invent facts that are not present in the available context?
  • Answer completeness: Does the response fully address the user’s request?
  • Context relevance: Is the retrieved or supplied context useful for answering the question?
  • Factual accuracy: Does the response match an expected or known-correct answer?
  • Summarization quality: Does the summary preserve the important information?
  • Conciseness: Is the response direct with no unnecessary detail?
  • Fluency: Is the response clear and readable?
  • Toxicity: Does the response contain harmful or abusive content?
  • Bias: Does the response show unfair or inappropriate bias?
  • PII leakage: Does the response expose sensitive personal information?
  • Prompt injection: Did the input or response show signs of instruction manipulation?
  • User frustration: Does the interaction suggest the user is blocked or dissatisfied?
  • Drift: Are scores changing meaningfully compared with prior behavior?

A Retrieval Augmented Generation (RAG) application might focus on faithfulness, hallucination, context relevance, and answer completeness. A customer-facing support agent might focus on relevance, fluency, bias, toxicity, and prompt-injection risk. An internal assistant might add custom checks for tone, policy compliance, or whether the answer includes required next steps.

Add custom evaluations

Built-in metrics are useful, but most production AI systems also need checks that are specific to the business, domain, or workflow.

Custom evaluations let teams define their own judge prompts, scoring rules, labels, and thresholds. For example, a support team can create a custom evaluator that checks whether an answer includes a required troubleshooting step before recommending escalation. A financial services team can check whether responses include the required disclaimers. A platform team can check whether an agent uses the correct tool before answering.

The critical point is that custom evaluators run through the same pipeline as built-in evaluators. They can produce the same structured results, appear alongside other scores, and be used in dashboards, alerts, and release checks.

A typical configuration defines the target service, judge provider, sampling strategy, enabled metrics, and thresholds:

schemaVersion: 1
name: support-agent-prod

dynatrace:
  environmentUrl: https://your-env.apps.dynatrace.com
  platformToken: dt0s16.xxxxx

judge:
  provider: openai
  model: gpt-5.5

scope:
  service: support-agent
  since: 1h
  sampling:
    strategy: random
    percent: 10

metrics:
  enabled:
    - faithfulness
    - hallucination
    - relevance
    - drift

alerts:
  thresholds:
    faithfulness: 0.7
    relevance: 0.7

Evaluation results in the AI Observability app

Evaluation results appear directly in the AI Observability app, so teams don’t have to jump between a trace view, an eval report, and a separate dashboard to understand what happened.

In the Prompts view, teams can filter for prompts with evaluation scores and inspect row-level verdicts. Score badges such as relevance, fluency, bias, faithfulness, or toxicity make response quality easy to scan without opening every trace.

This is useful when triaging a regression. Instead of starting with a generic failure count, teams can quickly see which prompts failed, which evaluator failed them, and whether the issue is isolated or widespread.

AI Observability App Prompts stream with evaluation results
Figure 2: AI Observability App Prompts stream with evaluation results

From an individual prompt or trace, the Evaluations tab shows run-level details, including the evaluation name, score, provider, judge model, method, and supporting metadata.

That detail matters because a failed score is only useful if teams can explain it. Engineers and evaluation owners can move from a low score to the exact prompt, response, trace, model, evaluator, and rationale that produced it.

Prompt detail view with evaluation results and trace context in Dynatrace AI Observability
Figure 3:  Prompt detail view with evaluation results and trace context in Dynatrace AI Observability

Query, trend, and alert on evaluation scores

Because dt-evals writes results back as structured events, evaluation scores can be analyzed with the rest of your telemetry.

Teams can ask questions such as:

  • Which evaluator has the lowest average score?
  • Which services are producing the most failed evaluations?
  • Did quality drop after a model or prompt change?
  • Are hallucinations increasing over time?
  • Is quality improving at the cost of latency or token usage?

For example, to get average score by evaluator you could write this query:

fetch bizevents
| filter event.type == "gen_ai.evaluation.result"
| summarize avg_score = avg(gen_ai.evaluation.score.value),
    by: { gen_ai.evaluation.name }
| sort avg_score asc 
Querying failed evaluations by service and evaluator in Dynatrace AI Observability
Figure 4: Querying failed evaluations by service and evaluator in Dynatrace AI Observability

Failed evaluations by service and metric can be determined with this query:

fetch bizevents
| filter event.type == "gen_ai.evaluation.result"
| filter gen_ai.evaluation.score.label == "fail"
| summarize failures = count(),
    by: { dt.service.name, gen_ai.evaluation.name }
| sort failures desc 
Average evaluation scores by evaluator in Dynatrace AI Observability
Figure 5: Average evaluation scores by evaluator in Dynatrace AI Observability

Trending is where evaluation data becomes more useful than a point-in-time report. A single failed score can show an issue. A trend can show whether quality is drifting slowly, whether a release caused a sudden drop, or whether a fix actually improved behavior over time.

On the AI Evaluation & LLM App Performance dashboard, teams can track trends in quality score, pass rate, failed evaluations, drift detections, evaluator health, run cadence, and pass/fail volume over time.

AI Evaluation &amp; LLM App Performance dashboard
Video 2: AI Evaluation & LLM App Performance dashboard

How to turn quality regressions into alerts

Evaluation results can also drive alerts. For example, a support agent team may want to notify the AI team when hallucinations appear in production, or when faithfulness drops for more than a few minutes.

name: support-agent-prod

alerts:
  notifications:
    - name: hallucination-detected
      metric: hallucination
      condition: count > 0
      window: 5m
      channel:
        type: slack
        connection: ai-observability-slack
        channel: "#ai-alerts"

    - name: faithfulness-regression
      metric: faithfulness
      condition: fail_rate > 10%
      window: 15m
      channel:
        type: email
        connection: ai-team-email
        to: [ai-team@example.com] 

Deploy the alerts with:

dt-evals alerts list ./support-agent-prod.yaml
dt-evals alerts apply ./support-agent-prod.yaml

Once applied, Dynatrace runs these checks continuously as Workflows. If hallucinations appear in the last five minutes, the team gets a Slack alert. If more than 10% of faithfulness checks fail over 15 minutes, the AI team receives an email. This turns LLM quality from something teams inspect manually into something Dynatrace can monitor and route automatically.

For continuous alerting, evaluation runs need to happen continuously or on a schedule. You can run dt-evals in CI for release checks (see our example here), run it manually during investigation, or deploy a scheduled runner for ongoing production evaluation. An alert is only as fresh as the evaluation results it carries.

Once configured, quality signals can be routed to the teams that need to act. If hallucinations appear in the last five minutes, the team can receive a Slack alert. If more than 10% of faithfulness checks fail over 15 minutes, the AI team can receive an email. This turns LLM quality from something teams inspect manually into something they can monitor and route automatically.

Close the loop in the AI software delivery lifecycle

Evaluation gates are most useful when they meet developers where they already work. Because dt-evals writes evaluation results back into the observability data layer, those results are not limited to dashboards or post-release reviews. They can be queried, inspected, and acted on from development workflows, CI/CD pipelines, and agentic coding environments.

For example, a team can run dt-evals after any change to a prompt, model, retriever, or agent tool, and then use dtctl (Dynatrace CLI tool for AI Agents) to query the resulting evaluation data, inspect related traces, review dashboards, or validate whether a release threshold was met. In an AI-assisted workflow, tools such as Claude Code, Cursor, GitHub Copilot, or an internal agent harness can leverage MCP or CLI access to bring that same observability context into the developer’s daily workflow.

That closes the loop of the AI software delivery lifecycle: teams can evaluate behavior, control rollout decisions, remediate regressions, and feed production learning back into the next development cycle. Quality signals are no longer in a separate report; they’ve become a part of how AI software is built, shipped, and operated.

Bring evaluations into the release process

Evaluation support is not just for inspection after something breaks. It can also help prevent regressions before they reach users.

Overview of a typical release workflow for an AI app with dt-evals
Figure 6: Overview of a typical release workflow for an AI app with dt-evals

This makes AI quality part of the release process. Teams can gate changes based on relevance, faithfulness, hallucination risk, prompt-injection risk, toxicity, or custom metrics, rather than relying solely on latency and error rate.

What makes this meaningful is that it’s the same pipeline teams already run. AI quality gates sit alongside the latency, error rate, and SLO gates teams have been using for years. There’s no second CI system, no second platform to learn, no second dashboard to monitor. Quality becomes one more dimension of the release decision, gated the same way performance is gated, by the same platform, in the same pipeline.

Coming next

The current experience makes evaluation results visible and actionable inside Dynatrace AI Observability. Next, the focus is on making evaluation workflows easier to run at scale and easier to compare across changes.

Planned improvements include targeted and bulk trace evaluations, custom evaluation libraries, evaluator versioning and lineage, baseline comparisons, experiment views, native quality gates, and deeper visibility into online evaluations.

These capabilities will help teams compare prompt and model variants, understand quality versus cost and latency tradeoffs, and detect sustained quality regressions before they affect more users.

Start today

To get started, check out the Git repository. You’ll need:

  • Node.js 20 or later
  • A Dynatrace environment with GenAI spans and the AI Observability app installed
  • Credentials for the judge provider you want to use
  • A service, trace sample, or CI workflow you want to evaluate

Then, install the CLI:

npm install -g @dynatrace-oss/dt-evals

Configure your service and judge provider:

dt-evals configure

Run your first evaluation:

dt-evals run --since 1h --sample 10

With Dynatrace AI Observability and dt-evals, teams can bring LLM and agent evaluations into the operational loop, where they can trace behavior, score outputs, trend results, alert on regressions, and gate releases before silent failures reach production.

The post Evaluate LLM and agent quality in Dynatrace AI Observability with dt-evals appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/evaluate-llm-and-agent-quality-in-dynatrace-ai-observability/feed/ 0
Dynatrace Release Radar 04.26 https://www.dynatrace.com/news/blog/dynatrace-release-radar-04-26-whats-new-and-why-it-matters/ https://www.dynatrace.com/news/blog/dynatrace-release-radar-04-26-whats-new-and-why-it-matters/#respond Wed, 27 May 2026 17:07:44 +0000 https://www.dynatrace.com/news/?p=74179 Release Radar

This series covers recent Dynatrace releases and updates, focusing on what’s new, what’s changed, and how these recent enhancements can benefit you and your organization. Each post covers newly available capabilities and points you toward where to explore them.

The post Dynatrace Release Radar 04.26 appeared first on Dynatrace news.

]]>
Release Radar

The April 2026 Dynatrace SaaS releases bring six updates aimed at a familiar problem: too much manual work between a signal and an answer. The updates focus on native cloud visibility, deeper Kubernetes insights, consistent severity handling, faster investigations, and smoother analytics work.

Explore all updates hands-on in the Release Radar launchpad.

A native Azure experience in Clouds

What does Dynatrace add for Azure practitioners?

Dynatrace now extends the enhanced Clouds experience to Microsoft Azure, putting Azure subscriptions on the same footing as AWS. Metrics, logs, metadata, and topology now sit in one managed view, so Azure teams can move from inventory to investigation without stitching the picture together by hand.

What you get out of the box

  • Opinionated insights and ready-made dashboards built from enriched Azure telemetry, so investigations start with answers instead of a blank canvas.
  • Pre-configured health alerts for Azure, created and managed directly in Clouds, with drill-down, search, and filtering directly in the Clouds app.
  • Broad metric coverage for any Azure Monitor native platform metric across Azure services.
  • A rich Azure topology inventory that periodically scans Azure environments and enriches resources with native metadata such as tags and subscription IDs, all queryable with Dynatrace Query Language (DQL).
  • Simple onboarding and lifecycle management that turns Azure subscriptions into native Dynatrace connections and manages them centrally.

Dynatrace also adds drilldowns from cloud entities into the relevant Dynatrace experiences, so teams can keep moving instead of bouncing between cloud and platform views.

The new Clouds experience for Azure lets you optimize cloud operations at scale.
The new Clouds experience for Azure lets you optimize cloud operations at scale.

Kubernetes visibility for autoscaling and custom resources

What’s new in Kubernetes observability?

Dynatrace extends Kubernetes visibility to two additional object types that SREs and platform teams rely on daily: Horizontal Pod Autoscalers (HPA) and Custom Resources (CRs).

Horizontal Pod Autoscaler as a first-class object

HPA is now a first-class object in enhanced Kubernetes visibility. You can see when scaling kicked in, what triggered it, and how desired and actual replica counts lined up next to the workloads involved.

Custom Resource insights

You can monitor up to five Custom Resources per cluster, surfaced the same way as built-in Kubernetes objects. This brings CRD-heavy ecosystems such as Argo, Istio, Cert-Manager, Kyverno, and operator-managed databases into the same investigation scope as the rest of your cluster.

For clusters connected through cloud integrations, the Kubernetes cluster details page now exposes the underlying cloud configuration (EKS, AKS, or GKE) in YAML or JSON, making cloud-side and cluster-side state accessible in one place.

HorizontalPodAutoscaler visibility in the Kubernetes app experience.
HorizontalPodAutoscaler visibility in the Kubernetes app experience.

A unified severity model for alerts and problems

What is event.severity in Dynatrace?

Dynatrace introduces a standardized event.severity field for alerts and problems, aligned with the ITIL Incident Management framework. Severity is stored in Grail as an integer from 1 (Critical) to 5 (Informational) and is shown as a human-readable label across the platform.

Severity levels at a glance

Value Label Description
1 Critical Major business disruption; service outage
2 Major Significant impact; workaround may exist
3 Minor Limited or non-critical impact
4 Warning Low impact; no business disruption
5 Informational No business impact

Severity automatically propagates from correlated alerts to the parent problem, with the highest severity always taking precedence. This gives teams one severity model to filter on, route with, and escalate from.

You can now:

  • Filter the problem feed by severity
  • Display a severity column with visual icons in problem lists
  • Use severity as a condition in Workflows for alert routing and notifications
Event severity in the Problems app experience.
Event severity in the Problems app experience.

Faster Investigations with Smartscape navigation

What changed in Smartscape?

Smartscape now offers all six ready-made views, such as vertical topology, horizontal topology, and visual resolution path, just a click away in a persistent side panel. You no longer need to return to the landing page in the middle of an investigation.

The new Recent views section shows your latest investigations, making it easy to reopen them, compare them, and keep working as you test a root-cause hypothesis.

The result is less backtracking in the middle of an incident.

The new sidebar navigation in Smartscape
The new sidebar navigation in Smartscape

Dashboards and notebooks: productivity improvements

What’s new for dashboard authors and analysts?

The latest release adds several practical upgrades for team members who build dashboards and work in notebooks every day.

  • Treemap visualization for identifying dominant categories in hierarchical data, such as requests per service by Kubernetes namespace.
    Treemap visualization example
    Treemap visualization example
  • Dashboard variables for dynamic coloring and thresholds, so visual conditions stay in sync with environment or team selectors.
    Use dashboard variables for dynamic coloring and threshold conditions
    Use dashboard variables for dynamic coloring and threshold conditions
  • Centralized tile indicator controls, allowing you to show or hide warnings, descriptions, and custom timeframes at the dashboard level.
    Select or clear tile indicators on a dashboard
    Select or clear tile indicators on a dashboard
  • Direct image upload in Markdown using a built-in image library shared across Dashboards, Notebooks, and the Launcher.
    Upload image directly in Markdown
    Upload image directly in Markdown
  • Row marker coloring for tables, making it easier to visually group related rows without sacrificing readability.
    Highlight table rows with color markers in Dashboards and Notebooks
    Highlight table rows with color markers in Dashboards and Notebooks

User experience improvements

Why does the platform feel faster?

This release smooths out the path from the first symptom to root cause analysis. Tracing and services workflows now handle high-span traces more reliably, show timing more clearly, and surface useful sample traces earlier.

Table-first workflows also benefit from richer entity-detail tables, better filtering, and clearer structure, helping teams answer more questions without switching views. Navigation patterns, overlays, and error messaging are now more consistent across the platform, reducing mental overhead when time is tight.

Explorer new table experience with entity details, alerts, and schema links
New Explorer table experience with entity details, alerts, and schema links

Why these updates matter

Taken together, these updates eliminate inefficiencies in the work that teams do every day. Cloud operations teams, Kubernetes SREs, on-call engineers, and analytics authors get richer context, faster paths to answers, and simpler ways to share what they find. This is where Dynatrace earns its keep under pressure.

Explore the updates live in the Release Radar launchpad.

The post Dynatrace Release Radar 04.26 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-release-radar-04-26-whats-new-and-why-it-matters/feed/ 0
Dynatrace Release Radar 01.26 https://www.dynatrace.com/news/blog/dynatrace-release-radar-2026/ https://www.dynatrace.com/news/blog/dynatrace-release-radar-2026/#respond Mon, 02 Mar 2026 17:29:50 +0000 https://www.dynatrace.com/news/?p=73224 Release Radar

This series covers recent Dynatrace releases and updates, focusing on what’s new, what changed, and how it applies to you and your organization. Each post outlines newly available capabilities and points to places where you can explore them directly, helping you understand what’s relevant and what to look at next.

The post Dynatrace Release Radar 01.26 appeared first on Dynatrace news.

]]>
Release Radar

We kicked off the new year with our annual customer event, Dynatrace Perform, and many new announcements. If you weren’t able to join us in person, you can watch all the mainstage keynotes, innovation sessions, and breakouts on demand on the Dynatrace Perform 2026 webpage.

In this blog, we’ll focus on brand-new product enhancements that accelerate service troubleshooting, provide richer cloud context for AWS, and deliver meaningful improvements that reduce friction in daily workflows.

If you want to jump straight to our curated sandbox environment for the capabilities mentioned below, head over to our dedicated playground launchpad.

Dynatrace Intelligence

Our biggest news is that Dynatrace Intelligence is now available. It’s the industry’s first agentic operations system that effectively fuses deterministic insights with agentic action to deliver reliable outcomes with autonomous prevention, remediation, and optimization at scale.

Dynatrace Intelligence Marketecture

Here are the new features and capabilities now available in Dynatrace Intelligence:

  • Dynatrace MCP Server: In addition to the local MCP server that was launched in May 2025, our remote MCP server is now generally available.
  • Dynatrace Assist: The evolution of Davis CoPilot puts Dynatrace Intelligence at your fingertips. Dynatrace Assist pulls context from Grail, maps relationships using Smartscape – our real-time dependency graph – and collaborates autonomously with Dynatrace agents using the tools provided by the Dynatrace MCP server.
  • Agentic ecosystem: Whether you aim to level up collaborative operations with SRE agents, enjoy closed-loop autonomous operations with ITSM agents, or want AI-powered code repair with developer and coding agents, we’ve got you covered. Maximize the value of your tool landscape by leveraging our agentic integrations.
  • Agentic workflows: Turn your workflows into agentic automations leveraging Dynatrace Intelligence. This program is currently available in a Preview program.

Smartscape: Real-time dependency graph

The new Smartscape experience delivers a real-time dependency graph that helps practitioners move from “watching signals” to understanding true entity health and cause-and-effect across fast-changing cloud, Kubernetes, on-premises, and hybrid environments. It adds major new capabilities:

  • An all-new Smartscape app with powerful visual analytics and domain-specific views,
  • Fully native cloud entities with complete metadata (including raw cloud/Kubernetes object JSON),
  • and agentless cloud data ingest for automatic enrichment of dependencies and policy context.

This allows teams to diagnose faster, reduce MTTR, and make better architectural and operational decisions with real production context.

Smartscape dashboard

Smartscape enables exploration across millions of relationships, strengthens incident collaboration via in-context workflows like Visual Resolution Path and “view topology” actions, and improves security and governance by visualizing exposure and attack paths with real blast radius and enriched cloud native semantics like tags, ownership, cost centers, and compliance attributes.

The new Smartscape extends far beyond a standard topology; it offers domain-specific views tailored to the unique requirements of your use cases—whether application performance, cloud infrastructure, or essential business services. The following pre-configured views are now available.

  • Smartscape on Grail: discover all entities and relationships in your environment
  • Infrastructure overview: gain insights into which components are running and how they’re connected
  • Service dependency graph: see how your services are connected
  • Problem graph: understand problem impact and blast radius
  • Kubernetes overview: map your Kubernetes environment, from clusters to components
  • AWS EC2 ecosystem overview: understand your entire EC2 ecosystem and resource relationships

Have a look at our recent Smartscape blog post to learn how these enhanced views help solve real-world challenges.

Cloud Operations for AWS

Dynatrace enhanced Cloud Platform Operations expands AI-powered observability into an operations-first experience for practitioners (cloud ops, SRE, and platform teams) by unifying cloud metrics, logs, and events across AWS, Azure (see Preview program), and Google Cloud (see Preview program) in a single platform, enriched with topology-aware context for faster troubleshooting and safer automation. It introduces:

  • fully managed cloud connections with a guided wizard (no extra infrastructure),
  • expanded ingest that captures more cloud service metrics plus richer cloud events (including hyperscaler-native security alerts),
  • and automatic reuse of existing cloud tags to drive access control, ownership, cost allocation, alert routing, and preventive workflows—so teams can move from fragmented signals to clear, actionable answers at enterprise scale.

Dynatrace Dashboards

This allows users to shift from reactive monitoring to proactive cloud operations built around three outcomes: prevention (predict anomalies and trigger workflows before user impact), remediation (AI-driven RCA plus self-healing automation to cut resolution time), and optimization (continuous cost and performance efficiency via real-time insights and recommendations).

For platform teams, the big win is operational simplicity: the onboarding flow is GitOps-ready and removes the need to maintain ActiveGates for CloudWatch ingest on this path. For practitioners, the win is troubleshooting speed: reimagined exploration, resource-rich metadata, and opinionated insights reduce the time from “something’s wrong” to “here’s why.”

Real User Monitoring experience

The new Real User Monitoring (RUM) experience adds modern frontend signals that match how today’s web and mobile apps behave—for example, soft navigation for Single Page Apps (SPA), user interactions (clicks/taps/scrolls), and background requests—alongside Core Web Vitals and key mobile performance signals (including troubleshooting enhancements like application not responding and symbolication). Out of the box, teams get task-focused workflows and dashboards that connect frontend symptoms to backend reality, so you can pinpoint what’s slow or broken and shorten the path from user complaint to verified cause and fix.

Dynatrace Real User Monitoring (RUM) experience

Achieve faster validation of real user impact and clearer prioritization: Users & Sessions grounds investigations in actual sessions, Error Inspector groups and prioritizes errors with the right context, and Experience Vitals helps identify which requests/assets drive slowdowns using redesigned analysis views—so teams can reduce friction, resolve complaints with confidence, and connect experience trends to business outcomes via custom dashboards, notebooks, and DQL exploration, with built-in privacy/permission controls, and optional extended retention for deeper historical analysis (currently available in a Preview program).

AI observability

Dynatrace has expanded agentic AI observability with a broader framework and protocol support, so teams can build, run, and debug autonomous agent systems with confidence across AWS, Azure, and Google Cloud. Support now includes popular agentic ecosystems such as Amazon Bedrock AgentCore, Amazon Bedrock Strands, LangChain Agents, Google Agent Development Kit (ADK), OpenAI Agents SDK, and Model Context Protocol (MCP)—with signals unified via OpenTelemetry and OpenLLMetry into a single correlated observability model for end-to-end visibility across agents, tools, models, and dependencies.

Agent topology visualizes agent execution flows, showing how they interact with one another.
Video: Agent topology visualizes agent execution flows, showing how they interact with one another.

Alongside this expanded support, the new AI Observability app delivers a purpose-built experience to observe AI workloads end-to-end—from agents and LLMs to orchestration layers and tools—so practitioners can validate changes faster, reduce risk, and ship AI features at scale. Key capabilities include end-to-end monitoring of agent interactions and tool usage, prompt/tool/model tracing and debugging across multi-step flows, cost visibility (token consumption, cost trends, caching impact), actionable dashboards and drill-downs (including faster validation via A/B testing across model/prompt variants), and enterprise-grade security, privacy, and governance views such as surfaced guardrail outcomes for auditability and trend monitoring.

Investigations: Transform how practitioners derive actionable insights

The Investigations app provides a central starting point for exploring analytical insights across Grail data. It gives practitioners immediate access to essential investigation capabilities—such as analyzing large DQL results, pivoting queries based on metadata, reviewing investigation history, and connecting logs, metrics, events, and traces—helping practitioners quickly uncover root causes and accelerate complex investigations.

Dynatrace investigations

Improved Dashboards experience

We’ve enhanced several ready-made dashboards that improve your dashboard experience and make insights clearer, faster, and more consistent. You can duplicate and adapt them to kick-start your own dashboards.

  • The Getting started dashboard demonstrates the major types of visualizations you can use and provides example tiles and layouts.
    Dynatrace Dashboards
  • The Page performance & errors dashboard serves as a starting point for investigating page performance and web front-end navigation. It surfaces the most important web performance and reliability KPIs at a glance, highlighting key metrics such as page load time, error count, navigations, LCP, INP, and CLS.
    Dynatrace Dashboards
  • The XHR & fetch performance dashboard includes core KPIs such as request duration, time to first byte (TTFB), and fetch failure rate. These help you quickly spot slow or failing back-end calls that affect the user experience.
    Dynatrace Dashboards

Where to start this week

We encourage you to take advantage of all the efficiencies and insights these new Dynatrace capabilities provide. Depending on your role, here are the recommended next steps for SREs, Cloud Owners, and Development teams seeking faster service troubleshooting loops, richer AWS cloud context, and other meaningful improvements that reduce friction in their daily workflows.

Get started: SREs

  1. Start with Dynatrace Intelligence for faster incident loops
    1. Open Dynatrace Assist during an active issue to pull context from Grail and map relationships via Smartscape, then let it collaborate with Dynatrace agents/tools (via MCP) to accelerate triage and next steps.
    2. If you use chat/agent tooling internally, connect via the Dynatrace MCP Server (remote if you want centralized access) to make Dynatrace context available in your agentic workflows.
  1. Make Smartscape your default “blast-radius + causality” view
    1. Use the new Smartscape app and Visual Resolution Path/view topology actions to validate true upstream/downstream impact and shorten MTTR.
    2. Leverage native cloud/Kubernetes metadata (including raw object JSON) to quickly confirm “what changed” vs. “what broke.”
  1. Automate closure with agentic workflows (Preview program)
    1. Convert recurring remediation steps into agentic automations by combining Dynatrace Intelligence with Workflows for closed-loop operations (start with a high-confidence, low-risk runbook).

Get started: Cloud owners

  1. Onboard AWS with enhanced Cloud Operations first
    1. Use the fully managed cloud connection and guided wizard to bring in unified metrics, logs, and events with richer AWS context, without maintaining ActiveGates for CloudWatch ingest on this path.
    2. Ensure your cloud tags are clean and meaningful, because they’ll automatically drive ownership, access control, cost allocation, and alert routing.
  1. Operationalize outcomes: prevention, remediation, and optimization
    1. Set up alerting and dashboards around the three outcomes:
    2. Prevention: anomaly prediction + proactive workflows
    3. Remediation: AI-driven RCA + self-healing actions
    4. Optimization: continuous cost/performance efficiency using context-rich insights
    5. Use Smartscape to validate dependencies and impacts across accounts, regions, clusters, and services.
  1. Plan for multicloud setup
    1. If you’re also on Azure or GCP, use what you learn on AWS to establish a standard operating model, then extend to preview programs when ready.

Get started: Development teams

  1. Start from user impact with the new RUM experience
    1. Use Users and Sessions to reproduce issues from real sessions, then jump to Error Inspector and Experience Vitals to identify which requests/assets/interactions drive pain.
    2. For SPAs and modern apps, validate soft navigation, user interactions, and background requests alongside Core Web Vitals to quickly pinpoint frontend bottlenecks.
  1. Connect frontend symptoms to backend issues
    1. From a slow, erroring session, follow the workflow to backend services and dependencies (Smartscape helps confirm causality), shortening the path from complaints to verified root causes.
  1. If you ship AI features, instrument them with AI Observability
    1. Adopt the AI Observability app for end-to-end tracing across agents, tools, and models; use cost visibility and A/B validation to safely iterate on prompts/models.
    2. Standardize telemetry via OpenTelemetry and OpenLLMetry, and if you use agent frameworks (LangChain Agents, OpenAI Agents SDK, Google ADK, Bedrock, or MCP), start by observing one representative production flow before scaling coverage.

The post Dynatrace Release Radar 01.26 appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-release-radar-2026/feed/ 0
The new Dynatrace Smartscape improves operational efficiency across clouds, Kubernetes, infrastructure, and more https://www.dynatrace.com/news/blog/the-new-dynatrace-smartscape-improves-operational-efficiency-across-clouds-kubernetes-infrastructure-and-more/ https://www.dynatrace.com/news/blog/the-new-dynatrace-smartscape-improves-operational-efficiency-across-clouds-kubernetes-infrastructure-and-more/#respond Wed, 18 Feb 2026 19:25:33 +0000 https://www.dynatrace.com/news/?p=73081 Smartscape graphic

The new Smartscape® real-time dependency graph gives teams a real‑time understanding of how their entire digital environment works. By unifying cloud resources, Kubernetes objects, services, and infrastructure into a single live topology, Smartscape removes the guesswork from operations. With a continuously updated view of production, enriched with full metadata and knowledge of all dependencies, teams can explore their environments visually in domain‑specific Smartscape views or analytically through the Grail® unified data lakehouse. In this blog, we highlight concrete new use cases across modern cloud‑native systems.

The post The new Dynatrace Smartscape improves operational efficiency across clouds, Kubernetes, infrastructure, and more appeared first on Dynatrace news.

]]>
Smartscape graphic

Unify cloud resources across accounts, regions, and services into a single, real-time dependency graph

As workloads continue to sprawl across AWS, Azure, Google Cloud, and on-premises data centers, teams are overwhelmed by massive volumes of telemetry and constant change. Simple questions like “What service depends on this?” or “Is this vulnerability exposed?” often turn into hours of manual investigation. Smartscape changes this dynamic by unifying every cloud asset, metadata field, and connectivity path into a single, real-time dependency graph, delivering instant answers and visualizing them in a continuously updated Smartscape view. Instead of hopping between AWS and Azure consoles, platform teams finally get a continuously updated picture of how their cloud environments are truly behaving.

Navigate across the AWS EC2 ecosystem view to instantly understand problems and their impact.
Figure 1. Navigate across the AWS EC2 ecosystem view to instantly understand problems and their impact. (video)

With new AWS integrations, Smartscape now also captures deep configuration data, such as VPCs, load balancers, security groups, subnets, network services, and compute metadata, and models these dependencies as native cloud entities. Unlike any other observability vendor, Dynatrace provides full access to the raw observability data in Grail via Dynatrace Query Language (DQL), unlocking powerful exploratory analytics use cases. Each entity includes the complete unprocessed definition of the cloud service as JSON, covering metadata, resources, configuration, security, and networking details, and tags, making this information fully transparent and directly queryable. This unified model delivers immediate customer value:

  • Security posture and exposure analysis: detect publicly reachable endpoints, analyze real security group and network policy paths, and prioritize fixes based on true blast radius and reachability.
  • IAM hygiene and drift control: uncover risky role sharing across Lambdas, identify configuration drift across accounts and regions, and validate whether access paths reflect intended policy.
  • Cost optimization: identify x86 vs ARM workloads, right-size EC2, RDS, and EBS based on real utilization, and connect cloud spend to actual service dependencies to make safer cost decisions.
  • Architecture & multi-account visibility: map cross-VPC and cross-region dependencies, unify runtime topology across all cloud accounts, and eliminate hidden or forgotten resources.
  • Operational readiness & risk reduction: understand how misconfigurations or outages propagate through infrastructure and into applications, improving impact assessment and response.

The Clouds app provides comprehensive insights and metadata, including metrics and logs for your services, deep insights into resource configurations and cloud topology, and the ability to leverage your cloud tags for access and visibility. With Clouds, teams can interactively explore and analyze their cloud estate, apply segment filters, follow connectivity paths, and compare environments.

The new Clouds app shows unified cloud resource details with configuration context.
Figure 2. The new Clouds app shows unified cloud resource details with configuration context.

Understand your entire setup at a glance through advanced visual analytics

The new Smartscape app’s domain-specific views turn complex, multi-layered cloud estates into something teams can understand instantly. Visual exploration makes it easier to:

  • Understand real, observed connectivity between workloads across VPCs and environments, enriched with cloud networking context such as subnets and security constructs.
  • Instantly understand problems and their blast radius with affected entities clearly highlighted.
  • Identify hidden relationships or unintended dependencies that spreadsheets or lists will never surface.
  • Validate migration plans, architectural assumptions, and segmentation strategies before changes go live.

Create a single source of production truth with flexible views and segmentation across cloud dimensions, including tags, accounts, regions, environments, and ownership.

This visual context is often where the “aha” moments happen, the point where teams finally see how their cloud is structured, where risks live, and where optimizations will have the greatest impact.

Smartscape visualizes a multicloud setup.
Figure 3. Smartscape visualizes a multicloud setup.

Utilize DQL for advanced insights customized and enriched with what matters to you

For deeper investigation or automation, DQL lets teams query relationships, join topology with logs and metrics, and run impact assessments programmatically. These queries can be operationalized through dashboards and notebooks. Learn more about how to utilize the new Smartscape DQL commands to query the AWS topology.

Use DQL to query all EC2 instances registered with a given Load Balancer's target group.
Figure 4. Use DQL to query all EC2 instances registered with a given Load Balancer’s target group.

Kubernetes: how Smartscape gives you clarity on fast-moving, complex clusters

Kubernetes environments evolve continuously: pods appear and disappear within seconds, configurations drift, and a single missing reference in a YAML file can cascade into service failures across namespaces, or even clusters. While traditional tools expose fragments of this reality, they fall short when teams need complete answers to foundational questions like what does this depend on?, what changed?, or why did this break?

Smartscape further enhances Dynatrace Kubernetes observability by unifying Kubernetes objects, relationships, and configurations across clusters and clouds into a single, real‑time dependency graph. Instead of jumping between kubectl commands, point‑in‑time UIs, and disconnected dashboards, teams gain a continuously updated, system‑level view of how their Kubernetes environments actually behave.

With enhanced ingest, Smartscape now captures all major Kubernetes object types, including ConfigMaps, Secrets, Ingress, PV/PVC, workloads, services, and namespaces, and stores their full YAML definitions and metadata directly in Grail. Teams can query configurations across clusters and clouds, trace live end-to-end dependency paths, and automatically surface misconfigurations, missing references, policy violations, and drift. What was previously scattered across files and tools becomes instantly explorable context, at a global scale. The value of Smartscape can be felt immediately:

  • Faster troubleshooting: trace live relationships across clusters, namespaces, workloads, and services to pinpoint drift or misconfigurations that cause runtime failures.
  • YAML misconfiguration detection: identify missing references, invalid fields, or policy violations with full YAML-in-context, and regenerate correct configurations using Dynatrace Intelligence.
  • Ephemeral awareness: retain visibility into short-lived workload changes or crashes that normally disappear before engineers can inspect them.
  • Policy and compliance enforcement: check networking, storage, config maps, resource quotas, and image standards at the object level for stronger governance.
  • Safer releases: segment clusters by team or namespace and visualize impact paths before and after deployments to reduce risk and improve deployment confidence.

All enhanced Kubernetes insights and YAML definitions are directly accessible within the Kubernetes app.

In Smartscape, access the Kubernetes domain view, where you can:

  • Visualize cluster topology for instant clarity on structure and relationships.
  • Follow real dependency chains across namespaces, workloads, services, and underlying infrastructure to understand impact paths.
  • Segment clusters dynamically by team, namespace, environment, or workload identity for precise context.
  • Isolate critical workloads or namespaces for focused investigation and remediation.
  • Validate architectural assumptions by comparing expected versus actual relationships.
Vertical topology for Kubernetes.
Figure 5. Vertical topology for Kubernetes.

For advanced analytics, DQL lets you query Kubernetes objects, relationships, and signals at scale. For actual use cases and examples, check out this notebook on the Dynatrace Playground.

Use the DQL traverse command to see which Kubernetes deployments communicate with each other. (video)
Figure 6. Use the DQL traverse command to see which Kubernetes deployments communicate with each other. (video)

Other domain-specific enhancements, from infrastructure to services

The new Smartscape unlocks a broader range of high-impact use cases across every layer of your IT environment, with topology-enriched information across apps; many new ways to explore your data via DQL, and several additional, use-case-optimized Smartscape views. Below are some additional examples and inspiration to help you get started:

Services

Smartscape now gives you deeper insight into how services connect and communicate in real time. By modeling upstream and downstream dependencies alongside KPIs and infrastructure anchors, Smartscape makes it easier than ever to understand how services interact, where failures originate, and how changes ripple across the stack.

With the new Service Dependency Graph view, teams can instantly visualize their service landscape. The interactive graph makes it easy to follow call flows, isolate a single service and its direct dependencies, highlight performance or error hotspots, and identify unexpected communication paths. Apply your own business context, for example, ownership, to help teams see how services come together to deliver business functionality.

Service Dependency Graph, visualizing a horizontal topology of services.
Figure 7. Service Dependency Graph, visualizing a horizontal topology of services.

Infrastructure

Smartscape expands visibility into infrastructure by mapping all running components, showing how they’re connected, and identifying how performance issues might impact other critical services. The Infrastructure Overview turns this into an intuitive, navigable map that lets teams focus on the data relevant to them and spot bottlenecks or drift patterns through topology shape. Building on this foundation, upcoming Dynatrace enhancements will allow teams to visually inspect host‑to‑process chains and explore network paths enriched with SNMP/LLDP data.

Open the Infrastructure Overview directly from the Infrastructure & Operations App. (video)
Figure 8. Open the Infrastructure Overview directly from the Infrastructure & Operations App. (video)

Problems

Smartscape enhances problem analysis by automatically connecting detected anomalies to the entities and dependencies they impact across your environment. This shows not only what is broken, but how issues propagate across services, workloads, and infrastructure, giving teams immediate clarity on root cause and blast radius. The Problem Graph highlights affected entities, correlates related anomalies, allows for impact isolation, and provides AI-powered insights in context.

End-to-end discovery

The Smartscape app also exposes your entire digital ecosystem as one coherent model, visualizing all dependencies and connecting cloud resources, Kubernetes clusters, infrastructure components, and services end-to-end in the Smartscape on Grail view. This allows teams to understand the real system structure, uncover hidden dependencies, and validate architectural assumptions with complete context rather than piecemeal data.

Figure 9: The Smartscape on Grail view visualizes all dependencies across all your digital systems.
Figure 9: The Smartscape on Grail view visualizes all dependencies across all your digital systems.

Experience the new Smartscape today

Smartscape changes how teams operate by providing automatic, real-time context across all domains, enabling faster troubleshooting, safer releases, stronger security posture, and more cost-efficient operations.

  • Explore domain-specific views for AWS EC2, Kubernetes, Infrastructure, and Services, with Azure coming soon.
  • Run impact analysis with DQL graph queries.
  • Combine topology with logs/metrics/traces/RUM for full stack insights.
  • Let Dynatrace Intelligence take safe, informed actions based on production truth.
The new Smartscape is available in all Dynatrace SaaS environments and on the Dynatrace Playground.

The post The new Dynatrace Smartscape improves operational efficiency across clouds, Kubernetes, infrastructure, and more appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-new-dynatrace-smartscape-improves-operational-efficiency-across-clouds-kubernetes-infrastructure-and-more/feed/ 0
The new Smartscape: Move faster and make better decisions with a real-time dependency graph of all your digital systems https://www.dynatrace.com/news/blog/new-smartscape-make-better-decisions-with-real-time-dependency-graph-of-digital-systems/ https://www.dynatrace.com/news/blog/new-smartscape-make-better-decisions-with-real-time-dependency-graph-of-digital-systems/#respond Wed, 28 Jan 2026 16:55:22 +0000 https://www.dynatrace.com/news/?p=72731 Smartscape graphic

Making the right decisions in modern IT can feel like changing a tire on a moving car. These environments span millions of rapidly changing entities across the cloud, on-premises, and Kubernetes, with adaptive AI agents that increase complexity as they evolve at runtime. IT leaders are often forced to act on incomplete knowledge, leading to misdirected investments, higher MTTR, architectural drift, and increased compliance risk. Dynatrace Smartscape®, a real-time dependency graph, closes this gap with an always-accurate view of your entire digital ecosystem. You always have full visibility into your entire end-to-end topology—see exactly what’s running in production, where it’s running, and the underlying infrastructure it depends on. The new Smartscape introduces powerful visual analytics, agentless cloud data ingestion, and complete domain-specific metadata, including raw cloud and Kubernetes objects. With richer context and stronger analytics, your teams will work more efficiently, diagnose issues sooner, and make better decisions.

The post The new Smartscape: Move faster and make better decisions with a real-time dependency graph of all your digital systems appeared first on Dynatrace news.

]]>
Smartscape graphic

A paradigm shift: Moving from signal monitoring to true entity health understanding

Digital systems continuously evolve and adapt. Making effective decisions requires real-time, always up-to-date insights into how everything is connected—from infrastructure to applications—and understanding the impact on users and the business when issues arise. When an anomaly occurs, you need to quickly pinpoint the root cause, understand why it happened, identify exposed components, and determine which critical services depend on them.

Autonomous AI agents raise expectations for speed, precision, and self-healing IT systems, but, like humans, AI agents rely on high-quality data and rich context to operate effectively. Basic observability signals fall short because they lack contextual information and a holistic view of IT entities and their dependencies, making it difficult to derive impact and causality. In a world of AI agents, an accurate, real-time production context becomes non-negotiable, as only then can you trust AI agents to make the right decisions.

Smartscape has been a core part of the Dynatrace platform for years, and it powers Dynatrace causal AI. Smartscape is a real-time dependency graph that visualizes how your IT components depend on each other, continuously updating as your topology changes. Smartscape maintains AI-ready context across ephemeral components in hybrid and multicloud environments. This real-time discovery and updating happens automatically across multiple sources, so the model always reflects reality and serves as your ultimate single source of truth.

Explore your IT systems and their dependencies visually with the new Smartscape app
Figure 1. Explore your IT systems and their dependencies visually with the new Smartscape app.

The new Smartscape was developed for cloud native, large-scale environments and comes with major new capabilities:

  • Powerful visual analytics in the all-new Smartscape app, including domain-specific views for cloud, Kubernetes, infrastructure, and more.
  • Fully native cloud entities with complete metadata and raw cloud/Kubernetes object JSON.
  • Agentless cloud data ingest that automatically adds all entity dependencies and policy context.
  • Exploration at scale with Dynatrace Query Language (DQL). Run native graph queries (for example, traverse) to multi-hop across millions of relationships with full Grail® context.

These enhancements unlock many high-value use cases, transforming the way you and your teams work.

  • End-to-end cloud visibility: Close cloud‑console gaps with full relationship context for faster decisions across multi‑cloud and hybrid environments.
  • Understand Kubernetes dependencies immediately: Diagnose issues quicker with full‑fidelity objects, YAML context, and cross‑cluster dependency tracing.
  • Stronger security posture: Visualize exposure and attack paths, prioritize by real blast radius, and enforce IAM and configuration best practices.
  • Accelerate incident response and collaboration: Use the new Visual Resolution Path to see upstream and downstream dependencies and get Dynatrace Intelligence insights in context. Route alerts to the right owners, cut handoffs, and reduce escalations.
  • Validate and optimize architecture and cost: Compare intended vs. runtime architecture, align ownership, and reduce costs using real dependency and utilization context.
  • Reliable CMDB/ServiceNow enrichment: Export precise, auto‑discovered dependencies with external IDs for accurate ITSM automation.
The new Visual Resolution Path in the Problems app.
Figure 2. The new Visual Resolution Path in the Problems app.

Spot and understand patterns instantly with the new Smartscape app’s powerful visual analytics capabilities

The new Smartscape app delivers rich, large-scale visual analysis tailored for modern cloud and AI-native environments, with a smooth and responsive experience at scale.

  • Explore thousands of entities interactively and quickly drill down into topology and entity details.
  • Navigate your entire environment with all dependencies, or use domain-focused, ready-made Smartscape views for clouds, Kubernetes, classic infrastructure, or services.
  • Work in context across the platform:
    • The Visual Resolution Path in the Problems app lets you jump directly into the Smartscape Problems Graph to analyze systemic patterns or blast radius.
    • View topology: an intuitive, in-context action that lets you explore vertical and horizontal topologies for any entity without leaving your app or workflow, which is perfect for understanding dependencies and drilling into details.
  • Align views to your business context with segments. Focus your analysis on entities by team ownership, environment, business unit, or region to reduce noise during change windows and reviews. This allows you to understand who owns what and accelerate collaboration with a shared understanding of your IT system.
  • Analyze patterns from every angle: switch topology layouts to reveal different insights. force exposes hidden clusters and dependency hubs for dynamic microservice exploration; horizontal traces workflows left to right for end-to-end transactions or CI/CD flows; and vertical highlights layered architectures top to bottom for dependency stacks or escalation paths.

Agentless topology discovery from clouds

Dynatrace can now directly discover entities and relationships from cloud environments, including security groups, VPCs, load balancers, subnets, and other network services, providing accurate connectivity, policy, and configuration context, even without deploying OneAgent. In addition to raw topology, Dynatrace automatically ingests, normalizes, and enriches cloud‑native metadata such as tags, labels, ownership properties, cost centers, compliance attributes, and compute metadata, creating a high-quality semantic layer.

All discovered entities and their metadata are unified in the new Smartscape, revealing practical and actionable insights, for example: Which instances are publicly accessible? How is traffic routed across accounts and regions? How do security groups influence exposure paths? And much more.

Infrastructure overview across different platforms and services
Figure 3. Infrastructure overview across different platforms and services

Experience the new Smartscape, now Generally Available, in your Dynatrace environment

Explore the new Smartscape app, which will be available in your Dynatrace SaaS environments during the first week of February 2026, and begin uncovering dependencies across clouds, Kubernetes, infrastructure, services, and problems.

For a deeper dive, check out our documentation and walk through the Playground notebook, which guides you step‑by‑step in using DQL to query Kubernetes entities.

And stay tuned: our next blog post in this series will take you even further, exploring the new domainspecific views in Smartscape.

Ready to get started? Start exploring your digital systems now on the Dynatrace Playground.

The post The new Smartscape: Move faster and make better decisions with a real-time dependency graph of all your digital systems appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/new-smartscape-make-better-decisions-with-real-time-dependency-graph-of-digital-systems/feed/ 0
Explore without friction: Deeper insights with Dynatrace expanded analytics app portfolio https://www.dynatrace.com/news/blog/deeper-insights-with-dynatrace-expanded-analytics-app-portfolio/ https://www.dynatrace.com/news/blog/deeper-insights-with-dynatrace-expanded-analytics-app-portfolio/#respond Wed, 28 Jan 2026 16:55:12 +0000 https://www.dynatrace.com/news/?p=72769 Dynatrace analytics app portfolio

Modern IT systems operate under constant pressure to deliver efficiency, resilience, security, agility, and business alignment simultaneously. That balance isn’t achieved overnight; it’s built through continuous improvement driven by learning and insight. This is where exploratory analytics becomes essential: analytics help you probe deeper, ask sharper questions, uncover patterns, anticipate issues, and optimize resources. With the Grail® unified data lakehouse, all live production data is unified in context and ready to explore. Discover the enhanced Dynatrace analytics app portfolio and see how embedding exploration into your processes requires little effort yet delivers transformative results.

The post Explore without friction: Deeper insights with Dynatrace expanded analytics app portfolio appeared first on Dynatrace news.

]]>
Dynatrace analytics app portfolio

Spark continuous improvement by making data exploration a habit

In practice, exploration often stalls. Logs live in one tool, metrics in another, traces somewhere else, and the broader business context lives with different teams. Answering a single question, such as “Did this pattern exist before the last release?” can mean switching contexts, running separate queries, and manually correlating results. The friction builds, and a deeper investigation is deferred until the next incident forces it.

There is more than one kind of exploration. Sometimes exploration is structured: embedded within workflows as part of postmortems, release validations, or SLO reviews, tracing signals across logs, metrics, traces, events, and business data to link impact to outcomes and prevent repeat failures. Other times, exploration is curiosity-driven: a quieter moment where you notice a memory pattern tied to a batch job, a timeout spike with specific clients, or a cloud spend anomaly you’d never have caught in a scheduled report.

Both of these exploration modes matter. Together, they foster learning and a culture of continuous improvement, both essential to modern enterprises.

Dynatrace Grail, a unified data lakehouse, makes exploration easy and rewarding: a single place for all your live production data across logs, metrics, traces, events, business, and security data, so you can follow the evidence wherever it leads without rigid queries or manual joins. With Dynatrace Query Language (DQL), every field and relationship is at your fingertips, enabling broad searches, contextual pivots, precise slicing, and rapid iteration to test and discard weak hypotheses.

Dynatrace Intelligence® makes exploration accessible to everyone. Use Assist to query and explore your data in natural language, or get support in interpreting and understanding your findings. Leverage AI-powered analysis to detect patterns and anomalies at scale, or forecast trends to predict future behavior.

To support different exploration needs, Dynatrace offers a portfolio of use-case optimized apps:

  • Dashboards for persistent, shared visibility and ongoing monitoring.
  • Smartscape for visual analytics of real-time topology and dependency context.
  • Notebooks for collaborative, ad hoc exploration and rapid hypothesis testing.
  • Investigations for sequential, forensic depth in complex scenarios.

Move seamlessly between apps without losing context. Start in Dashboards, drill into a data point, and continue to explore your data in Notebooks. From there, you might run a deep, focused analysis in Investigations, then pivot to Smartscape for a dependency graph. Finally, bring your findings back into a dashboard for continuous monitoring, making your entire exploration journey seamless.

Let’s look at the apps in more detail, starting with Dashboards, often a natural entry point for your exploration journey.

The exploratory apps portfolio, each app optimized for different use cases.
Figure 1. The exploratory apps portfolio, each app optimized for different use cases.

Dashboards: from real-time visualizations to taking action

Dashboards provide a powerful way to transform complex data visualizations into actionable insights, serving as the cornerstone of the Dynatrace exploratory analytics portfolio where exploration meets operational excellence. By offering real-time visibility into key metrics, dashboards help teams monitor performance, identify trends, and make informed decisions. With ready-made dashboards for common use cases, such as Kubernetes, infrastructure, and digital experience monitoring, teams gain immediate access to critical insights, allowing for faster and more proactive responses to their daily challenges.

Dashboards are designed to foster operational clarity with intuitive, interactive visualizations that allow you to drill down into metrics, apply filters, and segment data to uncover meaningful patterns. While Notebooks and Investigations are ideal for deep dives and custom analyses, Dashboards deliver concise, shareable, real-time views that keep teams aligned and informed.

Deeply integrated with Dynatrace’s AI-powered analytics, dashboards enhance visualizations with contextual explanations, anomaly detection, and forecasting. These capabilities allow teams not only to monitor what’s happening but also to understand why it’s happening and predict what might happen next. By making insights accessible to both technical and non-technical stakeholders, dashboards foster collaboration, break down silos, and empower teams to stay aligned and proactive.

Use Dashboards to:

  • Monitor KPIs and SLOs in real time
  • Identify anomalies and emerging trends early
  • Align teams with shared, role-based views and a single source of truth
  • Trigger deeper analysis via drill-downs into charts and entities
  • Track progress against goals and initiatives over time
  • Surface business and technical context side by side for informed decisions
Get instant insights into infrastructure health with ready-made dashboards.
Figure 2. Get instant insights into infrastructure health with ready-made dashboards.

Smartscape: visualize the topology and dependencies of your complete digital systems

Smartscape® is the latest addition to the Dynatrace exploratory analytics app portfolio, and it’s a game-changer for exploring highly dynamic IT systems. Purpose-built for real-time visual analytics, Smartscape gives you a dynamic, interactive view of your entire IT ecosystem—spanning all layers, including services, cloud, Kubernetes, and on-premises infrastructure. Unlike static diagrams or manual dependency maps, Smartscape updates continuously, so you can understand changes as they happen.

Smartscape’s visual analytics capabilities go far beyond simple mapping. It provides multidimensional, domain-specific views that allow teams to see how services, processes, and infrastructure interact in real time. This real-time visualization helps uncover hidden dependencies, assess the blast radius of outages, spot drift or misconfigurations, and validate architecture after deployments. Apply your business context by using Segments, and pivot from other apps like Problems, Kubernetes, or Clouds into Smartscape without losing context.

Visualize and explore dependencies across your IT systems at scale with the new Smartscape app.
Figure 3. Visualize and explore dependencies across your IT systems at scale with the new Smartscape app.

Use Smartscape to:

  • Visualize real-time dependencies and communication paths across services and infrastructure
  • Assess blast radius and map out highly connected and interdependent entities during incidents
  • Validate architecture and changes after deployment
  • Identify and understand hotspots, bottlenecks, and hidden dependencies
  • Navigate readymade domain views for clouds, Kubernetes, services, and infrastructure with zero setup
  • Align engineering, ops, and business teams with a shared, always-current understanding

Notebooks: collaborate, explore, and solve problems in real-time

Notebooks bridge the gap between the two modes of exploration and play an important role in both standardized processes and curiosity-driven exploration.

As a workspace for free exploration, Notebooks give you a playground to experiment with data, quickly visualize insights with a large set of chart types from a curated library, and iterate quickly. You can slice massive datasets in real time, pivot on context, and uncover patterns without constraints.

At the same time, Notebooks shine in collaborative workflows. Teams can work together to document and share findings during incident resolution or postmortems, create troubleshooting guides, and also generate automated reports from queries, all within the same space. Notebooks documenting incidents are automatically surfaced in the Problems app via vector search when similar issues occur, and snapshots of investigations can be preserved as long as needed outside of retention period settings, ensuring insights remain accessible.

Whether you’re just performing free-form discovery or creating documents within processes, Notebooks make it effortless to turn exploration into reusable assets.

Use Notebooks to:

  • Collaborate on incident investigations
  • Document postmortems for future reference
  • Report insights ad hoc or on a schedule
  • Analyze your data using generative AI
  • Prototype and validate DQL for alerts, workflows, and automation
  • Tell data stories with rich visuals and narrative
  • Build a reusable knowledge base to reduce MTTR
  • Extract data on demand
  • Transform and shape data on read
Notebooks are the perfect place for ad-hoc data exploration, collaboration, and data storytelling.
Figure 4. Notebooks are the perfect place for ad-hoc data exploration, collaboration, and data storytelling.

Investigations: dive deeper with sequential analysis and forensics

When exploration moves from curiosity to critical analysis, Investigations is your go-to tool. Built for structured, forensic deep dives, it’s the perfect complement to ad hoc exploration in Notebooks.

Investigations works with DQL across all data in Grail, including logs, events, metrics, traces, business data, and security signals. Teams can pivot from initial findings to comprehensive analysis without friction, comparing scenarios and following evidence trails wherever they lead. The query tree tracks your analytical path, letting you branch into parallel hypotheses and return to previous queries and results at any point.

Compare different scenarios and follow evidence trails with the query tree in Investigations.
Figure 5. Compare different scenarios and follow evidence trails with the query tree in Investigations.

Imagine this flow: a security team is alerted to unusual login attempts or wants to follow up on an anomaly spotted in Dashboards or Notebooks. With a single click, they transition to Investigations to trace lateral movement, simulate attack scenarios, and preserve evidence for future reference. Investigations support sequential workflows, allowing you to pivot queries based on metadata, visualize intricate patterns, and even enrich analysis with lookup tables and external data joins.

Pivot queries based on metadata, visualize patterns across multiple dimensions, and enrich your analysis with lookup tables and external data joins. Because all Grail data is accessible, you can seamlessly connect the dots, linking a suspicious error to its pod’s resource consumption, or tracing a payment failure back to the infrastructure event that caused it. Custom pivots let you select multiple findings and branch into separate queries automatically.

When you’ve found what you’re looking for, save the investigation, or just the relevant branches, as a Notebook to share with your team.

This isn’t just incident response, but everyday analytics for complex environments. Investigations help teams validate hypotheses, document findings, and strengthen resilience across IT systems.

Use Investigations to:

  • Conduct structured, multi-step analyses across services and systems
  • Correlate signals from applications, infrastructure, and user activity
  • Follow evidence trails to confirm or refute hypotheses
  • Explore parallel scenarios with branching query paths
  • Diagnose complex integration and dependency issues across environments
  • Collect and preserve evidence for auditability and knowledge reuse
  • Enrich analyses with context (for example, lookups, reputation data, metadata)

Ready to make more of your data with Dynatrace?

Chances are your IT organization is currently focused on increasing efficiency, strengthening resilience and security, increasing agility, or aligning more closely with business priorities. Within your Dynatrace data, there are likely far more insights waiting to be uncovered, helping you accelerate these goals. Start by formulating the right questions: Where are inefficiencies hiding? Which patterns precede incidents and impact uptime? How exposed are critical services?

Give yourself time and room to explore: Visualize dependencies in Smartscape. Use Assist to turn your questions into DQL queries in Notebooks; experiment and try things out. When findings require structured follow-up, Investigations will help you get a complete understanding. And finally, for anything interesting you see on a chart, dashboards let you drill down into the details.

Start exploring now and make exploration a cornerstone of continuous improvement.

For inspiration and an overview of available Exploratory Analytics resources, take a look at our Perform 2026 – Exploratory Analytics Launchpad.

The post Explore without friction: Deeper insights with Dynatrace expanded analytics app portfolio appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/deeper-insights-with-dynatrace-expanded-analytics-app-portfolio/feed/ 0
Unified observability delivers deeper insights with AI-driven analytics and automation https://www.dynatrace.com/news/blog/ai-driven-analytics-and-automation-for-unified-observability/ https://www.dynatrace.com/news/blog/ai-driven-analytics-and-automation-for-unified-observability/#respond Mon, 12 Feb 2024 21:23:39 +0000 https://www.dynatrace.com/news/?p=62396 Perform 2024: Make waves

Today’s organizations flock to multicloud environments for myriad reasons, including increased scalability, agility, and performance. However, these environments can drown enterprises in data, forcing them to adopt multiple tools and services to manage and secure it. This fragmented approach adds complexity and opens the door to security vulnerabilities. In fact, according to recent Dynatrace research, […]

The post Unified observability delivers deeper insights with AI-driven analytics and automation appeared first on Dynatrace news.

]]>
Perform 2024: Make waves

Today’s organizations flock to multicloud environments for myriad reasons, including increased scalability, agility, and performance. However, these environments can drown enterprises in data, forcing them to adopt multiple tools and services to manage and secure it. This fragmented approach adds complexity and opens the door to security vulnerabilities.

In fact, according to recent Dynatrace research, 85% of technology leaders say the number of tools, platforms, dashboards, and applications they use adds to the complexity of managing a multicloud environment. Further, 84% of technology leaders say multicloud complexity makes it harder to protect applications from security vulnerabilities and attacks.

With unified observability and security, organizations can protect their data and avoid tool sprawl with a single platform that delivers AI-driven analytics and intelligent automation.

During a Dynatrace Perform 2024 breakout session, Dynatrace colleagues Bipin Singh, product marketing director, and Markie Duby, principal solutions engineer, showed how organizations can bring together observability, security, and business data from cloud-native and multicloud environments with Dynatrace.

Update: We’ve expanded Dynatrace Intelligence, extending AI-powered insights across the Dynatrace platform. Dynatrace Intelligence is the evolution of Davis AI®, delivering deeper observability and more actionable intelligence.

The secret sauce of unified observability

Observability enables teams to measure a system’s state based on the data it generates. A unified observability approach takes it a step further, enabling teams to monitor and secure their full stack on an AI-powered data platform.

With the Davis AI engine, Grail data lakehouse, and Smartscape topology visualization at its core, the Dynatrace unified observability and security platform provides AI-driven analytics and automation capabilities.

An overview of the Dynatrace unified observability and security platform.
An overview of the Dynatrace unified observability and security platform.

“Grail handles data storage, data management, and processes data at massive speed, scale, and cost efficiency,” Singh said. “Smartscape contextualizes your entire environment and builds a real-time topology map that’s dynamic and stays up to date as your environment changes. And the Davis AI engine is continuously watching your environment and evaluating the emerging situation, automatically detecting problems, creating automated root-cause analysis for you and business impact analysis for prioritization.”

The importance of hypermodal AI to unified observability

Artificial intelligence is a critical aspect of a unified observability strategy. In fact, according to the recent Dynatrace report, “The state of AI 2024,” 83% of technology leaders say AI has become mandatory to keep up with the growing complexity of multicloud environments.

The Davis AI engine uses a hypermodal approach to bring together causal, predictive, and generative AI. Causal AI determines the underlying causes and effects of issues based on the system’s topology. Predictive AI, meanwhile, makes predictions about future events based on patterns from historical data. And generative AI, termed Davis CoPilot, creates queries, notebooks, and dashboards to simplify analytics, and provides workflow and automation recommendations.

By bringing together these AI types, organizations receive generative AI recommendations based on the precise context from predictive and causal AI. This coactive AI approach enables organizations to spend more time on innovation by simplifying and automating routine tasks.

A breakdown of how Grail, Smartscape, and Davis work together in the Dynatrace unified observability and security platform.
A breakdown of how Grail, Smartscape, and Davis work together.

How Davis tackles root cause for AI-driven analytics

Duby discussed how Dynatrace OneAgent, Smartscape, and Davis work together to take information from many different layers in a full stack to provide root-cause analysis.

“When [Davis is] going through and detecting anomalies within your environment, it’s using data both from the underlying interdependency link as well as that end-to-end trace from the end user all the way back in order to do things like root-cause analysis,” Duby said. “We’re using that causal AI to determine what is actually the underlying root cause.”

A visual representation of what Davis uses for its own analysis in the Dynatrace unified observability and security platform.
A visual representation of what Davis uses for its own analysis.

Davis enables users to go deeper into the details of the underlying processes running on a particular host. The hypermodal AI engine shows what’s happening in a system down to the data coming in, while presenting the information in context.

“It’s one thing to have the data; it’s another thing to have it in context,” Duby continued. “For performance, for security analytics, you have to have the data in context. You need to understand how these different pieces interact with each other and how those pieces are actually coming through.”

Once Davis has gathered all the necessary information throughout the different layers of the stack, it can determine what’s changing, what’s breaking, where the issues are, and how to resolve them. Additionally, it helps users prioritize which issues need immediate attention by providing the necessary context.

“[Davis is] looking at the business context—not just the IT, not just the individual metrics, but understanding the whole picture,” Duby said.

How Davis CoPilot takes AI further and promotes collaboration

A significant piece of the Dynatrace hypermodal AI approach is Davis CoPilot, the generative AI part of the hypermodal engine. Davis CoPilot enables users to create queries, dashboards, and notebooks using natural language input, while offering coding suggestions for workflow automation. Additionally, it simplifies the processes of onboarding, configuring, and adopting the Dynatrace unified observability and security platform.

“This is Davis CoPilot. This is your helper to make sure you can actually go in and take advantage of all of that underlying data,” Duby said. “So, you have the analytics and the performance tracking that Dynatrace is doing, and you also have the ability to build it for your own custom use case.”

A preview of Davis CoPilot returning results from a natural-language input.
A preview of Davis CoPilot returning results from a natural-language input.

In addition to creating queries, dashboards, and notebooks using natural language, Davis CoPilot enables users to share these notebooks with other team members across the organization, boosting collaboration efforts.

“[Davis CoPilot will] start building out queries. And now I can take this information that I just got back, and I can share this notebook with my colleague. And now they have the exact same information,” Duby continued. “I can reuse the same report next month and see if it changed. … This functionality allows me to collaborate with my team. It allows me to run with a new idea and see what comes back. And then I can take this information, and I can build on top of it to do more advanced analytics for my different teams.”

For an in-depth demonstration of using the Dynatrace unified observability and security platform, watch our on-demand session, “AI-driven analytics and automation for unified observability and security.” And for more coverage from Perform 2024, check out our guide.

The post Unified observability delivers deeper insights with AI-driven analytics and automation appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ai-driven-analytics-and-automation-for-unified-observability/feed/ 0
Intelligent, context-aware AI analytics for all your custom metrics https://www.dynatrace.com/news/blog/intelligent-context-aware-ai-analytics-for-all-your-custom-metrics/ https://www.dynatrace.com/news/blog/intelligent-context-aware-ai-analytics-for-all-your-custom-metrics/#respond Wed, 07 Oct 2020 16:02:08 +0000 https://www.dynatrace.com/news/?p=40174 Dynatrace Dashboard

For custom metrics ingested into Dynatrace via open API interfaces like StatsD, Telegraf, and Prometheus, you can now take advantage of the full power of Davis AI topology-aware anomaly detection and alerting.

The post Intelligent, context-aware AI analytics for all your custom metrics appeared first on Dynatrace news.

]]>
Dynatrace Dashboard

Dynatrace recently opened up the enterprise-grade functionalities of Dynatrace OneAgent to all the data needed for observability, including metrics, events, logs, traces, and topology data. Our breakthrough in augmenting open API interfaces like StatsD, Telegraf, and Prometheus now allows customers to feed third-party metrics into Dynatrace and map those metrics into our real-time Smartscape topology.

As your organization moves beyond the myriad of out-of-the-box technologies that are offered by Dynatrace and you begin to stream in third-party metrics, you need to apply the full power of the Dynatrace Davis® AI causation engine to these ingested metrics—think dependency detection, topology visualization, anomaly detection, auto-baselines, root cause analysis, or even business-impact analysis. This is exactly what Dynatrace now delivers.

Davis topology-aware anomaly detection and alerting for your custom metrics

We’re happy to announce that with the latest Dynatrace release, you can leverage the full power of Dynatrace Davis AI to detect and receive alerts on anomalies in your custom metrics. This allows you to:

  • Use auto-adaptive baselines for all your custom metrics.
  • Seamlessly report and be alerted on topology-related custom metrics.
  • Seamlessly report and be alerted on non-topology-related custom metrics, using Dynatrace as a metric database.
  • Convert non-topological custom metrics into topological metrics on the fly simply by adding semantic links to Smartscape topology.

Topology and non-topology metrics—what’s the difference?

Before diving deeper into anomaly detection for all custom metrics, let’s review the fundamentals. What’s the difference between topology and non-topology metrics in Dynatrace?

  • Topology metrics are related to specific entities in your Smartscape topology (for example, the number of successful and failed batch jobs processed by a host).
  • Non-topology metrics are not related to any Smartscape entity (for example, a retailer’s revenue numbers per store). Instead, the metric is related to the monitored environment as a whole.

Smartscape auto-detected topology is an important differentiator of the Dynatrace Software Intelligence Platform as compared to any other legacy monitoring solution. The Smartscape entity model plays an important role for Davis AI, as all built-in metrics are automatically linked to context-rich entities such as hosts, disks, processes, or services.

A topological link to an entity only makes sense, of course, if the measurement that’s sent to Dynatrace has a semantic relationship to that entity. This means that if a measurement is sent for a host, it must be logically linked to that specific host. The same is true for measurements that are sent for services or applications.

Choose your custom metric type

While, in the past, it was only possible to stream third-party metrics into Dynatrace through a custom device API (i.e., an entity), Dynatrace now also supports use cases for reporting, charting, and alerting on non-topological metrics.

When streaming custom metrics into your Dynatrace monitoring environment, you can now specify whether or not a metric has a topological relationship. Either way, you’re now able to seamlessly report, chart, and alert on these metrics, which fulfills a wide array of use cases across your organization that rely on time series metrics and alerting.

Now let’s take a look at anomaly detection for topology-related and non-topology-related custom metrics in action!

Let’s assume that you have an existing OneAgent instance running on a host and you want to stream measurements for the number of successful and failed batch jobs into your Dynatrace monitoring environment.

OneAgent comes with a new metric ingest channel already enabled. You can use a simple curl command to pipe these metrics into Dynatrace. Representative incoming measurements for each are shown below:

$ curl -d "batchjobs.execution.successes,jobname=payslip 5" http://127.0.0.1:14499/metrics/ingest
$ curl -d "batchjobs.execution.fails,jobname=payslip 1" http://127.0.0.1:14499/metrics/ingest

Ingest data via OneAgent rather than our REST API

Ingesting custom metrics through the OneAgent channel comes with a two major benefits as compared to using the same channel via the REST API:

  • Unlike the REST ingest channel, you don’t need an API token; OneAgent handles the secure connection for you.
  • Each OneAgent instance is already aware of the topology that it’s reporting on, so information about related hosts is automatically added to your ingested metrics.

Once you begin sending the two metrics through the OneAgent channel, they will automatically appear within the metric picker (shown below).

Metric picker for a custom chart showing ingested topology metrics

As mentioned above, each OneAgent instance adds its own topological information to each measurement sent to Dynatrace. You can see this in the image below where metrics have been split by host. This metric dimension was automatically added by OneAgent.

The job name is another dimension automatically reported by OneAgent for the two batch job metrics in this example. All metric dimensions, whether you report them or OneAgent adds them to enrich the topological information, can be transparently filtered and drilled into, as shown below.

Ingested custom metric with topological dimensions

Now let’s assume that we want Davis to trigger an alert whenever an anomaly is detected in the number of failed batch jobs. For this, we go to Settings > Anomaly detection > Custom events for alerting where we can select the metric for the number of failed batch jobs (batchjobs.execution.fails) using the metric picker.

Choose your monitoring strategy (i.e., either a Static threshold or an Auto-adaptive baseline), and define the event title and description for the resulting alert.

Set up alerts for ingested custom topology metrics

Our latest innovation for detecting anomalies in metrics, topology-aware Davis-AI auto-adaptive baselining, is unique in that it adapts to changing metric behavior over time, thereby helping you to avoid false-positive alerts.

Once configured, this event will be raised whenever an anomaly is detected in the number of failed batch jobs. As the metric is topology aware, it has a logical link to the host it is reported for. The event will be raised on this host.

Davis AI root cause detection is triggered based on your chosen event Severity level. Refer to Dynatrace Help to learn about which severity levels trigger Davis and which raise problems.

Now let’s see how Dynatrace ingests and alerts on non-topological metrics, which don’t have logical relationships with Smartscape entities.

Because you can now seamlessly report non-topological metrics, you can now use Dynatrace as a metric database. This gives you all the benefits of a metric storage system, including exploring and charting metrics, building dashboards, and alerting on anomalies.

Let’s take the example of a globally distributed retailer that collects revenue measurements every minute for all its shops worldwide. Revenue per shop isn’t really connected to any topological Smartscape entity, so we skip the association and simply stream the metric into Dynatrace.

Each shop sends its revenue measurements enriched with information about its region, country, and city.

See sample measurements below as they stream into Dynatrace (you can find the complete example on Github).

business.shop.revenue,country=us,region=useast,city=Charlotte,store=shop1 60
business.shop.revenue,country=us,region=useast,city=Jacksonville,store=shop2 24
business.shop.revenue,country=us,region=useast,city=Indianapolis,store=shop3 47
business.shop.revenue,country=us,region=useast,city=Columbus,store=shop4 44
business.shop.revenue,country=us,region=useast,city=NewYork,store=shop5 65
business.shop.revenue,country=us,region=uswest,city=SanFrancisco,store=shop6 95
business.shop.revenue,country=us,region=uswest,city=Seatle,store=shop7 83
business.shop.revenue,country=us,region=uswest,city=SanDiego,store=shop8 100
business.shop.revenue,country=us,region=uswest,city=Portland,store=shop9 100
business.shop.revenue,country=us,region=uswest,city=Anaheim,store=shop10 106

Once all the shops begin reporting their revenue, you can explore the data by slicing and dicing it based on multiple dimensions. The image below shows shop revenue by city.

Now, let’s set up a basic alert for one of these metrics: Go to global Settings > Anomaly detection > Custom events for alerting.

Choose the business.shop.revenue metric and select a dimension value, such as Anaheim, that you want to be alerted on if revenue drops for that city. Here, too, you can select a threshold (Monitoring strategy) and provide a name and description for the alert. Note in the image below that a Static threshold of 100 has been configured to immediately force an alert (for demonstration purposes).

Here’s the completed alert configuration for shops located in Anaheim:

Non-topological metric alert configuration

This intentionally low threshold generates an event and an alert within a few minutes, as shown below. With non-topological metrics, you have all the benefits in Dynatrace that a metric storage system can provide, such as exploring and charting metrics, building dashboards, and alerting on anomalies.

It’s important to note that the alert above was raised at the environment level, as no topological entity is linked to the incoming business metric.

What if a topological connection makes sense after all?

That’s easy! Say that, after some time, you discover that a topological connection to an entity (such as an application or a purchase service) makes perfect sense for a business metric. The benefit of the new Dynatrace metric ingestion functionality is the flexibility you get in adding semantic links to Smartscape topology on the fly. You can do this by adding dimensions to your measurements.

For example, let’s assume that you want to link all the business measurements in the example above to an existing application, easyTravel.

We can link a measurement to an application (by ID) in the job/script sending the metric to Dynatrace. This simply adds the reserved dimension dt.entity.application to the metric stream. Depending on your use case, you might want to link different applications to each individual shop’s business measurement or use one application for all, as we’ve done in the example below. The added dimension is highlighted in each measurement.

business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=Charlotte,store=shop1 73
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=Jacksonville,store=shop2 60
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=Indianapolis,store=shop3 90
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=Columbus,store=shop4 42
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=useast,city=NewYork,store=shop5 42
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=SanFrancisco,store=shop6 106
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=Seatle,store=shop7 122
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=SanDiego,store=shop8 137
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=Portland,store=shop9 98
business.shop.revenue,dt.entity.application=APPLICATION-A0641580EEB00D53,country=us,region=uswest,city=Anaheim,store=shop10 120

Now, if we set up the same custom alert as shown above, we’ll get a topologically enriched event and alert. The event will therefore be raised with reference to the specific application instead of the entire environment.

This also means that if you choose a dedicated event severity level, your configured event will be fully Davis enabled and trigger root cause detection on the auto-discovered topology.

See the Custom info metric event below that was raised for the easyTravel application based on an anomaly that was detected in the business metric for stores in Anaheim.

Info-level event for an application connected to an ingested custom metric

If you selected the Error severity level in the alert configuration, you will also get an alert and Davis root cause analysis will be triggered for your connected easyTravel application, as shown below:

Problem generated with Error severity level for ingested custom metric

Summary

Dynatrace has achieved a breakthrough in augmenting open API interfaces like StatsD, Telegraf, and Prometheus by allowing you to feed third-party data from these sources into Dynatrace and map the metrics into real-time Smartscape topology.

As you stream in third-party metrics, you now have the full power of Davis AI on these metrics—topology visualization, anomaly detection, auto-baselining, root cause analysis, and even business-impact analysis. You can either ingest these metrics with no topological connections (using Dynatrace as a metric storage system) or you can enrich the incoming metrics with semantic links to your autodiscovered Smartscape topology model.

With this advancement, Dynatrace is now the data-to-answers-to-actions processing engine of choice that relieves you of the burden of manual health and performance analysis. By leveraging automation over existing data sources, Davis AI enables proven, state-of-the-art AIOps, including auto-remediation workflows.

The post Intelligent, context-aware AI analytics for all your custom metrics appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/intelligent-context-aware-ai-analytics-for-all-your-custom-metrics/feed/ 0