Automation | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Tue, 19 May 2026 08:53:05 +0000 en hourly 1 Dynatrace AI agents begin working for you on day one, and are built to grow with you https://www.dynatrace.com/news/blog/dynatrace-ai-agents-begin-working-for-you-on-day-one-and-are-built-to-grow-with-you/ https://www.dynatrace.com/news/blog/dynatrace-ai-agents-begin-working-for-you-on-day-one-and-are-built-to-grow-with-you/#respond Fri, 03 Apr 2026 15:44:42 +0000 https://www.dynatrace.com/news/?p=73625 Agents graphic

AI agents are everywhere in tech conversations right now, but what agents can you actually use today to make your job easier? In Dynatrace, ready-made agents help developers, SREs, and IT operations teams investigate issues, understand system behavior, and reduce manual work using the data they trust every day. Dynatrace ready-made agents are not concepts or previews; they're available now, integrated into existing Dynatrace workflows, and designed to solve real operational problems. For teams ready to go further, Dynatrace agents lay the groundwork for autonomous operations.

This blog shows what Dynatrace ready-made agents are, how to get value from them quickly, and how to decide which agents are relevant for you, using concrete examples rather than promises.

The post Dynatrace AI agents begin working for you on day one, and are built to grow with you appeared first on Dynatrace news.

]]>
Agents graphic

From generic AI to task‑focused operational agents

Dynatrace ready‑made agents are purpose‑built capabilities that apply Dynatrace intelligence to specific, recurring operational tasks. Each agent focuses on a clearly defined problem, such as explaining why a service is slow, summarizing unusual behavior in an environment, or helping you understand what changed and why it matters. These agents are designed to take a question or a signal based on the exact data that is in your environment and organization and turn it into a useful answer you can act on.

Because Dynatrace agents are ready‑made, there is no need to define prompts, train models, or design behavior from scratch. Each agent already knows:

  • What type of input to expect,
  • Which Dynatrace signals and context it should use,
  • And what output types are most useful for each addressed problem type.

All available ready-made Dynatrace agents can be found in Dynatrace Hub.

Trigger agent actions with Dynatrace Workflows and the Dynatrace MCP Server

Ready‑made agents can be triggered automatically as part of Dynatrace Workflows or available wherever you already work via the Dynatrace MCP Server.

Using agents in Dynatrace Workflows

Dynatrace Workflows lets you run agents in response to events or on a schedule. Instead of manually asking questions about potential problems and remediation steps, the workflow autonomously responds to changes in your environment.

For example, the Kubernetes Troubleshooting Agent runs nine parallel queries for data enrichment, and Dynatrace Intelligence turns all the information into a structured diagnosis. Customize the agents to your needs, including instructions for human approval steps and automated remediation.

Dynatrace Kubernetes Troubleshooting Agent in action.
Figure 1. Dynatrace Kubernetes Troubleshooting Agent in action.

The fastest way to get started is with Dynatrace ready-made agentic workflow templates, currently available in a preview release. Instead of building from scratch, you get proven automations that summarize issues, suggest remediation, and deliver insights directly to the tools your teams already use.

Figure 2. Agentic workflow templates available in preview
Figure 2. Agentic workflow templates available in preview

Power users can go further by building their own agentic workflows that combine Dynatrace Intelligence actions with any trigger, data source, or integration in Workflows. Use cases range from auto-scaling Kubernetes clusters based on Dynatrace Intelligence forecasts to generating query-cost-optimization recommendations for stakeholders, to virtually any other automation your environment requires.

Using agents through the Dynatrace MCP Server

The Dynatrace MCP Server makes the agents available outside the Dynatrace web UI, without requiring you to deploy or operate any additional infrastructure. You can connect Dynatrace to any MCP‑compatible client in minutes, with no server to install, host, or maintain.

Through the tools exposed by the MCP Server, you can use natural language to query data in Grail®, check system health, and get problem analyses and remediation recommendations. This brings Dynatrace directly into the tools you already use, such as your IDE, Claude Code and Cowork, Microsoft Copilot, Slack, or automation platforms like n8n. The Dynatrace MCP Server also powers integrations with systems like Azure SRE, AWS DevOps, GitHub Copilot, Atlassian Rovo Ops, Amazon Q, and others.

Dynatrace MCP server in Visual Studio Code with GitHub Copilot
Figure 3. Dynatrace MCP server in Visual Studio Code with GitHub Copilot

This means agents are no longer tied to a single interface. You can ask Dynatrace questions and get grounded, production‑ready answers wherever you work, using the same agents and intelligence that power Assist and workflows.

Dynatrace Assist: a simple way to test ready-made agents

The quickest way to use a ready‑made agent and see how it works before you start creating a workflow is with . Dynatrace Assist lets you ask questions about your environment using natural language, without switching tools or setting anything up.

A simple way to start is with a real problem you already have. For example, when a service becomes slow, open Assist and ask a question such as “Summarize the open problems and highlight those that need immediate attention.” Assist interprets the question, evaluates the environment you’re working in, and pulls together relevant data and context using Dynatrace Intelligence. Instead of manually navigating metrics, traces, logs, and dependencies, you get an explanation grounded in what is actually happening in your system.

Continuing your conversation with Assist, you can refine the question or follow suggested drill‑downs. Assist supports this as a single flow, helping you move from an initial question to deeper analysis and, where applicable, to next steps. You’re not configuring an agent or defining behavior. You’re simply asking a question and letting Dynatrace coordinate the right intelligence and ready‑made agents behind the scenes.

Dynatrace Assist
Figure 5. Dynatrace Assist

This makes Assist your lowest‑friction entry point for using Dynatrace agents. You get a concrete result quickly, using the same data and context you already rely on in your daily work.

What’s next?

If you haven’t already, open Dynatrace Playground, or your Dynatrace tenant, and ask Dynatrace Assist a question to see the ready-made agents in action.

The post Dynatrace AI agents begin working for you on day one, and are built to grow with you appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-ai-agents-begin-working-for-you-on-day-one-and-are-built-to-grow-with-you/feed/ 0
Next-level batch job monitoring and alerting part 2: Using AI to automatically identify issues and workflows to remediate them https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-part-2-using-ai-to-automatically-identify-issues-and-workflows-to-remediate-them/ https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-part-2-using-ai-to-automatically-identify-issues-and-workflows-to-remediate-them/#respond Wed, 21 Jan 2026 14:38:37 +0000 https://www.dynatrace.com/news/?p=72554 Next-level batch job monitoring and alerting

Let’s say you have a nightly account update batch job that processes user transactions, recalculates balances, and synchronizes data across services in a financial application. If it fails or gets stuck and consumes so many resources that your service degrades, you’ll suffer significant business impact. Every second of such an incident can lead to customer […]

The post Next-level batch job monitoring and alerting part 2: Using AI to automatically identify issues and workflows to remediate them appeared first on Dynatrace news.

]]>
Next-level batch job monitoring and alerting

Let’s say you have a nightly account update batch job that processes user transactions, recalculates balances, and synchronizes data across services in a financial application. If it fails or gets stuck and consumes so many resources that your service degrades, you’ll suffer significant business impact. Every second of such an incident can lead to customer frustration, SLA breaches, and lost revenue.

In our previous blog, we demonstrated how you can use Dynatrace to observe batch jobs and how and why to treat them as first-class citizens in your monitoring strategy. We demonstrated how to uncover performance trends, pinpoint failures, and reveal bottlenecks. But detection isn’t enough.

This blog post takes you to the next level. We’re no longer just talking about detecting problems—we’re talking about fixing them automatically with the help of Dynatrace Davis® AI, Workflows, EdgeConnect, and ServiceNow integration.

We’ll walk you through the automated process of determining root cause, creating an incident in ServiceNow, and automatically remediating the problem.

Architecture overview: A modern, cloud-native setup

Our example environment is a Kubernetes-based architecture that mirrors many modern cloud-native enterprise deployments.

Here’s what the setup looks like:

  • A Kubernetes cluster orchestrating all workloads
  • A public-facing NGINX proxy handling incoming traffic
  • A React front-end delivering the user interface
  • A Java-based broker service responsible for backend processing and database interaction
  • A persistent database service for storing and retrieving application data
Architecture and flow of automatic detection and remediation
Figure 1 – Architecture and flow of automatic detection and remediation

This is automatically instrumented using the Dynatrace Kubernetes Operator, which injects monitoring capabilities directly into the cluster without manual configuration. This provides end-to-end visibility across services, infrastructure, and dependencies with no code changes or changes to existing containers.

Enhancing AI context: Tracking batch job events in Dynatrace

For custom use cases and unique business scenarios, Davis® AI needs to be made aware of specific events through custom approaches. While traditional monitoring captures standard metrics, enabling intelligent reasoning around specific workflows—like batch jobs—means providing targeted event tracking. This ensures Davis AI has full visibility into these activities, allowing it to understand and respond to them effectively.

Contextual events, such as the start of the batch job, could be sent to Dynatrace. This functionality allows you to maintain a clear timeline of all activities, including initiating and completing batch jobs, enhancing visibility and control over critical processes.

In our case, we use an HTTP call that follows this pattern, which utilizes the Dynatrace event types documented here:

An example Event API Payload for attaching batch job information to services:

{ 
  "entitySelector": "type(\"SERVICE\"),tag(service:broker-service 
  "eventType": "CUSTOM_INFO", 
  "properties": { 
    "batch-job-name": "update-account", 
    "process-id": "%s", 
    "workload-name": "account-updater", 
    "namespace": "easytrade" 
  }, 
  "title": "%s update-account batch-job" 
}

This payload attaches batch job information to services matching the entitySelector criteria. The event will appear in the timeline of services tagged with broker-service linking batch job execution data to the relevant services.

The Service with all events recorded about the start and completion of the batch job
Figure 2 – The Service with all events recorded about the start and completion of the batch job.

Davis AI in action: Instant root cause analysis

In this real-world example, the application team noticed latency spikes and increased error rates reported by end users. But before anyone raised a ticket or sent an alert, Dynatrace Davis AI had already detected the anomaly.

Here’s what Davis AI surfaced immediately:

  • Problem severity and duration
  • Number of users affected
  • Which services were degraded, and which business transactions were impacted
  • And most importantly: the root cause
AI-powered problem detection with affected users, events, SLOs, and the root cause of the problem
Figure 3 – AI-powered problem detection with affected users, events, SLOs, and the root cause of the problem.

In this case, the root cause was traced to a batch job that had started consuming excessive resources—CPU, memory, and I/O—all of which starved the Java broker service and caused downstream failures.

Without AI, isolating this root cause across layers could have taken hours. With Davis, it was done in minutes.

Identifying service failures due to batch job failures

The screenshot below displays distributed traces filtered to show batch job-related activity and reported failures. It reveals issues in other services, such as BrokerService, where failures from stuck batch jobs are impacting downstream services.

Screenshot showing BrokerService failures from stuck batch jobs
Figure 4 – Screenshot showing BrokerService failures from stuck batch jobs.

Tapping ServiceNow for instant remediation while Davis gave us the diagnosis, we still need a solution. This is where our integrations with ServiceNow come in. Dynatrace Workflows identifies the issue and checks ServiceNow for an existing ticket. If none exists, it leverages Dynatrace EdgeConnect to execute a Kubernetes API, suspending the affected batch job for remediation. If a ticket is already open, the Workflow adds comments to ServiceNow and follows the same remediation process.

Orchestrating recovery: Dynatrace Workflows

We created a custom Dynatrace Workflow specifically for this batch job scenario. It’s designed to react to Davis AI-detected problems that match a particular pattern—resource-intensive jobs degrading core services.

Here’s how the workflow functions:

  1. Trigger. Once the problem is detected and Davis confirms the root cause, the workflow is automatically triggered.
  2. ServiceNow integration. Without manual input, it instantly opens a ServiceNow incident with all contexts—affected services, severity, root cause.
  3. Live updates. As the situation evolves, the ticket is updated in real time, ensuring that both engineering and service management teams are in sync.
  4. Conditional logic. The workflow checks for remediation eligibility—e.g., is it safe to pause or reschedule the batch job?
  5. Remediation. The workflow now initiates the remediation by EdgeConnect by executing a Kubernetes API to suspend or delete the job
  6. Validation and closure. The causal AI validates successful remediation by confirming the batch job suspension, then automatically closes the problem ticket. Optionally, EdgeConnect can perform additional verification to ensure the root cause has been eliminated and system stability is restored.

All of this happens without anyone needing to log into a dashboard.

Closing the loop: Secure remediation with Dynatrace EdgeConnect

Now, let’s talk about the actual fix.

With the root cause identified and the ServiceNow ticket tracking the issue, the final step is to initiate remediation safely and securely inside the Kubernetes cluster.

This is achieved using Dynatrace EdgeConnect, a lightweight, secure connector that allows Dynatrace to command and interact with your private infrastructure; in this case, Kubernetes.

In our case, EdgeConnect was configured to:

  • Connect securely to the Kubernetes API
  • Execute a custom remediation action, such as scaling down the batch job, adjusting its priority, or rescheduling it to an off-peak window
  • Verify the effect of the remediation and confirm resolution back in Dynatrace and ServiceNow
Workflow that detects such custom problems and initiates remediation and notification to IT Service Management
Figure 5 – Workflow that detects such custom problems and initiates remediation and notification to IT Service Management.

All of this occurred without breaking security boundaries, using EdgeConnect’s outbound-only architecture.

Instant notifications with Slack integration

SysAdmins, DevOps, and SRE teams need to stay informed in real time to ensure system reliability. The Dynatrace integration with Slack delivers instant notifications for critical events, such as batch job failures or system errors, directly to your team’s Slack channels.

In this case, they are simply notified as all the action has already been taken care of by intelligent automation.

Notification seen in Slack, which updates the Problem in ServiceNow
Figure 6 – Notification seen in Slack, which updates the Problem in ServiceNow.

In this example, we conducted remediation by automatically suspending the batch job, as confirmed by running a relevant command to check the job status. Note that remediation approaches and use cases may vary, and we recommend tailoring solutions to your specific environment and requirements.

Demonstration of the automatic suspension of the cronjob through automation
Figure 7 – Demonstration of the automatic suspension of the cronjob through automation.

From reactive to proactive with AI-driven automation

Batch jobs have long been a blind spot in observability—often managed outside core APM tooling or considered too niche for automated handling.

But today, Dynatrace brings batch job management into the mainstream of modern observability and automation.

  • With Davis AI, you get root cause detection in real time.
  • With Workflows, you turn those insights into action.
  • With EdgeConnect, you push changes securely to production environments.
  • And with ServiceNow integration, you keep ITSM workflows up to date without lifting a finger.

This is a new era of autonomous cloud operations—where custom, complex issues like batch job failures are no longer exceptional cases, but standard parts of your AI-driven automation strategy.

How to get started

It’s considered best practice to send custom events to Dynatrace to enhance monitoring capabilities. These events help keep Dynatrace informed about key activities or changes in your environment.

If any of these events correlate with issues in the landscape, Davis AI automatically analyzes them and identifies the root cause. This insight can then be used to trigger Workflows that help remediate potential problems proactively.

Start optimizing your observability today!

The post Next-level batch job monitoring and alerting part 2: Using AI to automatically identify issues and workflows to remediate them appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-part-2-using-ai-to-automatically-identify-issues-and-workflows-to-remediate-them/feed/ 0
Write the future: Create your own agentic workflows https://www.dynatrace.com/news/blog/write-the-future-create-your-own-agentic-workflows/ https://www.dynatrace.com/news/blog/write-the-future-create-your-own-agentic-workflows/#respond Thu, 08 Jan 2026 08:00:10 +0000 https://www.dynatrace.com/news/?p=72347 Agentic workflows with Davis CoPilot

Imagine commissioning le Carré and Fleming to build your perfect undercover agent: quietly embedded in the system you’re watching. You hand in your mission brief, which includes the target, objective, and behaviors to track. Your agent observes without drawing attention, reporting insights back to you. On cue, the information flow you’ve carefully orchestrated turns signals […]

The post Write the future: Create your own agentic workflows appeared first on Dynatrace news.

]]>
Agentic workflows with Davis CoPilot

Imagine commissioning le Carré and Fleming to build your perfect undercover agent: quietly embedded in the system you’re watching. You hand in your mission brief, which includes the target, objective, and behaviors to track. Your agent observes without drawing attention, reporting insights back to you. On cue, the information flow you’ve carefully orchestrated turns signals into actionable intelligence that helps pre-empt risk.

Dynatrace doesn’t write spy fiction. However, even better, Dynatrace now lets you write your own smart agentic workflows that deliver intelligent reports and react to changes in your environment based on your objectives.

Adding generative AI to your workflow

Using the power of gen AI, Davis CoPilot® transforms your workflows into agentic instructions. Davis CoPilot lets you explore data using conversational language, translating complex data and queries into summaries, and provides intelligent recommendations across Dynatrace.

Integrated into Dynatrace Workflows, Davis CoPilot is your toolkit for building smart automations, bringing the power of generative AI into your mission-critical workflows.

Build conversational automation that adjusts to live data based on your instructions, sending summaries of your crash logs directly to Slack
Figure 1. Build conversational automation that adjusts to live data based on your instructions, sending summaries of your crash logs directly to Slack.

By embedding Davis CoPilot in your workflows, you can associate any automation with any number of conversational automations. Your workflows can even perform actions autonomously when combined with precise Davis® AI forecasting, for instance, scaling resources based on forecast demands.

In real time, these workflows monitor live data, summarize critical issues, and identify remediation paths or emerging threats. When scheduled, these smart workflows help you outsource routine tasks, such as alerting stakeholders of costly queries.

Let’s look at some examples of how these smart workflows can help you in your daily work.

Build agentic workflows that respond to critical events

Proactive guiding through complex problem remediation

Let’s assume you want to build an automation that cuts through alert noise and analyzes a problem as it occurs, guiding you through the remediation. When a new problem is detected, Davis CoPilot extracts the problem details, summarizes the situation, and provides tailored remediation guidance. By embedding it into a smart workflow, you can select your preferred automation to automatically syndicate this information, populate a ticket in ServiceNow or Jira, or post it to a dedicated Slack channel.

See how you can set up a workflow automation that automatically sends summaries and remediation guidance when a new problem is detected.

Monitor emerging threats to help you orchestrate a response

Next, you can build an agentic workflow that helps you monitor emerging threats and assess their risk to your environment as vulnerabilities are detected in your tenant. In plain language, you instruct your agent to extract IOCs, query security events in your environment, and correlate them with observability data in your environment. Information provided by the external threat feed is automatically matched against the live context in your tenant. The agent has now collected all the necessary information and provides a reliable risk assessment, along with a plan to orchestrate a response, directly in your Slack channel, ensuring around-the-clock visibility and a rapid response.

With a single workflow, you can monitor emerging security events as they occur, understand their impact, and determine the next steps.
Figure 2. With a single workflow, you can monitor emerging security events as they occur, understand their impact, and determine the next steps.
Example of a tailored and contextual analysis delivered to Slack as the issue arises
Figure 3. Example of a tailored and contextual analysis delivered to Slack as the issue arises

Write the future: Build an agentic workflow that autonomously auto-scales your resources

Dynatrace helps you build agents that reason autonomously. The key is to deliver data as precise as Dynatrace forecast capabilities. In this example, we linked the power of Davis AI to forecast demand, with generative AI and GitHub automations. Davis AI predicts the number of resources the hyperscaler infrastructure will need based on forecasted demand. When Davis AI notices a scaling need, Davis CoPilot interprets the data and autonomously edits the manifest using the GitHub automation. Giving you one end-to-end workflow that automatically scales resources up or down based on forecasted needs. To see this in action, watch how this workflow autonomously edits a manifest based on Davis AI suggestions to auto-scale a Kubernetes cluster.

Schedule agentic workflows to optimize routine tasks

Do you feel like sleeping in a little later? Maybe stretching your lunch break a little longer? Running that extra hill without sacrificing your productivity? Scheduling Davis CoPilot into your smart workflow is a great way to automate recurring tasks and save time.

Build an automation that predicts resource consumption

A recurring challenge for SREs is analyzing the full environment to predict future bottlenecks or over-resourcing and continuously translating the data to update stakeholders. Even with great observability in place, you need to ensure that you interpret the data and make timely decisions to inform future provisioning.

By combining Davis AI forecasting automation with Davis CoPilot, you can build an agent that answers key questions, such as which workloads are most resource-intensive, which resources show the most variance, and which require frequent scaling. This automation is capable of highly reliable forecasts, even when data points are limited. The automation interprets the data and emails actionable recommendations directly to you and anyone else who needs to stay informed.

To see this in action, watch the section of this video that explores predicting resource consumption.

Smart workflows that optimize query costs

Scheduling tasks can even help you keep costs lean and efficient. For admins or budget owners, staying within financial limits while maintaining performance is a constant challenge. In this example, we built a smart workflow that identifies the top 20 most expensive queries from the last 24 hours. Davis CoPilot analyzes each query and sends optimization recommendations directly to the query authors via email.

Smart workflow leveraging Davis CoPilot to recommend query optimizations tailored to your tenant
Figure 4. Smart workflow leveraging Davis CoPilot to recommend query optimizations tailored to your tenant
Example of an optimization suggestion delivered to the inbox of the query author
Figure 5. Example of an optimization suggestion delivered to the inbox of the query author

To implement this yourself, tailored to the most expensive queries executed on your tenant, go to our documentation

Conclusion: Adapt your workflows to any stage of your automation journey

These are just a few examples; the applications for it are endless. We’ve designed this workflow action to cater to your organization’s automation appetite. You may want to transform how you keep business stakeholders informed about what’s happening in your environment, leveraging Dynatrace’s highly accurate insights, which are translated into plain language and actionable next steps.

Alternatively, you may be ready to transition towards autonomous operations, where automation not only supports but also acts in a controlled and reliable manner. Davis CoPilot embedded into your workflows opens the door to your agentic journey.

Start your agentic journey and join the Davis CoPilot for Workflows Preview

Davis CoPilot for Workflows is available as a Preview. Sign up now and see how generative intelligence embedded into your workflows transforms your automation. Today, it helps you react faster, optimize more effectively, and collaborate seamlessly. Tomorrow, it will go even further: anticipating needs, orchestrating actions, and enabling truly autonomous reasoning.

Gain efficiency and have your agentic workflows do the work for you!

The post Write the future: Create your own agentic workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/write-the-future-create-your-own-agentic-workflows/feed/ 0
Reliable enterprise automation at scale: Accelerate the innovation loop with Dynatrace Workflows https://www.dynatrace.com/news/blog/reliable-enterprise-automation-at-scale-accelerate-the-innovation-loop-with-dynatrace-workflows/ https://www.dynatrace.com/news/blog/reliable-enterprise-automation-at-scale-accelerate-the-innovation-loop-with-dynatrace-workflows/#respond Mon, 05 Jan 2026 15:00:57 +0000 https://www.dynatrace.com/news/?p=72335 logs and traces

Enterprise automation rarely fails because teams lack ideas. It fails because scaling automation safely is challenging, and conventional workflows often remain static after they ship, while systems and conditions continue to evolve. The new update to Dynatrace Workflows helps teams accelerate their innovation loop and run automation like production software. Five improvements make it easier […]

The post Reliable enterprise automation at scale: Accelerate the innovation loop with Dynatrace Workflows appeared first on Dynatrace news.

]]>
logs and traces

Enterprise automation rarely fails because teams lack ideas. It fails because scaling automation safely is challenging, and conventional workflows often remain static after they ship, while systems and conditions continue to evolve.

The new update to Dynatrace Workflows helps teams accelerate their innovation loop and run automation like production software. Five improvements make it easier to iterate and operate at scale: workflow drafts, sub-workflows, approval requests, persistent execution data, and real-time notifications. Together, they support safer change, reusable standards, built-in governance, and data-driven refinement.

The innovation loop in enterprise automation

The innovation loop is a repeatable cycle for quickly and safely improving enterprise automation that treats workflows as living systems that evolve with your environment, self-improving as systems, teams, and requirements change, rather than remaining static:

  • Teams start by designing or updating a workflow and validating it safely before it goes live.
  • Next, they standardize the parts that work into reusable building blocks so future workflows are faster to create and more consistent.
  • Then they add governance where it matters, such as approval checkpoints for high-impact actions, so speed does not bypass policy.
  • After execution, they observe what actually happened using execution records and signals, including errors, latency, and outcomes, to understand reliability and impact.
  • Finally, they refine the workflow based on those findings by tuning logic, tightening controls, improving reuse, and adjusting notifications.

The result is a compounding impact: faster delivery over time, fewer incidents at scale, automation that integrates changing circumstances, and greater organizational trust, as transparency, auditability, and tight feedback loops back improvements.

How Dynatrace Workflows accelerates the innovation loop

Dynatrace Workflows helps you accelerate your innovation loop with powerful new features designed for adaptability, collaboration, transparency, and control. Each feature is designed to empower teams to safely iterate on automation in dynamic enterprise environments.

Faster iteration with safer testing and experimentation with workflow drafts

The new workflow drafts feature allows teams to create, edit, and refine workflows in a draft state, allowing them to validate logic and configuration without triggering live automation. Workflows that remain in draft mode don’t consume workflow hours.

Workflow drafts support the first step of the innovation loop: iterate safely. Teams can make changes, review them with peers, and confirm the workflow is ready before anything runs in production.

Key benefits

  • Reduce risk by testing configuration changes without production impact.
  • Allow faster iteration and cleaner change management.
  • Improve collaboration during the design and review stages.
  • Control costs with included workflow-hour consumption for draft-only workflows.
Create drafts to refine and safely test your workflows.
Figure 1. Create drafts to refine and safely test your workflows.

Standardize reusable blocks for reliability and speed with sub-workflows

Sub-workflows allow a workflow to be executed as a task inside another workflow, so teams can package repeated logic into reusable sub-workflows and compose larger automations from consistent building blocks.

This supports the “standardize” step of the innovation loop. When teams standardize and reuse proven components, automation becomes more consistent, easier to maintain, and quicker to expand across teams and environments.

Key benefits

  • Enhance reliability and consistency by reusing proven workflow components.
  • Speed up delivery by composing workflows from standard blocks.
  • Simplify maintenance by updating shared logic in a single location.
  • Keep complex automations understandable by breaking them down into smaller, manageable units.

Example use cases

  • Notification workflows: Automatically send notifications based on incident severity or timing.
  • Data processing workflows: Run filter/transform/store stages as distinct sub-workflows, processing large datasets in stages.
  • Incident response workflows: Automate repetitive tasks as reusable steps, such as creating tickets, logging incidents, notifying stakeholders, and triggering remediation steps.
Include a validated component in your workflows.
Figure 2. Include a validated component in your workflows.

Supervised human-in-the-loop control where it matters with approval requests

With Dynatrace Workflows, you can now insert an approval request into a workflow and pause execution until a human reviews and approves the next step.

This feature supports the “governance” step in the innovation loop. Approvals turn your workflows into supervised autonomous operations. This allows the automation of more processes while ensuring strict compliance with enterprise requirements for governance, risk management, and accountability, particularly for high-impact actions.

Key benefits

  • Add explicit human control at high-risk decision points.
  • Reduce errors in sensitive operations (changes, remediations, escalations).
  • Ensure compliance with organizational standards.
  • Increase trust in automation across security, ops, and platform teams.
  • Minimize operational risks.
Add human control to any of your workflow steps.
Figure 3. Add human control to any of your workflow steps.

End-to-end visibility with an audit trail and persistent execution data

Dynatrace Workflows now persists workflow execution data in DQL. All workflow activities are recorded as system events in the dt.system.events table, creating a queryable execution record for tasks, states, and outcomes as a comprehensive audit trail.

Persisting execution data supports the “observe” step of the innovation loop. When workflow execution is captured as data, teams can troubleshoot failures, provide transparency, prove what happened, and improve automation based on evidence rather than anecdotes.

Key benefits

  • Provide an audit trail for enterprise traceability and review.
  • Monitor workflow health and reliability through execution outcomes.
  • Speed up troubleshooting by pinpointing where failures occur.
  • Optimize automation using DQL-driven actionable insights (error rates, trends, hotspots, and more)

Example: Workflow health overview

This example demonstrates how to use execution data to support users in identifying those workflows that need improvement.

The following DQL query…

fetch dt.system.events
| filter event.kind == "WORKFLOW_EVENT" and event.provider == "AUTOMATION_ENGINE" and event.type =="TASK_EXECUTION"
| filter dt.automation_engine.state == "ERROR"
| summarize count = count() , by:{dt.automation_engine.workflow.title}

…counts task execution errors per workflow. Using a honeycomb visualization, it displays error distribution by workflow and provides visibility into those workflows that fail too often.

Audit and observe your workflows with persistent execution data.
Figure 4. Audit and observe your workflows with persistent execution data.

Instant alerts for failures and workflow changes with real-time notifications

Easily set up and customize real-time notifications by email for all workflow errors or changes made by others, ensuring you’re always informed.

These real-time notifications support the “respond and refine” step of the innovation loop. Proactive alerting allows you to manage workflows effectively, minimize disruptions, and maintain seamless operations.

Key benefits

  • Get immediate visibility into failures to address potential issues before they escalate.
  • Detect workflow changes early to reduce drift and surprises.
  • Improve operational coordination across teams that share workflows
  • Ensure uninterrupted operations with timely, actionable alerts.
See notifications in real time and get immediate visibility into your workflow changes.
Figure 5. See notifications in real time and gain immediate visibility into your workflow changes.

Empower your team: Build smarter, safer, and scalable workflows today

Scaling automation today is hard, changes are risky to test, automation becomes outdated and buggy as soon as one part of the organization makes changes, and when something inevitably breaks, you’re left hunting for answers across scattered logs and screenshots. But it doesn’t need to be like this.

Are you ready to make things easier, run your workflows like production software, and accelerate your innovation loop? Here’s how you get started:

Get started building smarter, safer, dynamic, and scalable workflows today!

The post Reliable enterprise automation at scale: Accelerate the innovation loop with Dynatrace Workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/reliable-enterprise-automation-at-scale-accelerate-the-innovation-loop-with-dynatrace-workflows/feed/ 0
Accelerate your autonomous IT operations journey with Dynatrace and ServiceNow integrations https://www.dynatrace.com/news/blog/accelerate-your-autonomous-it-operations-journey-with-dynatrace-and-servicenow-integrations/ https://www.dynatrace.com/news/blog/accelerate-your-autonomous-it-operations-journey-with-dynatrace-and-servicenow-integrations/#respond Mon, 17 Nov 2025 15:41:13 +0000 https://www.dynatrace.com/news/?p=71886 Dynatrace | ServiceNOW

For many teams, the path to automation starts with connecting data and workflows across platforms. Dynatrace and ServiceNow make that possible. Through a growing set of integrations, customers can seamlessly connect both platforms to create incidents, enrich CMDB data, and leverage AI-driven insights for faster, smarter decisions. Six ways to link Dynatrace data with ServiceNow […]

The post Accelerate your autonomous IT operations journey with Dynatrace and ServiceNow integrations appeared first on Dynatrace news.

]]>
Dynatrace | ServiceNOW

For many teams, the path to automation starts with connecting data and workflows across platforms. Dynatrace and ServiceNow make that possible. Through a growing set of integrations, customers can seamlessly connect both platforms to create incidents, enrich CMDB data, and leverage AI-driven insights for faster, smarter decisions.

Our first six integrations are now available in the ServiceNow store. Each one enriches context and helps you remediate issues faster than ever.

Dynatrace Incident Integration Application The Incident App will create incidents based on Dynatrace-identified problems enriched with relevant context and root-cause information. If the ServiceNow Service Graph Connector for Dynatrace is installed, then Configuration Items will automatically be associated within the incident.
Dynatrace Workflows for ServiceNow​ Dynatrace Workflows allow you to define custom processes based on a series of events. These workflows can trigger ticket creation, pull additional information, and more.
Service Graph Connector​ for Observability – Dynatrace The Service Graph Connector dynamically polls Dynatrace for updated entity information and dependencies. This information is then stored within the ServiceNow CMDB.​
Event Management Connector​ for Dynatrace The Event Management integration accepts Dynatrace problem events and transforms infrastructure events into actionable alerts and incidents for escalation and resolution.​
Service Observability Connector for Dynatrace​ The Service Observability integration displays Dynatrace information within the ServiceNow platform for ServiceNow analyst context and validation within ITOM. ​
Dynatrace Analysis AI Agent Connector​ The Dynatrace Analysis AI Agent connects ServiceNow AI agents to Dynatrace Davis® AI to analyze alert impact with agentic workflows. Once connected, the AI agent gathers information to help you investigate alerts.​

Advance your IT operations with Dynatrace and ServiceNow: Four steps to autonomous remediation

Every organization’s automation journey looks different, but the path typically unfolds in stages. Each step builds new capabilities and confidence in automation.

1. Context-rich incident creation

Start your journey by connecting Dynatrace® to ServiceNow ITSM to automatically create incidents with context enrichment and root-cause description to allow for uniform tracking and resolution of tickets. Dynatrace offers two methods for ticket creation within ServiceNow:

  • Dynatrace Incident Integration Application
  • Dynatrace Workflows for ServiceNow

Both methods can generate a ticket and populate the incident with Dynatrace-available context, such as:

  • Root-cause analysis identifying the exact component causing issues
  • Correlation identifiers for all affected hosts and infrastructure
  • Business impact assessment showing affected services and users
  • Dependency context explaining how the problem propagates

For advanced customization, Dynatrace Workflows for ServiceNow offers more flexibility than the standard Incident Integration Application, which is triggered based on Dynatrace-identified problems. This allows customers to further enrich incident descriptions, add additional information such as logs or metrics, or leverage predictive analytics to initiate remediation procedures even before a problem occurs.

This step helps teams eliminate manual triage and ensures every incident includes complete diagnostic context.

2. Real-time configuration management database (CMDB) enrichment

Once your organization is creating incidents, the next step is to add enrichment methods such as associated configuration items for impact analysis and dependency information.

Most organizations’ CMDB management is maintained manually and updated weekly at best. Connecting Dynatrace Smartscape® topology mapping to ServiceNow’s CMDB through ServiceNow Service Graph Connector (SGC) allows real-time entity identification, topology, and dependency mapping—all of which can be leveraged for automatic CI binding. This creates a living representation of an organization’s IT environment that updates continuously as the infrastructure evolves and helps teams make fast, informed decisions. The Service Graph Connector provides:

  • Automatic population of ServiceNow CMDB with entity information and real-time topology
  • Visibility of all dependencies across a cloud stack and Kubernetes
  • Service Mapping tree generation showing the holistic impact of incidents and events
  • Continuous synchronization as the environment changes

The result is a continuously accurate service map that evolves as fast as your environment.

3. Unified event management

To drive uniformity across your organization and reduce alert fatigue, organizations look to leverage ServiceNow as a central event management system with AIOps features enabled. By sending events to ServiceNow ITOM, teams can aggregate and prioritize alerts using AI-driven analytics, ensuring consistent visibility across platforms.

This provides organizations with a unified dashboard that embeds Dynatrace charts and metrics directly into ServiceNow dashboards, removing administrative silos for viewing information across platforms.

Once events and incidents are unified, the next step is intelligent assistance.

AI-powered assistance

As automation evolves, AI agents bridge the gap between insight and action. Once your organization has Dynatrace events flowing to ServiceNow ITOM, you can connect ServiceNow Now Assist with Dynatrace to allow administrators to request information from the Dynatrace console, such as recommended remediation actions or further details on root cause.

The Dynatrace AI Analysis Agent accelerates investigation by gathering key diagnostic details automatically.

4. Autonomous remediation

The journey toward autonomous operations culminates in closed-loop remediation, where systems detect, act, and validate automatically. Mature organizations rely on self-healing workflows to reduce manual effort and accelerate recovery. Whether triggered from Dynatrace or ServiceNow, each workflow can include validation steps that query Dynatrace APIs to confirm the problem is truly resolved, rather than just masked. Example workflows include the following:

  • Restart services when memory saturation is detected on application servers.
  • Scale infrastructure when capacity constraints are predicted.
  • Rollback deployments when anomalies correlate with recent code changes.
  • Clear caches when response time degradation is linked to cache performance.
  • Reset database connections when connection pool exhaustion is identified.
  • Create context-rich alerts or incidents for faster triaging (no remediation without incident).

Unlocking the future of autonomous IT operations

ServiceNow and Dynatrace together form a powerful foundation for intelligent, scalable, and future-ready IT operations. By combining real-time context with AI-driven insights, organizations can reduce noise, resolve issues faster, and operate more efficiently—around the clock. As agentic integrations mature, they pave the way for autonomous collaboration, delivering immediate value while laying the groundwork for a resilient, AI-powered future.

The post Accelerate your autonomous IT operations journey with Dynatrace and ServiceNow integrations appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/accelerate-your-autonomous-it-operations-journey-with-dynatrace-and-servicenow-integrations/feed/ 0
Transforming Azure Data Factory operations with Dynatrace https://www.dynatrace.com/news/blog/transforming-azure-data-factory-operations-with-dynatrace/ https://www.dynatrace.com/news/blog/transforming-azure-data-factory-operations-with-dynatrace/#respond Tue, 09 Sep 2025 15:01:53 +0000 https://www.dynatrace.com/news/?p=70939 Alerting and Analyzing

Data pipelines are critical to seamless operations and informed decision making in modern businesses, and efficiently managing and monitoring those pipelines is crucial for maintaining a competitive edge. Without visibility into pipeline performance, teams risk delays, data loss, and costly downtime. As data volumes grow and workflows become more complex, the stakes get higher—making intelligent […]

The post Transforming Azure Data Factory operations with Dynatrace appeared first on Dynatrace news.

]]>
Alerting and Analyzing

Data pipelines are critical to seamless operations and informed decision making in modern businesses, and efficiently managing and monitoring those pipelines is crucial for maintaining a competitive edge. Without visibility into pipeline performance, teams risk delays, data loss, and costly downtime. As data volumes grow and workflows become more complex, the stakes get higher—making intelligent observability a must-have, not a nice-to-have.

Azure Data Factory (ADF) is a powerful tool for orchestrating and automating data workflows. Whether you’re moving data across hybrid environments, transforming it for analytics, or syncing it between systems, ADF provides the flexibility and scalability needed for modern data operations. Data engineers who deploy Dynatrace with ADF gain valuable insights into performance, business analytics, and automation.

Identifying and resolving pipeline bottlenecks

Keeping data pipelines fast, reliable, and scalable is a huge challenge. When performance dips or failures occur, it’s often a scramble to pinpoint the issue. With Dynatrace, you gain clear visibility into pipeline behavior, resource usage, and failure patterns—making troubleshooting faster and optimization smarter.

Dynatrace addresses key questions such as:

  • What are my longest running pipelines?
    Identifying pipelines with extended runtimes helps teams spot inefficiencies in data processing or transformation logic so that you can optimize performance and reduce latency in downstream systems.
  • Do I have any failing pipelines?
    Immediate visibility into failures allows teams to respond quickly, minimizing data loss and avoiding disruptions to business-critical workflows.
  • Why are my pipelines failing?
    Understanding the root cause—whether it’s a misconfigured activity, resource constraint, or external dependency—enables faster resolution and helps prevent repeat incidents.
  • Which pipelines require optimization?
    By highlighting pipelines with high resource consumption or inconsistent performance, Dynatrace helps prioritize tuning efforts for maximum impact.
  • Do I need to scale resources or adjust concurrency settings?
    These insights guide infrastructure decisions, ensuring that pipelines run efficiently without overprovisioning or underutilizing resources.

By ingesting logs and metrics from Azure Monitor and correlating diagnostics from Azure Data Factory, Dynatrace delivers actionable insights and a comprehensive view of pipeline performance.

Azure Data Factory dashboard in Dynatrace screenshot

Azure Data Factory metrics dashboard in Dynatrace screenshot

Let’s take a closer look at the dashboard. In addition to status and duration, the captured logs and metrics allow us to thoroughly analyze ADF performance.

Azure Data Factory captured logs and metrics in Dynatrace screenshot

  • Time spent in queue: Indicates how long the pipeline waits before starting. Prolonged queue times might suggest a need to scale resources, improve scheduling, or adjust concurrency settings.
  • Time spent in progress: Reflects the actual execution time of the pipeline. Longer durations here may highlight opportunities to optimize pipeline logic or resource allocation.
  • Message: Displays detailed runtime information, including errors. Expanding this field provides additional insights for troubleshooting.

Azure Data Factory message in Dynatrace screenshot

Effective pipeline monitoring transforms reactive troubleshooting into proactive optimization. Leveraging these insights, organizations can not only resolve current bottlenecks but also establish best practices for sustainable, high-performing data environments.

Achieving real-time business insights

Unlocking real-time business analytics isn’t just about tracking technical metrics—it’s about connecting IT operations to business outcomes. With Azure Data Factory (ADF) and Dynatrace, practitioners can enrich pipeline observability by embedding business context directly into logs and metrics. This allows teams to monitor not just how pipelines are running, but what they’re delivering.

For example, imagine a pipeline processing a spreadsheet containing daily revenue figures. By using a Lookup activity in ADF, you can extract key business values—like total revenue or transaction count—and pass them as custom user properties into Dynatrace. This enables dashboards that show not only pipeline health, but also business impact: Was revenue successfully ingested today? Did a failure affect a critical report?

This integration empowers teams to:

  • Track business KPIs alongside technical metrics, making it easier to prioritize fixes based on impact.
  • Spot anomalies in business data early, such as missing values or unexpected drops in volume.
  • Align IT and business teams, fostering collaboration through shared visibility into what matters most.

By bridging the gap between data operations and business insights, Dynatrace helps practitioners move from reactive monitoring to strategic decision-making.

Dynatrace integration in Azure Data Factory

ADF Business Analytics metric in dashboard in Dynatrace screenshot

Ensuring pipeline reliability with automation

Harness the power of automation with Dynatrace Workflows, allowing you to streamline processes with Azure Data Factory. With the ADF REST API, you can configure automation to meet your unique business needs.

Consider the case where an administrator needed a way to automatically retry pipelines on failure for those managed by different teams. While native retry policies were available, this simple automation ensured that retries were executed even when the configuration was overlooked.

This is a perfect example of shifting from reactive troubleshooting to proactive reliability. Instead of waiting for failures to be manually addressed, Dynatrace enables automated responses that reduce downtime, improve consistency, and free up teams to focus on higher-value work. By embedding automation into pipeline operations, practitioners can build more resilient systems and ensure that critical workflows stay on track—even when things go wrong.

Azure Data Factory workflow in Dynatrace screenshot

Get started

By integrating Dynatrace with Azure Data Factory, practitioners gain more than just monitoring: they unlock a smarter, more proactive way to manage data pipelines. From identifying bottlenecks and failures to embedding business context and automating recovery, Dynatrace transforms pipeline operations into a strategic advantage.

The key benefits:

  • Faster troubleshooting with deep visibility into pipeline behavior and failure patterns
  • Smarter optimization through performance metrics and resource insights
  • Real-time business analytics by linking IT operations to business outcomes
  • Proactive reliability with automated workflows that reduce downtime and manual effort

Modern cloud environments need an expanded approach to observability. Learn more about how Dynatrace can help you say goodbye to cloud complexity. Or explore it for yourself in our public sandbox environment.

The post Transforming Azure Data Factory operations with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/transforming-azure-data-factory-operations-with-dynatrace/feed/ 0
Enhance efficiency and compliance with automated AWS tag change triggers: A step-by-step guide https://www.dynatrace.com/news/blog/srg-aws-tag-changes/ https://www.dynatrace.com/news/blog/srg-aws-tag-changes/#respond Wed, 02 Apr 2025 15:52:01 +0000 https://www.dynatrace.com/news/?p=68498 Site Reliability Guardian

Streamlining site reliability at scale can be daunting, particularly with large-scale AWS environments and architecture that rely on hundreds—or even thousands—of Amazon EC2 instances. However, you can simplify the process by automating guardians in the Site Reliability Guardian (SRG) to trigger whenever there are AWS tag changes, helping teams improve compliance and effectively manage system […]

The post Enhance efficiency and compliance with automated AWS tag change triggers: A step-by-step guide appeared first on Dynatrace news.

]]>
Site Reliability Guardian

Streamlining site reliability at scale can be daunting, particularly with large-scale AWS environments and architecture that rely on hundreds—or even thousands—of Amazon EC2 instances. However, you can simplify the process by automating guardians in the Site Reliability Guardian (SRG) to trigger whenever there are AWS tag changes, helping teams improve compliance and effectively manage system performance.

This step-by-step guide will show you how to configure your architecture to trigger guardians whenever EC2 tags are updated. Note that EC2 is an example; this guide can be made to work generically for tag changes on any AWS resource. By the end of this guide, you’ll be ready to automate guardians at scale and optimize Amazon EC2 management with ease.

Why automate guardians for AWS tag changes?

Before diving into the technical setup, here’s why automating guardians whenever EC2 tags change is beneficial for your organization:

  • Greater efficiency: Automatically triggering guardians removes the need for manual intervention, saving time for DevOps or site reliability engineering (SRE) teams and allowing for more efficient resource management at scale.
  • Better compliance: Automating guardians ensures critical policies and checks are consistently applied after changes across your architecture, improving security and compliance efforts.
  • Cost optimization: Immediate responses to tag changes lead to informed decisions about scaling, shutting down unused instances, or fine-tuning resource efficiency.
  • Proactive site reliability: Automated guardians can monitor the four golden signals, enabling proactive reliability measures.

Now, let’s get started with the setup!

Step 1: Create an API token

Step 1: Create an API token

First, create an API token to integrate AWS services with Dynatrace for guardian automation.

  1. Log into your Dynatrace tenant

Log in to your Dynatrace tenant and note the first part of the URL (for instance, “abc12345”), which is your tenant ID.

  1. Access token settings

Press Ctrl + K or CMD + K and search for “Access Tokens” within Dynatrace.

  1. Generate a new token

Create a new access token and assign it “bizevents.ingest” permissions.

  1. Save the token

Copy and securely store the token, which looks like “dt0c01.*****.*****”. You’ll use this later during configuration.

Step 2: Create the EventBridge connection

Create the EventBridge connection

Configure invocation

Create the EventBridge connection

Amazon EventBridge acts as the bridge between AWS and Dynatrace. Here’s how to set it up:

  1. Navigate to Amazon EventBridge

Log in to your AWS Management Console and go to EventBridge > Connections.

  1. Recreate the cURL command

You can use this cURL command as a reference to establish your connection:

curl -X POST \
'https://abc12345.live.dynatrace.com/api/v2/bizevents/ingest' \
-H 'Authorization: Api-Token dt0c01.*****.*****' \
-H 'Content-Type: application/cloudevent+json' \
-d '{…}'
  1. Set the Authorization Method

Create a new EventBridge connection with the Authorization Method set to “API Key” and use the API token from Step 1 as the value (i.e., “Api-Token dt0c01.*****.****”).

Reminder: The API token is a sensitive value and should be stored in an encrypted format using a tool like AWS Secrets Manager. When following this guide, AWS Secrets Manager is already used.

Step 3: Define and configure EventBridge rules

Event pattern

EventBridge rules define the exact conditions for triggering guardians:

  1. Specify the input template:

Create an input template to modify your event data:

{
  "specversion": "1.0",
  "id": "<id>",
  "source": "aws.<source>",
  "type": "ec2.tag.change",
  "time": "<time>",
  "aws.region": "<region>",
  "aws.eventbridge.rule.arn": "<aws.events.rule-arn>",
  "aws.resources": <resources>,
  "data": <detail>
}
  1. Set the event pattern

Create a rule in EventBridge with the following event pattern:

{
  "source": ["aws.tag"],
  "detail-type": ["Tag Change on Resource"],
  "detail": {
    "service": ["ec2"],
    "resource-type": ["instance"]
  }
}
  1. Apply targets and permissions

Apply targets and permissions

Apply targets and permissions

Assign targets and permissions to ensure successful data ingestion into Dynatrace. Use an IAM role to permit EventBridge to call your API destination.

Input transformer

The input transformer should be set as follows:

{"detail":"$.detail","id":"$.id","region":"$.region","resources":"$.resources","source":"$.source","time":"$.time"}

Step 4: Test tag changes on Amazon EC2 instances

To validate your configuration, perform the following:

  1. Change a tag

Modify a tag by going to your Amazon EC2 instances in the AWS Management Console. For instance, update the “Environment” tag with a new value.

  1. Verify event logging

Check the EventBridge console to ensure your tag change triggered the appropriate event.

  1. Confirm data in Dynatrace

Within Dynatrace, press CMD/Ctrl + K and search for “Notebooks.” Create a new notebook and run the following query:

fetch bizevents | filter event.type == "ec2.tag.change"

If the query returns results, your configuration is working correctly.

Test tag changes on Amazon EC2 instances

Step 5: Set up the guardian

  1. Create a new guardian

Set up the guardian

In Dynatrace, search for “Site Reliability Guardian” (`CMD/Ctrl + K`) and create a new guardian. For best practices, use the “Four Golden Signals” template.

  1. Automate the workflow

Set up the guardian

Either on the overview page showing all guardians or on the analysis page of a selected guardian, click the Automate button. This will generate a workflow that triggers the guardian based on incoming bizevents (Business events). Configure the event type as `bizevent` and set the filter query to:

event.type == "ec2.tag.change"

Set up the guardian

  1. Add a pause

Set up the guardian

To allow your systems to stabilize before triggering the guardian, add a “wait before” step. For example, set a delay of 600 seconds (10 minutes).

Example timeline:

  • 06:59: Tag changed on EC2 instance.
  • 07:00: EventBridge triggers the workflow.
  • 07:10: Guardian is executed after 10-minute pause.
  1. Save the workflow

Save your final workflow to activate the automation.

Step 6: Validate and monitor the setup

Perform end-to-end validation by changing an EC2 tag again. Confirm the following:

  • The tag change event reaches Dynatrace.
  • The workflow triggers the guardian.
  • The guardian results appear in Dynatrace (e.g., heatmaps or relevant logs).

Run the following query in Dynatrace for additional monitoring:

fetch bizevents | filter event.type == "ec2.tag.change"

You should see log entries confirming the successful execution of your guardian process.

Achieve more with Site Reliability Guardian

In this blog, we highlighted the significant benefits of automating Site Reliability Guardian  triggers for Amazon EC2 changes. With automation, SRG helps engineering teams achieve efficiency, improved compliance, and cost optimization.

Learn more about Site Reliability Guardian in our documentation page.
Looking for more insights and support? Join the Automation Guild.

The post Enhance efficiency and compliance with automated AWS tag change triggers: A step-by-step guide appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/srg-aws-tag-changes/feed/ 0
Create simple workflows to automate alerts during development https://www.dynatrace.com/news/blog/create-simple-workflows-to-automate-alerting/ https://www.dynatrace.com/news/blog/create-simple-workflows-to-automate-alerting/#respond Wed, 22 Jan 2025 17:29:22 +0000 https://www.dynatrace.com/news/?p=67426

Manual processes make it tough for engineering teams to keep up with modern software complexity. Simple Workflows, a new feature in Dynatrace, enables developers and SRE teams to quickly automate single-step tasks—like sending a Slack notification for a particular staging exception—without incurring any additional costs. This lightweight approach ensures teams catch issues early, reduce toil, and speed up their development cycle.

The post Create simple workflows to automate alerts during development appeared first on Dynatrace news.

]]>

Traditional monitoring approaches often require manual scripting and integration to get alerted about production-threatening issues in pre-production environments.
Dynatrace Simple Workflows make this process automatic and frictionless—there is no additional cost for workflows. By offering a single trigger and a single task, you can rapidly set up an action (like Slack or email notifications) to detect specific exceptions in your services. This means fewer surprises when deploying to production and more time spent delivering valuable features.

Why manual alerting falls short

As your product and deployments scale horizontally and vertically, the sheer volume of data makes it impossible for teams to catch every error quickly using manual processes. Your teams want to iterate rapidly but face multiple hurdles:

  • Increased complexity: Microservices and container-based apps generate massive logs and metrics.
  • Unstructured overview: Manually scanning logs or waiting for someone to notice an error in staging is time-consuming.
  • Siloed tooling: Stitching multiple tools and scripts together can create friction and decrease your deployment velocity.

As a result, minor issues in staging can slip into production, leading to potential downtime and a poor user experience. Developers need a way to quickly set up alerts for targeted pre-production exceptions without incurring steep costs or heavy overhead.

Introducing Dynatrace Simple Workflows for early alerting

Dynatrace Simple Workflows helps your team overcome the challenges of manual processes by offering easy, no-cost automation for single-step tasks.

Go to Workflows and start creating a new workflow. By default, the Simple Workflow will be selected to give you the most cost-effective experience.

You can select any trigger that’s available for standard workflows, including schedules, problem triggers, customer event triggers, or on-demand triggers.

After you’ve selected + Workflow, you can name your workflow and select a specific trigger.

Here, you can select a specific event or a timed trigger like a cronjob. Additionally, with an on-demand trigger, you can trigger the workflow through REST API from your CI/CD, scripts, or any other automation flows.

After specifying your trigger, you can select a specific action. This action will be executed each time an event occurs, a timer kicks in, or an on-demand request triggers the action. You can easily send Slack/email messages to your teams, create JIRA issues, or trigger PagerDuty alerts

Simple workflow to send Slack message in Dynatrace screenshot

Safeguard your pre-production environments.

Imagine you’re running a busy e-commerce platform composed of multiple containerized microservices running on Kubernetes. Each microservice is responsible for a specific and critical business function—managing user sessions, processing payment and authentication, and tracking inventory.

To orchestrate the different logging services, you use Fluent Bit to forward these logs to your centralized logging system, like Dynatrace.

Recently, in production, your team encountered an out-of-memory error exception that killed all the processes for Fluent Bit. This caused you to lose complete visibility of your containers’ logs, performance, and error data, and you could not tell if the system was down or not.

kubernetes.event.message": "Memory cgroup out of memory: Killed process 1146778 (fluent-bit) total-vm:952432kB, anon-rss:303872kB, file-rss:1108kB, shmem-rss:0kB, UID:1000 pgtables:1424kB oom_score_adj:-997",
kubernetes.event.reason": "OOMKilling",

The error message indicates that the Linux cgroup assigned to your Fluent Bit container has reached its memory limit. Because Kubernetes enforces memory constraints for each pod, the kernel OOM (OutOfMemory) killer forcibly terminates the Fluent Bit process once memory usage exceeds the threshold set in the container/pod specification.

You have fixed the error, and as a follow-up to your RCA documentation, you want to ensure you’re alerted of these errors early on staging, so you want to set up a quick alert on staging for your feature branch.

For this example, we go to Simple Workflows and select Trigger > Davis® event trigger to find these out-of-memory errors. You can learn more about event triggers in Dynatrace Documentation.

Simple workflow trigger in Dynatrace screenshot

When Davis AI event triggers find these, we will set a simple Slack or Email notification for our team so the team can ensure the fix is working and other future changes are caught early.
If you ever need more advanced capabilities—like custom code, multi-step logic, or conditional tasks—you can seamlessly upgrade to standard workflows with just a few clicks.

This ensures you’re never locked into an overly complex or expensive solution before you’re ready. It’s as simple as that!

Deep dive into Simple Workflows

How Simple Workflows work:

  • Simple Workflows provide core Dynatrace automation capabilities
  • You can use any trigger type, problem-event triggers, schedule-based triggers, or external webhooks.
  • They trigger a single, out-of-the-box action: for example, Sending a Slack Notification, Creating a Jira Issue, Sending an Email, or Executing an HTTP Request.
  • Unlike multi-step workflows, these don’t consume workflow hours.

Limitations:

  • You can’t add more than one task.
  • JavaScript actions, task conditions, and options are not allowed.

What’s next for Workflows?

Looking ahead, Dynatrace aims to expand the usability of Simple Workflows with more out-of-the-box actions and deeper integrations.

Both simple and standard workflows will benefit from future improvements like:

  • A growing catalog of out-of-the-box actions
  • Introduction of draft workflows, enabling you to work on your workflows without impact until you’re satisfied with your changes
  • Watch Workflows: Get notifications whenever a workflow fails, is updated, or throttled.

For now, Simple Workflows delivers a fast way for developers and SREs to receive automated alerts—exactly when needed—at no extra cost.

Ready to try Simple Workflows?

Check out the Simple Workflows documentation to explore more use cases. For advanced functionality—like multi-step actions or running JavaScript—it’s free to upgrade to a standard workflow and take advantage of the full Dynatrace AutomationEngine capabilities.

With Simple Workflows, you can quickly eliminate manual monitoring steps, reduce time-to-detection, and keep staging errors where they belong—out of production.

The post Create simple workflows to automate alerts during development appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/create-simple-workflows-to-automate-alerting/feed/ 0
Don’t just react: How executives can predict and prevent outages to maximize availability https://www.dynatrace.com/news/blog/dynatrace-for-executives-improved-availability/ https://www.dynatrace.com/news/blog/dynatrace-for-executives-improved-availability/#respond Thu, 03 Oct 2024 14:15:16 +0000 https://www.dynatrace.com/news/?p=65679 Dynatrace for Executives: Improved Availability

I’ve seen firsthand the sleepless nights and high-stress environments that come with keeping digital services up and running in production. The stakes are high, and the pressure to deliver fast while maintaining uptime and preventing outages is relentless. With Dynatrace, executives can now benefit from predicting and preventing issues before customers are impacted and reducing […]

The post Don’t just react: How executives can predict and prevent outages to maximize availability appeared first on Dynatrace news.

]]>
Dynatrace for Executives: Improved Availability

I’ve seen firsthand the sleepless nights and high-stress environments that come with keeping digital services up and running in production. The stakes are high, and the pressure to deliver fast while maintaining uptime and preventing outages is relentless. With Dynatrace, executives can now benefit from predicting and preventing issues before customers are impacted and reducing the need to react. And when outages do occur, Dynatrace AI-powered, automatic root-cause analysis can also help them to remediate issues as quickly as possible. The end goal, of course, is to optimize the availability of organizations’ software.

Key insights for executives:

  • Predict and prevent outages before they happen with a unique combination of causal, preventive, and generative AI—enhanced by agentic AI for autonomous action
  • Remediate faster with automatic root cause analysis fueled by deterministic AI
  • Prioritize incidents based on customer impact insights from end-to-end traces
  • Automate to scale proactively and self-heal systems before customers are impacted

I realized that automating root-cause analysis requires a comprehensive approach: observing end-to-end and full-stack with deep insights, unifying all data in real time with up-to-date topology, and applying causal AI that learns instantaneously to handle cloud-native dynamics. Combining multiple types of AI made Dynatrace even stronger, and enables auto-optimize, auto-prevention, and auto-remediation all in one. However, the ultimate goal goes beyond technical excellence. For executives, the real business need is understanding customer impact—which is why it never made sense to me to just monitor servers but make end-to-end observability essential. This is what we uniquely solved for our customers with Dynatrace.

Respond to issues before they impact your customers

For executives, IT outages are a major headache. For issues that can’t be prevented in the first place, the next best option is to resolve issues faster than customers notice. Being faster, however, requires automation.

As the name Dynatrace suggests, dynamic tracing is at the heart of what we do. Dynatrace traces end-user interactions deep into the full stack of server-side activity to understand dependencies, allowing the platform to quantify the impact, qualify the situation, and prioritize actions. A power-of-three approach to AI, complimented with agentic AI, fuels automatic root-cause analysis to pinpoint the culprit amongst millions of service interdependencies and lines of code faster than humans can grasp.

Cloud technology complexity with billions of dependencies has outgrown human ability to manage and requires AI to analyze and comprehend. Dynatrace AI increases efficiency by magnitudes and prevents alert storms. This means you can avoid finger–pointing and war rooms, and dev teams’ productivity and happiness improve, eliminating business risk alert fatigue. Session replay capabilities provide visual proof and incident context so that teams can more easily understand and act upon the root cause. Automatic root cause analysis with Dynatrace can ultimately reduce mean time to repair (MTTR) by 90% or more.

Dynatrace is widely recognized for its causal, predictive, and generative AI capabilities, which can predict and prevent issues and automatically identify root causes, maximizing availability.

As responsibilities shift left due to the increased use of cloud-native technologies, development teams take more control over production deployments. While I am excited that the people who create software are also responsible for it—in contrast to “throw over the wall” approaches—it poses consistency and compliance challenges in larger organizations. That’s why we have Dynatrace extended (not shifted) to the left to address both needs: developers have easy and safe access to staging and production deployments while central SRE and DevOps teams have the scalable and automatic observability they need to remain compliant, consistent, and resilient. Finally, a standardized approach to observability coupled with self-service for departmental users reduces tool sprawl and complexity.

Gone are the days when executives could afford for their teams to stare at dashboards 24/7 to manually interpret data and act on runbooks. By unifying observability data and applying advanced AI, Dynatrace progresses to a new generation of AIOps that can predict and prevent issues and leverage automation for self-healing.

Predict and prevent outages with AI

The 2025 State of Observability Report found that 100% of business leaders are now using AI in their operations, with the top anticipated benefits being real-time anomaly detection (41%) and improved detection and response to security risks (37%). In this journey, many organizations have investigated AIOps tools to improve pattern analysis and noise reduction, as most of these solutions provide only correlation, not true causation. Even worse, the idea that such systems learn from past outages is flawed, as training would require thousands of production outages that no executive can afford.

So, to truly predict and prevent issues, the complexity of systems must be captured instantaneously and continually assessed in full context, through AI that maps causation in real-time. Dynatrace addresses this need with causal, predictive, and generative AI capabilities in a single framework. This approach eliminates the need for learning from past outages and enables a highly automated software delivery process, maximizing resilience.

Moreover, along with the maturity of the market to use agentic AI and AI overall, the ability to make use of Dynatrace capabilities has expanded too – since we have pioneered automation of operations. On this path, we address the skepticism in AI usage by fusing deterministic AI and agentic AI, making AI more reliable. And we see executives are more willing to drive steps towards more proactive automation and open the doors to autonomous operations.

IT teams can also embed quality gates into their workflows so they continually meet the thresholds for user experience defined through service-level objectives (SLOs). As a result, they can predict capacity demands based on seasonal patterns and use causal dependencies to automatically capture and prevent problems as they emerge.

Improving availability to meet ever-growing customer expectations requires high grades of automation for scale, agility, and resilience. This includes auto-scaling, overload protection, auto-remediation, auto-rollback, auto-quality-gating, and more. Eventually, the goal is to arrive at self-healing through autonomous cloud operations.

Therefore, platform engineering emerges as a discipline for a holistic approach to software, infrastructure and delivery, with a relentless aim to automate. Automation, however, should not be done in isolation of tech. It needs to execute in the context of the business, which requires insights into business-impacting metrics including end-user experiences, public API call success rates, learning from seasonal changes, and strategic business considerations such as cost vs. performance goals.

That’s where observability from Dynatrace goes far beyond “observing systems.” Dynatrace observability provides AI, analytics, and automation that integrates with platform engineering, continuous delivery, and automated operations. This greatly offloads DevOps, SRE, and operations teams from manual tasks and allows them to shift their work to automation tasks. Note that the work doesn’t get reduced. The key benefit is increased availability and security, faster software delivery, improved productivity, and cloud cost optimization.

New certification and security legislation projects, such as the Digital Operational Resilience Act (DORA) in Europe, are emphasizing the heightened expectations for digital systems availability. DORA further requires continuous compliance and the ability to report on the status, placing a heavy burden on organizations. This is where Dynatrace provides additional help and automation with the new Compliance Assistant app.

Likewise, since availability is affected by not only technical issues but also security threats, observability, and cloud security must converge to minimize availability issues. That is where Dynatrace AI and analytics—on top of unified observability and security data—raise the bar to proactively prevent problems and remediate them faster.

Want to learn more about all nine use cases? See the overview on the homepage.
In case you missed it, we hosted a must-see streaming event unveiling the innovations that are powering a new era of possibility for customers all over the world. Watch the on-demand recording now.

The post Don’t just react: How executives can predict and prevent outages to maximize availability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-for-executives-improved-availability/feed/ 0
AIOps strategy unlocks new possibilities for automation, customer satisfaction https://www.dynatrace.com/news/blog/aiops-strategy-unlocks-new-possibilities-for-automation-customer-satisfaction/ https://www.dynatrace.com/news/blog/aiops-strategy-unlocks-new-possibilities-for-automation-customer-satisfaction/#respond Tue, 01 Oct 2024 14:36:07 +0000 https://www.dynatrace.com/news/?p=65865 How to implement an AIOps strategy at scale

From managing complex IT environments to ensuring seamless customer experiences, the demands on IT departments have never been greater. To manage these complexities, organizations are turning to AIOps, an approach to IT operations that uses artificial intelligence (AI) to optimize operations, streamline processes, and deliver efficiency. One Dynatrace customer, TD Bank, placed Dynatrace at the […]

The post AIOps strategy unlocks new possibilities for automation, customer satisfaction appeared first on Dynatrace news.

]]>
How to implement an AIOps strategy at scale

From managing complex IT environments to ensuring seamless customer experiences, the demands on IT departments have never been greater. To manage these complexities, organizations are turning to AIOps, an approach to IT operations that uses artificial intelligence (AI) to optimize operations, streamline processes, and deliver efficiency. One Dynatrace customer, TD Bank, placed Dynatrace at the center of its AIOps strategy to deliver seamless user experiences.

Why AIOps?

AI for IT operations (AIOps) uses AI for event correlation, anomaly detection, and root-cause analysis to automate IT processes. It plays a crucial role in managing complex multicloud environments by streamlining operations and enhancing efficiency, reducing costs, and driving innovation.

Paired with an observability platform, AIOps identifies and helps to resolve cloud application performance and security issues, preventing problems before they disrupt operations. Its adoption is growing rapidly, driven by the explosion of data complexity that accompanies modern cloud IT environments. Valued at $17 billion annually, the AIOps market reflects its importance as large companies increasingly integrate AIOps and digital experience monitoring tools, with adoption expected to rise significantly in the coming years.

As a leader and trailblazer in the AIOps space, Dynatrace uses AI-powered root-cause analysis to provide precise actionable insights, enabling businesses to automate operations across the enterprise.

TD Bank adopts an observability-based AIOps strategy

As one of the 10 largest banks in the U.S., with $1.4 trillion in assets and 27 million customers, TD Bank places customers at the center of everything it does. As its enterprise monitoring team modernized the bank’s digital ecosystem from legacy on-premises data centers to a hybrid multicloud environment, TD Bank faced significant challenges with complexities its traditional monitoring tools couldn’t handle.

TD Bank’s modernized technology stack became increasingly intricate, leading to operational inefficiencies. The bank had accumulated multiple monitoring tools, each providing fragmented insights. This disjointed approach made it difficult to collaborate effectively and resolve issues promptly.

The Dynatrace unified observability platform provided TD Bank with a single source for answers, offering end-to-end visibility across its entire technology stack. Using Dynatrace at the center of its AIOps strategy, the TD Bank team reduced the number of IT incidents they were experiencing, improving customer trust.

Faster responses for greater reliability

A standout feature of Dynatrace is its ability to deliver rapid and precise answers. For TD Bank, this meant significantly reducing the time to identify and resolve transaction failures. With AI-driven certainty, the bank could instantly pinpoint the root cause of issues, leading to a 25% increase in proactive incident identification and a 20% faster response rate. This efficiency translated to a dramatic reduction in the transaction failure rate, from 0.16% to just 0.06%.

Cost optimization and efficiency

Using Dynatrace, TD Bank was able to consolidate its observability tools and achieve substantial cost savings. With its platform-based approach to end-to-end observability, Dynatrace enabled the bank to eliminate up to seven redundant monitoring solutions, reducing infrastructure and licensing costs by up to 45%. Beyond cost savings, this consolidation freed up TD Bank’s teams to focus on innovation rather than routine maintenance, driving further efficiency.

Enhanced customer satisfaction

For TD Bank, customer satisfaction is paramount. With the efficiencies stemming from Dynatrace AI capabilities, the bank reduced customer irritants by more than 60% and sped up issue resolution by 20%. With precise answers, TD Bank’s teams can quickly understand issues and resolve customer calls , enhancing the overall customer experience and building trust in the bank’s digital services.

How Dynatrace delivers on AIOps

AI-powered root-cause analysis

At the heart of Dynatrace AIOps capabilities is its power-of-three AI engine, Davis®. Using causal, predictive, and generative AI, Davis delivers detailed insights into issues, including their root cause and impact. For TD Bank, the technology has been instrumental in quickly identifying and resolving issues, ensuring minimal disruption to customer services.

Automated remediation

By automating routine tasks and responses to common issues, Dynatrace helps businesses like TD Bank achieve zero-touch operations. This automation not only improves efficiency but also ensures teams can address critical issues promptly, minimizing downtime.

Predictive analytics

Dynatrace AI-driven predictive analytics provide foresight into potential issues before they occur. For enterprises, this means staying ahead of the curve, preventing disruptions, and ensuring seamless operations. TD Bank has used these capabilities to anticipate and mitigate risks, ensuring a smooth banking experience for its customers.

The broader effect of AIOps: Transforming IT operations

AIOps is not just a tool; it’s a transformation strategy. By integrating AI into IT operations, businesses can achieve unparalleled efficiency, agility, and resilience. It is clear that the future of IT operations lies in AI, and Dynatrace is leading the charge. With its deep-rooted AI expertise and innovative AIOps platform, Dynatrace is transforming the way businesses operate. The success TD Bank has achieved demonstrates how Dynatrace helps unlock the hidden value in customer data and maximize the tangible benefits of AI-driven operations.

Discover the power of Dynatrace to unlock the future of IT operations and transform your business. Sign up for a free trial today and experience the difference Dynatrace AI can make.

The post AIOps strategy unlocks new possibilities for automation, customer satisfaction appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/aiops-strategy-unlocks-new-possibilities-for-automation-customer-satisfaction/feed/ 0
Next-level batch job monitoring and alerting: Elevate performance and reliability https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-elevate-performance-and-reliability/ https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-elevate-performance-and-reliability/#respond Fri, 27 Sep 2024 15:45:45 +0000 https://www.dynatrace.com/news/?p=65730 Dynatrace Security

Batch jobs are the backbone of automated, scheduled processes that execute tasks in bulk, such as data processing, system maintenance, or report generation. These jobs, which typically run in the background without user interaction, are critical and indispensable for handling large-scale operations efficiently.

The post Next-level batch job monitoring and alerting: Elevate performance and reliability appeared first on Dynatrace news.

]]>
Dynatrace Security

As batch jobs run without user interactions, failure or delays in processing them can result in disruptions to critical operations, missed deadlines, and an accumulation of unprocessed tasks, significantly impacting overall system efficiency and business outcomes. The urgency of monitoring these batch jobs can’t be overstated.

Monitor batch jobs

Monitoring is critical for batch jobs because it ensures that essential tasks, such as data processing and system maintenance, are completed on time and without errors. Failures, delays, or resource issues can lead to operational disruptions, financial losses, or compliance risks. Continuous monitoring enables early detection of problems, allowing quick remediation and maintaining business continuity.

Most jobs provide detailed information about job execution, including status, errors, and processing times in logs. The first step in monitoring batch jobs is to ingest these logs into Dynatrace. This is achieved by identifying the log files generated by the batch job program.

Apply basic filtering to ensure the availability of batch job-related logs. In this case, filter the logs based on relevant phrases or keywords.
Figure 1. Apply basic filtering to ensure the availability of batch job-related logs. In this case, filter the logs based on relevant phrases or keywords.

In this case, batch job statuses are constantly written from the deployment name get-cc-status-*. Thus we can create a rule in Dynatrace to ingest these logs via OneAgent without making any changes to the container, cluster, or host. Logs can also be ingested from various sources, including OpenTelemetry and Fluentbit.

A great reference is our blog post, Leverage edge IoT data with OpenTelemetry and Dynatrace, in which we documented the required steps to parse and ingest a single JSON log file into Dynatrace via OpenTelemetry.

Once logs are ingested, parsing the key messages is crucial. Below is a sample query that demonstrates how batch jobs can be parsed to extract important fields:

fetch logs
| filter matchesPhrase(content, "JOBS") AND matchesPhrase(content, "RunID")
| filter matchesValue(dt.entity.host, "HOST-HOSTID12345678")
| parse content , "
LD 'JOBS.' WORD:Job
LD 'RunID ' STRING:RunId
LD:status"
| fields timestamp, Job, status, content, RunId
| filterOut status == "."
| fieldsAdd start_time=if(contains(content,"started."),timestamp)
| fieldsAdd end_time=if(contains(content,"ended normally."),timestamp)
| fields timestamp, content, Job, status, RunId,start_time,end_time
Parsing the log lines that have critical data related to batch job status
Figure 2. Parsing the log lines that have critical data related to batch job status

Now that we can parse critical information, we can make informed decisions. However, it’s important to know if a job that started has ended within the expected timeframe. When a batch job exceeds its allotted time, the issue must be quickly identified and remediated.

Capture the time difference between two log entities

We use JavaScript within Dynatrace Dashboards to determine whether a previously started job was successfully completed. This three-level approach helps track how long a job took to complete and identifies any stuck jobs.

  1. Identify the unique property of each job and initialize its structure.
    const batch = {};
    
    /* Reiterate through each record and populate the data-structure*/
    for (const record of recordSet) {
      const runId = record['RunId'];
      if (!batch[runId]) {
        batch[runId] = {
          Job: record["Job"],
          run_id: runId,
          Status: "",
          JobStarted: null,
          JobEnded: null,
          Duration: "NA"
        };
      }
    }
  2. Process each job’s start time, end time, and status from the DQL parsed output.
    if (record["start_time"]) batch[runId].JobStarted = utcToLocal(record["start_time"]);
    if (record["end_time"]) batch[runId].JobEnded = utcToLocal(record["end_time"]);
    
    if (!statusLocked[runId]) {
      let status = record["status"]?.trim() || "";
    
      if (status.toLowerCase().includes("ended with return code")) {
        batch[runId].Status = "Failed";
        statusLocked[runId] = true;
      } else if (status == "started.") {
        batch[runId].Status = "Running";
      } else if (status == "ended normally.") {
        batch[runId].Status = "Completed without errors";
        statusLocked[runId] = true;
      } else {
        batch[runId].Status = status;
      }
    }
  3. Update the job status based on specific conditions (running, failed, completed).
    /* Leverage pre-populated data to identify duration for the completed jobs*/
    for (const runId in batch) {
      const job = batch[runId];
    
      if (job.JobStarted && job.JobEnded) {
        const startTime = new Date(job.JobStarted);
        const endTime = new Date(job.JobEnded);
        const duration = endTime - startTime;
    
        job.Duration = `${duration / 1000} seconds`;
      }
    }

Resources for the dashboard and workflow mentioned above can be found in this GitHub repository.

Individual batch job status with processing times and status
Figure 3. Individual batch job status with processing times and status
Advanced statistics for further analysis of batch jobs (median duration and job by status)
Figure 4. Advanced statistics for further analysis of batch jobs (median duration and job by status)

Correlate the impact of batch jobs with the application

Batch jobs should not impact applications because they run in the background. While they consume resources, they shouldn’t impact resource usage or client-facing applications. We can use Dynatrace Grail™ data lakehouse for unified observability data.

Correlate batch job runs with Application and Service resource utilization
Figure 5. Correlate batch job runs with Application and Service resource utilization

Adjust log parsing to account for varying log patterns

No two batch jobs are the same, and the log patterns you encounter might differ from what you see here. You can achieve the same results by parsing the logs. Parsing logs, as shown above, can be done using DPL Architect.

DPL Architect is a handy tool, accessible through the Notebooks app, that helps you quickly extract fields from records. It helps create patterns, provides instant feedback, and allows you to save and reuse DPL patterns, for faster access to data analytics use cases. This blog post offers further details about DPL architect.

Alerting for long-running or failed batch jobs

Constantly monitoring a dashboard isn’t practical, so you need automated alerting. Dynatrace workflows can check the status of batch jobs every 15 minutes and send alerts for failures or long-running jobs. These alerts can trigger actions or notifications sent via Slack, Teams, or as a ticket in your IT service management tool. In this example, the notifications are sent via email.

Automate batch job alerting and reporting
Figure 6. Automate batch job alerting and reporting

Conclusion

Monitoring batch jobs is essential to ensure they run smoothly and within expected timeframes. We can effectively identify issues such as long-running or failed jobs by ingesting logs into Dynatrace, parsing critical job information, and using custom logic to track job completion times. Implementing automated alerts triggering actions and notifications ensures proactive management, allowing teams to quickly resolve problems and maintain operational efficiency. With these tools in place, organizations can improve the reliability and performance of their batch-processing systems.

Use the approach detailed in this blog post to implement advanced batch job monitoring in your environment. Download the dashboards and Notebooks from this GitHub repository and start your automation journey today.

What’s next

In a future blog post, we’ll show how batch job management can be efficiently orchestrated using workflows and predictive analysis to schedule and run jobs optimally. With Davis® AI identifying root causes, workflows can be used to stop erroneous batch executions.

Additionally, Davis® AI prediction analysis, in conjunction with workflows, can reschedule or pause jobs to ensure optimal resource utilization, preventing any negative impact on the application landscape.

Download the Dashboards and Notebooks used in this blog post from our GitHub repository.

The post Next-level batch job monitoring and alerting: Elevate performance and reliability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/next-level-batch-job-monitoring-and-alerting-elevate-performance-and-reliability/feed/ 0
Automate digital excellence with Dynatrace Synthetic Monitoring and Workflows https://www.dynatrace.com/news/blog/dynatrace-synthetic-monitoring-and-workflows/ https://www.dynatrace.com/news/blog/dynatrace-synthetic-monitoring-and-workflows/#respond Thu, 18 Jul 2024 19:52:17 +0000 https://www.dynatrace.com/news/?p=64828 Dynatrace Synthetic Monitoring graphic

Managing IT infrastructure in a rapidly evolving digital environment can feel like playing a Jenga game. Dynatrace® Synthetic Monitoring integrated with visual workflows offers a robust solution for maintaining stability and growth. Dynatrace now combines automation and synthetic monitoring to respond to events, execute monitors, and assess impacts on user experience. This integration supports use cases like automatic release validation, digital infrastructure change validation, and custom synthetic scheduling, ensuring that all modifications meet performance and functional criteria.

The post Automate digital excellence with Dynatrace Synthetic Monitoring and Workflows appeared first on Dynatrace news.

]]>
Dynatrace Synthetic Monitoring graphic

In today’s rapidly evolving digital environment, organizations face increasing pressure from customers and competitors to deliver faster, more secure innovations. The complexity of IT infrastructure management continues to grow, and each new application deployment and every change in infrastructure potentially impact user experience.

In some cases, managing your digital infrastructure is like playing Jenga, where the placement of every piece is critical, and the smallest change can destabilize the entire structure. On one hand, the complexity of systems demands precise control; on the other, staying competitive requires frequent updates and rapid service enhancements. This dynamic creates a challenge: keeping the tower stable while continuously adding and removing pieces.

Jenga game with hand

To keep up with current demands, DevOps and platform engineering teams need a solution that can fully embrace and understand complexity, delivering precise answers that enable the creation of trustworthy automation. The effectiveness of this automation relies on the quality of the underlying data. Observability, therefore, has become crucial in DevOps, offering insights into IT infrastructure stability, performance, and user experience. The DevOps Automation Pulse 2023 report notes that 78% of organizations employing observability-driven automation experience faster incident response and resolution times.

Synthetic monitoring enhances observability by enabling proactive testing and monitoring systems to identify potential issues before they quickly impact users. Returning to the Jenga metaphor, synthetic monitoring observes the tower from a distance, from the end user’s perspective, and triggers instability warnings immediately. Incorporating synthetic monitoring and observability-driven automation can greatly streamline the workflow for DevOps teams, allowing for continuous improvements in system reliability and efficiency.

Automation + Synthetic = Perfect match

This is why we integrated Dynatrace Synthetic Monitoring into Workflows. With this enhancement, Dynatrace can respond to any event and execute synthetic monitors within your workflows to assess the impact of events on user experience. Depending on the outcome, workflows can notify your teams by creating a Jira ticket, sending a Slack message, or initiating a remediation process. This enhancement improves reaction time to incidents impacting user experience and simplifies incident management. It reduces the time spent verifying the impact of changes and engages DevOps or platform engineers only when necessary, allowing them to focus on more strategic initiatives.

With the intuitive and easy-to-use web UI of Workflows, you can quickly build automation based on synthetic monitors with just a few steps, thanks to a dedicated Synthetic for Workflows task.

In the first step, you choose which monitors to execute. You have three options:

  • Select a fixed list of monitors.
  • Provide a list of tags; all monitors with the tags will be selected.
  • Provide a list of applications; all monitors associated with those applications will be selected.

Workflows diagram in Dynatrace screenshot

You can also provide a list of monitors, tags, or applications in the incoming event and extract the list using an expression, which allows you to build a generic workflow.

Jinja expression in synthetic monitors settings in Dynatrace screenshot

Use case: Automated release validation

Integrating Dynatrace Synthetic Monitoring into delivery pipelines is a strategic move that ensures all critical user journeys are validated as early as possible, whether a release involves minor bug fixes, updates to a microservice, or an entire monolithic application.

This integration is particularly valuable in progressive delivery models like canary deployments. In such scenarios, a new version of an application is initially rolled out to a small portion of users. By adding a synthetic user as “user zero,” the risk of negatively impacting real users is significantly reduced.

Combining Synthetic for Workflows with Site Reliability Guardian (which evaluates adherence to availability, performance, or security objectives based on synthetic results) automatically determines whether the given application version should proceed in the delivery pipeline. This integration facilitates the easy construction and implementation of solutions that enhance automatic release validation.

Workflows diagram with Dynatrace

Use case: Digital infrastructure change

The problem is not always in the application. As illustrated by the Jenga game metaphor, even minor changes in complex infrastructure can topple the entire tower. In 2021, Facebook and its other platforms, Instagram and WhatsApp, experienced a major outage. This global outage lasted for several hours and affected billions of users worldwide.

Interestingly, the applications were functioning perfectly; the root cause was a configuration change in the backbone routers.

Incorporating synthetic monitoring directly into deployment pipelines ensures that every infrastructure modification—regardless of scale—is automatically validated against performance and functional criteria.

This integration is particularly valuable within GitOps frameworks, where infrastructure changes are managed through code and deployed automatically. Synthetic monitors can simulate API clients and mimic user behavior, enabling IT teams to evaluate the impact of changes. Such preemptive testing significantly reduces potential downtime and mitigates performance degradation.

Synthetic for Workflows and Site Reliability Guardian can—continuously or on-demand—verify that all infrastructure modifications meet SLOs from the perspectives of users and API clients.

Use case: Custom synthetic scheduling

The traditional approach to scheduling synthetic monitors involves fixed intervals, which don’t always align with the varying operational demands of modern digital environments.

Workflows in Dynatrace provide a powerful mechanism for triggering tasks in response to various platform events. By leveraging Synthetic for Workflow combined with this event-driven architecture, you can achieve a level of customization in scheduling synthetic monitors that goes beyond the capabilities of the built-in synthetic scheduler.

This flexibility is particularly useful in scenarios where operational conditions can change rapidly. For instance, if a peak load period triggers a test failure, Synthetic for Workflow can initiate retries until the load stabilizes. This retry logic can help triage the problem and ensure that monitoring results are accurate and reliable.

Conclusion

By integrating Dynatrace Synthetic Monitoring with Workflows, Dynatrace offers a comprehensive solution for evaluating the impact of infrastructure changes, enhancing release validation, and improving custom scheduling. However, utilizing Synthetic for Workflows extends beyond these use cases, enabling the easy construction of any automation based on synthetic monitors, where assessing the impact on user experience is crucial.

For full details regarding Synthetic Monitoring for Workflows, go to Dynatrace Documentation.

The post Automate digital excellence with Dynatrace Synthetic Monitoring and Workflows appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-synthetic-monitoring-and-workflows/feed/ 0
Three ways Dynatrace can help to drive innovation through cloud modernization https://www.dynatrace.com/news/blog/dynatrace-for-executives-cloud-modernization/ https://www.dynatrace.com/news/blog/dynatrace-for-executives-cloud-modernization/#respond Thu, 18 Jul 2024 13:30:14 +0000 https://www.dynatrace.com/news/?p=64730 Dynatrace for Executives: Cloud Modernization

As executives, we drive change, balancing modernization speed with its risks. Technology—both a blessing and a curse—not only propels businesses forward but also adds complexity as developers introduce new innovations to enhance customer services and competitiveness. Anticipate future customers’ needs Anticipating customer needs three to five years ahead helps to reduce wasted investments into “wants” […]

The post Three ways Dynatrace can help to drive innovation through cloud modernization appeared first on Dynatrace news.

]]>
Dynatrace for Executives: Cloud Modernization

As executives, we drive change, balancing modernization speed with its risks. Technology—both a blessing and a curse—not only propels businesses forward but also adds complexity as developers introduce new innovations to enhance customer services and competitiveness.

Anticipate future customers’ needs

Anticipating customer needs three to five years ahead helps to reduce wasted investments into “wants” and directs them toward “needs” that future-proof the business.

This mentality has driven me to continuously innovate and reinvent Dynatrace®. My ongoing evaluation of how technology changes the way digital services are architected allowed me to recognize early on that change is on the horizon. The rise of cloud-native technologies, the convergence of observability and security, and the demand for actionable insights required a new approach to managing data at an exabyte scale, as existing databases could no longer keep up.

Change is constant

In our fast-paced world, success requires thinking big but acting small to create value quickly and sustainably. For cloud modernization, this means executives must change how software is built, operated, and secured; improve collaboration processes; and increase automation.

Dynatrace gives executives an indispensable platform for driving this change in the following three ways:

  • Enabling a modern AIOps strategy,
  • Accelerating software delivery, and
  • Making scarce engineering resources more productive.
Key insights for executives
  • Modern AIOps and AISecOps from Dynatrace get us closer to NoOps and NoSoc than ever with help of hypermodal AI
  • Early investment into automation pays off, and the 100 ready-made use cases  from Dynatrace accelerate software delivery with confidence
  • Extend to the left has become the modern shift left, and Dynatrace accelerates productivity with contextual analytics, AI, automation, and platform engineering

1. Go beyond traditional AIOps

The first wave of AIOps investment was about “noise reduction.” This has been helpful but falls short of the potential offered by the preventive NoOps and NoSOC approaches that many executives seek AI to enable.

With current hype causing a resurrection in AI investment, it is tempting to believe that this time, machine learning and generative AI will fulfill the promises of the past. However, while the advances in machine learning-based AI are a huge step up for many use cases, it is still problematic to apply it to prevent incidents and errors in IT systems. Why? Because training an AI requires errors, failures, and behaviors to occur many times to ‘learn’. While the exact numbers may have been reduced by the advances in generative AI, which executive wants to have service outages just to train AI to prevent them in the future? Even if it was possible to arrive at a trained model, it would quickly become obsolete as services get updated and new features introduced.

As we consider a way forward, I urge all executives to recognize that we are in the trough of disillusionment in the AI hype cycle. This is good news, as it allows us to think more rationally. We need to understand that there are multiple types of AI, each suited for different purposes.

Dynatrace is uniquely designed to help executives elevate their AIOps – and AISecOps strategy – to a different level by combining multiple types of AI in a single framework known as hypermodal AI: the power of predictive AI, causal AI, and generative AI for observability, security, and business use cases. Proven by thousands of customers in large-scale IT deployments, this approach delivers greater speed, automation, and precision.

Our hypermodal AI automatically infers the root cause of issues based on a real-time updated graph without needing to learn. Now, it is more feasible than ever to automate workflows for self-healing, security investigation, and preventive operations to deliver great software with confidence, all while enhancing security measures and boosting productivity.

2. Accelerate software delivery

One of the best features of the cloud and Kubernetes® is achieving most availability needs with minimal effort, a major improvement over the classic datacenter model. This allows executives to focus on accelerating software delivery. However, the inverse Pareto principle applies: achieving the final 20% of flawless, secure services requires 80% of the effort.

That’s why APIs have become my favorite feature of the cloud as the key to automate and orchestrate. This is where Dynatrace comes in. Dynatrace integrates with the cloud ecosystem and DevOps toolchain to enhance automation across software delivery, resilience, and security throughout the software lifecycle.<

Throughout the ten years since we embraced NoOps at Dynatrace, I understood the temptation to favor releasing new features over investing in automation. Automation always paid off. We have since developed over 100 ready-made use cases to support platform engineering across the software delivery lifecycle. From development and release to operation and flaw prevention, prediction, and resolution, Dynatrace offers a robust data analytics-driven automation platform.

We’ve seen the many benefits of investing in automation, including the following capabilities:

  • Releasing faster and securely with automated quality and security gates
  • Catching bugs earlier, before customers experience them
  • Preventing issues with predictive operations
  • Avoiding unnecessary high consumption and cost with causal and predictive auto-scaling
  • Empowering developers with context-rich insights derived from self-service observability and security
  • Orchestrating more intelligently with real-time user behavior and business data

In a nutshell, Dynatrace allows executives to accelerate software delivery with confidence.

Dynatrace Dashboards: visualize your complex hybrid cloud environments in real time, gaining insights into security and business performance.

3. Increase teams’ productivity

As Dynatrace CTO, one of the questions constantly on my mind is: how can I enable my team to be more productive?

Over the past 15 years, most of us have embraced the “shift left” ethos to empower software developers. The earliest iteration of this was the “you build it, you run it” mentality. However, given the responsibilities of creating enterprise-scale and secure software, the “extend left” ethos proves to be more successful and fitting for cloud modernization.

Extend to the left: The modern “shift left”

“Extend left” refers to sharing responsibility amongst developers and operations teams, through adding more self-service for developers while retaining consistency, tooling and knowledge management with central teams.

As neither full decentralization nor full centralization will be effective, a hybrid model, supported by platform engineering approaches, is much more likely to succeed. Centralizing the necessary expert knowledge within a platform engineering team enables rapid, secure, and safe software delivery. At the same time, this approach decentralizes innovation, making it accessible to many.

Dynatrace was created to enable precisely this approach, leveling up developer experience by providing self-service capabilities while allowing central safety and oversight maintenance. This gives executives the best of both worlds: decentralized autonomy supported by centralized governance and control.

Armed with the use cases across the three areas outlined here, executives can modernize their cloud operations faster and equip their teams with the capabilities they need to accelerate innovation confidently. As a result, they will be better placed to anticipate change and continuously reinvent their organization to stay ahead of the market.

Follow along the new “Dynatrace for Executives” blog series. In the coming weeks, I’ll dive deeper into each of the nine executive use case areas to help you unlock the potential of Dynatrace.
Want to learn more about all nine use cases? See the overview on the homepage.

The post Three ways Dynatrace can help to drive innovation through cloud modernization appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-for-executives-cloud-modernization/feed/ 0
Nine ways technology executives can get significant business value with the right observability platform https://www.dynatrace.com/news/blog/dynatrace-for-executives/ https://www.dynatrace.com/news/blog/dynatrace-for-executives/#respond Tue, 21 May 2024 12:00:10 +0000 https://www.dynatrace.com/news/?p=64050 Dynatrace for Executives

As a technology executive, you’re aware that observability has become an imperative for managing the health of cloud and IT services. You may not be aware of how much untapped value is waiting to be unlocked through the right observability platform. Data with context can improve your ability to deliver on your goals, modernize your […]

The post Nine ways technology executives can get significant business value with the right observability platform appeared first on Dynatrace news.

]]>
Dynatrace for Executives

As a technology executive, you’re aware that observability has become an imperative for managing the health of cloud and IT services. You may not be aware of how much untapped value is waiting to be unlocked through the right observability platform. Data with context can improve your ability to deliver on your goals, modernize your organization, and accelerate business transformation.

The Dynatrace platform enables executives to drive change faster, increase IT and R&D productivity, reduce business risks, optimize costs, and decrease carbon footprint. These outcomes are made easy through the platform’s unique ability to turn data into answers and action, in contextual, real-time, and cost-effective ways that were previously impossible.

Unearthing a goldmine of value

As founder and CTO of Dynatrace, I must constantly drive change. I also have the privilege of being “customer zero” for our platform, which enables me to continually discover where Dynatrace can deliver on more use cases to drive my team’s productivity and innovation. Change is my only constant.

Realizing that executives from other organizations are in a similar situation to my own, I want to outline three key objectives that Dynatrace’s powerful analytics can help you deliver, featuring nine use cases that you might not have thought possible.

Dynatrace for Executives: 3x3 use cases matrix

Drive innovation

To remain competitive, executives are seeking productivity gains while simultaneously driving modernization initiatives. Observability data presents executives with new opportunities to achieve this, by creating incremental value for cloud modernization, improved business analytics, and enhanced customer experience.

However, technology executives face a significant challenge getting answers in time, as their needs have evolved to real-time business insights that enable faster decision-making and business automation. Exploding volumes of data must be prepared, catalogued, stored in multiple, disconnected tools. The data must then be retrieved from data lakes and converted into rigid schemas. It can take data analysts months to extract insights and answer executives’ questions using these approaches.

With the latest advances from Dynatrace, this process is instantaneous. Unlike anything before, contextual analytics in Dynatrace provides answers to any question at any time, instantaneously. That’s because it does not require any pre-prepared schemas, and access to cold/hot storage is fully automatic and with zero latency. Moreover, it is fast, powered by its massively parallel processing data lakehouse.

As a result, organizations can reduce complexity, effort, and processing time to run powerful business analytics on exabytes of data in real time. Dynatrace enables executives to drive a stronger, data-driven organization by increasing automation and productivity.

Mitigate risk

To cope with serious business risks —including major outages, security breaches, or missing out on realizing AI’s value — executives require a modern, proactive approach. Dynatrace analytics capabilities, powered by hypermodal AI, enable executives to drive improved availability, strengthened security compliance, and heightened confidence in AI initiatives.

Executives are shifting to proactive risk management, aiming to prevent availability issues and expedite remediation. However, AI introduces new risks, such as increased software complexity, accelerated cyber-attacks, and potential regressions from rapid releases. Siloed teams and the reliance on disparate tools lead to manual intervention and delays, which are unsustainable given tightening regulations including DORA, NIS2, and the SEC’s four-day reporting rule.

Dynatrace uniquely solves this conundrum, enabling executives to use a new generation of AIOps and SecOps to predict and mitigate risk, rather than reacting to availability and security incidents. It does this by combining causal, predictive, and generative AI to uncover the deep context of issues using a unified source of observability and security data. Automated root-cause analysis and real-time risk analysis are only two examples that help executives get closer to the vision of self-healing operations and security.

Optimize cost

With the constant pressure to do more with less — or much more, much faster — executives must control cost and complexity. Dynatrace can help executives to achieve these goals by reducing tool sprawl, driving cost optimization, and meeting their sustainability goals.

Optimizing costs is a proven way to free up budgets for innovation. Young talent (our future executives) has a valid interest beyond making more money, as sustainability and green coding are vital to protecting both their own and our future.

As new waves of technology roll over us, executives are struggling to keep tool sprawl under control. Tool sprawl not only goes deep into our pockets, but also hampers consistency and productivity. Tens or even hundreds of DIY and commercial tools are being used to handle logs, metrics, traces, security events, and vulnerabilities all in their own way.

Insights are therefore dispersed in a multitude of data lakes, storage systems, and reporting platforms. This is inefficient and creates avoidable risks. The principle of “keep it simple, stupid” is more important than ever, translating to consolidating tools and making processes more consistent at higher grades of scalability and automation.

Dynatrace is uniquely placed to meet this need as it consolidates tools, storage, data, processing, and automation capabilities together in a single, unified platform. This reduces the number of moving parts and eliminates process inconsistencies, driving team productivity and increasing software delivery quality and security.

As a result, organizations can streamline processes by moving towards platform engineering and developer self-service portals to unburden engineers while increasing software quality and security at a higher consistency.

In the coming weeks, I’ll dive deeper into each of the executive use cases outlined above to help you unlock the potential of Dynatrace. In the meantime, find more at Dynatrace for Executives.

The post Nine ways technology executives can get significant business value with the right observability platform appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-for-executives/feed/ 0
Context-aware security incident response with Dynatrace Automations and Tetragon https://www.dynatrace.com/news/blog/context-aware-security-incident-response/ https://www.dynatrace.com/news/blog/context-aware-security-incident-response/#respond Fri, 03 May 2024 08:00:37 +0000 https://www.dynatrace.com/news/?p=63832 security incident response with Dynatrace Automation Engine and Tetragon

For the most severe threat scenarios, you want multiple layers of automated defenses, and not have to rely on humans to analyze the traces of an attack weeks after your system got compromised. Many security teams use runbooks to glue together tools, processes, events, and actions for security incident response. A runbook lays out the […]

The post Context-aware security incident response with Dynatrace Automations and Tetragon appeared first on Dynatrace news.

]]>
security incident response with Dynatrace Automation Engine and Tetragon

For the most severe threat scenarios, you want multiple layers of automated defenses, and not have to rely on humans to analyze the traces of an attack weeks after your system got compromised. Many security teams use runbooks to glue together tools, processes, events, and actions for security incident response. A runbook lays out the step-by-step instructions to follow when a security incident happens, when an emerging threat surfaces, or when your security tool reports suspicious behavior.

But runbooks that stitch together glamorous security tooling are merely decorations without automated workflows for incident detection and response.

The security community agrees on many high-level best practices in such situations, but we need a single platform solution to orchestrate application security, observability, and DevOps practices. Because every situation is a little unique, Dynatrace makes it easy to create custom runbooks using Dynatrace Automations, fine-tuned to your individual business risks.

In this blog post, we’ll demonstrate how to use Dynatrace Automations to build a runbook that combats sophisticated security incidents with honeytokens and eBPF-based detection. We show an end-to-end solution, starting with deploying policies in a Kubernetes cluster and ending in a pull request assigned to the responsible team, all without manual intervention.

To demonstrate the integration of external security tools into the Dynatrace platform, we use Tetragon for eBPF-based security monitoring. Using Kyverno, we can automatically kick attackers out of our cluster with network policies and harden our configuration with a GitOps workflow to prevent the same incident from happening again.

Better, faster application protection and security investigation

Defending against threats in cloud-native environments requires advanced capabilities that can provide insight into runtime application details and speed up evidence-driven security investigations, such as:

  • Runtime Vulnerability Analytics detects and analyzes vulnerabilities in the third-party dependencies of your applications, assisting you in reducing your overall risk profile.
  • Runtime Application Protection defends your applications from the inside against injection attacks, even novel yet unseen zero-day attacks.
  • The Security Investigator app helps security operators and security analysts to speed-up threat hunting and incident resolution, without losing data context.

Incident response and security investigations often require numerous manual steps. There is sufficient potential to automate many elements of this process. To demonstrate, we explore the following security nightmare:

  • An attacker exploited a vulnerability, breached our perimeter without setting off any alerts, and managed to get access to the file system of one of our application containers.
  • The attacker is now roaming freely in our container, searching for secrets in sensitive configuration files and abusing them to escalate privileges, move laterally to neighboring systems, or plant malware in our container.
  • Unfortunately, we forgot to follow Kubernetes security best practices, such as using role-based access control (RBAC), deploying network policies, and setting the security context.

We’re going to write an automated workflow to detect and respond to this incident as it unfolds.

How an attacker gains access to a file system
Figure 1: How an attacker gains access to our application container

Step 1: Automating the placement of honeytokens to create strong indicators of compromise

In this sophisticated threat scenario, we don’t know exactly what events or indicators of compromise we should be looking for. Thus, we turn to a slightly different strategy: Setting up traps for attackers. We know that an attacker eventually looks for interesting files in our container. So, let’s provide some interesting files as a lure. This honeytoken strategy levels the playing field. Usually, defenders must attempt to close all security holes while attackers only need to find one open hole. With our honeytoken strategy, attackers must watch their steps not to trip over our traps.

As a lure, we place an s3_token file in path /run/secrets/eks.amazonaws.com/, which is one location for secrets that attackers commonly visit when attacking Amazon EKS clusters. When they try to access the file, it will generate an alert.

We use Kyverno, a popular policy management solution for Kubernetes, to automatically generate and mount this honeytoken into every new container in our Kubernetes cluster. Here’s the Kyverno policy YAML file we can use:

# kyverno-policy.yaml 
# 
apiVersion: kyverno.io/v1 
kind: ClusterPolicy 
metadata: 
  name: generate-honeytoken 
spec: 
  # policy also applies to existing namespaces 
  generateExisting: true 
  rules: 
    - name: generate-honeytoken 
      match: 
        any: 
          - resources: 
              kinds: 
                - Namespace 
      exclude: 
        any: 
          - resources: 
              namespaces: 
                - default 
                - kyverno 
                - kube-system 
                - kube-public 
      context: 
        - name: random-token 
          variable: 
            jmesPath: random('[a-z0-9]{6}') 
      generate: 
        kind: Secret 
        apiVersion: v1 
        name: honeytoken 
        namespace: "{{request.object.metadata.name}}" 
        # if the policy is deleted, also delete secrets 
        synchronize: true 
        data: 
          data: 
            token: "{{ random('[a-z0-9]{16}') | base64_encode(@) }}" 
--- 
apiVersion: kyverno.io/v1 
kind: ClusterPolicy 
metadata: 
  name: mount-honeytoken 
spec: 
  rules: 
    - name: mount-honeytoken 
      match: 
        any: 
          - resources: 
              kinds: 
                - Pod 
      mutate: 
        patchStrategicMerge: 
          spec: 
            volumes: 
              # depends on the generate-honeytoken policy 
              - name: honey-volume 
                secret: 
                  secretName: honeytoken 
            containers: 
              # match any image 
              - (image): "*" 
                volumeMounts: 
                  - name: honey-volume 
                    readOnly: true 
                    subPath: token 
                    mountPath: /run/secrets/eks.amazonaws.com/s3_token

Save this to file kyverno-policy.yaml and apply it in your cluster:

kubectl apply -f kyverno-policy.yaml

After applying those two cluster policies, we already solved a big pain of deception technology: The highly labor-intensive and tedious deployment of honeytokens that is often configured and maintained manually.

Step 2: Alerting with automated context enrichment

Placing honeytokens is only half of the story. We also need to receive alerts when attackers access our honeytokens. In our scenario, we chose Tetragon to detect read attempts on our honeytoken. Tetragon is a popular eBPF-based security tool for composing sophisticated security policies. We only need two policies for our case, but software engineers Kornilios Kourtis and Anastasios Papagiannis illustrate advanced techniques for file monitoring with eBPF and Tetragon in their blog post.

In our scenario, we only need two tracing policies: One to detect access attempts to the honeytoken file and one to track TCP connections, so we also find the attacker’s IP address.

# tetragon-policy.yaml 
# 
# This file was originally authored by Tetragon developers and adapted by Dynatrace. 
# - https://github.com/cilium/tetragon/blob/main/examples/tracingpolicy/filename_monitoring.yaml 
# - https://github.com/cilium/tetragon/blob/main/examples/tracingpolicy/tcp-connect.yaml 
# 
apiVersion: cilium.io/v1alpha1 
kind: TracingPolicy 
metadata: 
  name: monitor-honeytoken 
spec: 
  kprobes: 
    - call: security_file_permission 
      syscall: false 
      return: true 
      args: 
        # (struct file *) used for getting the path 
        - index: 0 
          type: file 
        # 0x04 is MAY_READ, 0x02 is MAY_WRITE 
        - index: 1 
          type: int 
      returnArg: 
        index: 0 
        type: int 
      returnArgAction: Post 
      selectors: 
        - matchArgs: 
            - index: 0 
              operator: Prefix 
              values: 
                - /run/secrets/eks.amazonaws.com/s3_token 
--- 
apiVersion: cilium.io/v1alpha1 
kind: TracingPolicy 
metadata: 
  name: monitor-tcp-connect 
spec: 
  kprobes: 
    - call: tcp_connect 
      syscall: false 
      args: 
        - index: 0 
          type: sock

Save this to file tetragon-policy.yaml and apply it in your cluster:

kubectl apply -f tetragon-policy.yaml

Enable the automatic ingest of Tetragon alerts to Dynatrace

We already installed Dynatrace on Kubernetes so that OneAgent automatically picks up logs from our cluster. Tetragon emits alerts using a shell process, which is typically noisy and not monitored in Dynatrace by default. To enable Dynatrace to automatically capture Tetragon logs in Dynatrace, navigate to Settings > Processes and containers > Declarative process grouping and add a new monitored technology for the Tetragon export-stdout process group:

  • Process group display name: export-stdout
  • Process group identifier: export-stdout
  • Report process group: Always
  • Detection rule > Select process property: Command line
  • Detection rule > Condition: $contains(export-stdout)
Tetragon process group defined in Dynatrace
Figure 2: Enable Dynatrace to automatically capture Tetragon logs

To test if this setup is working, we open a shell in any of our containers and try reading the honeytoken that Kyverno previously placed. The honeytoken is only available in pods created after we apply the Kyverno policy.

kubectl exec -it -n your-namespace-name your-pod-name -- /bin/bash 
$ cat /run/secrets/eks.amazonaws.com/s3_token 
dyxt57xkazpw1cj5

Then, we open the Security Investigator app or the Notebooks app to track down the expected policy violation in Dynatrace:

fetch logs, from: now() - 1h 
| filter k8s.deployment.name == "tetragon" 
| parse content, "JSON:content" 
| filter content[process_kprobe][policy_name] == "monitor-honeytoken" 
| fields timestamp, content
Honeytoken monitoring expressed in DQL
Figure 3: Query Tetragon events using DQL

Great, we spotted this indicator of compromise!

Step 3: Auto-remediate with network policies and GitOps

We want to automatically react to such intrusion alerts to reduce the burden on the incident response team. With the Dynatrace Platform, you can manage such security runbooks in a central and unified way.

Let’s assume an attacker tripped over our trap by reading the honeytoken. As a next step, we leverage the deep, context-aware observability insights Dynatrace has collected, block the attacker’s IP address, and inform the security team. Here’s an automation workflow that executes the following in order:

  1. Listen for recent token access alerts using Dynatrace Query Language (DQL).
  2. Correlate the alerts with logs on TCP connection events to find the attacker’s IP address.
  3. Identify the owners of the affected resource and its associated source code repository.
  4. Create a pull request in that repository that sets a new security policy to block the attacker’s IP address.
  5. Create a ticket for the security team to follow-up on that incident.
GitOps automation workflow for security incident response
Figure 4: Context aware security incident response

Correlating token access alerts and network connections

The Tetragon alert on honeytoken access can’t provide the IP address of the related network session. But we can look for the parent process that established the TCP connection. Attackers rarely connect “directly” to your infrastructure; no port would be open to accept such a connection anyway. Instead, connections are established the other way around: Attackers connect from inside your infrastructure to their remote server, often referred to as the command-and-control (CnC) server. This attack technique is called opening a “reverse shell”. The resulting process tree typically consists of three processes: The malware itself, the shell it spawns, and the process that accesses the honeytoken.

Reverse shell attack technique
Figure 5: Reverse shell attack technique process tree

With just three DQL queries, we find the attacker’s IP address. Note that the following queries target the structure of this attack technique and don’t hard-code process names, which attackers can easily change.

  1. The first query finds the low-level process that accessed the honeytoken, for example, /bin/cat. We return the execution ID of its parent process, that is, the shell that called /bin/cat.
    fetch logs, from: now() - 10m 
    | filter k8s.deployment.name == "tetragon" 
    | parse content, "JSON:content" 
    | filter content[process_kprobe][policy_name] == "monitor-honeytoken" 
    |  fields timestamp, {content[node_name], alias: node_name}, 
            {content[process_kprobe][parent][exec_id], alias: parent_exec_id}
  2. The second query finds the shell (for example, /bin/sh) by its execution ID. The only reason for this query is to get the execution ID of the shell’s parent process, that is, the process that initially spawned the shell.
    fetch logs, from: now() - 10m 
    | filter k8s.deployment.name == "tetragon" 
    | parse content, "JSON:content" 
    | filter content[node_name] == {{ result("fetch_tetragon_token_logs").records[0].node_name }} 
    | filter content[process_exec][process][exec_id] == {{ result("fetch_tetragon_token_logs").records[0].parent_exec_id }} 
    | fields timestamp, {content[node_name], alias: node_name}, 
             {content[process_exec][parent][exec_id], alias: parent_exec_id}
  3. The third query finds processes that established network connections (for example, /usr/bin/python if the malware was written in Python) and filters on the execution ID returned from the previous query. The malware that spawned the shell also established the network connection. The Tetragon event also holds the remote IP address of the attacker’s server.
    fetch logs, from: now() - 10m 
    | filter k8s.deployment.name == "tetragon" 
    | parse content, "JSON:content" 
    | filter content[node_name] == {{ result("fetch_tetragon_shell_logs").records[0].node_name }} 
    | filter content[process_kprobe][policy_name] == "monitor-tcp-connect" 
    | filter content[process_kprobe][function_name] == "tcp_connect" 
    | filter content[process_kprobe][process][exec_id] == {{ result("fetch_tetragon_shell_logs").records[0].parent_exec_id }} 
    | fields timestamp,  
             {content[process_kprobe][args][0][sock_arg][daddr], alias:daddr},  
             {content[process_kprobe][process][pod][name], alias:pod_name}, 
             {content[process_kprobe][process][pod][namespace], alias:pod_namespace}, 
             {content[process_kprobe][process][pod][pod_labels], alias:pod_labels}, 
             {dt.entity.cloud_application_namespace, alias:namespace_id}

At this point, we got the intrusion event, and we got the attacker’s IP address. A typical security solution might create a generic security problem and assign it to – well, who? – some overworked security engineer who then has to dig up the component’s owner.

Automatically assigning tickets and pull requests to the correct developer teams

There is a better way: In the Kubernetes app, we can set ownership tags on our workloads. A best practice is to label your workloads already when you deploy them. A workflow action then finds the team responsible for this component and their source code repository. For demonstration, we close the loop with two more workflow actions:

  • We open a pull request in the affected repository to create a network policy that blocks the attacker from accessing the cluster again based on the remote IP address Tetragon detected.
  • We create a security ticket, notably, with full context information, and assigned to the correct team.
Security ticket with full context information
Figure 6: Ownership tags on the Kubernetes namespace

Bonus step: Deploying the security policy into the live cluster

The previous workflow follows the GitOps pattern for security remediation. That means we auto-create a pull request with a new security policy and let the developers review it. This creates a well-defined audit trail and avoids unintended disruption of live workloads.

Under some circumstances, you might want to kick out attackers as soon as you can, without any human intervention. We can utilize the newly introduced Kubernetes actions to perform cluster operations directly. We adapt the workflow as follows:

  1. Directly deploy the security policy to the cluster, in addition to creating the pull request.
  2. Possibly delete the pod on which we spotted the attacker to avoid further harm.

Deleting a pod sounds drastic, but if your environment is configured in a resilient manner, for example, if your application can tolerate a brief outage of some of its services, it’s a “cloud native strategy” to kick out an attacker. The Kubernetes control plane automatically re-creates a fresh (and uninfected) pod within seconds.

Kubernetes control plane automatically recreates a fresh and uninfected pod within seconds for a resilient security incident response
Figure 7: Kubernetes control plane automatically recreates a fresh and uninfected pod

Workflows for security incident response on the Dynatrace platform

We demonstrated how to create a security runbook with Dynatrace Automations, orchestrating external security alerts, and custom processes:

  • Tetragon as our source for eBPF-based security events.
  • Kyverno to manage policies in Kubernetes.
  • Grail as our storage for logs and events, which we query using the Dynatrace Query Language (DQL).
  • Ownership tags to identify responsible developer teams.
  • Jira and Slack integrations for workflows.
  • Available soon: GitHub actions to automatically create pull requests with new security policies in the correct repository.
  • Available soon: Kubernetes actions to make urgent changes to the live cluster directly.

Your tech stack might be different, your security runbooks might be different, your risk profile might be different, but with the unique, context-specific insights and automation capabilities of the Dynatrace platform, you can implement and automate it all easily and flexibly.

The post Context-aware security incident response with Dynatrace Automations and Tetragon appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/context-aware-security-incident-response/feed/ 0
The right person at the right time makes all the difference: Best practices for ownership information https://www.dynatrace.com/news/blog/the-right-person-at-the-right-time-makes-all-the-difference-best-practices-for-ownership-information/ https://www.dynatrace.com/news/blog/the-right-person-at-the-right-time-makes-all-the-difference-best-practices-for-ownership-information/#respond Wed, 27 Mar 2024 17:13:39 +0000 https://www.dynatrace.com/news/?p=63337 Import teams into Dynatrace

This blog post covers a best practice approach for linking ownership information with observability data, enabling the automation of incident triaging and reducing MTTR.

The post The right person at the right time makes all the difference: Best practices for ownership information appeared first on Dynatrace news.

]]>
Import teams into Dynatrace

Knowledge is power. Knowing who is responsible for specific areas and services is a vital skill that saves time and hassle. This can quickly become overwhelming and hard to manage in today’s complex microservice environments. Combining services and responsible teams increases transparency, simplifies the identification of the right people, enables automation, and saves time and money.

As always, work must be done before you can reap the rewards. This includes preparing the environment with the required metadata to provide the ownership information transparently and when needed. These efforts include:

  • Extending entities with ownership metadata.
  • Enriching ownership information with desired metadata.
  • Ensuring ownership coverage in the environment.
  • Automating incident triage via targeted remediation or notification tasks.

What to consider when adding ownership information

Introducing ownership information requires several simple steps and typically depends on the desired application. In any case, three building blocks are required to connect the right information to software artifacts.

  • Linking team ownership with the specific service or component they own.
  • Enriching team ownership information with desired metadata.
  • Providing easy access to team ownership information.

The first challenge is the assignment of team ownerships to the appropriate entities. At first glance, this sounds easy: Team A takes care of component B. However, as software changes continuously due to new deployments and releases, so can responsibilities. Revisiting an environment to ensure the assignments are still correct is almost impossible when done manually. The automated extraction of ownership information, for example, from Kubernetes annotations, is therefore essential.

Secondly, knowing who is responsible is essential but not sufficient, especially if you want to automate your triage process. Essential metadata about team ownership helps in addressing the right team with the appriopriate responsibilities via the desired channels. As teams and their structure and metadata are often maintained in a dedicated database, such as Microsoft Entra ID (formerly Azure Active Directory) or ServiceNow. This information needs to be reusable, avoiding the need to maintain the same information twice and ensuring data is not out of sync.

Keeping ownership teams and their properties up to date is essential, as is having the right contact information available when needed.

Finally, the best information is still useless if users can’t retrieve it quickly when needed and use it accordingly. Be it a visual representation or an automated task in a workflow, easily resolving ownership information for services is where the full value becomes noticeable.

How to efficiently introduce team ownerships

Dynatrace provides different ways of associating team ownership with entities and adding desired team metadata, such as contact details, to your environments.

Import teams

It is necessary to get ownership team information into the system and keep it updated. Dynatrace offers several ways to ingest ownership team information. Besides supporting UI and API input for ownership teams, a dedicated workflow action for importing, storing, and updating ownership teams is available.

The import_teams workflow action can be used in either an on-demand or a regularly triggered workflow to get ownership team information and store it accordingly. As data sources can vary, the import_teams action supports different options. Besides the generic import option of accepting JSON objects in the ownership schema (find details on how to use the import teams workflow actions here), two often-used databases for storing and maintaining team information are available to you:

  • Microsoft Entra (formerly Azure Active Directory)
  • ServiceNow (sys_user_groups)
Dynatrace workflow showing two ownership team-import workflows for ServiceNow and Microsoft Entra ID
Figure 1. Dynatrace workflow showing two ownership team-import workflows for ServiceNow and Microsoft Entra ID

The ownership team information can be queried regularly from the databases, maintaining them as the source of truth for the team and its properties. A dedicated workflow allows automatic synchronization of the database information with Dynatrace and keeps the ownership team information up-to-date.

Dynatrace ownership functionality supports configuration-as-code via its proprietary Monaco (Monitoring as code) CLI or Terraform. A workflow template for importing teams can be found in this public GitHub repository.

Assign teams to services

This correlation between people and software is crucial; if a problem occurs or a new security vulnerability is detected, it’s critical to quickly and easily know who is responsible and owns each specific software artifact.

As the ownership of certain components is highly dependent on the organization structure and possibly quite dynamic or heterogeneous, the linking is realized via key-value pairs added to the entities. This ensures flexibility in adding the ownership information when needed.

Since adding key-value pairs to entities is highly flexible, there are several ways of adding ownership metadata to an environment. However, to increase reliability, it’s recommended to add the ownership information depending on the environment characteristics. In cloud-native environments Kubernetes annotations or labels are recommended; these can be further used to propagate certain information within the software topology. Dedicated environment variables or custom properties are additional options to add ownership related metadata. More details on the supported ways of enriching your environment are described in Best practices for ownership information documentation.

Figure 2. Example of a Namespace Definition with an ownership team assigned via annotations
Figure 2. Example of a Namespace Definition with an ownership team assigned via annotations

When adding ownership teams, the key must start with one of five customizable indicators to be recognized as ownership metadata. By default, Dynatrace supports dt.owner or owner as the key prefix, while the value needs to reflect the unique team identifier. This guarantees an unambiguous identification of the right owners per component. The team identifier can be set only when creating the ownership team for the first time, either through the previously explained ownership importer or manually via the web UI or API.

If certain naming conventions or key-value pairs are already used to tag services or apps with responsible teams, it’s possible to reuse the same keys easily. For example, if team:myTeamName is already used to mark selected components, then team can be added as a supported key. Dynatrace will automatically recognize the existing tags as ownership metadata. Otherwise, it’s recommended that you reuse the default key prefix to stay consistent throughout the environment.

Illustration how Dynatrace automatically adds ownership team information of a work load after successful ingest and assignment
Figure 3. Illustration how Dynatrace automatically adds ownership team information of a work load after successful ingest and assignment

To keep an overview of which areas in an environment are already covered and where there might be potential blind spots, ownership information (or the lack thereof) can be easily visualized in notebooks and dashboards. The dashboard shown in figure 4 below can be found in the publicly accessible Github Repository:
https://github.com/dynatrace-perfclinics/platform-engineering-demo/blob/main/dynatraceassets/dashboards/team-ownership-dashboard.json

The dashboard serves as an example and likely needs to be adapted to your specific user needs.

Dynatrace dashboard presenting an overview of ownership-team coverage within the environment
Figure 4. Dynatrace dashboard presenting an overview of ownership-team coverage within the environment

Ownership  

If it comes to the worst case, and a security vulnerability is introduced or a severe incident happens, it’s important to automatically inform the responsible teams and start remediation actions.

With Dynatrace Workflows, triaging time can be minimized and eliminated. Based on the detected issue or event, a workflow can automatically retrieve the associated ownership teams during the execution. The dedicated get_owner workflow action, queries all teams of affected entities. Further the ingested ownership metadata, such as contact details can be used to prepare Jira tickets or notify the right teams with all the crucial information they need to immediately start issue resolution.

Some examples of how ownership information enables automated notification and remediation workflows are listed below.

Sample workflow for an automated release validation
Figure 5. Sample workflow for an automated release validation
  • Problem Remediation Automation utilizing Red Hat Ansible Automation Platform
Sample illustrating a problem remediation with a Red Hat Ansible Automation Platform integration
Figure 6. Sample illustrating problem remediation with a Red Hat Ansible Automation Platform integration
  • Kubernetes Workload Optimization
Sample workflow to automate Kubernetes workload optimization
Figure 7. Sample workflow to automate Kubernetes workload optimization
  • Security Vulnerability Processing
Sample workflow of a security vulnerability processing automation
Figure 8. Sample workflow of a security vulnerability processing automation

Another example of utilizing targeted notifications via Dynatrace Workflows can be found in this public GitHub repository. These examples can be extended to cover similar use cases as above.

What’s next

Dynatrace already provides the possibility to reference ownership team information from entities, ingest crucial ownership team metadata such as contact information in an efficient way, and offers automated access to this information either via the user interface or Dynatrace Workflows. This enables a wide possibility of use cases where automated and targeted notifications play a significant role in reducing MTTR and increasing efficiency and reliability.

To further improve the user handling of the ownership functionality, we are currently working on making this information easily accessible via DQL (Dynatrace Query Language). Making use of all relations and dependencies will further simplify the retrieval of necessary information reliably when needed.

Follow these steps to increase an environment’s transparency and prepare environments to automatically inform and notify the right people at the right time:

  • Ingest ownership metadata, such as contact details, into a Dynatrace environment.
    • Using the teams-importer to keep ownership data automatically in sync.
    • An example of how to set up an ownership import workflow is given in our workflow-samples repository.
  • Extend components and services with ownership information.
    • Use object-specific and reliable association methods such as Kubernetes annotations, environment variables, host metadata, etc. More details can be found in Dynatrace Documentation.
  • Detect blind spots and ownership gaps in environments.
    • Use dashboards and notebooks to keep an overview of the current state and proactively detect blind spots.
    • Set up automation workflows, tailoring and targeting ticketing and remediation activities based on the affected entities and occurring events, problems, or security vulnerabilities. Checkout the already available examples within our publicly available Configuration as Code GitHub repository or our Dynatrace Discover tenant.

Contact us to schedule a demo. We’ll walk you through the various workflows and dashboards discussed in this blog post.

The post The right person at the right time makes all the difference: Best practices for ownership information appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-right-person-at-the-right-time-makes-all-the-difference-best-practices-for-ownership-information/feed/ 0
How platform engineering and IDP observability can accelerate developer velocity https://www.dynatrace.com/news/blog/how-platform-engineering-can-accelerate-developer-velocity/ https://www.dynatrace.com/news/blog/how-platform-engineering-can-accelerate-developer-velocity/#respond Wed, 06 Mar 2024 20:40:49 +0000 https://www.dynatrace.com/news/?p=62869 Data privacy by design, CrowdStrike

At Dynatrace Perform 2024, Dynatrace colleagues Andreas Grabner and Adam Gardner discussed how platform engineering accelerates developer velocity.

The post How platform engineering and IDP observability can accelerate developer velocity appeared first on Dynatrace news.

]]>
Data privacy by design, CrowdStrike

As organizations look to expand DevOps maturity, improve operational efficiency, and increase developer velocity, they are embracing platform engineering as a key driver. Indeed, recent research found that 54% of organizations are investing in platforms to enable easier integration of tools and collaboration between teams involved in automation projects.

Platform engineering creates and manages a shared infrastructure and set of tools, such as internal developer platforms (IDPs), to enable software developers to build, deploy, and operate applications more efficiently. The goal is to abstract away the underlying infrastructure’s complexities while providing a streamlined and standardized environment for development teams. As a result, teams can focus on writing code and building features, rather than dealing with infrastructure nuances.

During a breakout session at Dynatrace Perform 2024, Dynatrace DevSecOps activist Andreas Grabner and staff engineer Adam Gardner demonstrated how to use observability to monitor an IDP for key performance indicators (KPIs). The pair showed how to track factors, including developer velocity, platform adoption, DevOps research and assessment metrics, security, and operational costs.

Recent Dynatrace research has found that only 40% of a typical engineer’s time is spent on productive tasks, and 36% of developers resign because of a bad developer experience, Grabner noted. “If your developers are leaving the company, the IDP may have something to do with it,” he said.

Platform engineering: Build for self-service

Self-service deployment is a key attribute of platform engineering. It gives developers the means to create environments and toolsets unique to their projects.

“[An IDP] must be a product that developers want to use because it helps them get the job done,” Grabner said. “It makes them more productive . . . and reduces the complexity of things such as reading a new app or service. They shouldn’t worry about the platform; they should just start writing code.”

Because of their versatility, teams can use IDPs for all types of software engineering projects, not just those in cloud-native scenarios. IDPs can eliminate much of the administrative minutiae that stalls development projects. Grabner gave the example of one Dynatrace banking customer who built an IDP that enables developers to provision new Microsoft Azure machines or Chef policies without administrative help. “IDPs are not constrained to building microservices or a new serverless app,” Grabner noted.

Before putting an IDP in place, organizations must encourage their platform engineering teams to adopt a product mindset with feedback loops between developers and users. They should also establish milestones to ensure the built product solves a defined business problem.

Reference IDP with Dynatrace

The Dynatrace IDP encompasses platform services, delivery services, and access to observability and automation tools. The Dynatrace Operator automatically ingests all observability data from OpenTelemetry and Prometheus. Furthermore, OneAgent® software observes and gathers all remaining workload logs, metrics, traces, and events.

Automate deployment for faster developer velocity

Additionally, the IDP used during the session connects to the open source Backstage developer portal platform and a library of templates stored in a GitLab repository. The templates can deploy automatically into the development environment with just a few clicks.

Argo works in a GitOps fashion to automate the deployment of files stored in Git. “Argo has an eagle eye on the Git repository,” Gardner said. “Every time something changes, it’s synced to Kubernetes.”

Backstage holds many of an organization’s critical development resources that must be treated with the same respect as business-critical data. Observability is not only about measuring performance and speed but also about capturing granular business analytics to support data-driven decision-making. These metrics can include how many people are using the IDP, how quickly the tasks are running in the IDP, and more. “That means making it available, resilient, and secure,” Grabner said.

Intelligent monitoring is also crucial. “If you don’t monitor, you risk building a product that nobody needs,” Grabner continued.

Observability is a critical component of an IDP. It illuminates the activity of components such as Backstage, GitHub, Argo, and other tools. Service-level objectives (SLOs) are similarly important. SLOs help developers to accelerate their velocity and remain productive with an optimally functioning platform.

Test continuously

Synthetic testing simulates user behaviors within an application or service to pinpoint potential problems. This process is vital to an IDP’s effectiveness. An observability solution can monitor both synthetic and real-user tests to verify an application is on track.

GitLab, a source code repository and collaborative software development platform for DevOps and DevSecOps projects, is populated with a set of pre-filled templates. The combination gives developers a unique set of tools they can deploy on a self-service basis with full monitoring by Dynatrace.

“Every time [developers] pick a template in Backstage, they get their own version of the Git repository based on the template. Then, Argo deploys the app,” Grabner said. “It has worked kind of flawlessly.”

Observability at the core

How we built the IDP

Platform engineering is about being responsible for making sure platforms are available,” Gardner said. “Dynatrace can tell us whether Argo is up and whether it’s killing GitHub with too many syncs. It lets us see events such as starts and traces in a standardized manner.” This certainty can accelerate developer velocity and improve the developer experience, resulting in better software and happier, more productive developers.

Dynatrace has made the reference IDP architecture available on GitHub for anyone to use. It includes a notebook with configuration and deployment instructions.

“It explains every single step that was involved in building the IDP, creating the configuration, and setting up Argo,” Gardner said. “You can launch a code space that starts a container that shows you everything about how an app was built and deployed.”

Curious to learn more about observability to optimize KPI success? Check out the Perform 2024 session: Observability guide to platform engineering.

FAQs about platform engineering

What is the primary role of an internal developer platform (IDP)?

An IDP is a shared infrastructure and set of tools created and managed by platform engineering teams. By providing a self-service, standardized environment, IDPs enable software developers to more efficiently build, deploy, and operate applications. IDPs reduce the need for developers to manage underlying infrastructure nuances, thereby simplifying workflows and accelerating application delivery.

How does platform engineering improve software development?

Platform engineering accelerates developer velocity by providing a streamlined, standardized environment that abstracts away infrastructure complexities. This allows developers to focus on writing code and building features more efficiently.

It improves the developer experience by offering self-service tools and automating tasks, reducing the cognitive load and administrative minutiae that can otherwise hinder productivity and frustrate developers.

Why is observability crucial for successful platform engineering?

Observability is crucial for successful platform engineering because it provides deep insights into the performance, health, and activity of the internal developer platform and its components. It allows platform teams to monitor key performance indicators (KPIs) like platform adoption and operational costs, helping to support the platform’s availability, resiliency, and security. Continuous monitoring helps verify the effectiveness of the IDP and ensures developers maintain optimal productivity.

What are the key principles for building an effective internal developer platform?

Building an effective internal developer platform involves adopting a product mindset, treating developers as internal customers, and incorporating their feedback through continuous loops. Here are some principles to keep in mind as you build:

  • Provide clear self-service capabilities.
  • Automate repetitive tasks.
  • Focus on solving common developer pain points.

The platform should also emphasize scalability, security, and compliance, while fostering a culture of collaboration and knowledge sharing within the engineering organization.

Discover how unified observability unlocks platform engineering success in the free ebook: Driving DevOps and platform engineering for digital transformation.

The post How platform engineering and IDP observability can accelerate developer velocity appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-platform-engineering-can-accelerate-developer-velocity/feed/ 0
Automate CI/CD pipelines with Dynatrace: Part 4, Validation stage https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-4-validation-stage/ https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-4-validation-stage/#respond Wed, 28 Feb 2024 17:58:04 +0000 https://www.dynatrace.com/news/?p=62539 Services Response Rate

In the previous blog post of this series, we discussed the crucial role of Dynatrace as an orchestrator that steps in to stop the testing phase in case of any errors. Additionally, Dynatrace equips SREs and application teams with valuable insights powered by Davis® AI. In this blog post of the series, we will explore […]

The post Automate CI/CD pipelines with Dynatrace: Part 4, Validation stage appeared first on Dynatrace news.

]]>
Services Response Rate

In the previous blog post of this series, we discussed the crucial role of Dynatrace as an orchestrator that steps in to stop the testing phase in case of any errors. Additionally, Dynatrace equips SREs and application teams with valuable insights powered by Davis® AI. In this blog post of the series, we will explore the use of Site Reliability Guardian (SRG) in more detail.

SRG is a potent tool that automates the analysis of release impacts, ensuring validation of service availability, performance, and capacity objectives throughout the application ecosystem by examining the effect of advanced test suites executed earlier in the testing phase.

Dynatrace observability in validation stage

Validation stage overview

The validation stage is a crucial step in the CI/CD (Continuous Integration/Continuous Deployment) process. It involves carefully examining the test results from the previous testing phase. The main goal of this stage is to identify and address any issues or problems that were detected. Doing so reduces the risk of production disruptions and instills confidence in both SREs (Site Reliability Engineers) and end-users. Depending on the outcome of the examination, the build is either approved for deployment to the production environment or rejected.

Challenges of the validation stage

In the Validation phase, SREs face specific challenges that significantly slow down the CI/CD pipeline. Foremost among these is the complexity associated with data gathering and analysis. The burgeoning reliance on cloud technology stacks amplifies this challenge, creating hurdles due to budgetary constraints, time limitations, and the potential risk of human errors. Additionally, another pivotal challenge arises from the time spent on issue identification. Both SREs and application teams invest substantial time and effort in locating and rectifying software glitches within their local environments. These prolonged processes not only strain resources but also introduce delays within the CI/CD pipeline, hampering the timely release of new features to end-users.

Mitigate challenges with Dynatrace

With the support of Dynatrace Grail™, AutomationEngine, and the Site Reliability Guardian, SREs and application teams are assisted in making informed release decisions by utilizing telemetry observability and other insights. Additionally, the Visual Resolution Path within generated problem reports helps in reproducing issues in their environments. The Visual Resolution Path offers a chronological overview of events detected by Dynatrace across all components linked to the underlying issue. It incorporates the automatic discovery of newly generated compute resources and any static resources that are in play. This view seamlessly correlates crucial events across all affected components, eliminating the manual effort of sifting through various monitoring tools for infrastructure, process, or service metrics. As a result, businesses and SREs can redirect their manual diagnostic efforts toward fostering innovation.

Promoting or rejecting the build for production deployment with Dynatrace workflow

  1. Configure an action for the Site Reliability Guardian in the workflow. The action should focus on validating the guardian’s adherence to the application ecosystem’s specific objectives (SLOs). Additionally, align the action’s validation window with the timeframe derived from the recently completed test events.
    Leveraging SRG task to validate the newly build code with Dynatrace Workflow
  2. As the action begins, the Site Reliability Guardian (SRG) evaluates the set objective by analyzing the telemetry data produced during advanced test runs. At the same time, SRG uses DAVIS_EVENTS to identify any potential problems which could result in one of two outcomes.

    Outcome #1: Build promotion

    Once the newly developed code is in line with the objectives outlined in the Guardian—and assuming that Davis AI doesn’t generate any new events—the SRG  action activates the successful path in the workflow. This path includes a JavaScript action called promote_jenkins_build, which triggers an API call to approve the build being considered, leading to the promotion of the build deployment to production.
    SRG assessment - approve the build with Dynatrace Workflow
    Outcome #2: Build rejection
    If Davis AI generates any issue events related to the wider application ecosystem or if any of the objectives configured from the defined guardian are not met, the build rejection workflow is automatically initiated. This triggers the disapprove_jenkins_build  JavaScript action, which leads to the rejection of the build.
    SRG assessment - rejectthe build with Dynatrace Workflow
    Moreover, by utilizing helpful service analysis tools such as Response Time Hotspots and Outliers, SREs can easily identify the root cause of any issues and save considerable time that would otherwise be spent on debugging or taking necessary actions.  SREs can also make use of the Visual Resolution Path to recreate the issues on their setup or identify the events for different components that led to the issue. In both scenarios, a Slack message is sent to the SREs and the impacted app team, capturing the build promotion or rejection.The telemetry data’s automated analytics, powered by SRG and Davis AI, simplify the process of promoting builds. This approach effectively tackles the challenges that come with complex application ecosystems. Additionally, the integration of service tools and Visual Resolution Path helps to identify and fix issues more quickly, resulting in an improved mean time to repair (MTTR).

Validation in the platform engineering context

Dynatrace—essential within the realm of platform engineering—streamlines the validation process, providing critical insights into performance metrics and automating the identification of build failures. By leveraging SRG and Visual Resolution Path, along with Davis AI causal analysis, development teams can quickly pinpoint issues, and further rectify them ensuring a fail-smart approach. The integration of service analysis tools further enhances the validation phase by automating code-level inspections and facilitating timely resolutions. Through these orchestrated efforts, platform engineering promotes a collaborative environment, enabling more efficient validation cycles and fostering continuous enhancement in software quality and delivery.

In conclusion, the integration of Dynatrace observability provides several advantages for SREs and DevOps, enabling them to enhance the key DORA metrics:

  • Deployment Frequency: Improved deployment rate through faster and more informed decision-making. SREs gain visibility into each stage, allowing them to build faster and promptly address issues using the Dynatrace feature set.
  • Change Lead Time: Enhanced efficiency across stages with Dynatrace observability and security tools, leading to quicker postmortems and fewer interruption calls for SREs.
  • Change Failure Rate: Reduction in incidents and rollbacks achieved by utilizing “Configuration Change” events or deployment and annotation events in Dynatrace. This enables SREs to allocate their time more effectively to proactively address actual issues instead of debugging underlying problems.
  • Time to restore service: While these proactive approaches can help improve Deployment Frequency and Change Lead Time, telemetry observability data with Dynatrace AI causation engine Davis AI can aid in improving Time to restore service.

In addition, Dynatrace can leverage the events and telemetry data that it receives during the Continuous Integration/Continuous Deployment (CI/CD) pipeline to construct dashboards. By using JavaScript and DQL, these dashboards can help generate reports on the current DORA metrics. This method can be expanded to gain a better understanding of the SRG executions, enabling us to pinpoint the responsible guardians and the SLOs managed by various teams and identify any instances of failure. Addressing such failures can lead to improvements and further enhance the DORA metrics. Below is a sample dashboard that provides insights into DORA and SRG execution.

DORA metrics and SRE validation insights with Dynatrace workflow

In the next blog post, we’ll discuss the integration of security modules into the DevOps process with the aim of achieving DevSecOps. Additionally, we’ ll explore the incorporation of Chaos Engineering during the testing stage to enhance the overall reliability of the DevSecOps cycle. We’ll ensure that these efforts don’t affect the Time to Restore Service turnaround build time and examine how we can improve the fifth key DORA metric, Reliability.

What’s next?

Curious to see how it all works? Contact us to schedule a demo and we’ll walk you through the various workflows, JavaScript tasks, and the dashboards discussed in this blog series.

Contact us to schedule a demo and we’ll walk you through the various workflows, JavaScript tasks, and the dashboards discussed in this blog series.

If you’re an existing Dynatrace Managed customer looking to upgrade to Dynatrace SaaS, see How to start your journey to Dynatrace SaaS.

The post Automate CI/CD pipelines with Dynatrace: Part 4, Validation stage appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-4-validation-stage/feed/ 0
Automate CI/CD pipelines with Dynatrace: Part 3, Testing stage https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-3-testing-stage/ https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-3-testing-stage/#respond Mon, 18 Dec 2023 17:24:42 +0000 https://www.dynatrace.com/news/?p=61124 Services graphic

In the last blog post of this series, we delved into how Dynatrace, functioning as a deploy-stage orchestrator, solves the challenges confronted by Site Reliability Engineers (SREs) during the early of automating CI/CD processes. Having laid the foundation during the deployment stage, we’ll now explore the benefits of Dynatrace visibility and orchestration during the testing […]

The post Automate CI/CD pipelines with Dynatrace: Part 3, Testing stage appeared first on Dynatrace news.

]]>
Services graphic

In the last blog post of this series, we delved into how Dynatrace, functioning as a deploy-stage orchestrator, solves the challenges confronted by Site Reliability Engineers (SREs) during the early of automating CI/CD processes. Having laid the foundation during the deployment stage, we’ll now explore the benefits of Dynatrace visibility and orchestration during the testing phase.

Dynatrace Observability in Testing Stage

The testing stage plays a crucial role in ensuring the quality of newly built code through the execution of automated test cases. Testing includes integration tests, which assess whether the code functions as intended when interacting with other services and application functionalities. It can also include performance testing to determine if the application can effectively handle the demands of the production environment.

Mitigate challenges during the testing stage

Like the build stage, the testing stage is time-sensitive and can consume a significant amount of time for execution, in turn creating a substantial waiting period for SREs to determine the success or failure of tests and the need for retesting. This slow feedback and time spent rerunning tests can hinder the overall software deployment process.

To mitigate these challenges, building upon the groundwork established during the deploy phase, Dynatrace can effectively pinpoint any issues encountered during testing. Moreover, by configuring alert notifications through native features such as ownership and alerting profiles, teams can receive prompt alerts in the event of failures. This proactive strategy significantly minimizes wait times and empowers SREs to redirect their focus toward innovative endeavors.

Test Workflow in Dynatrace screenshot

The steps outlined below show you how to achieve such a proactive solution.

  1. Pass annotation events to Dynatrace, leveraging the Ingest events API, at the beginning and end of each test. These events provide additional context to the Davis® AI causation engine in case of issues and function as logic operators for the execution of advanced testing, such as soak, integration, or chaos engineering. Load test information in Dynatrace
  2. Integrate performance test tools with Dynatrace by adding headers to HTTP requests. Further, harness request attributes and telemetry data gathered from requests to gain insights into the requests.
  3. (Optional) Set web request naming rules using the earlier configured request attributes to facilitate easy identification of requests across different releases or test suites.
    Integrate load test tools with Dynatrace
  4. During the execution of tests, the telemetry data generated by the newly built code is transmitted to Grail through OneAgent.
    Web-request naming rules to disThe data is examined and can fall into one of the below two categories:

    If there are errors in the telemetry data

    If there are errors in the telemetry data, the task immediately marks the stage as problematic and triggers the process to stop the ongoing job. The stop_the_build task utilizes the CI/CD API (in this case, Jenkins) to halt the build. This task also retrieves additional job metadata to include in the generated message.Monitoring the spans and orchestrating the testing page

    In tandem with stopping the job, another task in the workflow, send_load_test_failure, is triggered: a Slack message is dispatched to the team responsible for initiating the job. The immediate alerting mechanism ensures that SREs are promptly informed about test errors, thereby circumventing the high wait times.

    If there are no errors and the outcome is successful

    If there are no errors reported in the data, the pipeline job will proceed to conduct integration and reliability tests, which will be validated in the final validation stage by the Site Reliability Guardian (SRG). During this phase, SRG will validate the test coverage to assess the impact of the newly introduced code and determine its suitability for promotion to the production environment.

Optional best practices

Optionally, the following best practices are recommended to achieve the most effective outcomes.

  • During the testing stage, generate on-demand synthetic monitors to monitor the performance of the application as experienced by your end users in different geolocations.
  • Generate SLOs, leveraging the Dynatrace API, that will be validated in the next stage to determine if the build should be rejected or promoted to production.

Test stage orchestration in a platform engineering context

Dynatrace, integral to platform engineering, offers invaluable visibility into test results and accelerates the test cycle by swiftly identifying and resolving underlying issues. Utilizing service tools like response hotspots and distributed traces, along with Davis causal AI, dev teams can achieve fast turnarounds in the event of test failures. Dynatrace root cause analysis enhances collaboration, providing a comprehensive solution for optimizing testing phases so that they allow for a more agile and productive development cycle, fostering a culture of continuous improvement and accelerating overall software delivery.

Also, to optimize the workflow for both SREs and application teams, leverage Dynatrace dashboard capabilities to expand your data visibility. This allows for easy comparisons of telemetry data across different releases and enables the utilization of service tools for in-depth analysis, ultimately contributing to service optimization.

Conclusion

In conclusion, this blog post examined how Dynatrace capabilities and telemetry data can enhance the success of the testing stage. In the next blog post in this series, we delve deeper into how Dynatrace can further streamline the promotion of new code to production.

The post Automate CI/CD pipelines with Dynatrace: Part 3, Testing stage appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-3-testing-stage/feed/ 0
Automate CI/CD pipelines with Dynatrace: Part 2, Deploy stage https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-2-deploy-stage/ https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-2-deploy-stage/#respond Tue, 28 Nov 2023 13:55:19 +0000 https://www.dynatrace.com/news/?p=60859 Automate CI-CD pipelines with Dynatrace Part 2, Deploy stage

In the previous installment of this blog series, we explored how to set up Dynatrace as a build-stage orchestrator to effectively address the challenges faced by Site Reliability Engineers (SREs). In this blog post, we’ll transition to the next pipeline stage, the Deploy stage, and examine the visibility advantages that Dynatrace provides during this critical […]

The post Automate CI/CD pipelines with Dynatrace: Part 2, Deploy stage appeared first on Dynatrace news.

]]>
Automate CI-CD pipelines with Dynatrace Part 2, Deploy stage


In the previous installment of this blog series, we explored how to set up Dynatrace as a build-stage orchestrator to effectively address the challenges faced by Site Reliability Engineers (SREs). In this blog post, we’ll transition to the next pipeline stage, the Deploy stage, and examine the visibility advantages that Dynatrace provides during this critical phase.

Deploy stage

In the deployment stage, the application code is typically deployed in an environment that mirrors the production environment. This step is crucial as this environment is used for the final validation and testing phase before the code is released into production. This stage ensures the code meets the required quality standards before it goes live.

Deployment challenges

Configuration drift

Configuration drift is a challenge in the deployment stage. Configuration drift occurs when a staging environment’s configuration deviates from the production environment. Such deviations often result from undocumented or un-versioned environmental changes, potentially causing unexpected behavior and production outages upon deployment.

Even when the staging environment closely mirrors the production environment, achieving a complete replication of all potential scenarios, such as simulating extremely high traffic volumes to assess software performance, remains challenging. This can lead to a lack of insight into how the code will behave when exposed to heavy traffic.

Creating a staging environment that faithfully replicates the production setup is crucial to effectively addressing such challenges. Furthermore, augmenting test coverage to mirror the scenarios encountered in production is imperative. These strategies can play a vital role in the early detection of issues, helping you identify potential performance bottlenecks and application issues during deployment for staging. With this approach, you gain increased confidence that a release will succeed when promoted to the production environment.

Leverage OneAgent functionality

Ingesting configuration changes into Dynatrace through the events API call and utilizing OneAgent® to detect configuration changes for supported technologies help maintain close alignment between your staging and production environments. This approach effectively combats configuration drift. Furthermore, the Dynatrace Davis® AI engine can accurately identify whether such changes have resulted in any issues and allows you to roll back configurations if necessary.

By harnessing OneAgent native features, the generated telemetry data can be effectively combined with Davis AI prediction and forecasting capabilities. This enables the prediction of trends that might not have been thoroughly tested due to variations in different environments. This proactive approach proves instrumental in anticipating and mitigating potential performance issues before they result in adverse impacts.

Combat configuration drift

Dynatrace actively tracks modifications to the standard configuration files of supported technologies. Whenever a change is detected, Dynatrace automatically generates a Deployment change event for the corresponding process and the host on which the process runs. The predefined set of files monitored for configuration alterations is maintained within ruxitagentproc.conf. In the event of any incidents stemming from these configuration adjustments, Davis AI utilizes the Deployment change data to pinpoint potential root causes.

As an illustration, in the screenshot below, when a modification was made to the nginx.conf file, Dynatrace promptly detected the configuration change and reported it as a Deployment change event. In this specific instance, the misconfiguration of nginx led to a critical issue where nginx crashed and could not restart, prompting an alert from Dynatrace. Davis AI efficiently identified the deployment change as the potential root cause for the malfunctioning of nginx.

Dynatrace worklflow

Dynatrace worklflow

By harnessing these deployment events, SREs can effectively track all configuration changes that have occurred over a specified time. This capability empowers them to replicate the staging environment to match the production setup, proactively combating configuration drift.

Although Dynatrace can detect configuration changes, the best practice is to systematically ingest a configuration change or deployment change event each time the team modifies a configuration file. This approach ensures proper record-keeping and assists in achieving consistency between the staging and production environments, effectively mitigating configuration drift.

Predictive traffic analysis

Deploying OneAgent within the staging environment facilitates the availability of telemetry data for analysis by Davis AI. Davis AI can leverage this data to enable predictive analysis. This enables a scenario where—even if test cases can’t replicate identical traffic to production due to various limitations—predictive analysis can offer insights into the capacity of the new build to withstand production-level loads.

To illustrate this concept, consider the scenario below. During testing with limited generated traffic, Davis AI predictive analysis offered valuable insights into how Elastic Book Storage (EBS) might perform and how the rate of Input/Output Operations Per Second (IOPs) might change under heavier traffic load.

Orchestrating Deploy Stage

Other proven strategies

We recommend following best practices to achieve the most effective observability outcomes.

  • At the onset of the deployment stage, send a deployment event to Dynatrace notifying it about the new deployment. These deployment events are important as they provide additional context to the Davis AI causation engine, aiding in evaluating deployment quality.Orchestrating Deploy Stage
  • To further aid with postmortem analysis, use the DT_TAGS metatag to clearly distinguish between different build and service requests.
  • Additionally, consider using the DT_CUSTOM_PROP environment variable to include extra metadata about the build, providing valuable information for monitoring and analysis.

Orchestrating Deploy Stage

The framework outlined above provides a comprehensive view of the deployment process and facilitates comparisons across different releases. For instance, in the animation below, we utilized deployment events and DT_TAGS to assess the quality of releases and capture golden signals across various builds. This approach can be a valuable resource for developers in understanding underlying issues and identifying opportunities for enhancements.

Video thumbnail

Deploy stage orchestration in a platform engineering context

Deployment events automatically identified by OneAgent or manually triggered by app owners play a pivotal role in ensuring consistency across all environments. When issues are detected, causal AI determines if the problem is linked to deployment changes. This contributes significantly to increased productivity for development teams, as they spend less time troubleshooting and identifying discrepancies across environments. Additionally, the integration of predictive analysis, fueled by Davis AI, instills confidence in developers that their application’s behavior will withstand the anticipated load, including potentially high traffic.

Dynatrace as a deploy-stage orchestrator

You’ll find all the Jenkins pipeline and other scripts you need to set up Dynatrace as a deploy-stage orchestrator, including the referenced dashboard, in thisGit repository.

What’s next?

This blog post delved into how Dynatrace capabilities and telemetry data from a staging environment can effectively align staging with the production setup. In the next blog post in this series, we’ll further explore how Dynatrace can aid in streamlining the testing phase within an SRE pipeline.

The post Automate CI/CD pipelines with Dynatrace: Part 2, Deploy stage appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-2-deploy-stage/feed/ 0
Automate CI/CD pipelines with Dynatrace: Part 1, Build stage https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-1-build-stage/ https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-1-build-stage/#respond Fri, 17 Nov 2023 17:54:05 +0000 https://www.dynatrace.com/news/?p=60770 Automate CI-CD pipelines with Dynatrace, Build stage_ high res version

In the first blog post of this series, we explored how the Dynatrace® observability and security platform boosts the reliability of Site Reliability Engineers (SRE) CI/CD pipelines and enhances their ability to focus on innovation. This blog post guides you through configuring Dynatrace to automate CI/CD processes to achieve these objectives. Dynatrace observability architecture can […]

The post Automate CI/CD pipelines with Dynatrace: Part 1, Build stage appeared first on Dynatrace news.

]]>
Automate CI-CD pipelines with Dynatrace, Build stage_ high res version


In the first blog post of this series, we explored how the Dynatrace® observability and security platform boosts the reliability of Site Reliability Engineers (SRE) CI/CD pipelines and enhances their ability to focus on innovation. This blog post guides you through configuring Dynatrace to automate CI/CD processes to achieve these objectives.

Site Reliability Engineering Architecture

Dynatrace observability architecture can be classified into three layers:

  1. Orchestration (Dynatrace)
  2. CI/CD toolset (Jenkins / Chef / Puppet / Bamboo, etc.)
  3. Infrastructure layer (Kubernetes Cluster, GCP, Standalone server, etc.)

A conventional pipeline has the following stages:

  1. Build
  2. Deploy
  3. Test
  4. Validation

Now, let’s explore the challenges that arise in the different stages of the pipeline and examine how the role of Dynatrace as the orchestrator can contribute to overcoming these challenges. In this blog post, we’ll focus on the first stage of the pipeline, the Build stage.

Build stage

The build stage involves pulling the latest code, compiling it, and ensuring its proper functionality by running feature tests.

Dynatrace Observability in Build Stage

Mitigate challenges

While the application is building, you might see compilation problems that necessitate debugging, and there might be considerable delay before the root causes of such issues are identified. The build process must be restarted once a solution is identified and deployed. A high rate of build failures indicates that the team is investing significant time into diagnosing and rectifying issues. This waiting and problem-solving can hinder your overall software development process.

Enhancing visibility during the build stage can help identify compilation and integration errors earlier in the development cycle. This proactive approach reduces wait times and allows SREs to redirect their efforts toward innovation.

Leverage Dynatrace features

The instrumentation flow currently operates as follows:

Dynatrace Observability in Build Stage

  1. After code commits, a GitHub action is employed to trigger the workflow, utilizing the workflow API. Within this workflow, a designated task, powered by JavaScript, initiates the build stage of the pipeline job, commencing a new build. The below screenshot shows the code initiating the Jenkins job.
    Dynatrace Workflow
  2. During the job pipeline execution, logs are extracted and pushed to Grail. Log reading can be facilitated by installing OneAgent or employing a log ingestion system to ingest the logs into Dynatrace. Subsequently, a workflow DQL task named look_for_build_errors will scrutinize the build stage for errors.
  3. Based on the outcome of the log analysis, one of two paths is pursued:
    1. If build errors are detected
      Should the job logs reveal errors, the task promptly flags the build as erroneous and initiates the process to halt the ongoing job. The task stop_the_jenkins_job leverages the CI/CD API (in this case, Jenkins) to stop the build, as indicated in the output from step 2. This task also extracts additional job metadata to incorporate in the generated alert.
      Dynatrace Workflow
      In tandem with stopping the job, another task in the workflow, send_alert_to_team, activates, sending a Slack message to the team responsible for initiating the job. The immediate alerting mechanism ensures that SREs are promptly informed about build errors, thereby circumventing high wait times.
      Dynatrace Workflow
    2. If no errors are detected
      If no errors are identified, the workflow seamlessly advances the job to the Deploy stage. As a good practice, we recommend that you incorporate a workflow with a JavaScript task that periodically retrieves job details and ingests them as business events to enhance visibility into the pipeline stages and provide insights into current trends.Dynatrace Workflow
      Above is a sample JavaScript workflow that extracts data from the Jenkins pipeline via API calls and the dashboard capturing current trends.Video thumbnailThe Jenkins Build Stage Insights dashboard provides visualizations of the build stages represented as business events.

Build stage orchestration in a platform engineering context

In the realm of platform engineering, the focus is on relieving the cognitive burden on development teams and boosting their efficiency by offering predefined pathways, essentially treating the development platform as a product. 

With build-log monitoring and the initiation of diverse workflows within Dynatrace based on log outcomes, automation is now a reality for many development teams. Furthermore, the insights gained from the build stage provide SREs and app teams with valuable information for optimizing existing pipelines and areas for improvement. 

What’s next

Now, it should be clear how Dynatrace orchestration can be introduced early in the pipeline process and how it can significantly decrease wait times during the build stage by notifying the teams of any errors during the application build process.

In the next blog post in this series, we’ll delve into how this orchestration further aids SRE teams in addressing the challenges encountered during the Deploy stage.

All the resources you need to set up Dynatrace as a build-stage orchestrator

In the meantime, you’ll find all the resources you need to set up Dynatrace as a build-stage orchestrator (workflows, JavaScript tasks, and the referenced dashboard) in this Git repository.

The post Automate CI/CD pipelines with Dynatrace: Part 1, Build stage appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/automate-ci-cd-pipelines-with-dynatrace-part-1-build-stage/feed/ 0
Ensure safe and secure releases at scale by providing Golden Paths https://www.dynatrace.com/news/blog/ensure-safe-and-secure-releases-at-scale-by-providing-golden-paths/ https://www.dynatrace.com/news/blog/ensure-safe-and-secure-releases-at-scale-by-providing-golden-paths/#respond Tue, 14 Nov 2023 16:17:41 +0000 https://www.dynatrace.com/news/?p=60743 Unlocking the Power of Quality Gating by enabling Self-Service

Golden Paths for rapid product development Modern software development aims to streamline development and delivery processes to ensure fast releases to the market without violating quality and security standards. DevOps practices have been established in the last decade to accomplish this goal and deal with the dynamics of modern, cloud-native software architectures. To bring these […]

The post Ensure safe and secure releases at scale by providing Golden Paths appeared first on Dynatrace news.

]]>
Unlocking the Power of Quality Gating by enabling Self-Service

Golden Paths for rapid product development

Modern software development aims to streamline development and delivery processes to ensure fast releases to the market without violating quality and security standards. DevOps practices have been established in the last decade to accomplish this goal and deal with the dynamics of modern, cloud-native software architectures. To bring these practices to life within an organization at scale, the discipline of platform engineering has gained popularity. From a high-level point of view, platform engineering aims to:

  • Reduce the cognitive load on development teams.
  • Improve reliability and resiliency of products that rely on platform capabilities.
  • Accelerate product development and delivery by reusing and sharing platform tools and knowledge.
  • Reduce risk of security, regulatory, and functional issues in products and services.
  • Enable cost-effective and productive use of services.

While it takes multiple capabilities to achieve these goals, implementing “Golden Path” templates is a fundamental ingredient. The Cloud Native Computing Foundation defines a Golden Path as a “templated composition of well-integrated code and capabilities for rapid project development.” Simply put, a Golden Path is a self-service template for common tasks that allows for autonomy while providing guardrails that safeguard production environments.

Imagine that instead of development teams fending for themselves amidst a sea of tools and infrastructure, well-defined and enterprise-wide templates are provided for the development of all new product services. Such a template should contain a get-started tutorial, sample source-code framework, policy guardrails, CI/CD pipeline, infrastructure-as-code templates, and reference documentation. This approach helps you quickly integrate best practices within your organization and provides cloneable artifacts for rapid product development.

No developer is left in the dark to fend for themselves, Golden Paths light the way.

Ensure governance across your organization

While Golden Paths are key to bringing DevOps best practices to development teams, distributing them is challenging. This challenge is partially addressed by internal development platforms (IDP), which have been adopted by many, but not all, organizations.

Consequently, Dynatrace provides templates in the context they are needed—with or without the internal use of an IDP. This ensures governance across your organization with the proper templates in the right place. At the same time, this does not restrict you from managing them in your IDP, as explained below.

Leverage opinionated, self-service, and optional templates

The latest enhancements of the Site Reliability Guardian have introduced opinionated Golden Path templates. For now, they concentrate on Kubernetes, host, and security objectives but they will grow to cover even more reliability, performance, and security aspects. These templates enable development teams to instantiate a guardian for automating release validation in a self-service manner. Along the journey, monitored entities can be selected to provide the context to fetch the right data from Dynatrace Grail™.

Two-step approach to creating a guardian using a template.
A two-step approach to creating a guardian using a template.

After completing this two-step process, a ready-to-use guardian is created. It’s possible to add new objectives if needed or to tailor the objectives and their thresholds.

Fully functional and pre-configured guardian for a Kubernetes workload.
Fully functional and pre-configured guardian for a Kubernetes workload.

Allow for flexibility

Custom query variables are available to fine-tune guardian objectives and maintain flexibility in fetching data from Grail. An example of such a variable is the version number of a service. In many cases, you want to retrieve the logs, metrics, or traces of a particular version, which is unknown upfront. Consequently, the custom variable version can be defined in the DQL query, which retrieves its value before executing.

Custom query variables can be defined by starting with the $ sign followed by the variable name in the DQL query editor of a guardian objective. An overlay component then helps you add this variable and define the default value in the context of the guardian.

Leverage variables to parameterize the query of an objective.
Leverage variables to parameterize the query of an objective.

While a variable will fall back to its default if no value is provided, there are three ways of ingesting the value:

  1. When triggering a guardian validation from within the UI by selecting the Validate button, set a value for the configured variables using the Set variables option.
  2. The Site Reliability Guardian workflow action allows you to set a variable value. In this case, it is possible to consume data from the triggering event or a previous workflow action using Jinja expressions.
  3. If a workflow is used to automate the validation process, the workflow trigger event can provide automatically mapped properties to variables. Therefore, the properties must be within the execution_context property, as shown below. Given this example, the version number 0.1.0 is set where the variable $version is defined.

Code snippet to set variables

Since variables provide important context for validation, you can view validation-result values over time. The following screenshots depict the version number associated with the last 10 validation results.

Site Reliability Guardian Carts validation history in Dynatrace screenshot

Integrate with existing internal developer platforms

If you have an internal developer platform that serves as a central product engineering hub, you need Golden Path templates incorporated there so that you can scale out the platform for use by different teams. With Dynatrace, this can be achieved by extracting an instantiated template using Dynatrace Configuration as Code or cloning it from our public repository. This declarative format of a template, consisting of a guardian and workflow configuration, as depicted by the sample below, can then be added to the template repository managed and distributed by an IDP.

Dynatrace Configuration as Code sample

What’s next?

Site Reliability Guardian is built to ensure and maintain reliability and resiliency of products by leveraging automated release validations. Therefore, Golden Path templates in a self-service manner are offered for platform engineers to enable development teams to get started easily. The next enhancements of the Site Reliability Guardian will bring Davis® AI even closer to the Site Reliability Guardian by recommending relevant objectives and baselines for comparison.

  • The new functionality has already been released and is available for your use. If you’re using the Site Reliability Guardian already, update to the latest version. Otherwise, navigate to Dynatrace Hub and install it from there.
  • If the Site Reliability Guardian is new to you and you’re curious about it, go to Dynatrace Discovery and explore its capabilities in a playground environment.

We’d love to hear your feedback about Site Reliability GuardianLet us know your thoughts or share your ideas about further improvements.

The post Ensure safe and secure releases at scale by providing Golden Paths appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/ensure-safe-and-secure-releases-at-scale-by-providing-golden-paths/feed/ 0
Dynatrace EdgeConnect securely connects your local systems to Dynatrace SaaS https://www.dynatrace.com/news/blog/dynatrace-edgeconnect-securely-connects-your-local-systems-to-dynatrace-saas/ https://www.dynatrace.com/news/blog/dynatrace-edgeconnect-securely-connects-your-local-systems-to-dynatrace-saas/#respond Tue, 10 Oct 2023 18:19:55 +0000 https://www.dynatrace.com/news/?p=59974 EdgeConnect

EdgeConnect enables secure connections between the Dynatrace® analytics platform and third-party systems within protected networks, whether they're running on-premises or in the cloud. EdgeConnect is designed to address automation use cases where accessibility to internet-restricted systems is crucial but impossible due to security policies.

The post Dynatrace EdgeConnect securely connects your local systems to Dynatrace SaaS appeared first on Dynatrace news.

]]>
EdgeConnect

Bridge the gaps while staying secure

While more and more applications are moving to the cloud, the need for secure environments isn’t going away. EdgeConnect provides a secure bridge for SaaS-heavy companies like Dynatrace, which hosts numerous systems and data behind VPNs. EdgeConnect facilitates seamless interaction, ensuring data security and operational efficiency.

With the increasing adoption of SaaS platforms and escalating security concerns caused by the proliferation of cyber threats, companies are becoming increasingly aware of the importance of safeguarding their systems and data. In this hybrid world, IT and business processes often span across a blend of on-premises and SaaS systems, making standardization and automation necessary for efficiency. Enterprises seek solutions that enable these processes to interact across their entire system landscape without compromising security.

Due to a variety of challenges, organizations face significant difficulties in achieving seamless interaction between on-premises and SaaS systems. Security concerns make it risky to bridge these systems, leaving enough flexibility for users to achieve their goals and keeping IT in control of infrastructure and deployments.

A look behind the curtain: EdgeConnect connects to the outside world

An EdgeConnect instance establishes a WebSocket secure connection (WSS/443) to the Dynatrace platform that doesn’t require opening ports or inbound connections. EdgeConnect acts as a bridge between Dynatrace and the network where it’s deployed

Figure 1: Visualization of an EdgeConnect connection to the Dynatrace platform.
Figure 1: Visualization of an EdgeConnect connection to the Dynatrace platform.

Setting up an EdgeConnect is simple. You can manage your EdgeConnect configuration in the Settings app under General > External requests. Once the initial setup is complete, Dynatrace routes all HTTP(s) service requests that match your configured rules via EdgeConnect.

Figure 2: Editing the host pattern of an existing EdgeConnect configuration.
Figure 2: Editing the host pattern of an existing EdgeConnect configuration.

Efficiency and control

EdgeConnect boasts a range of features designed for efficiency and control. It’s built for container-based deployments, operating either as a standalone container or within your Kubernetes or Openshift clusters. EdgeConnect is designed to forward HTTP(s) requests exclusively, ensuring secure data transmission. It supports multi-instance high availability and load balancing, providing robust performance and reliability.

EdgeConnect empowers operators with an optional allow list for domains during deployment. This feature enables IT to maintain control without the need for any configuration in Dynatrace, further simplifying the process.

The versatility of EdgeConnect enables many use cases:

  • Automated creation and update of tickets in your local Jira instance using Dynatrace Workflows whenever Davis® AI identifies problems.
  • Integration of data from your internal data warehouse to Dynatrace dashboards, blended with data from Grail.
  • Flexible and adaptable connections to internet-based services like ServiceNow via a dedicated host and IP that you control.

What’s next

EdgeConnect is available for all Dynatrace SaaS environments, version 1.275+. Have a look at our documentation to learn more about how to configure and deploy EdgeConnect. Once deployed, you can manage your EdgeConnect configuration in the Settings app under General > External requests.

We’d love to hear your feedback about EdgeConnect. Let us know your thoughts or share your ideas about how we can improve EdgeConnect in the future.

The post Dynatrace EdgeConnect securely connects your local systems to Dynatrace SaaS appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-edgeconnect-securely-connects-your-local-systems-to-dynatrace-saas/feed/ 0
Accelerate and empower Site Reliability Engineering with Dynatrace observability https://www.dynatrace.com/news/blog/accelerate-and-empower-site-reliability-engineering-with-dynatrace-observability/ https://www.dynatrace.com/news/blog/accelerate-and-empower-site-reliability-engineering-with-dynatrace-observability/#respond Tue, 10 Oct 2023 17:53:49 +0000 https://www.dynatrace.com/news/?p=59950 Observability graphic

In this blog series, we explore how organizations can connect Dynatrace to their release processes and leverage Davis® AI to attain full automation, giving Site Reliability Engineering (SRE) teams more time to focus on innovation and other important business goals.

The post Accelerate and empower Site Reliability Engineering with Dynatrace observability appeared first on Dynatrace news.

]]>
Observability graphic

In the realm of SRE, time and effort allocation planning are crucial factors that involve a delicate balance between operational management and project improvements. This intricate allocation strategy can be categorized into two main domains. In this blog post, we’ll delve deeper into these categories to gain a comprehensive understanding of their significance and the challenges they present.

Planned effort

Site Reliability Engineering (SRE) effort and time allocation planning typically fall into two domains:

  • Operations Management (50%)
    Operations Management includes on-call responsibilities, post-mortem assessments, addressing other interruptions, and buffer time. These tasks collectively ensure uninterrupted production service.
  • Process Improvements (50%)
    The allocation for process improvements is devoted to automation and continuous improvement SREs help to ensure that systems are scalable, reliable, and efficient. This improves the current project and paves the way for future innovation.

Reality

In practice, while both these categories have equal attention, project improvements hold paramount importance for business outcomes. SREs invest significant effort in enhancing software reliability, scalability, and dependability. Regrettably, recent reports indicate that SREs spend a substantial portion of their time addressing build issues and managing production incidents. This challenge escalates with the growing complexity of cloud systems and organizational aspirations for digital transformation, often leaving minimal time for substantial project improvements. Consequently, organizations grapple with various issues that impact software reliability.

Process Improvements

Organizations are strategically integrating observability into the initial stages of their release processes to tackle these challenges. As they embark on this initiative, SREs are tasked with identifying an observability solution that aligns seamlessly with their application teams and seamlessly integrates with their existing toolset.

Outcome

Dynatrace plays a pivotal role in this endeavor by empowering application teams through its seamless integration. The Dynatrace integration leverages native features and events that pass through the pipeline. Events serve as logic operators that can trigger or stop subsequent tasks within the pipeline. Additionally, the Site Reliability Guardian serves as the governing entity, making decisions on whether to proceed or halt a specific build based on observable telemetry data supplied by OneAgent during the CI/CD pipeline process.  This proactive strategy significantly enhances the chances of success for SREs, providing them with more time to focus on substantial project improvements (50%) and broaden the buffer zone (30%). This empowers them to spearhead innovations that ensure the business is prepared for future expansions.

Site Reliability Engineering with Dynatrace

Integrate DevSecOps with Dynatrace

Software delivery is structured around CI/CD pipelines, which play a critical role in the SRE process and represent the first step toward effective automation. Automated CI/CD pipelines greatly reduce the occurrence of manual errors. They provide continuous feedback to developers and enable rapid product iterations. Elevating the efficiency of release pipelines is one key to producing high-quality software; it mitigates the need for post-incident analyses and on-call duties for the SRE team. It also returns valuable time back to the SRE team.

A few avenues for elevating CI/CD pipelines are:

  • Enhancing the extent of automated test coverage during the testing phase.
  • Embracing the tenets of DevOps and DevSecOps methodologies anchored in engineering principles.
  • Identifying and automating the validation of business/application Service Level Objectives (SLOs) during release cycles.
  • Streamlining the CI/CD process to ensure optimal efficiency.

To realize these goals, SRE teams can seamlessly follow a three-step process within their Dynatrace environment:

  1. Commencement (optional): If desired, initiate a Dynatrace workflow task to establish a connection with the DevOps tool using HTTP. This step is only necessary if you intend to control the build creation process through Dynatrace.
  2. Push events: Configure your DevOps tool to dispatch deployment events at the inception of the deployment process. Additionally, introduce annotation events to notify Dynatrace of the progress within your testing phase. These events serve as logical operators that dictate the course of the release process.
  3. Automated validation and progression: Depending on the configuration of your tasks, Site Reliability Guardian (SRG) validation can be automatically activated to promote or disapprove the advancement of a build towards production.

Accelerate Delivery Pipelines - empowered by Dynatrace

This integration leads to complete automation with end-to-end pipeline visibility, thereby reducing the heavy lifting of release management for SRE teams and empowering them to focus on innovation, which catalyzes organizational growth.

Designing systems for reliability

Engineering teams must extend their focus beyond functional and load tests to instill assurance in the software release process via automated cycles. The rationale is that, during actual production, variables can induce outcomes that are different from those predicted in standardized tests. Thus, more comprehensive testing becomes essential to embrace unpredictability and mirror real-world conditions. These practices are commonly known as “chaos engineering.

By embracing chaos engineering practices, development teams cultivate a higher degree of confidence in the robustness of their applications within specific production scenarios. However, the very nature of chaos engineering introduces a challenge: identifying the responsible service when a failure occurs. This is where Dynatrace Davis AI comes into play, leveraging telemetry data from your services. Davis AI automatically establishes a baseline of each service’s behavior, learned during the development and functional testing phases, and subsequently identifies the underlying causes of failures, thereby eliminating the uncertainty of the root cause and “war room” scenarios.

To illustrate, a CI/CD pipeline setup has a functional test phase followed by a chaos-engineering test. The chaos engineering test was structured to randomly slow down a Docker container and, thereby, a critical service for the application. During this phase, Davis AI automatically identifies that a key business request has breached its automated baseline due to a significant Mongo database slowdown. Moreover, Davis AI identifies the reason behind the slowdown and pinpoints the exact query that caused the problem.

Chaos engineering dashboard in Dynatrace

With Davis AI’s contextual capabilities, embracing chaos engineering and making application code robust and better prepared for production deployments is easy.

Reduce operations overhead

Ideally, no bugs are reported once software is deployed to production. However, this is highly unlikely. Therefore, it’s important to have a process in place to minimize downtime in the event of a failure.

Davis AI assists with automated root cause analysis, providing details of the underlying services, traces, logs, and user sessions that caused the failure. This can save SRE teams from the time and effort of debugging or a war room scenario. By identifying the exact code or trace that causes a failure, Davis AI helps teams fix problems quickly and significantly reduces MTTR. In addition, if you configure an automated remediation workflow, Dynatrace invokes it and restores applications with virtually no downtime.

Automate Operations dashboard in Dynatrace

In the above screenshot, one of the requests in the application reports errors. Davis AI identifies that this anomaly was reported in the newer releases. Davis AI captures and relays the details of previous and current builds to the remediation workflow. The remediation workflow uses this information to roll back the build, leaving Davis AI to automatically validate and close detected problems.

Dynatrace Observability empowers the SRE team, propelling them towards LowOps/NoOps. This allows them to prioritize innovation, strategic goals, and minimize downtime while optimizing the delivery of perfect software. In our next blog post in this series, we’ll delve into the step-by-step process of automating the release pipeline using Dynatrace best practices.

The post Accelerate and empower Site Reliability Engineering with Dynatrace observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/accelerate-and-empower-site-reliability-engineering-with-dynatrace-observability/feed/ 0