cloud native | Dynatrace news The tech industry is moving fast and our customers are as well. Stay up-to-date with the latest trends, best practices, thought leadership, and our solution's biweekly feature releases. Mon, 09 Mar 2026 14:15:41 +0000 en hourly 1 The runtime reckoning: How the agentic evolution is reshaping security https://www.dynatrace.com/news/blog/the-runtime-reckoning-how-the-agentic-evolution-is-reshaping-security/ https://www.dynatrace.com/news/blog/the-runtime-reckoning-how-the-agentic-evolution-is-reshaping-security/#respond Thu, 05 Mar 2026 18:50:47 +0000 https://www.dynatrace.com/news/?p=73311 Observability graphic

AI is fundamentally reshaping the speed and scale of cyberattacks. Financially and geopolitically motivated threat actors are executing sophisticated attacks against Fortune 500 enterprises that frequently bypass perimeter defenses.

The post The runtime reckoning: How the agentic evolution is reshaping security appeared first on Dynatrace news.

]]>
Observability graphic

AI is compressing attack timelines

The organizations pulling ahead are those with runtime visibility into the applications that directly generate and impact revenue.

Stats about AI usage

What AI is changing

In a report published in November 2025, Anthropic documented the first large-scale AI-orchestrated cyberattack campaign, in which AI autonomously performed 80–90% of attack operations with minimal human intervention. This signaled a turning point.

The threat actors causing the most damage today aren’t traditional nation-state actors or advanced persistent threats (APTs)—they’re financially motivated collectives like Scattered Spider. AI hasn’t invented fundamentally new attack techniques yet, but it has significantly increased attack volume and velocity—CrowdStrike observed an 89% increase in attacks by AI-powered adversaries in 2025 alone.1 Techniques once reserved for state-sponsored groups are now accessible to cybercriminals through AI-assisted tooling.

The same automation accelerating attacks also allows for faster defense. But with an average time of only 29 minutes from initial access to lateral movement, defenders relying on periodic assessments are structurally disadvantaged. The capability that matters now is continuous exposure assessment—analyzing real-time signals across all cloud assets and acting before adversaries complete their objectives.

Two battlefronts: One adversary

Traditional attack vectors haven’t disappeared—supply chains, endpoints, physical security, and human-targeted attacks are increasing just as rapidly. But business-critical applications have become a focal point of attack operations because lower entry barriers make them accessible to more threat actors. The applications that process transactions, manage customer data, and orchestrate supply chains are precisely where runtime compromise has the greatest business impact.

Organizations must drive parallel initiatives: one focused on people, identity governance, and communications; another on runtime protection, continuous exposure management, and application-layer visibility. These require different tools and skills—but critically, connected processes and connected insights. Siloed security functions are precisely what sophisticated attackers exploit.

Why runtime is the new frontline

Initiatives like Anthropic’s Claude Code Security and OpenAI’s reasoning-based vulnerability detection are collapsing scanning stages and shortening developer feedback loops. This is real progress. But production remains the definitive validation point. AI-generated code and accelerated release cycles introduce risks that surface only at runtime—pipeline scanners offer helpful signals but lack the reliability and context of a unified runtime view.

In an agentic ecosystem, detection and response can’t be focused on the perimeter—they must be intrinsic to the application and infrastructure. Environment-aware malware and prompt injection against enterprise AI agents aren’t visible at the perimeter. They’re visible at runtime. The XZ Utils backdoor (CVE-2024-3094) demonstrated this perfectly: malicious code passed all static analysis, activating only when loaded by sshd on targeted Linux distributions in production.7

Autonomous workflows that lack visibility into security risk operate with an incomplete picture. Integrating security context into agentic decision-making—rather than treating it as a separate operational domain—will be essential as enterprises scale these capabilities.

– IDC Link, Dynatrace Perform 2026: From Observability to Supervised Autonomous Operations (Doc #lcUS54307526, February 2026)

Security that’s embedded, not bolted on

The convergence of observability and security isn’t theoretical—it’s operational. IDC analysis notes that organizations expect the next generation of security to be embedded in solutions rather than bolted on. Dynatrace application security capabilities—runtime vulnerability analytics and runtime application protection—operate within Grail alongside observability and business data, sharing the same contextual mapping and causal dependency graph.

Resilience as a competitive advantage

Fortune 500 organizations that embed security into procurement, prioritize supplier maturity assessments, and integrate threat intelligence into operations are positioned to protect both infrastructure and the applications that drive revenue. Forward-thinking organizations recognize the agentic evolution as an opportunity to align security investments with business outcomes and build operational resilience that allows confident growth.

In a world where adversaries move from initial access to lateral movement in 29 minutes or less and autonomous agents make decisions at machine speed, the organizations that thrive will be those with unified runtime visibility.

Clarifying the competitive narrative

There is a prevailing belief that frontier AI labs—Anthropic with Claude Code Security, OpenAI with Codex5—are making traditional security tools obsolete. This belief is partially correct, but it fundamentally misunderstands which tools are being displaced.

What frontier labs are changing

These capabilities collapse scanning stages in the IDE and CI/CD pipeline. They find vulnerabilities through reasoning rather than pattern matching, uncovering business-logic flaws that static analysis misses. This is genuine progress—and it is expected to reduce certain vulnerability classes over time. SQL injection, for example, may become less prevalent as LLM-generated code matures and shift-left tools are integrated throughout the agent development lifecycle.

What frontier labs don’t address

Frontier lab security tools operate before deployment. They don’t see how code behaves in production—outside of API integrations that push context to them. They can’t detect configuration drift, environment-aware malware that activates only at runtime, or prompt injection against live AI agents. Anthropic describes Claude Code Security as an evolution of static analysis: rather than matching known patterns, it “reads and reasons about your code the way a human security researcher would.”4 Static analysis operates on code before it runs. These tools sit at completely different points in the security lifecycle from runtime protection.

What Dynatrace surfaces today

These capabilities deliver findings that frontier lab tools can’t—evidence of what is actually happening in production, not predictions based on code analysis.

Runtime Vulnerability Analytics

Identifies vulnerabilities in the context of actual execution—not theoretical exposure, but real risk based on how code runs in production.

Runtime Application Protection

Detects and blocks exploit attempts as they happen—the defensive layer that shift-left tools structurally can’t provide.

Security Posture Management

Surfaces misconfigurations and compliance gaps in live cloud infrastructure, including Kubernetes environments.

The evolving threat landscape

The Open Worldwide Application Security Project (OWASP) Top 10 for Agentic Applications (2026) report, provides consensus-based guidance and introduces new categories—Agent Behavior Hijacking, Tool Misuse and Exploitation, and Identity and Privilege Abuse6—directly relevant for organizations deploying autonomous agents. For any organization running business-critical applications, runtime visibility remains the validation layer that confirms whether security controls actually work.

What this means in practice

Organizations using Dynatrace already have the core capabilities required for agentic security.

  • Security findings flow directly to development and SRE teams through workflows that include clear remediation guidance.
  • Natural-language queries on security findings are already available through Dynatrace Intelligence.
  • Dynatrace Intelligence allows agentic remediation, allowing automation to act within goals and constraints defined by humans and for humans to retain the flexibility to stay in the loop (see Agentic workflows in Dynatrace documentation).

What sets the Dynatrace approach apart is the combination of deterministic runtime findings with agentic workflows—enabling remediation that is reliable, repeatable, and grounded in real production evidence.


Article citations

  1. CrowdStrike 2026 Global Threat Report. Average eCrime breakout time (initial access to lateral movement) fell to 29 minutes in 2025, a 65% increase in speed from 2024. Fastest observed: 27 seconds. CrowdStrike also observed a 89% increase in attacks by AI-enabled adversaries compared with 2024. crowdstrike.com/en-us/blog/crowdstrike-2026-global-threat-report-findings
  2. Fortune, “Feds are hunting teenage hacking groups like ‘Scattered Spider’ who have targeted $1 trillion worth of the Fortune 500 since 2022,” January 2026.
  3. Anthropic, “Disrupting the first reported AI-orchestrated cyber espionage campaign,” published November 2025. Attack detected mid-September 2025. AI executed 80–90% of tactical operations independently.
  4. Anthropic, “Claude Code Security,” anthropic.com/news/claude-code-security, 2026.
  5. OpenAI, “Codex,” platform.openai.com/docs/codex.
  6. OWASP GenAI Security Project, “Top 10 for Agentic Applications 2026,” genai.owasp.org, December 2025.
  7. CVE-2024-3094 (XZ Utils). Backdoor activated only when loaded by sshd on targeted Linux distributions — bypassing all static analysis. Wired, April 2024.
  8. Dynatrace Intelligence: “supports in-context natural language for investigation and guided next steps.” docs.dynatrace.com/docs/dynatrace-intelligence

The post The runtime reckoning: How the agentic evolution is reshaping security appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/the-runtime-reckoning-how-the-agentic-evolution-is-reshaping-security/feed/ 0
Redefining cloud operations: Dynatrace brings intelligence to observability https://www.dynatrace.com/news/blog/redefining-cloud-operations-dynatrace-brings-intelligence-to-observability/ https://www.dynatrace.com/news/blog/redefining-cloud-operations-dynatrace-brings-intelligence-to-observability/#respond Wed, 28 Jan 2026 16:55:20 +0000 https://www.dynatrace.com/news/?p=72671 Hyperscalers: Azure, AWS, and Google Cloud

Managing cloud environments has never been more complex; Dynatrace is redefining cloud operations to make them simple and easy. As organizations adopt hyperscaler technologies and cloud native architectures, traditional monitoring tools fall short, leaving teams with fragmented data, manual troubleshooting, and slow incident resolution. Dynatrace’s newly enhanced AI-powered Cloud Platform Operations for AWS, Azure, and Google Cloud eliminates these challenges by unifying all observability signals into a single platform, delivering real-time visibility, proactive insights, and automated remediation. This approach reduces risk, accelerates recovery, and optimizes costs, allowing enterprises to move from reactive firefighting to autonomous cloud operations.

The post Redefining cloud operations: Dynatrace brings intelligence to observability appeared first on Dynatrace news.

]]>
Hyperscalers: Azure, AWS, and Google Cloud

Enterprise cloud has outpaced traditional reactive monitoring. As workloads multiply across AWS, Azure, and Google Cloud—and extend into data centers—teams find themselves contending with vast amounts of telemetry, ephemeral infrastructures, and constant change. Visibility is tough; clarity of impact, ownership, and root cause is tougher.

Today’s tech stacks are distributed by design, featuring containers, serverless architectures, managed services, and data planes that can spin up and down in seconds. High-cardinality signals, fragmented logs, and partial traces create blind spots across accounts, subscriptions, and projects. In hybrid setups, network overlays, identity boundaries, and platform services blur the lines between where issues begin and how they propagate.

Organizationally, the challenge is even bigger. Platform, SRE, Dev, SecOps, and FinOps each hold a piece of the truth. However, inconsistent tagging and governance, manual runbooks, and handoffs across time zones can slow MTTR, fuel tool sprawl, and inflate budgets. The goal is to provide clear risk and cost insights, reduce noise, and deliver faster, context-rich answers.

However, visibility alone isn’t enough to achieve this goal. AI-driven context, such as topology-aware analytics, causal correlation, and safe automation, shifts operations from reactive firefighting to proactive prevention that’s aligned with SLOs, compliance, and budgets.

AWS Business Resilience dashboard in Dynatrace screenshot

Today, we’re introducing Dynatrace enhanced cloud operations for AWS, Azure, and Google Cloud, delivering complete visibility across all your cloud and hybrid environments with deeper, actionable insights. If you’re ready to take the guesswork out of monitoring your cloud environments, read on to see how Dynatrace helps you collaborate faster, resolve issues earlier, and run at enterprise scale with confidence.

What is Dynatrace Cloud Platform Operations?

Dynatrace Cloud Platform Operations takes the guesswork out of monitoring by redefining how organizations manage complex cloud environments. By unifying all cloud signals—metrics, logs, and events—into a single AI-powered platform, Dynatrace delivers complete visibility and actionable insights at scale.

Every data point is enriched with context and analyzed by Dynatrace AI, enabling proactive automation and informed decision-making. This ensures faster troubleshooting, better resource efficiency, and simplified onboarding—all without the need for additional infrastructure components.

Cloud overview dashboard in Dynatrace

Organizations eliminate blind spots and accelerate resolution by gaining real-time visibility into the health, performance, and configuration of their cloud resources. This allows them to maintain the correct governance through tag-based control, ownership, and cost allocation, thereby aligning teams while reducing operational noise.

Services list in Dynatrace

Additionally, the latest advancements in cloud observability provide real-time visibility into resource health, performance, and configurations, allowing teams to act quickly and decisively while reducing the time to resolution.

Start monitoring your cloud environment

Spin up monitoring without spinning up your infrastructure. The new cloud connections are fully managed by Dynatrace and guided by a simple wizard. It gets you from setup to actionable telemetry in minutes; no agents to wrangle, no custom pipelines to maintain.

Once connected, you gain instant visibility across all your cloud environments. Connect your AWS, Azure, or GCP accounts once, and Dynatrace auto‑ingests everything from compute and databases to storage, networking, and security configurations. Platform coverage has been expanded to capture more data types than ever, including metrics for any supported cloud service and a richer set of cloud events, such as hyperscaler‑native security alerts.

New AWS connection in Dynatrace

The new cloud connections are:

  • Secure by design: native authentication offers least‑privilege access and auditable permissions.
  • Practitioner‑friendly: a step‑by‑step wizard with built‑in checks and defaults works across accounts and regions.
  • Zero overhead: the Dynatrace platform manages the connection end‑to‑end, so you focus on insights, not maintenance.

With all your cloud data unified, Dynatrace automatically applies the tags you already use in your cloud environments, bringing your existing operational model directly into the platform. Your existing cloud tags instantly drive access control, ownership, cost allocation, alert routing, and preventive workflows, with no manual tagging or re‑mapping required. Additionally, tag‑enriched signals keep operations aligned by allowing you to filter everything by owner, app, or environment for precise alert routing and actionable insights. This paves the way for more advanced analytics and clear insight into what’s happening across every environment.

The enhanced cloud ingest not only unifies telemetry and context; it also captures all configuration details (VPCs, load balancers, security groups, subnets, network services, compute metadata, and more). The new Smartscape® uses this information, along with cloud monitoring data, to provide a comprehensive, always-accurate topology of your cloud infrastructure. You can use the new Smartscape app to navigate your cloud topology, visualize dependencies, and eliminate hidden or forgotten resources. Ready-made views for AWS and Azure provide a comprehensive, automatically discovered cloud inventory across services, databases, networking layers, and security controls.

Infrastructure overview

Additionally, you get ready-made dashboards, essential metrics, and expanded cross-cloud visibility, providing you with immediate clarity, especially when issues originate on the provider side, thanks to built-in AWS Health event integration.

All of this comes together in our newly enhanced Clouds app, a shared place for teams to analyze services, metrics, events, and logs. Pre-built dashboards and alerts automatically highlight unhealthy resources, helping teams stay ahead of issues, while the streamlined onboarding process eliminates the need for additional collectors or components. It’s fast, safe, and modern—exactly how cloud onboarding should feel. Ready to see it in action? Keep reading to see the magic.

Transform from reactive monitoring to proactive cloud operations

Dynatrace transformed cloud monitoring into proactive cloud operations by combining AI-powered observability with intelligent automation. This allows organizations to move beyond simply identifying problems to actively solving and preventing problems.

With Dynatrace, you can remediate issues before they impact your users, prevent future issues, and optimize your cloud environments to ensure they operate at maximum efficiency and resiliency.

Prevention

Leave reactive firefighting behind and gain the foresight needed to stay ahead of issues. Dynatrace’s AI-driven automation delivers proactive insights that allow you to identify and resolve potential problems before they impact your users. By predicting anomalies and triggering automated workflows, Dynatrace helps maintain high availability and optimal performance across your cloud environments. This proactive approach minimizes downtime, protects user experience, and ensures your teams can focus on strategic initiatives rather than crisis management.

Remediation

Instead of wasting valuable time on lengthy resolution cycles that disrupt business operations, accelerate recovery with intelligent automation. Dynatrace automates root cause analysis and remediation, enabling self-healing workflows that dramatically reduce resolution times. By eliminating manual troubleshooting and streamlining incident response, your teams can focus on driving innovation rather than firefighting. With AI-driven insights and automated corrective actions, Dynatrace ensures issues are resolved quickly and efficiently, thus minimizing impact, improving reliability, and keeping your business moving forward.

Optimization

Stop overspending on infrastructure due to a lack of visibility into resource usage and performance. With Dynatrace, you can continuously optimize both cost and efficiency across your environment. Real-time insights into resource consumption and application performance allow you to identify waste and prevent unnecessary expenses. By leveraging AI-driven analytics and automated recommendations, you can improve cost management and drive peak performance in your applications.

AWS EBS workflow

Ready to experience the new world of Cloud Operations for yourself?

With the introduction of enhanced cloud operations for AWS, Azure, and GCP, you can achieve smarter collaboration, faster issue resolution, and streamlined operations at scale using the Dynatrace Cloud Platform Operations solution.

AWS enhanced cloud operations

Discover how easy it is to get started using proactive monitoring capabilities for AWS. AWS cloud operations are now generally available for all Dynatrace SaaS customers.

Azure enhanced cloud operations

Interested in seeing the magic for your Azure environments?
Join the Azure Preview program

GCP enhanced cloud operations

Are you a GCP customer looking for a cloud operations solution?
Join the GCP Preview program

The post Redefining cloud operations: Dynatrace brings intelligence to observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/redefining-cloud-operations-dynatrace-brings-intelligence-to-observability/feed/ 0
Database monitoring made easy with Dynatrace’s AI-powered solution https://www.dynatrace.com/news/blog/database-monitoring-made-easy-with-dynatraces-ai-powered-solution/ https://www.dynatrace.com/news/blog/database-monitoring-made-easy-with-dynatraces-ai-powered-solution/#respond Wed, 21 Jan 2026 16:21:06 +0000 https://www.dynatrace.com/news/?p=72573 Business process observability graphic

Modern organizations face growing challenges in managing thousands of distributed databases across clouds, regions, and teams, often leading to blind spots and reactive firefighting. Dynatrace’s new Database Monitoring app changes the game by providing unified visibility and AI-powered insights across your entire database landscape. With proactive health scoring, query-level analytics, and seamless integration with application […]

The post Database monitoring made easy with Dynatrace’s AI-powered solution appeared first on Dynatrace news.

]]>
Business process observability graphic

Modern organizations face growing challenges in managing thousands of distributed databases across clouds, regions, and teams, often leading to blind spots and reactive firefighting. Dynatrace’s new Database Monitoring app changes the game by providing unified visibility and AI-powered insights across your entire database landscape. With proactive health scoring, query-level analytics, and seamless integration with application monitoring, teams can quickly identify and resolve issues before they impact users. This solution empowers developers, SREs, and platform engineers to collaborate effectively, reduce complexity, and take ownership of database performance with confidence.

Organizations today rely on thousands of databases spread across clouds, regions, and teams; traditional monitoring tools simply can’t keep up.

In the AI era, databases are increasingly ephemeral; workloads are massively data-intensive, and debugging efforts span across complex, distributed systems. Simply put, traditional tools were never designed for this pace or scale. Fragmented tooling and data silos add complexity, while usage-based costs limit observability, creating blind spots and slowing issue detection. Teams often end up firefighting problems, such as slow queries and deadlocks, leading to operational fatigue and delayed resolutions that impact performance and innovation.

At the same time, as database reliability shifts from DBAs to engineering teams, visibility gaps and expertise challenges grow. Developers often lack feedback on how their code impacts performance, which slows down releases, increases risk, and makes collaboration with SRE and platform teams more challenging. With many developers lacking in-depth database knowledge, it’s crucial to deliver tools that simplify management and provide actionable insights, allowing teams to innovate faster without compromising reliability.

It’s time to transform database observability with an AI-native approach designed specifically for developers, SREs, and platform engineers. Dynatrace’s new Database Monitoring solution is revolutionizing database observability with unified visibility across your entire database estate, from on-premises to hybrid and multicloud environments.

Unified observability across your entire database estate

The new Dynatrace Databases app, powered by the Dynatrace® unified observability platform, takes database monitoring to the next level with deep, actionable insights.

You can now monitor entire database clusters, gain deep AI-powered insights with remediation plans and alerts, and access granular visibility to reduce outages and downtime. Essentially, this means you can now go beyond basic health checks and gain visibility into:

  • Query-level analytics and transaction tracing across services to pinpoint performance bottlenecks.
  • Schema and index behavior analysis to understand structural inefficiencies.
  • Preventive alerts and insights that help you act before issues impact users.
  • Historical and real-time performance trends for smarter optimization and capacity planning.

With all critical data in one place, practitioners get a holistic view of their entire ecosystem. What used to take significant time, resources, and tools is now available in one central place.

Database monitoring Explorer dashboard in Dynatrace

The Databases application was designed to monitor a diverse range of database technologies across on-premises, public cloud, and hybrid cloud environments. With native support for cloud-specific services, Dynatrace delivers comprehensive coverage regardless of your infrastructure strategy.

From reactive to proactive database monitoring

Proactive database monitoring starts here. Instead of waiting for problems to surface and reacting to outages and slowdowns, teams can now prioritize the necessary adjustments to keep their databases healthy and performing optimally, and to ensure their applications run smoothly. This proactive approach drives performance improvements beyond query optimization, allowing you to address issues before they impact users.

With granular database views and new metrics, including replication status, cache hit ratios, lock contention, and server pulse overviews, practitioners gain the clarity needed to make informed decisions.

With real-time intelligence, Databases eliminates silos to help you optimize performance, reduce time to resolution, and deliver a seamless user experience. Combined with APM integration, Dynatrace delivers end-to-end visibility from services down to individual database queries, so you can understand how the layers interact.

How it works

After a seamless and quick integration, Databases immediately provides you with an overview of everything happening across all your databases.

Database monitoring Explorer dashboard in Dynatrace

You can dive deeper into the instances you’re responsible for. Go to the Overview tab and view the Health score, which indicates the instance’s health.

The health score is an innovative feature that allows you to be proactive with your database fleets and make sure they always meet your application needs. In this example, you can see how a low health score indicates a database health issue related to the instance’s Cache Hit Ratio.

Database monitoring health score dashboard in Dynatrace

Once you have a clear indication of what’s preventing a database from maximizing its potential, you can dive deeper into that specific database and the database calls reported as the root cause of the problem, all from the health score.

Database monitoring Activity metrics dashboard in Dynatrace

In the above graph, we can see that the problem is within the airbases. Now that we know which database is causing the problem, we can access it and debug the queries that are contributing to the low database Cache Hit Ratio.

Seamless debugging across databases and services

Experience the full power of Dynatrace’s unified platform by taking your database monitoring a step further. Monitoring your databases and services together in Dynatrace gives you a holistic view of your application ecosystem. This integration allows you to quickly get to the root cause, identify the issues, and transition effortlessly from the Database app to the Services app for deeper analysis, quick and efficient debugging, and resolution.

Let’s take a look at an example. Let’s say one of your databases experiences a Cache Hit Ratio problem, as depicted in the screenshot below.

Database monitoring Activity metrics dashboard in Dynatrace

From the database host overview page, go to the Calling Services tab. Here, you can pinpoint which service is causing, or is impacted by, the issue based on certain metrics:

Database monitoring Calling services dashboard in Dynatrace

Once you identify the problematic service in the Databases app, go to the Services app. There, you can debug the root cause by analyzing the problematic queries in detail.

Database monitoring Database queries dashboard in Dynatrace

Get started

By breaking down silos and delivering intuitive, actionable insights, the new Dynatrace Database Monitoring solution helps teams collaborate seamlessly, reduce complexity, and take ownership of database performance with confidence. From unified observability and AI-powered analytics to proactive health scoring and end-to-end debugging across databases and services, Dynatrace transforms how you manage and optimize your database landscape.

Ready to experience Dynatrace database monitoring for yourself? Explore the full capabilities today and see how effortless it can be.

The post Database monitoring made easy with Dynatrace’s AI-powered solution appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/database-monitoring-made-easy-with-dynatraces-ai-powered-solution/feed/ 0
Dynatrace acquires DevCycle to strengthen feature delivery for modern cloud and AI-native workloads https://www.dynatrace.com/news/blog/dynatrace-acquires-devcycle-to-strengthen-feature-delivery/ https://www.dynatrace.com/news/blog/dynatrace-acquires-devcycle-to-strengthen-feature-delivery/#respond Tue, 13 Jan 2026 20:59:17 +0000 https://www.dynatrace.com/news/?p=72427 Dynatrace and DevCycle

Safer feature releases with real-time insight.

The post Dynatrace acquires DevCycle to strengthen feature delivery for modern cloud and AI-native workloads appeared first on Dynatrace news.

]]>
Dynatrace and DevCycle

Modern software teams are shipping faster than ever. With cloud-native architectures, AI-powered development, fast releases, and continuous delivery pipelines, new features flow continuously from commit through development and testing into production in minutes.

While release velocity has accelerated, release risk has not disappeared. As modern software environments grow more complex, releases have become harder to understand and riskier to manage, placing increasing strain on developer experience while forcing organizations to balance rapid innovation with the reliability that customers expect.

That’s why Dynatrace has acquired DevCycle, a feature management platform built on the OpenFeature standard, to help developers, SREs, and platform teams bring progressive delivery for AI-native applications directly into the Dynatrace platform.

Why feature management matters in modern delivery

Feature flags are a core control plane for progressive delivery including:

  • Canary deployments
  • Blue-green releases
  • Experimentation
  • Rapid mitigation when something breaks

Feature flags are frequently used as standalone control mechanisms, allowing teams to change application behavior without clear visibility into the consequences of those changes. While teams can toggle features instantly, they lack direct insight into how each decision affects performance, reliability, and user experience in real time.

What DevCycle adds to Dynatrace

DevCycle brings enterprise-grade feature management built on OpenFeature, the open standard that Dynatrace helped establish. Together, Dynatrace and DevCycle will enable teams to:

  • Reduce risk: Release features to small cohorts, validating behavior, and scaling gradually based upon real-time performance and error monitoring.
  • Experimentation: Compare feature variants, models, or prompts using real traffic and real telemetry instead of offline assumptions.
  • Improve MTTR: Use Dynatrace causal analysis to quickly pinpoint when a specific feature is driving an incident, and act immediately.
  • Improve developer experience and flow: Give developers a fast, intuitive way to ship, test, and control features with immediate feedback from production telemetry.

Built on OpenFeature

Dynatrace recognized the importance of progressive delivery early by initiating and leading the OpenFeature standard in 2022. It is now a CNCF project and the industry’s open vendor-neutral standard for feature flagging.

Because DevCycle is OpenFeature-native, customers can use Dynatrace with DevCycle or with any OpenFeature-compliant feature flag system. This preserves openness, avoids lock-in, and gives teams flexibility as their platforms evolve.

Feature flags and observability in practice

Dynatrace brings together development, operations, and business teams to deliver software more efficiently, securely, and reliably. By integrating feature flagging, teams gain a single, contextual view of intent, execution, and outcome across the software development lifecycle.

Practical examples include:

  • Roll out a new checkout flow only to premium users and immediately observe performance, errors, and conversion impact.
  • Automatically disable a feature for a specific region if Dynatrace detects degradation.
  • Test new AI prompts to compare latency, quality, and user behavior in real time.

This integration enables advanced use cases such as real-time, health-driven feature control and experimentation while integrating naturally into AI-assisted workflows. For example, an AI assistant can query the Dynatrace MCP Server to understand blast radius and KPI impact, then safely reduce exposure or disable a feature without redeploying code. By unifying feature management and experimentation with real-time observability, AI-powered insights, and automation, teams can make precise changes to application behavior without redeployments or manual intervention.

What’s next

This acquisition further advances observability into an active system of control. By integrating the precise runtime controls and centralized feature concept from DevCycle with comprehensive insights from Dynatrace, we will move closer to realizing our vision of an intelligent resilience platform capable of self-remediation, prevention, and continual optimization.

We are already working to integrate DevCycle capabilities into Dynatrace. Stay tuned for updates, and thank you for being part of the Dynatrace community!

The post Dynatrace acquires DevCycle to strengthen feature delivery for modern cloud and AI-native workloads appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/dynatrace-acquires-devcycle-to-strengthen-feature-delivery/feed/ 0
Want to catch the AI native wave? Learn the lessons of the cloud-native shift https://www.dynatrace.com/news/blog/want-to-catch-the-ai-native-wave-learn-the-lessons-of-the-cloud-native-shift/ https://www.dynatrace.com/news/blog/want-to-catch-the-ai-native-wave-learn-the-lessons-of-the-cloud-native-shift/#respond Mon, 24 Nov 2025 18:35:51 +0000 https://www.dynatrace.com/news/?p=72006 PurePerformance podcast

From the personal computer to the internet to mobile and cloud native, transformative technologies bring new challenges and opportunities. While every innovation is different, Pini Reznik says there are patterns that can be applied when adopting any new technology. In the latest PurePerformance podcast, Reznik talks to hosts Andi Grabner and Brian Wilson about his […]

The post Want to catch the AI native wave? Learn the lessons of the cloud-native shift appeared first on Dynatrace news.

]]>
PurePerformance podcast

From the personal computer to the internet to mobile and cloud native, transformative technologies bring new challenges and opportunities. While every innovation is different, Pini Reznik says there are patterns that can be applied when adopting any new technology.

In the latest PurePerformance podcast, Reznik talks to hosts Andi Grabner and Brian Wilson about his new book From Cloud Native to AI Native: Catching the Next Wave of Innovation, and how organizations can take a pragmatic approach to AI adoption.

The AI Native transformation process

Here’s how Reznik describes the process of adoption AI, or any other transformational technology:

  • Experiment: Start with a small, skunkworks-style team operating in a sandbox. Their mission isn’t to deliver production-ready systems but to experiment, learn, and identify viable business cases.
  • Find a small win: Once you find a promising use case, build a minimal viable product that demonstrates tangible value. Success here justifies incremental investment.
  • Build a foundation: Build the infrastructure, develop the team structure, and foster the new culture necessary to adopt the technology at scale.
  • Scale up: Expand as you grow, transitioning from legacy systems to new solutions carefully.

The biggest mistake organizations make is skipping the first two phases and jumping straight to large-scale initiatives. That’s why you hear so much about failed AI pilots: Organizations try to scale before they validate their use cases.

“Transformation isn’t a one-time project; It’s a way of thinking.”

— Pini Reznik, CEO and Co-Founder, re:cinq

AI for the little guy

In the previous episode of PurePerformance, Laura Tacho made the case for AI’s true potential in software development being not in code generation but in speeding up feedback loops and helping ensure that developers build the right things. Where Tacho focused on developer experience and the software development lifecycle, Reznik focuses on organization-wide potential, particularly for companies that haven’t traditionally built much, if any, software in-house. AI could make it affordable for smaller companies to build custom software based around their unique value propositions.

Our perspective: Start small, validate value early, and ensure every step is measurable and observable. When teams can see the impact of AI in real time—on performance, cost, and outcomes—they make smarter decisions and scale with purpose.

To to dive into the world of software performance and innovation, listen to the latest episode of PurePerformance.

The post Want to catch the AI native wave? Learn the lessons of the cloud-native shift appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/want-to-catch-the-ai-native-wave-learn-the-lessons-of-the-cloud-native-shift/feed/ 0
Unified observability: Why storing OpenTelemetry signals in one place matters https://www.dynatrace.com/news/blog/unified-observability-why-storing-opentelemetry-signals-in-one-place-matters/ https://www.dynatrace.com/news/blog/unified-observability-why-storing-opentelemetry-signals-in-one-place-matters/#respond Mon, 09 Jun 2025 22:55:31 +0000 https://www.dynatrace.com/news/?p=69431 OpenTelemetry signals

Thinking of observability in terms of the traditional “three pillars” limits the value of signals coming from data sources, including OpenTelemetry. By unifying these telemetry signals in a single analytics backend, you can connect all these details in context.

The post Unified observability: Why storing OpenTelemetry signals in one place matters appeared first on Dynatrace news.

]]>
OpenTelemetry signals

In observability’s early days, we often talked about the “three pillars.” That is, traces, logs, and metrics, which gave us the information to make our systems observable. The problem with referring to these three signals as “pillars” is that it implies they’re siloed and therefore independent of each other, when in fact, the exact opposite is true.

To harness the true power of observability, you need to treat these signals not as pillars, but as three strands that make up a braid, as OpenTelemetry (OTel) co-founder Ted Young so aptly put it. While observability as a whole also encompasses user behavior and security data among other signals, the main OpenTelemetry signals–traces, logs, and metrics–each serve a different and important purpose and contribute to the observability story, giving us the full picture of what’s happening in our systems. We achieve even greater value from our OpenTelemetry signals through context, the glue that connects related information, revealing patterns and relationships, and making raw data more meaningful and actionable.

Stronger together!

And yet, many organizations practicing observability still tend to send different OpenTelemetry signals to different backends for storage and analysis. The problem with this separation is two-fold. First, you’re not storing all of the signals in one place. This siloes the data, making it impossible to correlate and get insights from the data. Second, you’re having to go back and forth between different tools to look at your signals, try to correlate them, and understand what’s going on. How can you effectively analyze all of your telemetry data if it’s not all stored in the same place?

This problem is further amplified when you consider how some organizations send telemetry data from different applications to different vendors. For example, an organization might have App A send metrics to SaaS Tool X, and traces and logs to SaaS Tool Y. App B sends traces to SaaS Tool J, logs to SaaS Tool L, and metrics to self-hosted Tool M. Let’s not forget the teams that go rogue and decide to do their own thing. See that tower under Bob’s desk? It’s running a whole suite of self-hosted open-source observability tools, and App C is sending its telemetry signals there.

diagram showing the connections among applications, OpenTelemetry signals, and different vendors
Figure 1. Sending telemetry data to multiple backends silos data, and makes it harder to analyze.

Moving from “swivel chair” observability to “unified observability”

To borrow a term coined by my husband, these types of organizations are practicing “swivel chair” observability, and to be honest, calling it “observability” at this point is being very generous. You don’t have a braid or even pillars. You have islands of pillars.

To leverage the true power of observability, we need a single pane of glass that provides “unified observability.” One single platform for storing, viewing, correlating, and analyzing your telemetry signals.

This enables teams to improve data analysis, streamline workflows, and innovate. It also allows them to focus on delivering better results with less effort. At the same time, organizations get the support they need to take on whatever challenges are thrown at them.

architecture diagram showing the connections among applications and OpenTelemetry signals consolidating them to unified observability with Dynatrace
Figure 2. Sending telemetry data to a single backend allows us to view and analyze our data in context.

Remember, though, that observability must go beyond tooling if it’s to be truly effective and sustainable. Observability must be treated as a team sport, where everyone in an organization plays a role in making systems observable. This should be complemented by enterprise oversight to guide tooling choices, patterns, and best practices.

Unified observability with Dynatrace

By using the Dynatrace observability and security platform as your OpenTelemetry analytics backend, however, you can connect all these details in context. Dynatrace stores all data in Grail™, a unified and purpose-built data lakehouse optimized for storing and analyzing not just traces, logs, and metrics, but also security, business events, and RUM/behavioral data.

Consider some examples.

Distributed tracing

The first example shows Dynatrace Distributed Tracing. Note how we can see a distributed trace and its associated logs and metrics in Dynatrace. Together, these pieces of information all help paint a picture of what is happening in our system.

Screenshot of Dynatrace Distributed Tracing dashboard showing the ability to trace OpenTelemetry signals in context.
Figure 3. The Distributed Tracing App showing related log entries (on the bottom) and metrics (on the right) in context when analyzing a trace.

Searching logs

The next example shows how you can use Dynatrace Query Language (DQL) to search for logs with a specific error message and to join them to the traces associated with those logs.

Screenshot showing Dynatrace query language query on unified observability data
Figure 4: Joining logs and traces to find spans associated with logs having a specific error message.

Investigating problems

Finally, you can use the Dynatrace Problems app for root cause analysis, showing a problem that was detected, the impacted entities, and its associated metrics and logs. And for those who only occasionally look at observability data, Dynatrace also provides Davis CoPilot™ using natural language queries to quickly explain what this data means, and propose concrete remediation steps.

Screenshot of Dynatrace Problems app showing a problem detected using unified observability analysis.
Figure 5. Dynatrace’s automated anomaly and root cause detection – showing all relevant data in context.

All of this is possible because of the integrated storage approach of Grail.

By storing all OpenTelemetry signals in context under one roof, users can ask meaningful questions, get useful answers, and act effectively on what they learn. That is, Dynatrace makes unified observability possible.

Unified observability leads to better outcomes

Unifying OpenTelemetry signals in a single analytics backend helps you gain critical context that identifies patterns and relationships among all observability signals. This context helps you spot the meaning in raw data, making it easier and faster to anticipate problems, take action, and automate responses.

To learn more about how Dynatrace augments, amplifies, and accelerates actionable answers from OpenTelemetry, check out these resources.

Also check out our playlist featuring the video series, Dynatrace Can Do THAT with OpenTelemetry?, in which my teammate Andi Grabner teaches me all sorts of cool things that Dynatrace can do with OpenTelemetry data ingested into Dynatrace using the OTel Collector.

The post Unified observability: Why storing OpenTelemetry signals in one place matters appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/unified-observability-why-storing-opentelemetry-signals-in-one-place-matters/feed/ 0
Kubernetes logging made easy: Comprehensive Kubernetes observability with Dynatrace https://www.dynatrace.com/news/blog/kubernetes-logging-made-easy-comprehensive-kubernetes-visibility-with-dynatrace/ https://www.dynatrace.com/news/blog/kubernetes-logging-made-easy-comprehensive-kubernetes-visibility-with-dynatrace/#respond Thu, 10 Apr 2025 15:47:09 +0000 https://www.dynatrace.com/news/?p=68822 Abstract image with Kubernetes logo for KubeCon EU 2025 and Kubernetes misconfiguration

The Dynatrace Log Module for Kubernetes is a fully automated, full-service approach to Kubernetes log management and analytics. Dynatrace is your one-stop shop and go-to solution for all Kubernetes needs. Logging is integral to Kubernetes monitoring In the ever-evolving software development landscape, logs have always been—and continue to be—one of the most critical sources of […]

The post Kubernetes logging made easy: Comprehensive Kubernetes observability with Dynatrace appeared first on Dynatrace news.

]]>
Abstract image with Kubernetes logo for KubeCon EU 2025 and Kubernetes misconfiguration

The Dynatrace Log Module for Kubernetes is a fully automated, full-service approach to Kubernetes log management and analytics. Dynatrace is your one-stop shop and go-to solution for all Kubernetes needs.

Logging is integral to Kubernetes monitoring

In the ever-evolving software development landscape, logs have always been—and continue to be—one of the most critical sources of insight. Log data is essential, whether you’re troubleshooting and conducting forensics into past problems, investigating potential security issues, or debugging in real time.

However, the distributed and ephemeral nature of Kubernetes means that logs are scattered across multiple nodes and pods, making it difficult to ensure that all logs are preserved, easy to access, and enriched with necessary context for future analytics.

A robust log collection solution must ensure:

  • Central log management. Logs must be centralized, preserved, easily accessible, and connected with other signals from your Kubernetes environments to allow effective monitoring and troubleshooting.
  • Context-aware and topology-rich logs. Logs must be enriched with metadata and contextual information, allowing powerful log analytics of all distributed traces, metrics, and events within the Kubernetes topology. This allows analysis of logs from out-of-memory containers or pods with application errors.
  • Powerful access controls. The ability to define fine-grained access control for logs based on Kubernetes namespace or cluster level is essential. This ensures that users can only access the necessary logs, helping maintain the highest security, compliance, and operational efficiency.

At Dynatrace, we support whichever integration you use to collect logs from your systems, whether it’s OpenTelemetry, Fluent Bit, or other tools. Moreover, we’re continuously looking to improve our customers’ experience, which is why we’ve enhanced our fully automated, full-service log streaming solution with the Dynatrace Log module.

The updated solution allows you to:

  • Stream and analyze logs from your Kubernetes environments without running OneAgent on each Kubernetes node.
  • Flexibly choose the level of observability you need. For example, you can start small with only Kubernetes platform monitoring and log analytics and grow into comprehensive observability maturity with distributed tracing, security analytics, real-user monitoring, and more. This flexibility extends to privileges, ensuring you can grant the least possible number of privileges in Kubernetes to get the needed data.
  • Get insights into logs from short-lived containers and pods such as InitContainers or Jobs.
  • Easily onboard log analytics within the Kubernetes app and control log ingestion and management centrally to ensure an optimal experience. Rather than configuring log collection locally, as you might have done with open source log shippers you’ve used, you can now centrally define and manage which logs Dynatrace collects.
  • Effortlessly get logs in context as the Dynatrace Log Module for Kubernetes is fully managed for you; Dynatrace handles all upkeep, configuration, and lifecycle management.

New Kubernetes logging capabilities integrated the Dynatrace platform

Let’s look at how the new capabilities integrate with the Dynatrace platform.

One platform, complete log value

Not having to lift a finger to stream your logs is awesome, but that’s not the only value the Dynatrace Log Module provides.

Central log management

You have the ability to fully manage the lifecycle and configuration of your Kubernetes log collection within the Dynatrace platform. Gone are the days when you needed to manage the configuration of a log shipper across all your clusters. Dynatrace allows you to do this on one platform. You can centrally manage and control data collection, data masking, data dropping, data transformation, and data retention vs distributed configurations at the source. Additionally, the ability to dynamically define retention periods and compliance settings gives you much more flexibility when working and allows you to get the data you need to resolve issues easily. This ensures the smooth and reliable operation of applications running on Kubernetes.

Troubleshooting and remediation

Logs are critical to understanding and maintaining system health and performance. Dynatrace helps you quickly troubleshoot and easily remediate in a variety of ways, allowing you to:

  • Inspect and understand logs from crashing application containers and other workloads in the Dynatrace Kubernetes app.
  • Slice and dice log data with traces and Kubernetes topology in Notebooks with DQL.
  • Drive evidence-based investigations for security issues and data forensics.
  • Easily derive metrics and events from logs for dashboarding and prediction.
  • Automate remediations with log data in the Workflows app.
  • Dive into log data to explore surrounding logs and patterns using the Dynatrace Logs app.

Workload error logs to trace video thumbnail

Cost management and allocation

Dynatrace allows you to manage costs by controlling the amount of ingested logs, the retention period for which logs are kept for analysis, and query timeframes for your analytics. You can also manage ingest and retention by filtering log sources, managing buckets, and picking the right license model for your needs.

Additionally, you can leverage existing labels or annotations to enrich log data with cost-allocation metadata, making it easier to allocate expenses accurately. This allows for precise cross-charging, helping ensure that bills are charged to the correct departments.

Log enrichment

Dynatrace enriches every log line with additional Kubernetes metadata. In addition to cost-allocation use cases, you can enrich logs with security context and other common Kubernetes context metadata.

This allows you to define IAM rules to control access to log data so that only users with specific roles or permissions can query certain logs. This ensures that teams see only the logs they need, helping your organization maintain the highest levels of privacy and security. Further, if you also monitor your applications with Dynatrace, your log lines will be enriched with span and trace IDs for context-rich analytics of log, trace, and metrics data.

Get started

Seamlessly manage your logs with the new Dynatrace Log Module. Get started today to see how Dynatrace can help you with your observability and security needs.

If you’re not yet a Dynatrace customer, a Dynatrace Playground environment can provide you with ready-to-use data that allows you to troubleshoot, remediate, manage costs, and more.

Dynatrace and the Dynatrace logo are trademarks of the Dynatrace, Inc. group of companies. All other trademarks are the property of their respective owners.

The post Kubernetes logging made easy: Comprehensive Kubernetes observability with Dynatrace appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/kubernetes-logging-made-easy-comprehensive-kubernetes-visibility-with-dynatrace/feed/ 0
Observability as Code: DIY with Crossplane https://www.dynatrace.com/news/blog/observability-as-code-diy-with-crossplane/ https://www.dynatrace.com/news/blog/observability-as-code-diy-with-crossplane/#respond Fri, 14 Feb 2025 15:16:49 +0000 https://www.dynatrace.com/news/?p=67882 Observability as code graphic

In today’s rapidly evolving cloud-native landscape, managing observability efficiently can be a game-changer for any platform. During our recent tech talk at #KCDAustria, Observability as Code – DIY with Crossplane, we demonstrated how the Upjet project can streamline the creation of a custom Crossplane Dynatrace provider, unlocking robust observability capabilities in Kubernetes environments. This blog […]

The post Observability as Code: DIY with Crossplane appeared first on Dynatrace news.

]]>
Observability as code graphic

In today’s rapidly evolving cloud-native landscape, managing observability efficiently can be a game-changer for any platform. During our recent tech talk at #KCDAustria, Observability as Code – DIY with Crossplane, we demonstrated how the Upjet project can streamline the creation of a custom Crossplane Dynatrace provider, unlocking robust observability capabilities in Kubernetes environments. This blog post covers what we did and how we did it.

Crossplane and Dynatrace

In this blog post, we will deploy a monitoring dashboard with alerts and notifications as one simple Kubernetes resource. To help us achieve our goal, we have to set up a Kubernetes cluster and infrastructure components. But before that, we need to discuss some important concepts and patterns.

Configuration as Code

Managing vast amounts of configurations for organizational setups at scale is a hard problem to solve since they span over many tools and providers and have many different contributors.

However, there is a pattern for remedying many of these problems: treating these configurations as declarative code instead of applying changes manually.

Originally dubbed Infrastructure as Code, the pattern can be generalized and used for anything that provides a proper interface—simply put, Configuration as Code.

While many tools and their respective approaches exist to write and apply such configuration, one of the most interesting recent developments is extending Kubernetes and using its readily available REST API and reconciliation loops.

The operator pattern in Kubernetes

The mechanism that Kubernetes provides for interface extension is called the operator pattern. This is a powerful mechanism for automating the management of complex applications. It extends the Kubernetes API via custom resource definitions (CRDs), enabling new object types to be created. These are constantly watched by custom controllers, so-called operators.

Operators are designed with a “reconciliation loop,” meaning they continuously compare a resource’s real state against its desired state. When a resource deviates, the operator brings it back into alignment. This is the essence of Kubernetes automation and declarative infrastructure.

Crossplane and the operator pattern

Crossplane builds on the operator pattern and extends Kubernetes beyond managing containerized workloads. Through the Kubernetes API, it enables you to define and provision—among other things—cloud infrastructure resources such as databases, compute instances, and networking components.

Crossplane providers implement the operator pattern for external systems (for example, AWS, GCP, Azure). When you install one, it installs its external resources as Kubernetes native custom resource definitions (CRDs). The provider’s controller watches for changes in the desired state of objects—instanced from these CRDs, reconciling them to ensure they match your expectations.

Compositions

In Kubernetes, low-level resources are managed by high-level resources. Crossplane also allows you to build high-level resources using the Composition pattern.

An example would be an Application resource that abstracts away details like database, network, and compute needs. A Crossplane composition enables you to build something like this, giving you control over the new interface (also a Kubernetes CRD) and the implementation (which low-level resources are created and how they are created).

Using the Upjet project

Now that we have an overview of the concepts, let’s look at how we implement our demo. When building a custom Crossplane provider, we can take two approaches: build the provider from scratch or leverage an existing tool. We opted for the latter, by using the Upjet project, which automates the creation of Crossplane providers based on existing Terraform providers. Here’s why:

  1. Speed and simplicity: Writing a provider from scratch requires a deep understanding of the external system API and how Crossplane manages resources. Upjet allows us to generate a provider much faster by transforming Terraform provider schemas into Crossplane CRDs, significantly reducing the development time.
  2. Reuse of Terraform providers: Upjet allows us to tap into the vast ecosystem of Terraform providers. Since there are already many well-established Terraform providers for various cloud platforms and services, using Upjet means we don’t have to reinvent the wheel.

Upjet project diagram with Crossplane and Dynatrace

Generating a new Crossplane provider with Upjet

Creating a new Crossplane provider using Upjet is a streamlined process that allows you to extend Crossplane’s capabilities with minimal setup. Follow the official Upjet documentation to get started.

In our demonstration at the KCD Austria tech talk, we showcased how to build a Dynatrace provider. Below are the detailed steps we followed.

Step 1: Adjust the Makefile

The Makefile needs to reference the official Dynatrace Terraform module. This adjustment allows Upjet to use the correct source when generating the provider.

Here’s an example configuration:


export TERRAFORM_PROVIDER_SOURCE ?= dynatrace-oss/dynatrace
export TERRAFORM_PROVIDER_REPO ?= https://github.com/dynatrace-oss/terraform-provider-dynatrace
export TERRAFORM_PROVIDER_VERSION ?= 1.66.0
export TERRAFORM_PROVIDER_DOWNLOAD_NAME ?= terraform-provider-dynatrace
export TERRAFORM_PROVIDER_DOWNLOAD_URL_PREFIX ?= https://releases.hashicorp.com/$(TERRAFORM_PROVIDER_DOWNLOAD_NAME)/$(TERRAFORM_PROVIDER_VERSION)
export TERRAFORM_NATIVE_PROVIDER_BINARY ?= terraform-provider-dynatrace_v1.66.0
export TERRAFORM_DOCS_PATH ?= docs/resources

These settings specify the source, version, and download paths for the Dynatrace Terraform provider that Upjet will wrap as a Crossplane provider.

Step 2: Configure Provider Resources

Set Up the Provider Config

To configure the connection details, we need to modify internal/clients/dynatrace.go to reference the secret structure expected for the provider. In this case, define tenantURL and apiToken for Dynatrace connectivity:


const (
   tenantURL = "dt_env_url"
   apiToken  = "dt_api_token"
)

Then, reference these credentials in the TerraformSetupBuilder:


// TerraformSetupBuilder builds a Terraform setup function, returning provider configuration.
func TerraformSetupBuilder(version, providerSource, providerVersion string) terraform.SetupFn {
    return func(ctx context.Context, client client.Client, mg resource.Managed) (terraform.Setup, error) {
        ...
        // Set credentials in the provider configuration.
        ps.Configuration = map[string]any{}
        if v, ok := creds[tenantURL]; ok {
            ps.Configuration[tenantURL] = v
        }
        if v, ok := creds[apiToken]; ok {
            ps.Configuration[apiToken] = v
        }
    }
}

Define External Name Configurations

To identify external names for resources, update config/external_name.go by adding mappings for the Dynatrace resources:


// ExternalNameConfigs contains all external name configurations for this provider.
var ExternalNameConfigs = map[string]config.ExternalName{
    "dynatrace_alerting":           config.IdentifierFromProvider,
    "dynatrace_email_notification": config.IdentifierFromProvider,
    "dynatrace_json_dashboard":     config.IdentifierFromProvider,
    "dynatrace_metric_events":      config.IdentifierFromProvider,
}

This setup ensures that each resource is correctly identified using the provider’s unique identifier.

Add Custom Configurations for Resources

For each resource, create a corresponding config subfolder and add a config.go file with a Configure function. This function customizes the resource’s configuration and short group name as needed:

config/alerting/config.go


package alerting
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_alerting", func(r *config.Resource) {
    r.ShortGroup = "alerting"
  })
}

Repeat this for other resources, such as Dashboard, Event, and Notification.

config/dashboard/config.go


package dashboard
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_json_dashboard", func(r *config.Resource) {
    r.ShortGroup = "dashboard"
  })
}

config/event/config.go


package event
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_metric_events", func(r *config.Resource) {
    r.ShortGroup = "event"
  })
}

config/notification/config.go


package notification
import "github.com/crossplane/upjet/pkg/config"
// Configure customizes individual resources.
func Configure(p *config.Provider) {
  p.AddResourceConfigurator("dynatrace_email_notification", func(r *config.Resource) {
    r.ShortGroup = "notification"
  })
}

Register Custom Configurations

To ensure these custom configurations are applied, register each Configure function in config/provider.go:


import (
    "github.com/xoanmi/provider-dynatrace/config/alerting"
    "github.com/xoanmi/provider-dynatrace/config/event"
    "github.com/xoanmi/provider-dynatrace/config/dashboard"
    "github.com/xoanmi/provider-dynatrace/config/notification"
)
for _, configure := range []func(provider *ujconfig.Provider){
    alerting.Configure,
    event.Configure,
    notification.Configure,
    dashboard.Configure,
} {
    configure(pc)
}

This setup allows Upjet to apply each resource’s configuration during the provider generation process, customizing each resource group as specified.

Step 3: Generate the Code

Once all the necessary configurations are in place, you’re ready to generate the provider code by running the following command:

The make generate command will use the settings specified in the previous steps to:

  • Generate the Crossplane provider code based on the Terraform provider configurations.
  • Create the necessary Kubernetes Custom Resource Definitions (CRDs) for each resource, allowing Crossplane to manage them.

Running make generate will produce output similar to the following, showing the installation of required tools and the generation of the provider schema and resource CRDs:


➜ make generate
11:38:21 [ .. ] installing terraform darwin-arm64
…
11:38:22 [ OK ] installing terraform darwin-arm64
11:38:22 [ .. ] generating provider schema for dynatrace-oss/dynatrace 1.66.0
11:38:24 [ OK ] generating provider schema for dynatrace-oss/dynatrace 1.66.0
11:38:26 [ .. ] go generate linux_arm64

Generated 4 resources!
11:38:59 [ OK ] go generate linux_arm64
11:38:59 [ .. ] go mod tidy
11:39:00 [ OK ] go mod tidy

➜ tree package/crds
package/crds
├── alerting.crossplane.io_alertings.yaml
├── dashboard.crossplane.io_dashboards.yaml
├── dynatrace.crossplane.io_providerconfigs.yaml
├── dynatrace.crossplane.io_providerconfigusages.yaml
├── dynatrace.crossplane.io_storeconfigs.yaml
├── event.crossplane.io_events.yaml
└── notification.crossplane.io_notifications.yaml

With the generated code and CRDs in place, you can now deploy the provider and begin managing Dynatrace resources in your Kubernetes environment.

Step 4: Deploy and run

In the setup, we’re going to run the operator locally while applying and connecting to a Kubernetes cluster (where this cluster runs doesn’t matter, as long as it’s reachable).

First, we apply the CRDs make generate has created.


kubectl apply -f package/crds

Once this is done, you can run the operator itself.


make run

The last missing part is adding the credentials the operator needs to connect to the chosen Dynatrace tenant. The necessary token is in the Access management documentation.

Create the namespace and a secret containing your brand-new access token.


apiVersion: v1
kind: Secret
metadata:
  name: example-creds
  namespace: crossplane-system
type: Opaque
stringData:
  credentials: |
    {
     "dt_env_url": "https://my-tenant.com",
     "dt_api_token": "my-secret-token"
    }

kubectl create namespace crossplane-system
kubectl apply -f example-creds.yaml

Finally, the last puzzle piece, the ProviderConfig, can be created:


apiVersion: dynatrace.crossplane.io/v1beta1
kind: ProviderConfig
metadata:
  name: default
  namespace: crossplane-system
spec:
  credentials:
    source: Secret
    secretRef:
      name: example-creds
      namespace: crossplane-system
      key: credentials

All done! The previously generated CRDs are now available and their objects in your Kubernetes cluster will result in entities and changes in your Dynatrace tenant. Try it out with a dashboard resource!


apiVersion: dashboard.crossplane.io/v1alpha1
kind: Dashboard
metadata:
  name: example-dashboard
  namespace: crossplane-system
spec:
  forProvider:
    contents: |
      {
        "dashboardMetadata": {
          "name": "Our small example dashboard",
          "owner": "my@mail.com",
          "preset": true,
          "hasConsistentColors": true
      },
      "tiles": [
        {
          More config…
        }
     }

kubectl apply -f example-dashboard.yaml

FAQ

Several key questions were raised during the talk and in the following discussions. We’ve summarized the main points:

Q: Why use Crossplane when Terraform can do the same thing?

A: It’s not a matter of “should” versus “shouldn’t.” If you’re already invested in Kubernetes, Crossplane allows you to manage cloud resources while leveraging the same tooling you use to deploy, maintain, and monitor your applications. This makes Crossplane highly convenient for teams already embedded in the Kubernetes ecosystem.

Additionally, these tools don’t exclude each other. One is used to build platforms, and the other is a command-line tool. Their potential use cases differ quite a lot.

Q: How is state management handled?

A: With the Upjet approach, you’re essentially bridging two worlds. Kubernetes manages the state of each object through its etcd system. Simultaneously, the Crossplane operator runs Terraform in the background, continuously reconciling the state between the Kubernetes Custom Resource (CR) and the Terraform-managed infrastructure.

Q: Can I use the Upjet approach in production?

A: Yes, you can, but remember that the provider uses Terraform under the hood. This means that during each reconciliation loop, a terraform plan and terraform apply run. Due to the nature of these continuous operations, managing a large number of resources this way could demand significant resources.

While the Upject project is very good at translating the provider, some things need to be added manually. The concept of Kubernetes labels simply doesn’t exist in Terraform. If you want to utilize them, you need to implement them yourself.

Q: How does the mapping between Terraform objects and Kubernetes Custom Resources (CRs) work?

A: The mapping is defined in the `/config` folder when configuring the provider. Here, we specify the relationship between the Terraform object and the corresponding Kubernetes CR. Running the `make generate` command triggers the generation of all necessary code, including the API, client, provider, and CRD (Custom Resource Definition). This allows Kubernetes to manage the Terraform-defined resource seamlessly.

Q: I read about the proposal for Crossplane v2.0. Do you know if that will impact the described provider creation process?

A: Recently, the Crossplane developers created a draft for the next version of Crossplane. Here, they talk in-depth about how they want to change composite resources and their structure. This will impact provider creation since they reconcile aforementioned resources. This section discusses the proposed changes. The developers also plan on keeping things backward compatible. For now, we
have to wait and see what the final implementation looks like.

Get started

Crossplane enables a seamless cloud-native approach for managing any cloud resource by extending the Kubernetes API. By leveraging Kubernetes as a control plane and using Crossplane compositions, you can declaratively define and automate your entire observability stack.

It’s easy to get started, all you need to start is

If you’re interested in diving deeper, you can check out the following resources from our session:

We hope this talk inspired you to explore Crossplane for your infrastructure automation needs and provided valuable insights into building observability solutions using the power of Kubernetes and declarative infrastructure.

The post Observability as Code: DIY with Crossplane appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/observability-as-code-diy-with-crossplane/feed/ 0
Let’s learn how to send OpenTelemetry data to Dynatrace together! https://www.dynatrace.com/news/blog/send-opentelemetry-data-to-dynatrace/ https://www.dynatrace.com/news/blog/send-opentelemetry-data-to-dynatrace/#respond Tue, 21 Jan 2025 20:38:50 +0000 https://www.dynatrace.com/news/?p=67385 OpenTelemetry trends

This blog post will help new and existing customers get started with Dynatrace support for OpenTelemetry. Learn how to send OpenTelemetry data to Dynatrace from an OTel veteran and Dynatrace newbie.

The post Let’s learn how to send OpenTelemetry data to Dynatrace together! appeared first on Dynatrace news.

]]>
OpenTelemetry trends

One of the things I love most about OpenTelemetry (OTel) is that it’s vendor-neutral, which means you can send the same OpenTelemetry data to different vendors. In fact, most of the major Observability vendors out there not only support ingesting OpenTelemetry data but also actively contribute to the project, including Dynatrace. Check out the 2023 OpenTelemetry Journey Report for more info.

Why does this matter? I used to work at another Observability vendor, and many of the OpenTelemetry examples that I played with and blogged about in the last 2 years or so featured sending OTel data to that backend.

Now that I work at Dynatrace, which, by the way, ingests OTLP natively, I wanted to educate myself on how to send OpenTelemetry data to Dynatrace. A great way to learn is to try to run my go-to examples using Dynatrace as the Observability backend. Luckily for me, since OTel is vendor-neutral, all I had to do was reconfigure my OTel Collector to point to Dynatrace to get my examples to work.

Want to learn how? Let’s do it together!

Note: If you’re evaluating multiple vendors, you can send the same data to different vendors at the same time (à la “vendor bake-off”) to help you determine which vendor best suits your organization’s needs.

Prerequisites for sending data to Dynatrace

To send OpenTelemetry data to Dynatrace, you need two pieces of information:

Dynatrace tenant: Each user (or, more likely, organization) is assigned a tenant. When sending OpenTelemetry data from your application to Dynatrace, you need to specify the Dynatrace OTLP endpoint (used by Dynatrace to receive data), which includes your tenant name.

Access token: The access token allows you to send OTel data to your Dynatrace instance. It also specifies what kind of data you’re allowed to send to Dynatrace. You can find more on Dynatrace access tokens here.

Before we get to any of that, you first need a Dynatrace account. If you already have a Dynatrace account, feel free to skip the following section.

Create a Dynatrace account

If you don’t have a Dynatrace account, you can create a free trial account, which is valid for 15 days.

  1. Go here, and select the Free trial button.

    Free trial signup button at dynatrace.com
    Free trial signup button at dynatrace.com
  2. Enter your email, select the Terms of Use checkbox, and then select Continue.
    Enter your email and accept the Terms of Use
    Enter your email and accept the Terms of Use
  3. Enter the rest of the info and select Start free trial.

    Fill in the rest of the form fields
    Fill in the rest of the form fields

    You will receive an email once your account has been created. You will also see a page that looks like the one below. Select Launch Dynatrace to get started.

    Your Dynatrace tenant is ready!
    Your Dynatrace tenant is ready!

    This takes you to the Dynatrace login page.

    Dynatrace login page
    Dynatrace login page

Your Dynatrace tenant

To find your Dynatrace tenant, log into Dynatrace here, and select the Login button at the top right of the page.

This takes you to the sign-in page. Once you sign in, take note of the URL. It should look something like this:

https://<your-dynatrace-tenant>.apps.dynatrace.com

Take note of the value of <your-dynatrace-tenant>, because we’ll need that later.

Create a Dynatrace access token

After confirming that you’re logged into Dynatrace, press ctrl+k, and then type access token. Next, select Access Tokens from the top of the search results.

Access token search
Access token search

On the Access tokens page, select Generate new token.

Dynatrace Access tokens page
Dynatrace Access tokens page

On the Generate new token page, enter:

  • Token name: be sure to give it a descriptive name
  • Expiration date: this is optional
  • Template: Kubernetes Data Ingest

Even if we’re not necessarily using Kubernetes, the Kubernetes Data Ingest template has the token scopes (permissions) that we need to send OpenTelemetry data to Dynatrace, namely:

  • Ingest logs (ingest)
  • Ingest metrics (ingest)
  • Ingest OpenTelemetry traces (ingest)

Find more information on these and other Dynatrace token scopes here.

Once you’re done, select Generate token at the bottom of the page.

Access token generation page
Access token generation page

The next page shows your access token. Be sure to copy and store it somewhere for safekeeping (for example, a secrets manager such as HashiCorp Vault or your cloud provider’s secrets manager) before selecting Done, because after that, it’s gone forever. If you lose that token information, you should delete the old one (not necessary, but highly recommended), and create a new one.

Access token page showing generated token
Access token page showing generated token

Configure the OTel Collector for Dynatrace

Now that you have your tenant info and access token, you can plug this information into your OpenTelemetry Collector configuration.

Note: There are two ways to send OTel data to an Observability backend: (1) direct from the application, or (2) via the OTel Collector. There’s a time and place for each, and you can check out my blog post on OTel Collector Anti-patterns on the OTel Blog to learn more.

Your OTel Collector config YAML file should look something like this:

receivers:
  otlp:
    protocols:
      grpc:
      http:

processors:
  cumulativetodelta:
  batch:

exporters:
  otlphttp:
    endpoint: "https://${DT_TENANT}.live.dynatrace.com/api/v2/otlp"
    headers:
      Authorization: "Api-Token ${DT_API_TOKEN}"
  debug:
    verbosity: detailed

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlphttp,debug]
    metrics:
      receivers: [otlp]
      processors: [cumulativetodelta,batch]
      exporters: [otlphttp,debug]
    logs:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlphttp,debug]

Dynatrace accepts data in the native OpenTelemetry Protocol (OTLP) format via HTTP (gRPC is not yet supported). You need to specify the Dynatrace OTLP endpoint (used by Dynatrace to receive data), which includes your tenant name, ${DT_TENANT}:

https://${DT_TENANT}.live.dynatrace.com/api/v2/otlp

${DT_TENANT} is the value of your Dynatrace tenant name, which you hopefully jotted down and stored in a secrets manager for safekeeping.

Finally, “Dynatrace requires metrics data to be sent with delta temporality, not cumulative temporality”. This means that you’ll need to include the cumulativetodelta processor in:

  • Your Collector configuration (cumulativetodelta)
  • Your metrics pipeline (pipelines.metrics)

Never store your Dynatrace token and tenant name in plain text. Instead, store them in a secrets manager and pull them from the secrets manager at runtime.

Dynatrace and the OTel Operator

If you’re using the OpenTelemetry Operator to send OpenTelemetry data to Dynatrace, you’ll need to configure your OpenTelemetryCollector resource as follows:

apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: otelcol
  namespace: opentelemetry
spec:
  mode: statefulset
  image: ghcr.io/dynatrace/dynatrace-otel-collector/dynatrace-otel-collector:0.7.0
  env:
    - name: DT_API_TOKEN
      valueFrom:
        secretKeyRef:
          key: DT_API_TOKEN
          name: otel-collector-secret
    - name: DT_TENANT
      valueFrom:
        secretKeyRef:
          key: DT_TENANT
          name: otel-collector-secret
  config:
    receivers:
      otlp:
        protocols:
          grpc: {}
          http: {}

    processors:
      cumulativetodelta: {}
      batch: {}

    exporters:
      otlphttp:
        endpoint: "https://${DT_TENANT}.live.dynatrace.com/api/v2/otlp"
        headers:
          Authorization: "Api-Token ${DT_API_TOKEN}"
      debug:
        verbosity: detailed

    service:
      pipelines:
        traces:
          receivers: [otlp]
          processors: [batch]
          exporters: [otlphttp,debug]
        metrics:
          receivers: [otlp]
          processors: [cumulativetodelta,batch]
          exporters: [otlphttp,debug]
        logs:
          receivers: [otlp]
          processors: [batch]
          exporters: [otlphttp,debug]

Notice that the spec.config looks the same as what we defined in the otelcol-config.yaml file we saw earlier. The only added thing here is that DT_API_TOKEN and DT_TENANT are environment variables pulled from a Kubernetes Secret. The secret YAML definition looks like this:

apiVersion: v1
kind: Secret
metadata:
 name: otel-collector-secret
 namespace: 
data:
 DT_API_TOKEN: 
 DT_TENANT: 
type: "Opaque"

Both DT_API_TOKEN and DT_TENANT values must be base64 encoded before being added to the secrets YAML. To base64 encode a value and copy the encoded value to your buffer, use this action for easy copy/paste:

echo <value_to_encode> | base64

Remember that storing secrets in Kubernetes (or storing a secrets YAML in version control, for that matter), is not recommended because base64 does not encrypt your data. You should instead consider using sealed secrets or the Kubernetes Secrets Store CSI Driver + your favorite secrets provider. For more info on these better alternatives check out this article.

The Dynatrace OTel Collector Distribution

Many vendors have their own OTel Collector Distributions. These distributions are curated with Collector components that are specific to that vendor. They can be a combination of vendor-developed custom components and components from Collector Core and Contrib. Using vendor-specific distributions ensures you’re using just the Collector components you need, reducing overall bloat. You can learn more here.

Dynatrace also has its own Collector distribution. It features a set of Collector components for sending Observability data to Dynatrace from various sources. It stays up-to-date with upstream components of the opentelemetry-collector and opentelemetry-collector-contrib repositories.

In addition, the Dynatrace Collector distribution offers the following advantages:

  • It is covered by Dynatrace support
  • Collector components are verified by Dynatrace
  • Security patches are independent of OpenTelemetry Collector releases

Try it out!

Want to try sending data to Dynatrace yourself? Then check out my example repo. I created this repo for a talk on the Target Allocator that I gave at KubeCon in March 2024. It has been updated to include instructions on configuring the OpenTelemetryCollector resource to send OTel data to Dynatrace.

OTel Data in Dynatrace

And if you’re curious to see what OTel data looks like in Dynatrace, here are some screenshots of the web UI.

I’m not going in-depth on how to navigate the Dynatrace UI because there are already some great videos on the Dynatrace YouTube channel. I encourage you to check them out for a more in-depth look.

Dynatrace Distributed Tracing UI
Dynatrace Distributed Tracing UI
Dynatrace Logs UI
Dynatrace Logs UI
Dynatrace Notebooks UI showing a metric called ”some_counter_total”
Dynatrace Notebooks UI showing a metric called ”some_counter_total”

Final thoughts

As someone with experience sending OpenTelemetry data to various backends, I found that getting OpenTelemetry data into Dynatrace was fairly straightforward. My only personal hiccup was in generating the application token, but I got that sorted out, and now I’ve passed on my knowledge and highly detailed screenshots along to you.

I have to say that it’s always fun to use a product with fresh eyes, a fresh perspective, and a newbie point of view. There’s nothing quite like it. And, having worked at another observability vendor before, it’s always fun to see the similarities and differences. It’s like learning a new programming language and comparing it to another one that you already know. What a blast!

One final point that I want to make. I don’t want to trivialize things and give you the impression that moving from one observability vendor to another is simply a matter of repointing your OTel Collector from one vendor backend to another. That is only one aspect of a vendor migration, no matter what vendor you’re moving to/from. You also must consider the fact that you’ll likely have a bunch of dashboards, alerts, and whatnot that you created with your original vendor. When you migrate to another vendor, there won’t be a 1:1 translation; so keep that in mind.

But that may be a sacrifice that you’re willing to make because OTel’s vendor neutrality means that all vendors supporting OpenTelemetry are ingesting the same data. What sets them apart is what they do with your data. And if one vendor does something with your data better than another, well, don’t you owe it to yourself to check that out?

What’s Next?

If you’re interested in learning first-hand what Dynatrace and OpenTelemetry can do together, then dive into Dynatrace! You can do so in one of two ways:

  • Check out the Dynatrace Playground to explore Dynatrace using pre-populated OpenTelemetry data
  • Get started with a free trial and ingest your own OpenTelemetry data today!

The post Let’s learn how to send OpenTelemetry data to Dynatrace together! appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/send-opentelemetry-data-to-dynatrace/feed/ 0
When things go sideways: Troubleshooting the OpenTelemetry Operator https://www.dynatrace.com/news/blog/troubleshooting-the-opentelemetry-operator/ https://www.dynatrace.com/news/blog/troubleshooting-the-opentelemetry-operator/#respond Fri, 13 Dec 2024 16:41:22 +0000 https://www.dynatrace.com/news/?p=67050 Kubernetes and OpenTelemetry

Learn the basics of the OpenTelemetry (OTel) Operator and how to troubleshoot when things don’t go according to plan.

The post When things go sideways: Troubleshooting the OpenTelemetry Operator appeared first on Dynatrace news.

]]>
Kubernetes and OpenTelemetry

This blog post was co-written with Reese Lee.

If you already have an application running in Kubernetes and are exploring using OpenTelemetry to gain insights into the health and performance of your app and cluster, you might be interested in an implementation of the Kubernetes Operator called the OpenTelemetry Operator.

As you’ll learn shortly, due to its range of capabilities, the Operator is your go-to for (almost) hassle-free OpenTelemetry management. But, as with any powerful tool, what happens when things go sideways?

In this blog post, you’ll learn about the OpenTelemetry Operator (hereafter referred to as “the Operator”), along with issues commonly encountered across installation, Collector deployment, and auto-instrumentation. You’ll also learn how to resolve these issues, and be better prepared for running the Operator.

Overview of the Operator

Let’s take a closer look at the Operator’s main capabilities.

Managing the Collector

The Operator automates the deployment of your Collector, and makes sure it’s correctly configured and running smoothly within your cluster. The Operator also manages configurations across a fleet of Collectors using Open Agent Management Protocol (OpAMP), which is a network protocol for remotely managing large fleets of data collection agents. Since the protocol is vendor-agnostic, this helps ensure consistent observability settings and simplifies management across agents from different vendors.

Managing Auto-Instrumentation in Pods

The Operator automatically injects and configures auto-instrumentation for your applications, which enables you to collect telemetry data without modifying your source code. If your application isn’t already instrumented with OpenTelemetry, this is a fantastic option to feed two birds with one scone, and start generating and collecting application telemetry.

Installing the Operator

This might seem obvious, but before installing the Operator, you must have a Kubernetes cluster you can install it into, running Kubernetes 1.23+. Check the compatibility matrix for specific version requirements. You can spin up a cluster on your machine using a local Kubernetes tool such as minikube, k0s, or KinD, or use a cluster running on a cloud provider service.

Next, and this is less obvious: You must have a component called cert-manager already installed in that cluster. This piece manages certificates for Kubernetes by making sure the certificates are valid and up to date. You can install both the cert-manager and the Operator via kubectl or a Helm chart.

Note that in either case, you have to wait for cert-manager to finish installing before you install the Operator; otherwise, the operator installation will fail.

Using kubectl

To install cert-manager, run the following command:

kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.10.0/cert-manager.yaml

Next, install the Operator:

kubectl apply -f https://github.com/open-telemetry/opentelemetry-operator/releases/latest/download/opentelemetry-operator.yaml

Using Helm

To install cert-manager, first add the Helm repository:

helm repo add jetstack https://charts.jetstack.io --force-update

Next, install the cert-manager Helm chart:

helm install \
cert-manager jetstack/cert-manager \
--namespace cert-manager \
--create-namespace \
--version v1.16.1 \
--set crds.enabled=true

Expect the preceding step to take up to a few minutes. You can verify your installation of cert-manager by following the steps in this link, or check the deployment status by running:

kubectl get pods -namespace cert-manager

To install the Operator, note that Helm 3.9+ is required. First, add the repo:

helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts 
helm repo update

Then, install the Operator:

helm install --namespace opentelemetry-operator-system \  
  --create namespace \  
  opentelemetry-operatoropen-telemetry/opentelemetry-operator 

Deploying the OpenTelemetry Collector

Once you have cert-manager and Operator set up in your cluster, you can deploy the Collector. The Collector is a versatile component that’s able to ingest telemetry from a variety of sources, transform the received telemetry in a number of ways based on its configuration, and then export that processed data to any backend that accepts the OpenTelemetry data format (also referred to as OTLP, which stands for OpenTelemetry Protocol).

The Collector can be deployed in several different ways, referred to as “patterns.” Which pattern or patterns you deploy is dependent on your telemetry needs and organizational resources. This topic is out of scope for this blog post, but you can read more about them via this link.

Collector Custom Resource

A custom resource (CR) represents a customization of a specific Kubernetes installation that isn’t necessarily available in a default Kubernetes installation; CRs help make Kubernetes more modular.

The Operator has a CR for managing the deployment of the Collector, called OpenTelemetryCollector. The following is a sample OpenTelemetryCollector resource:

apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: otelcol
  namespace: opentelemetry
spec:
  mode: statefulset
  config:
    receivers:
      otlp:
        protocols:
          grpc: {}
          http: {}
      prometheus:
        config:
          scrape_configs:
            - job_name: 'otel-collector'
              scrape_interval: 10s
              static_configs:
              - targets: [ '0.0.0.0:8888' ]

    processors:
      batch: {}

    exporters:
      logging:
        verbosity: detailed

    service:
      pipelines:
        traces:
          receivers: [otlp]
          processors: [batch]
          exporters: [logging]
        metrics:
          receivers: [otlp, prometheus]
          processors: []
          exporters: [logging]
        logs:
          receivers: [otlp]
          processors: [batch]
          exporters: [logging]

There are many configuration options for the OpenTelemetryCollector resource, depending on how you plan on instantiating it; however, the basic configuration requires:

  • mode, which should be one of the following: deployment, sidecar, daemonset, or statefulset. If you leave out mode, it defaults to deployment.
  • config, which may look familiar, because it’s the Collector’s YAML config.

Common Collector deployment issues and troubleshooting tips

If you’re not seeing the data you expect, or you suspect something isn’t working right, try the following troubleshooting tips.

Check that the Collector resources deployed properly

When an OpenTelemetryCollector YAML is deployed, the following objects are created in Kubernetes:

1. OpenTelemetryCollector

2. Collector pod:

  • If you specified non-sidecar mode, look for Deployment, StatefulSet, or DaemonSet resources named <collector_CR_name>-collector-<unique_identifier>).
  • If you specified the mode as sidecar, a Collector sidecar container will be created in an app pod, named otc-container.

3. Target Allocator pod:

  • If you enabled the Target Allocator, look for a resource named <collector_CR_name>-targetallocator-<unique_identifier>.

4. ConfigMap of Collector configurations:

  • If you specified non-sidecar mode, look for Deployment, StatefulSet, or DaemonSet resources named <collector_CR_name>-collector-<unique_identifier>.
  • If you specified the mode as sidecar, note that the Collector config is included as an environment variable.

Thus, when you deploy the OpenTelemetryCollector resource, make sure that the preceding objects are created.

First, confirm that the OpenTelemetryCollector resource was deployed:

kubectl get otelcol -n <namespace>

When you deploy the Collector using the OpenTelemetryCollector resource, it creates a ConfigMap containing the Collector’s configuration YAML. Confirm that the ConfigMap was created in the same namespace as the Collector, and that the configurations themselves are correct.

List your ConfigMaps:

kubectl get configmap -n <namespace> | grep <collector-cr-name>-collector

We also recommend checking your Collector pods by running the appropriate command based on the Collector’s mode:

  • deployment, statefulset, daemonset modes:
kubectl get pods -n <namespace> | grep <collector_cr_name>-collector
  • sidecar mode:
kubectl get pods <pod_name> -n opentelemetry -o jsonpath='{.spec.containers[*].name}'

This will list all the containers created in the pod, including the Collector sidecar container, which includes the Collector config as an environment variable.

Check the Collector CR version

Take a look at the OpenTelemetryCollector CR version you’re using. There are two versions available: v1alpha1:

apiVersion: opentelemetry.io/v1alpha1 
kind: OpenTelemetryCollector 
metadata: 
  name: otelcol 
  namespace: opentelemetry 
spec: 
  mode: statefulset 
  config: | 
    receivers: 
      otlp: 
        protocols: 
          grpc: 
          http: 
 
    processors: 
      batch: 
 
    exporters: 
      otlp: 
        endpoint: "<my_o11y_backend>" 
      logging: 
        verbosity: detailed 
 
    service: 
      pipelines: 
        traces: 
          receivers: [otlp] 
          processors: [batch] 
          exporters: [otlp/ls, logging] 
        metrics: 
          receivers: [otlp, prometheus] 
          processors: 
          exporters: [otlp/ls, logging] 
        logs: 
          receivers: [otlp] 
          processors: [batch] 
          exporters: [otlp/ls, logging] 

and v1beta1:

apiVersion: opentelemetry.io/v1beta1 
kind: OpenTelemetryCollector 
metadata: 
  name: otelcol 
  namespace: opentelemetry 
spec: 
  mode: statefulset 
  config: 
    receivers: 
      otlp: 
        protocols: 
          grpc: {} 
          http: {} 
 
    processors: 
      batch: {} 
 
    exporters: 
      otlp: 
        endpoint: "<my_o11y_backend>" 
      logging: 
        verbosity: detailed 
 
    service: 
      pipelines: 
        traces: 
          receivers: [otlp] 
          processors: [batch] 
          exporters: [otlp/ls, logging] 
        metrics: 
          receivers: [otlp, prometheus] 
          processors: [] 
          exporters: [otlp/ls, logging] 
        logs: 
          receivers: [otlp] 
          processors: [batch] 
          exporters: [otlp/ls, logging] 

There are two main differences between these two API versions:

1. The config sections are different; for v1beta1, the config values are key-value pairs that are part of the CR configuration, whereas for v1alpha1, the config value is one long text string. Keep in mind that the text string still needs to follow YAML formatting.

2. If you’re using v1beta1, you can’t leave the Collector config values empty. You must specify either empty curly braces ({}) for scalar values or empty brackets ([ ]) for arrays. This isn’t necessary if you’re using v1alpha1.

Check the Collector base image

By default, the OpenTelemetryCollector CR uses the core distribution of the Collector. The core distribution is a bare-bones distribution of the Collector for OpenTelemetry developers to develop and test. It contains a base set of components: extensions, connectors, receivers, processors, and exporters.

If you want access to more components than the ones offered by core, you can use the Collector’s Kubernetes distribution instead. This distribution is made specifically to be used in a Kubernetes cluster to monitor Kubernetes and services running in Kubernetes. It contains a subset of components from the core and contrib distributions. Alternatively, you can build your own Collector distribution.

You can set the Collector’s base image by specifying the image attribute in spec.image, as in the following example:

apiVersion: opentelemetry.io/v1beta1 
kind: OpenTelemetryCollector  
metadata: 
  name: otelcol 
  namespace: opentelemetry 
spec: 
  mode: statefulset  
  image: otel/opentelemetry-collector/contrib:0.102.1 
config:  
  receivers:  
    otlp: 
      protocols:  
      grpc: {} 
       http: {} 
  processors:  
    batch: {} 
  exporters:  
    otlp: 
      endpoint: "<olly_backend_endpoint>"

Check your backend vendor’s access requirements

If you’re using a backend vendor to ingest your telemetry data, you’ll likely need to configure an account license key or some kind of access token, which you’ll want to keep confidential.

To store it as a secret and prevent it from appearing as plain text, first create a Kubernetes secret, and Base64-encode it:

apiVersion: v1  
kind: Secret  
metadata: 
  name: otel-collector-secret  
  namespace: opentelemetry 
data: 
  ACCESS_TOKEN: <base64_encoded_token> 
type: "Opaque"

Check your exporter configuration

Confirm that you’ve configured the correct endpoint according to your region in your exporter configuration.

When all else fails…check Kubernetes events

Kubernetes events provide detailed and chronological information about what’s happening within various components of your cluster. To view events for a specific namespace, use:

kubectl get events -n <namespace>

Replace <namespace> with the actual namespace where your OpenTelemetry Operator and resources are deployed.

Instrumentation

Instrumentation is the process of adding code to software to generate telemetry signals–logs, metrics, and traces. You have several options for instrumenting your code with OpenTelemetry, the primary two being code-based and zero-code solutions.

Code-based solutions require you to manually instrument your code using the OpenTelemetry API. While it can take time and effort to implement, this option enables you to gain deep insights and further enhance your telemetry, as you have a high degree of control over what parts of your code are instrumented and how.

To instrument your code without modifying it (or if you’re unable to modify the source code), you can use zero-code solutions (or auto-instrumentation agents). This method uses shims or bytecode agents to intercept your code at runtime or at compile-time to add tracing and metrics instrumentation to the third-party libraries and frameworks you depend on. At the time of publication, auto-instrumentation is currently available for Java, Python, .NET, JavaScript, PHP, and Go. Learn more about zero-code instrumentation at this link.

You can also use both options simultaneously. Some end users opt to start with a zero-code agent and manually insert additional instrumentation, such as adding custom attributes or creating new spans. Alternatively, OpenTelemetry also provides options beyond code-based and zero-code solutions. Learn more at this link.

Zero-code Instrumentation with the Operator

The Operator has a CR called Instrumentation that can automatically inject and configure OpenTelemetry instrumentation into your Kubernetes pods, providing the benefit of zero-code instrumentation for your application. This is currently available for the following: Apache HTTPD, .NET, Go, Java, nginx, Node.js, and Python.

The following is a sample Instrumentation resource definition for a Python service:

apiVersion: opentelemetry.io/v1alpha1  
kind: Instrumentation  
metadata: 
  name: python-instrumentation  
  namespace: application 
spec: 
  env: 
    - name: OTEL_EXPORTER_OTLP_TIMEOUT 
      value: "20" 
    - name: OTEL_TRACES_SAMPLER 
      value: parentbased_traceidratio 
    - name: OTEL_TRACES_SAMPLER_ARG 
      value: "0.85" 
  exporter: 
    endpoint: http://localhost:4317 
  propagators: 
    - tracecontext 
    - baggage  
  sampler: 
    type: parentbased_traceidratio  
    value: "0.25" 
  python:  
    env: 
      - name: OTEL_METRICS_EXPORTER 
        value: otlp_proto_http 
      - name: OTEL_LOGS_EXPORTER 
        value: otlp_proto_http 
      - name: OTEL_PYTHON_LOGGING_AUTO_INSTRUMENTATION_ENABLED 
        value: "true" 
      - name: OTEL_EXPORTER_OTLP_ENDPOINT 
        value: http://localhost: 4318 

You can use a single auto-instrumentation YAML to serve multiple services written in different languages (provided they are supported for auto-instrumentation). List your global environment variables under spec.env, and list language-specific environment variables under spec.<language_name>.env. You can mix and match language-specific environment variable configurations in the same Instrumentation resource.

In order to use the Operator’s auto-instrumentation capability, deploying an Instrumentation resource alone isn’t enough. The auto-instrumentation configuration must be associated with the code being instrumented. This is done by adding an auto-instrumentation annotation in your application’s Deployment YAML, in the template definition section, such as in the following example:

apiVersion: apps/v1 
kind: Deployment  
metadata: 
  name: my-deployment-with-sidecar  
spec: 
  replicas: 1  
  selector: 
    matchLabels: 
      app: my-pod-with-sidecar 
  template: 
    metadata:  
      labels: 
        app: my-pod-with-sidecar 
      annotations: 
        sidecar.opentelemetry.io/inject: "true" 
        instrumentation.opentelemetry.io/inject-python: "true" 
spec: 
  containers: 
    - name: py-otel-server 
      image: otel-python-lab:0.1.0-py-otel-server ports: 
    - containerPort: 8082 
      name: py-server-port 

When the annotation called instrumentation.opentelemetry.io/inject-python is set to true, it tells the Operator to inject Python auto-instrumentation (in this case) into the containers running in this pod. For other languages, simply replace python with the appropriate language name (for example, instrumentation.opentelemetry.io/inject-javafor Java apps). You can disable instrumentation by setting this value to false.

If you have multiple Instrumentation resources, you need to specify which one to use, otherwise the Operator won’t know which one to pick. You can do this as follows:

  • By name. Use this if the Instrumentation resource resides in the same namespaces as the Deployment. For example, opentelemetry.io/inject-java: my-instrumentation will look for an Instrumentation resource called my-instrumentation.
  • By namespace and name. Use this if the Instrumentation resource resides in a different namespace. For example: opentelemetry.io/inject-java: my-namespace/my-instrumentation will look for an Instrumentation resource called my-instrumentation in the namespace my-namespace.

You must deploy the Instrumentation resource before the annotated application; otherwise, your code won’t be automatically instrumented. The Operator injects auto-instrumentation by adding an init container to the application’s pod when it starts up, which means that if the Instrumentation resource isn’t available by the time your service is deployed, the auto-instrumentation will fail.

Common instrumentation issues and troubleshooting tips

If your Collector doesn’t seem to be processing data or if you think the auto-instrumentation isn’t working, try the following steps to troubleshoot and resolve the problem.

Check that the instrumentation resource deployed properly

Run the following command to make sure the Instrumentation resource(s) was created in your Kubernetes cluster:

kubectl describe otelinst -n <namespace>

Confirm the resource deployment order

Double check that your Instrumentation CR is deployed before your Deployment. As we learned earlier, if you’re auto-instrumenting via the Operator, you must deploy the Instrumentation resource before deploying your service’s Deployment resource, because the Deployment will create an init-container for the auto-instrumentation. You should therefore see an auto-instrumentation init-container when you run the following command:

kubectl get pod  -n  \ 
  -o jsonpath='{.spec.initContainers[*].name}' 

Check your auto-instrumentation CR annotations

1- Confirm that there are no typos in the annotations.

2- Confirm that they are in the pod’s metadata definition (spec.template.metadata.annotation), not the deployment’s metadata definition (metadata.annotation), as in the following example:

apiVersion: apps/v1 
kind: Deployment 
metadata: 
  name: py-otel-server 
  namespace: opentelemetry 
  labels: 
    app: my-app 
    app.kubernetes.io/name: py-otel-server 
spec: 
  replicas: 1 
  selector: 
    matchLabels: 
      app: my-app 
      app.kubernetes.io/name: py-otel-server 
  template: 
    metadata: 
      labels: 
        app: my-app 
        app.kubernetes.io/name: py-otel-server 
      annotations: 
        instrumentation.opentelemetry.io/inject-python: "true" 
    spec: 
      containers: 
      - name: py-otel-server 
        image: otel-target-allocator-talk:0.1.0-py-otel-server 
        imagePullPolicy: IfNotPresent 
        ports: 
        - containerPort: 8082 
          name: py-server-port 
        env: 
          - name: OTEL_RESOURCE_ATTRIBUTES 
            value: service.name=py-otel-server,service.version=0.1.0 

Check your endpoint configurations

The endpoint, configured in the following example under spec.exporter.endpoint, refers to the destination for your telemetry within your Kubernetes cluster:

apiVersion: opentelemetry.io/v1alpha1 
kind: Instrumentation 
metadata: 
  name: python-instrumentation 
  namespace: opentelemetry 
spec: 
  exporter: 
    endpoint: http://otelcol-collector.opentelemetry.svc.cluster.local:4318 
  env: 
  propagators: 
    - tracecontext 
    - baggage 
  python: 
    env: 
      - name: OTEL_METRICS_EXPORTER 
        value: console,otlp_proto_http 
      - name: OTEL_LOGS_EXPORTER 
        value: otlp_proto_http 
      - name: OTEL_PYTHON_LOGGING_AUTO_INSTRUMENTATION_ENABLED 
        value: "true" 

The spec.exporter.endpoint configuration in the Instrumentation resource allows you to define the destination for your telemetry data. If you omit it, it defaults to http://localhost:4317.

If you’re sending out your telemetry to a Collector, the value of spec.exporter.endpoint must reference the name of your Collector Service.

Looking at the example above, otel-collector is the name of the OTel Collector Kubernetes Service.

In addition, if the Collector is running in a different namespace, you must append opentelemetry.svc.cluster.localto the Collector’s service name, where opentelemetry is the namespace in which my Collector happens to be deployed to. It can be any namespace of your choosing.

Finally, make sure that you are using the right Collector port. Normally, you can choose either 4317 (gRPC) or 4318(HTTP); however, for Python auto-instrumentation, you can only use 4318. Confirm whether there are similar caveats for the language(s) you’re using.

Note: If you’re deploying your Collector as a Sidecar, your endpoint needs to be  http://localhost:4317 or http://localhost:4318 (remember: it has to be 4318 for Python).

When all else fails…check the Operator logs

Run the following command to check the Operator logs for any occurrences of error in the log messages:

kubectl logs -l app.kubernetes.io/name=opentelemetry-operator \ 
  --container manager \ 
  -n opentelemetry-operator-system --follow 

Note that the above only applies if you have admin access to your Kubernetes cluster. If you don’t, you can still tell what’s going on by checking your Kubernetes event log, just like we did when troubleshooting issues with the OpenTelemetryCollector resource:

kubectl get events -n <namespace>

Summary

The OpenTelemetry Operator manages the deployment and configuration of one or more Collectors, and injects and configures zero-code instrumentation solutions into your Kubernetes pods. This enables you to get started with OpenTelemetry instrumentation, and you can further enhance your telemetry by adding manual instrumentation to your application.

In this blog post, you learned the ins and outs of the Operator, from common installation hurdles to resolving auto-instrumentation and Collector deployment issues. With detailed installation steps and troubleshooting tips, you’re now equipped to leverage the Operator effectively for the deployment, configuration, and management of your Collectors and auto-instrumentation of supported libraries.

This blog post is based on a talk that Adriana and Reese did at KubeCon North America’s 2024 co-located event, Observability Day. You can check out the recording of the talk here:

When Things Go Sideways: Troubleshooting the OTel Operator – Adriana Villela & Reese Lee

The post When things go sideways: Troubleshooting the OpenTelemetry Operator appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/troubleshooting-the-opentelemetry-operator/feed/ 0
Cloud-native observability made seamless with OpenPipeline and AI-driven observability https://www.dynatrace.com/news/blog/simplify-cloud-native-environments-ai-driven-observability/ https://www.dynatrace.com/news/blog/simplify-cloud-native-environments-ai-driven-observability/#respond Thu, 03 Oct 2024 14:08:48 +0000 https://www.dynatrace.com/news/?p=65815 cloud-native observability made seamless with OpenPipeline

Dynatrace helps cloud-native teams monitor, manage, and troubleshoot complex systems built on Kubernetes® and cloud technologies. With new bespoke apps for Ops and SRE teams, Dynatrace provides familiar experiences with capabilities beyond those offered by popular open-source solutions. Immediate observability insights into workloads and platforms without configuration are complemented with automatic security assessments and AI analytics for predictive operations.

The post Cloud-native observability made seamless with OpenPipeline and AI-driven observability appeared first on Dynatrace news.

]]>
cloud-native observability made seamless with OpenPipeline

The complexity of modern cloud-native environments is ever-increasing. The latest State of Observability 2024 report shows that 86% of interviewed technology leaders see an explosion of data beyond humans’ ability to manage it. On average, organizations have twelve different platforms and services embedded in their multicloud environments.

To avoid drowning in data, it’s critical to ensure that organizations can seamlessly collect data from any source in a single place and in context. Visualizing data in context while supporting and automating decisions with causal, predictive, and generative AI—all while providing a seamless experience—is where the future of cloud-native observability lies.

With new and updated experiences for Ops and SRE teams, proactively managing your cloud environments has never been easier.

Cloud-native observability for monitoring your cloud

OpenPipeline™ is the Dynatrace platform data-handling solution designed to seamlessly ingest and process data from any source, regardless of scale or format. With OpenPipeline, you can effortlessly collect data from Dynatrace OneAgent®, open-source collectors such as OpenTelemetry, or other third-party tools. OpenPipeline then filters and preprocesses the data.

OpenPipeline also incorporates data contextualization technology, enriching data with metadata and linking it to other relevant data sources. By contextualizing data, OpenPipeline enhances the Dynatrace platform’s ability to offer AI-driven insights, analytics, and automation across observability, security, software lifecycle, and business domains.

Furthermore, OpenPipeline is designed to collect and process data securely and in compliance with industry standards. It features high-performance filtering, masking, routing, and encryption capabilities that are easy to configure and operate.

In the latest enhancements of Dynatrace Log Management and Analytics, Dynatrace extends coverage for

  • Native Syslog support: Use Dynatrace ActiveGate to automatically add context and optimize network traffic to your Syslog messages.
  • Seamless integration with AWS Data Firehose: address high-impact issues quickly through real-time, high-frequency log analytics. Dynatrace support for AWS Data Firehose includes AWS Lambda logs, Amazon Virtual Private Cloud (VPC) flow logs, Amazon S3 logs, and Amazon CloudWatch.
  • Kubernetes log monitoring with Fluent Bit

In an effort to further democratize data, Dynatrace provides a curated and supported OpenTelemetry collector. The Dynatrace Otel Collector includes collector components verified by Dynatrace for seamless operation. This removes the burden of manually validating each component and use case. The list of use cases is actively extended and includes batching and ingesting data from Syslog, Fluentd®, Jaeger™, Prometheus®, StatsD, and more.

Dynatrace OTel Collector
Figure 1. Dynatrace OTel Collector

Understand and secure your cloud

It’s critical to have easy and intuitive access to contextually relevant answers when working within complex cloud-native environments and the gold mine of information they provide. Dynatrace is essential for unlocking that gold mine of data, allowing you to enhance application performance, deliver better experiences, and optimize operational efficiency.

Kubernetes

The Dynatrace Kubernetes experience for Site Reliability Engineers (SREs) and Platform Engineers focuses on providing insights into the health and performance of multicloud Kubernetes environments in the tailor-made Kubernetes app. Powered by Davis® AI, the Kubernetes app offers proactive monitoring and analysis, allowing automated monitoring and optimization of health and performance and providing simple and easy troubleshooting. In addition, ready-made dashboards are available for a quick and easy overview, allowing you to see the Kubernetes data you want alongside the cloud-native observability data you need, all in one place.

Kubernetes Cluster Dashboard
Figure 2. Kubernetes Cluster Dashboard

Clouds

Quickly onboard and manage cloud monitoring in one place across different cloud providers and observe multiple cloud environments, including their instances, resources, cost analysis, health, and optimization. Ingest data remotely through cloud integrations covering Amazon CloudWatch, Azure Monitor, Azure Liftr, and Google Cloud™ Kubernetes with GKE™ AutoPilot cluster.

Vulnerabilities

Prioritize and sort vulnerabilities based on Davis Security Score, detection time, or the number of affected entities. Quickly identify if public exploits, internet exposure, or reachable data assets are exposed. Get powerful insights into the true impact of vulnerabilities in your environment. The embedded Davis Security Advisor can help you enhance remediation for third-party vulnerabilities with AI-assisted and precise recommendations on remediation actions, allowing you to address several critical vulnerabilities simultaneously. Davis AI uses security intelligence and runtime context to determine risk and remediation based on criteria like internet exposure and access to sensitive data.

Prioritization of vulnerabilities and recommendations by Davis Security Advisor
Figure 3. Prioritization of vulnerabilities and recommendations by Davis Security Advisor

Service-Level Objectives

Measuring a system’s reliability can be complex and overwhelming. Indicators that provide insights into the environment’s health and performance vary across industries, products, and applications. Applying Service Level Objectives (SLO) to track these indicators is a common best practice within site reliability engineering. Still, an SLO’s quality lies in the significance of the underlying service-level indicator.

Dynatrace guides you in quickly setting up the most valuable SLOs—considering the typically used metrics but providing the freedom to use any data stored in Dynatrace Grail™ data lakehouse, such as logs and events.

It doesn’t matter if you need the typically used failure rate or response-time metrics to ensure your system’s availability and performance or if you need to rely on abnormal log drops to gain insights into raising problems—SLOs leveraged with Grail provide all the information you need.

Automate your cloud

Answer-driven Dynatrace Automation is further extended by providing direct interaction with your cloud and cloud-native ecosystem:

  • Kubernetes: Read manifests, free up resources (for example, delete failed terminations), and restart deployments.
  • AWS: Automate your AWS infrastructure with actions across EC2, S3, Lambda, and more.
  • GitHub®: Integrate with your GitHub repositories. Automate issues and pull requests (for example, to change configuration files for sizing).
  • GitLab™: Integrate with your GitLab projects. Automate issues and merge requests (for example, to change configuration files for sizing).

Example: Predictive auto-scaling for Kubernetes workloads

Kubernetes provides flexibility but also introduces complexity. For instance, manual scaling is time-consuming, reactive, and prone to errors. Using Dynatrace Automation and Davis AI helps you predict bottlenecks and automatically open pull requests to scale applications. This proactive, automated approach minimizes downtime, ensures your applications perform at their best, and helps optimize resource utilization and cost.

The Auto-Scaling tutorial provides a step-by-step guide to automatically scaling Kubernetes workloads up or down, horizontally or vertically.

Workflow in Dynatrace video thumbnail

Try out cloud-native observability yourself

Proactively manage your environments to increase performance and reduce cost. It’s never been easier to truly own your cloud!

The capabilities highlighted in this blog post will be available in Dynatrace SaaS environments in the coming weeks.

The post Cloud-native observability made seamless with OpenPipeline and AI-driven observability appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/simplify-cloud-native-environments-ai-driven-observability/feed/ 0
How to automate version aware distributed trace analysis https://www.dynatrace.com/news/blog/how-to-automate-version-aware-distributed-trace-analysis/ https://www.dynatrace.com/news/blog/how-to-automate-version-aware-distributed-trace-analysis/#respond Mon, 13 Sep 2021 09:17:52 +0000 https://www.dynatrace.com/news/?p=46075 Version Aware Diagnostics Dashboard

Distributed Traces are “the source of truth” for developers and architects as they capture the true end-to-end execution path for each individual request processed by your applications and services. When running unit or API tests in a local dev environment, IDE (Integrated Development Environment) and test tool integrations with observability platforms make it easy for […]

The post How to automate version aware distributed trace analysis appeared first on Dynatrace news.

]]>
Version Aware Diagnostics Dashboard

Distributed Traces are “the source of truth” for developers and architects as they capture the true end-to-end execution path for each individual request processed by your applications and services. When running unit or API tests in a local dev environment, IDE (Integrated Development Environment) and test tool integrations with observability platforms make it easy for engineers to run distributed trace analysis on the distributed traces they just generated. Answering questions like “Does my code correctly call the new backend service version for my specific use case?” becomes easy as the outgoing call with request and response details will be shown in the distributed trace.

The following screenshot shows a distributed trace in Dynatrace with detailed version information of every service call involved. In the rest of the blog, you will learn more about how to capture version information and use it to answer your version-specific questions:

Dynatrace gives developers full version and request details of every service involved in end-2-end distributed trace
Dynatrace gives developers full version and requests details of every service involved in end-2-end distributed trace

If you’re not already capturing distributed traces in your development environments, I hope this blog enables you to make the case for it.

From a handful to millions of distributed traces requires automation

When distributed traces are captured in shared testing or production environments, where hundreds of services/microservices are deployed in one or multiple versions, you potentially end up with millions of captured traces in a short time span. If you then figure out how to automate the analysis of those traces you can empower DevOps and SRE teams as they need answers to questions such as:

  • “Which versions of our services are currently processing our critical transactions?”
  • “How does an overloaded backend service impact the SLOs of the frontend service?”
  • “Which frontend services are responsible for the changed traffic behavior on the backend services?”
  • “Is there a different behavior between two versions of a service? If so – shall we stop the rollout into production?”

I personally keep hearing those questions more frequently these days, which is somewhat worrying. Many members in our Dynatrace community are moving towards k8s and microservices which allows DevOps and SREs to leverage zero-downtime update strategies (also known as Progressive Delivery), such as Blue/Green, Rolling Updates, Canary Deployments or Feature Flags more easily. To answer version-specific questions like the one above it’s not only necessary to capture version information on every distributed trace but you must also automate the analysis as no one can dig through millions of traces manually to end up with answers that lead to better delivery and release decisions.

The good news is that Dynatrace PurePath, our leading automated distributed trace technology for the past 15+ years, is version aware by default meaning that it automatically captures the version information on every PurePath and provides automated analysis options to answer those version specific questions. But it’s not just the raw data that matters – it’s what Dynatrace does with the data, which I’ll delve into now. , which I’ll delve into now.

To learn more about Dynatrace’s version aware analysis capabilities I invited Thomas Rothschaedl, Product Manager at Dynatrace, to my latest Performance Clinic where he explained how version aware PurePaths are captured, how Dynatrace provides real-time release overview, and how Dynatrace provides automated answers to DevOps & SRE based on version-aware PurePath data.

While I encourage you to watch the full 30 minutes recording on YouTube or Dynatrace University I captured the key learnings in the remainder of this blog:

Dynatrace real-time release and version overview

Thomas started of reminded me about the Dynatrace Releases screen, which gives our users a live overview of all deployed releases in every monitored environment, even providing release lifecycle events (deployment, tests, quality gate, promote, rollouts, rollbacks, problems, etc.) as well as direct access to any open development or support tickets:

Dynatrace provides a real-time version-aware release overview answering critical questions for DevOps, SREs and Release Managers
Dynatrace provides a real-time version-aware release overview answering critical questions for DevOps, SREs, and Release Managers

If you’d like to learn more about the releases overview make sure to watch my Performance Clinic on Risk-Free Delivery with Dynatrace Cloud Automation Release Management.

Analyzing rolling version updates through multi-dimensional analysis

At Dynatrace we’re proud to use Dynatrace on Dynatrace which allowed Thomas to show an internal example of analyzing rolling software updates. The screenshot below shows a 72-hour analysis window of the rolling updates we do in our end-to-end testing environment. Here, we continuously roll out the latest Dynatrace versions that come out of our build system. The environment is constantly under load and it’s therefore great to see how the rolling update is truly and smoothly updating from one version to the next:

Dynatrace version-aware PurePath analysis automates the validation of successful rolling updates or canary deployments
Dynatrace version-aware PurePath analysis automates the validation of successful rolling updates or canary deployments

Analyzing request count (=throughput) is just one option, as you can see in the next section.

Automatic regression detection across versions

While the above example nicely shows the rolling updates that happened during constant load on the system, it also only focuses on throughput. What’s more interesting is if you switch to a different metric that can be extracted from PurePaths such as “Number of Exceptions Thrown”, “Number of Database Calls Made”, “Number of Database Rows Fetched”, or “Time Spent in I/O”.

Thomas demonstrated this in the performance clinic I mentioned above, where he wanted to know if any of our builds introduced a regression that caused more exceptions to be thrown. Exceptions are by default captured by Dynatrace OneAgent as part of the version aware PurePath. As you can see from the below screenshot, he immediately found a regression that was introduced in one of the builds that were rolled out earlier that day:

Automatically detect regressions introduced with a particular version by focusing on different metrics extracted from version-aware PurePaths
Automatically detect regressions introduced with a particular version by focusing on different metrics extracted from version-aware PurePaths

This is the true power of analyzing a massive amount of PurePaths in an automated way. No one would be able to identify those problems quickly by manually digging through millions of PurePaths. That’s why Dynatrace’s automation is valued by our users as it finds these issues automatically without any manual effort. But there’s more than what Thomas showed us.

Diagnostics, metrics, dashboards, and alerting

In the 30 minutes I had, Thomas walked me through the use cases he additionally demonstrated how to:

  • Drill to the offending line of code of the exception regression
  • Create metrics, put them on a dashboard and roll those out across all teams
  • Get alerted on version-specific anomalies
Dynatrace dashboard including host, service and version specific request metrics. Easy to share between teams
Dynatrace dashboard including host, service, and version-specific request metrics. Easy to share between teams

If you want to see all these demos, then check out the Performance Clinic recording. The live demo piece starts at the 15:35 timestamp.

Make better release decisions through version aware distributed traces

Whether you use Dynatrace or any other tool to capture distributed traces, it should be clear that you must make sure to capture version information on each trace and have a way to analyze large volumes of distributed traces to make better release and deployment decisions. If you don’t yet have a distributed tracing option, or if your current tooling doesn’t support what Thomas has shown in his demo, then feel free to sign up for a Dynatrace trial and try it out yourself.

To end this, I’d like to say THANK YOU Thomas for your great preparation of the Performance Clinic content. You did an amazing job in showing the value of the latest capabilities in Dynatrace. I also want to say THANK YOU to Dynatrace engineering, which has not only built a great platform but also uses Dynatrace on Dynatrace and with that, makes our demos and storytelling even easier.

The post How to automate version aware distributed trace analysis appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/how-to-automate-version-aware-distributed-trace-analysis/feed/ 0
Kubernetes: Challenges for observability platforms https://www.dynatrace.com/news/blog/kubernetes-challenges-for-observability-platforms/ https://www.dynatrace.com/news/blog/kubernetes-challenges-for-observability-platforms/#respond Mon, 23 Nov 2020 08:23:15 +0000 https://www.dynatrace.com/news/?p=40853 person shines a light on the Kubernetes logo to discover the root cause of OOMKilled out of memory errors, Kubernetes adoption and Kubernetes survey

“The thought experiment, Schroedinger’s cat, introduces a cat that is both alive and dead, at the same time. No better analogy exists for describing the complexity of monitoring a platform like Kubernetes, where things come, go, live, and die, dozens of times, every minute.” — Matt Reider, Dynatracer and Kubernetes Wizard

The post Kubernetes: Challenges for observability platforms appeared first on Dynatrace news.

]]>
person shines a light on the Kubernetes logo to discover the root cause of OOMKilled out of memory errors, Kubernetes adoption and Kubernetes survey

Kubernetes is the de-facto standard for container orchestration as it solves many problems, like distributing workloads across machines, achieving fault tolerance, and re-scheduling workloads when problems occur. While speeding up development processes and reducing complexity does make the lives of Kubernetes operators easier, the inherent abstraction and automation can lead to new types of errors that are difficult to find, troubleshoot, and prevent.
Typically, Kubernetes monitoring is managed using a separate dashboard (like the Kubernetes Dashboard or the Grafana App for Kubernetes) that shows the state of the cluster and alerts when anomalies occur. Monitoring agents installed on the Kubernetes nodes monitor the Kubernetes environment and give valuable information about the status of nodes. Nevertheless, there are related components and processes, for example, virtualization infrastructure and storage systems (see image below), that can lead to problems in your Kubernetes infrastructure.

Kubernetes Dashboard as a part of a comprehensive monitoring solution
Kubernetes Dashboard as a part of a comprehensive monitoring solution
This article introduces the key building blocks to consider when designing your Kubernetes monitoring solution.

As a platform operator, you want to identify problems quickly and learn from them to prevent future outages. As an application developer, you want to instrument your code to understand how your services communicate with each other and where bottlenecks cause performance degradations. Fortunately, monitoring solutions are available to analyze and display such data, provide deep insights, and take automated actions based on those insights (for example, alerting or remediation).

The Kubernetes experience

When using managed environments like Google Kubernetes Engine (GKE), Amazon Elastic Kubernetes (EKS), or Azure Kubernetes Service it’s easy to spin up a new cluster. After applying the first manifests (which are likely copied and pasted from a how-to tutorial), a web server is up and running within minutes.
However, when extending the configuration for production, with your growing expertise, you may discover that:

  • Your application isn’t as stateless as you thought it was.
  • Configuring storage in Kubernetes is more complex than using a file system on your host.
  • Storing configurations or secrets in the container image may not be the best idea.

You overcome all these obstacles, and after some time, your application is running smoothly. During the adoption-phase, some assumptions about the operating conditions were made, and the application deployment is aligned to them. Even though Kubernetes has built-in error/fault detection and recovery mechanisms, unexpected anomalies can still creep in, leading to data loss, instability, and negative impact on user experience. Additionally, the auto-scaling mechanisms embedded in Kubernetes can have a negative impact on costs if your resource limits are set to high (or not set at all).
To protect yourself from this, you want to instrument your application to provide deep monitoring insights. This enables you to take actions (automatically or manually) when anomalies and performance problems occur that have an impact on end-user experience.

What does observability mean for Kubernetes?

When designing and running modern, scalable, and distributed applications, Kubernetes seems to be the solution for all your needs. Nevertheless, as a container orchestration platform, Kubernetes doesn’t know a thing about the internal state of your applications. That’s why developers and SRE’s rely on telemetry data (i.e., metrics, traces, and logs) to gain a better understanding of the behavior of their code during runtime.

  • Metrics are a numeric representation of intervals over time. They can help you find out how the behavior of a system changes over time (for example, how long do requests take in the new version compared to the last version?).
  • Traces represent causally related distributed events on a distributed system, showing, for example, how a request flows from the user to the database.
  • Logs are easy to produce and provide data in plain-text, structured (JSON, XML), or binary format. Logs can also be used to represent event data.
  • Apart from the three pillars of observability (i.e., logs, metrics, and traces), more sophisticated approaches can add topology information, real user experience data, and other meta-information.

Monitoring makes sense of observability data

To make sense of the firehose of telemetry-data provided by observability, a solution for storing, baselining, and analyzing is needed. Such analysis must provide actionable answers with anomaly root-cause detection and automated remediation actions based on collected data. A wide range of monitoring products with distinct functions, alerting methods, and integrations are available. Some of these monitoring products follow a declarative approach, where the hosts and services to be monitored must be specified exactly. Others are almost self-configuring — they automatically detect entities to be monitored or the monitored entities register themselves when the monitoring agent is rolled out.

A layered approach for monitoring Kubernetes

“Kubernetes is only as good as the IaaS layer it runs on top of. Like Linux, Kubernetes has entered the distro era.” — (Kelsey Hightower via Twitter, 2020)
Although a Kubernetes system may run perfectly by itself, with no issues reported by your monitoring tool, you may run into errors outside of Kubernetes that can pose a risk.
Example:
Your Kubernetes deployment uses virtual machines that run inside dynamically provisioned virtual machine image files (best practice or not). You haven’t detected that there’s a shortage of disk space on the virtualization host (or its shared storage). Now, when the machine image attempts to scale up, your hypervisor stops the virtual machine, thereby making one of your nodes unavailable.
In this example, neither the Kubernetes monitoring itself or the OS agent installed on the Kubernetes node caught the problem. As there are many examples of such cross-cutting anomalies, a more comprehensive approach should be considered for Kubernetes monitoring. Such a solution can be broken down into smaller, more focused domains as shown in the following illustration.

Layers of a Kubernetes Monitoring Solution
Layers of a Kubernetes Monitoring Solution
Let’s look at these layers one at a time.

Cloud provider/infrastructure layer

Depending on the deployment model you’re using, problems can occur in the infrastructure of your cloud provider or in your on-premise environment. When running in a public cloud environment, you want to ensure that you don’t run out of resources while also ensuring that you don’t use more resources than will be needed when your Kubernetes cluster starts to scale. Therefore, keeping track of the quotas configured on the cloud provider, but also monitoring the usage and costs of the resources you consumed will help you reduce costs while not running out of resources. Additionally, problems can be caused by changes in the cloud infrastructure. Therefore, audit logs can be exported or directly imported into a monitoring system.
When running your Kubernetes environment on-premise, it’s necessary to monitor all infrastructure components that could affect it. Some examples are the network (switches, routers), storage systems (especially when using thin provisioning), and virtualization infrastructure. Some typical measures for the network would be the throughput, error rates on network interfaces, and even dropped/blocked packets on security devices. Log file analysis helps you proactively detect problems (for example, over-provisioning) before they occur.

Operating system / Instance layer

If you do not run your Kubernetes cluster on a managed service, you are responsible for keeping your operating system up-to-date and maintained. When doing so, it makes sense to check the status of your Kubernetes services (such as kubelet, api-server, scheduler, controller-manager) and the container-runtime (for example, containerd or Docker). Additionally, checking if security updates/patches are available and automatically installing them at the next update cycle should be on your list. Even in this layer, log entries will help you find out if there is something wrong on your system and are a useful source for auditing the system.

Cloud platform layer

By designing the lower layers of your monitoring solution, you ensured that the infrastructure of your Kubernetes environment is stable and observable. Many problems which happen in Kubernetes arise because of misconfigurations in the manifests or because the number of applications in the cluster is growing without expanding the infrastructure.
As described in many guides, you can check if all nodes in your cluster are schedulable (for example, kubectl get nodes). The same can be done with pods, deployments, and any other Kubernetes object type. Especially, when you keep onboarding new services on your cluster, you might figure out that some Pods are in a “PENDING” state. This indicates that the scheduler is unable to do its work properly. Using kubectl describe <kind> <object> will give you more insights and will print out the events related to this object. Often, this information is useful and will give you an idea about what is missing (or misconfigured).
Tip: As for every other clustered system, keep in mind that a node can fail intentionally (for example, following an update) or accidentally. In this case, you want to be able to schedule your workloads on the remaining nodes, so if you are configuring threshold values on your monitoring infrastructure you should plan for such a reserve.
In some cases, you might configure a deployment and find out that no pod will be created (for example, by using a service account that does not exist). One indication that there is something wrong is a diverging value of desired and available pods of a deployment (kubectl get deployment).
Finally, when configuring low memory requests, there might be the situation that your pods restart many times which can happen because they run out of memory. This is simply fixed by adapting the container requests and limits.

Application layer

Even if your infrastructure runs perfectly and Kubernetes shows no errors, you still might hit problems on the application layer.
Example:
You are running a multi-tier application (web server, database) and you configured a HTTP health check which simply prints out “OK” on the application server and one which is running a simple database query on the Database. Both health checks are running perfectly fine and it seems that there is no problem. When a customer tries to run a specific action on the web interface, a white page is shown, and the page loads infinitely.
The application problem in the above example does not force the system to crash and cannot be detected using simple check mechanisms. As described in one of the first sections, however, you can instrument your application to detect such anomalies. For example, traces can show that requests are passed to the application server, the database gets queried, but doesn’t produce a response. After some time, it might happen that additional requests are queued on the application (or some other customers try to do the same thing) and the connection pool fills up (could be a metric). After a while, it might be the case that the HTTP Server itself is no longer able to get requests and after that, the health check will fail.
Using synthetic monitoring techniques, customer behavior is simulated and as a result, the availability from a customer’s view can be validated. In the case of the example described before, a check that simulates this behavior can be set up and it will report an error, as the request does not finish (in time).
Real user monitoring (RUM) gives you insights into the behavior and experience of your users with your application. RUM helps you identify errors, but also find usability issues, like many customers leaving your site at a specific point of an ordering process.

Business layer

The best new feature can be unsuccessful if the customer is unable to use it or does not like it. Therefore, mapping changes or new features in your application to business-related metrics like revenue or conversion rate can help quantify the results of your software development efforts. For instance, an updated version of an application can be identified using tags, and the metrics (for example, orders per hour) compared to the previous version. If you see a negative impact on these metrics, rolling back to the last successful version is to be a good remediation option.

Possible solutions

Many products exist to support you on your Kubernetes monitoring journey. If you prefer to use open source products, there are a lot of CNCF (Cloud Native Computing Foundation) projects, like Prometheus for storing and scraping monitoring data. Additionally, OpenTelemetry helps with instrumenting software and Jaeger is used to represent tracing data. Other projects, as Zabbix and Icinga can help you to monitor your services and infrastructure. Grafana is a tool for representing data collected by almost any monitoring tool.
While many open-source tools are specialized to fulfil their use cases well, commercial solutions typically cover a broader set of infrastructure, application, and real user monitoring use cases. For example, Dynatrace covers all the functionality described in this article out of the box, with the smallest setup effort. If you want to know more about Dynatrace features for Kubernetes monitoring, have a look at these four blog posts:

Conclusion

When designing a monitoring solution for Kubernetes, there are many things to consider beyond Kubernetes itself. By breaking up your monitoring into smaller chunks, teams are able to focus on and maintain their areas of responsibility. For instance, application deployments can be switched from a on-premises to a public cloud infrastructure deployment without affecting the monitoring of the application itself.
As Kubernetes is a highly dynamic orchestration platform and application instances can come and go in a short time, a monitoring solution that can deal with this behavior ensures a smoother monitoring experience.
There are a lot of approaches and tools that can support you on your monitoring journey. Commercial observability platforms, such as Dynatrace, enable you to monitor your Kubernetes infrastructure comprehensibly with little setup effort.

Building your Kubernetes monitoring solution with these things in mind will help you keep your customers satisfied and your systems stable. Using Dynatrace, you will get a solution out-of-the-box. To get familiar and started with Dynatrace, you can check out the free trial.

The post Kubernetes: Challenges for observability platforms appeared first on Dynatrace news.

]]>
https://www.dynatrace.com/news/blog/kubernetes-challenges-for-observability-platforms/feed/ 0