Skip to technology filters Skip to main content
Dynatrace Hub

Extend the platform,
empower your team.

Popular searches:
Home hero bg
NVIDIA GPUNVIDIA GPU
NVIDIA GPU

NVIDIA GPU

Monitor base parameters of the GPU, including load, memory and temperature.

Extension
Free trialDocumentation
Dashboard showing NVIDIA GPUs
  • Product information
  • Release notes

Overview

This extension monitors base parameters of NVIDIA GPUs, tracking load, memory and resource utilization of the GPUs. The extension leverages Python access to NVIDIA toolset to provide details on GPU utilization.

This is intended for users, who:
Want to expand monitoring of their hosts onto GPU and have an overview of their utilization.

Use cases

This extension enables you to:

  • Monitor utilization of the GPU across your environment
  • Locate bottlenecks in GPU memory usage

Get started

For more information on the installation and configuration, please see NVIDIA GPU extension in the Dynatrace Documentation.

Details

Compatibility information

This extension relies on following external libraries, that need to be supported by your GPU (card and driver):

  • gpustat
  • nvidia-ml-py
Dynatrace
Documentation
By Dynatrace
Dynatrace support center
Subscribe to new releases
Copy to clipboard

Feature sets

Below is a complete list of the feature sets provided in this version. To ensure a good fit for your needs, individual feature sets can be activated and deactivated by your administrator during configuration.

Feature setsNumber of metrics included
Metric nameMetric keyDescriptionUnit
GPU utilizationnvidia.gpu.utilizationUtilization of GPU unit as provided by NVIDIA librariesPercent
GPU processesnvidia.gpu.processesNumber of processes running on GPU unit as provided by NVIDIA librariesCount
Metric nameMetric keyDescriptionUnit
GPU temperaturenvidia.gpu.temperatureTemperature of GPU unit as provided by NVIDIA librariesCelsius
GPU power drawnvidia.gpu.power_drawPower currently used by GPU unit as provided by NVIDIA librarieswatt
Metric nameMetric keyDescriptionUnit
GPU used memorynvidia.gpu.memory_usedUsed memory of GPU unit as provided by NVIDIA librariesMebiByte
GPU total memorynvidia.gpu.memory_totalTotal memory of GPU unit as provided by NVIDIA librariesMebiByte

Full version history

To have more information on how to install the downloaded package, please follow the instructions on this page.
ReleaseDate

Full version history

Version 2.0.2

Adds recommended feature sets.

Introduces three recommended feature sets to allow selective metric collection:

  • GPU Performance — utilization rate and active process count.
  • GPU Memory — total capacity and current memory usage.
  • GPU Power and Thermal — power draw and temperature.

All three feature sets are recommended and enabled by default.

Full version history

Version 3.0.0

Adds Smartscape on Grail support.

This version introduces Smartscape on Grail support, with the following nodes and edges:

  • NVIDIA_GPU (was nvidia:gpu in the classic topology)
    • runs_on HOST

Notes

  • This feature is only relevant to SAAS customers.
  • This version requires a minimum Dynatrace version of 1.340 and a minimum ActiveGate/OneAgent version of 1.338.
  • You must have migrated OpenPipeline configurations to the Settings API, the extension can't be installed otherwise.
  • The existing classic entities still exist and are supported

Full version history

Version 1.1.3

  • Fix an issue where the extension would fail to reinitialize the driver after an initial failure, requiring an extension restart

Full version history

🐛 Bugs fixed in this version:

  • Configure Extension link on classic dashboard points to correct extension

🚀 Improved in this version:

  • Error codes added to logged errors

Full version history

  • Add support for the dt.security_context attribute
  • Add new platform and classic dashboards
  • Add platform screens

Full version history

Version 1.0.3

🐛Bugfixes

  • Fix an issue where the extension would fail if a GPU did not report certain metrics.

Full version history

v1.0.2

  • Report 0 when there are no processes for a specific GPU

Full version history

Added power consumption metric

Full version history

Support for Nvidia GPU monitoring based on latest Extension Framework. Monitors temperature, memory, number of processes and utilization. Creates host screens for these metrics.

Dynatrace Hub
Hub HomeGet data into DynatraceBuild your own app
Dynatrace Intelligence - Agentic Operations SystemThe Dynatrace Agentic AI ecosystem
All (914)Log Management and AnalyticsKubernetesInfrastructure ObservabilitySoftware DeliveryApplication ObservabilityApplication SecurityBusiness ObservabilityDigital Experience
Filter
Type
Built and maintained by
Deployment model
SaaS
  • SaaS
  • Managed
Partner FinderBecome a partnerDynatrace Developer

AI and LLM Observability

Achieve complete visibility and insights across every layer of your AI and LLM ecosystem – from data ingestion and vector stores to agentic frameworks and prompt engineering – ensuring optimal performance, cost efficiency, compliance, and system reliability at scale.

Essentials

AI Observability logo

AI Observability

End-to-end observability for your Agentic AI and LLM workloads.

OneAgent for GenAI logo

OneAgent for GenAI

Monitor and trace your AI workloads and apps automatically with OneAgent.

OpenTelemetry for GenAI logo

OpenTelemetry for GenAI

Ingest and analyze OTel GenAI traces & metrics for your AI workloads.

OpenInference logo

OpenInference

Instrument your AI agents, services and apps with OpenInference.

AI Agents

Get visibility into Agentic AI workloads: trace execution paths, tool invocations, and inter-agent communication. Monitor and debug Agent interactions (function calling, tool-use, RAG), and resolve performance, latency, cost, and reliability issues.

OpenAI Agents logo

OpenAI Agents

Monitor and trace your OpenAI Agents.

Amazon Bedrock AgentCore logo

Amazon Bedrock AgentCore

Monitor and trace your Amazon Bedrock AgentCore Agents.

Google ADK logo

Google ADK

Monitor your Google Agent Development Kit.

LangGraph logo

LangGraph

Monitor and trace your LangChain Agents.

Pydantic AI logo

Pydantic AI

Monitor and trace your Pydantic AI agents.

MCP AI Agent monitoring logo

MCP AI Agent monitoring

Monitoring and tracing of agents communicating via MCP.

Model providers and platforms

Monitor and gain insights into the performance, consumption, latency, availability, response time, and health of the platforms used for pre-trained foundational models, agentic frameworks, and specialized AI APIs for building, training, and deploying machine learning models.

See more (16)
OpenAI logo

OpenAI

Monitoring your OpenAI & Azure OpenAI services such as GPT, o1, DALL-E, ChatGPT.

Amazon Bedrock logo

Amazon Bedrock

Observe end-to-end generative AI models provided by Amazon Bedrock.

CrewAI logo

CrewAI

Monitor CrewAI workloads and AI Agents.

Azure AI Foundry logo

Azure AI Foundry

End-to-end observability for GenAI & LLM applications build with Azure.

Anthropic logo

Anthropic

Monitor end-to-end your Anthropic services such as Haiku, Sonnet, and Opus.

Gemini logo

Gemini

Observe end-to-end multimodal AI models provided by Google Gemini.

AI Coding Agent Monitoring

Monitor coding AI agents with end‑to‑end distributed tracing, cost and performance insights, and full‑stack context, so teams can reduce token spend, resolve agent failures faster, and confidently run AI‑driven code workflows in production.

Claude Code Agent monitoring logo

Claude Code Agent monitoring

Monitor your Claude Code coding agents with OTel.

Gemini CLI Monitoring logo

Gemini CLI Monitoring

End‑to‑end visibility into Gemini CLI usage, tokens, and performance.

OpenAI Codex Monitoring logo

OpenAI Codex Monitoring

Visibility into OpenAI Codex CLI usage, tokens, and performance in Dynatrace.

Github Copilot SDK Monitoring logo

Github Copilot SDK Monitoring

End-to-end insight into Copilot SDK model calls, performance, and token usage.

OpenCode Monitoring logo

OpenCode Monitoring

Visibility into OpenCode sessions, LLM calls, tools, and performance.

OpenClaw Monitoring logo

OpenClaw Monitoring

Monitor OpenClaw agent activity with AI observability in Dynatrace.

Data management and vector stores

Monitor, optimize, and manage data ingestion, preprocessing, and storage for traditional and vector-based workflows.

See more (2)
Pinecone logo

Pinecone

Gain insight into your Pinecone vector databases to build knowledgeable AI.

LanceDB logo

LanceDB

Monitor the performance of your multimodal AI database powered by LanceDB.

Chroma logo

Chroma

Gain insights into the health of your vector and embedding databases from Chroma.

Milvus logo

Milvus

Gain insights about vector database resource utilization and cache behavior.

Weaviate logo

Weaviate

Observe your semantic cache efficiency to reduce cost and latency for LLM apps.

Qdrant logo

Qdrant

Gain insights about your Qdrant semantic vector collections.

Orchestration and Prompt Engineering Frameworks

Automate multi-step LLM workflows, manage prompt chaining, agent-based systems, and retrieval-augmented generation (RAG).

LangChain logo

LangChain

Monitor your generative AI LLM applications built by LangChain framework.

Haystack logo

Haystack

Observe your LLM applications at scale, with RAG pipeline models by Haystack.

LlamaIndex logo

LlamaIndex

Monitor your LLM-powered agents and workflows built with LlamaIndex framework.

Infrastructure and Compute Resources

Manage and monitor hardware and compute environments for training, fine-tuning, costs, and inference acceleration.

Google Cloud Tensor Processing Units logo

Google Cloud Tensor Processing Units

Observe and monitor your machine learning models built on top of Tensor Units.

TensorFlow Keras logo

TensorFlow Keras

Observe the training progress of TensorFlow Keras AI models.

NVIDIA GPU logo

NVIDIA GPU

Monitor base parameters of the GPU, including load, memory and temperature.

vLLM logo

vLLM

Monitor your services built with vLLM's inference and LLM serving solution.

Security, Governance and Traffic Management

Ensure secure, compliant, and well-routed AI traffic with transparent governance and policy enforcement.

Kong AI logo

Kong AI

Automatic, intelligent observability for Kong AI and LLM API traffic.

LiteLLM logo

LiteLLM

Automatic, intelligent observability for your LLM Gateway traffic.

More resources

Groq logo

Groq

Monitor your services built with Groq AI inference models.

Microsoft Agent Framework logo

Microsoft Agent Framework

Observe your Microsoft Agent Framework AI agents with built-in OpenTelemetry.

Are you looking for something different?

We have hundreds of apps, extensions, and other technologies to customize your environment

More resources

AI and LLM Observability

AI and LLM Observability

Leverage best-in-class observability to improve the performance, explainability, and compliance of your Generative AI applications, LLMs, and agents.
Read more
Deliver secure, safe GenAI apps with Dynatrace

Deliver secure, safe GenAI apps with Dynatrace

Amazon Bedrock, equipped with Dynatrace Davis® AI and LLM observability, gives you end-to-end insight into the Generative AI stack, from code-level visibility and performance metrics to GenAI-specific guardrails.
Read more
AI and LLM Observability Solution

AI and LLM Observability Solution

Leverage best-in-class observability to improve the performance, explainability, and compliance of your Generative AI applications, LLMs, and agents.
Read more
AI Observability Documentation

AI Observability Documentation

Documentation