investigate_latency
Correlate latency metrics, error logs, and alerts for a vLLM model
How to use it
investigate_latency is exposed by the Rhoai Observability MCP MCP server. Add the server to your MCP client (Claude Desktop, Cursor, Windsurf and others), and the investigate_latency tool becomes available to the model automatically. See the full listing for setup details and every tool this server provides.
Install Rhoai Observability MCP
uv pip install -e ".[dev]"Other tools in Rhoai Observability MCP (20)
Get detailed description of a Kubernetes resource
Get alerts grouped by their routing labels
Get active alerts from Alertmanager, filterable by severity and labels
Get panels and their queries from a Grafana dashboard
List Kubernetes events, filterable by resource and reason
List KServe InferenceService resources
Get node status, capacity, and GPU allocation info
Get logs for a specific pod by namespace and name
List pods in a namespace with status, restarts, and creation time
Fetch a distributed trace by its trace ID
Get a summary of key vLLM metrics (TTFT, TPOT, E2E, cache, queue) for a model
Correlate error logs, alerts, and Kubernetes events in a namespace
Correlate GPU utilization, KV cache, queue depth, and pod status
List available Grafana dashboards, filterable by tag or title
List available Prometheus metric names, optionally filtered by regex
List available trace tag names for building TraceQL queries
Execute a LogQL query against OpenShift LokiStack
Execute a raw PromQL query against ThanosQuerier
Execute a PromQL range query to get time-series data (trends, spikes, correlations)
Search for traces using TraceQL expressions