Skip to content

Set Up Service Monitoring

This guide covers how to monitor MCP servers, services, and workflows in muster using built-in health checks, events, and CLI commands.

Overview

muster provides several monitoring capabilities out of the box:

  • Health checks for MCP servers and service instances
  • Event tracking for resource lifecycle changes and failures
  • CLI commands for inspecting status, connectivity, and tool availability

Prerequisites

  • muster installed and running
  • At least one MCP server or service configured

Quick Start

1. Check Resource Health

Use muster check to verify the health of any resource:

# Check a specific MCP server
muster check mcpserver kubernetes

# Check a workflow (validates steps, tools, and parameters)
muster check workflow deploy-app

Exit codes indicate health status: - 0 -- Available and healthy - 3 -- Unavailable or unhealthy - 4 -- Degraded / partial availability - 5 -- Connection error

2. List Resources with Status

# List all MCP servers with their current status
muster list mcpserver

# Wide output for more details (tool counts, connection state)
muster list mcpserver -o wide

3. Monitor Events

# View recent events
muster events

# Filter to warnings from the last hour
muster events --type Warning --since 1h

# Filter by resource
muster events --resource-type mcpserver --resource-name kubernetes

Configure MCP Server Health Checks

Enable periodic health checks in the MCPServer spec to detect failures automatically:

apiVersion: muster.giantswarm.io/v1alpha1
kind: MCPServer
metadata:
  name: kubernetes
  namespace: default
spec:
  type: stdio
  autoStart: true
  command: ["mcp-kubernetes"]
  description: "Kubernetes cluster management MCP server"
  healthCheck:
    enabled: true
    interval: "30s"
    timeout: "10s"

Health check status values:

Status Meaning
healthy Running and responsive
unhealthy Running but not responding correctly
error Not running or returning errors
starting In startup phase
unknown Status cannot be determined

Monitor with Events

muster emits events for key lifecycle transitions. Use these for alerting and diagnostics.

MCP Server Events

Event Meaning
MCPServerHealthCheckFailed Health checks are consistently failing
MCPServerRecoveryStarted Automatic recovery process began
MCPServerRecoverySucceeded Recovery restored the server
MCPServerRecoveryFailed Recovery failed
MCPServerToolsDiscovered Tools were discovered from the server
MCPServerToolsUnavailable Tools became unavailable

Service Instance Events

Event Meaning
ServiceInstanceHealthy Health checks passing
ServiceInstanceUnhealthy Health checks failing
ServiceInstanceHealthCheckFailed Individual health check failed
ServiceInstanceHealthCheckRecovered Recovered after failures

Example: Watch for Warnings

# Stream warning events
muster events --type Warning --since 5m

# JSON output for scripting
muster events --type Warning --output json

Monitoring Scripts

Check All MCP Servers

#!/bin/bash
echo "=== MCP Server Health ==="
muster list mcpserver --output json | jq -r '.[].name' | while read server; do
  STATUS=$(muster check mcpserver "$server" --output json | jq -r '.status')
  echo "  $server: $STATUS"
done

Pre-Deployment Validation

#!/bin/bash
# Verify all critical servers are healthy before deployment
CRITICAL_SERVERS=("kubernetes" "prometheus" "github")

for server in "${CRITICAL_SERVERS[@]}"; do
  STATUS=$(muster check mcpserver "$server" --output json | jq -r '.status')
  if [ "$STATUS" != "Healthy" ]; then
    echo "FAIL: $server is $STATUS"
    exit 1
  fi
  echo "OK: $server"
done
echo "All critical servers healthy."

Prometheus Integration

If you have a Prometheus MCP server configured, AI agents can query metrics directly:

apiVersion: muster.giantswarm.io/v1alpha1
kind: MCPServer
metadata:
  name: prometheus
  namespace: default
spec:
  type: stdio
  autoStart: true
  command: ["mcp-prometheus", "--config", "/etc/prometheus/config.yaml"]
  env:
    PROMETHEUS_URL: "http://localhost:9090"
  description: "Prometheus metrics collection and querying"

Once configured, agents can query Prometheus metrics through natural language (e.g., "show me the error rate for the API service over the last hour").

Troubleshooting

Health Check Failing

# Get detailed server info
muster get mcpserver <server-name>

# Check recent events for the server
muster events --resource-type mcpserver --resource-name <server-name>

# Verify the server binary is available
which <server-command>

Events Not Appearing

  • Verify the resource name and type are correct
  • Check time range: muster events --since 24h
  • Use --output json for machine-readable diagnostics

Service Stuck in Unhealthy State

  • Verify health check thresholds aren't too aggressive (low failureThreshold with short interval)
  • Review service events: muster events --resource-type service --resource-name <name>