Mitsuki API Observability with Grafana and Prometheus
DEV Community

Mitsuki API Observability with Grafana and Prometheus

Running the Stack

From the example directory:

cd examples/instrumentation_demo
docker compose up -d --build

Once it is up:

Service Address
Mitsuki application http://localhost:8000
Prometheus http://localhost:9090
Grafana http://localhost:3000 (no login)

The Compose File

# docker-compose.yml
services:
  mitsuki:
    build:
      context: ../..
      dockerfile: examples/instrumentation_demo/Dockerfile
    container_name: instrumentation-demo-app
    ports:
      - "8000:8000"
    networks:
      - monitoring

  # Emits the Grafana dashboard bundled with Mitsuki into a shared volume,
  # which Grafana provisions from. Runs once and exits.
  dashboard-init:
    build:
      context: ../..
      dockerfile: examples/instrumentation_demo/Dockerfile
    container_name: instrumentation-demo-dashboard-init
    command: ["mitsuki", "grafana-dashboard", "-o", "/dashboards"]
    volumes:
      - grafana-dashboards:/dashboards
    networks:
      - monitoring

  prometheus:
    image: prom/prometheus:latest
    container_name: instrumentation-demo-prometheus
    ports:
      - "9090:9090"
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml
      - prometheus-data:/prometheus
    networks:
      - monitoring

  grafana:
    image: grafana/grafana:latest
    container_name: instrumentation-demo-grafana
    ports:
      - "3000:3000"
    environment:
      - GF_AUTH_ANONYMOUS_ENABLED=true
      - GF_AUTH_ANONYMOUS_ORG_ROLE=Admin
    volumes:
      - ./grafana/provisioning:/etc/grafana/provisioning
      - grafana-dashboards:/var/lib/grafana/dashboards
      - ./grafana/dashboards:/var/lib/grafana/demo-dashboards:ro
      - grafana-storage:/var/lib/grafana
    networks:
      - monitoring
    depends_on:
      prometheus:
        condition: service_started
      dashboard-init:
        condition: service_completed_successfully

networks:
  monitoring:
    driver: bridge

volumes:
  grafana-storage:
  grafana-dashboards:
  prometheus-data:

The parts that matter:

  • dashboard-init runs mitsuki grafana-dashboard -o /dashboards once and exits. The framework dashboard ships inside the mitsuki package, and this command writes it into the volume Grafana reads.
  • condition: service_completed_successfully makes Grafana wait until the dashboard file exists. service_started would race it.
  • Two dashboard sources: grafana-dashboards holds the framework dashboard; ./grafana/dashboards holds a second dashboard for this demo's custom metrics.
  • GF_AUTH_ANONYMOUS_* lets you open Grafana without creating an account. It gives every visitor admin rights, so keep it to local use.

Configuring Prometheus

# prometheus.yml
global:
  scrape_interval: 5s
scrape_configs:
  - job_name: 'mitsuki'
    metrics_path: '/metrics/prometheus'
    static_configs:
      - targets: ['mitsuki:8000']
  • metrics_path points at /metrics/prometheus, the text-format endpoint. /metrics is the JSON summary, which Prometheus can't parse.
  • The target uses the Compose service name, mitsuki, which resolves on the monitoring network. Prometheus scrapes from inside that network, so the application's metrics.allowed_ips has to include Docker's address range, 172.16.0.0/12. The demo's application.yml already does - see Securing the Endpoints for the full list.
  • To check that Prometheus is scraping, open Status → Target health at http://localhost:9090: the mitsuki target should be UP. If it's down, the address is usually refused by allowed_ips or the path is wrong.

Generating Traffic

Create some data first:

curl -X POST http://localhost:8000/api/users \
  -H "Content-Type: application/json" \
  -d '{"username": "alice", "email": "a****@example.com"}'

curl -X POST http://localhost:8000/api/orders \
  -H "Content-Type: application/json" \
  -d '{"user_id": 1, "product_type": "digital", "amount": 99.99, "region": "us-east"}'

Then keep traffic flowing for ten minutes: two list endpoints, and a lookup of a user that doesn't exist, which adds a 404 series to Status Code Distribution:

for i in $(seq 1 600); do
  curl -s -o /dev/null http://localhost:8000/api/users
  curl -s -o /dev/null http://localhost:8000/api/orders
  curl -s -o /dev/null http://localhost:8000/api/users/99
  sleep 1
done

The Framework Dashboard

Open http://localhost:3000 → Dashboards → Mitsuki Application Metrics. This is the dashboard mitsuki grafana-dashboard writes. It only queries metrics Mitsuki itself emits, so it works unchanged for any Mitsuki application.

Row Panels

Overview Statistics

  • Total Requests
  • Requests/Second
  • Avg Response Time
  • Error Rate
  • P95 Latency
  • P99 Latency

HTTP Performance - Per Route Breakdown

  • Response Time by Route (a table with average, p95, p99, rate and total per route)

HTTP Metrics Over Time

  • Request Rate by Route
  • Response Time Percentiles
  • Status Code Distribution

Component Performance

  • Component Metrics (a table per component)

Component Metrics Over Time

  • Component Call Rate
  • Component Duration (P95)
  • Component Success vs Failure

Component Performance - Per Method

  • Call Rate by Method
  • P95 Duration by Method
  • Failure Rate by Method

Scheduler

  • Task Execution Rate
  • Task Duration (P95)
  • Task Failure Rate
  • Running Tasks

System Resources

  • Memory Usage
  • CPU Usage

The HTTP rows break everything down by route template. Requests that match no route are grouped under <unmatched>:

The component rows go one level deeper than most HTTP instrumentation - into your own controllers, services and repositories:

And then per method, so a slow repository query shows up by name:

How the Panels Are Built

Three of the queries show how the rest are built.

Request Rate by Route is the per-second rate of the request counter over the last minute, summed per route template:

sum(rate(http_requests_total{path!~"/metrics.*"}[1m])) by (path)

P95 Latency estimates the 95th percentile from the histogram buckets. The buckets have to be summed by le, the bucket boundary label, before histogram_quantile can use them:

histogram_quantile(0.95, sum(increase(http_request_duration_seconds_bucket{path!~"/metrics.*"}[5m])) by (le)) * 1000

Why percentiles rather than the average? A mean hides the one request in a hundred that takes ten times as long - which is usually the one a user notices.

Error Rate is the share of requests that returned a 5xx:

((sum(http_requests_total{status=~"5.*",path!~"/metrics.*"}) or vector(0)) / sum(http_requests_total{path!~"/metrics.*"})) * 100

The or vector(0) keeps the panel at 0 instead of empty while no 5xx has happened yet. The 404s from the traffic loop don't count - they show up in Status Code Distribution instead.

Scheduled Tasks

Background jobs fail quietly: nobody gets an error page when a nightly job dies. The demo runs one, OrderReconciliationService.reconcile_orders, every ten seconds:

# src/services/order_reconciliation_service.py
from src.repositories.order_repository import OrderRepository
from mitsuki import Scheduled, Service

@Service()
class OrderReconciliationService:
    """
    Periodically checks stored orders for invalid amounts.
    As a @Scheduled task, every run is recorded in the scheduler metrics:
    executions, failures and duration.
    Creating an order with a non-positive amount makes every following run fail,
    which shows up on the dashboard's task failure panel.
    """

    def __init__(self, order_repo: OrderRepository):
        self.order_repo = order_repo

    @Scheduled(fixed_rate=10000)
    async def reconcile_orders(self):
        orders = await self.order_repo.find_all()
        invalid = [order.id for order in orders if order.amount <= 0]
        if invalid:
            raise ValueError(f"Orders with non-positive amounts: {invalid}")

The Scheduler row of the framework dashboard shows its execution rate, p95 duration, failure rate and whether a run is in progress:

The execution rate holds at 0.1 per second, one run every ten seconds. Grafana's automatic axis stretches the small jitter around it into a zigzag. Task Failure Rate shows "No data" because no run has failed yet.

To see a failure, create an invalid order:

curl -X POST http://localhost:8000/api/orders \
  -H "Content-Type: application/json" \
  -d '{"user_id": 1, "product_type": "digital", "amount": -1}'

The next run fails and logs the error with its traceback:

2026-09-29 07:32:45,991 - mitsuki - ERROR- Scheduled task OrderReconciliationService.reconcile_orders failed with error: Orders with non-positive amounts: [4]
Traceback (most recent call last):
  File "/usr/src/app/mitsuki/core/scheduler.py", line 218, in task_loop

The failure is counted in both the scheduler metrics and the component metrics, since the service is instrumented:

component_calls_total{component="OrderReconciliationService", method="reconcile_orders", status="success"} 2.0
component_calls_total{component="OrderReconciliationService", method="reconcile_orders", status="failure"} 1.0
scheduler_task_executions_total{status="success", task="OrderReconciliationService.reconcile_orders"} 2.0
scheduler_task_executions_total{status="failure", task="OrderReconciliationService.reconcile_orders"} 1.0

The Task Failure Rate panel rises and stays up: every run fails until the invalid order is gone. That makes it the panel to alert on. Its query is the rate of failed executions per task:

sum(rate(scheduler_task_executions_total{status="failure"}[5m])) by (task)

A Grafana alert rule on that query with the condition "is above 0" notifies you when any scheduled task starts failing.

The Demo Dashboard

The demo's OrderService also records custom metrics through InstrumentationProvider. Those only exist in this application, so they get their own dashboard: Dashboards → Demo → Demo Custom Metrics, provisioned from grafana/dashboards/demo-custom-metrics.json.

Panel Query

  • Orders Created per Minute: sum(rate(orders_created_total[5m])) by (product_type, region) * 60
  • Database Write Operations: rate(database_writes_total[5m])
  • Rows Returned per Second: sum(rate(database_rows_returned_total[5m])) by (table, query)
  • Full Table Scans: rate(database_full_scan_total[5m])
  • Expensive Aggregations: rate(expensive_aggregation_total[5m])

How those metrics are recorded is the subject of Creating Custom Business Metrics in Mitsuki.

Using the Dashboard in Your Own Application

The framework dashboard is not tied to the demo. Write it into your own Grafana provisioning directory:

mitsuki grafana-dashboard -o ./grafana/dashboards

Output:

Wrote Grafana dashboard to grafana/dashboards/dashboard.json

Then point a Grafana dashboard provider at that directory, or import the file through the Grafana UI. The demo's provider for the framework dashboard:

# grafana/provisioning/dashboards/mitsuki.yml
apiVersion: 1
providers:
  # Framework dashboard, generated by `mitsuki grafana-dashboard` (dashboard-init).
  - name: Mitsuki
    orgId: 1
    folder: ''
    type: file
    disableDeletion: false
    updateIntervalSeconds: 10
    allowUiUpdates: true
    options:
      path: /var/lib/grafana/dashboards
      foldersFromFilesStructure: false

Note: The dashboard's panels don't name a datasource - they query whichever Prometheus datasource is Grafana's default. If your Grafana has several, make the one scraping your application the default, or change the panels' datasource after importing.

Cleaning Up

# stop the stack, keep Prometheus and Grafana data
docker compose down

# stop the stack and delete its volumes
docker compose down -v

Next Steps

  • Record business metrics of your own and add them to a dashboard: Creating Custom Business Metrics in Mitsuki - demo on GitHub.
  • Before running this beyond your machine, read the per-process and reverse-proxy caveats in Introducing Automatic Instrumentation.
  • For every configuration option, see the Instrumentation & Metrics documentation.
Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.