Extend your Checkmk monitoring with packages created by community members. Or create your own packages and share the here with the rest of the community.
Docker reports the CPU usage of a container like `docker stats`: **100 % per fully used CPU core**. On a host with many cores a single container can show hundreds or thousands of percent, which breaks every threshold written for "total CPU in %". On Checkmk versions older than 2.5.0p12 the built-in service also shows a bogus "under high load for ... years" duration (Werk 20187). This package adds the service **Docker CPU utilization** for Docker container hosts. It reports the utilization in percent of the available CPUs (always 0 to 100 %) and the number of CPU cores in use. ## Features - Normalized CPU utilization per container, e.g. `Total CPU: 1.48%, 1.07 of 72 CPUs in use` - Metrics `util` (shown in the standard CPU graph) and `docker_cpu_cores_used` - Levels in the WebUI, with levels over an extended time period as default (WARN after 1 h 30 min, CRIT after 3 h at 90 %), so short peaks do not alarm - Works with the **standard Docker data** of `mk_docker.py`, no agent plug-in has to be rolled out - Optional agent plug-in for **Linux and Windows** (Agent Bakery rule) that also reports the container CPU limit - Counter resets after a container restart are ignored instead of producing negative values - The built-in "CPU utilization" service is not changed - Safe to install: no service is created until a discovery rule exists ## Requirements - Checkmk **2\.5.0 or newer** - Docker hosts monitored with the Docker agent plug-in `mk_docker.py` (container hosts as piggyback hosts) ## Notes - With the standard Docker data the CPU limit of a container is not known, the value is related to all CPUs of the Docker host. The optional agent plug-in also takes the limit into account. - Full documentation and changelog are included in the package (`doc` folder, in English). ## Tested Tested live on Checkmk 2.5.0p10 in a distributed setup with the standard Docker data on Linux Docker hosts. The optional agent plug-in was checked on a Linux Docker host up to the delivery of its data. The Windows variant has not been tested on a real Windows Docker host. --- ## Deutsch Docker meldet die CPU-Nutzung eines Containers wie `docker stats`: **100 % pro voll ausgelastetem CPU-Kern**. Auf einem Host mit vielen Kernen kann ein einzelner Container hunderte oder tausende Prozent zeigen, wodurch alle Schwellwerte für "Gesamt-CPU in %" nicht mehr funktionieren. Bei Checkmk-Versionen vor 2.5.0p12 zeigt der eingebaute Service außerdem eine falsche Dauer "under high load for ... years" (Werk 20187). Dieses Paket fügt den Service **Docker CPU utilization** für Docker-Container-Hosts hinzu. Er zeigt die Auslastung in Prozent der verfügbaren CPUs (immer 0 bis 100 %) und die Zahl der genutzten CPU-Kerne. ### Funktionen - Normalisierte CPU-Auslastung pro Container, z. B. `Total CPU: 1.48%, 1.07 of 72 CPUs in use` - Metriken `util` (im Standard-CPU-Graphen sichtbar) und `docker_cpu_cores_used` - Schwellwerte in der WebUI, standardmäßig über einen Zeitraum (WARN nach 1 h 30 min, CRIT nach 3 h bei 90 %), kurze Lastspitzen alarmieren nicht - Arbeitet mit den **Standard-Docker-Daten** von `mk_docker.py`, es muss kein Agent-Plugin ausgerollt werden - Optionales Agent-Plugin für **Linux und Windows** (Agent-Bakery-Regel), das auch das CPU-Limit des Containers meldet - Zählerrücksetzungen nach einem Container-Neustart werden ignoriert, statt negative Werte zu erzeugen - Der eingebaute Service "CPU utilization" wird nicht verändert - Sicher zu installieren: Ohne Discovery-Regel wird kein Service angelegt ### Voraussetzungen - Checkmk **2\.5.0 oder neuer** - Docker-Hosts, die mit dem Docker-Agent-Plugin `mk_docker.py` überwacht werden (Container-Hosts als Piggyback-Hosts) ### Hinweise - Bei den Standard-Docker-Daten ist das CPU-Limit eines Containers nicht bekannt, der Wert bezieht sich auf alle CPUs des Docker-Hosts. Das optionale Agent-Plugin berücksichtigt auch das Limit. - Die vollständige Dokumentation und das Changelog sind im Paket enthalten (Ordner `doc`, auf Englisch). ### Getestet Live getestet auf Checkmk 2.5.0p10 in einer verteilten Umgebung mit den Standard-Docker-Daten auf Linux-Docker-Hosts. Das optionale Agent-Plugin wurde auf einem Linux-Docker-Host bis zur Lieferung seiner Daten geprüft. Die Windows-Variante wurde auf keinem echten Windows-Docker-Host getestet.
by Django01
# Apprise Notifications for Checkmk Send Checkmk host and service notifications to an **Apprise API server** and let Apprise handle the final routing to your notification services. This extension adds a native **Apprise** notification method to Checkmk Setup. Checkmk sends a provider-independent notification to Apprise, while destinations such as Signal, Matrix, Discord, Slack, Gotify, email and other services remain configured centrally in Apprise. ## Features - Native Checkmk notification method and Setup form - Checkmk 2.5.x - Host and service notifications - Problem and recovery notifications - Acknowledgements - Downtime start, end and cancellation - Flapping and custom notifications - Apprise tag-based routing - HTTP Basic authentication - Checkmk Password Store support - Plain text and rich-text messages - TLS certificate verification enabled by default - Optional private CA certificate - Configurable request timeout - Checkmk-native retry behavior for temporary delivery failures - No Apprise package or additional Python dependencies required on the Checkmk server ## Requirements - Checkmk 2.5.x - Apprise API with Apprise 2.0 or newer - A saved Apprise configuration reachable from the Checkmk site The extension uses the stateful Apprise API endpoint: `POST /notify/{config_id}` Provider configuration stays entirely on the Apprise side. ## Routing A Checkmk notification rule can optionally pass an Apprise tag expression. For example: - `ops` - `network` - `ops night` - `ops, oncall` This makes it possible to use Checkmk rules for deciding **when** a notification is generated while Apprise decides **where** it is delivered. ## Security - TLS verification is enabled by default. - Private CA certificates are supported and must reside inside the Checkmk site directory. - Credentials can be stored in the Checkmk Password Store. - Credentials, server URL, configuration ID and response bodies are not written to notification output. - Redirects and environment proxies are intentionally not used. ## Compatibility Verified with: - Checkmk 2.5.x - Checkmk Community Edition - Checkmk Enterprise Edition - Python 3.13 on the Checkmk site - Apprise 2.0 - End-to-end delivery through Apprise to Signal The MKP declares compatibility with Checkmk 2.5.x only. ## Installation Install the MKP and enable it, then create a notification rule using the **Apprise** notification method. No additional Python modules are required on the Checkmk server. Full installation, configuration and troubleshooting documentation is available in the project repository: [https://github.com/Django1982/CheckMK\_Apprise\_MKP](https://github.com/Django1982/CheckMK_Apprise_MKP) ## License GPL-2.0-only
by avaccaro
A Checkmk agent plugin that discovers host capabilities and reports each one as a host label, so you can see at a glance which plugins are worth installing on that host. From the Capability Scout service's detail you can then add the bakery rules to deploy the corresponding plugins or access user guide pages, when available. Coverage today includes databases, web and application servers, containers and virtualization, clustering and HA, backup and DR, messaging and monitoring engines, collaboration platforms, and cloud VM provisioning. **Prerequisites** For the Capabilities Scout service to render correctly, set these two things in Checkmk: 1. Do not escape HTML in service output: add a rule under Setup → Services → Service monitoring rules → "Escape HTML in service output", set it to "Don't escape HTML", limit it to the Capabilities Scout service, and activate changes. Without it, the logos and "Add rule" links show up as raw HTML text. 2. Increase maximum long output size: on hosts with many capabilities, raise Setup → Global settings → "Maximum long output size" above its default of 2000 bytes. Otherwise the service details are cut off.
by otAAAh
Monitor any JSON API in Checkmk without writing a line of code. Point it at a /health, /status, or metrics endpoint, pick the fields you care about, and get a Checkmk service for each — thresholds, graphs, and alerts included. One rule. Any API. Done. # Generic JSON API ## What you get - 🎯 **Any endpoint, unmodified** — Spring Boot, Kubernetes, vendor appliances, your own apps. No special response format required. - 🧭 **Pick fields by path** — `components.db.status`, `items[0].count`, done. - 🧾 **Monitor the response headers too** — prefix a path with `@header.` and watch what the API says *outside* the JSON: `@header.X-RateLimit-Remaining` alerts before you exhaust your quota, `@header.Last-Modified` (as a timestamp) tells you how stale the data is. - 🔁 **Auto-discover arrays *and* objects** — `nodes[*].status` becomes one service per array element, and `components[*].status` one per object key (e.g. a Spring Boot Actuator `/health` map), automatically. - 🔢 **Aggregate a collection** — where `[*]` fans out one service per element, an aggregation collapses the whole collection into one service: the **number of elements** (queue length, unhealthy nodes) or the **sum / average / min / max** of the values (`queues[*].depth`). The result is a number, so units, WARN/CRIT levels and a metric all apply. - 🔍 **Filter elements by a condition** — restrict a `[*]` wildcard or an aggregation to the elements that match: one service per node whose `status` is *not* `ok`, or a count of only the pods that aren't `Running` (equals / not-equals / regex / not-regex). - ⏱️ **Counters and timestamps, done right** — mark a field as a **counter** and monitor its per-second **rate** instead of an ever-growing total (`requests_total`); mark it as a **timestamp** and monitor its **age**, so upper levels alert on stale data (`last_backup` older than 26 h → WARN). - 🖧 **One Checkmk host per element** — point a `[*]` field at a field holding a host name and every element becomes a **host of its own**, not just another service. An API describing a fleet gives you hosts with their own downtimes, contact groups and availability, instead of one host with a hundred services. - 📋 **Facts into the HW/SW inventory** — a version, a build, a region, a licence tier is not a state worth a service that is OK forever. Send it to the **inventory tree** instead, where it is searchable *across* hosts ("which hosts still run a version below 4.2?") and keeps its own change history. A `[*]` wildcard becomes a **table**, one row per element. - 💬 **Show the reason next to the value** — an optional summary text with `{path}` placeholders, so a service on `status` reads `Value: DEGRADED, replica lag 42s (leader db-3)` instead of needing a second service for the message the API already returned. - 🩺 **Every endpoint monitors itself** — each endpoint also gets a `JSON API #name#` service with the **HTTP status, response time** (thresholds optional), response size and, for HTTPS, the **TLS certificate's remaining validity** — read from the connection it is already making, so no second check against the same URL. Zero configuration; it comes with the rule. - 📄 **See the response that caused the state** — optionally report the raw body and the response headers in that endpoint's own service details, including for a **rejected** response (an unexpected status, or a body that is not JSON) — which is where an API explains itself: `HTTP 403` sends you to the credentials, `tenant disabled` sends you to the right place. Capped at a byte budget you set, with `Set-Cookie`, authorization headers and the endpoint's own secret masked first. Off by default. - 🐢 **Rate-limited API? Cache it** — give an endpoint a TTL and the agent reuses its last response instead of asking again, so monitoring cannot exhaust a request quota. It never caches an error and never answers a failed request from an expired cache, so a real outage still shows up. - 🔁 **Retry a blip instead of alerting on it** — an optional per-endpoint retry with backoff for the failures a repeat can fix (connection reset, timeout, 429/5xx), so a load balancer dropping connections during a rolling restart is not a CRIT and a notification. A 4xx or a non-JSON body is never retried, and the service reports when a retry *was* needed — it cannot hide a degrading API. - 🔗 **Many endpoints, one rule** — poll several APIs together, each with its own method, auth, and fields; an unreachable one only affects its own services. - 🧩 **One service for several fields** — where a service per field is too much: name a shared service on each field and they become **lines** of one service, which takes the **worst** of their states. `status`, `component` and `timestamp` become one service that is OK while all three are fine, each line keeping its own levels, matching, transform and unit. A `[*]` wildcard fans out into lines too, so a whole collection can be one service that goes CRIT if any element does. - 🔖 **Services named after the endpoint** — two endpoints of the same shape extract the same fields, which would give you `JSON STATUS` and `JSON STATUS (2)`. Name the endpoints instead and every service of one application shares its prefix — `JSON app1-health STATUS` — which sorts them together and makes them addressable as a **group** in service rules and notification conditions. The endpoint's own status service joins that group too, as `JSON app1-health API`. - 📈 **Thresholds & graphs in Checkmk** — WARN/CRIT and metrics live in *your* rule, not upstream in the API, and can be retuned per folder, host or service from a normal check-parameters rule without touching the connection. - 🏷️ **Labels from the response** — attach Checkmk **host** and **service labels** built from fields (`json_api/version`, `json_api/region`), so views, rules and filters can key off what the API says about itself. The hosts a `[*]` rule creates can carry labels from **their own element**, so 50 generated hosts are addressable by region or role instead of being an anonymous crowd. - 🧮 **Transform the numeric value** — apply a small arithmetic expression like `value / 1024 / 1024` (bytes→MiB) or `(value - 32) * 5 / 9` (°F→°C) before levels and the metric; safely evaluated, no `eval`. A **second field** can join in as `other`, which is what turns the used/total pair most APIs actually return into a percentage: `value / other * 100`, resolved per `[*]` element so every disk is measured against its own capacity. - 🔤 **String matching, two ways** — require a value to match a regex (pick the state when it doesn't, default CRIT), or map values like `ready` / `degraded` / `failed` straight to OK / WARN / CRIT. - 🔐 **Secure by default** — basic auth, bearer tokens, **API keys** (in a header of the API's choosing, or a query parameter) **and OAuth 2.0 client credentials** all come from the Checkmk **password store**, never from clear text in the rule or on a command line. For OAuth the agent exchanges the client ID and secret for a short-lived token itself and caches it until shortly before it expires, so monitoring does not hammer your identity provider once a minute. TLS verification is on by default, with a **custom CA bundle** for a private CA and **client certificates** for mutual TLS. Non-2xx status codes can be opted in per endpoint (read a `/health` that reports its problems with a 503), and an endpoint reachable only through a corporate egress **proxy** is supported. - 🧰 **Bonus field picker** — paste your JSON in the bundled explorer, click what to monitor, copy the ready-made rule. On Checkmk 2.5\+, install the optional companion package **Generic JSON API – Explorer (extra)** for a guided in-site wizard that builds the rule for you from a live API response. ## In 30 seconds `GET /actuator/health` → `{"status": "UP", "components": {"db": {"status": "UP"}}}` Tick `status` (expect `UP`) and `components.db.status` → instant services `JSON Health` and `JSON Database`. That's the whole setup. ## Details - **Checkmk 2.4\+**, any edition. Tested on real 2.4 and 2.5 sites. - Install via `mkp add` / `mkp enable`, or **Setup → Extension packages**. - GPL-2.0-only
by otAAAh
A guided setup wizard for the Generic JSON API agent — build a monitoring rule from your API's real response, right inside Checkmk. This is the optional companion to the **Generic JSON API** package. It adds an in-site wizard under **Setup → Quick setup** that walks you from a live API response to a finished rule — no `rules.mk`, no `curl`, no leaving Checkmk. ## Requires - The **Generic JSON API** (`json_api`) package must be installed and enabled first — this Explorer only *builds* rules for that agent; it does not monitor anything on its own. - **Checkmk 2.5 or newer**, any edition. The wizard is built on Checkmk's native Quick-Setup UI, which does not exist on 2.4. ## What it does - 🧭 **Guided, step by step** — choose the target folder and host, define one or more endpoints (URL, method, auth, headers, TLS/redirect options), then pick the fields to monitor. - 🔎 **Fetches the real response** — the wizard calls each endpoint from the site and shows you the actual JSON, so you click the fields that exist instead of guessing paths. The call is fully authenticated, OAuth 2.0 included: the wizard performs the token exchange itself, so the preview is exactly what the agent will see. - 🧾 **Body *and* headers** — a tab beside the field picker lists the response headers, so a rate-limit budget or a `Last-Modified` age is one click away instead of a path typed from memory. - 🎯 **Point-and-pick fields** — select values by path, set WARN/CRIT thresholds, units, a numeric transform, an aggregation over a collection, a counter's rate or a timestamp's age, string matching, or turn each element of a `[*]` collection into a Checkmk host of its own — the same options the agent supports. - 🔖 **Names that survive contact with a second endpoint** — an endpoint can put its own name in front of its field service names, so two applications monitored by one rule do not both produce `JSON STATUS`; the endpoint's own status service joins that group as well. The wizard offers it in the connection step and applies Setup's own rule: the prefix needs a name. - 🧩 **One service for several fields** — a field can report into a shared service instead of getting one of its own, so a small API becomes one service whose state is the worst of its fields. The wizard offers it on each field, and the review step previews the result like any other. - 📄 **Keep the response that caused the state** — an endpoint can report its raw body and response headers in its own service details, so the response that made a service CRIT is readable from the service itself instead of from a URL your browser may not even reach. Credentials are masked first. - ✅ **Live preview before you commit** — the review step evaluates every chosen field against the fetched sample and shows the resulting service state, so you catch a wrong path or threshold before the rule exists. - 🔐 **Secure by default** — credentials are stored in the Checkmk password store and referenced, never written in clear text; TLS verification stays on. That covers basic auth, bearer tokens, API keys and OAuth 2.0 client credentials alike. - 🚀 **One click to create** — the wizard writes the finished Generic JSON API rule for you. ## In short Install the **Generic JSON API** agent, then install this Explorer. Open **Setup → Quick setup → Generic JSON API**, point it at an endpoint, tick the fields you care about, and press create. The services appear on your host. ## Details - Extra/companion package — install alongside, and after, the Generic JSON API agent. - Checkmk **2\.5\+**, any edition. - Install via `mkp add` / `mkp enable`, or **Setup → Extension packages**. - GPL-2.0-only.
by rsander
Agent Plugin to check SSL certificates in specified directories Now with support to check signature algorithm Windows Plugin added JSON data in agent data section Is now able to ignore certificates with a short lifetime
by simonmeggle
Robotmk integrates Robot Framework results into Checkmk.
by lgbff
This check monitors the status of ssl cert checks on ssllab.com. Changelog: - 3\.2.0 no systemtime in agent output - 3\.x version for cmk 2.2 - 2\.1.0 change to python3 - 2\.0.1 fix datasource\_programm group definition - 2\.0.0 insert cache for api response - 1\.6.0 insert Agent status (line\[3\]) - 1\.5.3 change default status - 1\.5.2 Insert Agent Header infos - 1\.5.1 cache default settings - 1\.5.0 new data seperator (59) - 1\.4.9 fix display JSON Error on check result. - 1\.4.8 fix connect exeption - 1\.4.7 add default value for timeout - 1\.4.6 change timeout handling - 1\.4.5 fix inventory issue while JSON issues - 1\.4.4 change urlopen error handling - 1\.4.3 Rename check file - 1\.4.2 Fix check manpage - 1\.4.1 delete modul calls from check - 1\.4.0 changes for cmk version 1.4 - 1\.3.1 improve error handling.
by rsander
Plugin to gather Ceph statistics. ### Starting with Checkmk 2.4, this plug-in is distributed with Checkmk. No need to download this MKP. Use this MKP only if you want to monitor Ceph with Checkmk 2.2 or 2.3.