A Kubernetes remediation tool that holds no cluster credentials. Its only write is a git commit, so every fix it makes can be undone with git revert.
Plenty of tools will read your cluster and tell you what they think is wrong. Almost none are trusted to act on it, because nobody has a convincing answer to what stops the tool doing something catastrophic at 3am. "The model is usually careful" isn't an answer.
kubemend's answer is that every constraint on it is code with tests. It diagnoses freely and acts narrowly.
Detect
Eight read-only rules (crash loops, OOM kills, image pull failures, bad config, stuck rollouts, replica shortfalls, unschedulable pods, flapping) run against a live cluster or a recorded snapshot.
Plan
A deterministic planner picks one plan per incident from a closed set of six typed actions, each carrying the state it found and the state it intends. When no safe action exists, it abstains.
Gate
A pure policy function checks protected namespaces, deny-by-default action kinds, a computable undo, blast radius, a rate limit and an autonomy ceiling, before anything is written.
Emit
The fix becomes a one-line diff in the GitOps repo. In apply mode it commits to mainline, and in propose mode it opens a branch and pull request for a human.
Verify
A fix counts as healthy only after two consecutive clean reads. If the workload is still failing, kubemend reverts its own commit.
Record
An append-only SQLite journal tracks every incident, the revert rate and repeat offenders, shown in a read-only local console.
The planner is deterministic and there's no model in the loop yet. That's on purpose. The safety layer has to hold before anything smarter gets to propose changes.
Wiring it up to a real Argo CD setup exposed three bugs that a full green test suite had missed:
Separately, the console showed "no incidents" because the SQLite connection was bound to the thread that created it, and the journal's fail-safe swallowed the error. Each fix shipped with a test that would have caught it.
| Result | What it measures | Source |
|---|---|---|
| 26 s | Time for a good fix to recover in the live demo, after which the commit was kept | demo transcript, real k3d cluster |
| 76 s | Time before a bad fix was judged still failing and its commit reverted automatically | same demo run |
| 13 to 4 | Findings on a recorded broken cluster, reduced to plans, with 5 findings deliberately producing no action | EVALUATION.md, recorded fixture |
| 2, 1, 1 | Plans applied, proposed for review, and refused under the conservative policy | EVALUATION.md |
| 195 | Tests, with zero runtime dependencies, and a CI job that runs the real-cluster demo on every push | pytest collection and EVALUATION.md |
26 s
Time for a good fix to recover in the live demo, after which the commit was kept
demo transcript, real k3d cluster
76 s
Time before a bad fix was judged still failing and its commit reverted automatically
same demo run
13 to 4
Findings on a recorded broken cluster, reduced to plans, with 5 findings deliberately producing no action
EVALUATION.md, recorded fixture
2, 1, 1
Plans applied, proposed for review, and refused under the conservative policy
EVALUATION.md
195
Tests, with zero runtime dependencies, and a CI job that runs the real-cluster demo on every push
pytest collection and EVALUATION.md
These come from injected incidents on a local k3d cluster and a recorded snapshot. They show the mechanism working end to end. They aren't production accuracy numbers.
There's no production deployment yet, no measured precision or recall for the planner, and only Deployments are covered, so no StatefulSets, DaemonSets or Jobs. The console has no auth and the journal isn't tamper-evident. The Argo CD end-to-end run is still unfinished. After that comes a model layer that can propose actions, kept behind the same policy gate.
Next project
ECI Pipeline (SENTINEL)Watches Android security, CVE and policy feeds, works out what actually changed, and turns the significant changes into cited risk tickets.