ARGUS
A hundred eyes, and never all of them closed
Same idea, pointed at every Kubernetes cluster you run.
Self-hosted, AI-assisted incident detection and repair for Kubernetes fleets. ARGUS watches every cluster you run, works out what actually broke, and proposes a fix with the evidence attached — then stops, because nothing changes until a person approves it.
The loop
Four questions, answered in order
Notice
Every cluster, continuously. A problem has to persist before ARGUS calls it an incident, so nobody is woken for noise.
Group
One bad config behind twelve crash-looping pods is one incident, not twelve alerts.
Explain
The root cause in plain words, with the evidence that led to it — so you can check the reasoning rather than trust it.
Fix
A pull request, a change to the cluster, or both. Proposed, and never applied on its own.
See it before you talk to us
One incident, followed end to end
- A fleet view that leads with what it cannot see
- A root cause with the read-only calls it came from
- Why this fix is a pull request and that one is not
- The approval screen, and the dry run behind it
- The incident closing itself — and the undo it refuses
Sample data — the clusters, namespaces and workloads on those screens are invented. The product’s own wording is not: the root cause, the routing decision, the approval gate and the undo refusal are the strings ARGUS actually emits.
ARGUS runs on your own infrastructure, and every cluster dials out to it. Nothing in your network listens for us.
Safety
Nothing changes until you say so
Hands-off mode is off by default, and most teams leave it that way.
GitOps first
The fix lands where it will survive
- A pull request against your sourceWhole files, as a diff. You review it and merge it. Nothing merges itself.
- Or a change straight to the clusterRestart, scale, resize, roll an image back — for what GitOps does not own.
- Then it watches it landAnd reopens the incident if the symptom comes back.
Commercially
It comes with our container platform services
- No licence
No extra cost, per cluster or per node
It is part of the engagement, so the tenth cluster arrives without a conversation about cost.
- Self-hosted
It runs where your platform runs
On your infrastructure, against your database, behind your own sign-in.
- Supported
Run by the people who built the platform
The engineers who put the cluster together are the ones who tune the detectors on it.
Common questions
Worth asking early
Documentation
Everything, written down
Point it at one cluster and see what it finds
A single cluster, a fortnight, and a read-only start — no autonomy, no write access, just the notices and the explanations. It is included with our container platform work, so the only thing it costs is the afternoon it takes to install.

