ARGUS

A hundred eyes, and never all of them closed

Argus Panoptes was the watchman of Greek myth — the giant who slept a few eyes at a time so that something was always awake. The mark is an aperture that never fully opens, which is the one detail of the myth worth keeping.

Same idea, pointed at every Kubernetes cluster you run.

Self-hosted, AI-assisted incident detection and repair for Kubernetes fleets. ARGUS watches every cluster you run, works out what actually broke, and proposes a fix with the evidence attached — then stops, because nothing changes until a person approves it.

Included free with

  • Kubernetes
  • OpenShift
  • Rancher
ARGUS4 clustersCrashLoopBackOffpayments-api · prod-eu · 3 pods · 11 mingrouped ×3lifecycledetectedinvestigatingdiagnosedawaiting youroot causeConfigMap payments-config has no DATABASE_URL,which the container reads on start.7 evidenceApprove fixOpen a PRdry run passed

The loop

Four questions, answered in order

  1. Notice

    Every cluster, continuously. A problem has to persist before ARGUS calls it an incident, so nobody is woken for noise.

    Always on

  2. Group

    One bad config behind twelve crash-looping pods is one incident, not twelve alerts.

    Fleet-wide

  3. Explain

    The root cause in plain words, with the evidence that led to it — so you can check the reasoning rather than trust it.

    Read-only

  4. Fix

    A pull request, a change to the cluster, or both. Proposed, and never applied on its own.

    Your approval

nubestack.com/demo/argussample dataONE INCIDENT01Signing in02The fleet03Connecting a cluster04One cluster05Incidents06What broke07Why it broke08The recommendation09Where the fix belongs10The gate11Merged and closed12What changed13The recordwhere ARGUS stopsIncidents / CrashLoopBackOffread-onlyROOT CAUSEPROPOSED FIX · PULL REQUESTApproveDry run0707 / 13

See it before you talk to us

One incident, followed end to end

A guided tour of the real console, running in your browser. It is a single incident from first detection to automatic close — not a feature list — so you can see exactly where ARGUS acts on its own and where it stops and waits for you. Jump to any chapter; each one has its own link.
  • A fleet view that leads with what it cannot see
  • A root cause with the read-only calls it came from
  • Why this fix is a pull request and that one is not
  • The approval screen, and the dry run behind it
  • The incident closing itself — and the undo it refuses
Open the incident tour

Sample data — the clusters, namespaces and workloads on those screens are invented. The product’s own wording is not: the root cause, the routing decision, the approval gate and the undo refusal are the strings ARGUS actually emits.

Your AI providerread-only accessYour Git hostpull requests onlyARGUS hubself-hosted, behind your own sign-innoticegroupexplainfixnothing executes without a human approvalYour databaseincidents, evidence and the audit trailprod-euargus agentprod-usargus agentstagingargus agentoutbound

ARGUS runs on your own infrastructure, and every cluster dials out to it. Nothing in your network listens for us.

Safety

Nothing changes until you say so

Every proposed change goes through the same checks, in the same order, every time.

Hands-off mode is off by default, and most teams leave it that way.

proposed change1The AI reads, and only reads2Its suggestion is checked by ordinary code3A person approves it4Rehearsed against your live cluster5Your cluster decides what it may touchdone, logged, and reversible

GitOps first

The fix lands where it will survive

A change that isn’t in the repository is one the next sync undoes — so that is where ARGUS puts it.
  • A pull request against your sourceWhole files, as a diff. You review it and merge it. Nothing merges itself.
  • Or a change straight to the clusterRestart, scale, resize, roll an image back — for what GitOps does not own.
  • Then it watches it landAnd reopens the incident if the symptom comes back.

Commercially

It comes with our container platform services

ARGUS is a NubeStack product, included at no extra cost with our Kubernetes, OpenShift and container platform work.
  • No licence

    No extra cost, per cluster or per node

    It is part of the engagement, so the tenth cluster arrives without a conversation about cost.

  • Self-hosted

    It runs where your platform runs

    On your infrastructure, against your database, behind your own sign-in.

  • Supported

    Run by the people who built the platform

    The engineers who put the cluster together are the ones who tune the detectors on it.

Common questions

Worth asking early

No. The AI reads, and what it produces is a proposal. A person approves it, and a separate component with narrowly scoped write access is the only thing that ever changes anything.
Yours. A commercial provider you already hold a contract with, or a model you host yourself — either works, and switching between them is a settings change rather than a rebuild. Your incident data goes where you have already decided it may go.
No. For anything Argo or Flux is syncing, ARGUS opens a pull request instead of patching — because a change that isn’t in the repository is one the next reconcile undoes.
Only if you switch it on, within limits you set, and it is off by default. Plenty of teams leave it off and use ARGUS purely for what it finds and explains.

Documentation

Everything, written down

Install, architecture, cluster agents, incident flow and the approval model — the operational reference for running ARGUS.
Open the documentation

Point it at one cluster and see what it finds

A single cluster, a fortnight, and a read-only start — no autonomy, no write access, just the notices and the explanations. It is included with our container platform work, so the only thing it costs is the afternoon it takes to install.