Free
Cluster Triage
The free first step after a cluster incident. One read-only script across your Hyper-V, Failover Cluster or Azure Local environment. You receive the triage list: the five heaviest signals around your incident window, sorted by severity. Free, without subscription and without obligation.
WhatsApp · hans@cldlbs.com
Example of the form; the content comes from your own measurement.
How it works
Request the script
Fill in the form and send it via WhatsApp and e-mail; you receive the script with its instructions.
Run it from your management server
It only reads: the event channels of all nodes over the incident window. It first shows the plan of the run, with one yes-or-no question. Memory dumps stay out of scope in this free variant.
Hand over the output
You place the file on the encrypted Proton share we set up for you.
You receive the triage list by e-mail.
Request
Five details are enough to send you the right script and instructions. The number of hours sets the measurement window of the run.
What the triage list is, and is not
What you get
Five signals, sorted by severity, each with the measured value, the node and the timestamp. Whatever could not be read is listed too, with the reason.
What it is not
No cause and no reconstruction: triage sorts, the diagnosis takes deeper investigation. No report, no review session, and no subscription attached.
More than five found? The triage list also states the total number of signals found; the five heaviest are always in this free list.
Safe and read-only
- Nothing is installed and nothing is changed; nothing is left behind afterwards.
- No passwords are processed or stored. Before the file is written, the script checks its own output for values that look like credentials.
- Memory dumps do not travel in the free variant.
- What you hand over is used solely to produce the triage list.
If the list points at something
Triage sorts the signals, it does not explain them. When the list points at a real problem, the next step is a proper root-cause investigation: the reconstruction across all nodes, with timeline, gaps and an RCA document.
If the environment has never been measured end to end, that investigation is strongest combined with a baseline assessment through ClusterTriage Assurance: the same specialists, a read-only method, and a report built from measured state rather than a checklist.
Recognise this?
What Cluster Triage is for
These sentences come from findings written down in real engagements. If you recognise one, the triage list is the first step: it sorts the problem before you start searching.
Cluster and nodes
- Cluster is down
- Cluster node keeps failing
- Node not added to the cluster
- Cluster went down during patching
- Cluster went down twice last weekend
- Bluescreen on a cluster node
- No crash dump after an outage
- Event 1135: node removed from the cluster
- Cluster validation reports errors
Quorum and storage
- Quorum lost
- Witness unreachable
- Storage Spaces Direct volume degraded
- S2D repair job does not finish
- CSV went offline
- Cluster Shared Volume in redirected mode
- Event 5120: CSV I/O pause
- High latency on storage
- MPIO path lost
Network and migration
- Live migration fails
- Live migration is slow
- RDMA does not work
- Virtual machines unreachable after a restart
- VM freezes randomly
Azure Local and backup
- Azure Local update fails
- Solution update hangs
- Readiness check failed on Azure Local
- Backup fails on the cluster
- Veeam backup on the cluster is slow
If your situation is not listed, triage is still the safest first step: it only reads and changes nothing. What is needed after that emerges from the list.
Is your cluster down?
Do not wait until the evidence is gone.
Request the read-only triage script. You receive the five heaviest signals free of charge and know which next step makes sense.