Bubble Unreachable Kubelet Alert to the very top

Question

Bubble Unreachable Kubelet Alert to the very top

michaelgugino opened this issue 5 years ago · 13 comments

If a kubelet goes unreachable, you get like 57 alerts of all different kinds firing. It's very difficult to get a sense of what's happening. And that is just for one random worker kubelet, not even a master.

There should be a separate section on the alerting dashboard for just these types of alerts.

After just doing an oc delete node, almost all of the alerts cleared (Obviously, this was just to test, don't just oc delete node, use the machine-api to replace a broken node).

michaelgugino commented 4 years ago

/reopen

Answer 1 · 2020-07-23T13:47:24.000Z

Another idea is to always show a simple status (pie chart?) of healthy nodes at the top of the alerting screen.

Answer 2 · 2020-10-30T01:29:44.000Z

Issues go stale after 90d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle stale.
Stale issues rot after an additional 30d of inactivity and eventually close.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle stale

Answer 3 · 2020-11-29T03:24:49.000Z

Stale issues rot after 30d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle rotten.
Rotten issues close after an additional 30d of inactivity.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle rotten
/remove-lifecycle stale

Answer 4 · 2020-12-29T05:14:49.000Z

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

Answer 5 · 2020-12-29T05:15:04.000Z

@openshift-bot: Closing this issue.

In response to this:

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository.

Answer 6 · 2021-02-09T23:49:50.000Z

@michaelgugino: Reopened this issue.

In response to this:

/reopen

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository.

Answer 7 · 2021-02-09T23:53:17.000Z

/remove-lifecycle rotten

Answer 8 · 2021-02-10T00:08:56.000Z

Idea:

We have alerts for all kinds of daemonsets. If an individual kubelet goes unreachable, you get like 37 warning alerts about various pods and daemonsets that can't find a home.

It would be nice if we could somehow reconcile on a per-node basis that one or more daemonsets are pending/not making it to the node, etc.

Answer 9 · 2021-05-11T01:23:35.000Z

Issues go stale after 90d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle stale.
Stale issues rot after an additional 30d of inactivity and eventually close.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle stale

Answer 10 · 2021-06-10T05:05:59.000Z

Stale issues rot after 30d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle rotten.
Rotten issues close after an additional 30d of inactivity.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle rotten
/remove-lifecycle stale

Answer 11 · 2021-07-10T07:51:28.000Z

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

Answer 12 · 2021-07-10T07:51:36.000Z

@openshift-bot: Closing this issue.

In response to this:

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository.