# Reschedule builds on other agents rather than Fail builds when agents time out or are killed (machine shut down or put to sleep)

**URL:** <https://forum.buildkite.community/t/reschedule-builds-on-other-agents-rather-than-fail-builds-when-agents-time-out-or-are-killed-machine-shut-down-or-put-to-sleep/1388>\
**Category:** Features Requests\
**Created:** [December 15, 2020, 6:31pm UTC](https://forum.buildkite.community/t/reschedule-builds-on-other-agents-rather-than-fail-builds-when-agents-time-out-or-are-killed-machine-shut-down-or-put-to-sleep/1388 "2020-12-15T18:31:03Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![harisekhon](https://avatars.discourse-cdn.com/v4/letter/h/6de8d8/32.png) [@harisekhon](https://forum.buildkite.community/u/harisekhon)\
**Post date:** [December 15, 2020, 6:31pm UTC](https://forum.buildkite.community/t/reschedule-builds-on-other-agents-rather-than-fail-builds-when-agents-time-out-or-are-killed-machine-shut-down-or-put-to-sleep/1388/1 "2020-12-15T18:31:03Z")

</div>

Please can you implement build retries on other agents to handle when the BuildKite agent goes away due to the agent or machine being shut down or machine put to sleep.

In distributed processing systems it is common to retry a task 4 times on 4 different hosts before declaring a task as actually failed, because a lot of the time those failures are due to temporary issues or machines dying or in my case being shut down or put to sleep because I run BuildKite agents on my laptop.

This will also make BuildKite more suitable for use on Kubernetes where pods get evicted or on Cloud where preemptible instances can be killed with short notice, not enough time to wait for builds to finish cleanly and which will also result in false negative build failures and red failed badges on projects that shouldn’t happen but currently does (hence how I found out to raise this ticket).

---

<div class="post-metadata">

**Author:** ![moensch](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.buildkite.community/moensch/32/33_2.png) [@moensch](https://forum.buildkite.community/u/moensch)\
**Post date:** [December 16, 2020, 7:01pm UTC](https://forum.buildkite.community/t/reschedule-builds-on-other-agents-rather-than-fail-builds-when-agents-time-out-or-are-killed-machine-shut-down-or-put-to-sleep/1388/2 "2020-12-16T19:01:22Z")

</div>

Isn’t this already achieved by [https://buildkite.com/docs/pipelines/command-step#automatic-retry-attributes](https://buildkite.com/docs/pipelines/command-step#automatic-retry-attributes)?

---

<div class="post-metadata">

**Author:** ![Jason](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.buildkite.community/jason/32/742_2.png) [@Jason](https://forum.buildkite.community/u/Jason)\
**Post date:** [December 17, 2020, 12:41am UTC](https://forum.buildkite.community/t/reschedule-builds-on-other-agents-rather-than-fail-builds-when-agents-time-out-or-are-killed-machine-shut-down-or-put-to-sleep/1388/3 "2020-12-17T00:41:07Z")

</div>

@moensch suggestion is the way that we would recommend you handle these situations.

As in the example provided, you can utilise the exit statuses to check for particular failures:

> A job will fail with an exit status of -1 if communication with the agent has been lost (e.g. the agent has been forcefully terminated, or the agent machine was shut down without allowing the agent to disconnect). See the section on [Exit Codes](https://buildkite.com/docs/agent/v3#exit-codes) for information on other exit codes.

---

<div class="post-metadata">

**Author:** ![harisekhon](https://avatars.discourse-cdn.com/v4/letter/h/6de8d8/32.png) [@harisekhon](https://forum.buildkite.community/u/harisekhon)\
**Post date:** [December 17, 2020, 1:34pm UTC](https://forum.buildkite.community/t/reschedule-builds-on-other-agents-rather-than-fail-builds-when-agents-time-out-or-are-killed-machine-shut-down-or-put-to-sleep/1388/4 "2020-12-17T13:34:15Z")

</div>

@moensch / @Jason ah yes you are correct, thanks I didn’t see that.

I guess I’d change this request to make retries automatic when it is a buildkite/agent issue since I want my CI pipelines to reflect the state of the code being tested.

For now, is it possible to set this automatic exit status retry at the global or pipeline level rather than having to repeat this for every step in each pipeline, ballooning out my pipeline.yml? (I have 4 such steps in each pipeline so this is a lot of redundancy)

```
- command: make
  retry:
    automatic:
      - exit_status: -1 # Agent was lost
        limit: 2
      - exit_status: 255 # Forced agent shutdown
        limit: 2
```

---

<div class="post-metadata">

**Author:** ![Jason](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.buildkite.community/jason/32/742_2.png) [@Jason](https://forum.buildkite.community/u/Jason)\
**Post date:** [December 18, 2020, 1:55am UTC](https://forum.buildkite.community/t/reschedule-builds-on-other-agents-rather-than-fail-builds-when-agents-time-out-or-are-killed-machine-shut-down-or-put-to-sleep/1388/5 "2020-12-18T01:55:15Z")

</div>

Currently, it’s not possible to set it at a global level at the moment, sorry.

But as you could use YAML Anchors here:

```
anchors:
  std_retries: &std_retries
    retry:
      automatic:
        - exit_status: -1 # Agent was lost
          limit: 2
        - exit_status: 255 # Forced agent shutdown
          limit: 2
      
steps:
  - command: exit -1
    <<: [*std_retries]
  - command: exit 255
    <<: [*std_retries]
```

---

<div class="post-metadata">

**Author:** ![harisekhon](https://avatars.discourse-cdn.com/v4/letter/h/6de8d8/32.png) [@harisekhon](https://forum.buildkite.community/u/harisekhon)\
**Post date:** [December 19, 2020, 3:21pm UTC](https://forum.buildkite.community/t/reschedule-builds-on-other-agents-rather-than-fail-builds-when-agents-time-out-or-are-killed-machine-shut-down-or-put-to-sleep/1388/6 "2020-12-19T15:21:54Z")

</div>

Thanks, I’ll try that!
