# Post mortem: failing checks us-west-1

> Particularly the AWS SNS in the us-west-1 region seems to have been the issue. Checks were retried but still failed.

Source: https://www.checklyhq.com/blog/post-mortem-failing-checks-us-west-1/

---

[Blog](https://www.checklyhq.com/blog/)

# Post mortem: failing checks us-west-1

[Tim Nolet](https://www.checklyhq.com/blog/author/tim-nolet/)[Umut Uzgur](https://www.checklyhq.com/blog/author/umut/)

April 29, 2020 · Updated October 17, 2024

On 28-04-2020 from ~02:00 AM - 02:30 AM CET Checks failed to schedule due to downstream problems. Particularly the AWS SNS in the us-west-1 region seems to have been the issue. Checks were retried but still failed. This reports the check as a failure for that region.

**Impact**
Around 200 checks failed to run and reported an error. This was across all customers. The error message was similar to

```
503: null
at Request.extractError (/var/task/node_modules/aws-sdk/lib/protocol/query.js:55:29)
at Request.callListeners
```

**Root Causes**
The root cause seems to be very slow responses and eventual 503 errors returning from calls to AWS SNS.

**Trigger**
No specific trigger.

**Resolution**
The issue resolved itself.

**Detection**
Our logging reported elevated errors in the alerting daemon because it was missing specific check data. We did not detect the error directly.

## What Are We Doing About This?

- Handle infrastructure errors differently from "user" errors, e.g. scripts with programming mistakes.
- Start alerting on elevated error rates for this scheduling logic to alert us on similar issues.

## Timeline

28-04-2020

- 02:04 AM - cron daemon starts retrying and failing to schedule checks
- 02:24 AM - cron daemon reverts to normal behaviour. Three customers notice strange behaviour and notify us through support channels.
- 05:58 AM - First employee in CET timezone sees emails and manually triggers alert.
- 06:05 AM - Primary does not acknowledge due to missing phone/voice alerting Secondary is paged.
- 06:30 AM - Analysis shows the issue has solved itself.
- 08:12 AM - Customers are notified by a retro active incident report.

[Tim Nolet Chief Evangelist](https://www.checklyhq.com/blog/author/tim-nolet/)

[Umut Uzgur Senior Engineering Team Lead](https://www.checklyhq.com/blog/author/umut/)

Share on social

## Related Articles

[Changelog: Extended Handlebars template functions & Groups API May 6, 2020](https://www.checklyhq.com/blog/changelog-template-variables-group-api/)[Post mortem: outage browser check results & alerting May 19, 2020](https://www.checklyhq.com/blog/post-mortem-outage-browser-check-results-alerting/)[How we monitor Checkly's API and Web App (updated) June 18, 2020](https://www.checklyhq.com/blog/how-we-monitor-checkly/)[Checkly Runtime 2021.10 with faker.js and updated Playwright November 23, 2021](https://www.checklyhq.com/blog/checkly-runtime-2021-10/)[How to run Checkly in your infrastructure - our new private locations May 23, 2022](https://www.checklyhq.com/blog/new-private-locations/)
