# Alerts: know when your queues stop working

> Twelve built-in checks with severities, recovery messages and a quiet period after deploys, over mail, Slack, SMS or a webhook.

Source: https://boring-observability.dev/skyline/docs/alerts
Section: Observability — Skyline for Laravel documentation
Updated: 2026-09-27

---

Horizon ships one notification, `LongWaitDetected`, for a queue whose wait crosses a threshold, and it never says when the wait ended. Everything else a queue can do to you happens in silence: workers crash-looping, a supervisor killed for memory, jobs failing in bulk, a unique lock that blocks every dispatch of a job, and the one that costs the most, Horizon not running at all.

Skyline adds an alert pipeline with twelve built-in checks, a severity on every alert, and a recovery message when the condition clears. It runs inside your application against your own Redis, with no agent and no outside service, and it sends over the mail, Slack and SMS channels Horizon already knows plus a JSON webhook.

## Turning it on

Alerts are off until you enable them, and an `alerts` block with no `enabled` key counts as off. A `config/horizon.php` published before 1.5 needs no changes: the package supplies the whole block. Set the flag and at least one destination:

```ini
HORIZON_ALERTS=true
HORIZON_ALERT_SLACK="https://hooks.slack.com/services/..."
HORIZON_ALERT_MAIL="ops@example.com"
HORIZON_ALERT_SMS="+15555550100"
HORIZON_ALERT_WEBHOOK="https://events.example.com/queue-alerts"
```

If your application already calls `Horizon::routeSlackNotificationsTo()`, `routeMailNotificationsTo()` or `routeSmsNotificationsTo()` in its service provider, those destinations are used when the matching variable is empty, so a working long-wait setup starts receiving alerts without being moved. SMS goes through Laravel's Vonage notification channel, as Horizon's own SMS notification does.

Then schedule the one check that cannot run inside Horizon, wherever the rest of your scheduled commands run:

```php
// routes/console.php
Schedule::command('horizon:check')->everyMinute();
```

Every other check is evaluated by a running supervisor, once a minute across the whole fleet. See [Noticing that Horizon is not running](#horizon-down) for why this one is different.

## What it watches

| Check | Fires when | Default | Severity |
| --- | --- | --- | --- |
| `horizon_down` | No master supervisor has reported in | after 2 min | critical |
| `queue_wait` | The oldest job on a queue has waited longer than its threshold | your `waits` thresholds (60s), after 1 min | warning |
| `queue_stalled` | Jobs are waiting and workers are assigned, but nothing on the queue finishes | after 5 min | critical |
| `queue_not_draining` | The backlog kept growing across the window and the time to clear it is rising past its threshold | 15 min to clear, over a 10 min window | warning |
| `queue_paused` | A master, supervisor or queue was paused and never resumed, including `queue:pause --all` | after 15 min | warning |
| `failure_rate` | Jobs are failing in bulk on one queue, or on one job class | 10 failures in 5 min | critical |
| `worker_crash_loop` | Workers are dying without reporting why | 5 crashes in 5 min | critical |
| `stranded_lock` | A unique or overlap lock outlived its job, so every dispatch of that job is skipped | after 5 min | critical |
| `misconfiguration` | A supervisor's `timeout` is not below its connection's `retry_after`, so long jobs run twice | immediately, repeats daily | warning |
| `supervisor_out_of_memory` | A supervisor exceeded its memory limit | immediately | warning |
| `master_out_of_memory` | The master process exceeded `memory_limit` | immediately | critical |
| `process_launch_failed` | A worker process would not start | immediately | critical |

Each check lives under `horizon.alerts.checks` and takes `enabled`, plus `for`, `threshold` and `window` where they apply. Every check also accepts `severity` and `cooldown`, which override the defaults above for that check alone.

Two checks lean on data another feature records. `worker_crash_loop` needs `HORIZON_INSIGHTS=true`, which is what records worker restarts in the first place (see [Insights](https://boring-observability.dev/skyline/docs/insights)). `stranded_lock` needs `lock_insights`, which is on by default and feeds the [Locks & Limits](https://boring-observability.dev/skyline/docs/locks-and-limits) screen.

### Counting failures per job class

`failure_rate` counts per queue by default. One class failing every run on a queue shared by twenty others is diluted into nothing that way, so set `'per' => 'job'` to count each job class on its own. `'mode' => 'percent'` compares the share of finished jobs that failed instead of the count, and a class needs at least `min_jobs` finished jobs in the window before its percentage counts.

```php
'failure_rate' => [
    'enabled' => true,
    'threshold' => 25,     // percent, because of the mode below
    'window' => 300,
    'per' => 'job',
    'mode' => 'percent',
    'min_jobs' => 20,
],
```

### Moving from the long wait notification

With alerts enabled, `queue_wait` takes over from Horizon's `LongWaitDetected` notification. It reads the same `horizon.waits` thresholds unless you give a queue its own under `alerts.checks.queue_wait.queues`, keyed `connection:queue`, where a zero turns the check off for that queue. The difference is in how it behaves: it holds before it pages, says when the wait cleared, stays quiet after a deploy and goes over every channel including the webhook.

Horizon's own long wait notification stands down while alerts are enabled, so you are not told about one queue twice. The `LongWaitDetected` event still fires for anything of yours that listens to it.

## Why it does not page you at 4am

- **Conditions have to hold.** Each check has a `for` window. A backlog that clears inside five minutes is a burst being absorbed, which is what the queue is for.
- **It says it once.** A still-firing alert repeats at most every `cooldown` seconds, 15 minutes by default, with "Still firing" in the subject. Set it to `0` to hear about it exactly once.
- **It tells you when it stopped.** Every condition sends a recovery message saying how long it lasted, so a fixed problem and a muted one don't look the same.
- **Deploys are quiet.** A deploy replaces every worker and leaves a backlog behind it, so notifications are withheld for `suppress_after_deploy` seconds (300) afterwards. The checks keep running and their timers keep counting, so anything the deploy really broke still announces itself when the window closes.

Rate checks such as `failure_rate` sample the counters Horizon already maintains rather than keeping a tally of their own. An evaluator that restarts or moves to another host loses a reading, not the count, and a counter that went down is read as `horizon:clear-metrics` rather than as a recovery.

## Sending each severity somewhere different

Every alert has a severity: `critical` means work is not getting done and someone should get up, `warning` can wait for office hours, and `info` is for the record. It opens the subject line and colours the Slack message:

```text
[critical] Acme: The "default" queue has stopped moving
[critical] Still firing: Acme: The "default" queue has stopped moving
[critical] Resolved: Acme: The "default" queue has stopped moving
```

By default every alert goes to every configured channel. `routes` narrows that per severity, so only a critical alert sends the SMS:

```php
'alerts' => [
    'routes' => [
        'critical' => ['slack', 'sms', 'webhook'],
        'warning' => ['slack'],
        'info' => ['slack'],
    ],
],
```

A severity you leave out still goes everywhere. A recovery goes to the same channels as the alert it recovers from. A route that names a channel with no destination is skipped, and `horizon:alerts` warns about it.

## Noticing that Horizon is not running

Every other alert is raised by a supervisor that is still looping, which is exactly what is missing when Horizon is down. So `horizon_down` runs only from `horizon:check`, outside the fleet, from the scheduler you set up [above](#enabling). Its two-minute `for` window covers the gap while a deploy replaces the master supervisor.

`horizon:check` also works as a plain health probe for a load balancer or an uptime monitor. It exits non-zero whenever a check is unhappy, whether or not it has been unhappy long enough to notify anyone yet.

A paused master is still running, so a paused fleet is reported by `queue_paused` from inside the fleet, per machine and with the backlog it is holding. `horizon_down` reports pauses only while `queue_paused` is turned off.

## Checking your setup

```bash
php artisan horizon:alert:test                      # send a test alert down every channel
php artisan horizon:alert:test --severity=critical  # test one severity's route
php artisan horizon:alerts                          # what is firing right now
php artisan horizon:alerts --history                # what was sent recently
php artisan horizon:alerts --clear
```

`horizon:alert:test` ignores the enabled flag, the sustain windows, the cooldown and deploy suppression on purpose. The point is to find out that the Slack webhook was revoked before an outage does. The history keeps the last 200 delivered alerts, set by `alerts.history`.

## Sending alerts to PagerDuty, Opsgenie or your own service

`HORIZON_ALERT_WEBHOOK` posts each alert, and each recovery, as JSON to any URL:

```json
{
  "application": "Acme",
  "environment": "production",
  "state": "firing",
  "repeat": false,
  "seconds": 420,
  "alert": {
    "key": "queue_wait:redis:default",
    "check": "queue_wait",
    "severity": "warning",
    "title": "Acme: Jobs on the \"default\" queue are waiting too long",
    "summary": "The oldest job on the \"default\" queue has been waiting 7 minutes, over its threshold of 5 minutes. 4,120 jobs are waiting, estimated to clear in 12 minutes.",
    "context": {"Queue": "redis:default", "Oldest job": "7 minutes", "Threshold": "5 minutes", "Waiting": "4,120", "Workers": 6},
    "url": "https://acme.test/horizon/queues/default"
  }
}
```

`state` is `firing` or `resolved`, and `alert.key` stays the same across both, so an incident service can open and close one incident per condition. Add authentication headers under `horizon.alerts.channels.webhook_headers`; the request times out after `webhook_timeout` seconds (5).

To change how alerts are worded on every channel, bind your own notification class:

```php
$this->app->bind(
    \Laravel\Horizon\Contracts\AlertNotification::class,
    \App\Notifications\QueueAlert::class
);
```

## Alerting on something only you know about

Thresholds cover the general cases. "The `invoices` queue should never sit idle on a weekday" is a rule only your application knows, so write it as a check and register it:

```php
use Laravel\Horizon\Alerts\Alert;
use Laravel\Horizon\Alerts\Checks\Check;

class InvoiceQueueIdleCheck extends Check
{
    public function slug()
    {
        return 'invoice_queue_idle';
    }

    public function run()
    {
        if (! $this->isIdle()) {
            return [];
        }

        return [new Alert(
            'invoice_queue_idle',
            $this->slug(),
            Alert::CRITICAL,
            'No invoices have been processed today',
            'The invoices queue has run nothing since midnight, which has never happened on a weekday.'
        )];
    }

    protected function defaultSeverity()
    {
        return Alert::CRITICAL;
    }
}

// In a service provider:
Horizon::alertCheck(InvoiceQueueIdleCheck::class);
```

Return every alert your check owns on every call. Anything you stop returning is treated as recovered, which is what sends the resolve message. If you cannot read what you need, throw instead of returning an empty list: the manager then leaves your previous alerts alone rather than announcing a recovery it cannot vouch for. A custom check gets the same sustain window, cooldown, deploy quiet period and severity routing as the built-in ones.

## Alerts and Prometheus

If you already run Prometheus and Alertmanager, the [Prometheus endpoint](https://boring-observability.dev/skyline/docs/prometheus-metrics) exports the queue lengths, wait times, failures and worker counts that several of these checks read, and you can keep alerting there. The built-in alerts are for the teams that don't run that stack, and for the conditions a scrape can't see well: a stranded lock, a worker that would not launch, or a supervisor whose timeout will run jobs twice.


## Common questions

### How do I get alerted when Laravel Horizon stops processing jobs?

Install Skyline, set HORIZON_ALERTS=true with a Slack, mail, SMS or webhook destination, and schedule php artisan horizon:check every minute. That command runs outside the fleet, so it notices when no master supervisor has reported in for two minutes, which a check running inside Horizon never could. The other checks, such as a stalled queue, a failure spike or crash-looping workers, run inside the supervisors on their own.

### Does Skyline replace Horizon's LongWaitDetected notification?

Yes, while alerts are enabled. The queue_wait check reads the same horizon.waits thresholds, but it waits a minute before it pages, sends a recovery message when the wait clears, stays quiet for five minutes after a deploy and goes over every alert channel including the webhook. Horizon's own notification stands down so you are not told twice. The LongWaitDetected event still fires for your own listeners.

### How does Skyline avoid alert fatigue?

Each condition has to hold for its for window before anyone is told, a still-firing alert repeats at most every 15 minutes, and every alert sends a recovery message saying how long it lasted. Notifications are withheld for five minutes after a deploy, while the checks keep running, so a problem the deploy really caused still reports when the window closes. Severity routing lets only critical alerts reach the SMS or the pager.

### Can Skyline send queue alerts to PagerDuty or Opsgenie?

Yes. HORIZON_ALERT_WEBHOOK posts every alert and every recovery as JSON, with a stable key per condition and a state of firing or resolved, so an incident service can open and close one incident per problem. Headers for authentication go under horizon.alerts.channels.webhook_headers.
