Alerts: know when your queues stop working
Twelve built-in checks with severities, recovery messages and a quiet period after deploys, over mail, Slack, SMS or a webhook.
Horizon ships one notification, LongWaitDetected, for a queue whose wait crosses a threshold, and it
never says when the wait ended. Everything else a queue can do to you happens in silence: workers crash-looping, a
supervisor killed for memory, jobs failing in bulk, a unique lock that blocks every dispatch of a job, and the one
that costs the most, Horizon not running at all.
Skyline adds an alert pipeline with twelve built-in checks, a severity on every alert, and a recovery message when the condition clears. It runs inside your application against your own Redis, with no agent and no outside service, and it sends over the mail, Slack and SMS channels Horizon already knows plus a JSON webhook.
Turning it on#
Alerts are off until you enable them, and an alerts block with no enabled key counts as
off. A config/horizon.php published before 1.5 needs no changes: the package supplies the whole block.
Set the flag and at least one destination:
HORIZON_ALERTS=true
HORIZON_ALERT_SLACK="https://hooks.slack.com/services/..."
HORIZON_ALERT_MAIL="ops@example.com"
HORIZON_ALERT_SMS="+15555550100"
HORIZON_ALERT_WEBHOOK="https://events.example.com/queue-alerts"
If your application already calls Horizon::routeSlackNotificationsTo(),
routeMailNotificationsTo() or routeSmsNotificationsTo() in its service provider, those
destinations are used when the matching variable is empty, so a working long-wait setup starts receiving alerts
without being moved. SMS goes through Laravel's Vonage notification channel, as Horizon's own SMS notification does.
Then schedule the one check that cannot run inside Horizon, wherever the rest of your scheduled commands run:
// routes/console.php
Schedule::command('horizon:check')->everyMinute();
Every other check is evaluated by a running supervisor, once a minute across the whole fleet. See Noticing that Horizon is not running for why this one is different.
What it watches#
| Check | Fires when | Default | Severity |
|---|---|---|---|
horizon_down |
No master supervisor has reported in | after 2 min | critical |
queue_wait |
The oldest job on a queue has waited longer than its threshold | your waits thresholds (60s), after 1 min |
warning |
queue_stalled |
Jobs are waiting and workers are assigned, but nothing on the queue finishes | after 5 min | critical |
queue_not_draining |
The backlog kept growing across the window and the time to clear it is rising past its threshold | 15 min to clear, over a 10 min window | warning |
queue_paused |
A master, supervisor or queue was paused and never resumed, including queue:pause --all |
after 15 min | warning |
failure_rate |
Jobs are failing in bulk on one queue, or on one job class | 10 failures in 5 min | critical |
worker_crash_loop |
Workers are dying without reporting why | 5 crashes in 5 min | critical |
stranded_lock |
A unique or overlap lock outlived its job, so every dispatch of that job is skipped | after 5 min | critical |
misconfiguration |
A supervisor's timeout is not below its connection's retry_after, so long jobs run twice |
immediately, repeats daily | warning |
supervisor_out_of_memory |
A supervisor exceeded its memory limit | immediately | warning |
master_out_of_memory |
The master process exceeded memory_limit |
immediately | critical |
process_launch_failed |
A worker process would not start | immediately | critical |
Each check lives under horizon.alerts.checks and takes enabled, plus for,
threshold and window where they apply. Every check also accepts severity and
cooldown, which override the defaults above for that check alone.
Two checks lean on data another feature records. worker_crash_loop needs
HORIZON_INSIGHTS=true, which is what records worker restarts in the first place (see
Insights). stranded_lock needs
lock_insights, which is on by default and feeds the
Locks & Limits screen.
Counting failures per job class#
failure_rate counts per queue by default. One class failing every run on a queue shared by twenty
others is diluted into nothing that way, so set 'per' => 'job' to count each job class on its own.
'mode' => 'percent' compares the share of finished jobs that failed instead of the count, and a class
needs at least min_jobs finished jobs in the window before its percentage counts.
'failure_rate' => [
'enabled' => true,
'threshold' => 25, // percent, because of the mode below
'window' => 300,
'per' => 'job',
'mode' => 'percent',
'min_jobs' => 20,
],
Moving from the long wait notification#
With alerts enabled, queue_wait takes over from Horizon's LongWaitDetected notification. It
reads the same horizon.waits thresholds unless you give a queue its own under
alerts.checks.queue_wait.queues, keyed connection:queue, where a zero turns the check off
for that queue. The difference is in how it behaves: it holds before it pages, says when the wait cleared, stays quiet
after a deploy and goes over every channel including the webhook.
Horizon's own long wait notification stands down while alerts are enabled, so you are not told about one queue
twice. The LongWaitDetected event still fires for anything of yours that listens to it.
Why it does not page you at 4am#
- Conditions have to hold. Each check has a
forwindow. A backlog that clears inside five minutes is a burst being absorbed, which is what the queue is for. - It says it once. A still-firing alert repeats at most every
cooldownseconds, 15 minutes by default, with "Still firing" in the subject. Set it to0to hear about it exactly once. - It tells you when it stopped. Every condition sends a recovery message saying how long it lasted, so a fixed problem and a muted one don't look the same.
- Deploys are quiet. A deploy replaces every worker and leaves a backlog behind it, so
notifications are withheld for
suppress_after_deployseconds (300) afterwards. The checks keep running and their timers keep counting, so anything the deploy really broke still announces itself when the window closes.
Rate checks such as failure_rate sample the counters Horizon already maintains rather than keeping a
tally of their own. An evaluator that restarts or moves to another host loses a reading, not the count, and a
counter that went down is read as horizon:clear-metrics rather than as a recovery.
Sending each severity somewhere different#
Every alert has a severity: critical means work is not getting done and someone should get up,
warning can wait for office hours, and info is for the record. It opens the subject line and
colours the Slack message:
[critical] Acme: The "default" queue has stopped moving
[critical] Still firing: Acme: The "default" queue has stopped moving
[critical] Resolved: Acme: The "default" queue has stopped moving
By default every alert goes to every configured channel. routes narrows that per severity, so only a
critical alert sends the SMS:
'alerts' => [
'routes' => [
'critical' => ['slack', 'sms', 'webhook'],
'warning' => ['slack'],
'info' => ['slack'],
],
],
A severity you leave out still goes everywhere. A recovery goes to the same channels as the alert it recovers from.
A route that names a channel with no destination is skipped, and horizon:alerts warns about it.
Noticing that Horizon is not running#
Every other alert is raised by a supervisor that is still looping, which is exactly what is missing when Horizon is
down. So horizon_down runs only from horizon:check, outside the fleet, from the scheduler
you set up above. Its two-minute for window covers the gap while a deploy
replaces the master supervisor.
horizon:check also works as a plain health probe for a load balancer or an uptime monitor. It exits
non-zero whenever a check is unhappy, whether or not it has been unhappy long enough to notify anyone yet.
A paused master is still running, so a paused fleet is reported by queue_paused from inside the fleet,
per machine and with the backlog it is holding. horizon_down reports pauses only while
queue_paused is turned off.
Checking your setup#
php artisan horizon:alert:test # send a test alert down every channel
php artisan horizon:alert:test --severity=critical # test one severity's route
php artisan horizon:alerts # what is firing right now
php artisan horizon:alerts --history # what was sent recently
php artisan horizon:alerts --clear
horizon:alert:test ignores the enabled flag, the sustain windows, the cooldown and deploy suppression on
purpose. The point is to find out that the Slack webhook was revoked before an outage does. The history keeps the
last 200 delivered alerts, set by alerts.history.
Sending alerts to PagerDuty, Opsgenie or your own service#
HORIZON_ALERT_WEBHOOK posts each alert, and each recovery, as JSON to any URL:
{
"application": "Acme",
"environment": "production",
"state": "firing",
"repeat": false,
"seconds": 420,
"alert": {
"key": "queue_wait:redis:default",
"check": "queue_wait",
"severity": "warning",
"title": "Acme: Jobs on the \"default\" queue are waiting too long",
"summary": "The oldest job on the \"default\" queue has been waiting 7 minutes, over its threshold of 5 minutes. 4,120 jobs are waiting, estimated to clear in 12 minutes.",
"context": {"Queue": "redis:default", "Oldest job": "7 minutes", "Threshold": "5 minutes", "Waiting": "4,120", "Workers": 6},
"url": "https://acme.test/horizon/queues/default"
}
}
state is firing or resolved, and alert.key stays the same across
both, so an incident service can open and close one incident per condition. Add authentication headers under
horizon.alerts.channels.webhook_headers; the request times out after
webhook_timeout seconds (5).
To change how alerts are worded on every channel, bind your own notification class:
$this->app->bind(
\Laravel\Horizon\Contracts\AlertNotification::class,
\App\Notifications\QueueAlert::class
);
Alerting on something only you know about#
Thresholds cover the general cases. "The invoices queue should never sit idle on a weekday" is a rule
only your application knows, so write it as a check and register it:
use Laravel\Horizon\Alerts\Alert;
use Laravel\Horizon\Alerts\Checks\Check;
class InvoiceQueueIdleCheck extends Check
{
public function slug()
{
return 'invoice_queue_idle';
}
public function run()
{
if (! $this->isIdle()) {
return [];
}
return [new Alert(
'invoice_queue_idle',
$this->slug(),
Alert::CRITICAL,
'No invoices have been processed today',
'The invoices queue has run nothing since midnight, which has never happened on a weekday.'
)];
}
protected function defaultSeverity()
{
return Alert::CRITICAL;
}
}
// In a service provider:
Horizon::alertCheck(InvoiceQueueIdleCheck::class);
Return every alert your check owns on every call. Anything you stop returning is treated as recovered, which is what sends the resolve message. If you cannot read what you need, throw instead of returning an empty list: the manager then leaves your previous alerts alone rather than announcing a recovery it cannot vouch for. A custom check gets the same sustain window, cooldown, deploy quiet period and severity routing as the built-in ones.
Alerts and Prometheus#
If you already run Prometheus and Alertmanager, the Prometheus endpoint exports the queue lengths, wait times, failures and worker counts that several of these checks read, and you can keep alerting there. The built-in alerts are for the teams that don't run that stack, and for the conditions a scrape can't see well: a stranded lock, a worker that would not launch, or a supervisor whose timeout will run jobs twice.