Skyline

Alerts: know when your queues stop working

Twelve built-in checks with severities, recovery messages and a quiet period after deploys, over mail, Slack, SMS or a webhook.

Horizon ships one notification, LongWaitDetected, for a queue whose wait crosses a threshold, and it never says when the wait ended. Everything else a queue can do to you happens in silence: workers crash-looping, a supervisor killed for memory, jobs failing in bulk, a unique lock that blocks every dispatch of a job, and the one that costs the most, Horizon not running at all.

Skyline adds an alert pipeline with twelve built-in checks, a severity on every alert, and a recovery message when the condition clears. It runs inside your application against your own Redis, with no agent and no outside service, and it sends over the mail, Slack and SMS channels Horizon already knows plus a JSON webhook.

Turning it on#

Alerts are off until you enable them, and an alerts block with no enabled key counts as off. A config/horizon.php published before 1.5 needs no changes: the package supplies the whole block. Set the flag and at least one destination:

HORIZON_ALERTS=true
HORIZON_ALERT_SLACK="https://hooks.slack.com/services/..."
HORIZON_ALERT_MAIL="ops@example.com"
HORIZON_ALERT_SMS="+15555550100"
HORIZON_ALERT_WEBHOOK="https://events.example.com/queue-alerts"

If your application already calls Horizon::routeSlackNotificationsTo(), routeMailNotificationsTo() or routeSmsNotificationsTo() in its service provider, those destinations are used when the matching variable is empty, so a working long-wait setup starts receiving alerts without being moved. SMS goes through Laravel's Vonage notification channel, as Horizon's own SMS notification does.

Then schedule the one check that cannot run inside Horizon, wherever the rest of your scheduled commands run:

// routes/console.php
Schedule::command('horizon:check')->everyMinute();

Every other check is evaluated by a running supervisor, once a minute across the whole fleet. See Noticing that Horizon is not running for why this one is different.

What it watches#

Check Fires when Default Severity
horizon_down No master supervisor has reported in after 2 min critical
queue_wait The oldest job on a queue has waited longer than its threshold your waits thresholds (60s), after 1 min warning
queue_stalled Jobs are waiting and workers are assigned, but nothing on the queue finishes after 5 min critical
queue_not_draining The backlog kept growing across the window and the time to clear it is rising past its threshold 15 min to clear, over a 10 min window warning
queue_paused A master, supervisor or queue was paused and never resumed, including queue:pause --all after 15 min warning
failure_rate Jobs are failing in bulk on one queue, or on one job class 10 failures in 5 min critical
worker_crash_loop Workers are dying without reporting why 5 crashes in 5 min critical
stranded_lock A unique or overlap lock outlived its job, so every dispatch of that job is skipped after 5 min critical
misconfiguration A supervisor's timeout is not below its connection's retry_after, so long jobs run twice immediately, repeats daily warning
supervisor_out_of_memory A supervisor exceeded its memory limit immediately warning
master_out_of_memory The master process exceeded memory_limit immediately critical
process_launch_failed A worker process would not start immediately critical

Each check lives under horizon.alerts.checks and takes enabled, plus for, threshold and window where they apply. Every check also accepts severity and cooldown, which override the defaults above for that check alone.

Two checks lean on data another feature records. worker_crash_loop needs HORIZON_INSIGHTS=true, which is what records worker restarts in the first place (see Insights). stranded_lock needs lock_insights, which is on by default and feeds the Locks & Limits screen.

Counting failures per job class#

failure_rate counts per queue by default. One class failing every run on a queue shared by twenty others is diluted into nothing that way, so set 'per' => 'job' to count each job class on its own. 'mode' => 'percent' compares the share of finished jobs that failed instead of the count, and a class needs at least min_jobs finished jobs in the window before its percentage counts.

'failure_rate' => [
    'enabled' => true,
    'threshold' => 25,     // percent, because of the mode below
    'window' => 300,
    'per' => 'job',
    'mode' => 'percent',
    'min_jobs' => 20,
],

Moving from the long wait notification#

With alerts enabled, queue_wait takes over from Horizon's LongWaitDetected notification. It reads the same horizon.waits thresholds unless you give a queue its own under alerts.checks.queue_wait.queues, keyed connection:queue, where a zero turns the check off for that queue. The difference is in how it behaves: it holds before it pages, says when the wait cleared, stays quiet after a deploy and goes over every channel including the webhook.

Horizon's own long wait notification stands down while alerts are enabled, so you are not told about one queue twice. The LongWaitDetected event still fires for anything of yours that listens to it.

Why it does not page you at 4am#

  • Conditions have to hold. Each check has a for window. A backlog that clears inside five minutes is a burst being absorbed, which is what the queue is for.
  • It says it once. A still-firing alert repeats at most every cooldown seconds, 15 minutes by default, with "Still firing" in the subject. Set it to 0 to hear about it exactly once.
  • It tells you when it stopped. Every condition sends a recovery message saying how long it lasted, so a fixed problem and a muted one don't look the same.
  • Deploys are quiet. A deploy replaces every worker and leaves a backlog behind it, so notifications are withheld for suppress_after_deploy seconds (300) afterwards. The checks keep running and their timers keep counting, so anything the deploy really broke still announces itself when the window closes.

Rate checks such as failure_rate sample the counters Horizon already maintains rather than keeping a tally of their own. An evaluator that restarts or moves to another host loses a reading, not the count, and a counter that went down is read as horizon:clear-metrics rather than as a recovery.

Sending each severity somewhere different#

Every alert has a severity: critical means work is not getting done and someone should get up, warning can wait for office hours, and info is for the record. It opens the subject line and colours the Slack message:

[critical] Acme: The "default" queue has stopped moving
[critical] Still firing: Acme: The "default" queue has stopped moving
[critical] Resolved: Acme: The "default" queue has stopped moving

By default every alert goes to every configured channel. routes narrows that per severity, so only a critical alert sends the SMS:

'alerts' => [
    'routes' => [
        'critical' => ['slack', 'sms', 'webhook'],
        'warning' => ['slack'],
        'info' => ['slack'],
    ],
],

A severity you leave out still goes everywhere. A recovery goes to the same channels as the alert it recovers from. A route that names a channel with no destination is skipped, and horizon:alerts warns about it.

Noticing that Horizon is not running#

Every other alert is raised by a supervisor that is still looping, which is exactly what is missing when Horizon is down. So horizon_down runs only from horizon:check, outside the fleet, from the scheduler you set up above. Its two-minute for window covers the gap while a deploy replaces the master supervisor.

horizon:check also works as a plain health probe for a load balancer or an uptime monitor. It exits non-zero whenever a check is unhappy, whether or not it has been unhappy long enough to notify anyone yet.

A paused master is still running, so a paused fleet is reported by queue_paused from inside the fleet, per machine and with the backlog it is holding. horizon_down reports pauses only while queue_paused is turned off.

Checking your setup#

php artisan horizon:alert:test                      # send a test alert down every channel
php artisan horizon:alert:test --severity=critical  # test one severity's route
php artisan horizon:alerts                          # what is firing right now
php artisan horizon:alerts --history                # what was sent recently
php artisan horizon:alerts --clear

horizon:alert:test ignores the enabled flag, the sustain windows, the cooldown and deploy suppression on purpose. The point is to find out that the Slack webhook was revoked before an outage does. The history keeps the last 200 delivered alerts, set by alerts.history.

Sending alerts to PagerDuty, Opsgenie or your own service#

HORIZON_ALERT_WEBHOOK posts each alert, and each recovery, as JSON to any URL:

{
  "application": "Acme",
  "environment": "production",
  "state": "firing",
  "repeat": false,
  "seconds": 420,
  "alert": {
    "key": "queue_wait:redis:default",
    "check": "queue_wait",
    "severity": "warning",
    "title": "Acme: Jobs on the \"default\" queue are waiting too long",
    "summary": "The oldest job on the \"default\" queue has been waiting 7 minutes, over its threshold of 5 minutes. 4,120 jobs are waiting, estimated to clear in 12 minutes.",
    "context": {"Queue": "redis:default", "Oldest job": "7 minutes", "Threshold": "5 minutes", "Waiting": "4,120", "Workers": 6},
    "url": "https://acme.test/horizon/queues/default"
  }
}

state is firing or resolved, and alert.key stays the same across both, so an incident service can open and close one incident per condition. Add authentication headers under horizon.alerts.channels.webhook_headers; the request times out after webhook_timeout seconds (5).

To change how alerts are worded on every channel, bind your own notification class:

$this->app->bind(
    \Laravel\Horizon\Contracts\AlertNotification::class,
    \App\Notifications\QueueAlert::class
);

Alerting on something only you know about#

Thresholds cover the general cases. "The invoices queue should never sit idle on a weekday" is a rule only your application knows, so write it as a check and register it:

use Laravel\Horizon\Alerts\Alert;
use Laravel\Horizon\Alerts\Checks\Check;

class InvoiceQueueIdleCheck extends Check
{
    public function slug()
    {
        return 'invoice_queue_idle';
    }

    public function run()
    {
        if (! $this->isIdle()) {
            return [];
        }

        return [new Alert(
            'invoice_queue_idle',
            $this->slug(),
            Alert::CRITICAL,
            'No invoices have been processed today',
            'The invoices queue has run nothing since midnight, which has never happened on a weekday.'
        )];
    }

    protected function defaultSeverity()
    {
        return Alert::CRITICAL;
    }
}

// In a service provider:
Horizon::alertCheck(InvoiceQueueIdleCheck::class);

Return every alert your check owns on every call. Anything you stop returning is treated as recovered, which is what sends the resolve message. If you cannot read what you need, throw instead of returning an empty list: the manager then leaves your previous alerts alone rather than announcing a recovery it cannot vouch for. A custom check gets the same sustain window, cooldown, deploy quiet period and severity routing as the built-in ones.

Alerts and Prometheus#

If you already run Prometheus and Alertmanager, the Prometheus endpoint exports the queue lengths, wait times, failures and worker counts that several of these checks read, and you can keep alerting there. The built-in alerts are for the teams that don't run that stack, and for the conditions a scrape can't see well: a stranded lock, a worker that would not launch, or a supervisor whose timeout will run jobs twice.

Common questions

How do I get alerted when Laravel Horizon stops processing jobs?

Install Skyline, set HORIZON_ALERTS=true with a Slack, mail, SMS or webhook destination, and schedule php artisan horizon:check every minute. That command runs outside the fleet, so it notices when no master supervisor has reported in for two minutes, which a check running inside Horizon never could. The other checks, such as a stalled queue, a failure spike or crash-looping workers, run inside the supervisors on their own.

Does Skyline replace Horizon's LongWaitDetected notification?

Yes, while alerts are enabled. The queue_wait check reads the same horizon.waits thresholds, but it waits a minute before it pages, sends a recovery message when the wait clears, stays quiet for five minutes after a deploy and goes over every alert channel including the webhook. Horizon's own notification stands down so you are not told twice. The LongWaitDetected event still fires for your own listeners.

How does Skyline avoid alert fatigue?

Each condition has to hold for its for window before anyone is told, a still-firing alert repeats at most every 15 minutes, and every alert sends a recovery message saying how long it lasted. Notifications are withheld for five minutes after a deploy, while the checks keep running, so a problem the deploy really caused still reports when the window closes. Severity routing lets only critical alerts reach the SMS or the pager.

Can Skyline send queue alerts to PagerDuty or Opsgenie?

Yes. HORIZON_ALERT_WEBHOOK posts every alert and every recovery as JSON, with a stable key per condition and a state of firing or resolved, so an incident service can open and close one incident per problem. Headers for authentication go under horizon.alerts.channels.webhook_headers.

Queue control, not just queue monitoring.

Skyline is a drop-in replacement for Laravel Horizon that lets you act on what you see — pause a queue, jump a job to the front, drain a backlog. $99 once, every app you run it on.

Buy Skyline — $99 once

Secure checkout by Anystack. 30 days to change your mind.