Prometheus metrics & Grafana
Scrape the measurements behind the dashboard, and import the bundled Grafana dashboard.
Skyline can export the measurements behind the dashboard's metric graphs in the Prometheus text exposition format, so the numbers you watch during an incident are also the numbers your alerts fire on. Throughput, failures, retries, runtime, wait time, queue length, worker processes and job states — labelled by queue and by job class.
A Grafana dashboard covering all of it ships with the package. There is nothing to build.
Enabling the endpoint#
The endpoint is off by default, and off means off: no route is registered at all until you turn it on, so an application that never opts in cannot leak its queue topology to anyone who guesses the URL.
HORIZON_PROMETHEUS_ENABLED=true
That is the whole setup. Metrics are then served at /horizon/prometheus — or
{HORIZON_PATH}/prometheus if you moved the dashboard, because the scrape endpoint follows it. Point
Prometheus at it:
scrape_configs:
- job_name: horizon
metrics_path: /horizon/prometheus
scrape_interval: 30s
static_configs:
- targets: ['127.0.0.1:8000']
Counters are folded forward by horizon:snapshot, so keep it on its usual every-minute schedule and
prefer a scrape interval at or below that. See
Metrics & trends for what the snapshot command
does and why the metrics page is empty without it.
Access control#
A Prometheus server has no session to authenticate with, so this endpoint cannot sit behind the
viewHorizon gate the way the rest of the
dashboard does. It is deliberately registered outside the dashboard's middleware group and guarded by an
IP allowlist instead — which, out of the box, admits nothing but loopback:
'prometheus' => [
'enabled' => env('HORIZON_PROMETHEUS_ENABLED', false),
// Defaults to "{horizon.path}/prometheus" when left empty.
'path' => env('HORIZON_PROMETHEUS_PATH'),
'domain' => env('HORIZON_PROMETHEUS_DOMAIN'),
'allowed_ips' => array_filter(array_map('trim', explode(
',', (string) env('HORIZON_PROMETHEUS_ALLOWED_IPS', '127.0.0.1,::1')
))),
'middleware' => [],
'prefix' => env('HORIZON_PROMETHEUS_PREFIX', 'horizon'),
],
Entries may be exact addresses or CIDR ranges, in either IP family, and the single entry * allows any
address:
HORIZON_PROMETHEUS_ALLOWED_IPS="127.0.0.1,::1,10.0.4.0/24"
Read it as "which addresses this application sees", not "which machines may scrape". The address checked is the one Laravel resolves for the request, so behind a load balancer or reverse proxy — including nginx in front of php-fpm, or Octane, on the same host — every request arrives from the proxy, and the loopback default then admits anyone who can reach that proxy.
Configure the application's trusted proxies so the real client address is what gets checked, and treat the
allowlist as one layer rather than the whole of your access control. Anything listed under
middleware runs after it, which is where a token check or a rate limiter belongs.
Setting enabled back to false removes the route entirely — there is no disabled-but-present
state to get wrong.
Exported metrics#
Throughput, failures and retries are counters; everything else is a gauge. Runtime and wait time are reported in seconds here, where the dashboard's own API reports milliseconds.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
horizon_job_processed_total |
counter | job_class | Jobs of this class processed |
horizon_job_failed_total |
counter | job_class | Jobs of this class that failed |
horizon_job_retried_total |
counter | job_class | Jobs of this class released back onto a queue |
horizon_job_runtime_seconds |
gauge | job_class | Moving average runtime |
horizon_job_wait_seconds |
gauge | job_class | Moving average wait before processing |
horizon_job_memory_peak_bytes † |
gauge | job_class | Worker's peak memory while running this class, worst run this window |
horizon_job_memory_growth_bytes_total † |
counter | job_class | Heap left behind after jobs of this class returned |
horizon_job_cpu_seconds † |
gauge | job_class | Moving average CPU time one job consumed |
horizon_queue_processed_total |
counter | queue | Jobs processed on this queue |
horizon_queue_failed_total |
counter | queue | Jobs that failed on this queue |
horizon_queue_retried_total |
counter | queue | Jobs released back onto this queue |
horizon_queue_runtime_seconds |
gauge | queue | Moving average runtime |
horizon_queue_wait_seconds |
gauge | queue | Moving average wait before processing |
horizon_queue_memory_peak_bytes † |
gauge | queue | Worker's peak memory while running jobs on this queue |
horizon_queue_memory_growth_bytes_total † |
counter | queue | Heap left behind after jobs on this queue returned |
horizon_queue_cpu_seconds † |
gauge | queue | Moving average CPU time one job consumed |
horizon_queue_length |
gauge | queue, connection, group | Jobs waiting on the queue |
horizon_queue_oldest_pending_seconds |
gauge | queue, connection, group | Age of the oldest waiting job |
horizon_queue_processes |
gauge | queue, connection, group | Worker processes able to pick this queue up — the whole shared pool, not a per-queue share |
horizon_queue_paused |
gauge | queue, connection, group | 1 while the queue is paused |
horizon_queue_time_to_clear_seconds |
gauge | queue, connection, group | Estimated time to drain the queue |
horizon_jobs |
gauge | status | Jobs retained per status — pending, reserved, delayed, completed, failed, silenced |
horizon_recent_jobs |
gauge | — | Jobs inside the recent retention window |
horizon_recent_failed_jobs |
gauge | — | Failed jobs inside the recent retention window |
horizon_supervisor_processes |
gauge | supervisor, connection, group | Processes per supervisor pool |
horizon_processes |
gauge | — | Total worker processes |
horizon_worker_restarts_total † |
counter | supervisor, connection, group, reason | Worker processes that died and were replaced |
horizon_supervisor_start_time_seconds |
gauge | supervisor | Unix time the supervisor started; subtract from time() for uptime |
horizon_master_supervisor_start_time_seconds |
gauge | master | Unix time the master started |
horizon_master_supervisor_paused |
gauge | master | 1 while a master supervisor is paused |
horizon_up |
gauge | — | 1 while a master supervisor is reporting in |
horizon_info |
gauge | name | Always 1; carries the Horizon name |
horizon_scrape_duration_seconds |
gauge | — | Time spent collecting the scrape |
† Exported only while insights are enabled
(HORIZON_INSIGHTS=true); the families disappear from the scrape entirely when they are not. Several
panels in the bundled Grafana dashboard read them and stay empty until insights are on. The two
start_time_seconds gauges are deliberately unmarked: uptime is process state rather than
instrumentation, so it is always exported.
The per-class label is job_class rather than job, because Prometheus reserves
job for the scrape config's own job_name and renames any exposed one to
exported_job — which would collapse every class in the export onto the target's name.
Queues that a supervisor balances as one pool are split apart in the export, one series per queue, each carrying a
group label naming the pool it shares. That has two consequences for queries. Collapse the pools with
max by (queue, connection) before summing queue length or time to clear, so a queue served by two
supervisors is not counted twice. And horizon_queue_processes is not additive across queues — the whole
pool can pick up every queue in it — so take the fleet total from horizon_processes instead of summing
that metric.
Change the horizon_ prefix with HORIZON_PROMETHEUS_PREFIX if it collides with something
else you scrape — and remember to change it in the Grafana queries too.
Why throughput is a counter#
Horizon's metric graphs are windowed: horizon:snapshot resets the counters every time it runs.
That is the opposite of what Prometheus wants, where a counter must only ever go up so that rate() can
tell the difference between a quiet minute and a restart.
So throughput, failures and retries are exported as true counters. Each snapshot folds the window it closes into a
running total, and a scrape reads that total plus the window still open. rate() and
increase() behave correctly across snapshots, and nothing is double counted — the total is read before
the open window on purpose, so a snapshot landing mid-scrape makes the figure briefly under-count, which
self-corrects, rather than over-count, which Prometheus would read as a counter reset.
Runtime and wait time stay moving averages, so they remain gauges.
There is no jobs-per-minute gauge#
Deliberately. Horizon derives that figure from the time since the last snapshot, and reading it writes — so a scrape would move the window the dashboard's own figure is measured against. Scraping this endpoint never writes to Redis. Compute the same thing from the counters instead:
sum(rate(horizon_queue_processed_total[5m])) * 60
Alerts worth having#
If you don't run Alertmanager, Skyline can send these itself. Its built-in alerts cover Horizon being down, a queue waiting too long or left paused, and failure rates, and deliver them to mail, Slack, SMS or a webhook. The rules below are for teams that already route everything through Prometheus.
The two failure modes that page you at 3am are a queue nobody is draining and a worker fleet that has quietly died:
groups:
- name: horizon
rules:
- alert: HorizonDown
expr: max(horizon_up) == 0
for: 2m
- alert: QueueBacklogAgeing
expr: horizon_queue_oldest_pending_seconds > 300
for: 5m
- alert: QueueFailureRate
expr: |
sum by (queue) (rate(horizon_queue_failed_total[5m]))
/ sum by (queue) (rate(horizon_queue_processed_total[5m])) > 0.05
for: 10m
- alert: QueueLeftPaused
expr: horizon_queue_paused == 1
for: 30m
horizon_queue_oldest_pending_seconds is the one to alert on rather than queue length: a queue of ten
thousand fast jobs is fine, and a queue of one job nobody has picked up in an hour is not.
QueueLeftPaused exists because pausing a queue is a deliberate act that somebody eventually forgets to
undo.
The Grafana dashboard#
Publish the bundled dashboard into your project:
php artisan vendor:publish --tag=horizon-grafana
That writes grafana/horizon-dashboard.json, plus a short README with a sample scrape config. In Grafana
choose Dashboards → New → Import, upload the file, and pick your Prometheus data source.
Or take it straight from here — this is the same file, byte for byte:
The Queue and Job class variables at the top filter every panel, so one dashboard
serves a fleet of queues rather than needing a copy per queue. Every rate panel uses $__rate_interval,
which means the graphs stay correct when you change the time range or the scrape interval.
| Row | Panels |
|---|---|
| Overview | Horizon up, jobs/min, failures/min, worker processes, pending jobs, recent failures, jobs by state, throughput vs. failures |
| Queues | Throughput, failures, releases, average runtime, average wait, queue length, oldest pending job, estimated time to clear, worker processes, paused queues |
| Job classes | Throughput, failures, releases, average runtime, average wait, and tables of the slowest and the most failing job classes |
| Supervisors | Processes by supervisor, paused master supervisors |
The dashboard's queries are written against the default horizon_ prefix. If you set
HORIZON_PROMETHEUS_PREFIX to something else, find and replace horizon_ in the JSON
before importing.