Horizon keeps its metrics as windows. Every horizon:snapshot reads them and resets them. An exporter
that reads those figures at scrape time gets a throughput graph that notches every five minutes and reads low,
and a counter built from them drops to zero so often that Prometheus thinks the app keeps restarting.
This exporter keeps its own counters instead. Workers count each outcome into Redis, the counts only ever go up,
and a web request serves them in Prometheus format. So rate() and increase() work the
way they do for any other service, and you can alert on them.
What it exports
Counters for processed, failed and retried jobs, both per queue and per job class. A summary of how long jobs waited before a worker picked them up. Gauges for how deep each queue is, how long its oldest ready job has been waiting, Horizon's time-to-clear estimate and how many workers each supervisor runs.
# average wait per queue over the last five minutes
sum by (queue) (rate(horizon_queue_wait_seconds_sum[5m]))
/ sum by (queue) (rate(horizon_queue_wait_seconds_count[5m]))
A Grafana dashboard ships with the package: an overview, then rows per queue, per job class and per supervisor.
Publish it with vendor:publish and import it onto any Prometheus datasource.
Labels that add up
Job classes are labelled job_class, because Prometheus reserves job for the scrape
config and would rename yours to exported_job. A supervisor that balances high,default
as one pool is exported as two queues, each carrying group="high,default", so the gauges line up
with the per-queue counters instead of arriving as one series named after the pool.
A retry is a job released after an exception. A job that calls $this->release() itself, to wait
on a rate limit or a lock, is not counted as one. Wait time is measured per attempt, from when the job last
became ready.
What it costs your workers
One Redis round trip per job outcome, and one more per pickup for the wait time. Scrapes never write, which is
why there is no jobs-per-minute gauge: reading Horizon's own figure can write a timestamp and shift the window
its dashboard measures. Use rate() on the processed counter instead.
Two things to know
- It has to be enabled on every machine, workers included. The web servers serve the endpoint, and the workers keep the counters it reads.
-
The endpoint sits outside Horizon's
viewHorizongate, since Prometheus has no session. An IP allowlist protects it, and it defaults to loopback only. Behind a proxy, configure trusted proxies or every request will look like it came from the proxy.
Scrape one web server, not all of them. Every instance reads the same Redis, so several targets return identical series. Why the other Horizon exporters read low is in our post on Horizon in Prometheus.
Every configuration key, and how to run the test suite, is in the README on GitHub. Bugs and questions go in its issues.