Events
Sixteen detectors, from connections running out to a redundant index, with the threshold each one uses and how to change it.
sailfish:detect runs at the end of every collection. Each detector opens an event when its condition
starts to hold, keeps that one event open per table, index or server while it does, and resolves it once it clears.
Whether an event is sent as it opens or in the daily digest is set per type, and so is its severity.
The server#
| Event | When | Severity | Sent |
|---|---|---|---|
Connections running outconnections_saturated |
Any connection refused at max_connections in the last hour, or 80% of them open at a collection. | critical | Immediately |
Row-lock surgerow_lock_surge |
Time waiting for InnoDB row locks over the last hour is at least 3× the usual for that hour, and at least a minute. | critical | Immediately |
Slow-query surgeslow_query_surge |
Statements over long_query_time in the last hour are at least 3× the usual for that hour, and at least 60. | warning | Immediately |
Error surgeerror_surge |
One MySQL error, such as a deadlock (1213) or a lock wait timeout (1205), raised at least 3× as often as usual for that hour, and at least 100 times. An error the server never raised before counts from none. | warning | Immediately |
History list growinghistory_list_growing |
InnoDB's history list is at least 100,000 and longer than an hour ago, which means a transaction was left open. Needs the PROCESS privilege. | warning | Immediately |
Buffer pool missesbuffer_pool_misses |
Over the last hour, under 99% of reads served from memory, with 100,000 pages or more read from disk. Quiet after a restart, while the pool warms up. | warning | Daily digest |
Temp tables on disktmp_disk_surge |
Internal temporary tables written to disk over the last hour are at least 3× the usual for that hour, and at least 1,000. | warning | Daily digest |
Tables and indexes#
| Event | When | Severity | Sent |
|---|---|---|---|
Table-scan surgescan_surge |
Rows scanned without an index over the last hour are at least 3× the average for that hour over the previous 7 days, and at least 100,000. | critical | Immediately |
Index went coldindex_cold |
Read in the 14 days before the last 7, and not fetched since. | warning | Daily digest |
Never usedindex_never_used |
No fetch in 30 days of observed history. Says how big the index is, what its writes cost and whether another index covers it. | warning | Daily digest |
Redundant indexindex_redundant |
Another index makes it redundant: a duplicate, a left prefix of a wider one, or on InnoDB one spelling out the primary key every index already ends with. The index page shows how to drop it. | warning | Daily digest |
Table growthtable_growth |
Grew at least 1 GiB in the last 7 days, and at least 2× what it gained in an average week over the 4 before. | warning | Daily digest |
Unindexed referenceunindexed_reference |
*_id columns no index starts with, on a table of at least 10,000 rows that was read by full scans in the last 7 days. A heuristic aimed at foreignId() without constrained(). | info | Daily digest |
No primary keyno_primary_key |
A table without one. | info | Daily digest |
Reclaimable spacetable_free_space |
An InnoDB table's tablespace is at least 25% free, and at least 1 GiB. | info | Daily digest |
Schema changeschema_change |
A table or index was added or dropped. | info | Daily digest |
Detectors that need history#
The surges, the cold and never-used index detectors, table growth and unindexed references compare against what Sailfish has recorded, so they stay quiet until the history covers their window. Until then the Events page says how many more days each one needs. A surge needs 7 days, a cold index 21, a never-used one 30 and table growth 35.
The PRIMARY key is left out of the cold and never-used detectors. A table's clustered index is read through every other index, whatever its own counters say.
How the surges are measured#
Every surge compares the last hour with the average of the same hour of the day over the previous 7 days, so a nightly batch job that always scans at 02:00 is its own baseline. It fires when the hour is at least 3× that average and above a floor, so a table going from 10 rows scanned to 40 stays quiet. Each takes the same keys:
'scan_surge' => [
'enabled' => true,
'history_days' => 7,
'ratio' => 3,
'min_rows_per_hour' => 100_000,
],
'row_lock_surge' => [
'enabled' => true,
'history_days' => 7,
'ratio' => 3,
'min_per_hour' => 60_000, // milliseconds waited
],
'tmp_disk_surge' => [ /* ... */ 'min_per_hour' => 1_000 ],
'slow_query_surge' => [ /* ... */ 'min_per_hour' => 60 ],
The error surge is measured per error number, and takes two more keys to narrow it, each a list of error numbers. An application that relies on duplicate keys (1062) for its upserts might ignore that one:
'error_surge' => [
'enabled' => true,
'history_days' => 7,
'ratio' => 3,
'min_per_hour' => 100,
'only' => [], // default: every error
'ignore' => [], // e.g. [1062]
],
Index thresholds#
'index_cold' => [
'enabled' => true,
'baseline_days' => 14, // read during these...
'quiet_days' => 7, // ...and not since
'min_baseline_reads' => 1,
],
'index_never_used' => [
'enabled' => true,
'days' => 30,
],
'index_redundant' => [
'enabled' => true,
'days' => 7, // window its writes are measured over
],
A never-used event says how big the index is, what its writes cost and whether another index covers it. A redundant index needs no history, since its definition is enough, and its page on the dashboard shows how to drop it.
Table thresholds#
'table_growth' => [
'enabled' => true,
'days' => 7,
'baseline_periods' => 4,
'ratio' => 2,
'min_bytes' => 1 << 30, // 1 GiB
],
'table_free_space' => [
'enabled' => true,
'min_share' => 0.25,
'min_bytes' => 1 << 30,
],
'unindexed_reference' => [
'enabled' => true,
'days' => 7,
'min_rows' => 10_000,
],
'schema_change' => [
'enabled' => true,
'open_hours' => 24, // informational: closes itself after a day
],
Server thresholds#
'connections_saturated' => [
'enabled' => true,
'ratio' => 0.8, // of max_connections open at a collection
],
'history_list_growing' => [
'enabled' => true,
'min_length' => 100_000,
],
'buffer_pool_misses' => [
'enabled' => true,
'hours' => 1,
'min_hit_ratio' => 0.99,
'min_disk_reads_per_hour' => 100_000,
],
Acknowledging and muting#
On the Events page, acknowledging an event stops its reminders. It still resolves, and a recovery is still sent, when its condition clears. Muting holds every message about it until the mute ends, or for good. A muted event stays muted if it clears and comes back.
To change what is sent for a type, rather than whether it is detected, see notifications.