Turn raw metrics into alerts that page someone only when it actually matters.
Start from a group rather than a single device, WAN routers, for example, so the rule scales as you add more of the same device type.
Require a threshold to hold for a few minutes, not just one bad poll, to avoid paging on a single transient blip.
Reserve Critical for conditions that need immediate action; use Warning for things worth knowing about but not worth an interrupt.
Point the rule at a channel, not just "all admins", so the right team sees it without digging through a shared firehose.