Your Monitoring Needs Monitoring Too
Infrastructure monitoring should be boring. A dashboard stays green, a few graphs move around, and maybe a certificate-renewal warning appears before it becomes a problem. The whole point is to find trouble before a customer does, not to create a new source of excitement.
There’s a weird point in infrastructure work where the thing that’s supposed to tell you something is broken becomes the thing that’s broken.
That’s when monitoring gets interesting.
But monitoring is infrastructure too. It has dependencies, assumptions, network paths, authentication, software updates, rate limits, alerting channels, and failure modes of its own. If you maintain systems long enough, eventually you have to troubleshoot the thing you built to help you troubleshoot everything else.
And that’s exactly what happened recently.
Infrastructure monitoring is part of the job, not a decoration
When you are responsible for infrastructure, “the site loads when I look at it” isn’t a monitoring strategy.
You need to know whether the web service is responding, whether the application is doing the right thing, whether the server underneath it is healthy, whether backups are happening, whether certificates are valid, and whether the outside services you depend on are having a bad day.
That responsibility has changed a lot since the early internet.
Back then, if you needed a service, there was a decent chance you rolled your own. Need mail? Build a mail server. Need DNS? Run DNS. Need a web server? Configure one. Need a monitoring tool? Find some open-source software, install it somewhere, and spend a weekend convincing it to send an email when another machine stopped answering.
That world was complicated, but it was also oddly straightforward. There were fewer layers between you and the thing that was broken. If the server stopped responding, you logged into the server. If DNS was wrong, you checked DNS. If Apache was angry, Apache told you why in a log file.
Cloud services have made that enormously easier.
They’ve also made it more complicated.
Today a perfectly ordinary website can depend on a hosting provider, a CDN, a DNS provider, a web application firewall, a SaaS ecommerce platform, an email provider, an identity provider, several APIs, and a handful of services you forgot were dependencies until one of them stopped working.
You maintain fewer physical boxes. You maintain more relationships between systems.
Technical discovery & auditing
The public page doesn’t tell you much about the machinery behind it. Raymond Tec audits inherited and long-running projects to uncover the plugins, integrations, data, dependencies, and old decisions that determine what the next change will really involve.
When the monitor says “down” and the site says “nope”
A recent problem with one of my monitoring checks is a good example.
A couple of HTTP checks started flipping between down and up. The public page checks were failing, recovering, failing again, and generally behaving like an unreliable smoke detector. Another lightweight endpoint on the same service stayed healthy.
The obvious first suspicion was the security layer in front of the site. That would’ve made sense. Automated checks look like automated traffic because, well, they are automated traffic.
Except there was no matching security event.
So I slowed the checks down. The heartbeat interval went from three minutes to five. Retries became less aggressive. That was worth doing anyway because hammering a service harder while it is rate limiting you is approximately the opposite of troubleshooting.
The alarms kept coming.
At that point the pattern mattered more than the individual failure. The HTML page checks were getting HTTP 429 responses, which means “too many requests,” while the simpler endpoint was not. The response headers also told the monitor to back off.
The site itself was fine. The monitoring request was the thing the platform no longer liked.
The hosted platform was now treating scripted requests to normal storefront pages differently. The fix was not a firewall exception or a bigger server. The monitor needed to identify and authenticate itself using the crawler-signing headers the platform expected.
I added those headers.
The problem disappeared immediately.
That’s a satisfying fix, but the useful part is everything that happened before it. One endpoint working while two others failed was evidence. No corresponding firewall event was evidence. A 429 response instead of a timeout was evidence. The retry header was evidence.
The monitor was not telling me, “the website is down.”
It was telling me, “this particular request is no longer welcome.”
Those are very different problems.
Monitoring the monitors
Once you accept that monitoring is infrastructure, the next question is obvious: who watches the watcher?
You don’t need an infinite philosophical loop here, but you should have at least one independent way to know whether your monitoring stack itself is alive.
That can be an external check, a dead-man switch that expects regular check-ins, or a second system watching the first. It can also be as simple as regularly testing the notification path instead of assuming the absence of alerts means everything is wonderful.
Because a dead monitoring server can produce the most beautiful dashboard in the world.
Nobody can see it, but technically nothing is red.
This is one reason I like free and open-source monitoring tools. Uptime Kuma is very good at answering the simple question: can I reach this thing, and did I get the response I expected? Beszel gives a lightweight view of what the hosts themselves are doing. Neither one needs to become a giant enterprise observability project just because I want to know whether a server is swapping itself to death at 2:00 in the morning.
More importantly, they are understandable. I can see what they are checking, change the request, adjust the interval, inspect the failure, and decide which dependency matters enough to monitor separately. That transparency matters when the monitor breaks.
Of course, self-hosting the monitoring tools means I also own them. Updates are my problem. Backups are my problem. Notification credentials are my problem. If the monitoring host runs out of disk space, that’s also, somewhat irritatingly, my problem.
Free software is not free maintenance.
Reliable automation
Automation is wonderful until it quietly stops working three Tuesdays ago. Raymond Tec builds integrations with logging, monitoring, and failure handling in mind so the boring work stays automated without becoming mysterious.
Cloud administration is simpler right up until it is not
The modern system administrator does less of some work and more of other work.
I don’t need to build a global CDN or write a DDoS mitigation system. I can turn those services on. I can rent compute, databases, storage, mail delivery, authentication, and dozens of other services from people whose entire business is operating them.
That’s fantastic.
But every abstraction hides machinery.
The CDN has caching rules. The firewall has bot rules. The SaaS platform has rate limits. The API has authentication requirements. The managed database has maintenance windows. The cloud provider has an outage page. The outage page can, occasionally, have an outage.
We traded a lot of low-level maintenance for dependency management.
I think that’s usually a good trade. But it changes what “knowing your infrastructure” means.
Knowing the server isn’t enough. You need to know the request path.
When a monitor fails, ask what actually failed. DNS? TLS? The TCP connection? The HTTP request? Authentication? A bot check? The application? A database query? The exact content assertion the monitor expected?
“Down” is a summary.
Troubleshooting starts when you stop treating the summary as the diagnosis.
Business IT goes well beyond the website
Your business also depends on workstations, cloud accounts, browsers, Wi-Fi, remote access, collaboration tools, and all the other technology that quietly becomes infrastructure. Raymond Tec works across that whole stack, whether the problem lives on a server, on a desk, or somewhere in between.
Maintenance includes the warning system
Monitoring needs maintenance for the same reason every other piece of infrastructure does: the world around it changes.
Endpoints move. Vendors change bot policies. Certificates rotate. APIs deprecate authentication methods. Alert destinations change. A check that made sense two years ago may now be noisy, redundant, or accidentally testing the wrong thing.
So I periodically look at the monitors themselves.
Are they checking something meaningful? Are the intervals sane? Are retries making a temporary hiccup look like a disaster? Is a monitor causing the rate limit it is complaining about? Does the alert actually reach somebody? Are there dependencies that can fail independently and deserve their own checks?
And, just as important, are there alerts I have mentally learned to ignore?
A monitor that cries wolf every day is not monitoring. It’s wallpaper.
The responsibility is not to create the maximum possible number of checks. The responsibility is to create checks you trust.
That means enough coverage to catch real failures, enough independence to know when the monitoring system is sick, and enough maintenance that an old assumption does not keep paging you for a problem that is not real.
The goal is boring again
The best infrastructure monitoring system is not the one with the fanciest dashboard.
It’s the one that quietly tells the truth.
When something breaks, it gives you useful evidence. When nothing is broken, it shuts up. When the monitoring system itself has a problem, there is another path that tells you about that too.
Then you fix it, test it, and go back to not thinking about it.
That is the part people sometimes miss about infrastructure work. Success is not a screen full of activity. Success is a system that does its job so consistently that nobody needs to think about it.
Boring is still underrated.
Keeping a website running is its own job. Raymond Tec provides website security and maintenance, including updates, backups, monitoring, access cleanup, recovery planning, and help when something has already gone sideways.
