It was supposed to be a routine 10-minute upgrade. Instead, 40 devices fell into a connect–disconnect loop, the automation we trusted kept making things worse, and a quiet office evening turned into the longest 85 minutes of the quarter.
This is the real story. All identifying details have been removed — but every technical detail is true.
The Setup: Automation That "Just Worked"
Every sysadmin has that one automation they're quietly proud of. Ours managed Wi-Fi access control.
The setup: every authorized device has its MAC address registered with a RADIUS server. When a device connects, RADIUS checks the list and decides where it belongs — corporate network, management network, or the guest network. A neat, scripted playbook read a CSV of authorized devices and pushed the list to the RADIUS server automatically. It had worked flawlessly for months.
There was just one problem, and it was hiding in plain sight:
We had two lists of truth.
- The firewall's GUI list — 14 devices, typed in by hand long ago.
- The automation-managed file — 54 devices. The real list. Updated by the playbook every time someone joined or left the company.
Here's the detail that would ruin our evening: the RADIUS package rebuilds its live configuration from the GUI list every time it's upgraded or reinstalled. Our automation had been writing directly to the live file — completely bypassing the GUI list.
The 14-entry GUI list had been stale for months. Nobody noticed. It never mattered... until the one night it mattered more than anything.
Lesson in the making: If two systems can edit the same data, you don't have redundancy. You have a time bomb with a countdown nobody can see.
18:22 — The Nightmare Begins
The firewall upgrade ran. Version to version, smooth as always. The RADIUS package updated alongside it — and with that update, the live file holding all 54 devices was wiped and rebuilt from the stale list of 14.
The results were immediate and brutal:
- ~40 devices could no longer authenticate. Access denied. No corporate Wi-Fi.
- Our two newest access points had also been added by automation instead of the GUI — so the upgrade erased them too. The RADIUS server now silently ignored every request coming from those access points. 315 authentication requests vanished into the void.
- The system never crashed. No alarm bells. No error page. It just quietly said "no" — and devices started churning.
Connect. Disconnect. Connect. Disconnect. For the people affected, it looked like the Wi-Fi itself was broken.
The worst outages don't announce themselves. They whisper "no" — and wait for you to notice.
Why Some People Were Completely Fine (and Why That Made It Worse)
This is the part that confused everyone — including us.
Some people worked perfectly all evening. "What's the problem?" they asked. "My Wi-Fi is fine!"
They weren't wrong. Wi-Fi MAC authentication is only checked at the moment a device connects. Devices that stayed connected never asked RADIUS for permission again. They sailed straight through the outage untouched.
So the office split into two worlds:
- Never disconnected? Totally fine. You'd have no idea anything was wrong.
- Woke a laptop from sleep, walked between rooms, or connected to the new access points? Stuck in an endless loop.
A device roaming between access points could get different behavior within minutes. Same office, same network, two completely different realities. From the outside it looked random. It wasn't — it was one broken rule colliding with a dozen everyday situations.
19:31 — When Automation Fought Back
The team spotted the problem and did what any sensible engineer would do: re-ran the automation to fix it.
And the automation failed. Again. And again. And again.
The restart routine used a classic pattern — a hard kill (kill -9) followed by starting the server again. But here's the trap: kill -9 doesn't let a process clean up after itself. The server's PID file stayed behind — like a "reserved" sign on a table where nobody is sitting.
The next start attempt saw the sign and refused to move:
Failed creating PID file /var/run/radiusd.pid: File exists
Every retry hit the same wall. The logs showed a storm of about ten short-lived processes in five minutes — our automation slamming into the same closed door, over and over, like a cartoon character who never learns.
The fix, once we found it, was one line: delete the stale PID file before starting. But finding that line at 7:30 PM, under pressure, with users asking for updates every two minutes, felt like defusing a bomb with oven mitts.
Automation doesn't get frustrated, doesn't improvise, and doesn't learn from failure. It repeats your mistake with perfect loyalty — forever.
19:48 — The Recovery
By 19:48, the team had:
- Restored the full 54-device list.
- Re-registered the two missing access points.
- Fixed the restart routine to clean up after itself before starting.
- Verified everything with live authentication tests — all passed.
Final score
| Metric | Result |
|---|---|
| Degraded window | ~85 minutes |
| Devices affected | ~40 of 54 |
| Requests silently ignored | 315 |
| Authentication rejects logged | 63 |
| Data lost | None |
| Evenings ruined | At least one |
Service was fully restored the same night. And then the real work began — making sure it could never happen again.
The 7 Lessons (This Is Why You're Here)
1. Two sources of truth is one source of disaster
Our automation and our GUI both claimed to own the device list. One survived upgrades; the other didn't. Nothing warned us they disagreed — until the upgrade rebuilt everything from the wrong one.
Do this: Pick ONE source of truth for every critical list. Everything else reads from it. Document it. If two systems can edit the same data, fix that today.
2. Automation amplifies whatever you give it — including your mistakes
Our playbook wasn't broken. It did exactly what it was told. The problem was it was told to write to a file the platform considered disposable. Wrong automation doesn't just fail — it fails confidently, repeatedly, and at scale.
Do this: Automate through supported interfaces, not around them. If a platform rebuilds a file from its own config, your automation must write to that config — never the artifact.
3. Silent failures are the worst failures
Those 315 ignored requests triggered no alert, no error page, no ticket. The access points simply got no answer, and devices retried forever. Silent rejection is a nightmare because there's nothing to Google.
Do this: Make failures loud. Alert on "zero responses" and "ignored requests" as seriously as you alert on errors.
4. kill -9 is not a restart strategy
Hard-killing a process leaves debris behind — a stale PID file, locked resources, corrupted state. A restart recipe that skips cleanup is a trap that fires exactly when you're already having a bad day.
Do this: Restart through the platform's service manager. If you must hard-kill, clean up the leftovers before starting again — and verify with retries.
5. You don't have a backup until you've restored from it
We had backups — taken automatically at the worst possible moment, of the worst possible file: a 2 KB snapshot of the truncated list. The real save was the automation's own source CSV.
Do this: Know which file is the master copy. Test restores. A backup of a broken file is just a broken file with a timestamp.
6. Every upgrade exposes hidden dependencies
The 14-entry list had been stale for months. The upgrade wasn't the bug — it was the event that finally made the hidden dependency visible. Every infrastructure has these landmines.
Do this: Treat every upgrade as an archaeology expedition. Before you touch anything: what depends on what, and which file gets rebuilt from where?
7. Monitor outcomes, not just services
"Service is running" was true for parts of the outage. Devices still couldn't connect. Uptime dashboards lied to us by telling the truth about the wrong thing.
Do this: Monitor the outcome your users care about — can devices authenticate and reach the network? — not just whether the daemon is alive.
The Boring Checklist That Would Have Saved Us
- One source of truth for every critical list — written down and enforced
- Automation writes through the platform, never around it
- Restart = clean shutdown → service-manager start → verify with retries
- Post-change verification for every dependent system, not just the one you upgraded
- Alerts for silent failures: ignored requests, zero-response patterns
- A written upgrade runbook, even if you're a one-person team
Print it. Tape it to the wall. Future-you will be grateful.
Final Thoughts: Respect the Machine
I love automation. When it works, it's magic — consistent, tireless, fast. But that night taught me something I won't forget:
Automation is an amplifier, not a brain.
It doesn't know your intentions, your history, or your hidden dependencies. It executes your assumptions with perfect loyalty — including the wrong ones. And when wrong automation meets a routine upgrade, a small mistake becomes an 85-minute nightmare.
The good news? Every trap in this story is avoidable with boring, unglamorous discipline. Single sources of truth. Loud failures. Tested restores. Clean restarts. Runbooks nobody reads until the night they save everyone.
Boring is what keeps the lights on.
After that night, I'll take boring every single time.
Have your own automation horror story? I'd love to hear it in the comments. And if you run MAC-based authentication anywhere in your network, do one thing for me today: go find your source of truth. If you can't point to it in ten seconds, you've just found your next project.
.jpg)

Seems like a nightmare of the night! Great story Saugat. Thanks for sharing!
ReplyDelete