Rotate One Secret on Purpose: The AI Agent Key Rotation Drill
Most self-hosted agent setups share one quiet weakness: nobody knows what happens when a key changes.
You pasted an API key into an .env file months ago. A cron job reads it. A script copied it somewhere else. A skill cached it in a config you forgot about. The agent works, so you assume the wiring is clean. Then a key leaks, a provider forces a rotation, or a token expires on a Saturday night, and you discover how many places the old value was hiding.
The fix is not a policy document. It is a drill. Rotate one secret on purpose, while nothing is on fire, and write down what breaks.
Why rotation is a reliability problem, not just a security one
Security teams frame rotation as hygiene. For anyone running unattended automation, it is mostly a reliability question: can I replace this credential in under ten minutes without guessing?
If the answer is “I think so,” you do not have a rotation process. You have a hope.
Three things go wrong in practice:
- Duplicate copies. The same key lives in
.env, a shell profile, a systemd unit, and a script someone hardcoded during testing. You rotate one and the other three keep using the revoked value. - Silent failure. A job gets a 401, logs it somewhere nobody reads, and quietly stops doing its work. Nothing alerts. You notice three days later because the daily report did not arrive.
- Unknown consumers. You do not remember which jobs use the key at all. A webhook, a backup script, and an old experiment all turn out to depend on it.
A drill surfaces all three while the stakes are low.
Pick the right first secret
Do not start with your payment processor or your production database. Pick a secret that is:
- Low blast radius. A read-only analytics token, a search API key, a free-tier model provider key.
- Easy to reissue. The provider dashboard lets you create a new one in a minute.
- Used by more than one job. You want at least two or three consumers so the drill teaches you something.
A good first target is something like a web search or scraping API key. It is cheap, replaceable, and probably touched by several research or content jobs.
The 30-minute drill
Set a timer. The goal is a complete rotation with a written record, not perfection.
Step 1: Find every copy (10 minutes)
Before changing anything, search for the current value. Not the variable name, the value itself, using the first and last few characters so you do not paste the whole key into your terminal history.
Check these places:
- Your main env file and any per-project env files
- systemd units and timers, including user-level services
- Crontab entries and wrapper scripts
- Shell profiles and exported variables
- Agent workspace config, skill folders, and tool settings
- Backup directories and old checkouts, which often hold stale copies
Write each location in a note. The count is the first finding. If you expected one location and found five, you just learned why rotation would have hurt.
Step 2: List every consumer (5 minutes)
For each location, ask which job or agent reads it. Build a small table:
| Consumer | Where it reads the key | What it does if the key fails |
|---|---|---|
| Morning research cron | env file | Posts an error, or silently skips? |
| Content pipeline | systemd unit | Retries forever? |
| Old test script | hardcoded | Nobody knows |
The third column is the one that matters. “Nobody knows” is the finding.
Step 3: Rotate and replace (5 minutes)
Create the new key in the provider dashboard. Do not revoke the old one yet. Update every location from step 1, then restart whatever needs a restart: the gateway, the services, any long-running process that read the value at startup.
Keeping the old key alive briefly gives you a rollback path. If something breaks, you can revert one location instead of scrambling.
Step 4: Verify each consumer actually works (5 minutes)
This is where most people stop too early. A restart that comes up cleanly proves nothing about whether the job uses the new key.
For each consumer, trigger one real run or a safe dry run and check the output. Look for the provider’s usage page too: after you revoke the old key, does traffic show up under the new one?
Step 5: Revoke the old key and watch (5 minutes)
Revoke the old value. Then wait for the next scheduled runs of your consumers and check that none failed. Anything that still held the old key will now surface as an authentication error, which is exactly what you want to learn today rather than during an incident.
What the drill usually finds
Run this once and you will almost certainly hit at least one of these:
- A forgotten script with a hardcoded key
- A job that swallows authentication errors and reports success
- A service that never restarted and kept using the old value in memory
- A backup or archive folder containing the live secret in plain text
- Two jobs sharing one key that should have had separate ones
Each of these is a real defect. None of them show up until a key changes.
Turn the findings into three permanent fixes
A drill only pays off if it changes the system. Aim for these:
One canonical home per secret. Every secret lives in exactly one file, and everything else reads from it. No copies in scripts, no values in unit files. When you rotate, you edit one place.
Loud authentication failures. Any job that gets a 401 or 403 should fail visibly: a message in your alert channel, not a line in a log file. A rotated key should produce noise within one run, not silence for three days.
A secrets register. A plain list of each secret, its owner, which jobs use it, when it was last rotated, and where to reissue it. It does not need to be fancy. A markdown file in your ops notes is enough. The point is that the next rotation starts from a list instead of a grep.
Set a cadence that you will actually keep
Quarterly rotation of everything sounds responsible and never happens. Instead, rotate one secret per month, starting with low-risk ones. Rotate immediately after any suspected exposure, like a key pasted in a chat or committed to a repo.
The real payoff
The first drill is slightly annoying. The third is boring. Boring is the goal.
When a provider emails you about a leaked key, you want a procedure you have already run, not an archaeology project. Autonomy without recoverability is a liability, and being able to swap a credential without anything quietly dying is what makes it safe to leave agents running overnight.
Pick one low-risk secret this week. Set a timer for thirty minutes. Rotate it on purpose, and write down everything that surprised you. That list is your roadmap.
More from the build log
Suggested
Want the full MarketMai stack?
Get the core MarketMai guides and operator playbooks in one premium bundle for $49.
View Bundle