Your AI Agent Has a Credential Expiry Date. Put It on a Calendar.
Most agent outages are not model problems. They are not bad prompts, rate limits, or a clever failure mode someone will write a conference talk about. They are a token that quietly expired on a Tuesday, and nobody noticed until Friday.
If you run self-hosted automation, you are holding a pile of credentials with different lifetimes: API keys, OAuth refresh tokens, deploy tokens, webhook secrets, bot tokens, app passwords. Each one has an expiry behavior. Some die on a date. Some die after inactivity. Some die when the person who created them changes a password. Almost none of them page you when they go.
This post is a practical routine for handling that: a credential expiry calendar, a cheap weekly probe, and a renewal procedure you can run half asleep.
Why expired credentials are the worst kind of failure
A crashed process is loud. A full disk is loud. An expired credential is often silent, because well-meaning code catches the 401 and moves on.
Here is what silent looks like in practice:
- A publishing job gets a 401 from your CMS, logs a warning, and exits zero. Your content calendar stops moving.
- A reporting agent fails to pull analytics, falls back to cached data, and sends a confident summary of last month.
- A calendar sync loses its OAuth refresh token and your morning briefing quietly omits every meeting.
- A deploy token expires, and the “daily deploy” job reports success because it only checked that the script ran.
Step 1: Build the inventory
You cannot track what you have not listed. Spend twenty minutes and write down every credential your agents and scripts use. A plain table is enough:
| Credential | Used by | Type | Expires | Owner | Renewal steps |
|---|---|---|---|---|---|
| Model provider key | Main agent | API key | No expiry, revocable | You | Dashboard, create key, update env |
| Hosting deploy token | Deploy script | Scoped token | 2027-01-15 | You | Dashboard, roll token, update env |
| Google OAuth | Calendar and indexing | Refresh token | 6 months idle | You | Re-run auth flow |
| Discord bot token | Chat gateway | Bot token | Until reset | You | Developer portal, reset |
| GitHub PAT | Git pushes | PAT | 90 days | You | Settings, regenerate |
Keep this in the repo that holds your runbooks, not in your head and not in a chat thread. Never put the secret values in it. The inventory records where a credential lives and how to replace it, not what it is.
Step 2: Classify by failure behavior
Not all expiry is the same, and the differences change how you monitor.
Hard-dated. A fixed expiry date, like a 90-day personal access token. These are the easiest to handle. Put the date on a calendar with a reminder two weeks ahead.
Idle-based. OAuth refresh tokens that die after a period of non-use, or sessions that lapse when a job stops running. These punish you for pausing a workflow. If you shut an agent down for a month, expect to re-authenticate it when you bring it back.
Event-based. Tokens that die when something else changes: a password reset, a revoked app grant, a rotated organization secret, an employee offboarding. You cannot calendar these. You can only detect them with a probe.
No expiry, but revocable. Many API keys never expire but can be revoked by the provider, by a leak scanner, or by you. Treat these as event-based.
Hard-dated credentials get calendar reminders. Everything else gets a probe.
Step 3: The weekly credential probe
A probe is a tiny, read-only call that proves a credential still works. It is not the real job. It is the cheapest authenticated request you can make against each service.
Good probes are boring:
- Model provider: list available models, or send a one-token request.
- Hosting platform: verify the token or list projects.
- Git remote:
git ls-remoteagainst the repo. - Chat bot: fetch the bot’s own user record.
- Google APIs: request a fresh access token from the refresh token.
Wrap them in one script that prints a single line per credential:
OK model-provider
OK hosting-token
OK git-remote
FAIL google-oauth (401 invalid_grant)
OK discord-bot
Run it weekly, and also before any high-stakes scheduled run. The rule that makes it useful: a FAIL must exit non-zero and notify a human. A probe that logs to a file nobody reads is decoration.
Step 4: Make the real jobs fail loudly
Treat every 401 or 403 from a dependency as a hard failure: non-zero exit, message naming the credential. And define success as the outcome, not the script run. A publishing job should verify the post exists; a deploy job should check the live URL.
Step 5: A renewal routine you can run tired
When a credential fails, you want a procedure, not an investigation. Write the same five steps for every entry in the inventory:
- Confirm it is the credential. Run the probe for that service alone.
- Create the replacement first. Generate the new key or token before revoking the old one, when the provider allows both to coexist.
- Update the single source. Credentials should live in one environment file or secret store, not copied across five scripts. One place to update means one place to get wrong.
- Restart only what reads it. Some processes load environment variables at startup. Know which ones do, and restart those.
- Re-run the probe, then the real job. Do not call it fixed until the actual workflow has completed once.
What to put on the calendar
Keep the calendar simple. Three recurring items cover most self-hosted setups:
- Weekly: run the credential probe and read the output.
- Monthly: skim the inventory for anything expiring within 30 days.
- Quarterly: audit for credentials that no longer have a job. Delete the ones nothing uses. Unused credentials still carry risk and still clutter the inventory.
The payoff
None of this is glamorous. That is the point. The agents that keep working month after month are rarely the cleverest ones. They are the ones whose operators know, in advance, which keys are about to die and what to do about it.
A credential expiry calendar costs an afternoon to build and a few minutes a week to maintain. It converts the most common silent failure in unattended automation into a scheduled, boring, five-minute chore. Do it before your agent goes quiet, not after you notice.
More from the build log
Suggested
Want the full MarketMai stack?
Get the core MarketMai guides and operator playbooks in one premium bundle for $49.
View Bundle