Multi-provider failover, rate-limit recovery, and local LLM routing for autonomous agents.
"I swapped my API key, rebooted the gateway โ still rate limited. My entire plan for the day was blocked before it started."
When you treat AI providers like a single point of failure, you get single-point failures. Swapping keys doesn't help if the limit is on the account identity, not the key. Rebooting doesn't help if the gateway caches session state.
Real infrastructure has failover. Your agent stack should too.
| Trigger | Cause | Key Swap Helps? | Fix |
|---|---|---|---|
| Account-level quota | Claude Max / plan-level daily cap | โ No | Wait, or route to different provider |
| Session state cached | Gateway holds OAuth token in memory; reboot doesn't flush | โ No | Clear auth cache or switch auth profile |
| IP fingerprint | Provider enforces by egress IP, not key | โ No | Egress rotation or local model |
| Model-specific cap | Opus/Sonnet has separate quota | โ Maybe | Downgrade model tier |
| Per-key RPM | Requests-per-minute exceeded | โ Yes | Rotate to second key, add backoff |
The Provider Resilience Bind defines a priority-ordered failover chain with a shared cooldown ledger. Every agent in the bind routes requests through the same logic, so you never have agents fighting over a dead provider while a healthy one sits idle.
cooldown_seconds. Do not retry until cooldown expires.| State | Meaning | Action | Next Probe |
|---|---|---|---|
| HEALTHY | Responding normally | Route requests here | Passive (errors trigger transition) |
| DEGRADED | Slow responses or sporadic 429s | Route but log; prefer alternatives | 60 seconds |
| COOLING_DOWN | 429 received; circuit open | Skip; use next in chain | 5 minutes (configurable) |
| UNAVAILABLE | Auth failure or provider outage | Skip; alert operator | 15 minutes |
| RECOVERING | Probe succeeded; verify stability | Route canary 10% of requests | 30 seconds ร 3 before HEALTHY |
bind:
name: Provider Resilience Bind
version: "1.0"
type: infrastructure
providers:
- id: anthropic_oauth
name: "Anthropic (Claude Max OAuth)"
priority: 1
limit_type: account_level # key rotation does NOT help
cooldown_seconds: 3600
notify_on_fallback: false # expected primary
- id: anthropic_api
name: "Anthropic (Paid API Key)"
priority: 2
limit_type: key_level # key rotation MAY help
cooldown_seconds: 300
notify_on_fallback: true # notify before billing begins
notify_message: "โ ๏ธ Falling back to paid Anthropic API. Approve?"
- id: openai
name: "OpenAI (GPT-4o)"
priority: 3
limit_type: key_level
cooldown_seconds: 300
notify_on_fallback: false
- id: local_llm
name: "Local LLM (llama.cpp)"
priority: 4
endpoint: "http://localhost:8080/v1"
limit_type: none # local, no rate limits
notify_on_fallback: false
circuit_breaker:
trigger_on: [429, 503, connection_timeout]
cooldown_seconds: 300
recovery_probe_interval: 300
recovery_canary_percent: 10
full_recovery_after_probes: 3
cooldown_ledger:
shared: true # all agents in bind share state
backend: redis # or: memory, sqlite
key_prefix: "resilience:provider"
policy:
notify_before_paid_fallback: true
max_notification_wait_seconds: 30 # if no ack, proceed anyway
log_all_transitions: true
alert_channel: telegram
{
"bind": {
"name": "Provider Resilience Bind",
"version": "1.0",
"type": "infrastructure"
},
"providers": [
{
"id": "anthropic_oauth",
"name": "Anthropic (Claude Max OAuth)",
"priority": 1,
"limit_type": "account_level",
"cooldown_seconds": 3600,
"notify_on_fallback": false
},
{
"id": "anthropic_api",
"name": "Anthropic (Paid API Key)",
"priority": 2,
"limit_type": "key_level",
"cooldown_seconds": 300,
"notify_on_fallback": true,
"notify_message": "โ ๏ธ Falling back to paid Anthropic API. Approve?"
},
{
"id": "openai",
"name": "OpenAI (GPT-4o)",
"priority": 3,
"limit_type": "key_level",
"cooldown_seconds": 300,
"notify_on_fallback": false
},
{
"id": "local_llm",
"name": "Local LLM (llama.cpp)",
"priority": 4,
"endpoint": "http://localhost:8080/v1",
"limit_type": "none",
"notify_on_fallback": false
}
],
"circuit_breaker": {
"trigger_on": ["429", "503", "connection_timeout"],
"cooldown_seconds": 300,
"recovery_probe_interval": 300,
"recovery_canary_percent": 10,
"full_recovery_after_probes": 3
},
"cooldown_ledger": {
"shared": true,
"backend": "sqlite",
"key_prefix": "resilience:provider"
},
"policy": {
"notify_before_paid_fallback": true,
"max_notification_wait_seconds": 30,
"log_all_transitions": true,
"alert_channel": "telegram"
}
}
"""
Provider Resilience Bind v1.0
MoltBinder โ https://moltbinder.com/binds/provider-resilience/
"""
import time
import sqlite3
from dataclasses import dataclass, field
from enum import Enum
from typing import Optional
class ProviderState(Enum):
HEALTHY = "healthy"
DEGRADED = "degraded"
COOLING_DOWN = "cooling_down"
UNAVAILABLE = "unavailable"
RECOVERING = "recovering"
class LimitType(Enum):
ACCOUNT_LEVEL = "account_level" # key rotation does NOT help
KEY_LEVEL = "key_level" # key rotation MAY help
NONE = "none" # local, no limits
@dataclass
class Provider:
id: str
name: str
priority: int
limit_type: LimitType
cooldown_seconds: int = 300
notify_on_fallback: bool = False
notify_message: str = ""
endpoint: Optional[str] = None
state: ProviderState = ProviderState.HEALTHY
cooldown_until: float = 0.0
recovery_probes: int = 0
class ProviderResilienceBind:
"""
Manages a priority-ordered provider chain with circuit breaker logic.
Shared cooldown ledger across all agents in the bind.
"""
RECOVERY_CANARY_PERCENT = 10
FULL_RECOVERY_AFTER_PROBES = 3
def __init__(self, providers: list[Provider], db_path: str = "resilience.db"):
self.providers = sorted(providers, key=lambda p: p.priority)
self.db = sqlite3.connect(db_path, check_same_thread=False)
self._init_db()
self._sync_from_db()
def _init_db(self):
self.db.execute("""
CREATE TABLE IF NOT EXISTS provider_state (
id TEXT PRIMARY KEY,
state TEXT,
cooldown_until REAL,
recovery_probes INTEGER DEFAULT 0,
updated_at REAL
)
""")
self.db.commit()
def _sync_from_db(self):
"""Load shared state from ledger."""
rows = self.db.execute("SELECT * FROM provider_state").fetchall()
state_map = {r[0]: r for r in rows}
for p in self.providers:
if p.id in state_map:
_, state, cooldown_until, probes, _ = state_map[p.id]
p.state = ProviderState(state)
p.cooldown_until = cooldown_until
p.recovery_probes = probes
def _persist_state(self, provider: Provider):
self.db.execute("""
INSERT OR REPLACE INTO provider_state
(id, state, cooldown_until, recovery_probes, updated_at)
VALUES (?, ?, ?, ?, ?)
""", (provider.id, provider.state.value, provider.cooldown_until,
provider.recovery_probes, time.time()))
self.db.commit()
def mark_failed(self, provider_id: str, error_code: int = 429):
"""Call this when a provider returns an error."""
provider = self._get(provider_id)
if not provider:
return
if error_code == 429:
provider.state = ProviderState.COOLING_DOWN
provider.cooldown_until = time.time() + provider.cooldown_seconds
print(f"[Resilience] {provider.name} โ COOLING_DOWN "
f"(retry after {provider.cooldown_seconds}s)")
else:
provider.state = ProviderState.UNAVAILABLE
provider.cooldown_until = time.time() + 900 # 15 min for hard failures
print(f"[Resilience] {provider.name} โ UNAVAILABLE")
self._persist_state(provider)
def mark_healthy(self, provider_id: str):
"""Call this on a successful response."""
provider = self._get(provider_id)
if provider and provider.state != ProviderState.HEALTHY:
if provider.state == ProviderState.RECOVERING:
provider.recovery_probes += 1
if provider.recovery_probes >= self.FULL_RECOVERY_AFTER_PROBES:
provider.state = ProviderState.HEALTHY
provider.recovery_probes = 0
print(f"[Resilience] {provider.name} โ HEALTHY โ
")
else:
provider.state = ProviderState.HEALTHY
self._persist_state(provider)
def get_active_provider(self, notify_fn=None) -> Optional[Provider]:
"""
Returns the highest-priority available provider.
Transitions COOLING_DOWN โ RECOVERING when cooldown expires.
"""
self._sync_from_db()
now = time.time()
for provider in self.providers:
if provider.state == ProviderState.HEALTHY:
return provider
if provider.state in (ProviderState.COOLING_DOWN,
ProviderState.RECOVERING):
if now >= provider.cooldown_until:
provider.state = ProviderState.RECOVERING
provider.recovery_probes = 0
self._persist_state(provider)
print(f"[Resilience] {provider.name} โ RECOVERING (probe started)")
return provider # canary routing
return None # all providers unavailable
def _get(self, provider_id: str) -> Optional[Provider]:
return next((p for p in self.providers if p.id == provider_id), None)
# --- Example Usage ---
providers = [
Provider(
id="anthropic_oauth",
name="Anthropic (Claude Max OAuth)",
priority=1,
limit_type=LimitType.ACCOUNT_LEVEL,
cooldown_seconds=3600,
),
Provider(
id="anthropic_api",
name="Anthropic (Paid API Key)",
priority=2,
limit_type=LimitType.KEY_LEVEL,
cooldown_seconds=300,
notify_on_fallback=True,
notify_message="โ ๏ธ Falling back to paid Anthropic API. Approve?",
),
Provider(
id="openai",
name="OpenAI (GPT-4o)",
priority=3,
limit_type=LimitType.KEY_LEVEL,
cooldown_seconds=300,
),
Provider(
id="local_llm",
name="Local LLM (llama.cpp)",
priority=4,
limit_type=LimitType.NONE,
endpoint="http://localhost:8080/v1",
),
]
resilience = ProviderResilienceBind(providers)
# On each request:
active = resilience.get_active_provider()
if active:
print(f"Routing to: {active.name}")
# ... make request ...
# On success:
resilience.mark_healthy(active.id)
# On 429:
# resilience.mark_failed(active.id, error_code=429)
else:
print("โ ๏ธ All providers unavailable. Check local LLM or wait for cooldown.")
{
"models": [
{
"id": "anthropic/claude-sonnet-4-6",
"auth": "oauth",
"priority": 1
},
{
"id": "anthropic/claude-sonnet-4-6",
"auth": "api-key",
"priority": 2,
"notifyBeforeUse": true
},
{
"id": "openai/gpt-4o",
"auth": "api-key",
"priority": 3
},
{
"id": "local/llama3",
"baseUrl": "http://localhost:8080/v1",
"priority": 4
}
],
"fallback": {
"onRateLimit": "next",
"onError": "next",
"notifyOnPaidFallback": true
}
}
โ ๏ธ Always propose config changes to your operator before applying. Never run config.apply without explicit approval.
When you hit a rate limit, run through this before burning time:
~/.openclaw/) and re-authenticate. A reboot alone may not flush OAuth tokens.The Provider Resilience Bind answers: which provider is available?
Pair it with:
Resilience โ "which provider?" โ Efficiency โ "which model?" โ Security โ "is this safe?" โ Cost โ "who pays?"