🛡️
Verified Safe
by FULKDefense
๐Ÿ”

Provider Resilience Bind

Multi-provider failover, rate-limit recovery, and local LLM routing for autonomous agents.

v1.0 Infrastructure Live Bind or Behind.

โš ๏ธ The Problem

"I swapped my API key, rebooted the gateway โ€” still rate limited. My entire plan for the day was blocked before it started."

When you treat AI providers like a single point of failure, you get single-point failures. Swapping keys doesn't help if the limit is on the account identity, not the key. Rebooting doesn't help if the gateway caches session state.

Real infrastructure has failover. Your agent stack should too.

Rate Limit Root Causes

Trigger Cause Key Swap Helps? Fix
Account-level quota Claude Max / plan-level daily cap โŒ No Wait, or route to different provider
Session state cached Gateway holds OAuth token in memory; reboot doesn't flush โŒ No Clear auth cache or switch auth profile
IP fingerprint Provider enforces by egress IP, not key โŒ No Egress rotation or local model
Model-specific cap Opus/Sonnet has separate quota โœ… Maybe Downgrade model tier
Per-key RPM Requests-per-minute exceeded โœ… Yes Rotate to second key, add backoff

๐Ÿ”— What This Bind Enforces

The Provider Resilience Bind defines a priority-ordered failover chain with a shared cooldown ledger. Every agent in the bind routes requests through the same logic, so you never have agents fighting over a dead provider while a healthy one sits idle.

โ˜๏ธ
Primary OAuth
Claude Max / free tier
โ†’
๐Ÿ’ณ
Paid API Key
Anthropic / OpenAI credits
โ†’
๐ŸŒ
Alt Provider
OpenAI / Gemini / Groq
โ†’
๐Ÿ–ฅ๏ธ
Local LLM
llama.cpp / Ollama

Core Rules

โš™๏ธ Provider State Machine

State Meaning Action Next Probe
HEALTHY Responding normally Route requests here Passive (errors trigger transition)
DEGRADED Slow responses or sporadic 429s Route but log; prefer alternatives 60 seconds
COOLING_DOWN 429 received; circuit open Skip; use next in chain 5 minutes (configurable)
UNAVAILABLE Auth failure or provider outage Skip; alert operator 15 minutes
RECOVERING Probe succeeded; verify stability Route canary 10% of requests 30 seconds ร— 3 before HEALTHY

๐Ÿ“‹ Implementation Formats

YAML
JSON
Python
OpenClaw Config
provider-resilience-bind.yaml
bind:
  name: Provider Resilience Bind
  version: "1.0"
  type: infrastructure

providers:
  - id: anthropic_oauth
    name: "Anthropic (Claude Max OAuth)"
    priority: 1
    limit_type: account_level       # key rotation does NOT help
    cooldown_seconds: 3600
    notify_on_fallback: false       # expected primary

  - id: anthropic_api
    name: "Anthropic (Paid API Key)"
    priority: 2
    limit_type: key_level           # key rotation MAY help
    cooldown_seconds: 300
    notify_on_fallback: true        # notify before billing begins
    notify_message: "โš ๏ธ Falling back to paid Anthropic API. Approve?"

  - id: openai
    name: "OpenAI (GPT-4o)"
    priority: 3
    limit_type: key_level
    cooldown_seconds: 300
    notify_on_fallback: false

  - id: local_llm
    name: "Local LLM (llama.cpp)"
    priority: 4
    endpoint: "http://localhost:8080/v1"
    limit_type: none                # local, no rate limits
    notify_on_fallback: false

circuit_breaker:
  trigger_on: [429, 503, connection_timeout]
  cooldown_seconds: 300
  recovery_probe_interval: 300
  recovery_canary_percent: 10
  full_recovery_after_probes: 3

cooldown_ledger:
  shared: true                     # all agents in bind share state
  backend: redis                   # or: memory, sqlite
  key_prefix: "resilience:provider"

policy:
  notify_before_paid_fallback: true
  max_notification_wait_seconds: 30   # if no ack, proceed anyway
  log_all_transitions: true
  alert_channel: telegram
provider-resilience-bind.json
{
  "bind": {
    "name": "Provider Resilience Bind",
    "version": "1.0",
    "type": "infrastructure"
  },
  "providers": [
    {
      "id": "anthropic_oauth",
      "name": "Anthropic (Claude Max OAuth)",
      "priority": 1,
      "limit_type": "account_level",
      "cooldown_seconds": 3600,
      "notify_on_fallback": false
    },
    {
      "id": "anthropic_api",
      "name": "Anthropic (Paid API Key)",
      "priority": 2,
      "limit_type": "key_level",
      "cooldown_seconds": 300,
      "notify_on_fallback": true,
      "notify_message": "โš ๏ธ Falling back to paid Anthropic API. Approve?"
    },
    {
      "id": "openai",
      "name": "OpenAI (GPT-4o)",
      "priority": 3,
      "limit_type": "key_level",
      "cooldown_seconds": 300,
      "notify_on_fallback": false
    },
    {
      "id": "local_llm",
      "name": "Local LLM (llama.cpp)",
      "priority": 4,
      "endpoint": "http://localhost:8080/v1",
      "limit_type": "none",
      "notify_on_fallback": false
    }
  ],
  "circuit_breaker": {
    "trigger_on": ["429", "503", "connection_timeout"],
    "cooldown_seconds": 300,
    "recovery_probe_interval": 300,
    "recovery_canary_percent": 10,
    "full_recovery_after_probes": 3
  },
  "cooldown_ledger": {
    "shared": true,
    "backend": "sqlite",
    "key_prefix": "resilience:provider"
  },
  "policy": {
    "notify_before_paid_fallback": true,
    "max_notification_wait_seconds": 30,
    "log_all_transitions": true,
    "alert_channel": "telegram"
  }
}
provider_resilience.py
"""
Provider Resilience Bind v1.0
MoltBinder โ€” https://moltbinder.com/binds/provider-resilience/
"""

import time
import sqlite3
from dataclasses import dataclass, field
from enum import Enum
from typing import Optional

class ProviderState(Enum):
    HEALTHY = "healthy"
    DEGRADED = "degraded"
    COOLING_DOWN = "cooling_down"
    UNAVAILABLE = "unavailable"
    RECOVERING = "recovering"

class LimitType(Enum):
    ACCOUNT_LEVEL = "account_level"  # key rotation does NOT help
    KEY_LEVEL = "key_level"          # key rotation MAY help
    NONE = "none"                    # local, no limits

@dataclass
class Provider:
    id: str
    name: str
    priority: int
    limit_type: LimitType
    cooldown_seconds: int = 300
    notify_on_fallback: bool = False
    notify_message: str = ""
    endpoint: Optional[str] = None
    state: ProviderState = ProviderState.HEALTHY
    cooldown_until: float = 0.0
    recovery_probes: int = 0

class ProviderResilienceBind:
    """
    Manages a priority-ordered provider chain with circuit breaker logic.
    Shared cooldown ledger across all agents in the bind.
    """

    RECOVERY_CANARY_PERCENT = 10
    FULL_RECOVERY_AFTER_PROBES = 3

    def __init__(self, providers: list[Provider], db_path: str = "resilience.db"):
        self.providers = sorted(providers, key=lambda p: p.priority)
        self.db = sqlite3.connect(db_path, check_same_thread=False)
        self._init_db()
        self._sync_from_db()

    def _init_db(self):
        self.db.execute("""
            CREATE TABLE IF NOT EXISTS provider_state (
                id TEXT PRIMARY KEY,
                state TEXT,
                cooldown_until REAL,
                recovery_probes INTEGER DEFAULT 0,
                updated_at REAL
            )
        """)
        self.db.commit()

    def _sync_from_db(self):
        """Load shared state from ledger."""
        rows = self.db.execute("SELECT * FROM provider_state").fetchall()
        state_map = {r[0]: r for r in rows}
        for p in self.providers:
            if p.id in state_map:
                _, state, cooldown_until, probes, _ = state_map[p.id]
                p.state = ProviderState(state)
                p.cooldown_until = cooldown_until
                p.recovery_probes = probes

    def _persist_state(self, provider: Provider):
        self.db.execute("""
            INSERT OR REPLACE INTO provider_state
            (id, state, cooldown_until, recovery_probes, updated_at)
            VALUES (?, ?, ?, ?, ?)
        """, (provider.id, provider.state.value, provider.cooldown_until,
              provider.recovery_probes, time.time()))
        self.db.commit()

    def mark_failed(self, provider_id: str, error_code: int = 429):
        """Call this when a provider returns an error."""
        provider = self._get(provider_id)
        if not provider:
            return

        if error_code == 429:
            provider.state = ProviderState.COOLING_DOWN
            provider.cooldown_until = time.time() + provider.cooldown_seconds
            print(f"[Resilience] {provider.name} โ†’ COOLING_DOWN "
                  f"(retry after {provider.cooldown_seconds}s)")
        else:
            provider.state = ProviderState.UNAVAILABLE
            provider.cooldown_until = time.time() + 900  # 15 min for hard failures
            print(f"[Resilience] {provider.name} โ†’ UNAVAILABLE")

        self._persist_state(provider)

    def mark_healthy(self, provider_id: str):
        """Call this on a successful response."""
        provider = self._get(provider_id)
        if provider and provider.state != ProviderState.HEALTHY:
            if provider.state == ProviderState.RECOVERING:
                provider.recovery_probes += 1
                if provider.recovery_probes >= self.FULL_RECOVERY_AFTER_PROBES:
                    provider.state = ProviderState.HEALTHY
                    provider.recovery_probes = 0
                    print(f"[Resilience] {provider.name} โ†’ HEALTHY โœ…")
            else:
                provider.state = ProviderState.HEALTHY
            self._persist_state(provider)

    def get_active_provider(self, notify_fn=None) -> Optional[Provider]:
        """
        Returns the highest-priority available provider.
        Transitions COOLING_DOWN โ†’ RECOVERING when cooldown expires.
        """
        self._sync_from_db()
        now = time.time()

        for provider in self.providers:
            if provider.state == ProviderState.HEALTHY:
                return provider

            if provider.state in (ProviderState.COOLING_DOWN,
                                  ProviderState.RECOVERING):
                if now >= provider.cooldown_until:
                    provider.state = ProviderState.RECOVERING
                    provider.recovery_probes = 0
                    self._persist_state(provider)
                    print(f"[Resilience] {provider.name} โ†’ RECOVERING (probe started)")
                    return provider  # canary routing

        return None  # all providers unavailable

    def _get(self, provider_id: str) -> Optional[Provider]:
        return next((p for p in self.providers if p.id == provider_id), None)


# --- Example Usage ---

providers = [
    Provider(
        id="anthropic_oauth",
        name="Anthropic (Claude Max OAuth)",
        priority=1,
        limit_type=LimitType.ACCOUNT_LEVEL,
        cooldown_seconds=3600,
    ),
    Provider(
        id="anthropic_api",
        name="Anthropic (Paid API Key)",
        priority=2,
        limit_type=LimitType.KEY_LEVEL,
        cooldown_seconds=300,
        notify_on_fallback=True,
        notify_message="โš ๏ธ Falling back to paid Anthropic API. Approve?",
    ),
    Provider(
        id="openai",
        name="OpenAI (GPT-4o)",
        priority=3,
        limit_type=LimitType.KEY_LEVEL,
        cooldown_seconds=300,
    ),
    Provider(
        id="local_llm",
        name="Local LLM (llama.cpp)",
        priority=4,
        limit_type=LimitType.NONE,
        endpoint="http://localhost:8080/v1",
    ),
]

resilience = ProviderResilienceBind(providers)

# On each request:
active = resilience.get_active_provider()
if active:
    print(f"Routing to: {active.name}")
    # ... make request ...
    # On success:
    resilience.mark_healthy(active.id)
    # On 429:
    # resilience.mark_failed(active.id, error_code=429)
else:
    print("โš ๏ธ All providers unavailable. Check local LLM or wait for cooldown.")
openclaw.json โ€” models section (patch, don't replace)
{
  "models": [
    {
      "id": "anthropic/claude-sonnet-4-6",
      "auth": "oauth",
      "priority": 1
    },
    {
      "id": "anthropic/claude-sonnet-4-6",
      "auth": "api-key",
      "priority": 2,
      "notifyBeforeUse": true
    },
    {
      "id": "openai/gpt-4o",
      "auth": "api-key",
      "priority": 3
    },
    {
      "id": "local/llama3",
      "baseUrl": "http://localhost:8080/v1",
      "priority": 4
    }
  ],
  "fallback": {
    "onRateLimit": "next",
    "onError": "next",
    "notifyOnPaidFallback": true
  }
}

โš ๏ธ Always propose config changes to your operator before applying. Never run config.apply without explicit approval.

๐Ÿฉบ Rate Limit Diagnostic Checklist

When you hit a rate limit, run through this before burning time:

  1. Is it account-level or key-level? โ€” Swap to a key from a different account. If still limited, it's account-level (key rotation won't help).
  2. Is it model-specific? โ€” Try a smaller model (Haiku, GPT-3.5-turbo). Some quotas are per-model-tier.
  3. Is session state cached? โ€” Clear your auth cache (~/.openclaw/) and re-authenticate. A reboot alone may not flush OAuth tokens.
  4. Is it IP-based? โ€” Check if a different network (mobile hotspot, VPN egress) succeeds with the same key.
  5. How long is the window? โ€” Most Anthropic limits reset hourly or daily. Check your provider dashboard for the reset time.
  6. Is your local LLM running? โ€” If llama.cpp or Ollama is available, route there immediately. Don't block on cloud recovery.

๐Ÿ”— Composing with Other Binds

The Provider Resilience Bind answers: which provider is available?

Pair it with:

Composition pattern:
Resilience โ†’ "which provider?" โ†’ Efficiency โ†’ "which model?" โ†’ Security โ†’ "is this safe?" โ†’ Cost โ†’ "who pays?"