Anyone running multi-provider routing for LLM inference as the default, not just failover? We’re running a multi-region gateway (us-east1/us-west2) with circuit breakers, token-aware rate limiting, and hedged requests, but a provider latency surge last Tuesday blew our 200 ms P95 — what’s kept your throughput and cost steady during partial outages?