Endpoint rate limits

MCP backends and HTTP tool groups accept rate_limit; the default is unlimited. Each configured ID has its own allowance shared by all users. Manual and OpenAPI tools in an HTTP group share that group's allowance. Edit Rate limits in the admin UI, or include the object in backend/tool-group management API input.

# Add to a backends entry; configure HTTP groups through the admin UI/API.
rate_limit:
  requests_per_second: 20
  burst: 40
  max_concurrent: 8
  • requests_per_second: average admission rate, including fractions; 0 or omission means unlimited.
  • burst: token bucket capacity. With a positive rate, 0 or omission uses capacity 1. A positive burst requires a positive rate.
  • max_concurrent: active request cap; 0 or omission means unlimited. Can be used independently of a rate limit.

Limits apply after authentication and scope checks to tools/call, prompts/get, resources/read, resources/subscribe, completion/complete, and resource-bearing subscriptions/listen. Streaming responses hold a concurrency slot until completion or cancellation. A subscription spanning endpoints atomically checks all their budgets and counts once per endpoint. Initialization, catalog reads, health checks, cancellation, unsubscribe, admin probes and background refresh do not consume this allowance.

Rejected requests never reach the backend: HTTP 429, Retry-After seconds and a JSON-RPC error preserve the original request ID. Retry timing for a concurrency cap is only a hint. mcpbridge connect reports the limit and retry hint without replaying calls. Configuration saves, SIGHUP and HTTP tool edits preserve unchanged budgets and active counts; restarting the process resets in-memory counters. Policies persist in SQLite/PostgreSQL; counters do not use the database and are single-instance only. Global ingress, IP and per-user quotas require separate policies; limit unauthenticated traffic at the reverse proxy.