Semogram Docs
Workspace managementUsage and activity

Request limits

Manage workspace and caller request windows separately from spending budgets

Consumer request limits constrain traffic during a 60-second window. The workspace bucket is shared across consumers; each API key and signed-in MCP user also has a caller bucket. Requests must fit both limits.

You need a Semogram account as owner/admin or an org:manage key to read/change these settings. They are separate from model token budgets, pipeline concurrency and monthly spend alerts.

Values and behavior

SettingDefaultScope
workspaceRequests600Shared workspace requests per window
apiKeyRequests120Each API key or signed-in assistant user per window
windowSeconds60Fixed window duration

These are fallback defaults; inspect the saved workspace settings rather than assuming them. Positive whole numbers up to 2,147,483,646 are accepted. Updating a limit does not reset the current window's usage.

Example: a bounded integration

For a small integration choose a reviewed workspace limit of 300 and caller limit of 60. Several keys share the 300 total; creating more keys does not bypass it. Another signed-in assistant consumes its own caller bucket plus the same shared workspace bucket.

Open Workspace settings → Consumer API limits. Set Organization requests to 300 and Requests per key or assistant user to 60. Save API limits and read the values back. Review existing traffic before lowering a limit.

Ask the administrator assistant to read organization_api_limits_get and explain the shared/caller limits before proposing a change. The MCP tool is read-only; save changes in the UI or API.

Set reviewed limits
curl --request PUT "https://platform.semogram.com/api/v1/api-limits" \
  --header "Authorization: Bearer ${SEMOGRAM_ADMIN_KEY}" \
  --header "Content-Type: application/json" \
  --data '{"workspaceRequests":300,"apiKeyRequests":60}'

GET the same route to inspect values and windowSeconds. Both actions require org:manage.

Call organization_api_limits_get with empty arguments on the workspace connection. Inspect workspaceRequests, apiKeyRequests and windowSeconds.

Handle throttling

For a 429 response honor Retry-After. Public responses can include X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset for the constraining dimension. Stop immediate retry loops and spread bounded requests over time. Retain operation identity for an uncertain mutation; rate-limit retries do not make writes idempotent.

A 402 usage_limit_reached is a different condition: the applicable plan's included credit is exhausted and new paid work is paused. Increasing request limits does not add credit. See Usage and budgets.