# Configure capacity

> Configure capacity - For deployed models, you can access the Capacity tab to edit throughput, usage
> limits, and reserved capacity for agents and other entities.

This Markdown file sits beside the HTML page at the same path (with a `.md` suffix). It summarizes the topic and lists links for tools and LLM context.

Companion generated at `2026-10-09T14:05:25.082037+00:00` (UTC).

## Primary page

- [Configure capacity](https://docs.datarobot.com/en/docs/workbench/console/settings/quota-settings.html.md): Full documentation for this topic (Markdown sidecar).

## Sections on this page

- [Capacity configuration](https://docs.datarobot.com/en/docs/workbench/console/settings/quota-settings.html.md#capacity-configuration): In-page section heading.
- [Reserved capacity](https://docs.datarobot.com/en/docs/workbench/console/settings/quota-settings.html.md#reserved-capacity): In-page section heading.
- [Set rate limits](https://docs.datarobot.com/en/docs/workbench/console/settings/quota-settings.html.md#set-rate-limits): In-page section heading.
- [Per-entity exceptions](https://docs.datarobot.com/en/docs/workbench/console/settings/quota-settings.html.md#per-entity-exceptions): In-page section heading.
- [Next steps](https://docs.datarobot.com/en/docs/workbench/console/settings/quota-settings.html.md#next-steps): In-page section heading.

## Related documentation

- [NextGen UI documentation](https://docs.datarobot.com/en/docs/workbench/index.html.md): Linked from this page.
- [Console](https://docs.datarobot.com/en/docs/workbench/console/index.html.md): Linked from this page.
- [Deployment settings](https://docs.datarobot.com/en/docs/workbench/console/settings/index.html.md): Linked from this page.
- [agent API key](https://docs.datarobot.com/en/docs/platform/acct-settings/api-key-mgmt.html.md#agent-api-keys): Linked from this page.
- [Usage](https://docs.datarobot.com/en/docs/workbench/console/monitoring-tools/usage.html.md): Linked from this page.
- [Service health](https://docs.datarobot.com/en/docs/workbench/console/monitoring-tools/service-health.html.md): Linked from this page.
- [Manage prediction environments](https://docs.datarobot.com/en/docs/classic-ui/mlops/deployment/prediction-env/index.html.md): Linked from this page.

## Documentation content

The Capacity tab provides controls for managing and enforcing usage on deployments. Deployment owners can protect shared deployment infrastructure and guarantee minimum throughput for critical agents and users when multiple consumers share one deployment.

Set capacity and the utilization threshold for the deployment as a whole; those are global to the deployment.Quotas —the default rules and optional per-entity limits below—define what happens when utilization reaches that threshold:

- Default throughput configuration: Configure a deployment's capacity, utilization threshold, and baseline usage rules that apply to any entity that can access the deployment. Entities without their own overrides use these defaults.
- Entity rate limits: Rate limits are optional settings that provide a higher priority for specific deployments, users, or groups. Use reserved capacity to guarantee a share of deployment capacity for each entity, or per-entity rate limits to control the total deployment throughput.

> [!TIP] Learn more
> For decision guidance, load test examples, and sizing recommendations—including when rate limits alone are sufficient—see [Rate limiting vs. quota reservations: a practical guide for platform teams](https://www.datarobot.com/blog/rate-limiting-quota-reservations/) on the DataRobot blog.

> [!NOTE] Rate limit application
> Rate limit changes may take up to 5 minutes to apply. This delay occurs because the gateway updates its quota cache every 5 minutes.

## Capacity configuration

Capacity is the throughput you expect a deployment to sustain expressed as units per time window (e.g., requests per minute or tokens per minute). It defines the baseline “pipe size” used for deployment-wide quota enforcement and for sizing reservations.

When choosing capacity values, common inputs include:

- Load tests that measure how the deployment behaves under target traffic.
- Model or hosting limits imposed by the model, runtime, or infrastructure.
- Latency budgets you need to meet at expected concurrency and payload sizes.
- Operational experience from comparable deployments or historical usage.

Throughput configuration governs how a deployment applies limits:

- It sets capacity as the overall ceiling for requests or tokens.
- It sets the utilization threshold as how full that capacity can get before the deployment enforces its default quota behavior.
- Below the threshold , it relaxes enforcement so the gateway can allow short bursts and treat traffic more permissively.
- Above the threshold , it applies the deployment's quota rules dynamically as utilization rises.
- With reserved capacity , it guarantees entitled entities a share when consumers compete for the deployment.
- Under sustained overload , it can reject excess traffic to protect the model and shared infrastructure.

To configure capacity:

1. ClickSet throughputto configure the capacity settings for a deployment.
2. Choose a metric to track (requests per minute or tokens per minute).
3. Define the capacity of requests or tokens per minute by providing a value. These values are not inferred automatically by DataRobot, so plan these values accordingly based on deployment usage.
4. Set the utilization threshold as a percentage of the capacity. DataRobot recommends setting thresholds at 70–80% as a common starting point. This leaves room for bursts of usage before enforcement tightens.
5. After configuring each capacity setting, clickSave.

### Reserved capacity

Reserved capacity is configured per entity (agent deployment, user, or group). It defines how much of the deployment’s capacity you guarantee to the selected entity when utilization is above the utilization threshold and consumers compete for the deployment.

- Floor, not a ceiling : A reservation guarantees a minimum share; an entity can often use more than its reserved portion when spare capacity exists.
- Leave unreserved headroom : Keep part of deployment capacity unreserved (often 10–20%) so ad-hoc traffic, new consumers, and overflow still have room.

To configure reserved capacity, you must already have the Capacity settings configured.

1. Once capacity settings are configured, clickAdd entity.
2. Select an entity from theDeployments,Users, orGroupslist. Deployments require their own API keyIf you reserve capacity for aDeploymententity, that deployment must call in using its ownagent API key, not the personal API key of the user who owns or triggers it. Agent API keys are generated automatically for agentic workflow deployments and are listed on theAPI keys and toolspage. Calls authenticated with a personal API key aren't attributed to the deployment entity, so its reserved capacity isn't applied.
3. Set the percentage of the capacity to reserve for the selected entity.
4. Perform this process for one or more entities (depending on your organization's needs) and clickSave.

## Set rate limits

On the Capacity page, manage per-entity settings in the Rate limits section:

1. ClickAdd policyto modify the rate limit settings for the deployment.
2. ClickAdd metricto begin configuration. Adding metricsA new policy row appears each time you clickAdd metric, until a row is present for every metric available.
3. In the new row, select aMetric, enter aLimit, and choose a timeInterval. The selected resolution applies to each metric-based policy defined here. The policy settings allow defining limits on three key metrics: MetricDescriptionRequestsControls the number of prediction requests a deployed model can handle in the selected time window, defined by the resolution setting. The default is 300 requests per minute.TokensControls how many tokens a deployed model can process in the selected time window, defined by the resolution setting. This limit includes all types of tokens (input and output).Input sequence lengthControls the number of tokens in the prompt or query sent to the model.Concurrent requestsControls the number of prediction requests a deployed model can process at the same time. The default is 50 concurrent requests.
4. Perform this process for one or more metrics (depending on your organization's needs) and clickSave.

### Per-entity exceptions

You can make exceptions to rate limits for specific entities.

To configure per-entity exceptions:

1. ClickAdd entity.
2. Select an entity from theDeployments,Users, orGroupslist. Deployments require their own API keyRate limit exceptions for aDeploymententity only apply when that deployment calls in using its own agent API key. SeeReserved capacityfor details.
3. ClickAdd metricto begin configuration.
4. In the new row, select aMetric, enter aLimit, and choose a timeInterval. The selected resolution applies to each metric-based quota defined here. For more information, seeSet rate limits.
5. Perform this process for one or more metrics (depending on the entity's required configuration) and clickSave.

## Next steps

After configuring capacity and rate limits, monitor how the deployment is using its allotted resources.

- Usage : Review prediction volume and actuals processing against the capacity and rate limits you configured.
- Service health : Check response time and error rates to see how the deployment performs under its configured throughput.
- Manage prediction environments : Review the prediction environment hosting this deployment, since infrastructure limits can also affect achievable throughput.
