Open Chat Interfacedocs

Budgets and limits

Usage budgets in messages, tokens or cost, per-person overrides, rate limits, storage allowances, and how reservations keep concurrent use honest.

Three kinds of limit, for three different jobs:

LimitJobWhere
Usage budgetsHow much someone may consumeModels → Usage budgets
Rate limitsStop one client overwhelming the instancePeople → Roles & access
Storage allowanceHow much someone may hold in filesPeople → Roles & access

People see their remaining budget and storage as a percentage, never as figures (Limits and usage).

Usage budgets

Models → Usage budgets (/admin/quotas). A budget (a quota policy) caps consumption and applies to one or more roles. A role can carry several and every one is enforced; each person gets the full amount on their own.

Models → Usage budgets: budgets with their measure, window, model scope and the roles they apply to.

What to measure

MeasureSuits
MessagesSimplest to explain: "500 messages a month".
TokensCloser to real cost, harder to explain to someone who hits it.
CostExact, and the only one that means anything across models of different prices. Needs prices on the models; unpriced models cost nothing.

Measure cost when models differ in price, messages when they do not.

Which window

  • Rolling: the last n hours, moving continuously. Capacity returns gradually.
  • Calendar: daily, weekly or monthly, resetting at midnight in a time zone you choose. Easier to explain; produces a rush at the start of each period.

Which models

A budget can cover every model or specific ones. Scoping lets you be generous with an inexpensive model and strict with an expensive one. A model in no budget is unlimited.

People are warned as they approach a limit, at 80% and again at 95%. The thresholds are fixed.

A worked example

A department wants staff using a costly model without an open-ended bill:

  • Measure: cost. Window: calendar month. Scope: the expensive model only.
  • Limit: look at Usage for a fortnight first rather than guessing.

Per-person overrides

Give one person more without changing the policy, from Limits beside them in the user list or Adjust limits on their page, with an optional expiry and reason.

How enforcement works

A reply reserves an amount before it starts and settles to the actual usage when it ends, so concurrent requests see each other and cannot together overshoot a limit. Stopped and failed replies still settle. The amounts held are Budget held per response (default $0.25) and Tokens held per response (default 4,000) under Instance-wide on Roles & access.

If a provider does not report usage, the estimate stays held within the budget's window rather than counting as free; it is released when complete usage arrives or the request ages out of the window. A reservation left by a crashed process is swept after 15 minutes and marked as unknown usage, not forgiven. These are estimates: use provider-side spending controls where a hard financial cap is required.

Besides replies, conversation summaries, embeddings and reranking are usage too (summaries and embeddings count tokens; none counts messages). Imported conversations do not count.

Rate limits

Per role, on People → Roles & access:

Limitadminauditoruserrestricted
Replies streaming at once10331
Messages per minute120303010
Uploads per minute12020205

Starting a conversation counts against the same messages-per-minute value, separately: a person can start as many conversations a minute as they can send messages. A person who already has ten conversations from the last minute that are still untitled and empty is refused another until one is used or the minute passes. Both answer 429 with Retry-After, and together they stop a looping client from filling an account with empty conversations.

These are the built-in defaults. An environment variable such as RATE_LIMIT_CHAT_PER_MINUTE replaces the default for every role; a value saved here wins over both. Each value shows where it came from.

A person working normally should never meet these. The concurrency cap also bounds how far a budget can be overshot by simultaneous requests. Without Redis, rate limits are enforced per API replica.

Sign-in attempts

Sign-in attempts per minute, under Instance-wide on Roles & access (RATE_LIMIT_AUTH_PER_MINUTE, default 10), limits how fast passwords and tokens can be tried. It counts every request to sign in with a password or single sign-on, sign up, request or complete a password reset, and send or follow a verification link, successful or not, in two counters with the same limit:

  • per client address, so one machine cannot try many accounts;
  • per account (the email address in the request), so many machines cannot try one account.

Past the limit the request is refused with 429 Too Many Requests and a Retry-After header until the minute ends, and the sign-in page says Too many attempts. Wait a minute and try again. The first refusal each minute for an address or account is audited as auth.rate_limited, so a flood of refused requests does not flood the audit log too. Refused requests never reach the sign-in code, so they are not also recorded as failed sign-ins. Signing out, reading the session and changing a password while signed in are not counted.

The counters live in Redis and are shared by every API replica; without Redis each replica counts on its own. The client address is the one the web container reports: behind a load balancer or ingress, set TRUSTED_PROXIES, or everybody shares the proxy's address and its one counter. Many people behind one real address, such as a campus network, share a counter too: raise the value if they meet it at the start of a class.

Before v0.10 this setting had no effect.

Storage allowance

Per role, on People → Roles & access: total storage, number of stored files, and the largest single file. A blank total or file count means no limit; a blank per-file size falls back to the instance's upload limit on Storage. Enforce allowance switches it off without losing the values. A role with nothing saved is unlimited.

Storage is a gauge, not a flow: it measures what someone holds now. Uploads in progress reserve space. Deleting files, or moving their conversation to the trash, frees allowance at once; restoring needs it again. Artifact versions count towards total storage (not the file count), and so do project files.

Audit

Budget changes, overrides and role settings are audited. Role changes, including bulk ones, are kept regardless of audit retention.

On this page