The mechanism
What happens between your application and the provider
Outcap is a proxy: your application talks to it the way it talked to OpenAI or Anthropic, and it forwards. In between, it makes decisions, in a fixed order, without ever reading a database. This page describes that order and what each step is allowed to do.
The path of one request
The eight checks run in the same synchronous turn, before a single byte reaches the provider. A refusal consumes no tokens and writes nothing at the provider: it is logged on your side, with its reason.
| Check | What it looks at | What it returns if it refuses |
|---|---|---|
| Outcap key | The project key, its state, its owner. | 401 |
| Kill switch | The project switch, before any work proportional to the body. | 429 outcap_kill_switch |
| Input guardrails | API keys, personal data, invisible text in the prompt. | 400, or redaction before sending |
| Margin | The end customer's spend this month against the ceiling derived from their revenue. | Cheaper model, then 429 if you chose it |
| Budgets | Counters per project, key, route or end customer, in memory. | 429 outcap_budget_exceeded, not retryable |
| Rate limits | A token bucket, in requests or in tokens. | 429 retryable, with retry-after |
| Cache | Exact body match, never semantic. | The answer you already paid for, at 0 tokens |
| Dispatch | Served model, applied cap, backup key if yours is refused. | The provider's answer |
The order is not decorative. The kill switch comes before any work proportional to the body size, and the cache comes after the limits, otherwise a limit you set would not be honoured to the letter.
Why the counters live in memory
A budget checked with a SQL query would cost a database round trip on every call. Outcap keeps its counters in memory and hydrates them at startup from your own logs. At dispatch it RESERVES the worst-case cost, then writes the real cost on return. Two requests fired at the same time therefore see each other's exposure, instead of both slipping under the ceiling.
The trade-off is written down: the counters live in a single process. Two instances share nothing and would let twice the ceiling through. The proxy detects its peers, shouts it in its logs and shows it in its health endpoint.
The learned output cap
For each route, Outcap measures how long answers are and sets a cap at the 99th percentile times a margin. A cap needs at least ten answers before it exists, and it never goes below the max_tokens you send yourself. If the cut rate goes above two percent, the cap is raised automatically: a route whose answers grow is not strangled.
Never a cap on a reasoning model, nor on a request that declares tools: a cut in the middle of a tool call would produce invalid JSON, and repair cannot fix that.
The clean cut
When an answer hits our cap, it is returned at a sentence boundary, never mid-word, and a truncated JSON body is repaired before it leaves. In streaming, repair on the fly is impossible: the cap is widened instead of delivering a broken object. A cut caused by your own max_tokens is left untouched: it is not ours.

Margin, decided during the request
You declare the monthly revenue of each end customer. Outcap derives a ceiling from it, compares it to that customer's spend this month, and acts before blocking: past the degrade threshold, the request goes to the cheaper model with a shorter output cap. The x-outcap-margin header says what was applied, so your application can see it.
With no declared revenue for a customer, no decision is taken. Outcap does not guess a ceiling, and does not read your billing system.

When things go wrong
A proxy is a place everything passes through: it has to behave when the provider, your database or your client misbehave.
The provider refuses
On a 429, 502, 503, 504 or 529, the request can go out again on a fallback model of the same family, one you chose. Never on a 500: it can happen after generation, so after billing, and replaying would mean paying twice.
Your provider key is refused
You can pass several keys. Outcap switches only when the refusal is about the key itself: revoked, out of credit, spend limit reached. Never to get around a provider rate limit, and it adds no capacity.
Your database goes down
Keys already known keep being served, even expired, and logs are buffered. Your traffic flows. What stops is the statistics, not the service.
Your client walks away
A client that drops its connection does not interrupt the call to the provider, deliberately. Cutting it would make the spend unknowable while the provider still bills. The request runs to the end and its cost is counted.
How we check that all this holds
Every guarantee on this page has a test. And because a passing test proves nothing until you have seen what makes it fail, each fix is verified by putting the defect back into the code: if no test falls over, it is the test we fix, not the code.
Two successive adversarial reviews found defects inside the fixes themselves, including one that let one tenant reset another tenant's limits. The changelog names them.
The default mode changes nothing
Plug it in, watch what it would have done on your real traffic, then switch on what you want, route by route.