Security and data

What we keep, what we cannot keep, and what happens when we go down.

A proxy puts itself on the critical path of your production. Three questions deserve a written answer rather than a ticked box: what do you see of my data, what happens to my traffic if you go down, and how do I leave if I change my mind. Every answer on this page comes with the procedure that would prove it wrong.

The answer that does not flatter us is on this page too: if Outcap goes down entirely, your traffic stops. It is in the section about things going down, not in a footnote.

WHAT COMES INprompt bodyresponse bodyprovider keyrequested modelroute tagend customer idOUTCAPWRITTEN TO THE LEDGERtokens, cost, latencymodel, status, timestampdetected types, as countsNO COLUMN CAN RECEIVE ITthe text of your promptsthe text of the answersyour provider keyseverything crosses, almost nothing stays: the proxy counts, it does not archivecheckable in ten minutes, the procedure is two sections below

The data

What crosses the proxy

Outcap measures counters, not content. The body of your requests and answers is written to no database, and no column exists to receive it. The distinction matters: "we do not store it" is a promise, "no column can receive it" is a fact about the schema.

DataOn our sideDetail
The body of your promptsneverCrosses memory for the duration of the request, then goes. No database column can receive it.
The body of the answersneverSame, with one exception we would rather write down: if you turn the cache on for a route, the answer is held in RAM for the duration you chose, then swept within 30 seconds of expiring. It is never written to disk. Never cached: an error, an answer cut by our cap, an answer served by a fallback model.
Your OpenAI or Anthropic keysneverThey travel in the x-provider-key header on every request, and are neither logged nor persisted. If our servers were compromised, there would be no key to steal.
Your backup keys in memoryneverIf you pass several keys, the proxy remembers a fingerprint of the ones that were just refused, computed with a secret drawn at startup and never written down. It is erased at most 61 minutes after that key was last used, and nothing survives a restart: enough to recognise a burnt key, not enough to recover it. Logs and alerts only ever name a key by its position in your header.
Tokens, cost, latency, model, statuskeptThis is the product itself. Without these counters, nothing can be capped.
Route identifierkeptEither the tag you pass yourself in x-outcap-route, or a SHA-256 fingerprint of the first 256 characters of your system prompt. A fingerprint, never the text, and it is never sent in the OpenTelemetry export.
End customer identifierkeptExactly the value you hand us in x-outcap-user. If you put an email in it, we store an email: pass an opaque identifier.
Sensitive data that was detectedkeptOnly if you turn input guardrails on, and only the type and the count: 2 emails, 1 OpenAI key. It lives in a column of whole numbers, which cannot hold any text. Never the value found, never an excerpt, never a position.
Revenue you declare per customerkeptA monthly amount you type in yourself. We never read your billing, and no payment connector exists in the product.
Telemetry sent to your collectorif you turn it onOnly if you configure an OpenTelemetry export. One trace per logged request: model, tokens, cost, latency, Outcap's decisions. Never prompt or answer content, never your keys, never your system prompt fingerprint, never your customers' revenue. The end customer identifier appears only if you enable it.
Your collector's auth headersif you turn it onStored on our side, unencrypted, so they can be sent again on every export. The API never reads them back: only their names can be listed, and you have to type them again if you change the destination URL. We would rather say it than let you picture a vault.

The check

Do not take our word, go looking for the sentence

The only proof worth anything is the one you produce yourself. Send a sentence you will recognise, export your logs, search for it. Ten minutes, no special access, and the result is a number.

check that the text is nowhere

# 1. a sentence you will recognise, in a real call
curl -s "$OUTCAP/v1/chat/completions" \
  -H "authorization: Bearer $OUTCAP_KEY" \
  -H "x-provider-key: $OPENAI_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"pineapple-sextant-9134"}]}'

# 2. export your logs as CSV from the dashboard, Logs page

# 3. search for the sentence in the export
grep -c "pineapple-sextant-9134" outcap-logs.csv
0

The same test works on the OpenTelemetry export: point it at a collector in debug mode and search its output for the sentence. It is not there either, and the traces carry no prompt, no answer and no key.

Input guardrails

What is recognised in a prompt, and what is not

Input guardrails read the request before it leaves, and can redact or refuse. They are off by default. Here is the exact list of the twenty detectors, because a promise of "sensitive data detection" without a list means nothing.

FamilyWhat is recognisedWhat is done
Secrets (10)OpenAI, Anthropic, AWS, GitHub, Slack, Stripe and Google keys, PEM private keys, JWTs, and our own Outcap keys.redact or refuse
Personal data (6)Email, phone, card number validated by its Luhn check digit, IBAN, US social security number, French NIR.redact or refuse
Invisible text (2)Unicode tag characters and runs of invisible characters: the simplest way to hide an instruction in text your user pasted in.strip or flag
Injection markers (2)Chat template tokens misplaced inside a user message, and a public, versioned list of known override phrases.flag only

What is never recognised

Not people's names, not postal addresses, not your in-house secrets with no checkable shape. Not images, not PDFs: only the text of the request is read. Recognising a name takes a language model, which would have to run in a container alongside or at a third party. The first adds latency and memory, the second cancels the promise of this page.

What a redaction looks like

The value found is replaced by a pseudonym computed with a key of your own project, something like [EMAIL_3f9a1c7e5b02]. The same value gives the same pseudonym from one request to the next: your conversations stay coherent and the provider's prompt cache stays valid. In exchange, the provider can see that two of your requests contain the same value, without being able to recover it.

The detectors use no regular expression at all, and a test checks that on the syntax tree of all three modules: a crafted text cannot blow up the scan time. Shadow mode computes everything and changes nothing, so you can look before you decide.

Failure

What happens when something goes down

A guardrail that turns its own outage into an outage of your production is worse than no guardrail at all. Here are the real cases, one by one, including the one that does not flatter us.

Our database goes down
Traffic keeps flowing. Keys already known are still served from the memory cache even when expired, unknown routes fall back to ephemeral state, and logs are buffered then written later. The honest caveat: during that outage, a budget scoped to a route not yet warm in memory does not apply, and that spend does not enter the per-route aggregates.
The provider goes down or refuses
If you configured a fallback on the route, the request is replayed on the next model, but never on an ambiguous error such as a 500: a request that may have been processed may have been billed, and we refuse to risk making you pay twice. With no fallback configured, you get the provider's status and error body unchanged, along with its retry headers.
One of your keys is refused
With a single key, nothing changes: you get the provider's answer. With several, the request moves to the next one before a single byte reaches your client, and only when the refusal is about the key or the account behind it: revoked key, credit exhausted, billing problem, a spend limit you set yourself, no access to the model. A limit imposed by the provider never causes a key change. If all of them are refused, you get one of their real answers, chosen by a fixed rule: the proxy never invents a refusal.
Your trace collector goes down
Your traffic is not blocked: the export runs in the background and never makes a request wait. Traces queue in memory, then the oldest are dropped and counted in the dashboard. When at least half the sends of the last five minutes failed, with at least five failures and the first one a minute old or more, the export suspends itself, retries at most once a minute and notifies your alert channels. Whatever is still queued is lost if the proxy restarts: a trace is not a guardrail, and we do not put it ahead of your traffic.
An answer is cut by our cap
It is finished at the last sentence boundary, and if it was JSON, it is repaired before it reaches you. A tested invariant guarantees that no invalid JSON ever leaves the proxy. A cut caused by your own max_tokens is never rewritten: it is not ours.
Outcap goes down entirely
Your traffic stops, like with any proxy. That is the honest answer and we will not pretend otherwise. The two real defences are being able to leave in one line, just below, and self-hosting in time. No written uptime commitment replaces those two, least of all from a beta.

The exit

Leaving takes one line

The category consolidated in 2026: Stripe bought OpenRouter, Palo Alto Networks bought Portkey, Helicone went into maintenance. "Will you still be here in eighteen months?" is a fair question for everyone, us included. The useful answer is not a promise, it is an exit we show you before you come in.

in, then out

# Under Outcap
client = OpenAI(
    base_url="https://proxy.outcap.tech/v1",
    api_key=os.environ["OUTCAP_KEY"],
    default_headers={"x-provider-key": os.environ["OPENAI_API_KEY"]},
)

# To leave: drop the base URL and the headers.
client = OpenAI()

The line in question is the one in your client's configuration: the one you changed to come in is the only one to undo. No data migration, no proprietary format to take with you, no application code to rewrite. Your keys never moved from the provider, and your logs export to CSV at any time.

The surface

Nine direct dependencies

On 24 March 2026, two LiteLLM releases published to PyPI contained a credential stealer, live for about forty minutes, with official advice to rotate every secret. That incident moved the question: it is no longer "who hosts it" but "how much third party code runs on the path of your requests". Here are the nine, and how to recount them.

Counted on 16 September 2026: 9 direct dependencies, of which 7 are third party npm packages and 2 are ours. We publish no transitive total here, because it moves with every update and a stale number is worth nothing: the command is just below and will give you yours.

PackageWhat it does
fastifyThe HTTP server.
undiciThe HTTP client towards OpenAI and Anthropic.
drizzle-ormSQL queries, typed at compile time.
postgresThe Postgres driver.
@electric-sql/pgliteEmbedded Postgres, for tests and for self-hosting without a separate database.
@fastify/rate-limitThe per IP limit on public endpoints.
jsonrepairRepairing JSON cut by our cap.
@outcap/dbOur own schema and migrations.
@outcap/sharedOur own per model price table.

recount it yourself

pnpm --filter @outcap/proxy list --prod --depth Infinity

# 9 direct: 7 third party npm packages, 2 of our own.
# No agent, no provider SDK, no third party telemetry.

The limits

What we cannot claim yet

A product's limits always get discovered in the end. Better here, before the purchase, than in production on a Tuesday night.

No SOC 2

It would be a line on a page, not a report you could download. We would rather answer directly the questions SOC 2 is meant to cover, which is what this page does. If your purchase depends on it, tell us and let us talk about it openly.

One instance, deliberately

Budget counters live in memory, which is what makes the check instant and safe under concurrency inside one instance. Two instances would share nothing, so spend could reach twice the limit. The proxy detects a peer, writes it as an error in its logs and shows it on its health page.

No status page, so no latency figure

You will find no latency overhead figure on this site, and that is not an oversight: without several months of public history, such a number reads as bluff. The only one worth having is yours: the proxy records overhead_ms on every request, and you see its p50 and p95 on your own traffic, in the overview and in the CSV export. The benchmark method is published (pnpm bench, percentile against percentile, alternating series, warm-up excluded, nothing discarded), and the script itself prints that a fake upstream flatters us.

No SSO, no roles, no signed DPA

The admin API is protected by a single shared secret, with no roles and no rotation. That is exactly why every action that changes a setting is written to an audit log that nothing erases, with its author and their address: to the question "who disabled that budget?", it is the only honest answer we have today. No SSO, no roles, no signed DPA to send you.

More keys do not add capacity

OpenAI applies its rate limits per organisation and project, Anthropic per organisation and workspace: keys that share those share their ceiling. The proxy never switches key on a rate limit, because OpenAI's terms forbid working around them. Switching key also sends your prompts to another account, whose data retention settings may differ.

Prompt injection is flagged, never blocked

Input guardrails spot invisible text, misplaced chat template tokens and a handful of known phrases, and they only flag them. A reworded attack gets through. We do not sell protection against prompt injection, because nobody knows how to provide a reliable one today: OWASP and the UK NCSC say so themselves.

A beta, one person, zero paying customers

The product is young, it is written by one person, it has no funding and no paying customer yet. That is a real risk and you are entitled to weigh it, and better read here than guessed. It is also why everything starts in shadow mode, with not one of your requests modified until you decide otherwise, and why the exit is shown before the entrance.

A vendor questionnaire to fill in?

Send it as it is. Answers that are not on this page will be written here afterwards, including when they are "no".

The headers and error codes named here are detailed in the technical reference

Security and data · Outcap