Changelog

What changed, why, and what we had broken.

A changelog that lists features teaches nothing about how a product is held together. This one gives the reason for each change, and when a fix repairs our own mistake, it says so.

Entries taken from the journal shipped with the code. Dates are delivery dates, not announcement dates.

Numbering

What a number means

v0.x
One increment per delivered batch. Before a v1, nothing is promised about the admin API staying compatible. The OpenAI and Anthropic formats on the hot path do not move.
One version online
The hosted beta runs the latest version listed here. There is no version to pick, no stable channel and no test channel.
What is not in it
Neither comfort fixes nor wording changes. An entry exists when a behaviour changes on your side, or when we found a defect that concerned you.

v0.29

14 September 2026

Backup provider keys: a refused key no longer cuts the service

  • The x-provider-key header takes up to five keys separated by commas. When the provider refuses a key for itself, the request restarts on the next one with the same model, before the first byte is written to your client.
  • Never on a limit the provider imposes. A rate limit or a tier's monthly ceiling does not change key: OpenAI's terms forbid working around its limits, and a refused request still counts against the per-minute limit.
  • Keys are still neither logged nor stored. To remember a refusal, the proxy keeps a fingerprint computed with a secret drawn at startup, erased at most 61 minutes after the key was last used.
  • A memory reorders, it never removes. A key that is invalid, out of credit or at the ceiling you set is tried last for 15 minutes, and its first success clears the memory.
  • No capacity promised: keys from the same organisation share its limits and its credit. The phrase load balancing was cut from the design before the first line of code, along with the mode that would have spread your prompts across several accounts.

Our mistake

Our own texts promised more than the code. They said a refused key was tried last for 15 minutes whatever the refusal. That is only true of a key that is invalid, out of credit or at the ceiling you set, and at OpenAI of a key without access to the model, for that model. The texts say it now.

v0.29

14 September 2026

OpenTelemetry export: your traces in your tool, without the content

  • One trace per logged request, in OTLP/HTTP JSON, to at most three destinations per project. Every send to the provider gets its own span with the standard gen_ai attributes: a model fallback or a key rotation adds one, which says why the request left again.
  • Never any content, and no option to add some: no prompt, no response, no key, no fingerprint of your system prompt, no customer revenue. A network error enters only through its class, never through its message, which sometimes copies an address or a header value.
  • No dependency added: the encoding and the sending fit inside the proxy's own code. The security page still shows 9 direct dependencies.
  • A broken collector makes nothing wait. Past half of the sends failing over five minutes, the export suspends itself, keeps the 64 most recent requests, retries at most once a minute and warns your alert channels.
  • A block is not an error. A 429 decided by a budget, by margin or by a rate limit leaves with a status that is not marked as an error, otherwise your dashboards would confuse a guardrail doing its job with an outage.
  • What the export does not do: no metrics, no protobuf, no OTLP ingestion, no guaranteed delivery. The queue lives in memory, and a restart loses whatever was waiting.

Our mistake

An ingestion key could end up in clear text in our logs. When saving an export failed in the database, the error copied every parameter of the query, headers included, and our handler logged the whole error. The same path exposed the secret URL of a Slack webhook. An error now enters a log only through its type, its codes and its call stack, never through its message.

v0.29

14 September 2026

Input guardrails: what should not have left for the provider

  • API keys, personal data and invisible text are looked for in every request, then flagged, redacted or refused before the call. Nothing is on by default, and everything can run in observation first: the logs say what would have been redacted, and requests do not change.
  • Deterministic, in memory, without a model. Sending your prompts to a third party to inspect them would contradict the security page. In exchange, names and postal addresses are not recognised, and that is written wherever the feature is described.
  • No regular expressions. An email pattern written the way one usually writes it takes more than a second on twenty-eight characters built for it, and that time doubles with every extra character. The proxy has one thread: a single tenant would have frozen all the others.
  • A value that is found is never written down. Logs, headers, alerts, audit log, traces and error messages carry only the type and the count. A guardrail that copied into its logs the data it protects would be worse than none.
  • Prompt injection signals never block. Nobody knows how to stop prompt injection reliably, and the list of phrases we look for is public in the dashboard: hiding it would stop no one and would stop you from judging what it is worth.

Our mistake

A private key longer than 16 KB left half in clear text, under a response that announced it as redacted. And redacting invisible text could glue back together the two halves of a card number that those characters had split: the proxy was building a valid number in clear text itself. Both are fixed and locked by a test.

v0.28.1

14 September 2026

A client that left too early made its spend disappear

  • When a client cut a streaming request before the provider had started answering, the request stayed suspended on our side: never settled, never logged. The provider generated and billed anyway.
  • That spend therefore entered neither the budgets nor the logs, not even after a restart. A client could stay under a hard budget by abandoning its streams at the right moment.
  • The request now runs to the end, its real cost is counted, and a test locks that behaviour.
  • One decision became explicit along the way: a client that leaves does not cut the call to the provider. Cutting would sometimes save a little, but would make the cost unknowable, and a cost control tool has to count what is spent first.

Our mistake

The code looked like it cut the call when the client left. It did not: the departure listener was inert, since Node 24 closes the request as soon as its body is read. The hole was found while preparing the OpenTelemetry export, not reported by a user, and reproduced before being fixed.

v0.28

14 September 2026

Rate limits: protecting the pace, not just the total

  • Requests or tokens per minute, per hour or per day, on the project, a key, a route or an end customer. A budget in dollars protects the month's bill, it does not stop anyone from spending it in five minutes.
  • With no specific target, the limit applies to each one separately: twenty requests per minute per end customer slows the one who abuses without touching the others, and without declaring your customers one by one.
  • Tokens are reserved at worst case then refunded at the real figure, like dollars. Concurrent requests see what is already in flight, and none slips through on a counter that has not caught up.
  • The refusal is retryable, unlike an exhausted budget: the 429 carries retry-after, retry-after-ms and x-should-retry.
  • Two sentences in the dashboard were wrong and have been rewritten. A token bucket is not a sliding window and lets through up to about twice the limit over one window; the official SDKs only wait on their own if the delay is one minute or less.
  • The counters live in the memory of a single instance and start full again after a restart, unlike budgets.

Our mistake

One tenant could reset another tenant's limits. To bound memory, the engine forgot the oldest counters across all projects: those were precisely the abusers' counters, long since drained, and their next request started on a full one. The bound is now per project, and beyond it new objects share a common counter, stricter, never more permissive.

Earlier

From v0.27 back to v0.20

Eight batches, one line each. Three of them contain nothing but fixes to our own code.

VersionWhat changed
v0.27The fixes had defects of their ownA second review was told to break the v0.26 fixes: thirteen defects, none refuted. Above all it showed that our protection against double billing could be removed without a single one of the 357 tests of the day failing. Since then, every fix is verified by deliberately putting the defect back.
v0.26Seventeen defects found in our own codeAdversarial review of batches v0.20 to v0.24. Any signup could have our proxy probe our internal network through a webhook URL, a fallback could bill you twice, the cache could serve another request's response, a webhook signing secret travelled inside a URL, and three dashboard figures were wrong.
v0.25OpenAPI specificationThe admin API is described in OpenAPI, with a continuous integration check that fails if a route exists on one side and not the other. It found an undocumented endpoint on its first run.
v0.24Security page, rerunnable benchmark, self-hostingThe security page says the response cache keeps bodies in memory when it is on, that the end customer identifier is stored exactly as you pass it to us, and that an unavailable Outcap stops your traffic. No overhead figure is shown while there is no status page with several months of history behind it.
v0.23Exact response cacheExact match, never semantic: a semantic cache changes the answer without being asked. The cache runs after budgets and margin, because a limit that has been set is respected to the letter even when a hit costs nothing.
v0.22Outbound alertsSigned webhooks and Slack, a margin alert per end customer, and spike detection whose method is written in plain sight: baseline over seven rolling days, current day excluded, trigger at five times the baseline. It blocks nothing, because blocking on a moving average would cut a legitimate customer on a busy day.
v0.21Audit log and the glass ceilingA budget scoped to a key or a route was never checked as belonging to the project: a typo created a budget that never fired. The audit log is read-only, and the single-instance limit is written to the logs and exposed by /healthz.
v0.20Model fallbackThe request is replayed on the next model when the provider refuses to serve, never on an ambiguous status such as a 500: a request that may have been processed may have been billed. A fallback is not counted as a routing saving, otherwise a day of outage would look like a day of optimisation.

Check rather than believe

A changelog can be read, not verified. The security page says what is stored, what never is, and how to check it yourself; the documentation says what each header does.

Every guarantee quoted here has its test, and every test was verified by deliberately putting back the defect it protects against.

Changelog · Outcap