API relay audit

An API relay audit asks what a third-party LLM proxy can change between your client and the real model provider. The relay sees your API key, your prompts, and often your tool definitions. It can rewrite any of them before the upstream call and again on the way back.

Founders route traffic through relays for price, region, or a single OpenAI-compatible base URL. That convenience moves trust to an operator you may not know. A clean latency test does not prove the bytes you receive match what Anthropic or OpenAI would have returned.

Empyre Relay at relay.empyre.dev is a different category. It is OAuth for AI agents, not an HTTP gateway in the model path. The sections below cover LLM proxy risks, manual tests, a checklist, and the open-source api-relay-audit project checked 2026-10-07.

What an LLM API relay sits on

In this context, a relay is an intermediary that accepts OpenAI-compatible or Claude-compatible HTTPS calls and forwards them upstream. Resellers, mirrors, and "one key for every model" gateways fit here.

The relay terminates TLS from your app. It can log bodies, trim context, attach hidden system text, pick a cheaper model, or edit streaming chunks. Your SDK still thinks it talked to the vendor.

That shape is not the same as an OAuth identity broker. Brokers mint scoped tokens for your app. They do not need to parse every completion stream. Confusing the two buys the wrong controls.

For auth brokers versus HTTP gateways in agent integrations, read API relay for AI agents on this site. LLM proxy audits belong here; OAuth broker choice belongs there.

What a relay can do without telling you

Each item is a mechanism. A relay that does none of them on day one can still change behavior later:

Hidden system prompts and injection

Extra instructions prepended to your messages inflate token usage and steer answers. The model may follow relay-owned rules you never wrote.

Model substitution

The response may claim one model while routing to another. Latency, stream metadata, and self-reported model names are signals, not proof by themselves.

Context truncation

A relay may drop middle or tail context to save cost. Long agent threads look fine until recall fails on facts you sent earlier.

Tool-call and package-command rewriting

Proxies can edit install commands or tool JSON in the body. That is a supply-chain risk when agents run shell or package steps.

SSE stream tampering

Streaming responses can reorder events, strip usage fields, or alter thinking blocks. Clients that only check HTTP 200 miss it.

Key and error-response leakage

Malformed requests sometimes echo env vars, file paths, upstream keys, or proxy internals in error JSON.

Logging and retention

Even an honest operator may store prompts for abuse review. Policy matters, but you still need technical probes because policy can change.

Manual probes before you automate

Keep a notebook of request ids, timestamps, and hashes. Compare relay output to a first-party call with the same model id, temperature, and messages when your contract allows direct access.

Send a canary user message with a unique nonce. Ask the model to repeat the nonce verbatim. If the nonce is missing, something between you and the model altered the user turn or the answer.

Pin a trivial completion: fixed system text, zero temperature, and a one-token answer you expect. Repeat ten times. Drift in wording or model id across runs is a substitution signal.

Measure input token counts on a minimal user message. A large delta versus the first-party API suggests hidden system content at the relay.

For tool-using agents, send a harmless package install instruction you control. Compare the exact command string returned through the relay and through the vendor API.

Deliberately break requests: wrong model name, oversize payload, invalid JSON. Read error bodies for secrets. Never paste production keys into a web form that forwards them to an unknown host.

Pre-production checklist

Use this before coding agents or wallets against a new relay base URL:

Contract

Written policy on logging, retention, subprocessors, and model routing. Vague marketing copy is not a contract.

Identity

Document which model ids you pay for. Run pinned-output tests on each id you will use in production.

Injection

Token delta tests and extraction-style probes on your relay URL. Treat inconclusive runs as open questions, not passes.

Tools

If agents install packages or call HTTP tools, compare tool-related strings relay versus direct.

Streams

On Anthropic-style SSE, verify event types, usage monotonicity, and terminal completeness for long answers.

Errors

Catalog error shapes from bad requests. Redact and rotate if anything looks like a secret.

Exit

Plan how you switch base URL or key without rewriting every agent if the relay fails audit.

The api-relay-audit open-source tool

The project toby-bridges/api-relay-audit on GitHub is a local audit runner for AI API relays and LLM proxies. Docs live at toby-bridges.github.io/api-relay-audit/, read 2026-10-07.

The site states execution stays on your machine. Your API key is sent only to the relay URL you choose. It is not sent to API Relay Audit servers or an extra web checker.

A run produces a Markdown report with findings labeled LOW, MEDIUM, or HIGH. The site read 2026-10-07 states the tool does not certify that a relay is safe, does not replace manual security review, and does not treat inconclusive as clean.

The published site describes fourteen audit steps covering prompt injection signals, model identity checks, context truncation probes, tool-call rewriting checks, error leakage tests, stream integrity, and optional Web3-oriented steps behind a profile flag.

You can run a standalone CLI from your terminal, or use a DeepSeek Harness plugin bundle pinned in their docs. Either way, treat the report as evidence to investigate, not a badge.

Relay API spend during a run is billed by your provider. The project homepage quotes roughly twenty to fifty cents of API usage per audit on 2026-10-07. Your invoice may differ.

When local audit helps and when it does not

SituationWhat audit gives youWhat it cannot prove
New reseller URLStructured probes and a saved Markdown trailFuture behavior after you ship
Coding agent with toolsSignals on rewritten install commandsSafety of every tool your agent invents later
Wallet-adjacent agentsProfile-gated Web3 refusal tests on the site, read 2026-10-07Protection against a compromised client device
Regulated dataEvidence for your own risk fileLegal sign-off without your counsel

Empyre Relay is not in the model request path

Empyre Relay implements OAuth for AI agents: hosted consent, mandatory S256 PKCE, one-time authorization codes, refresh rotation with family revocation, and fail-closed revoke and introspect. Feature-frozen since 2026-07-10 except scoped GET /relay/oauth/userinfo added 2026-08-06.

The npm package is @empyre/relay-sdk version 1.0.0 on registry.npmjs.org, read 2026-10-07. Relay tokens authorize an agent against third-party APIs on a user's behalf. Model calls go from your app to the provider and never pass through Relay.

Endpoint and billing detail live in Relay API and Relay agents. Secret handling for keys that must never leave storage belongs in AI agent secret storage without server private keys and How to store private keys for an AI agent app.

Where Empyre the company builder fits

Empyre on empyre.dev turns a brief into a live company with eight agents that keep operating after launch. That is a business operator, not an LLM proxy auditor and not a coding seat.

If your problem is only whether a discount gateway alters prompts, stay with relay audit work above. If you need a product, deploy, and ongoing operations without living in an IDE, compare operators instead.

Common questions

Is api-relay-audit the same as Empyre Relay?

No. api-relay-audit is an open-source security probe for third-party LLM HTTP relays. Empyre Relay is OAuth for agents at relay.empyre.dev.

Does a LOW report mean the relay is safe?

The tool docs read 2026-10-07 say it does not certify safety. LOW only means that step did not flag HIGH or MEDIUM under its rules.

Should I paste my production key into a web relay checker?

Prefer local runners that send the key only to your chosen base URL. Web checkers add another party that sees the secret.

How is this different from the API relay pattern article?

API relay for AI agents sorts OAuth brokers and HTTP gateways for integrations. Auditing an LLM proxy in the model path means probes, checklists, and tools like api-relay-audit, not OAuth broker selection.

What was verified on 2026-10-07?

toby-bridges.github.io/api-relay-audit/ and github.com/toby-bridges/api-relay-audit repository description; registry.npmjs.org @empyre/relay-sdk version 1.0.0. No search rankings, star counts, or Empyre customer figures appear here.

Try Empyre free for 3 days

Describe a business in plain words and watch eight AI agents build and deploy it. Starter is free for the first 3 days.

Start your free trial

Related

Last updated 2026-10-07. Competitor descriptions reflect each product's publicly documented capabilities at that date; they change often, so check the source before relying on a detail.