An API relay audit asks what a third-party LLM proxy can change between your client and the real model provider. The relay sees your API key, your prompts, and often your tool definitions. It can rewrite any of them before the upstream call and again on the way back.
Founders route traffic through relays for price, region, or a single OpenAI-compatible base URL. That convenience moves trust to an operator you may not know. A clean latency test does not prove the bytes you receive match what Anthropic or OpenAI would have returned.
Empyre Relay at relay.empyre.dev is a different category. It is OAuth for AI agents, not an HTTP gateway in the model path. The sections below cover LLM proxy risks, manual tests, a checklist, and the open-source api-relay-audit project checked 2026-10-07.
What an LLM API relay sits on
In this context, a relay is an intermediary that accepts OpenAI-compatible or Claude-compatible HTTPS calls and forwards them upstream. Resellers, mirrors, and "one key for every model" gateways fit here.
The relay terminates TLS from your app. It can log bodies, trim context, attach hidden system text, pick a cheaper model, or edit streaming chunks. Your SDK still thinks it talked to the vendor.
That shape is not the same as an OAuth identity broker. Brokers mint scoped tokens for your app. They do not need to parse every completion stream. Confusing the two buys the wrong controls.
For auth brokers versus HTTP gateways in agent integrations, read API relay for AI agents on this site. LLM proxy audits belong here; OAuth broker choice belongs there.
What a relay can do without telling you
Each item is a mechanism. A relay that does none of them on day one can still change behavior later:
Hidden system prompts and injection
Extra instructions prepended to your messages inflate token usage and steer answers. The model may follow relay-owned rules you never wrote.
Model substitution
The response may claim one model while routing to another. Latency, stream metadata, and self-reported model names are signals, not proof by themselves.
Context truncation
A relay may drop middle or tail context to save cost. Long agent threads look fine until recall fails on facts you sent earlier.
Tool-call and package-command rewriting
Proxies can edit install commands or tool JSON in the body. That is a supply-chain risk when agents run shell or package steps.
SSE stream tampering
Streaming responses can reorder events, strip usage fields, or alter thinking blocks. Clients that only check HTTP 200 miss it.
Key and error-response leakage
Malformed requests sometimes echo env vars, file paths, upstream keys, or proxy internals in error JSON.
Logging and retention
Even an honest operator may store prompts for abuse review. Policy matters, but you still need technical probes because policy can change.
Manual probes before you automate
Keep a notebook of request ids, timestamps, and hashes. Compare relay output to a first-party call with the same model id, temperature, and messages when your contract allows direct access.
Send a canary user message with a unique nonce. Ask the model to repeat the nonce verbatim. If the nonce is missing, something between you and the model altered the user turn or the answer.
Pin a trivial completion: fixed system text, zero temperature, and a one-token answer you expect. Repeat ten times. Drift in wording or model id across runs is a substitution signal.
Measure input token counts on a minimal user message. A large delta versus the first-party API suggests hidden system content at the relay.
For tool-using agents, send a harmless package install instruction you control. Compare the exact command string returned through the relay and through the vendor API.
Deliberately break requests: wrong model name, oversize payload, invalid JSON. Read error bodies for secrets. Never paste production keys into a web form that forwards them to an unknown host.
Pre-production checklist
Use this before coding agents or wallets against a new relay base URL:
Contract
Written policy on logging, retention, subprocessors, and model routing. Vague marketing copy is not a contract.
Identity
Document which model ids you pay for. Run pinned-output tests on each id you will use in production.
Injection
Token delta tests and extraction-style probes on your relay URL. Treat inconclusive runs as open questions, not passes.
Tools
If agents install packages or call HTTP tools, compare tool-related strings relay versus direct.
Streams
On Anthropic-style SSE, verify event types, usage monotonicity, and terminal completeness for long answers.
Errors
Catalog error shapes from bad requests. Redact and rotate if anything looks like a secret.
Exit
Plan how you switch base URL or key without rewriting every agent if the relay fails audit.
The api-relay-audit open-source tool
The project toby-bridges/api-relay-audit on GitHub is a local audit runner for AI API relays and LLM proxies. Docs live at toby-bridges.github.io/api-relay-audit/, read 2026-10-07.
The site states execution stays on your machine. Your API key is sent only to the relay URL you choose. It is not sent to API Relay Audit servers or an extra web checker.
A run produces a Markdown report with findings labeled LOW, MEDIUM, or HIGH. The site read 2026-10-07 states the tool does not certify that a relay is safe, does not replace manual security review, and does not treat inconclusive as clean.
The published site describes fourteen audit steps covering prompt injection signals, model identity checks, context truncation probes, tool-call rewriting checks, error leakage tests, stream integrity, and optional Web3-oriented steps behind a profile flag.
You can run a standalone CLI from your terminal, or use a DeepSeek Harness plugin bundle pinned in their docs. Either way, treat the report as evidence to investigate, not a badge.
Relay API spend during a run is billed by your provider. The project homepage quotes roughly twenty to fifty cents of API usage per audit on 2026-10-07. Your invoice may differ.
When local audit helps and when it does not
| Situation | What audit gives you | What it cannot prove |
|---|---|---|
| New reseller URL | Structured probes and a saved Markdown trail | Future behavior after you ship |
| Coding agent with tools | Signals on rewritten install commands | Safety of every tool your agent invents later |
| Wallet-adjacent agents | Profile-gated Web3 refusal tests on the site, read 2026-10-07 | Protection against a compromised client device |
| Regulated data | Evidence for your own risk file | Legal sign-off without your counsel |
Empyre Relay is not in the model request path
Empyre Relay implements OAuth for AI agents: hosted consent, mandatory S256 PKCE, one-time authorization codes, refresh rotation with family revocation, and fail-closed revoke and introspect. Feature-frozen since 2026-07-10 except scoped GET /relay/oauth/userinfo added 2026-08-06.
The npm package is @empyre/relay-sdk version 1.0.0 on registry.npmjs.org, read 2026-10-07. Relay tokens authorize an agent against third-party APIs on a user's behalf. Model calls go from your app to the provider and never pass through Relay.
Endpoint and billing detail live in Relay API and Relay agents. Secret handling for keys that must never leave storage belongs in AI agent secret storage without server private keys and How to store private keys for an AI agent app.
Where Empyre the company builder fits
Empyre on empyre.dev turns a brief into a live company with eight agents that keep operating after launch. That is a business operator, not an LLM proxy auditor and not a coding seat.
If your problem is only whether a discount gateway alters prompts, stay with relay audit work above. If you need a product, deploy, and ongoing operations without living in an IDE, compare operators instead.
Common questions
Is api-relay-audit the same as Empyre Relay?
No. api-relay-audit is an open-source security probe for third-party LLM HTTP relays. Empyre Relay is OAuth for agents at relay.empyre.dev.
Does a LOW report mean the relay is safe?
The tool docs read 2026-10-07 say it does not certify safety. LOW only means that step did not flag HIGH or MEDIUM under its rules.
Should I paste my production key into a web relay checker?
Prefer local runners that send the key only to your chosen base URL. Web checkers add another party that sees the secret.
How is this different from the API relay pattern article?
API relay for AI agents sorts OAuth brokers and HTTP gateways for integrations. Auditing an LLM proxy in the model path means probes, checklists, and tools like api-relay-audit, not OAuth broker selection.
What was verified on 2026-10-07?
toby-bridges.github.io/api-relay-audit/ and github.com/toby-bridges/api-relay-audit repository description; registry.npmjs.org @empyre/relay-sdk version 1.0.0. No search rankings, star counts, or Empyre customer figures appear here.
Try Empyre free for 3 days
Describe a business in plain words and watch eight AI agents build and deploy it. Starter is free for the first 3 days.
Related
Last updated 2026-10-07. Competitor descriptions reflect each product's publicly documented capabilities at that date; they change often, so check the source before relying on a detail.