Migrating an app to the Anthropic API
The order of operations and the gotchas for moving an existing app or agent onto claude.hep.gg: key per project, system blocks with caching, strict turn order, no temperature, tool loop cap, error meanings, a reference client.
Migrating an app to the Anthropic API
This page is written for whoever does the port, human or agent. It is the order of operations and the handful of things that are different from calling Anthropic directly or from calling an OpenAI-compatible endpoint. Three Team Hydra projects were moved this way in one day and every item below came from one of them.
Checklist, in order
- Mint the keyDashboard, AI, then Anthropic API, then Create key. Name it after the project (the name shows up in the usage tables, so make it the thing you will search for). Restrict models to the family the app needs, usually
sonnet. Set a daily spend cap that is clearly above normal use and clearly below a runaway loop. Copy the key into the project's secret store, never into the repo. - Swap the clientReplace the OpenAI or Ollama client with an Anthropic Messages call. The official SDK works with
baseURLandapiKeyset. A rawfetchis about forty lines and avoids a dependency, see the reference client below. - Reshape the promptThe OpenAI
systemrole message becomes the first block of asystemarray, withcache_controlon it. Per-request context (the user's record, retrieved documents) goes in later blocks or in the user turn, so the cached prefix stays stable. - Normalize the historyAnthropic rejects two user turns in a row, two assistant turns in a row, and a history that starts with the assistant. Merge consecutive same-role turns, drop a leading assistant turn, and skip empty turns before you send.
- Remove sampling knobsDelete
temperature,top_p,frequency_penaltyandpresence_penalty. Sonnet 5.5 returns400fortemperatureoutright and the others are OpenAI-only. - Run it once, then read the logMake one real request from the deployed app, then open the key's Recent requests on the Anthropic API page. If the app swallowed an upstream error, that log is where it shows first, with the status and the model.
- Retire the old keyPause the gateway or vendor key the app used before. Do not delete it until the new path has carried real traffic for a while, so a rollback is a config change.
The five things that bite
1. temperature is a 400
Sonnet 5.5 answers 400 invalid_request_error with "temperature is deprecated for this model" the moment the field is present, even at 1. Clients that copy an OpenAI request body over usually carry it. Strip it before you send anything else, or your first request fails and looks like an auth problem.
2. The system prompt is an array, and the gateway adds a line in front
Requests are shaped the way Claude Code shapes them, which means a short fixed sentence is placed at the start of system before your content. Your prompt follows it unchanged. Two consequences:
- Send
systemas an array of text blocks, not a string. A string still works (the gateway converts it) but you lose control of caching. - Put
cache_control: {"type": "ephemeral"}on your main prompt block. Cached reads cost a tenth of a fresh read and theusageobject reportscache_read_input_tokensso you can see it working.
"system": [
{ "type": "text", "text": "<your master prompt>", "cache_control": { "type": "ephemeral" } },
{ "type": "text", "text": "<per-request context: records, retrieved docs>" }
]3. Turns must alternate and start with the user
Anthropic validates the messages array strictly. A DM history or a chat log that was fine against an OpenAI endpoint will often fail with 400 invalid_request_error about roles. The normalizer is small:
- drop any
systemrole entries (they go insystemabove) - drop assistant turns until the first user turn
- merge consecutive same-role turns with a blank line between them
- skip turns whose content is empty after trimming
4. Tool use is a loop, so cap it
If you give the model tools (document search, a lookup), a response with stop_reason: "tool_use" means you run the tool and send the result back as a tool_result block in a new user turn, then call again. Bound it: four rounds is plenty, and on the last round send the request without tools so the model has to answer in words.
5. What the errors mean
| Status | error.type | What it means | What to do |
|---|---|---|---|
401 | authentication_error | Missing, wrong, or paused key | Check the header name and whether the key is paused on the dashboard |
403 | permission_error | The key's model policy does not allow this model | Change the model or edit the key's policy |
429 | rate_limit_error with a cap message | The key hit its daily spend or request cap | Wait for midnight UTC or raise the cap |
429 | rate_limit_error "Usage limit reached on all your accounts" | Every pooled account is at a usage limit | Back off for minutes, not seconds. Tell the user the assistant is busy |
503 | api_error | No account connected, or the serving account needs re-login | Fix it on the Accounts page |
400 | invalid_request_error | Your body: temperature, role order, a bad block | Read the message, it names the field |
A 429 from the pool is not a per-second rate limit. Retrying immediately burns requests against your daily cap and never succeeds sooner. One retry after a short pause, then surface a friendly failure.
Reference client
Plain fetch, no SDK. System blocks, turn normalization, the tool loop with a cap, and the usage object returned so the caller can meter cost. Drop it in and replace the two TODOs.
const BASE = process.env.CLAUDE_BASE_URL ?? "https://claude.hep.gg";
const KEY = process.env.CLAUDE_API_KEY!; // sk-hepgg-...
const MODEL = process.env.CLAUDE_MODEL ?? "claude-sonnet-5-5";
const MAX_TOOL_ROUNDS = 4;
type Turn = { role: "user" | "assistant" | "system"; content: string };
type Block = Record<string, unknown>;
type Tool = { name: string; description: string; input_schema: Record<string, unknown>; run: (input: any) => Promise<string> };
export function normalizeTurns(history: Turn[]): { role: "user" | "assistant"; content: string }[] {
const out: { role: "user" | "assistant"; content: string }[] = [];
for (const t of history) {
if (t.role === "system") continue;
const content = (t.content ?? "").trim();
if (!content) continue;
if (out.length === 0 && t.role !== "user") continue; // must start with the user
const last = out[out.length - 1];
if (last && last.role === t.role) last.content += "\n\n" + content; // no two in a row
else out.push({ role: t.role, content });
}
return out;
}
export async function ask(opts: { system: string; context?: string; history: Turn[]; tools?: Tool[]; maxTokens?: number }) {
const system: Block[] = [{ type: "text", text: opts.system, cache_control: { type: "ephemeral" } }];
if (opts.context) system.push({ type: "text", text: opts.context });
const messages: Block[] = normalizeTurns(opts.history);
const usage = { input: 0, output: 0, cacheWrite: 0, cacheRead: 0 };
const toolDefs = (opts.tools ?? []).map(({ run, ...t }) => t);
for (let round = 0; ; round++) {
const lastRound = round >= MAX_TOOL_ROUNDS;
const body: Block = { model: MODEL, max_tokens: opts.maxTokens ?? 1024, system, messages };
if (toolDefs.length && !lastRound) body.tools = toolDefs; // no temperature, ever
const res = await fetch(`${BASE}/v1/messages`, {
method: "POST",
headers: { "x-api-key": KEY, "anthropic-version": "2023-06-01", "content-type": "application/json" },
body: JSON.stringify(body),
});
const data: any = await res.json();
if (!res.ok) throw Object.assign(new Error(data?.error?.message ?? `HTTP ${res.status}`), { status: res.status, type: data?.error?.type });
usage.input += data.usage?.input_tokens ?? 0;
usage.output += data.usage?.output_tokens ?? 0;
usage.cacheWrite += data.usage?.cache_creation_input_tokens ?? 0;
usage.cacheRead += data.usage?.cache_read_input_tokens ?? 0;
const text = data.content.filter((b: any) => b.type === "text").map((b: any) => b.text).join("");
const calls = data.content.filter((b: any) => b.type === "tool_use");
if (data.stop_reason !== "tool_use" || calls.length === 0) return { text, usage };
messages.push({ role: "assistant", content: data.content });
const results: Block[] = [];
for (const c of calls) {
const tool = opts.tools!.find((t) => t.name === c.name);
const content = tool ? await tool.run(c.input).catch((e) => `error: ${e.message}`) : `unknown tool ${c.name}`;
results.push({ type: "tool_result", tool_use_id: c.id, content });
}
messages.push({ role: "user", content: results });
}
}The error thrown carries status and type, so the caller can map 429 to "busy, try later" and everything else to a logged failure. Add streaming only if the product needs it; the shape is Anthropic's SSE and the SDK handles it.
Metering cost inside the app
The dashboard already prices every request, per key and per account. If the app needs its own budget (a per-user allowance, a daily stop), price the usage object at list rates per million tokens. For claude-sonnet-5-5:
| Field | USD per MTok |
|---|---|
input_tokens | 3.00 |
output_tokens | 15.00 |
cache_creation_input_tokens | 3.75 |
cache_read_input_tokens | 0.30 |
Keep that figure as an estimate in your own storage (Redis is enough for a counter that is read before and written after each reply). The subscription accounts absorb the actual usage, so this is a guard rail, not a bill.
Verifying the port
- The first request from the deployed app appears under the key in Recent requests with status
200and the model you expected. usage.cache_read_input_tokensis non-zero from the second request on, which proves the system block is cached.- A deliberate bad model name returns
403 permission_errorif you restricted the key, which proves the policy is attached. - The old key shows no new requests after the switch.