Migrating an app to the Anthropic API

The order of operations and the gotchas for moving an existing app or agent onto claude.hep.gg: key per project, system blocks with caching, strict turn order, no temperature, tool loop cap, error meanings, a reference client.

Migrating an app to the Anthropic API

This page is written for whoever does the port, human or agent. It is the order of operations and the handful of things that are different from calling Anthropic directly or from calling an OpenAI-compatible endpoint. Three Team Hydra projects were moved this way in one day and every item below came from one of them.

Checklist, in order

  1. Mint the key
    Dashboard, AI, then Anthropic API, then Create key. Name it after the project (the name shows up in the usage tables, so make it the thing you will search for). Restrict models to the family the app needs, usually sonnet. Set a daily spend cap that is clearly above normal use and clearly below a runaway loop. Copy the key into the project's secret store, never into the repo.
  2. Swap the client
    Replace the OpenAI or Ollama client with an Anthropic Messages call. The official SDK works with baseURL and apiKey set. A raw fetch is about forty lines and avoids a dependency, see the reference client below.
  3. Reshape the prompt
    The OpenAI system role message becomes the first block of a system array, with cache_control on it. Per-request context (the user's record, retrieved documents) goes in later blocks or in the user turn, so the cached prefix stays stable.
  4. Normalize the history
    Anthropic rejects two user turns in a row, two assistant turns in a row, and a history that starts with the assistant. Merge consecutive same-role turns, drop a leading assistant turn, and skip empty turns before you send.
  5. Remove sampling knobs
    Delete temperature, top_p, frequency_penalty and presence_penalty. Sonnet 5.5 returns 400 for temperature outright and the others are OpenAI-only.
  6. Run it once, then read the log
    Make one real request from the deployed app, then open the key's Recent requests on the Anthropic API page. If the app swallowed an upstream error, that log is where it shows first, with the status and the model.
  7. Retire the old key
    Pause the gateway or vendor key the app used before. Do not delete it until the new path has carried real traffic for a while, so a rollback is a config change.

The five things that bite

1. temperature is a 400

Sonnet 5.5 answers 400 invalid_request_error with "temperature is deprecated for this model" the moment the field is present, even at 1. Clients that copy an OpenAI request body over usually carry it. Strip it before you send anything else, or your first request fails and looks like an auth problem.

2. The system prompt is an array, and the gateway adds a line in front

Requests are shaped the way Claude Code shapes them, which means a short fixed sentence is placed at the start of system before your content. Your prompt follows it unchanged. Two consequences:

  • Send system as an array of text blocks, not a string. A string still works (the gateway converts it) but you lose control of caching.
  • Put cache_control: {"type": "ephemeral"} on your main prompt block. Cached reads cost a tenth of a fresh read and the usage object reports cache_read_input_tokens so you can see it working.
"system": [
  { "type": "text", "text": "<your master prompt>", "cache_control": { "type": "ephemeral" } },
  { "type": "text", "text": "<per-request context: records, retrieved docs>" }
]

3. Turns must alternate and start with the user

Anthropic validates the messages array strictly. A DM history or a chat log that was fine against an OpenAI endpoint will often fail with 400 invalid_request_error about roles. The normalizer is small:

  • drop any system role entries (they go in system above)
  • drop assistant turns until the first user turn
  • merge consecutive same-role turns with a blank line between them
  • skip turns whose content is empty after trimming

4. Tool use is a loop, so cap it

If you give the model tools (document search, a lookup), a response with stop_reason: "tool_use" means you run the tool and send the result back as a tool_result block in a new user turn, then call again. Bound it: four rounds is plenty, and on the last round send the request without tools so the model has to answer in words.

5. What the errors mean

Statuserror.typeWhat it meansWhat to do
401authentication_errorMissing, wrong, or paused keyCheck the header name and whether the key is paused on the dashboard
403permission_errorThe key's model policy does not allow this modelChange the model or edit the key's policy
429rate_limit_error with a cap messageThe key hit its daily spend or request capWait for midnight UTC or raise the cap
429rate_limit_error "Usage limit reached on all your accounts"Every pooled account is at a usage limitBack off for minutes, not seconds. Tell the user the assistant is busy
503api_errorNo account connected, or the serving account needs re-loginFix it on the Accounts page
400invalid_request_errorYour body: temperature, role order, a bad blockRead the message, it names the field

A 429 from the pool is not a per-second rate limit. Retrying immediately burns requests against your daily cap and never succeeds sooner. One retry after a short pause, then surface a friendly failure.

Reference client

Plain fetch, no SDK. System blocks, turn normalization, the tool loop with a cap, and the usage object returned so the caller can meter cost. Drop it in and replace the two TODOs.

claude.ts
const BASE = process.env.CLAUDE_BASE_URL ?? "https://claude.hep.gg";
const KEY = process.env.CLAUDE_API_KEY!; // sk-hepgg-...
const MODEL = process.env.CLAUDE_MODEL ?? "claude-sonnet-5-5";
const MAX_TOOL_ROUNDS = 4;
 
type Turn = { role: "user" | "assistant" | "system"; content: string };
type Block = Record<string, unknown>;
type Tool = { name: string; description: string; input_schema: Record<string, unknown>; run: (input: any) => Promise<string> };
 
export function normalizeTurns(history: Turn[]): { role: "user" | "assistant"; content: string }[] {
  const out: { role: "user" | "assistant"; content: string }[] = [];
  for (const t of history) {
    if (t.role === "system") continue;
    const content = (t.content ?? "").trim();
    if (!content) continue;
    if (out.length === 0 && t.role !== "user") continue; // must start with the user
    const last = out[out.length - 1];
    if (last && last.role === t.role) last.content += "\n\n" + content; // no two in a row
    else out.push({ role: t.role, content });
  }
  return out;
}
 
export async function ask(opts: { system: string; context?: string; history: Turn[]; tools?: Tool[]; maxTokens?: number }) {
  const system: Block[] = [{ type: "text", text: opts.system, cache_control: { type: "ephemeral" } }];
  if (opts.context) system.push({ type: "text", text: opts.context });
  const messages: Block[] = normalizeTurns(opts.history);
  const usage = { input: 0, output: 0, cacheWrite: 0, cacheRead: 0 };
  const toolDefs = (opts.tools ?? []).map(({ run, ...t }) => t);
 
  for (let round = 0; ; round++) {
    const lastRound = round >= MAX_TOOL_ROUNDS;
    const body: Block = { model: MODEL, max_tokens: opts.maxTokens ?? 1024, system, messages };
    if (toolDefs.length && !lastRound) body.tools = toolDefs; // no temperature, ever
    const res = await fetch(`${BASE}/v1/messages`, {
      method: "POST",
      headers: { "x-api-key": KEY, "anthropic-version": "2023-06-01", "content-type": "application/json" },
      body: JSON.stringify(body),
    });
    const data: any = await res.json();
    if (!res.ok) throw Object.assign(new Error(data?.error?.message ?? `HTTP ${res.status}`), { status: res.status, type: data?.error?.type });
    usage.input += data.usage?.input_tokens ?? 0;
    usage.output += data.usage?.output_tokens ?? 0;
    usage.cacheWrite += data.usage?.cache_creation_input_tokens ?? 0;
    usage.cacheRead += data.usage?.cache_read_input_tokens ?? 0;
 
    const text = data.content.filter((b: any) => b.type === "text").map((b: any) => b.text).join("");
    const calls = data.content.filter((b: any) => b.type === "tool_use");
    if (data.stop_reason !== "tool_use" || calls.length === 0) return { text, usage };
 
    messages.push({ role: "assistant", content: data.content });
    const results: Block[] = [];
    for (const c of calls) {
      const tool = opts.tools!.find((t) => t.name === c.name);
      const content = tool ? await tool.run(c.input).catch((e) => `error: ${e.message}`) : `unknown tool ${c.name}`;
      results.push({ type: "tool_result", tool_use_id: c.id, content });
    }
    messages.push({ role: "user", content: results });
  }
}

The error thrown carries status and type, so the caller can map 429 to "busy, try later" and everything else to a logged failure. Add streaming only if the product needs it; the shape is Anthropic's SSE and the SDK handles it.

Metering cost inside the app

The dashboard already prices every request, per key and per account. If the app needs its own budget (a per-user allowance, a daily stop), price the usage object at list rates per million tokens. For claude-sonnet-5-5:

FieldUSD per MTok
input_tokens3.00
output_tokens15.00
cache_creation_input_tokens3.75
cache_read_input_tokens0.30

Keep that figure as an estimate in your own storage (Redis is enough for a counter that is read before and written after each reply). The subscription accounts absorb the actual usage, so this is a guard rail, not a bill.

Verifying the port

  • The first request from the deployed app appears under the key in Recent requests with status 200 and the model you expected.
  • usage.cache_read_input_tokens is non-zero from the second request on, which proves the system block is cached.
  • A deliberate bad model name returns 403 permission_error if you restricted the key, which proves the policy is attached.
  • The old key shows no new requests after the switch.