Anthropic API

An Anthropic-compatible Messages API at claude.hep.gg, served by your own pooled Claude accounts, with per-key model, account and daily-cap policy and per-key usage tracking.

Anthropic API

Use your pooled Claude subscription accounts from your own code. https://claude.hep.gg is an Anthropic-compatible Messages API: point the official Anthropic SDK (or any client that speaks the Anthropic API) at it with a Hep.gg API key, and each request is served by one of the Claude accounts connected on your Accounts page. Account selection, usage-limit rotation and keep-warm behave exactly as they do for Claude Code through the proxy, so your apps and harnesses share the same pool logic without installing ccas.

Create a key

  1. Open AI, then Anthropic API in the dashboard and choose Create key.
  2. Name it, then optionally restrict it: which model families it may use, which of your pooled accounts it should prefer, and a daily spend or request cap.
  3. Copy the key once. It starts with sk-hepgg- and is also re-viewable from the key card later.

A key with no restrictions may use every model your accounts can run and draws from your whole pool.

Call it

The Anthropic SDKs need two settings: the base URL and the key. Everything else, including streaming, tool use and prompt caching, works as it does against Anthropic directly.

curl
curl https://claude.hep.gg/v1/messages \
  -H "x-api-key: $HEPGG_CLAUDE_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-5-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Hello"}]
  }'

The key is accepted as x-api-key (what the SDKs send) or as Authorization: Bearer.

Endpoints

POSThttps://claude.hep.gg/v1/messagesAuth required
Create a message. Same request and response shape as Anthropic's Messages API, streaming included.
POSThttps://claude.hep.gg/v1/messages/count_tokensAuth required
Count the tokens a request would use. Does not count toward the key's daily caps.
GEThttps://claude.hep.gg/v1/modelsAuth required
The models this key may use, in Anthropic's list shape.

Other paths return a 404 in Anthropic's error envelope. Batches and the Files API are not served here.

What a key can be limited to

Key policy
Models
listoptional
Model families (sonnet, opus, haiku, fable) or exact model ids. A request for anything else gets a 403 permission_error. Empty means every model your accounts can run.
Accounts
listoptional
Pooled accounts the key should use first. When all of them are at a usage limit the rest of your pool serves, and when nothing can serve you get the usual 429. Empty means the whole pool, picked the same way the proxy picks for Claude Code.
Daily spend cap
USDoptional
Estimated spend for the UTC day, priced at public list rates. Once reached the key returns 429 rate_limit_error until midnight UTC.
Daily request cap
numberoptional
Requests for the UTC day. Same 429 once reached.

Pausing a key stops it within seconds without deleting it. Rotating it mints a new secret and keeps the name, policy and history.

How the pool is used

  • Every request runs on one of your own accounts. Keys never use anyone else's capacity, and nobody else's keys can use yours.
  • A usage limit on the serving account rotates the request to another account that can serve, inside the same call, so your client sees a normal response rather than a 429. Only when nothing can serve does the 429 reach you.
  • Key traffic never changes the account your own Claude Code sessions are using. It rotates for its own requests only.
  • If you have keep-warm enabled, a key request starts the five-hour windows on your other accounts the same way Claude Code traffic does.

Usage and cost

Every request is recorded against your account, the key that made it and the pooled account that served it. The Anthropic API page shows requests, tokens and estimated cost per key with a per-account and per-model breakdown, a daily chart, and a live log of recent calls. The same usage also appears in your overall AI spend everywhere it is shown: the usage dashboard, ccas, the status line and the usage widgets.

Estimated cost uses public list prices for the model that served the request. It is a way to watch what your own apps are doing, not a bill: the usage comes out of the subscription accounts you connected.