Commands, settings and headers
The five commands people use and the settings each is most often given, the headers your code can send, and the environment variables and settings file that can stand in for a setting. Dollar amounts here are examples.
run: one agent, start to finish
Starts a gateway and the dashboard for one agent, runs the agent through them, and prints a receipt when it finishes. Put your own agent's command after the two dashes; it keeps using its own provider key.
# watch only: record every call, refuse none axigate-finops run -- claude # stop the run once it has spent $5 or made 200 calls axigate-finops run --max-spend 5 --max-calls 200 -- claude # let every call through, and mark the ones a $5 limit would have stopped axigate-finops run --max-spend 5 --shadow -- claude # a scheduled job: one $20 limit a day per key, across every run on this computer axigate-finops run --max-spend-per-key-day 20 -- ./nightly.sh # send a short message to a web address if the run is stopped axigate-finops run --max-spend 5 --stop-alert-url https://example.com/alerts/axigate -- claude # see it work first: invented data and a runaway it stops, no agent or key needed axigate-finops run --demo
serve: a gateway and dashboard that stay running
Starts the gateway on 127.0.0.1:8787 and the dashboard on 127.0.0.1:8788, and keeps them running for your own code or for several agents. Point your code's AI base URL at the gateway's address.
# watch only axigate-finops serve # stop any run at $5 or 200 calls axigate-finops serve --max-spend-per-run 5 --max-calls-per-run 200 # pause a key once it has spent $50 in a day (UTC) axigate-finops serve --max-spend-per-key-day 50 # try a limit first: every call goes through, and the ones it would stop are marked axigate-finops serve --max-spend-per-run 5 --shadow # send a short message to a web address when a run or key is stopped axigate-finops serve --stop-alert-url https://example.com/alerts/axigate
gateway: the gateway on its own
The same gateway without the dashboard: to report to a hosted workspace, or as one shared gateway for a team. It takes the same limits serve does.
# report every call to your hosted workspace AXIGATE_INGEST_KEY=<the key from Connect a gateway> axigate-finops gateway --ingest-url https://<your workspace>/api/ingest # one shared gateway for the team: open to your network, one key per person axigate-finops gateway --listen 0.0.0.0:8787 --gateway-keys ./keys.toml # the same limits as on serve axigate-finops gateway --max-spend-per-run 5 --max-spend-per-key-day 50
analyze: compare a bill with the record
Reads the provider's usage or cost files together with the gateway's ledger file, on your computer, and prints what was recorded, what was billed and the difference. Nothing is uploaded.
# the provider's file and the gateway's ledger file, in any order axigate-finops analyze ~/Downloads/claude_api_cost_2026_08.csv <ledger file> --out ~/Downloads # also write a FOCUS file for other FinOps tools axigate-finops analyze ~/Downloads/claude_api_cost_2026_08.csv <ledger file> --out ~/Downloads --focus ~/Downloads/focus.csv # say who owns each provider key or project axigate-finops analyze ~/Downloads/claude_api_cost_2026_08.csv --owners owners.csv
It writes report.txt and statement.csv into the folder --out names, or the current folder if you leave it off. Reading the same file twice never counts anything twice.
version: which build you have
Prints which version of AxiGate you have. To get a newer one, see Keep AxiGate current.
axigate-finops version
Headers your code can send
Headers on your AI calls tell the gateway who each call belongs to. It reads them and takes them off before the call goes on, so the provider never sees them; when run sends your calls through a shared AxiGate gateway, the run id and these headers go on to that gateway, never to a provider. Claude Code needs none: each session is a run on its own.
X-AxiGate-Run: the run the call belongs to. Per-run limits count by it.X-AxiGate-Agent: which agent made the call.X-AxiGate-Team,X-AxiGate-Project,X-AxiGate-Customer: who the spend belongs to.X-AxiGate-Max-Spend: a dollar limit for the run named byX-AxiGate-Run, set from your code with no restart; if the gateway also has a per-run limit, the lower of the two applies. A value that isn't a positive dollar amount gets the call refused.X-AxiGate-Key: a name for the provider key, shown on Keys. The key's limits follow the key itself, so a new name can't get around them.X-AxiGate-Gateway-Key: your own key for a shared gateway. When you start your agent with run and your gateway key, run adds it for you.
client = OpenAI(base_url="http://127.0.0.1:8787/v1")
resp = client.chat.completions.create(
model="gpt-4o", messages=msgs,
extra_headers={
"X-AxiGate-Run": run_id,
"X-AxiGate-Agent": "planner",
"X-AxiGate-Max-Spend": "5.00",
})Environment variables
Three of the settings above can also come from the environment, so a key or an address never has to be typed on the command line. A setting on the command line wins over its variable.
AXIGATE_INGEST_KEYstands in for--ingest-key, on gateway and serve.AXIGATE_STOP_ALERT_URLstands in for--stop-alert-url, on gateway, serve and run. Either one wins over an address saved on the Limits page.AXIGATE_GATEWAY_KEYstands in for--gateway-key, on run.
run keeps every AXIGATE_ variable out of your agent's environment, AXIGATE_INGEST_KEY included: run sends the stop message and adds the gateway key itself, so the agent never holds either, and it needs none of the others.
export AXIGATE_INGEST_KEY=<the key from Connect a gateway> axigate-finops gateway --ingest-url https://<your workspace>/api/ingest export AXIGATE_GATEWAY_KEY=<your key> axigate-finops run --gateway http://<server>:8787 -- claude
Settings in a file
serve and gateway also read an axigate.toml from the folder they start in: limits and loop settings, written without the dashes. Commit it next to your agent and everyone's gateway starts with the same limits.
max-spend-per-run = 5 max-spend-per-key-day = 50 max-calls-per-run = 400
A setting on the command line wins over the file.
Your own prices
Prices ship with each release; the gateway never goes online to learn them. For a rate you negotiated, or a model that came out after your release, put a prices.json in the same folder as the ledger (/data/prices.json in the Docker image). gateway, serve and run read it once when they start, and your prices win over the built-in list for those models.
{"claude-opus-5-5": {"input": 3.60, "output": 18, "cache_read": 0.18}}Prices are dollars per million tokens: input and output, and optionally cache_read, cache_write_5m and cache_write_1h. They are the standard price: a call on one of those models that ran in fast mode or another tier is priced at your price and marked as an estimate. A file that can't be read stops the start and says why. Calls it priced show the price list's version with +local on the end.
A model on neither list is still served, and its cost is an estimate. On the provider's own API it is priced at the highest current price its model family has in this release's list, or at the provider's flagship line when its name matches no family, so a limit is less likely to stop late, and its call names what it was priced as. A new model priced above all of those still counts low until you add it to prices.json. On any other server, such as a local model server, it is priced at $3 per million tokens.