Stop a runaway agent
An agent stuck in a loop can keep spending all night. Give each run a limit: once a run reaches it, its next call is refused before it reaches the provider, so it costs nothing.
You will learn
- How to put a spending limit on each agent run
- How to change limits while AxiGate runs, on the Limits page
- How to set a tighter one from your own code
- How to try a limit before it refuses anything
- How to get a message the moment a run is stopped
- How to set where that message goes on the Limits page
- Where a run stuck repeating a step would have been stopped
Where this works
You don't need a plan to read on.| Free editionyour own machines | Hosted Freehosted | Teamhosted | Businesshosted | |
|---|---|---|---|---|
| Limit what each run can spend | ||||
| Change limits on the Limits page, no restart | — | — | — | |
| Set a limit from your code | ||||
| Try a limit without enforcing it | ||||
| Get a message when a run is stopped | ||||
| Set where stop messages go on the Limits page | — | — | — | |
| See where a stuck step would have stopped a run |
Give each run a limit
A run is one piece of an agent's work, such as one Claude Code session. The $5 below is an example: pick what a normal run of yours costs, with room to spare.
The limit goes on the gateway that reports to your workspace. Add it to the command you started the gateway with, start the gateway again, then start your agent through it.
AXIGATE_INGEST_KEY=<the key from Connect a gateway> axigate-finops gateway --ingest-url https://<your workspace>/api/ingest --max-spend-per-run 5 axigate-finops run --gateway http://127.0.0.1:8787 -- claude
Change limits while AxiGate runs
Free editionOn your own machines, the dashboard has a Limits page. It lays the limits out in the order the gateway checks each call, shows what they stopped today, and changes them with no restart: click a figure, type the new one, and Apply. The next call follows it.
Under Spend per run, the page shows what your recent runs of each command cost (median, 95th percentile from 10 runs, largest) and what the figure you are typing would have stopped. It is a place to start a limit from your own runs, not a guess.
Under Runs that repeat themselves, the stuck-step check's rows are marked Recording only: they are not settings and do not count as limits that are on. Each says how many runs it would have stopped in the last 7 days and what those runs spent after that point, estimated, and names the latest runs, which open their receipts.
A change is saved in limits.json, in the same folder as the ledger. Every serve, gateway and run whose ledger is in that folder reads it when it starts and looks again every two seconds, so the change reaches agents already running too. Runs and keys already stopped stay stopped until you resume them. Switching Watch only off is the exception for what it only marked: those runs and keys are checked against the limits then in force at their next call.
A limit given when AxiGate started, on the command line or in axigate.toml, always wins: the page shows it locked, with the option that set it, and does not change it. Start without the option to change that limit on the page. A limit given at start without --shadow also keeps Watch only off, so it always refuses.
# spend per run and calls per run can both be changed on the page axigate-finops serve # calls per run is locked at 200 here; the rest can be changed on the page axigate-finops serve --max-calls-per-run 200
Set a limit from your code
If your own code starts the runs, it can give one run a tighter limit with the X-AxiGate-Max-Spend header. No restart, no change to the gateway.
Send it together with X-AxiGate-Run, because a limit needs a run to belong to.
client = OpenAI(base_url="http://127.0.0.1:8787/v1")
resp = client.chat.completions.create(
model="gpt-4o", messages=msgs,
extra_headers={
"X-AxiGate-Run": run_id,
"X-AxiGate-Max-Spend": "5.00",
})Try a limit before it refuses anything
Not sure where to set it? Add --shadow: every call still goes through, and the ones the limit would have stopped are marked.
AXIGATE_INGEST_KEY=<the key from Connect a gateway> axigate-finops gateway --ingest-url https://<your workspace>/api/ingest --max-spend-per-run 5 --shadow
Watch for a week, then start again without it.
The marks show on the run's page, on Needs a look, and in a list on the Keys page. With run, the receipt counts them too, and --shadow works only alongside --max-spend or --max-spend-per-key-day.
See where a stuck step would have stopped a run
A stuck agent often calls the same tool again and again: it polls a build that is still running, runs the same failing install, reads the same log. The gateway counts that on every run, whether or not you set a limit.
- The same tool call gets the same answer 10 times in a row.
- A tool fails the same way 3 times in a row.
- The same tool call is made 20 times in a row, whatever comes back: for an answer that changes only because it carries the time.
Each counts only once the run has spent at least $1 on those repeated rounds: a status check that costs a fraction of a cent is not worth a stop. What counts is what the calls that made the rounds cost, not what the run's other agents spent meanwhile.
Each time a repeated step crosses one of these, the call where it happened is marked on its row. The run's page, and the receipt run prints when the agent exits, then say where the first one would have stopped the run, for example "AxiGate would have stopped this run at call 12: the same tool call got the same answer 10 times in a row", and what was spent after that point.
Get a message when a run is stopped
Give the gateway a web address (URL), and the moment it stops a run or a key it sends a short message there: which agent, which run, why, and the spend so far.
AXIGATE_INGEST_KEY=<the key from Connect a gateway> axigate-finops gateway --ingest-url https://<your workspace>/api/ingest --max-spend-per-run 5 --stop-alert-url https://example.com/alerts/axigate
run sends the message before it exits: it waits a few seconds at most, and the receipt says so if the message was still on its way or did not arrive. When run sends your agent through a shared gateway with --gateway, that gateway stops the run, so give it the address instead.
Set where stop messages go on the Limits page
Free editionOn your own machines you can set the address on the dashboard instead of when AxiGate starts. On the Limits page, under When a limit is reached, paste the same kind of address --stop-alert-url takes into Send stop messages to, and Save. The bell at the top of the dashboard has a Stop messages row that leads there.
There is no restart: the next stop sends its message to the new address. Send a test message posts one line, "AxiGate test: stop messages reach this channel", and nothing else, then says whether it was delivered, and if not, why. Change and Remove sit beside the address.
The address is saved in limits.json beside the ledger, readable only by you, and every serve, gateway and run whose ledger is in that folder picks it up within two seconds. The address is a secret, since anyone who has it can post to your channel, so once it is saved the page shows only its last four characters, and it never appears in a log line or a stop message. It must start with https://; http:// works only on 127.0.0.1 or localhost, for testing, and only saved on that computer's own dashboard. A message is never sent on to wherever the address redirects: a redirect counts as not delivered.
An address given at start, with --stop-alert-url or AXIGATE_STOP_ALERT_URL, still wins: the page shows it locked by that name and by its last four characters, and the test goes to it. axigate.toml may not set it. With --shared-counter a saved address is not used, like the saved limits: give each gateway --stop-alert-url. The demo sends no stop messages.
When a run is stopped
The agent gets an error back instead of an answer. The run's page names the limit it hit and the call where it happened, and the run shows on Needs a look.
If the limit was right, there is nothing to do: the next run starts fresh. If it was too tight, raise it where you set it and start the run again; when run set the limit itself, its receipt suggests one twice as high, and when your code set a lower one with X-AxiGate-Max-Spend, the receipt says to raise it there.
When run stops an agent for repeating itself, or for starting too many sub-agents, the stop was not about what it spent, and a higher spend or call limit would only let it go on longer. The receipt says why it was stopped and asks you to check the agent instead. If the agent was doing what you meant, --stop-loops=false turns those checks off.
Recap
- A limit on each run refuses its next call before it costs anything.
- On your own machines, the Limits page changes limits with no restart; a limit given at start always wins.
- Your code can make one run's limit tighter with a header, never looser.
- Shadow mode marks what a limit would stop, so you can watch before you enforce.
- A stop message tells you the moment a run or a key is stopped; on your own machines, the Limits page sets where it goes.
- A run stuck repeating a step is recorded where it would have been stopped; nothing is stopped for it yet.