Going to production
A demo agent can assume the happy path. A production agent can't: runs end for many reasons, tools cost money, and some actions are dangerous. This guide is the checklist for shipping an agent you can trust to run unattended.
Handle every run status
A run doesn't always complete. result.status tells you how it ended, and you should branch on it rather than assuming success:
const result = await agent.runDetailed(messages)
switch (result.status) {
case 'completed':
return result.outputText
case 'max_steps':
case 'budget_tokens':
case 'budget_calls':
// Hit a limit you set — decide whether to continue or report a partial answer.
return handlePartial(result)
case 'too_many_failures':
case 'repeated_tool_call':
// The agent got stuck. Log it; don't silently retry forever.
throw new Error(`Agent stalled: ${result.status}`)
case 'aborted':
return null // you (or a timeout) cancelled it
case 'failed':
throw new Error('Agent run failed')
}
Cap cost and loops with budgets
Pass limits in the run options so a single run can't run away — in tokens, tool calls, or turns:
const result = await agent.runDetailed(messages, {
maxSteps: 20, // agentic turns
maxTokens: 100_000, // total token budget
maxCalls: 40, // total model/tool calls
maxConsecutiveFailures: 5, // stop after repeated failures
maxRepeatedCalls: 3, // stop if it calls the same thing over and over
})
When a budget stops a run, the status reflects which one (budget_tokens, budget_calls, max_steps). Set these deliberately — they're your ceiling on cost and your protection against loops.
Control spend
- Read
result.usageafter every run —{ inputTokens, outputTokens, cachedInputTokens, calls }— and track it.cachedInputTokensare cheaper, so prompt caching is already working for you when it's non-zero. - Pick the right tier.
poolot-miniandpoolot-standardare inexpensive; reservepoolot-profor work that needs it. See model tiers. - Check the balance before a big batch with
await agent.getMe()— it returns your tier and remainingbalance.
Gate risky actions with approvals
By default a custom tool runs as soon as the model calls it. For anything that changes the world — places an order, sends a message, spends money — mark it approval: "ask" and install a handler that decides:
import { setApprovalHandler, approvalRules } from '@poolot/lily-web'
setApprovalHandler(
approvalRules({
allow: [{ tool: 'run_command', argMatches: /^git (status|diff|log)\b/ }],
deny: [{ tool: 'run_command', argMatches: /\brm\s+-rf\b/ }],
otherwise: (tool) => false, // default-deny anything not explicitly allowed
}),
)
approvalRules reads the tool's arguments, not just its name — so git status and rm -rf get different answers. deny always wins over allow, regardless of order. Return a promise from otherwise if you need to ask a human. Default-deny (returning false) is the safe posture for unattended runs.
Make custom tools robust
inputSchema— pass a Zod schema; arguments are validated before your handler runs, and the model is told what was wrong so it can fix the next call.timeoutMs— bound each call (default 30s). A timeout is never retried.retries— extra attempts after a throw (default 0), for flaky I/O.- Throwing is a normal signal, not a crash — it marks the tool result failed and lets the model recover.
Answer questions without a human
If the agent asks a question mid-run (the ask_user tool) and nobody's there, the run stalls. For headless jobs, install an input handler so questions are answered programmatically:
import { setInputHandler } from '@poolot/lily-web'
setInputHandler((question) => {
// Return a default, or look the answer up. Return null to decline.
return 'Proceed with the safe default.'
})
See what's happening
- Hooks — pass
hookstocreateLilyAgentto observe the run at each stage (run_start,tool_call_before,tool_call_after,error,run_end). Use them for logging, metrics, and tracing. - Error reporting — configure a Sentry
dsnin the load options to capture the SDK's own errors, and callcaptureException(err, { component, op })to report your own. - Cancellation — call
agent.abort()to stop a run in flight; long-running tool handlers should checkcontext.signaland bail when it aborts.
Keep your key server-side
The biggest production mistake is shipping a provider key to the browser.
- Never embed a long-lived programmatic key in client code — it's visible to anyone who opens dev tools.
- In the browser, pass a short-lived signed-in user token as
apiKey, or route requests through your own same-origin proxy withbaseUrl(default/api/agent-server) that holds the key server-side. - Bring-your-own-key (
provider) is a Node/harness capability for exactly this reason — see Browser vs Node.
Production checklist
- Branch on every
result.status, not justcompleted. - Set
maxSteps/maxTokens/maxCallsbudgets. - Track
result.usage; pick the cheapest model tier that works. -
approval: "ask"on every world-changing tool, with a default-deny handler. - Validate tool inputs (
inputSchema) and bound them (timeoutMs). - Install
setInputHandlerfor unattended runs. - Wire
hooksand a Sentrydsnfor observability. - Key stays server-side; browser uses a user token or a proxy.
See also
- Building an agent — the walkthrough these practices harden.
- createLilyAgent reference — every option in detail.
- Permissions & safety — how the agent's own guardrails work.