Skip to main content

Going to production

A demo agent can assume the happy path. A production agent can't: runs end for many reasons, tools cost money, and some actions are dangerous. This guide is the checklist for shipping an agent you can trust to run unattended.

Handle every run status

A run doesn't always complete. result.status tells you how it ended, and you should branch on it rather than assuming success:

const result = await agent.runDetailed(messages)

switch (result.status) {
  case 'completed':
    return result.outputText
  case 'max_steps':
  case 'budget_tokens':
  case 'budget_calls':
    // Hit a limit you set — decide whether to continue or report a partial answer.
    return handlePartial(result)
  case 'too_many_failures':
  case 'repeated_tool_call':
    // The agent got stuck. Log it; don't silently retry forever.
    throw new Error(`Agent stalled: ${result.status}`)
  case 'aborted':
    return null // you (or a timeout) cancelled it
  case 'failed':
    throw new Error('Agent run failed')
}

Cap cost and loops with budgets

Pass limits in the run options so a single run can't run away — in tokens, tool calls, or turns:

const result = await agent.runDetailed(messages, {
  maxSteps: 20,               // agentic turns
  maxTokens: 100_000,         // total token budget
  maxCalls: 40,               // total model/tool calls
  maxConsecutiveFailures: 5,  // stop after repeated failures
  maxRepeatedCalls: 3,        // stop if it calls the same thing over and over
})

When a budget stops a run, the status reflects which one (budget_tokens, budget_calls, max_steps). Set these deliberately — they're your ceiling on cost and your protection against loops.

Control spend

  • Read result.usage after every run — { inputTokens, outputTokens, cachedInputTokens, calls } — and track it. cachedInputTokens are cheaper, so prompt caching is already working for you when it's non-zero.
  • Pick the right tier. poolot-mini and poolot-standard are inexpensive; reserve poolot-pro for work that needs it. See model tiers.
  • Check the balance before a big batch with await agent.getMe() — it returns your tier and remaining balance.

Gate risky actions with approvals

By default a custom tool runs as soon as the model calls it. For anything that changes the world — places an order, sends a message, spends money — mark it approval: "ask" and install a handler that decides:

import { setApprovalHandler, approvalRules } from '@poolot/lily-web'

setApprovalHandler(
  approvalRules({
    allow: [{ tool: 'run_command', argMatches: /^git (status|diff|log)\b/ }],
    deny: [{ tool: 'run_command', argMatches: /\brm\s+-rf\b/ }],
    otherwise: (tool) => false, // default-deny anything not explicitly allowed
  }),
)

approvalRules reads the tool's arguments, not just its name — so git status and rm -rf get different answers. deny always wins over allow, regardless of order. Return a promise from otherwise if you need to ask a human. Default-deny (returning false) is the safe posture for unattended runs.

Make custom tools robust

  • inputSchema — pass a Zod schema; arguments are validated before your handler runs, and the model is told what was wrong so it can fix the next call.
  • timeoutMs — bound each call (default 30s). A timeout is never retried.
  • retries — extra attempts after a throw (default 0), for flaky I/O.
  • Throwing is a normal signal, not a crash — it marks the tool result failed and lets the model recover.

Answer questions without a human

If the agent asks a question mid-run (the ask_user tool) and nobody's there, the run stalls. For headless jobs, install an input handler so questions are answered programmatically:

import { setInputHandler } from '@poolot/lily-web'

setInputHandler((question) => {
  // Return a default, or look the answer up. Return null to decline.
  return 'Proceed with the safe default.'
})

See what's happening

  • Hooks — pass hooks to createLilyAgent to observe the run at each stage (run_start, tool_call_before, tool_call_after, error, run_end). Use them for logging, metrics, and tracing.
  • Error reporting — configure a Sentry dsn in the load options to capture the SDK's own errors, and call captureException(err, { component, op }) to report your own.
  • Cancellation — call agent.abort() to stop a run in flight; long-running tool handlers should check context.signal and bail when it aborts.

Keep your key server-side

The biggest production mistake is shipping a provider key to the browser.

  • Never embed a long-lived programmatic key in client code — it's visible to anyone who opens dev tools.
  • In the browser, pass a short-lived signed-in user token as apiKey, or route requests through your own same-origin proxy with baseUrl (default /api/agent-server) that holds the key server-side.
  • Bring-your-own-key (provider) is a Node/harness capability for exactly this reason — see Browser vs Node.

Production checklist

  • Branch on every result.status, not just completed.
  • Set maxSteps / maxTokens / maxCalls budgets.
  • Track result.usage; pick the cheapest model tier that works.
  • approval: "ask" on every world-changing tool, with a default-deny handler.
  • Validate tool inputs (inputSchema) and bound them (timeoutMs).
  • Install setInputHandler for unattended runs.
  • Wire hooks and a Sentry dsn for observability.
  • Key stays server-side; browser uses a user token or a proxy.

See also