Choosing an AI Model

Every reply your agent writes is generated by an AI model, and you choose which one on the agent's Agent screen under AI Model. The model you pick is the biggest single factor in how much each reply costs, and one of the bigger factors in how careful and nuanced those replies are.

There's no single right answer. A busy agent answering routine questions can run happily on a light model, while a sales agent where every reply matters may be worth a more capable one. This article explains the options and how to choose between them with evidence rather than guesswork.

The models and what they cost

What it is

The AI Model picker lists every model you can choose, with the number of credits each response costs shown next to it. OpenAI models are listed first, then Anthropic models, each ordered from most to least expensive.

How it works

  • GPT-5.6 Sol - 10 credits per response (the most expensive)
  • GPT-5.6 Terra - 3 credits per response (the default for new agents)
  • GPT-5.6 Luna - 1 credit per response (the fastest and cheapest)
  • Claude 4.8 Opus - 6 credits per response
  • Claude 4.6 Opus - 5 credits per response
  • Claude 4.6 Sonnet - 3 credits per response

The cost applies to every turn where the model runs - including a notify-moderators interaction card where the agent looks and decides to stay silent. Your monthly allowance is shared across all your agents, so a pricier model on one agent leaves fewer credits for the others. See Plans, Credits & Usage.

Best practices

Start with the default, GPT-5.6 Terra. It's a good balance of quality and cost for most agents, and it gives you a baseline to compare against before you change anything.


Picking a model for the job

What it's best for

  • Lighter models (Luna) - high-volume agents answering routine, well-documented questions, where speed and cost matter more than nuance
  • Mid-range models (Terra, Claude 4.6 Sonnet) - most agents; solid answers at a moderate cost
  • Top-tier models (Claude 4.6 Opus, Claude 4.8 Opus, GPT-5.6 Sol) - lower-volume agents where a sharper, more careful answer is worth several times the credits, such as sales conversations or complex product questions

How it works

Changing the model takes effect for the agent's next replies. It doesn't change what the agent knows - training, Instructions, and Grounding strictness all stay the same - only which model writes the answer. If answers are wrong because training is missing or outdated, a bigger model won't fix that; add or correct Q&A instead.

Best practices

  • Estimate the cost before you switch: multiply your typical monthly reply count by the model's credits per response, and compare it with your plan's allowance.
  • Don't upgrade the model to paper over thin training - fix the training first, then decide whether a better model still helps.
  • Revisit the choice after big changes in volume; what made sense for a quiet agent may be expensive once traffic grows.

Testing a switch before you commit

What it is

Before moving a live agent to a different model, you can measure the difference instead of guessing.

How it works

Try a handful of real questions in Playground first to get a feel for tone. Then run the same Evaluate set on the current model and on the one you're considering, and compare the scores. Evaluation runs use credits too, based on what the run actually cost, so a comparison on a pricier model costs more.

Best practices

Decide ahead of time how much of a score change you'd accept for a given saving. A small dip for a big drop in credits can be a good trade - but make it a deliberate choice, and check Chat Logs in the days after a switch to confirm real conversations still look right.