Skip to content

Profiles & voices

An agent profile carries two settings that answer two different questions, and they don’t move together:

  • Response tier — how smart (and how fast) the agent’s reasoning is.
  • Voice — what the agent sounds like.

Picking a fast response tier doesn’t limit which voice you can use, and vice versa — see the practical guide for the tier ladder and the voice list.

Four tiers, from fastest to most deliberate: Fast, Standard (the default), Smart, and Realtime (preview — bills like a Smart-tier call; until it is enabled for your environment it answers with Smart’s behavior, and nothing breaks either way).

  • Fast — lowest latency; the platform picks the fastest available model.
  • Standard — the balanced default between speed and reasoning depth.
  • Smart — trades latency for a single, more deliberate answer.

Every tier uses speech recognition with automatic fallback. A profile saved before the tier ladder existed keeps running exactly as it always did, with no forced migration.

Both the response tier and the voice you pick carry their own price multiplier, and the two multiply together rather than one overriding the other: a minute costs duration × response-tier factor × voice-tier factor. Picking a faster/cheaper response tier doesn’t force a cheaper voice tier, and vice versa — the “two independent knobs” model holds all the way through to billing. See Voice tiers below for which voices sit in which tier, and Web widget & text mode for how a WEB_TEXT session (no audio, so no voice cost) fits in.

The picker lists every voice available for your environment, grouped into three price tiers:

  • Standard (no extra cost) — Nairon voices.
  • Premium (available from the Growth plan) — premium voices.
  • HighRes (available from the Professional plan) — high-fidelity voices.

You pick a voice by name from the picker and its tier comes with it; a voice above your plan’s tier is shown locked rather than silently substituted. Custom voice cloning (using your own reference audio to create a new voice identity) is a separate feature and isn’t the subject of this page.

These rules apply to every call — phone, web, and test — in both Simple and Advanced projects.

  • Main language and timezone. Each profile has a main language (one of its languages) and a timezone (default: Central European Time, Zurich). Calls open in the main language, the greeting is written in it, and times the agent speaks (for example callback slots) are in the profile’s timezone. When you change the main language and your voice does not speak it, the picker moves to a voice that does, unless you picked the voice yourself.
  • Following the caller’s language. If the caller switches to another of the profile’s languages with a real sentence, the agent follows and stays in that language. Numbers, names, a single foreign word, greetings, or “yes”/“no” do not switch the call. Fixed lines such as the inactivity warning are spoken in the conversation’s current language. Messages you write in a say node are spoken as written; writing them per language is coming soon.
  • Interrupting the agent (barge-in). The caller can interrupt the agent with a single spoken word. Background noise that is not recognized as a word no longer pauses the agent. In browser sessions, what a caller says over the greeting is still heard.
  • Inactivity. After 120 seconds without the caller speaking, the agent says a short warning. If there is still no answer, the call ends 15 seconds after the warning and is recorded with an inactivity end reason.