A real-time speech model on the line and your production AI agent behind it — one brain answering support questions, analysing sales, troubleshooting devices and taking real, verified actions. In English and Arabic, in the accent you choose.
A live duplex call: barge-in, tool calls and spoken-number safety in action.
This is a real call: your browser connects to the engine's duplex voice WebSocket and the agent hears you through your microphone. Pick an accent, start the call, and ask anything — the topics on the side are ideas, not buttons.
Press start — the agent picks up in a couple of seconds and hears you through your mic.
Ideas, not buttons — say any of these in your own words on the call.
The same production agent behind your chat picks up the phone — with everything your business already gave it: tools, data, knowledge and policy.
Product and process questions are answered from an embedded, vector-indexed knowledge base — not from the model's imagination. Every answer carries the source it was fetched from.
Sales summaries, branch and terminal performance, payment-method breakdowns — pulled through MCP tools against live data while the caller waits, with figures pushed to the screen.
Step-by-step device and process troubleshooting from your own guides, escalating to a ticket with a technician visit when the steps don't resolve it.
Opens maintenance tickets, updates settlement details and more — with OTP verification held server-side, so a sensitive flow can never be talked past the agent.
Callers interrupt mid-sentence and are understood. Real speech is distinguished from background noise and "mm-hm" backchannels before the agent yields the floor.
Spoken output is rendered deterministically: IBANs are masked to their last four digits, long decimals are rounded, tables go to the screen — and the voice reads results verbatim, never paraphrasing.
A realtime speech model owns the ears and the mouth over one WebSocket. Your existing agent stays the brain — same prompt, same tools, same telemetry as chat.
The browser or phone streams raw audio over a single WebSocket. Semantic voice-activity detection ends turns in under a second; deep noise suppression and echo cancellation are built in.
A realtime speech model converses with exactly one tool: ask_agent. Watchdogs recover stalled turns, cancel obsolete answers by name, and re-deliver answers that never reached the caller.
ask_agent runs the identical turn your text chat runs — same profile, prompt, MCP tools, knowledge base, persistence and tracing. Voice and chat share one session history.
The model hears and generates speech itself — natural, dialect-aware Arabic and English with no text-to-speech ceiling. This powers our flagship merchant demo.
Azure speech recognition and neural voices around a fast LLM — in-region, roughly 4× cheaper per minute, with named voices configurable per language.
Google's realtime stack — it scored 100% exact transcription in our three-way provider bake-off. Provider choice is a config row, not a rebuild.
Switching provider, model, region, voice or accent is a database row edit in the admin portal — never a redeploy.
Every default in the voice stack is a measurement with a date on it. These are the ones you feel on a call.
| What we measure | Result | Why it matters |
|---|---|---|
| Interrupt cancellation | 33 ms | A named cancel stops obsolete audio almost instantly — versus 2.34 s of overhang when unnamed. |
| Stall recovery | ~0.4 s | A turn that ended on "one moment" is finished, not late — a reprompt lands the owed tool call in 0.37 s. |
| Answer delivery → first audio | 0.21–0.33 s | The gap the answer watchdog polices before re-delivering an answer that never made a sound. |
| In-region embeddings | ~100 ms | In-region versus ~600–800 ms cross-cloud — knowledge lookups stay off the caller's clock. |
| Streaming headroom | ~4× | 28 s of speech forwarded in 6 s — playback pacing, not generation, sets the tempo of the call. |
| Voice prompt budget | −51% | The voice persona was compressed from 38.7k to 19k characters with every bilingual rule preserved. |
| Daily voice budget | per tenant | Audio seconds are metered and capped per tenant per day, with concurrency limits per identity and per IP. |
| Latency telemetry | 5 stages | Every turn emits end-of-turn → first-audio, delegate time and more into the observability dashboard. |
Figures are live measurements from our duplex voice edge and demo deployments; your numbers depend on region, provider and tools.
The voice agent isn't a black box — it's an agent profile in the WAJ AI Engine, with the same lifecycle discipline as every other agent.
Provider, model, voices and accents per language, VAD timing, barge-in, watchdogs, budgets — around 25 live-voice settings on the agent profile, all editable in the admin portal without a deploy.
Deterministic voice scenarios place a real duplex call on every deploy and assert the transcript. Unit and turn-invariant suites cover every event ordering a call can physically produce.
A goal-driven LLM caller holds whole conversations with the voice agent — in simulation or over real audio — and a judge grades each rubric, with hard latency gates that clock the answer, not the holding line.
Promote an immutable profile snapshot, compare eval runs across variants, and let the drift check prove the live stack serves exactly the prompt and levers you think it does.
The engine's admin surface is itself an MCP server: ~48 tools to configure profiles, publish prompts, attach MCP servers and knowledge bases, trigger eval runs, check drift, promote and deploy. Your own agents — or Claude — can build, test and tune voice agents end to end.
Silent levers are the failure mode of voice operations — a disabled barge-in or a stale prompt looks fine until a live call. The drift check compares what every profile actually serves against source of truth, and flags incident levers left engaged.
The voice agent binds the same integration surface as every WAJ agent — attach a server in the portal and the tools are live on the next call.
Attach any MCP server and pick a tool subset per agent. Tenant credentials are injected server-side and stripped from the tool schemas — the model never sees a secret.
Point-and-click HTTP tools configured in the portal — the ticket and IBAN actions in our demo are exactly this, with OTP step-up held by the server, not the model.
Documents are ingested, chunked and embedded into a vector index; the voice agent searches them mid-call and cites what it fetched.
Long-term user memory, chat history shared between voice and text, and live business data through your systems — each togglable per profile and tuned for voice latency.
Not a translation layer: the voice agent is engineered per language down to the holding lines — and the Arabic accent is a setting you choose, not a hard-coded default.
With the native realtime model the agent speaks natural, colloquial Arabic itself — and the accent is configurable: Gulf, Egyptian, Levantine and more, from the accents Azure Voice Live supports.
The call language is fixed from your locale — a foreign brand or product name can't drag the agent into another language mid-call.
Holding lines, failure lines, "the figures are on your screen", empathy phrases, even hesitation sounds («ممم»، «طيب») exist per language — no English leaking into an Arabic call.
A generative test requires the English and Arabic rule sets to match rule for rule, so a one-sided prompt edit fails before it ships. Arabic speech-recognition variants are normalised, not enumerated.
This whole page — and the demo call — exists in Arabic with a natural Arabic voice. Switch over and try the same questions.
A 30-minute call: we dial the live demo agent together, walk through the admin portal, and scope what a voice agent over your data and tools would look like.
No commitment — and bring your hardest support questions.