Voice & Agent Core

380 milliseconds is the
whole product

Everything people dislike about voice bots comes from latency and rigid turn-taking. OrOn's core streams recognition, reasoning and synthesis continuously rather than in blocks, which is why callers interrupt it naturally instead of waiting for the beep.

Concentric voice core with a live waveform at its centre
490msMedian response
97%Recognition accuracy
32+Languages
Barge-inInterrupt any time
How the core works

Continuous, not turn-based

Older architectures wait for silence, transcribe, think, synthesise, then speak. Each stage adds a delay the caller feels. Ours overlap.

Streaming recognition

Transcription begins while the caller is still speaking, so understanding is nearly complete by the time they stop. No dead air waiting for a silence timer.

Barge-in and turn control

The caller can interrupt mid-sentence and the agent stops immediately, listens, and picks up the new thread, the single behaviour that most makes a voice agent feel human.

Grounded reasoning

Responses are grounded in your documents, your policies and live system data, with retrieval happening inside the same latency budget.

Natural synthesis

Neural voices with correct prosody, emphasis and pacing, including per-speaker cloned voices where the use case calls for it.

Locale-tuned, not translated

Each of the 32+ locales handles its own slang, formality registers and code-switching. Hebrew and Arabic are first-class, not afterthoughts.

Continuous learning

Real conversations feed evaluation sets. The agent improves on your traffic without you re-authoring scripts by hand.

How it works

One turn, end to end

What happens between the caller finishing a word and hearing a reply.

Capture

Audio streams in over SIP or WebRTC; voice activity detection isolates speech from noise in real time.

Understand

Streaming ASR plus intent and entity extraction, running while the caller is still talking.

Decide

The orchestrator applies context, session memory and your business rules, calling external systems as needed.

Speak

Neural synthesis streams the reply, starting before the full sentence has even been generated.

  • 490ms median end-to-end response latency
  • True barge-in, interruption handled at any point
  • Noise, accent and telephony-codec robust
  • Domain lexicons for names, products and jargon
  • Deterministic guardrails on what the agent may say
  • Runs on your GPUs for full on-premise operation
FAQ

Engine questions

Straight answers. Anything missing? Write to Sales@or-on.io.

Human conversation leaves only a short turn-taking gap. Past it, callers perceive the pause as a problem, they repeat themselves, talk over the agent, or assume the line dropped. Every architectural decision in the core exists to stay inside that window.

We are deliberately model-agnostic. The core selects speech recognition, reasoning and synthesis models per deployment based on your latency, cost and compliance requirements, including fully open-weight stacks for on-premise use where no external inference is permitted.

97% word accuracy on typical telephony audio, maintained through codec-aware processing and noise handling. Accented speech and code-switching are handled by the locale tuning rather than a single global model.

Yes, consent-based voice cloning is supported for brand voices and, in broadcast use cases, per-presenter voices. Cloned models are scoped to your deployment and never shared.

Yes, in on-premise configuration. The complete pipeline runs on your GPU hardware with no external calls, which is how government and defence deployments operate.

Contact

Hear the latency for yourself

Tell us the use case and the volume. We come back within one business day with a scoped pilot, a timeline and a number.

Contact Us

Accessibility