Benchmarks

The numbers, and
how we got them

Vendor benchmarks are usually measured under conditions that never occur on a real phone call. These are measured on production telephony, noisy, accented, interrupted, over real codecs, because that is the only number that predicts what your callers will experience.

Performance chart with rising bars and a trend line
490msMedian response latency
97%Word accuracy
99.5%Uptime SLA
60 daysMinimum measurement window
What we measure

Five numbers that decide
whether a deployment works

Each is defined precisely below, because the definition matters more than the figure. A vendor quoting latency without stating where they start and stop the clock is not quoting anything.

Response latency, 490ms median

Measured from the caller's final phoneme to the first audible phoneme of the reply, over real telephony. Not component sum, not local loopback.

Network latency

Round-trip from the caller's telephony edge to the inference layer, measured per deployment region. On-premise deployments effectively remove this term entirely.

Word accuracy, 97%

On production telephony audio including background noise, accents and codec compression. Clean-studio benchmarks run higher and are not reported here because they are not predictive.

Uptime, 99.5% SLA

Contractual availability with elastic capacity. Measured on successful call handling, not on infrastructure ping.

Throughput

Sustained concurrent request handling, scaling horizontally with stateless workers.

Time to production, 6 weeks

Median across deployments from scoping call to production traffic. Integration, not model tuning, is the usual critical path.

How it works

Measurement method

Full methodology is available to prospective customers under NDA, including the evaluation sets.

Production audio

Evaluation sets sampled from live telephony across markets, deliberately including poor-quality calls.

Customer baseline

Pre-deployment performance measured on the same call categories before the agent handles anything.

60-day window

Minimum measurement period once live, to capture peak days, edge cases and drift.

Reported in full

Results delivered to the customer including where the agent underperformed the baseline.

  • Latency measured end to end over real telephony
  • Accuracy measured on noisy production audio, not clean sets
  • Uptime measured on successful call handling
  • Every customer-facing metric traced to a customer baseline
  • Minimum 60-day measurement window
  • Methodology and evaluation sets reviewable under NDA
FAQ

Benchmark questions

Straight answers. Anything missing? Write to Sales@or-on.io.

Because of what we measure on. Word accuracy on clean studio audio is comfortably above 99% for most modern systems, including ours. On real telephony with noise, accents and codec loss it is 97%, and that is the number that predicts your callers' experience. We report the useful one.

From the caller's final phoneme to the first audible phoneme of the agent's reply, measured over real telephony transport. Not the sum of component processing times, and not measured on a local loopback, both of which produce flattering, meaningless numbers.

It removes network latency to the inference layer, which usually improves end-to-end response. Absolute performance then depends on the GPU hardware available, and we size that against your concurrency requirement during scoping.

Yes, and we encourage it. Pilots are structured so that you measure on your own traffic against your own baseline. Our numbers should be a hypothesis you test, not a claim you accept.

Successful call handling, a call answered and completed or cleanly escalated. Infrastructure availability figures are higher and less meaningful, because a reachable service that fails to complete calls is still an outage to your customers.

Contact

Test the numbers on your own calls

Tell us the use case and the volume. We come back within one business day with a scoped pilot, a timeline and a number.

Contact Us

Accessibility