Real-time translation latency benchmark
Aggregated, anonymized production and desktop-client timings. This report separates the delay people experience from internal server processing so the numbers remain clear and reproducible.
- Completed client timings
- 42
- Completed outgoing timings
- 277
- Incoming text timings
- 568
How fast is Loquora in actual use?
In recent desktop-client logs, the typical delay from the end of a phrase to the first translated audio was 920 milliseconds. In 95 out of 100 measurements, translation started within 1.33 seconds, and almost all completed measurements stayed within 1.5 seconds.
Server processing was faster: the typical time was 582 milliseconds across 277 completed spoken phrases. That measurement ends when the first translated audio is ready on the server, before delivery to the app.
The practical result: most completed translations in this sample began in roughly one second, while the exact delay varied with the network, phrase, language pair and device.
Delay measured at the desktop app
This is the most user-centered measurement in the report. It starts when Loquora marks the end of the spoken phrase and stops when the first translated audio chunk arrives back in the desktop application.
Based on 42 completed client measurements captured on July 16–17, 2026. Exact duplicate records were removed, while completed slow results were retained.
From speech end to synthesized translation
For outgoing speech, we count all time after the phrase ends: recognition, translation and creation of the first audio. The sample contains 277 completed voiced translations recorded on July 22, 2026.
Nearly all completed samples stayed below 1.5 seconds.
Time until the first translated words appear
Incoming speech is processed separately. Across 568 valid server measurements, the typical time from receiving a finished piece of speech to the first translated words was 784 milliseconds.
This is a server-processing metric. It does not claim to include client capture, network delivery back to the device or the start of optional subtitle voicing.
How the latency figures were calculated
The report uses technical timing records, not the text or audio content of conversations. App and server measurements are shown separately because they answer different questions.
- 01
Remove personal data and duplicates
User, call and session identifiers, IP addresses and conversation content are not published. Exact duplicate records are counted once.
- 02
Use fully measured phrases
A measurement is included only when the end of speech can be connected to the first translated audio or words. Without an end marker, the time cannot be calculated.
- 03
Measure client latency
For desktop logs, phrase duration is subtracted from the recorded time-to-first-audio, and the result is checked against the event timestamps.
- 04
Measure server latency
For outgoing speech, we count from the end of the phrase until the first audio is ready. For incoming speech, we count recognition and translation until the first words appear.
What these numbers do and do not prove
This is a transparent product measurement, not an independent laboratory comparison against another service.
- The results describe anonymized samples available through July 22, 2026; future infrastructure and model changes can alter the distribution.
- Internet route, computer load, microphone quality and calling application can add delay outside the measured server pipeline.
- Language pair, accent, speaking style and phrase complexity may affect recognition and translation time.
- The figures cover events with both a measurable start and finish. They are not a translation-accuracy score or a success-rate measurement.
Questions about Loquora latency
What does 920 ms measure?
It is the median desktop-client delay from the detected end of a spoken phrase until the first translated audio chunk reaches the Loquora app.
Why is the server time lower?
The 582 ms server figure stops when translated audio is ready on the server. The client figure also includes delivery of that first audio back to the desktop application.
Why did some translations take longer than 1.5 seconds?
On July 22 this happened in 5 of 277 server measurements. These were rare short processing spikes during heavier infrastructure load. We kept them in the results because users experience the delay regardless of its cause.
Does sentence length determine the delay?
Not by itself. Recognition completion, network conditions, language pair and audio creation all contribute. The current sample is not large enough to publish reliable separate figures for every duration and language pair.
Were slow results removed?
No. Completed slow results were retained. Records without a measurable finish were excluded because their delay cannot be calculated.
Is this an independent test?
No. An independent test requires a party with no connection to Loquora to collect the data under pre-published rules, using the same devices, networks, phrases and language pairs. It should also count failed attempts and publish anonymized results and the calculation method.
The most useful benchmark is your own call.
Try Loquora with free minutes in the calling app you already use and judge the delay, translation and voice together.
Download Loquora