the Verli blog
Hands-on write-ups from building live transcription and translation: benchmarks, tradeoffs, and the data behind the decisions.

Ten speech-to-text engines run through their real-time streaming APIs on identical audio: streaming WER against batch WER on matched clips with paired-bootstrap intervals, plus first word and finalization latency.

Compare 16 speech-to-text engines across 10 languages using the same audio, word error rate, character error rate, and paired-bootstrap tests.

Compare 17 text-to-speech APIs on pricing, streaming, voice cloning, languages, rate limits, training-data policies, latency, and quality.

See major speech-to-text engines tested on identical LibriSpeech and Earnings-22 audio with one scorer and paired-bootstrap confidence intervals.

A data-first look at 17 speech-to-text APIs plus the open-source models you can self-host: per-hour pricing, real-time streaming, built-in translation, languages, diarization, rate limits, deployment, and what open models actually cost to run. No winner named, just the tradeoffs, with sources.

A hands-on comparison of six real-time meeting translation tools (Verli, JotMe, Maestra, Transync, Whisperr, Notta) plus what Zoom, Teams, and Google Meet translate natively. Bot versus no-bot, captions versus spoken, and what each costs.