Skip to Content.
Sympa Menu

emacspeak - [Emacspeak] Omnivox: speech onset is damaged when playback starts after idle (1.12.0 and 1.13.0, Windows x64)

Subject: Emacspeak discussion list

List archive

[Emacspeak] Omnivox: speech onset is damaged when playback starts after idle (1.12.0 and 1.13.0, Windows x64)


Chronological Thread 
  • From: Ľuboš Pinteš <lubos.pintes AT gmail.com>
  • To: emacspeak AT emacspeak.net
  • Subject: [Emacspeak] Omnivox: speech onset is damaged when playback starts after idle (1.12.0 and 1.13.0, Windows x64)
  • Date: Mon, 28 Sep 2026 10:55:31 +0200

When Omnivox starts playing after a period of silence, the beginning of the utterance is corrupted. On 1.13.0, measured at the 44.1 kHz stereo pipeline rate:

1. About the first 256 frames of the utterance never reach the device.
2. The next 256 frames are played at half speed, an octave lower. They look like interleaved stereo samples being read as mono: 512 samples come out as 512 frames.
3. After that, playback is sample-exact.

This is about 12 ms, but it sits right on the onset. With short utterances such as character echo while typing, it is clearly audible: vowels sound rough and "thick" instead of clean.

The same half-speed stretch also happens with `omnivox --play-wav FILE` on the first sound after idle. That path loses nothing, but its first 256 frames are still played at half speed. So the problem seems to be in the playback stage, not in helper streaming. My unconfirmed guess is a channel-count mismatch in the output queue or mixer when a new source starts on an idle stream.

`--dump-wav` output is correct. The damage only shows in live playback, so I measured it with WASAPI process loopback of the Omnivox process tree and compared it with `--dump-wav` output.

**Reproduction:** Let the server go idle for about a second, send `l {e}` (any short utterance with a sharp onset), and repeat. Every repetition after idle shows the same pattern.

**A second observation from the workaround:** As a workaround, my helper prepends 20 ms of low-level lead-in to each utterance, so the damaged window falls on the lead-in instead of speech. The lead-in only survives if every sample of it is above the 0.01 silence-trim threshold. With a short tick followed by zeros, the live path collapses the lead-in to about 253 frames regardless of its length (20 or 60 ms): the zeros after the first non-zero sample get discarded. `--dump-wav` keeps the full lead-in. The workaround is therefore fragile, and a fix in Omnivox would let me remove it.




Archive powered by MHonArc 2.6.19+.

Top of Page