[Return] [Bottom]

Posting mode: Reply

Emotes
Kaomoji
Emoji
BBCode
(for deletion)
  • Allowed file types are: gif, jpg, jpeg, png, bmp, webp, swf, webm, mp4
  • Maximum file size allowed is 50000 KB.
  • Images greater than 200 * 200 pixels will be thumbnailed.
  • 19 unique users in the last 10 minutes (including lurkers)





Want your banner here? Click here to submit yours!

540871
I have implemented an improvement to the laryngograph onset detection procedure.

In typical laryngograph signals, like most signals, there is high-frequency noise. However, in addition to this high-frequency noise, there is also an abundance of low-frequency noise, usually much more than the high-frequency noise. This is due shifting of the neck and laryngograph. The standard DSP solution to this problem would be to use a highpass filter. However, the frequency of this noise, and thus the required transition band, are so low that it would render a standard single-stage FIR highpass filter impractical.

The complete procedure I have implemented is follows:
1. The signal is lowpass-filtered (to filter out the high-frequency noise).
2. Peaks (which correspond to glottal closure instants) are detected and interpolated using quadratic interpolation. The topographical prominence is also calculated for each peak, and peaks below a minimum prominence are discarded.
3. In between each adjacent pair of peaks, the minimum valley is found (also using quadratic interpolation).
4. An envelope is created using the natural cubic spline algorithm from this set of minimum valleys.
5. This envelope is subtracted from the original signal.
6. The signal is filtered again, replacing the old filtered signal.
7. Peaks are detected again, replacing the old set of peaks.
8. These peaks are grouped into regions using two heuristics: distance between subsequent peaks and change in subsequent distances.
9. Regions with only one peak are discarded.

Another approach would be to use two-stage downsampling. Then, it could be further lowpass-filtered to obtain fine-control (the new sample rate is chosen to be efficient, not necessarily precisely what is desired). These values could then be used in the same interpolation procedure; alternatively, they can be upsampled back to the original sampling rate (also using a two-stage process for efficiency reasons) and this signal subtracted from the laryngograph signal.

This approach is more correct in theory. On the other hand, this assumes a perfectly periodic signal. In reality, the laryngograph pulses vary. Furthermore, they are most definitely not period around the start and end of voiced regions. This aperiodicity will distort any filtering operation. On the other hand, the valleys in between pulse generally vary much less than the signal overall. Two stage downsampling is also less efficient than the spline approach.

Want your banner here? Click here to submit yours!

[Top]

Delete post: []
First
Last