[Return]

Report a post

Preview
540871
I have implemented an improvement to the laryngograph onset detection procedure.

In typical laryngograph signals, like most signals, there is high-frequency noise. However, in addition to this high-frequency noise, there is also an abundance of low-frequency noise, usually much more than the high-frequency noise. This is due shifting of the neck and laryngograph. The standard DSP solution to this problem would be to use a highpass filter. However, the frequency of this noise, and thus the required transition band, are so low that it would render a standard single-stage FIR highpass filter impractical.

The complete procedure I have implemented is follows:
1. The signal is lowpass-filtered (to filter out the high-frequency noise).
2. Peaks (which correspond to glottal closure instants) are detected and interpolated using quadratic interpolation. The topographical prominence is also calculated for each peak, and peaks below a minimum prominence are discarded.
3. In between each adjacent pair of peaks, the minimum valley is found (also using quadratic interpolation).
4. An envelope is created using the natural cubic spline algorithm from this set of minimum valleys.
5. This envelope is subtracted from the original signal.
6. The signal is filtered again, replacing the old filtered signal.
7. Peaks are detected again, replacing the old set of peaks.
8. These peaks are grouped into regions using two heuristics: distance between subsequent peaks and change in subsequent distances.
9. Regions with only one peak are discarded.

Another approach would be to use two-stage downsampling. Then, it could be further lowpass-filtered to obtain fine-control (the new sample rate is chosen to be efficient, not necessarily precisely what is desired). These values could then be used in the same interpolation procedure; alternatively, they can be upsampled back to the original sampling rate (also using a two-stage process for efficiency reasons) and this signal subtracted from the laryngograph signal.

This approach is more correct in theory. On the other hand, this assumes a perfectly periodic signal. In reality, the laryngograph pulses vary. Furthermore, they are most definitely not period around the start and end of voiced regions. This aperiodicity will distort any filtering operation. On the other hand, the valleys in between pulse generally vary much less than the signal overall. Two stage downsampling is also less efficient than the spline approach.
Post number No.196560
Board Off-Topic@Heyuri
Optional. Describe what's wrong with it.