all
a
/
b
/
c
/
f
/
h
/
j
/
jp
/
l
/
o
/
q
/
s
/
sw
/
lounge
cgi
up
wiki
Heyuri!
Bulletin Boards
2D Cute
2D Ero
2D Lolikon
3D Girls
Anime/Manga
Flash
Girl Talk
日本語/Japan
Lounge
Oekaki
Off-Topic
Site Discussion
Strange World
Overboard
Heyuri★CGI
Heyuri★CGI
@PartyII
Battle Royale R
Chat
Chinsouki★
Dating
DevChat
Drama Club
Hakoniwa Islands PvE
Hakoniwa Islands PvP
Polls
Slime Breeder
Web Banana
Web Shiritori
Yumemiru Gambler
Kakiko Checker
Other
Anime Nominations
Banners
Cytube
Heyuri Calendar
Heyuri Wiki
MAL Club
Museum
Steam Group
Uploader
[
Settings
]
[
Home
] [
Contact
] [
Catalog
] [
Search
] [
Thread list
] [
Stats
] [
Reports
] [
Watcher
] [
PMs
] [
Banners
] [
Admin
]
Submit board banners here.
Off-Topic@Heyuri
it's the place to be!
[
Return
]
Report a post
Preview
File:
Screenshot 2026-10-10 at 11.59.19 PM.png
[Animated PNG]
(528 KB, 1836x1556)
[
ImgOps
]
ImgOps
Hide image
▶
Improvement to algorithm for detecting glottal closure instants from laryngograph signal
QueueSevenM
◆Tnq5UWtkfs
[
PM
]
2026/10/11
(Sun)
04:17:13
No.
196560
[
Edit
]
[
Report
]
+
AS
[
Edit
] [
Report
]
▶
I have implemented an improvement to the laryngograph onset detection procedure.
In typical laryngograph signals, like most signals, there is high-frequency noise. However, in addition to this high-frequency noise, there is also an abundance of low-frequency noise, usually much more than the high-frequency noise. This is due shifting of the neck and laryngograph. The standard DSP solution to this problem would be to use a highpass filter. However, the frequency of this noise, and thus the required transition band, are so low that it would render a standard single-stage FIR highpass filter impractical.
The complete procedure I have implemented is follows:
1. The signal is lowpass-filtered (to filter out the high-frequency noise).
2. Peaks (which correspond to glottal closure instants) are detected and interpolated using quadratic interpolation. The topographical prominence is also calculated for each peak, and peaks below a minimum prominence are discarded.
3. In between each adjacent pair of peaks, the minimum valley is found (also using quadratic interpolation).
4. An envelope is created using the natural cubic spline algorithm from this set of minimum valleys.
5. This envelope is subtracted from the original signal.
6. The signal is filtered again, replacing the old filtered signal.
7. Peaks are detected again, replacing the old set of peaks.
8. These peaks are grouped into regions using two heuristics: distance between subsequent peaks and change in subsequent distances.
9. Regions with only one peak are discarded.
Another approach would be to use two-stage downsampling. Then, it could be further lowpass-filtered to obtain fine-control (the new sample rate is chosen to be efficient, not necessarily precisely what is desired). These values could then be used in the same interpolation procedure; alternatively, they can be upsampled back to the original sampling rate (also using a two-stage process for efficiency reasons) and this signal subtracted from the laryngograph signal.
This approach is more correct in theory. On the other hand, this assumes a perfectly periodic signal. In reality, the laryngograph pulses vary. Furthermore, they are most definitely not period around the start and end of voiced regions. This aperiodicity will distort any filtering operation. On the other hand, the valleys in between pulse generally vary much less than the signal overall. Two stage downsampling is also less efficient than the spline approach.
Post number
No.
196560
Board
Off-Topic@Heyuri
Reason
Optional. Describe what's wrong with it.
Style: