Detect voiced frames using Drugman's multi-branch VAD
trk_covarep_vad_drugman.RdEstimates per-frame speech activity as posterior probabilities using three independent ANN classifiers (MFCC-based, Sadjadi pitch-related, and CPP/SRH features) combined by geometric mean (Drugman et al. 2012) . Suitable for pre-filtering frames before pitch or voice quality analysis.
Usage
trk_covarep_vad_drugman(listOfFiles, beginTime = 0, endTime = 0, toFile = FALSE, explicitExt = "cvd", outputDirectory = NULL, verbose = TRUE)Arguments
- listOfFiles
Character vector of audio file paths. Any format supported by av is accepted; non-native inputs are transcoded automatically.
- toFile
Logical. If
TRUE, write SSFF output files and return the paths written invisibly. IfFALSE, return anAsspDataObj. DefaultFALSE.- explicitExt
Character. Output file extension. Default
"cvd".- beginTime
Start time for the extracted portion in seconds. Default: NULL (beginning of signal). Note: uses
beginTime/endTime(seconds) matching DSP function conventions, unlikeread_audio()which usesbegin/end.- endTime
The end time of the section of the sound files that should be analysed (in seconds). Use 0 for end of file.
- outputDirectory
The directory where the slice file should be stored. If not defiled (NULL), the sparse slice file will placed in the same folder as the media file.
- verbose
Logical. Show a progress bar (sequential path) or a progress-aware parallel apply (
pbapply/pbmcapply, if installed).
Value
If toFile = FALSE: an AsspDataObj with tracks:
vad_finalFLOAT, ensemble voicing posterior (geometric mean of three branches), 0–1, n_frames × 1. Recommended for downstream use.
vad_mfccFLOAT, MFCC-branch posterior, 0–1, n_frames × 1.
vad_sadjadiFLOAT, Sadjadi pitch-feature posterior, 0–1, n_frames × 1.
vad_newFLOAT, CPP/SRH-feature posterior, 0–1, n_frames × 1.
Frame rate: 200 Hz (5 ms hop, interpolated from 10 ms internal hop).
If toFile = TRUE: character vector of output file paths, returned invisibly.
Details
Audio is resampled to 16 kHz internally. A threshold of 0.5 on vad_final
gives a reasonable binary voiced/unvoiced decision; 0.7 is more conservative.
Each branch applies an 11-frame median filter to the posterior before combination.
References
Drugman T, Soria-Olivas E, Perez-Córdoba JL, Alwan A (2012). “A comparative study of different feature sets for acoustic voice quality assessment.” IEEE Transactions on Audio, Speech, and Language Processing, 20(6), 1690–1703. doi:10.1109/TASL.2012.2188377 . Multi-branch voice activity detection using MFCC and Sadjadi features.
Examples
if (FALSE) { # \dontrun{
trk_covarep_vad_drugman(
system.file("samples", "sustained", "a1.wav", package = "superassp"),
toFile = FALSE
)
} # }