Skip to contents

Pitch & F0 Tracking

Fundamental frequency estimation — 18 algorithms spanning C++ (fastest), classical signal processing, and deep learning.

trk_pitch_rapt()
Track fundamental frequency using RAPT (Robust Algorithm for Pitch Tracking)
trk_pitch_swipe()
Track fundamental frequency using SWIPE (Sawtooth Waveform Inspired Pitch Estimator)
trk_pitch_dio()
DIO Pitch Tracking (C++ implementation)
trk_pitch_harvest()
Harvest Pitch Tracking (C++ implementation)
trk_pitch_reaper()
Track fundamental frequency using REAPER (Robust Epoch And Pitch EstimatoR)
trk_pitch_yin()
Track fundamental frequency using the YIN algorithm
trk_pitch_pyin()
Track fundamental frequency using probabilistic YIN (pYIN)
trk_pitch_pda()
Track fundamental frequency using the ESTk PDA algorithm
trk_tandem()
Track pitch and voiced speech using the TANDEM-STRAIGHT algorithm
trk_pitch_swiftf0()
Track fundamental frequency using SwiftF0 (ONNX)
trk_pitch_crepe()
Track fundamental frequency and periodicity using CREPE (ONNX)
trk_pitch_cc()
Pitch tracking via Praat's cross-correlation method
trk_pitch_ac()
Pitch tracking via Praat's autocorrelation method
trk_pitch_shs()
Pitch tracking via Praat's subharmonic summation (SHS) method
trk_pitch_spinet()
Pitch tracking via Praat's SPINET method
trk_pitch_srh()
Track fundamental frequency using the Summation of Residual Harmonics (SRH)
trk_pitch_ksv()
Track fundamental frequency using the KSV periodicity detector
trk_pitch_mhs()
Track pitch using the Modified Harmonic Sieve algorithm
trk_pitch_snack()
Track fundamental frequency using the Snack/ESPS dp_f0 algorithm

Formant Analysis

Formant frequency and bandwidth tracking.

trk_formant_forest()
Track formant frequencies and bandwidths (FOREST)
trk_formant_deepformants()
Track formant frequencies using DeepFormants (ONNX)
trk_formant_tvwlp()
Track formants using Time-Varying Weighted Linear Prediction (TVWLP)
trk_formant_burg()
Formant frequencies and bandwidths via Praat's Burg method
trk_formant_cgdzp()
Track formants using Chirp Group Delay Zero-Phase (CGDZP) analysis
trk_formant_snack()
Track formants and bandwidths using the Snack/ESPS LPC tracker
trk_formant_formantnet()
Track formant frequencies and bandwidths using FormantNet (ONNX)

Spectral Analysis

Spectrum estimation, cepstral analysis, and spectral moments.

trk_dft_spectrum()
Track short-term DFT power spectrum
trk_css_spectrum()
Track cepstrally-smoothed spectrum
trk_lps_spectrum()
Track LP-smoothed spectrum
trk_cepstrum()
Track short-term cepstral coefficients
trk_spectral_moments()
Spectral moments (CoG, SD, skewness, kurtosis)
trk_mfcc()
Extract Mel-Frequency Cepstral Coefficients (MFCCs) via SPTK
trk_cheap_trick()
CheapTrick Spectral Envelope Estimation (WORLD vocoder, C++ implementation)

Energy & Amplitude

Signal energy, zero-crossings, autocorrelation, and intensity.

trk_rms()
Track short-term RMS amplitude
trk_zcr()
Track short-term zero-crossing rate
trk_acf()
Track short-term autocorrelation function
trk_intensity()
Sound intensity contour

Voice Quality & Aperiodicity

Comprehensive voice quality assessment — from single-track measures to 132-parameter dysphonia toolboxes.

trk_d4c()
Estimate band aperiodicity using the D4C algorithm (WORLD vocoder)
trk_cpps()
Cepstral Peak Prominence Smoothed (CPPS)
trk_vuv()
Voiced/unvoiced segmentation via two-pass adaptive pitch detection
trk_praatsauce()
Comprehensive voice quality feature set via PraatSauce
lst_covarep_vq()
Extract voice quality parameters (NAQ, QOQ, H1-H2, HRF, PSP) per utterance
lst_vq()
Voice Quality Measurements using pladdrr
lst_pharyngeal()
Pharyngeal Voice Quality Analysis
lst_voice_report()
Voice Report Analysis (pladdrr)
lst_voice_tremor()
Vocal Tremor Analysis Using pladdrr
lst_dsi()
Dysphonia Severity Index (DSI) Analysis (pladdrr)
lst_avqi()
Acoustic Voice Quality Index (AVQI) using pladdrr
trk_covarep_creak()
Detect creaky voice (vocal fry) per frame
trk_covarep_env_te()
Estimate spectral envelope using the True Envelope (Teager energy) method
trk_covarep_vad_drugman()
Detect voiced frames using Drugman's multi-branch VAD
trk_covarep_vq_gci()
Track GCI-anchored voice quality measures as a time series
trk_creak_vat()
Detect creaky voice using the Kane-Drugman VAT creak detector
trk_gci_vat()
Detect glottal closure instants (GCIs) using SE-VQ via voiceanalysis
trk_iaif_vat()
Estimate glottal flow via IAIF using voiceanalysis
trk_mdq_vat()
Track Maxima Dispersion Quotient (MDQ) for breathy/tense voice discrimination
trk_peakslope_vat()
Peak slope via voiceanalysis Daless wavelet bank
trk_peakslope()
Track spectral tilt using D'Alessandro PeakSlope (Morlet wavelet)
trk_pitch_vat()
Track fundamental frequency using SRH via the voiceanalysis package
trk_hmpd()
Extract harmonic model phase distortion features (HMPD): AE, PDM, PDD
lst_lf_vat_synthesis()
Synthesise an LF model glottal pulse via voiceanalysis
lst_vq_vat()
Per-GCI voice-quality summary via voiceanalysis
lst_polarity()
Signal polarity detection (RESKEW algorithm)

Prosody & Intonation

Prosodic features, rhythm, articulation complexity, and pitch modelling.

lst_dysprosody()
Extract Dysprosody Prosodic Features
lst_voxit()
Extract Voxit prosodic complexity features from audio files
lst_vowel_space()
Vowel space analysis (F1×F2 area ratio)
momel()
Run MOMEL algorithm on F0 values
intsint()
Run INTSINT algorithm on MOMEL targets
prosody_measures()
Compute prosodic measures from audio file or Sound object
voxit-analysis-stats
Voxit analysis statistics
voxit-dsp-utils
DSP utility functions from Voxit

Source-Filter Decomposition

Glottal source and vocal tract separation.

trk_gfmiaif()
Decompose speech into vocal tract, glottis, and lip radiation LP filters (GFM-IAIF)
trk_covarep_iaif()
Extract glottal flow waveform using Iterative Adaptive Inverse Filtering (IAIF)

OpenSMILE Feature Sets

Standardized acoustic feature extraction via OpenSMILE C++.

lst_GeMAPS()
Compute the GeMAPS openSMILE feature set (C++ Implementation)
lst_eGeMAPS()
Compute the eGeMAPS openSMILE feature set
lst_emobase()
Compute the emobase openSMILE feature set
lst_ComParE_2016()
Compute the ComParE 2016 openSMILE feature set

Epoch Detection

Glottal closure instants and pitch marks.

trk_pitchmark_estk()
Detect glottal closure instants in laryngograph signals using ESTk pitchmark
trk_pitchmark_reaper()
Detect glottal closure instants using REAPER (pitch marks)
lst_covarep_gci_sedreams()
SEDREAMS Glottal Closure Instant Detection

Psychoacoustics

Equal-loudness contours and loudness unit conversions (ISO 226, ISO 532).

iso226_phon
ISO 226:2023 Phon (Loudness Level) Conversions
iso532-sone
ISO 532 Sone (Loudness) Conversions

Unit Conversion

Psychoacoustic scale conversions (Hz, Bark, ERB, Mel, semitone, phon, sone).

ucnv_bark_to_hz()
Convert Bark Scale to Frequency
ucnv_db_and_hz_to_phon()
Convert Sound Pressure Level and Frequency to Loudness Level (Phon)
ucnv_db_and_hz_to_sone()
Convert dB and Hz Directly to Sone
ucnv_erb_to_hz()
Convert ERB-rate Scale to Frequency
ucnv_hz_to_bark()
Convert Frequency to Bark Scale
ucnv_hz_to_erb()
Convert Frequency to ERB-rate Scale
ucnv_hz_to_mel()
Convert Frequency to Mel Scale
ucnv_hz_to_semitone()
Convert Frequency to Semitones
ucnv_mel_to_hz()
Convert Mel Scale to Frequency
ucnv_phon_and_hz_to_db()
Convert Loudness Level (Phon) and Frequency to Sound Pressure Level
ucnv_phon_to_sone()
Convert Phon to Sone
ucnv_semitone_to_hz()
Convert Semitones to Frequency
ucnv_sone_and_hz_to_db()
Convert Sone and Hz to dB
ucnv_sone_to_phon()
Convert Sone to Phon

I/O — Audio & SSFF

Load audio, read/write SSFF signal files.

read_audio()
Read an audio file into an AsspDataObj
av_to_asspDataObj()
Convert audio file to AsspDataObj
avaudio_to_av()
Convert AVAudio to av::read_audio_bin Format
avaudio_to_tempfile()
Convert AVAudio to Temporary WAV File
read_ssff()
Read an SSFF or audio file into an AsspDataObj
write_ssff()
Write an AsspDataObj to an SSFF file
prep_recode()
Re-encode Media File with Custom Parameters

I/O — JSON Track Format (JSTF)

Create, read, write, and manipulate JSON Track Format files.

read_track()
Unified Track Reading Interface
write_track()
Write Track to File
create_json_track_obj()
Create a JsonTrackObj
append_json_track_slice()
Append a slice to JsonTrackObj
merge_json_tracks()
Merge multiple JsonTrackObj files
subset_json_track()
Subset JsonTrackObj
validate_json_track()
Validate JsonTrackObj
store_slice()
Provides the ability to store a multidimensional feature set related to a part of a signal.
json_track_core
JSON Track Object Core Functions
json_track_methods
JSON Track Conversion Methods
read_jstf()
Read JSTF File
write_jstf()
Write JSTF Object to File
jstf_io
JSTF (JSON Sparse Track Format) I/O Functions

Classes & Data Structures

Core data classes, S7 generics, and AsspDataObj accessors.

is_avaudio()
Check if Object is AVAudio
as_avaudio()
Convert to AVAudio Object
as.data.frame(<AsspDataObj>) print(<AsspDataObj>) as_tibble(<AsspDataObj>) cut(<AsspDataObj>)
AsspDataObj — ASSP Data Object
print(<JsonTrackObj>) as.data.frame(<JsonTrackObj>) as_tibble(<JsonTrackObj>) summary(<JsonTrackObj>)
JsonTrackObj — JSON Track Format Object
s7-methods
S7 Method System for DSP Functions
sample_rate() n_records() signal_duration() start_time() track_names() file_path() track_formats() dur() numRecs() rate() startTime() tracks()
Accessor methods for AsspDataObj and JsonTrackObj

ASSP Constants & Types

Enumerated types and constants from the ASSP signal processing library.

AsspFileFormats
list of possibly useful file formats for AsspDataObj corresponding to the first element of the fileInfo attribute
AsspLpTypes()
AsspLpTypes
AsspSpectTypes()
AsspSpectTypes
AsspWindowTypes()
AsspWindowTypes
isAsspLpType()
isAsspLpType
isAsspSpectType()
isAsspSpectType
isAsspWindowType()
isAsspWindowType
wrasspOutputInfos
list of default output extensions, track names and output type for each signal processing function in wrassp

Package Utilities

Introspection, visualisation, and media conversion helpers.

get_track_label()
Get track label for plotting
get_track_label_expr()
Get track label as expression for plotting
ggtrack()
Create ggplot with automatic track labels
differentiate()
Derivation of SSFF track objects

Python / pladdrr Integration Helpers

Audio loading and format conversion helpers for pladdrr backends.

av_load_for_pladdrr()
Load audio file as pladdrr Sound object
pladdrr_df_to_superassp()
Convert pladdrr data frame to superassp format
rmsana_memory()
Perform RMS analysis on AsspDataObj in memory
process_media_file()
Process audio from any media file format

Legacy Functions

Pre-v1.0 functions retained for backward compatibility. Prefer the trk_* equivalents for new code.

trk_pitch_mhs()
Track pitch using the Modified Harmonic Sieve algorithm
harmonics()
Compute the harmonic frequency structure from f0 measurements
trk_afdiff()
Differentiate an audio waveform
trk_affilter()
Apply a digital filter to audio signals
trk_arf()
Track LP-derived vocal tract area function coefficients
trk_lar()
Track LP-derived log area ratios
trk_lpc()
Track LP filter coefficients
trk_rfc()
Track LP reflection coefficients
useWrasspLogger
package variable to force the usage of the logger set to FALSE by default
read.AsspDataObj()
read.AsspDataObj from a signal/parameter file
write.AsspDataObj()
write.AsspDataObj to file