Voice-activity-detection (VAD) is a binary classification technique that determines whether a given audio segment contains human speech or non-speech (silence, background noise, music). Its parameters are: the decision threshold (probability cutoff for speech detection), the analysis frame size (typically 10-100ms windows), the feature set used (zero-crossing rate, spectral energy, spectral centroid, MFCCs), and the noise adaptation strategy (static threshold, adaptive threshold, or learned model). It persists through implementations in telephony protocols (SIP, WebRTC), speech recognition front-ends (Kaldi, Mozilla DeepSpeech), embedded voice assistants (Alexa, Google Assistant), and audio recording equipment — maintained through open-source libraries, proprietary SDKs, and engineering practice documented in standards bodies such as ITU-T and 3GPP. [formal: detectio-activitatis-voicalis | substrate: behavior | horizon: hours | explicit: yes | epoch: 0.02]
Accepted ontology entry
voice-activity-detection
Voice-activity-detection (VAD) is a binary classification technique that determines whether a given audio segment contains human speech or non-speech (silence, background noise, music). Its parameters are: the decision threshold (probabili…
Definition
Why it is in scope
A signal processing technique humans built to determine whether a given audio segment contains human speech or silence/noise. Implemented as a binary classification step in telephony systems, speech recognition pipelines, recording devices, and real-time audio processing — persisting through embedded firmware, software libraries, and communication protocol standards.
Names and aliases
- voice-activity-detectionen · CANONICAL
Relations from this entry
- cmspywurr06ztjlssi76045rfSERVES →
Voice-activity-detection is built or maintained for the sake of speech recognition — its designed purpose is to preprocess speech recognition pipelines by identifying speech-containing segments, reducing computational load, and improving recognition accuracy. Per Law 8d: the servant (VAD) points at the master (speech-recognition).
- cmsps9i0v06eejlssqrjcqyviINSTANCE_OF →
Voice activity detection IS a specific kind of signal processing task — it analyzes audio signals to determine whether speech is present. A competent speaker would call VAD 'a signal processing method.' Files against the nearest established kind (signal-processing, since speech-recognition DEPENDS_ON signal-processing is not transitive for INSTANCE_OF).
- cmss0y7r700vsh7yuppfym4fzDEPENDS_ON →
Voice activity detection operates by extracting and analyzing audio features (spectral energy, zero-crossing rate, spectral centroids, etc.) to determine speech presence. Remove audio-feature as a concept and VAD has no features to analyze — it cannot operate. The dependency is constitutive: VAD needs audio features to function, not merely to be described.
Relations to this entry
- cepstral-peak-prominence← SERVES
Cepstral peak prominence quantifies periodicity strength per frame and is built for the sake of voice-activity-detection tasks where periodicity discriminates speech from non-speech. The metric is used as a feature to drive VAD decisions; remove CPP and VAD loses a periodicity-based cue and its operation degrades. Servant points at master per Law 8d.
Record identity
- Created
- Aug 13, 2026, 11:22 PM UTC
- Content hash
- f38f6ea65443c242bf1e4eff331b87a133b243a8c0bd6cf0569cae94d4f04fb8