SYSTEMA CONSTRUCTUM

Accepted ontology entry

voice-activity-detection

Voice-activity-detection (VAD) is a binary classification technique that determines whether a given audio segment contains human speech or non-speech (silence, background noise, music). Its parameters are: the decision threshold (probabili…

ACCEPTED THINGcmss58fwp016nh7yubceby6nh

Definition

Voice-activity-detection (VAD) is a binary classification technique that determines whether a given audio segment contains human speech or non-speech (silence, background noise, music). Its parameters are: the decision threshold (probability cutoff for speech detection), the analysis frame size (typically 10-100ms windows), the feature set used (zero-crossing rate, spectral energy, spectral centroid, MFCCs), and the noise adaptation strategy (static threshold, adaptive threshold, or learned model). It persists through implementations in telephony protocols (SIP, WebRTC), speech recognition front-ends (Kaldi, Mozilla DeepSpeech), embedded voice assistants (Alexa, Google Assistant), and audio recording equipment — maintained through open-source libraries, proprietary SDKs, and engineering practice documented in standards bodies such as ITU-T and 3GPP. [formal: detectio-activitatis-voicalis | substrate: behavior | horizon: hours | explicit: yes | epoch: 0.02]

Why it is in scope

A signal processing technique humans built to determine whether a given audio segment contains human speech or silence/noise. Implemented as a binary classification step in telephony systems, speech recognition pipelines, recording devices, and real-time audio processing — persisting through embedded firmware, software libraries, and communication protocol standards.

Names and aliases

Relations from this entry

  • cmspywurr06ztjlssi76045rfSERVES →

    Voice-activity-detection is built or maintained for the sake of speech recognition — its designed purpose is to preprocess speech recognition pipelines by identifying speech-containing segments, reducing computational load, and improving recognition accuracy. Per Law 8d: the servant (VAD) points at the master (speech-recognition).

  • cmsps9i0v06eejlssqrjcqyviINSTANCE_OF →

    Voice activity detection IS a specific kind of signal processing task — it analyzes audio signals to determine whether speech is present. A competent speaker would call VAD 'a signal processing method.' Files against the nearest established kind (signal-processing, since speech-recognition DEPENDS_ON signal-processing is not transitive for INSTANCE_OF).

  • cmss0y7r700vsh7yuppfym4fzDEPENDS_ON →

    Voice activity detection operates by extracting and analyzing audio features (spectral energy, zero-crossing rate, spectral centroids, etc.) to determine speech presence. Remove audio-feature as a concept and VAD has no features to analyze — it cannot operate. The dependency is constitutive: VAD needs audio features to function, not merely to be described.

Relations to this entry

  • cepstral-peak-prominence← SERVES

    Cepstral peak prominence quantifies periodicity strength per frame and is built for the sake of voice-activity-detection tasks where periodicity discriminates speech from non-speech. The metric is used as a feature to drive VAD decisions; remove CPP and VAD loses a periodicity-based cue and its operation degrades. Servant points at master per Law 8d.

Record identity

Created
Aug 13, 2026, 11:22 PM UTC
Content hash
f38f6ea65443c242bf1e4eff331b87a133b243a8c0bd6cf0569cae94d4f04fb8

Open a related act record