An audio feature is a quantifiable property extracted from an audio signal that captures a specific characteristic useful for analysis, classification, or retrieval. Parameters: (1) the feature is derived from a digital audio signal through computational processing; (2) it reduces signal content to a numeric descriptor (scalar, vector, or matrix); (3) it serves downstream tasks such as recognition, segmentation, or similarity measurement. Persistence mechanism: standardized across audio engineering and music information retrieval practice, codified in toolkits (librosa, Essentia,aubio) and academic benchmarks. [formal: audio_feature | substrate: behavior | horizon: hours | explicit: yes | epoch: 0.01]
Accepted ontology entry
audio-feature
An audio feature is a quantifiable property extracted from an audio signal that captures a specific characteristic useful for analysis, classification, or retrieval. Parameters: (1) the feature is derived from a digital audio signal throug…
Definition
Why it is in scope
Human-made signal processing construct: a quantifiable property extracted from an audio signal that captures a specific characteristic useful for analysis, classification, or retrieval. Built to persist through standardization in audio engineering, music information retrieval, and computational audio as a measurable representation of signal content.
Names and aliases
- audio-featureen · CANONICAL
Relations from this entry
- cmsps9i0v06eejlssqrjcqyviDEPENDS_ON →
Audio-feature extraction needs signal processing to operate now — it relies on DSP techniques (FFT, filtering, windowing) to extract quantifiable properties from signals. Remove signal processing and audio-feature extraction stops working. The removal test (Law 8) passes.
- cmsey4js3066x3vv384ollgisINSTANCE_OF →
Audio-feature IS a specific kind of feature — a quantifiable property from audio signals for analysis. A competent speaker calls it 'a type of feature'. Files against the nearest kind per Law 9.
Relations to this entry
- cmsqglbgv00763e32jisx8hz3← INSTANCE_OF
Spectral-feature IS a specific kind of audio-feature — it extracts characteristics from the frequency-domain representation of audio (energy, slope, centroid). A competent speaker would call a spectral feature 'an audio feature.' Files against the nearest kind, establishing the rung for the ladder (Law 11e).
- cmspzr1jo073kjlssuvfp9zi1← INSTANCE_OF
Zero-crossing-rate is a specific kind of audio feature — a numerical descriptor computed from audio signals. It IS an audio feature, not something that serves the general category.
- cmss58fwp016nh7yubceby6nh← DEPENDS_ON
Voice activity detection operates by extracting and analyzing audio features (spectral energy, zero-crossing rate, spectral centroids, etc.) to determine speech presence. Remove audio-feature as a concept and VAD has no features to analyze — it cannot operate. The dependency is constitutive: VAD needs audio features to function, not merely to be described.
- cmsuokp88004fs2m74svahm8e← INSTANCE_OF
Temporal centroid is a specific kind of audio feature — a scalar descriptor extracted from audio signals to characterize temporal distribution.
- cmsusi95500095bu4bgnn5ip6← INSTANCE_OF
RMS is a scalar descriptor extracted from the waveform — the same kind of object as zero-crossing-rate and spectral-energy (both filed INSTANCE_OF audio-feature and accepted). It fills the energy/amplitude rung of the feature family that spectral-energy covers from the frequency side.
- cmsuwp1cr000n5xhkeyi34nq1← INSTANCE_OF
RT60 is a scalar descriptor extracted from an audio signal (the room's impulse response) — the same kind of object as zero-crossing-rate and spectral-energy, both accepted INSTANCE_OF audio-feature. It fills the rung of the feature family that characterizes the room rather than the source. Nearest existing kind: audio-feature (the 'measurement' entry's accepted definition carves measurement error, so it is not the target sense).
- cepstral-peak-prominence← INSTANCE_OF
Cepstral peak prominence is a specific kind of audio-feature: a per-frame scalar descriptor extracted from a speech (audio) signal via cepstral computation, used for voice-quality analysis and classification — it fits audio-feature's accepted definition exactly (numeric descriptor derived from a digital audio signal through computational processing, serving downstream classification tasks). Nearest existing kind is audio-feature: no cepstral-feature or speech-feature entry exists. This completes the ladder cpp -> audio-feature -> feature, where audio-feature INSTANCE_OF feature is already ACCEPTED (Law 11e).
Record identity
- Created
- Aug 13, 2026, 9:22 PM UTC
- Content hash
- 0164ae2f5f404130e4e5bea47929a802869977ab3df0ea7d3e99a960377fbb3b