Click or drag & drop an audio file hereWAV, MP3, FLAC, OGG — any format
Frame-level detection & tagging — every sound with its onset → offset. This is the model's primary, most reliable paralinguistic output.