Prosody-enhanced laughter detection for stand-up comedy audio.
audio-classificationWavLM + ProsodyMIT
| Variant | F1 |
|---|---|
| Prosody-only | 0.960 |
| Fusion (this model) | 0.879 |
| WavLM-only | 0.220 |
Key insight: prosody alone beats deep speech embeddings for laughter detection.
791-dim (768 WavLM + 23 prosody) โ 512 โ 256 โ 64 โ 1 sigmoid
Binary laughter probability per 5-second audio chunk.
Model: Hayasuki/chuckleNet-v2
Code: github.com/Das-rebel/ChuckleNet
from transformers import AutoModel
model = AutoModel.from_pretrained("Hayasuki/chuckleNet-v2", trust_remote_code=True)Built by Das-rebel ยท part of the Standup4AI project