Reading the mood of a song from its score, its lyrics, and its audio — then measuring which of the three actually knows.
The hard partAudio hears energy but not mood, so the two quadrants where valence and arousal disagree collapsed to near-zero F1. Predicting valence and arousal as separate regressions and reading the quadrant off them — plus SMOTE on the minority classes — took macro-F1 from 0.40 to 0.49 and stopped the model ignoring the hard classes entirely.
- 0.764 macro-F1 · GPT-5.5 zero-shot
- 0.743 · audio + lyrics fusion
- 5 approaches compared