LyricSense: A Multilingual Song Lyrics Corpus

Student Data Science Project

When it came time to select a topic for their Advanced Corpus Linguistics project, a group of UBC Master of Data Science (MDS) students (Kara Leier, Stuti Sheth. Saumi Rahnamay, and Spike Wang) wanted to combine their common love of music and embrace the multilingualism within their group to annotate music metadata by looking at the lyrics.

The students sourced their data through a combination of scraping from Genius, a well-known public site, and existing Kaggle datasets. Their total dataset had 10,000 songs, of which 1,000 were manually annotated. The team also focused on lyrics in these four languages: English, French, Mandarin and Hindi.

The students used eight broad narrative categories. The tags spanned a range of positive, neutral, and negative emotions. Through the lens of written text without auditory context, they analyzed the lyrics with tags such as ‘Nostalgia & Introspection’, ‘Party & Celebration’, ‘Pain & Surrender’, ‘Social Commentary’, and ‘Love & Relationships’.

One interesting finding was of the music analyzed, a lot of the music leaned heavily towards two of the eight tags: ‘Pain & Surrender’, and ‘Love & Relationships’.

One of the most striking findings was the dominance of Love & Relationships across all four languages, reinforcing the idea that romantic themes are central to global music cultures.
However, the students also observed meaningful cross-linguistic differences. For example, while English, Hindi, and Mandarin songs were heavily dominated by Love, French songs showed a more balanced distribution, with Social Commentary appearing nearly as frequently.

Additionally, the most common co-occurring tags were Love and Pain, suggesting that emotional complexity—particularly the overlap between affection and suffering—is a defining feature of lyrical storytelling.

These patterns highlight how music reflects both universal emotional themes and culturally specific priorities.
One key challenge was deciding on an annotation schema early in the process. In hindsight, conducting more extensive exploratory data analysis (EDA) before finalizing tags may have helped them better align categories with the underlying data distribution.

Another major challenge was meaning loss in translation. For non-English songs, annotators often relied on English translations, which sometimes failed to capture cultural nuance, idiomatic expressions, or emotional subtleties. This was reflected in their inter-annotator agreement (a statistical measure of consistency between different human annotators labelling the same data) results, where native speakers often disagreed with non-native speakers and AI systems.

Master of Data Science Computational Linguistics Lyrics
Screenshot of the LyricSense Dashboard

The final outcome was LyricSense, a multilingual annotated corpus and interactive search system for exploring song lyrics. LyricSense features a custom annotation schema for emotional and narrative tagging that demonstrates how lyrical analysis can enhance music understanding and recommendation systems.

With more time, the students would have liked to have improved the annotation system and schema. This could include refining categories into more granular sub-themes. They would also explore integrating audio features alongside lyrics to build a more comprehensive multimodal model. This would help them understand if lyrics and audio have a joint effect on music trends and analyze whether audio metadata is more consistent in a particular narrative category.

At the end of the day, the students believe that their project can contribute to improved music recommendation systems by incorporating lyrical meaning alongside traditional audio features. This enables more personalized recommendations based on emotional and thematic preferences, rather than just sound.

Second, the corpus can support cross-cultural analysis of music, helping researchers explore how emotions and narratives are expressed differently across languages. It offers a deeper analysis into art, culture and history

Finally, it provides a foundation for future work in multilingual NLP, particularly in areas involving subjective interpretation, such as sentiment and emotion analysis.

Explore Computational Linguistics Explore Other Data in Action Stories