An RNN-based Music Language Model for Improving Automatic Music Transcription


In this paper, we investigate the use of Music Language Models (MLMs) for improving AutomaticMusic Transcription performance. The MLMs are trained on sequences of symbolic polyphonic music from the Nottingham dataset. We train Recurrent Neural Network (RNN)-based models, as they are capable of capturing complex temporal structure present in symbolic music data. Similar to the function of language models in automatic speech recognition, we use the MLMs to generate a prior probability for the occurrence of a sequence. The acoustic AMT model is based on probabilistic latent component analysis, and prior information from the MLM is incorporated into the transcription framework using Dirichlet priors. We test our hybrid models on a dataset of multiple-instrument polyphonic music and report a significant 3% improvement in terms of Fmeasure, when compared to using an acoustic-only model.

Extracted Key Phrases

5 Figures and Tables

Cite this paper

@inproceedings{Sigtia2014AnRM, title={An RNN-based Music Language Model for Improving Automatic Music Transcription}, author={Siddharth Sigtia and Emmanouil Benetos and Srikanth Cherla and Tillman Weyde and Artur S. d'Avila Garcez and Simon Dixon}, booktitle={ISMIR}, year={2014} }