International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1704431

1704431 Vol 6 · Issue 11 Download Paper

A Survey on Automatic Music Transcription

Pranav Bhagwat Vishwajit Shelke Akhilesh Murugkar Krishiv Dakwala Shweta C. Dharmadhikari

Subject area: Science,Engineering and Technology  ·  Area of research: Automatic Music Transcription, Machine Learning

Abstract

Automatic Music Transcription (AMT) is a critical but less investigated problem in the field of music information retrieval. In this paper, we study different approaches for achieving Automatic Music Transcription using various methods based on pitch, timbre and note detection. Use of Convolutional Neural Network (CNN) and/or Long Short Term Memory Network (LSTM) is made to transcribe notes from the audio input. We also discuss source separation as a precursor to AMT and different approaches for the same.

Keywords

Automatic Music Transcription (AMT), Pitch Detection, Note Detection, Deep Learning, Convolutional Neural Network (CNN), Long Short Term Memory Network (LSTM), Source Separation

References

[1] E. Benetos, S. Dixon, Z. Duan, and S. Ewert, “Automatic music transcription,” IEEE Signal Processing Magazine, vol. 1053, no. 5888/19,2019.

[2] D. R. Tuohy and W. D. Potter, “A genetic algorithm for the automaticgeneration of playable guitar tablature,” in ICMC, pp. 499– 502, 2005.

[3] K. Yazawa, D. Sakaue, K. Nagira, K. Itoyama, and H. G. Okuno,“Audio-based guitar tablature transcription using multipitch analysisand playability constraints,” in 2013 IEEE International Conference onAcoustics, Speech and Signal Processing, pp. 196–200, IEEE, 2013.

[4] B. Gowrishankar and N. U. Bhajantri, “An exhaustive review of automatic music transcription techniques: Survey of music transcription techniques,” in 2016 International Conference on Signal Processing, Communication, Power and Embedded System (SCOPES), pp. 140–152,IEEE, 2016.

[5] B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, “Creating amultitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,” IEEE Transactions on Multimedia, vol. 21, no. 2, pp. 522–535, 2018.

[6] Y.-N. Hung, G. Wichern, and J. Le Roux, “Transcription is all you need: Learning to separate musical mixtures with score as supervision,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 46–50, IEEE, 2021.

[7] A. Jansson, E. Humphrey, N. Montecchio, R. Bittner, A. Kumar, and T. Weyde, “Singing voice separation with deep u-net convolutional networks,” 2017.

[8] L. Lin, Q. Kong, J. Jiang, and G. Xia, “A unified model for zero-shot music source separation, transcription and synthesis,” arXiv preprintarXiv:2108.03456, 2021.

[9] E. Manilow, P. Seetharaman, and B. Pardo, “Simultaneous separation and transcription of mixtures with multiple polyphonic and percussive instruments,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 771–775, IEEE,2020.

[10] Y.-N. Hung and A. Lerch, “Multitask learning for instrument activation aware music source separation,” arXiv preprint arXiv:2008.00616, 2020.

[11] K. A. Pati and A. Lerch, “A dataset and method for guitar solo detection in rock music,” in Audio Engineering Society Conference: 2017AES International Conference on Semantic Audio, Audio Engineering Society, 2017.

[12] D. Règnier, N. Martin, and L. Bigo, “Identification of rhythm guitar sections in symbolic tablatures,” in International Society for Music Information Retrieval Conference (ISMIR 2021), 2021.

[13] Y.-T. Wu, B. Chen, and L. Su, “Multi- instrument automatic music transcription with self-attention-based instance segmentation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28,pp. 2796–2809, 2020.

[14] K. Tanaka, T. Nakatsuka, R. Nishikimi, K. Yoshii, and S. Morishima, “Multi-instrument music transcription based on deep spherical clustering of spectrograms and pitchgrams.,” in ISMIR, pp. 327–334, 2020.

[15] F. Simonetta, S. Ntalampiras, and F. Avanzini, “Audio-to-score alignment using deep automatic music transcription,” in 2021 IEEE 23 rd International Workshop on Multimedia Signal Processing (MMSP), pp. 1–6, IEEE, 2021.

[16] N. Takahashi, N. Goswami, and Y. Mitsufuji, “Mmdenselstm: An efficient combination of convolutional and recurrent neural networks for audio source separation,” in 2018 16th International workshop on acoustic signal enhancement (IWAENC), pp. 106–110, IEEE, 2018.

[17] I. Barbancho, L. J. Tardon, S. Sammartino, and A. M. Barbancho, “In harmonicity-based method for the automatic generation of guitar tablature,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 20, no. 6, pp. 1857–1868, 2012.

[18] E. Mistler, “Generating guitar tablatures with neural networks,” Master of Science Dissertation, The University of Edinburgh, 2017.

[19] E. Manilow, G. Wichern, P. Seetharaman, and J. Le Roux, “Cutting music source separation some slakh: A dataset to study the impact of training data quality and quantity,” in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA),pp. 45–49, IEEE, 2019.

[20] Q. Xi, R. M. Bittner, J. Pauwels, X. Ye, and J. P. Bello, “Guitarset: Adataset for guitar transcription.,” in ISMIR, pp. 453–460, 2018.

How to cite this paper

Pranav Bhagwat, Vishwajit Shelke, Akhilesh Murugkar, Krishiv Dakwala, Shweta C. Dharmadhikari "A Survey on Automatic Music Transcription" Iconic Research And Engineering Journals Volume 6 Issue 11 2023 Page 268-274
Pranav Bhagwat, Vishwajit Shelke, Akhilesh Murugkar, Krishiv Dakwala, Shweta C. Dharmadhikari "A Survey on Automatic Music Transcription" Iconic Research And Engineering Journals, vol. 6, no. 11, May. 2023
Pranav Bhagwat, Vishwajit Shelke, Akhilesh Murugkar, Krishiv Dakwala, Shweta C. Dharmadhikari (2023). A Survey on Automatic Music Transcription. Iconic Research And Engineering Journals, 6(11).
Pranav Bhagwat, Vishwajit Shelke, Akhilesh Murugkar, Krishiv Dakwala, Shweta C. Dharmadhikari "A Survey on Automatic Music Transcription" Iconic Research And Engineering Journals, vol. 6, no. 11, May. 2023.
@article{1704431,
      author = {Pranav Bhagwat, Vishwajit Shelke, Akhilesh Murugkar, Krishiv Dakwala, Shweta C. Dharmadhikari},
      title = {A Survey on Automatic Music Transcription},
      journal = {Iconic Research And Engineering Journals},
      year = {2023},
      volume = {6},
      number = {11},
      pages = {268-274},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1704431.pdf},
      abstract = {Automatic Music Transcription (AMT) is a critical but less investigated problem in the field of music information retrieval. In this paper, we study different approaches for achieving Automatic Music Transcription using various methods based on pitch, timbre and note detection. Use of Convolutional Neural Network (CNN) and/or Long Short Term Memory Network (LSTM) is made to transcribe notes from the audio input. We also discuss source separation as a precursor to AMT and different approaches for the same.},
      keywords = {Automatic Music Transcription (AMT), Pitch Detection, Note Detection, Deep Learning, Convolutional Neural Network (CNN), Long Short Term Memory Network (LSTM), Source Separation},
      month = {May},
  }