학술논문

AudRandAug: Random Image Augmentations for Audio Classification

Document Type

Working Paper

Author

Kumar, Teerath; Turab, Muhammad; Mileo, Alessandra; Bendechache, Malika; Saber, Takfarinas

Source

Subject

Computer Science - Sound
Computer Science - Artificial Intelligence
Computer Science - Computer Vision and Pattern Recognition
Computer Science - Machine Learning
Electrical Engineering and Systems Science - Audio and Speech Processing

Language

Abstract

Data augmentation has proven to be effective in training neural networks. Recently, a method called RandAug was proposed, randomly selecting data augmentation techniques from a predefined search space. RandAug has demonstrated significant performance improvements for image-related tasks while imposing minimal computational overhead. However, no prior research has explored the application of RandAug specifically for audio data augmentation, which converts audio into an image-like pattern. To address this gap, we introduce AudRandAug, an adaptation of RandAug for audio data. AudRandAug selects data augmentation policies from a dedicated audio search space. To evaluate the effectiveness of AudRandAug, we conducted experiments using various models and datasets. Our findings indicate that AudRandAug outperforms other existing data augmentation methods regarding accuracy performance.
Comment: Paper has accepted at 25th Irish Machine Vision and Image Processing Conference

Online Access

Open Access (Arxiv) Find it@PNU

이메일

부산대학교 도서관

Online Access

메일 발송