Skip to main navigation menu Skip to main content Skip to site footer

Online Preference Elicitation for Personalized Singing Synthesis: A Comprehensive Framework for User-Centric Acoustic Adaptation

Language, Knowledge and Intelligent Systems, Volume 1, Issue 4, 2026 cover

Abstract

The rapid advancement of artificial intelligence in audio generation has significantly improved the naturalness and intelligibility of singing synthesis systems. However, a critical gap remains in aligning the generated vocal performances with the highly subjective and individualized preferences of different users. Conventional systems rely on static datasets and generalized models that output standardized acoustic features, largely ignoring idiosyncratic tastes regarding vibrato depth, breathiness, tension, and emotional delivery. This paper presents a comprehensive conceptual framework and methodology for online preference elicitation designed specifically for personalized singing synthesis. By integrating human-in-the-loop machine learning techniques, the proposed system dynamically interacts with the user, presenting iteratively generated vocal samples and capturing feedback in real time to update an underlying preference model. We explore the parameterization of acoustic features relevant to vocal aesthetics and detail the active learning strategies employed to minimize user fatigue while maximizing the convergence rate of the preference vector. Through detailed architectural design and extensive experimental validation involving both subjective perceptual evaluations and objective convergence metrics, we demonstrate that online preference elicitation significantly outperforms static personalization methods. The findings indicate that users achieve a higher degree of satisfaction and perceived control over the synthesis output. This research bridges the gap between sophisticated generative audio models and human-computer interaction, paving the way for intuitive, accessible, and deeply personalized creative tools in music production.

Keywords

Personalized Singing Synthesis, Online Preference Elicitation, Active Learning, Human-Computer Interaction

PDF

References

  1. 1. Yu, H., Zhang, J., Chen, C., Xiang, T., Fang, Y., Niebles, J. C., & Adeli, E. (2026, March). Socialgen: Modeling multi-human social interaction with language models. In 2026 International Conference on 3D Vision (3DV) (pp. 1-17). IEEE.
  2. 2. Zhang, B., Wang, S., Jiang, Y., Sui, D., Tu, Z., & Chu, D. (2025, July). Ask and Retrieve Knowledge: Towards Proactive Asking with Imperfect Information in Medical Multi-turn Dialogues. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 1055-1065).
  3. 3. Yao, J., Jia, Y., Wang, X., & Khan, A. N. (2026). The digital gaze in retail encounters: AI face dialogues and consumer preference in the circular economy. Journal of Retailing and Consumer Services, 93, 104915.
  4. 4. Zhou, J., Xiong, H., Lu, J., Lin, Z., & Feng, B. (2025, September). Cgtgait: Collaborative graph and transformer for gait emotion recognition. In 2025 IEEE International Joint Conference on Biometrics (IJCB) (pp. 1-11). IEEE.
  5. 5. Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., et al. (2016). Deep Speech 2: End-to-end speech recognition in English and Mandarin. In Proceedings of ICML.
  6. 6. Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2023). Robust speech recognition via large-scale weak supervision. In Proceedings of ICML.
  7. 7. Zhao, R., Tang, J., Zeng, W., Guo, Y., & Zhao, X. (2025). Towards human-like questioning: Knowledge base question generation with bias-corrected reinforcement learning from human feedback. Information Processing & Management, 62(3), 104044.
  8. 8. Zhao, R., Zeng, W., Tang, J., Li, Y., Ye, G., Du, J., & Zhao, X. (2025, May). Towards unsupervised entity alignment for highly heterogeneous knowledge graphs. In 2025 IEEE 41St international conference on data engineering (ICDE) (pp. 3792-3806). IEEE.
  9. 9. Sun, L., Hu, J., Li, M., & Peng, H. (2024, July). R-ode: Ricci curvature tells when you will be informed. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 2594-2598).
  10. 10. Gu, C., Zhang, W., Huang, Z., Kou, J., Liu, Z., Zhao, C., Liu, C., Zhang, L., Lin, W., Wang, Z., Deng, J., Xie, Y., Huang, G., Zhang, C., Lu, X., Wang, C., Zhang, Z., Yuan, H., Duan, X., & Fang, Y. (2024). LENS: Layers of evaluation of hallucination in GenAI systems. In 2024 7th International Conference on Universal Village (UV) (pp. 1–85). IEEE. https://doi.org/10.1109/UV63228.2024.11189150
  11. 11. Yao, Y., Zhu, Z., Miao, P., Cheng, X., Shu, F., & Wang, J. (2025). Optimizing hybrid RIS-aided ISAC systems in V2X networks: A deep reinforcement learning method for anti-eavesdropping techniques. IEEE Transactions on Vehicular Technology, 74 (6), 9224-9239.
  12. 12. Wei, M., Li, R., Wang, L., Chang, Z., Xu, L., & Han, Z. (2026). Toward Adaptive Tracking and Communication via an Airborne Maneuverable Bi-Static ISAC System. IEEE Transactions on Vehicular Technology.
  13. 13. Wang, Y., He, S., Chen, G., Chen, Y., & Jiang, D. (2022, December). XLM-D: Decorate cross-lingual pre-training model as non-autoregressive neural machine translation. In Proceedings of the 2022 conference on empirical methods in natural language processing (pp. 6934-6946).
  14. 14. Li, Q., Ye, Q., Zhang, N., Zhang, W., & Hu, F. (2025). Digital-twin-enabled industrial IoT: Vision, framework, and future directions. IEEE Wireless Communications, 32(6), 173-181.
  15. 15. Li, S. (2025). Momentum, volume and investor sentiment study for us technology sector stocks—A hidden markov model based principal component analysis. PloS one, 20(9), e0331658.
  16. 16. Wang, S., Zhang, X., Wang, P., & Li, X. (2026). Design of Magnetic Sensor Array-Based PUFs for IoT Security. IEEE Transactions on Instrumentation and Measurement, 75, 1-9.
  17. 17. Qu, W., Shao, Y., Meng, L., Huang, X., & Xiao, L. (2024, June). A conditional denoising diffusion probabilistic model for point cloud upsampling. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 20786-20795). IEEE.
  18. 18. Zhang, Q., Yang, M., Sun, G., Xiang, Y., Wang, S., Zhang, J., & Li, S. (2026). Representation learning accelerates the development of models for Li-ion battery health diagnostics and prognostics. Energy Storage Materials, 104897.
  19. 19. Yang, Z., Ji, W., Guo, Q., & Wang, Z. (2023, October). Javp: Joint-aware video processing with edge-cloud collaboration for dnn inference. In Proceedings of the 31st ACM International Conference on Multimedia (pp. 9152-9160).
  20. 20. Wang, Jiacheng, et al. "A Dynamic Factor Gating Architecture with Market Regime Awareness for Stock Return Forecasting." (2026).
  21. 21. Cao, X., Tao, J., Liu, Z., Lyu, R., & Li, J. (2026). Handling Missing Data in CALL: A Data Quality-Driven Imputation Framework for Learner Analytics. Future-Adaptive Intelligence and Lifelong Systems, 1 (1).
  22. 22. Abu Saleh, A., Siouris, S., Hughes, K., Yuan, R., & Pourkashanian, M. (2023). Assessment of mixtures of iso-pentanol and Jet A-1 for use in aviation gas turbine engines. In AIAA SCITECH 2023 Forum (p. 2341).