Code comments serve a crucial role in software development for documenting functionality, clarifying design choices, and assisting with issue tracking. They capture developers' insights about the surrounding source code, serving as an essential resource for both human comprehension and automated analysis. Nevertheless, since comments are in natural language, they present challenges for machine-based code understanding. To address this, recent studies have applied natural language processing (NLP) and deep learning techniques to classify comments according to developers' intentions. However, existing datasets for this task suffer from size limitations and class imbalance, as they rely on manual annotations and may not accurately represent the distribution of comments in real-world codebases. To overcome this issue, we introduce new synthetic oversampling and augmentation techniques based on high-quality data generation to enhance the NLBSE'26 challenge datasets. Our Synthetic Quality Oversampling Technique and Augmentation Technique (Q-SYNTH) yield promising results, improving the base classifier by 2.56\%.
- High-quality data augmentation for code comment classification
- Thomas BorsaniAndrea RosaniGiuseppe Di Fatta
- Proceedings: 2026 IEEE/ACM International Workshop on NL-based Software Engineering, NLBSE 2026, pp.38-41
- 9798400723964
- The 5th International Workshop on Natural Language-based Software Engineering (Rio de Janeiro, 12/04/2026–13/04/2026)
- ACM
- Online
- 4
- 979-840072396-4
(UNIBZ)93122322
991007292751101241 - 2-s2.0-105047266810
- This work is licensed under a Creative Commons Attribution 4.0 International License
- Faculty of Engineering
- English
- Conference proceeding
- Borsani T, Rosani A, Di Fatta G