Enhancing Semantic Matching and Text Robustness through Masked Language Modeling

Main Article Content

Zhen Bian
Xin Su

Abstract

This paper investigates the application of masked language models in semantic mining. It proposes a deep semantic modeling method based on selfsupervised learning. The method randomly masks parts of the input text and uses contextual information to predict the masked words. This enables effective modeling of deep semantic structures in text. A unified training framework is designed, integrating a cross-entropy loss function with a regularization mechanism. This improves both the semantic representation ability and training stability of the model. The model structure retains the core features of masked language model pretraining. It also incorporates task-specific classification layers to support various downstream semantic mining tasks. The experiments evaluate the model from multiple perspectives, including classification performance, semantic match, differences in text length, and robustness to text perturbations. Results show that the proposed method outperforms mainstream pretrained language models across several metrics. It is especially strong in semantic matching and contextual understanding. The study also constructs multiple semantic perturbation scenarios to analyze model robustness. It confirms the model's ability to preserve semantic integrity under complex variations. By integrating corpus structure, prediction behavior, and task adaptability, the proposed semantic mining method offers an effective solution to enhance semantic understanding in natural language processing systems.

Article Details

Section

Articles