Beam mismatch in polarization-sensitive bolometer pairs produces temperature-to-polarization (TP) leakage that contaminates cosmic microwave background (CMB) polarization maps and biases constraints on the tensor-to-scalar ratio. Conventional deprojection methods rely on the specific Gaussian beam model and filter the time-ordered data, introducing some extra EB mixing that must be corrected. In this work, we demonstrate that a deep learning approach can efficiently remove these systematics directly at the map level. Using an Ali CMB Polarization Telescope (AliCPT)-like telescope configuration and differential beam parameters drawn from BICEP/Keck measurements, we generate 4000 full-sky mock Q and U maps containing both the true primordial CMB signal and the TP leakage induced by beam mismatch. We divide each full-sky map into 48 square patches of size 512 × 512 and consider 12 of them that enclose the observation field of AliCPT. For each of these square regions, a multi-patch hierarchical convolutional neural network based on the U-Net architecture is trained on 3600 randomly selected sky patches extracted from the contaminated maps; the corresponding uncontaminated patches serve as ground truth. On an independent test set, the trained network recovers the true polarization maps with a mean absolute deviation of 0.024 μK in Q and 0.023 μK in U. The reconstructed BB and EE angular power spectra match the input cosmology to within residuals of order 10−4 μK2 and 10−3 μK2, respectively. Operating directly on maps, our neural network model avoids the filtering-induced EB mixing and achieves millisecond-level inference time per sample, making it a promising alternative to traditional techniques under the simulation assumptions considered here and potentially applicable to next-generation CMB experiments such as AliCPT and LiteBIRD.