Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR
Apple's new pseudo-labeling approach cuts through code-switching ASR's data scarcity problem — and it works.

Why it matters
Apple Research demonstrates that iterative pseudo-labeling can significantly improve speech recognition across language-switching scenarios, a capability gap that limits ASR deployment in multilingual markets. This is a training methodology advance, not a product feature.
The key facts
10 to knowFirst application of iterative pseudo-labeling to code-switching ASR
Three-phase approach: pseudo-label generation, two-stage bilingual training, iterative refinement
Addresses data scarcity in code-switching scenarios (Mandarin-English focus)
Leverages unlabeled data via semi-supervised learning
Published by Apple Machine Learning Research
Date: August 2026
Iterative pseudo-labeling applied to code-switching ASR for the first time
Three-phase approach: pseudo-label generation, bilingual model training, iterative refinement
Semi-supervised learning leverages unlabeled data to improve CS-ASR
Mandarin-English code-switching focus
Go to the source
Apple Machine Learningmachinelearning.apple.com
Publisher excerpt: Code-switching (CS), alternating languages within the same utterance, poses significant challenges for automatic speech recognition (ASR) due to limited CS training data. This paper applies an iterative pseudo-labeling training approach to CS-ASR for the first time, demonstrating its effectiveness…