FrontierThe story, in brief

Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR

Apple's new pseudo-labeling approach cuts through code-switching ASR's data scarcity problem — and it works.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Apple Research demonstrates that iterative pseudo-labeling can significantly improve speech recognition across language-switching scenarios, a capability gap that limits ASR deployment in multilingual markets. This is a training methodology advance, not a product feature.

The key facts

10 to know
  1. First application of iterative pseudo-labeling to code-switching ASR

  2. Three-phase approach: pseudo-label generation, two-stage bilingual training, iterative refinement

  3. Addresses data scarcity in code-switching scenarios (Mandarin-English focus)

  4. Leverages unlabeled data via semi-supervised learning

  5. Published by Apple Machine Learning Research

  6. Date: August 2026

  7. Iterative pseudo-labeling applied to code-switching ASR for the first time

  8. Three-phase approach: pseudo-label generation, bilingual model training, iterative refinement

  9. Semi-supervised learning leverages unlabeled data to improve CS-ASR

  10. Mandarin-English code-switching focus

Go to the source

Apple Machine Learningmachinelearning.apple.com

Publisher excerpt: Code-switching (CS), alternating languages within the same utterance, poses significant challenges for automatic speech recognition (ASR) due to limited CS training data. This paper applies an iterative pseudo-labeling training approach to CS-ASR for the first time, demonstrating its effectiveness…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier