Papers › English Intermediate-Task Training Improves Zero-Shot Cross-Lingual Transfer Too
English Intermediate-Task Training Improves Zero-Shot Cross-Lingual Transfer Too
Jason Phang, Iacer Calixto, Phu Mon Htut, Yada Pruksachatkun, Haokun Liu, Clara Vania, Katharina Kann, Samuel R. Bowman
Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tasks and moderate improvements on question-answering target tasks. MNLI, SQuAD and HellaSwag achieve the best overall results as intermediate tasks, while multi-task intermediate offers small additional improvements. Using our best intermediate-task models for each target task, we obtain a 5.4 point improvement over XLM-R Large on the XTREME benchmark, setting the state of the art as of June 2020. We also investigate continuing multilingual MLM during intermediate-task training and using machine-translated intermediate-task data, but neither consistently outperforms simply performing English intermediate-task training.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Zero-Shot Cross-Lingual Transfer | XTREME | X-STILTs | Avg | 73.5 | #20 of 25 | Archive leaderboard | report |
| Zero-Shot Cross-Lingual Transfer | XTREME | X-STILTs | Question Answering | 67.2 | #20 of 25 | Archive leaderboard | report |
| Zero-Shot Cross-Lingual Transfer | XTREME | X-STILTs | Sentence Retrieval | 76.5 | #20 of 25 | Archive leaderboard | report |
| Zero-Shot Cross-Lingual Transfer | XTREME | X-STILTs | Sentence-pair Classification | 83.9 | #20 of 25 | Archive leaderboard | report |
| Zero-Shot Cross-Lingual Transfer | XTREME | X-STILTs | Structured Prediction | 69.4 | #20 of 25 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections