Papers › Strong Baselines for Neural Semi-supervised Learning under Domain Shift
Strong Baselines for Neural Semi-supervised Learning under Domain Shift
Sebastian Ruder, Barbara Plank
Novel neural models have been proposed in recent years for learning under domain shift. Most models, however, only evaluate on a single task, on proprietary datasets, or compare to weak baselines, which makes comparison of models difficult. In this paper, we re-evaluate classic general-purpose bootstrapping approaches in the context of neural networks under domain shifts vs. recent neural approaches and propose a novel multi-task tri-training method that reduces the time and space complexity of classic tri-training. Extensive experiments on two benchmarks are negative: while our novel method establishes a new state-of-the-art for sentiment analysis, it does not fare consistently the best. More importantly, we arrive at the somewhat surprising conclusion that classic tri-training, with some additions, outperforms the state of the art. We conclude that classic approaches constitute an important and strong baseline.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Sentiment Analysis | Multi-Domain Sentiment Dataset | Multi-task tri-training | Average | 79.15 | #3 of 6 | Archive leaderboard | report |
| Sentiment Analysis | Multi-Domain Sentiment Dataset | Multi-task tri-training | Books | 74.86 | #3 of 6 | Archive leaderboard | report |
| Sentiment Analysis | Multi-Domain Sentiment Dataset | Multi-task tri-training | DVD | 78.14 | #3 of 6 | Archive leaderboard | report |
| Sentiment Analysis | Multi-Domain Sentiment Dataset | Multi-task tri-training | Electronics | 81.45 | #3 of 6 | Archive leaderboard | report |
| Sentiment Analysis | Multi-Domain Sentiment Dataset | Multi-task tri-training | Kitchen | 82.14 | #3 of 6 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections