Papers › Depth F₁: Improving Evaluation of Cross-Domain Text Classification by Measuring...

Depth F₁: Improving Evaluation of Cross-Domain Text Classification by Measuring Semantic Generalizability

20 Jun 2024arXiv:2406.14695archive 2025-07-28

Parker Seegmiller, Joseph Gatto, Sarah Masud Preum

Recent evaluations of cross-domain text classification models aim to measure the ability of a model to obtain domain-invariant performance in a target domain given labeled samples in a source domain. The primary strategy for this evaluation relies on assumed differences between source domain samples and target domain samples in benchmark datasets. This evaluation strategy fails to account for the similarity between source and target domains, and may mask when models fail to transfer learning to specific target samples which are highly dissimilar from the source domain. We introduce Depth F₁, a novel cross-domain text classification performance metric. Designed to be complementary to existing classification metrics such as F₁, Depth F₁ measures how well a model performs on target samples which are dissimilar from the source domain. We motivate this metric using standard cross-domain text classification datasets and benchmark several recent cross-domain text classification models, with the goal of enabling in-depth evaluation of the semantic generalizability of cross-domain text classification models.

PaperPDFCode

Code

pkseeg/df1 officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationCross-Domain Text ClassificationText ClassificationTransfer Learningtext-classification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections