Papers › CLMSM: A Multi-Task Learning Framework for Pre-training on Procedural Text

CLMSM: A Multi-Task Learning Framework for Pre-training on Procedural Text

22 Oct 2023arXiv:2310.14326archive 2025-07-28

Abhilash Nandy, Manav Nitin Kapadnis, Pawan Goyal, Niloy Ganguly

In this paper, we propose CLMSM, a domain-specific, continual pre-training framework, that learns from a large set of procedural recipes. CLMSM uses a Multi-Task Learning Framework to optimize two objectives - a) Contrastive Learning using hard triplets to learn fine-grained differences across entities in the procedures, and b) a novel Mask-Step Modelling objective to learn step-wise context of a procedure. We test the performance of CLMSM on the downstream tasks of tracking entities and aligning actions between two procedures on three datasets, one of which is an open-domain dataset not conforming with the pre-training dataset. We show that CLMSM not only outperforms baselines on recipes (in-domain) but is also able to generalize to open-domain procedural NLP tasks.

PaperPDFCode

Code

manavkapadnis/clmsm_emnlp_2023 officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Contrastive LearningLanguage ModellingMulti-Task LearningProcedural Text Understanding

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Contrastive LearningSET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections