Papers › InterMask: 3D Human Interaction Generation via Collaborative Masked Modelling

InterMask: 3D Human Interaction Generation via Collaborative Masked Modelling

13 Oct 2024arXiv:2410.10010archive 2025-07-28

Muhammad Gohar Javed, Chuan Guo, Li Cheng, Xingyu Li

Generating realistic 3D human-human interactions from textual descriptions remains a challenging task. Existing approaches, typically based on diffusion models, often generate unnatural and unrealistic results. In this work, we introduce InterMask, a novel framework for generating human interactions using collaborative masked modeling in discrete space. InterMask first employs a VQ-VAE to transform each motion sequence into a 2D discrete motion token map. Unlike traditional 1D VQ token maps, it better preserves fine-grained spatio-temporal details and promotes spatial awareness within each token. Building on this representation, InterMask utilizes a generative masked modeling framework to collaboratively model the tokens of two interacting individuals. This is achieved by employing a transformer architecture specifically designed to capture complex spatio-temporal interdependencies. During training, it randomly masks the motion tokens of both individuals and learns to predict them. In inference, starting from fully masked sequences, it progressively fills in the tokens for both individuals. With its enhanced motion representation, dedicated architecture, and effective learning strategy, InterMask achieves state-of-the-art results, producing high-fidelity and diverse human interactions. It outperforms previous methods, achieving an FID of $5.154$ (vs $5.535$ for in2IN) on the InterHuman dataset and $0.399$ (vs $5.207$ for InterGen) on the InterX dataset. Additionally, InterMask seamlessly supports reaction generation without the need for model redesign or fine-tuning.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

gohar-malik/intermask officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Motion Synthesis

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Motion Synthesis Inter-X InterMask FID 0.399 #1 of 6 Archive leaderboard report
Motion Synthesis Inter-X InterMask MMDist 3.705 #1 of 6 Archive leaderboard report
Motion Synthesis Inter-X InterMask MModality 2.261 #1 of 6 Archive leaderboard report
Motion Synthesis Inter-X InterMask R-Precision Top3 0.705 #1 of 6 Archive leaderboard report
Motion Synthesis InterHuman InterMask FID 5.154 #1 of 10 Archive leaderboard report
Motion Synthesis InterHuman InterMask MMDist 3.790 #1 of 10 Archive leaderboard report
Motion Synthesis InterHuman InterMask MModality 1.737 #1 of 10 Archive leaderboard report
Motion Synthesis InterHuman InterMask R-Precision Top3 0.683 #1 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformerVQ-VAE

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections