Papers › Weisfeiler-Leman in the BAMBOO: Novel AMR Graph Metrics and a Benchmark for AMR Graph...

Weisfeiler-Leman in the BAMBOO: Novel AMR Graph Metrics and a Benchmark for AMR Graph Similarity

26 Aug 2021arXiv:2108.11949archive 2025-07-28

Juri Opitz, Angel Daza, Anette Frank

Several metrics have been proposed for assessing the similarity of (abstract) meaning representations (AMRs), but little is known about how they relate to human similarity ratings. Moreover, the current metrics have complementary strengths and weaknesses: some emphasize speed, while others make the alignment of graph structures explicit, at the price of a costly alignment step. In this work we propose new Weisfeiler-Leman AMR similarity metrics that unify the strengths of previous metrics, while mitigating their weaknesses. Specifically, our new metrics are able to match contextualized substructures and induce n:m alignments between their nodes. Furthermore, we introduce a Benchmark for AMR Metrics based on Overt Objectives (BAMBOO), the first benchmark to support empirical assessment of graph-based MR similarity metrics. BAMBOO maximizes the interpretability of results by defining multiple overt objectives that range from sentence similarity objectives to stress tests that probe a metric's robustness against meaning-altering and meaning-preserving graph transformations. We show the benefits of BAMBOO by profiling previous metrics and our own metrics. Results indicate that our novel metrics may serve as a strong baseline for future work.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AMR Graph SimilarityGraph MatchingGraph SimilaritySentenceSentence Similarity

Datasets

Introduced by this paper, per the archive.

Benchmark for AMR Metrics based on Overt Objectives

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives WWLKΘ Pearson’s ρ (amean) 54.90 #2 of 9 Archive leaderboard report
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives S2MATCH Pearson’s ρ (amean) 52.22 #3 of 9 Archive leaderboard report
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives S2MATCH Spearman Correlation 50.89 #3 of 9 Archive leaderboard report
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives SMATCH Pearson’s ρ (amean) 51.28 #4 of 9 Archive leaderboard report
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives SMATCH Spearman Correlation 50.44 #4 of 9 Archive leaderboard report
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives WLK Pearson’s ρ (amean) 50.44 #5 of 9 Archive leaderboard report
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives WLK Spearman Correlation 49.64 #5 of 9 Archive leaderboard report
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives SEMA Pearson’s ρ (amean) 46.29 #6 of 9 Archive leaderboard report
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives SEMBLEU, k=4 Pearson’s ρ (amean) 46.15 #7 of 9 Archive leaderboard report
AMR Graph Similarity Benchmark for AMR Metrics based on Overt Objectives WWLK Pearson’s ρ (amean) 45.30 #8 of 9 Archive leaderboard report
Graph Matching RARE WLK Spearman Correlation 90.39 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections