Papers › Gendered Ambiguous Pronouns Shared Task: Boosting Model Confidence by Evidence Pooling

Gendered Ambiguous Pronouns Shared Task: Boosting Model Confidence by Evidence Pooling

1 Aug 2019WS 2019 8archive 2025-07-28

S Attree, eep

This paper presents a strong set of results for resolving gendered ambiguous pronouns on the Gendered Ambiguous Pronouns shared task. The model presented here draws upon the strengths of state-of-the-art language and coreference resolution models, and introduces a novel evidence-based deep learning architecture. Injecting evidence from the coreference models compliments the base architecture, and analysis shows that the model is not hindered by their weaknesses, specifically gender bias. The modularity and simplicity of the architecture make it very easy to extend for further improvement and applicable to other NLP problems. Evaluation on GAP test data results in a state-of-the-art performance at 92.5{\%} F1 (gender bias of 0.97), edging closer to the human performance of 96.6{\%}. The end-to-end solution presented here placed 1st in the Kaggle competition, winning by a significant lead.

PaperPDFCode

Code

sattree/gap officialmentioned in papertf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Coreference Resolutioncoreference-resolution

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Coreference Resolution GAP ProBERT Bias (F/M) 0.97 #2 of 5 Archive leaderboard report
Coreference Resolution GAP ProBERT Feminine F1 (F) 91.1 #2 of 5 Archive leaderboard report
Coreference Resolution GAP ProBERT Masculine F1 (M) 94.0 #2 of 5 Archive leaderboard report
Coreference Resolution GAP ProBERT Overall F1 92.5 #2 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections