Papers › Avicenna: a challenge dataset for natural language generation toward commonsense...

Avicenna: a challenge dataset for natural language generation toward commonsense syllogistic reasoning

24 Dec 2022Journal of Applied Non-Classical Logics 2022 12archive 2025-07-28

Zeinab Aghahadi, Alireza Talebpour

Syllogism is a type of everyday reasoning. For instance, given that "Avicenna wrote the famous book the Canon of Medicine " and " The Canon of Medicine has influenced modern medicine," it can be concluded that "Avicenna has influenced modern medicine." This study revolves around syllogistic natural language generation (NLG). The Avicenna corpus was developed as a benchmark for syllogistic NLG. In this respect, once the syllogistic relation between two premises is recognized, the Avicenna-trained models learn to generate the conclusion sentence (which is semantically unique). The experiments were performed using state-of-the-art pre-trained text generative models (TGMs). The state-of-the-art baseline was provided, and the accuracy was improved up to 32% when transfer learning was adopted. The models were evaluated using human and automatic procedures. The model’s confusion in detecting the middle-term (the duplicate part with the same meaning in the premises) was one of the main categories of errors that showed up in the error analysis. This issue indicates that the model learns how to extract new facts based on the premises, but it faces a challenge in commonsense reasoning. The outcomes of the study demonstrated that although TGMs are significantly powerful, they do not yet have sufficient reasoning capabilities to generate text, based on commonsense knowledge. The Avicenna dataset poses a new challenge of commonsense inference that is easy for humans (98.1%) while difficult for state-of-the-art TGMs (32%).

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

SentenceText GenerationTransfer Learning

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections