{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-language-from-a-large-unannotated","title":"Learning Language from a Large (Unannotated) Corpus","arxiv_id":"1401.3372","date":"2014-01-14","proceeding":null,"authors":["Linas Vepstas","Ben Goertzel"],"abstract":"A novel approach to the fully automated, unsupervised extraction of\ndependency grammars and associated syntax-to-semantic-relationship mappings\nfrom large text corpora is described. The suggested approach builds on the\nauthors' prior work with the Link Grammar, RelEx and OpenCog systems, as well\nas on a number of prior papers and approaches from the statistical language\nlearning literature. If successful, this approach would enable the mining of\nall the information needed to power a natural language comprehension and\ngeneration system, directly from a large, unannotated corpus.","url_abs":"http://arxiv.org/abs/1401.3372v1","url_pdf":"http://arxiv.org/pdf/1401.3372v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-language-from-a-large-unannotated","repo_url":"https://github.com/opencog/learn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}