{"url":"/dataset/lcstep","name":"LCStep","full_name":null,"description_markdown":"For our experiments, we collected a dataset of procedural knowledge of the LangChain Python library, unseen by many extant LLMs. We selected LangChain as the domain for our dataset because it was published in 2022, which is later than the knowledge cutoff date for many web-scale LLMs, including GPT-3.5, while also having plenty of documentation due to its popularity.\r\n\r\nThe LCStep dataset was collected from 180 tutorial pages in the Python section of the LangChain website. We used an LLM-enabled pipeline with human oversight and quality review to extract 276 procedures from these tutorials, representing each procedure in a format structure.","description_withheld":null,"homepage":"https://arxiv.org/abs/2409.01344","introduced_date":"2024-09-02","introduced_date_note":null,"introduced_by":{"paper":"/paper/pairing-analogy-augmented-generation-with","title":"Pairing Analogy-Augmented Generation with Procedural Memory for Procedural Q&A","first_author":"K Roth","url":null},"license":{"name":"MIT","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Procedure Learning","url":"/task/procedure-learning","datasets_with_task":"/datasets/task/procedure-learning"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["LCStep"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}