{"url":"/dataset/pathfinder-x2","name":"Pathfinder-X2","full_name":null,"description_markdown":"Pathfinder and Pathfinder-X have proven to be instrumental in training and testing Large Language Models with long-range dependencies.\r\nRecently, Meta's Moving Average Equipped Gated Attention model scored a 97% on the Pathfinder-X dataset, indicating a need for a larger,\r\nmore challenging dataset.  Whereas Pathfinder-X only went up to 256 x 256 pixel images (or a sequence length of 65,536 tokens), Pathfinder-X2 introduces images of 512 x 512 pixels, or 262,144 tokens.  \r\n\r\nEach image is meant to be read as a sequence of pixels.  A LLM's task is to segment out the one snake in each image with a circle at its tip.  The dataset includes 200,000 images and 200,000 segmentation masks, one for each image.","description_withheld":null,"homepage":"https://huggingface.co/datasets/Tylersuard/PathfinderX2","introduced_date":"2023-04-13","introduced_date_note":null,"introduced_by":null,"license":{"name":"C.C. B.Y. 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Long-range modeling","url":"/task/long-range-modeling","datasets_with_task":"/datasets/task/long-range-modeling"}],"languages":[],"variants":["Pathfinder-X2"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}