{"url":"/dataset/v2-train-pkl","name":"address_parser_data","full_name":null,"description_markdown":"This is a set of datasets containing three versions of data:\r\n\r\n- V0: the original sampled addresses with no augmentation.\r\n- V1: Given V0, we apply basic cleaning and address structure masking technique (rearrangement or removal, but no addition of augmented address parts).\r\n- V2: Given V0, we apply more advanced augmentation techniques (see paper for more info)\r\n\r\nFor each version we have three types of dataset:\r\n\r\n- Train: around 3M datapoints used for training\r\n- Test: around 100k datapoints used for testing\r\n- Zero-shot: around 300k datapoints used for validation, extracting from addresses in countries not seen in train and test","description_withheld":null,"homepage":"https://huggingface.co/datasets/hm-haitham/address_parser_data","introduced_date":"2024-04-08","introduced_date_note":null,"introduced_by":{"paper":"/paper/fighting-crime-with-transformers-empirical","title":"Fighting crime with Transformers: Empirical analysis of address parsing methods in payment data","first_author":"Haitham Hammami","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["address_parser_data"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}