{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/job-sdf-a-multi-granularity-dataset-for-job","title":"Job-SDF: A Multi-Granularity Dataset for Job Skill Demand Forecasting and Benchmarking","arxiv_id":"2406.11920","date":"2024-06-17","proceeding":null,"authors":["Xi Chen","Chuan Qin","Chuyu Fang","Chao Wang","Chen Zhu","Fuzhen Zhuang","HengShu Zhu","Hui Xiong"],"abstract":"In a rapidly evolving job market, skill demand forecasting is crucial as it enables policymakers and businesses to anticipate and adapt to changes, ensuring that workforce skills align with market needs, thereby enhancing productivity and competitiveness. Additionally, by identifying emerging skill requirements, it directs individuals towards relevant training and education opportunities, promoting continuous self-learning and development. However, the absence of comprehensive datasets presents a significant challenge, impeding research and the advancement of this field. To bridge this gap, we present Job-SDF, a dataset designed to train and benchmark job-skill demand forecasting models. Based on 10.35 million public job advertisements collected from major online recruitment platforms in China between 2021 and 2023, this dataset encompasses monthly recruitment demand for 2,324 types of skills across 521 companies. Our dataset uniquely enables evaluating skill demand forecasting models at various granularities, including occupation, company, and regional levels. We benchmark a range of models on this dataset, evaluating their performance in standard scenarios, in predictions focused on lower value ranges, and in the presence of structural breaks, providing new insights for further research. Our code and dataset are publicly accessible via the https://github.com/Job-SDF/benchmark.","url_abs":"https://arxiv.org/abs/2406.11920v3","url_pdf":"https://arxiv.org/pdf/2406.11920v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"job-sdf-a-multi-granularity-dataset-for-job","repo_url":"https://github.com/job-sdf/benchmark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"demand-forecasting","task_name":"Demand Forecasting"},{"task_slug":"self-learning","task_name":"Self-Learning"}],"methods":[{"method_slug":"align","method_name":"ALIGN"},{"method_slug":"self-learning","method_name":"Self-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2406.11920","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.11920"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/Job-SDF/benchmark","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/job-sdf/benchmark","reach":{"status":"ok"}}],"summary":{"ran":1,"ran_honours":1,"ran_draft_wrong":1,"ran_fixture":2,"unverified":2},"by_repo_kind":{"official":{"samples":7,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"2af0df7394e58e06","entry":"conv1d_fft","repo":"Job-SDF/benchmark","repo_kind":"official","path":"benchmark/multivariate_time_series/layers/ETSformer_EncDec.py","file_url":"https://github.com/Job-SDF/benchmark/blob/HEAD/benchmark/multivariate_time_series/layers/ETSformer_EncDec.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2af0df7394e58e06"}},{"code_sha256_prefix":"592ea8b254b006db","entry":"get_frequency_modes","repo":"Job-SDF/benchmark","repo_kind":"official","path":"benchmark/multivariate_time_series/layers/FourierCorrelation.py","file_url":"https://github.com/Job-SDF/benchmark/blob/HEAD/benchmark/multivariate_time_series/layers/FourierCorrelation.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"592ea8b254b006db"}},{"code_sha256_prefix":"f32036738135c8a8","entry":"get_phi_psi","repo":"Job-SDF/benchmark","repo_kind":"official","path":"benchmark/multivariate_time_series/layers/MultiWaveletCorrelation.py","file_url":"https://github.com/Job-SDF/benchmark/blob/HEAD/benchmark/multivariate_time_series/layers/MultiWaveletCorrelation.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f32036738135c8a8"}},{"code_sha256_prefix":"4ef26472da51e3aa","entry":"legendreDer","repo":"Job-SDF/benchmark","repo_kind":"official","path":"benchmark/multivariate_time_series/layers/MultiWaveletCorrelation.py","file_url":"https://github.com/Job-SDF/benchmark/blob/HEAD/benchmark/multivariate_time_series/layers/MultiWaveletCorrelation.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4ef26472da51e3aa"}},{"code_sha256_prefix":"a54c8c5c47c6a8b8","entry":"phi_","repo":"Job-SDF/benchmark","repo_kind":"official","path":"benchmark/multivariate_time_series/layers/MultiWaveletCorrelation.py","file_url":"https://github.com/Job-SDF/benchmark/blob/HEAD/benchmark/multivariate_time_series/layers/MultiWaveletCorrelation.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a54c8c5c47c6a8b8"}},{"code_sha256_prefix":"fd61ceb881a058a2","entry":"get_mask","repo":"Job-SDF/benchmark","repo_kind":"official","path":"benchmark/multivariate_time_series/layers/Pyraformer_EncDec.py","file_url":"https://github.com/Job-SDF/benchmark/blob/HEAD/benchmark/multivariate_time_series/layers/Pyraformer_EncDec.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fd61ceb881a058a2"}},{"code_sha256_prefix":"05bdfb126191dbb0","entry":"refer_points","repo":"Job-SDF/benchmark","repo_kind":"official","path":"benchmark/multivariate_time_series/layers/Pyraformer_EncDec.py","file_url":"https://github.com/Job-SDF/benchmark/blob/HEAD/benchmark/multivariate_time_series/layers/Pyraformer_EncDec.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"05bdfb126191dbb0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}