{"url":"/dataset/finsen","name":"FinSen","full_name":null,"description_markdown":"## Enhancing Financial Market Predictions: Causality-Driven Feature Selection\r\n\r\nThis paper introduces FinSen dataset that revolutionizes financial market analysis by integrating economic and financial news articles from 197 countries with stock market data. The dataset’s extensive coverage spans 15 years from 2007 to 2023 with temporal information, offering a rich, global perspective 160,000 records on financial market news. Our study leverages causally validated sentiment scores and LSTM models to enhance market forecast accuracy and reliability.\r\n\r\n\r\n# Our FinSen Dataset\r\n[![arXiv](https://img.shields.io/badge/stat.ML-arXiv%3A2006.08437-B31B1B.svg)](https://arxiv.org/abs/2408.01005)\r\n[![Pytorch 1.5](https://img.shields.io/badge/pytorch-1.5.1-blue.svg)](https://pytorch.org/)\r\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://github.com/EagleAdelaide/FinSen_Dataset/LICENSE)\r\n\r\nThis repository contains the dataset for [*Enhancing Financial Market Predictions:\r\nCausality-Driven Feature Selection*](https://arxiv.org/abs/2408.01005), which has been accepted in ADMA 2024.\r\n\r\nIf the dataset or the paper has been useful in your research, please add a citation to our work:\r\n\r\n```\r\n@article{liang2024enhancing,\r\n  title={Enhancing Financial Market Predictions: Causality-Driven Feature Selection},\r\n  author={Liang, Wenhao and Li, Zhengyang and Chen, Weitong},\r\n  journal={arXiv e-prints},\r\n  pages={arXiv--2408},\r\n  year={2024}\r\n}\r\n```\r\n\r\n### Datasets\r\n\r\n**[FinSen]** can be downloaded manually from the repository as csv file. Sentiment and its score are generated by FinBert model from the Hugging Face Transformers library under the identifier \"ProsusAI/finbert\".  (Araci, Dogu. \"Finbert: Financial sentiment analysis with pre-trained language models.\" arXiv preprint arXiv:1908.10063 (2019).)\r\n\r\n**We only provide US for research purpose usage, please contact w.liang@adelaide.edu.au for other countries (total 197 included) if necessary.**\r\n\r\nWe also provide other NLP datasets for text classification tasks here, please cite them correspondingly once you used them in your research if any.\r\n\r\n1. 20Newsgroups. Joachims, T., et al.: A probabilistic analysis of the rocchio algorithm with tfidf for\r\ntext categorization. In: ICML. vol. 97, pp. 143–151. Citeseer (1997)\r\n2. AG News. Zhang, X., Zhao, J., LeCun, Y.: Character-level convolutional networks for text\r\nclassification. Advances in neural information processing systems 28 (2015)\r\n3. Financial PhraseBank. Malo, P., Sinha, A., Korhonen, P., Wallenius, J., Takala, P.: Good debt or bad debt:\r\nDetecting semantic orientations in economic texts. Journal of the Association for\r\nInformation Science and Technology 65(4), 782–796 (2014)\r\n\r\n### Dataloader for FinSen\r\n\r\nWe provide the preprocessing file finsen.py for our FinSen dataset under dataloaders directory for more convienient usage.\r\n\r\n### Models - Text Classification\r\n\r\n1. DAN-3. \r\n\r\n2. Gobal Pooling CNN.\r\n\r\n### Models - Regression Prediction\r\n\r\n1. LSTM\r\n\r\n### Using Sentiment Score from FinSen Predict Result on S&P500\r\n\r\n### Dependencies\r\n\r\nThe code is based on PyTorch under code frame of https://github.com/torrvision/focal_calibration, please cite their work if you found it is useful.\r\n\r\n:smiley: ☺ Happy Research !","description_withheld":null,"homepage":"","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":{"name":"MIT","url":null},"modalities":[],"tasks":[{"name":"Text Classification","url":"/task/text-classification","datasets_with_task":"/datasets/task/text-classification"},{"name":"Stock Market Prediction","url":"/task/stock-market-prediction","datasets_with_task":"/datasets/task/stock-market-prediction"},{"name":"Time Series Regression","url":"/task/time-series-regression","datasets_with_task":"/datasets/task/time-series-regression"}],"languages":[],"variants":[],"data_loaders":[{"repo":"https://github.com/EagleAdelaide/FinSen_Dataset","url":"https://github.com/EagleAdelaide/FinSen_Dataset","frameworks":["pytorch"]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/time-series-regression-on-finsen","task":"Time Series Regression","dataset_variant":"FinSen","rows":1,"metrics":["Mean MSE"],"first_row_in_archive_order":{"model":"LSTM","paper":"/paper/enhancing-financial-market-predictions","metrics":{"Mean MSE":"0.01"},"code_links":[{"title":"EagleAdelaide/FinSen_Dataset","url":"https://github.com/EagleAdelaide/FinSen_Dataset"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/enhancing-financial-market-predictions","title":"Enhancing Financial Market Predictions: Causality-Driven Feature Selection","date":"2024-08-02","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}