{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/depression-scale-recognition-from-audio","title":"Depression Scale Recognition from Audio, Visual and Text Analysis","arxiv_id":"1709.05865","date":"2017-09-18","proceeding":null,"authors":["Shubham Dham","Anirudh Sharma","Abhinav Dhall"],"abstract":"Depression is a major mental health disorder that is rapidly affecting lives\nworldwide. Depression not only impacts emotional but also physical and\npsychological state of the person. Its symptoms include lack of interest in\ndaily activities, feeling low, anxiety, frustration, loss of weight and even\nfeeling of self-hatred. This report describes work done by us for Audio Visual\nEmotion Challenge (AVEC) 2017 during our second year BTech summer internship.\nWith the increase in demand to detect depression automatically with the help of\nmachine learning algorithms, we present our multimodal feature extraction and\ndecision level fusion approach for the same. Features are extracted by\nprocessing on the provided Distress Analysis Interview Corpus-Wizard of Oz\n(DAIC-WOZ) database. Gaussian Mixture Model (GMM) clustering and Fisher vector\napproach were applied on the visual data; statistical descriptors on gaze,\npose; low level audio features and head pose and text features were also\nextracted. Classification is done on fused as well as independent features\nusing Support Vector Machine (SVM) and neural networks. The results obtained\nwere able to cross the provided baseline on validation data set by 17% on audio\nfeatures and 24.5% on video features.","url_abs":"http://arxiv.org/abs/1709.05865v1","url_pdf":"http://arxiv.org/pdf/1709.05865v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"depression-scale-recognition-from-audio","repo_url":"https://github.com/serenera/Serenera","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"depression-scale-recognition-from-audio","repo_url":"https://github.com/yijiazh/DFER_Summer2019","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}