{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-perceptual-prediction-framework-for-self","title":"A Perceptual Prediction Framework for Self Supervised Event Segmentation","arxiv_id":"1811.04869","date":"2018-11-12","proceeding":"CVPR 2019 6","authors":["Sathyanarayanan N. Aakur","Sudeep Sarkar"],"abstract":"Temporal segmentation of long videos is an important problem, that has\nlargely been tackled through supervised learning, often requiring large amounts\nof annotated training data. In this paper, we tackle the problem of\nself-supervised temporal segmentation of long videos that alleviate the need\nfor any supervision. We introduce a self-supervised, predictive learning\nframework that draws inspiration from cognitive psychology to segment long,\nvisually complex videos into individual, stable segments that share the same\nsemantics. We also introduce a new adaptive learning paradigm that helps reduce\nthe effect of catastrophic forgetting in recurrent neural networks. Extensive\nexperiments on three publicly available datasets - Breakfast Actions, 50\nSalads, and INRIA Instructional Videos datasets show the efficacy of the\nproposed approach. We show that the proposed approach is able to outperform\nweakly-supervised and other unsupervised learning approaches by up to 24% and\nhave competitive performance compared to fully supervised approaches. We also\nshow that the proposed approach is able to learn highly discriminative features\nthat help improve action recognition when used in a representation learning\nparadigm.","url_abs":"http://arxiv.org/abs/1811.04869v3","url_pdf":"http://arxiv.org/pdf/1811.04869v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-perceptual-prediction-framework-for-self","repo_url":"https://github.com/CVPRUSFTampa/EventSegmentation","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"event-segmentation","task_name":"Event Segmentation"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"unsupervised-action-segmentation","task_name":"Unsupervised Action Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-action-segmentation-on-50-salads","task":"Unsupervised Action Segmentation","dataset":"50 Salads","model":"LSTM+AL","rank_in_archive_order":1,"of":3,"metrics":{"Acc":"60.6"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-action-segmentation-on-breakfast","task":"Unsupervised Action Segmentation","dataset":"Breakfast","model":"LSTM+AL","rank_in_archive_order":8,"of":8,"metrics":{"Acc":"42.9","mIoU":"46.9"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-action-segmentation-on-youtube","task":"Unsupervised Action Segmentation","dataset":"Youtube INRIA Instructional","model":"LSTM+AL","rank_in_archive_order":1,"of":8,"metrics":{"F1":"39.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.04869","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1811.04869"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/CVPRUSFTampa/EventSegmentation","reach":null}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"80659d51a96fc3a4","entry":"loadMiniBatch","repo":"CVPRUSFTampa/EventSegmentation","repo_kind":"official","path":"Zacks_VGG_RNN.py","file_url":"https://github.com/CVPRUSFTampa/EventSegmentation/blob/HEAD/Zacks_VGG_RNN.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"80659d51a96fc3a4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}