{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improved-mixed-example-data-augmentation","title":"Improved Mixed-Example Data Augmentation","arxiv_id":"1805.11272","date":"2018-05-29","proceeding":null,"authors":["Cecilia Summers","Michael J. Dinneen"],"abstract":"In order to reduce overfitting, neural networks are typically trained with\ndata augmentation, the practice of artificially generating additional training\ndata via label-preserving transformations of existing training examples. While\nthese types of transformations make intuitive sense, recent work has\ndemonstrated that even non-label-preserving data augmentation can be\nsurprisingly effective, examining this type of data augmentation through linear\ncombinations of pairs of examples. Despite their effectiveness, little is known\nabout why such methods work. In this work, we aim to explore a new, more\ngeneralized form of this type of data augmentation in order to determine\nwhether such linearity is necessary. By considering this broader scope of\n\"mixed-example data augmentation\", we find a much larger space of practical\naugmentation techniques, including methods that improve upon previous\nstate-of-the-art. This generalization has benefits beyond the promise of\nimproved performance, revealing a number of types of mixed-example data\naugmentation that are radically different from those considered in prior work,\nwhich provides evidence that current theories for the effectiveness of such\nmethods are incomplete and suggests that any such theory must explain a much\nbroader phenomenon. Code is available at\nhttps://github.com/ceciliaresearch/MixedExample.","url_abs":"http://arxiv.org/abs/1805.11272v4","url_pdf":"http://arxiv.org/pdf/1805.11272v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improved-mixed-example-data-augmentation","repo_url":"https://github.com/ceciliaresearch/MixedExample","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"image-augmentation","task_name":"Image Augmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.11272","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1805.11272"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ceciliaresearch/MixedExample","reach":null}],"summary":{"ran_honours":1,"ran_violates":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"b3633dbf0dcc00bc","entry":"first_example","repo":"ceciliaresearch/MixedExample","repo_kind":"official","path":"mixed_example.py","file_url":"https://github.com/ceciliaresearch/MixedExample/blob/HEAD/mixed_example.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b3633dbf0dcc00bc"}},{"code_sha256_prefix":"f3652e9749db4eb0","entry":"should_zero_mean","repo":"ceciliaresearch/MixedExample","repo_kind":"official","path":"mixed_example.py","file_url":"https://github.com/ceciliaresearch/MixedExample/blob/HEAD/mixed_example.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f3652e9749db4eb0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}