{"url":"/dataset/novic-caption-object-data","name":"NOVIC Caption-Object Data","full_name":null,"description_markdown":"This corpus contains data files that were generated as part of the NOVIC paper (see above). This includes the complete [Object Noun Dictionary](https://github.com/pallgeuer/object_noun_dictionary), the exact templates used for the multiset prompt templating strategy, and a large dataset of 1.8M LLM-generated and templated captions assorted by target noun. The captions were generated based on all of the target nouns in the Object Noun Dictionary.\r\n\r\nThe data is directly available at the following links:\r\n\r\n* [Object Noun Dictionary (JSON)](https://www2.informatik.uni-hamburg.de/wtm/corpora/novic/object_noun_dictionary.json)\r\n* [Multiset prompt templates](https://www2.informatik.uni-hamburg.de/wtm/corpora/novic/multiset_prompt_templates.json)\r\n* [LLM-generated captions dataset](https://www2.informatik.uni-hamburg.de/wtm/corpora/novic/captions_dataset.json)\r\n\r\nRefer to the [NOVIC code](https://github.com/pallgeuer/novic) and [Object Noun Dictionary code](https://github.com/pallgeuer/object_noun_dictionary) for examples of how the data can be used, as well as regenerated.","description_withheld":null,"homepage":"https://www.inf.uni-hamburg.de/en/inst/ab/wtm/research/corpora.html#novic","introduced_date":"2024-07-15","introduced_date_note":null,"introduced_by":{"paper":"/paper/unconstrained-open-vocabulary-image","title":"Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion","first_author":"Philipp Allgeuer","url":null},"license":{"name":"CC BY-NC-SA 4.0","url":"https://creativecommons.org/licenses/by-sa/4.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Classification","url":"/task/text-classification","datasets_with_task":"/datasets/task/text-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["NOVIC Caption-Object Data"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}