{"url":"/dataset/coloninst-v1","name":"ColonINST-v1","full_name":null,"description_markdown":"ColonINST is a large-scale instruction tuning dataset designed for multimodal analysis in colonoscopy. This dataset comprises 62 categories, and 303,001 colonoscopy images, including 128,620 positive and 174,381 negative cases collected from 19 publicly available datasets. We enhanced 128,620 colonoscopy images with detailed captions using a pipeline that interacts with GPT-4V through custom prompts, enriching the dataset for AI model training. We finally restructured 450,724 visual dialogues to guide the AI model through four downstream tasks critical for multimodal medical AI applications.\r\n\r\nPlease cite our work if you like it!\r\n```text\r\n@article{ji2024frontiers\r\n  author = {Ji, Ge-Peng and Liu, Jingyi and Xu, Peng and Barnes, Nick and Khan, Fahad Shahbaz and Khan, Salman and Fan, Deng-Ping},\r\n  title = {Frontiers in Intelligent Colonoscopy},\r\n  journal = {arXiv preprint arXiv:2410.17241},\r\n  year = {2024}\r\n}\r\n```","description_withheld":null,"homepage":"https://github.com/ai4colonoscopy/IntelliScope","introduced_date":"2024-10-22","introduced_date_note":null,"introduced_by":{"paper":"/paper/frontiers-in-intelligent-colonoscopy","title":"Frontiers in Intelligent Colonoscopy","first_author":"Ge-Peng Ji","url":null},"license":null,"modalities":[],"tasks":[{"name":"Image Classification","url":"/task/image-classification","datasets_with_task":"/datasets/task/image-classification"},{"name":"Referring Expression Comprehension","url":"/task/referring-expression-comprehension","datasets_with_task":"/datasets/task/referring-expression-comprehension"},{"name":"Referring expression generation","url":"/task/referring-expression-generation","datasets_with_task":"/datasets/task/referring-expression-generation"}],"languages":[],"variants":["ColonINST-v1","ColonINST-v1 (Seen)","ColonINST-v1 (Unseen)"],"data_loaders":[],"num_papers_in_archive":8,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}