Papers › Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression

Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression

20 Jul 2025archive 2025-07-28

Roy H. Jennings, Genady Paikin, Roy Shaul, Evgeny Soloveichik

Multimodal Large Language Models (MLLMs) show promise for image-based regression tasks, but current approaches face key limitations. Recent methods fine-tune MLLMs using preset output vocabularies and generic task-level prompts (e.g., "How would you rate this image?"), assuming this mimics human rating behavior. Our analysis reveals these approaches provide no benefit over image-only training. Models using preset vocabularies and generic prompts perform equivalently to image-only models, failing to leverage semantic understanding from textual input. We propose Regression via Transformer-Based Classification (RvTC), which replaces vocabulary-constrained classification with a flexible bin-based approach. Unlike approaches that address discretization errors through complex distributional modeling, RvTC eliminates manual vocabulary crafting through straightforward bin increase, achieving state-of-the-art performance on four image assessment datasets using only images. More importantly, we demonstrate that data-specific prompts dramatically improve performance. Unlike generic task descriptions, prompts containing semantic information about specific images enable MLLMs to leverage cross-modal understanding. On the AVA dataset, adding challenge titles to prompts improves correlations from 0.83 to 0.90, a new state-of-the-art. We demonstrate through empirical evidence from the AVA and AGIQA-3k datasets that MLLMs benefit from semantic prompt information surpassing mere statistical biases. This underscores the importance of incorporating meaningful textual context in multimodal regression tasks.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Aesthetics Quality AssessmentNo-Reference Image Quality Assessmentregression

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Aesthetics Quality Assessment Aesthetic Visual Analysis RvTC+ PLCC 0.901 #1 of 3 Archive leaderboard report
Aesthetics Quality Assessment Aesthetic Visual Analysis RvTC+ SRCC 0.899 #1 of 3 Archive leaderboard report
Aesthetics Quality Assessment Aesthetic Visual Analysis RvTC (image-only) PLCC 0.831 #2 of 3 Archive leaderboard report
Aesthetics Quality Assessment Aesthetic Visual Analysis RvTC (image-only) SRCC 0.833 #2 of 3 Archive leaderboard report
No-Reference Image Quality Assessment KADID-10k RvTC (image-only) PLCC 0.98 #1 of 9 Archive leaderboard report
No-Reference Image Quality Assessment KADID-10k RvTC (image-only) SRCC 0.98 #1 of 9 Archive leaderboard report
No-Reference Image Quality Assessment KonIQ-10k RvTC (image-only) PLCC 0.95 #1 of 1 Archive leaderboard report
No-Reference Image Quality Assessment KonIQ-10k RvTC (image-only) SRCC 0.94 #1 of 1 Archive leaderboard report
No-Reference Image Quality Assessment SPAQ RvTC (image-only) PLCC 0.93 #1 of 1 Archive leaderboard report
No-Reference Image Quality Assessment SPAQ RvTC (image-only) SRCC 0.93 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections