Start Date
1-5-2026 12:00 PM
End Date
1-5-2026 1:00 PM
Description
This project investigates how input modality affects mathematical reasoning in vision-language models (VLMs). Using 100 samples from the GSM8K benchmark, we evaluate Qwen2-VL-2B-Instruct and LLaVA-v1.6-Mistral-7B across three conditions: text-only, rendered image, and a modality mismatch condition where image and text describe different problems. Both models show substantial accuracy drops under visual input — Qwen2 falls from 55% to 30% and LLaVA from 40% to 21%. In the mismatch condition, both models follow the text modality in over 75% of cases, revealing strong text dominance when modalities conflict. These findings suggest that rendering math problems as images meaningfully impairs reasoning in current VLMs, even when visual content is clean and legible.
Investigating the effect of Input Modality on Reasoning in VLMs using Zero-shot evaluation
This project investigates how input modality affects mathematical reasoning in vision-language models (VLMs). Using 100 samples from the GSM8K benchmark, we evaluate Qwen2-VL-2B-Instruct and LLaVA-v1.6-Mistral-7B across three conditions: text-only, rendered image, and a modality mismatch condition where image and text describe different problems. Both models show substantial accuracy drops under visual input — Qwen2 falls from 55% to 30% and LLaVA from 40% to 21%. In the mismatch condition, both models follow the text modality in over 75% of cases, revealing strong text dominance when modalities conflict. These findings suggest that rendering math problems as images meaningfully impairs reasoning in current VLMs, even when visual content is clean and legible.