Qwen3-VL 8B

Small AI Model Beats GPT-5.6 on Tax Forms but Stumbles on Dates

A small 8-billion-parameter model just beat GPT-5.6 Terra on tax forms, 21 out of 32 W-2s fully correct versus 7. That's the kind of result that makes you rethink what "small" means. But Qwen3-VL stumbled on dates,…

4 min readMachine Learning
Small AI Model Beats GPT-5.6 on Tax Forms but Stumbles on Dates
Qwen3-VL 8B on a laptop vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats[R]
From Machine Learning

I benchmarked Qwen3-VL 8B Instruct (Q4_K_M, Ollama, M5 24GB, ~30s/doc) against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra on:

- receipts: CORD (Indonesia) and SROIE (Malaysia), 30 each

Read the original at Machine Learning