Can vision-language models turn free traffic-camera images into vehicle counts?
Testing whether off-the-shelf vision-language models can turn free public traffic-camera imagery into usable vehicle counts - zero-shot, no fine-tuning.
Zero-shot evaluation, no fine-tuning
Entirely public traffic-camera feeds
Released benchmark is the primary deliverable