Md Tanvirul Alam

PhD Candidate at Rochester Institute of Technology

I am Tanvir, a PhD candidate at Rochester Institute of Technology working on reasoning in language and vision-language models, with a particular focus on reinforcement learning with verifiable rewards. My research studies when these systems genuinely acquire new reasoning ability versus when they rely on shortcuts, semantic priors, or reward-driven heuristics, spanning multimodal benchmarks, self-play, and structured real-world domains such as cyber threat intelligence. More broadly, I am interested in building robust AI systems that can reason reliably under distribution shift and in high-stakes settings. Looking ahead, I am especially interested in world models and in how language, vision, and action can be integrated into more grounded reasoning systems; if you would like to collaborate, feel free to email me.

Examples from Trace across charts, games, geometry, graphs, icons, illustrations, pages, physics, puzzles, symbolic reasoning, and 3D scenes

Featured Project

Trace

A Taxonomy-Guided Environment for Multidomain Visual Reasoning

Trace is a procedural environment for broad, reproducible visual-reasoning training and evaluation. It generates images, questions, answers, and answer checks from the same underlying task state.

  • 1,000 tasks
  • 277 scene grammars
  • 11 visual domains

RLVR training with 64K Trace examples improved Qwen2.5-VL at both model scales across 24 external benchmarks.

Model Base Trace RLVR Gain
3B 39.34 42.85 +3.51
7B 47.93 51.99 +4.06

Research Interests

Reasoning Reinforcement Learning with Verifiable Rewards Large Language Models Vision-Language Models Benchmarking and Evaluation Robustness and Generalization World Models Cyber Threat Intelligence

Connect

Email: tanvirul.alam.research [at] gmail [dot] com