Md Tanvirul Alam
PhD Candidate at Rochester Institute of Technology
I am Tanvir, a PhD candidate at Rochester Institute of Technology working on reasoning in language and vision-language models, with a particular focus on reinforcement learning with verifiable rewards. My research studies when these systems genuinely acquire new reasoning ability versus when they rely on shortcuts, semantic priors, or reward-driven heuristics, spanning multimodal benchmarks, self-play, and structured real-world domains such as cyber threat intelligence. More broadly, I am interested in building robust AI systems that can reason reliably under distribution shift and in high-stakes settings. Looking ahead, I am especially interested in world models and in how language, vision, and action can be integrated into more grounded reasoning systems; if you would like to collaborate, feel free to email me.
Featured Project
Trace
A Taxonomy-Guided Environment for Multidomain Visual Reasoning
Trace is a procedural environment for broad, reproducible visual-reasoning training and evaluation. It generates images, questions, answers, and answer checks from the same underlying task state.
- 1,000 tasks
- 277 scene grammars
- 11 visual domains
RLVR training with 64K Trace examples improved Qwen2.5-VL at both model scales across 24 external benchmarks.
| Model | Base | Trace RLVR | Gain |
|---|---|---|---|
| 3B | 39.34 | 42.85 | +3.51 |
| 7B | 47.93 | 51.99 | +4.06 |
Research Interests
Connect
GitHub Google Scholar LinkedIn Stack Overflow
Email: tanvirul.alam.research [at] gmail [dot] com