Dawei Wang

Dawei Wang

PhD Student

Newcastle University

Research Interests

Large Multimodal Models
Reinforcement Learning
Computer Graphics
Video Game Development

About

I am a PhD student at the School of Computing, Newcastle University, advised by Dr. Rich Davison and Dr. Gary Ushaw.

Prior to this, I studied in the MSc Computer Game Engineering programme at Newcastle University, graduating with Distinction. I also worked as a Game Engine Developer at Tencent Games and as a Game Developer at Shengqu Games.

My current research focuses on Large Multimodal Models and Reinforcement Learning.

Selected Publications

View All

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation

Dawei Wang, Di Zhao, Xinyuan Liu, Marci Chi Ma, Xiaoyang Liu, Chengming Zhou, Gary Ushaw, Richard Davison

Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

MARS-RA improves credit assignment in cooperative MARL by using large multimodal models to rank agents through pairwise contribution comparisons, converting these rankings into robust reward-shaping signals for effective cooperation.

Tracing the Light of Thought: A Probabilistic Self- and Cross-Consistency Verification Mechanism Improving Mathematical Reasoning in LLMs

Xiaoyang Liu, Dawei Wang, Tian Li, Huizhi Liang, Gary Ushaw, Richard Davison

Findings of the Association for Computational Linguistics: ACL 2026

We explored bridging ray tracing and LLM reasoning by proposing inference algorithms that sample reasoning trajectories directly based on LLM confidence at multiple granularities.

Towermind: A tower defence game learning environment and benchmark for llm as agents

Dawei Wang, Chengming Zhou, Di Zhao, Xinyuan Liu, Marci Chi Ma, Gary Ushaw, Richard Davison

Proceedings of the AAAI Conference on Artificial Intelligence

We introduce TowerMind, a tower defense game benchmark for evaluating long-term planning and decision-making in agentic LLMs, characterized by low evaluation costs, multimodal inputs, and reinforcement learning compatibility.