Welcome to AI4GC Lab at Zhejiang University, led by Shengyu Zhang. We build efficient, deployable multimodal AI—compact multimodal LLMs and computer-use agents, accelerated image and video generation, and the device–cloud systems that carry them out of the datacenter and onto real phones and computers. Today's foundation models are remarkably capable yet costly in computation, memory, and energy, which keeps much of their power locked inside the cloud.
Our research closes that gap from both ends. We fine-tune small multimodal models and GUI agents to reason and act reliably under tight budgets; we attack the redundancy inside generation and the KV cache to make inference faster without sacrificing quality; and we design large–small collaboration so that powerful cloud models and constrained on-device models share intelligence and adapt to each user in real time. Alongside this, we take agent safety and evaluation seriously, stress-testing systems before they are trusted in the real world. We publish at leading venues including CVPR, ICML, ICLR, and ACL, and work with industry partners to put these ideas into production—aiming to make multimodal AI faster, more trustworthy, and accessible everywhere it is needed.