Portrait of Subhransu Sekhar Bhattacharjee

Introduction

Hi, I am Subhransu Sekhar Bhattacharjee (Bangla: শুভ্রাংশু-শেখর ভট্টাচার্য), nickname: Rudra (রুদ্র). I am an international research scholar at the ANU School of Computing since April 2023.

Education

I graduated with First Class Honours in Mechatronics at ANU, minoring in Electronic Communication Systems and Mathematics. My honours thesis, advised by Prof. Ian Petersen, was titled "Whiplash Inertial Gradient Descent Dynamics," and was later published as a journal paper J1. You can find the complete version here.

Research

I am a PhD candidate in the School of Computing at the Australian National University, working on autonomous spatio-semantic reasoning under uncertainty with Dr. Rahul Shome, Dr. Dylan Campbell, and Prof. Stephen Gould.

My thesis develops efficient generative world models of 3D indoor environments that turn partial, noisy observations into calibrated spatial–semantic priors for robotic planning and decision-making. Recent work appears at CVPR 2025 as Unobserved Object Detection.

My recent projects span CVPR 2025 unobserved object detection, ECCV 2026 generative floor-map completion with FlatLands, and MatterDoor for zero-shot spatio-semantic priors. Alongside the PhD, I have worked on applied AI, retrieval, and quantitative ML systems through Microsoft, Optiver, JP Morgan, and ANU industry-linked research.

My current research direction keeps computer vision at its core while extending into vision-language representation learning and token-efficient LLM/VLM architectures, with an emphasis on preserving strong spatial features under tighter memory, latency, and compute budgets.

My research interests include the following:

  1. Computer vision, 3D scene understanding, and generative world models: Learning uncertainty-aware geometry and semantics from partial observations for scene completion, spatial reasoning, and predictive world modelling.
  2. VLM representation learning and multimodal grounding: Learning dense and region-level visual features through cross-modal alignment, open-vocabulary perception, visual grounding, and language-conditioned spatial representations.
  3. Token-efficient LLM/VLM architectures and optimisation: Visual tokenisation, adaptive token pruning and merging, token compression, efficient attention, KV-cache optimisation, sparse or mixture-of-experts routing, and hardware-aware architecture–system co-design targeting inference latency, throughput, and memory efficiency.

If you are interested in collaborating with me, please reach out!

Hobbies

Outside research, I enjoy slow travel, theatre and long games of chess, and I am especially drawn to films with nonlinear narratives. I also love reading epics, modern classics and travelogues. I sketch and paint from time to time, experiment with cocktails for friends, look for unusual hiking routes, swim in the sea, and listen to audiobooks.