Featured Projects

Evaluated the robustness of Vision State Space Models against 21 image transformations (~15K images per transformation). Benchmarked VMamba-Base against Swin-Transformer and ConvNext, demonstrating consistent outperformance. Fine-tuned U-Mamba for road segmentation achieving >80% Dice score with custom gradient-freezing.

Computer Vision State Space Models

Analyzed how domain-specific knowledge is encoded in multimodal models like LLaVA and InstructBLIP. Implemented Domain Activation Probability Entropy (DAPE) and Logit Lens probing to track hidden representation evolution, finding specialized internal processing pathways for document understanding tasks.

VLLMs Mechanistic Interpretability