arxiv:2411.04923
Shehan Munasinghe
shehan97
AI & ML interests
Computer Vision, Multi-modal learning
Recent Activity
upvoted a paper about 5 hours ago
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding upvoted a paper 5 months ago
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework authored a paper 10 months ago
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in
Videos