|
Tapas Kumar Dutta
I am a researcher specializing in Computer Vision, Multimodal, Large Language Model.
I did my Master at the University of Surrey, where I have been worked on my thesis at the SketchX Lab under the supervision of Subhadeep Koley and Professor Yi-Zhe Song.
I was a Data Scientist at National Health Service, SWLEOC, and collaborated with Bagci Lab and INRIA, STARS Team on various research projects.
Alongside, I have held research and engineering positions at LearnOpenCV, BitsCrunch, VIVEN, Malaviya National Institute of Technology Jaipur and Indian Institute of Technology, Hyderabad.
|
|
|
Research Interests
I specialize in Deep Learning, focusing on Medical Imaging, Foundation Models, and Generative AI. My work includes AI-driven diagnostics, X-ray analysis, sketch understanding, GAN-based augmentation, and AI applications in NFT fraud detection and valuation.
|
|
DentiAsk: A VQA Benchmark for Multimodal Reasoning in Panoramic Dental Radiographs
New!
D.Jha, T. K. Dutta, R Paudel, O Susladkar, A Ghimire, A Dhakal, S Adhikari, N Yadav, D Nayak,P Pawar, W Chen, G Reynolds
arxiv 2025
[PDF]
/
[BibTeX]
/
[arXiv]
/
[Code]
/
|
Reviewer
- Conference: MICCAI, ACL
- Journals: Array (Elsevier), Expert Systems With Applications(Elsevier), IEEE Conference on Artificial Intelligence, Franklin Open, Image and Vision Computing (Elsevier)
|
|