Trilok Padhi
PhD Student, Computer Science · Georgia State University
I am a fifth-year PhD student in Computer Science at Georgia State University, advised by Prof. Ugur Kursuncu in the SWAN AI Research Group. My research focuses on making multimodal models and agents safer in real-world settings by improving their grounding, interpretability, and trustworthiness. Large vision-language models can describe images and act in the world, but they can still be confidently wrong in unpredictable ways and may fail midway through long-horizon tasks. My work on multimodal models and agents runs along three broad directions:
- Reliability. These systems fail quietly. They give confident answers the evidence does not support, and they drift partway through a long task without signaling it. I work on reading an agent’s internal state to catch failure before a trajectory ends, and on tying stated confidence to evidence that can actually be checked.
- Knowledge. Much of what a model needs in order to reason is stated in neither the image nor the text. I infuse commonsense and structured knowledge into compact vision-language models (under 1B parameters), so small models reach the accuracy of far larger ones while staying easier to inspect and steer.
- Safety. Models behave differently under pressure than they do on benchmarks. I red-team vision-language models and agents against realistic, multi-turn adversarial behavior, including harassment, toxicity, and manipulation, and study how they hold up in sensitive settings such as mental health support.
During my PhD, I was fortunate enough to contribute to creating high-quality synthetic data and pretraining for Physical AI Initiatives as a 2x Applied Scientist intern at Data & AI Team, Siemens GenAI R&D in Seattle. I also worked as a Research Intern at SRI International (formerly the Stanford Research Institute) in Menlo Park with the Neuro-Symbolic Computing and Intelligence group funded by ARPA-H to design UQ methods for health. Before starting my PhD, I worked as a Data Science Researcher at Rakuten AI Labs in Bengaluru. I have also held research internships at Bosch Research and NVIDIA Research. I am always happy to discuss multimodal reasoning, agent interpretability, and AI safety; feel free to get in touch. I actively collaborate with researchers and students. If you are interested in working together, please reach out via email at tpadhi1@student.gsu.edu or on LinkedIn.
news
| Aug 31, 2026 | Cross-Modal Grounding for Calibrated Confidence in Vision Language Models has been accepted at EMNLP 2026 GroundLM Workshop (Conference on Empirical Methods in Natural Language Processing). See you in Budapest, Hungary |
|---|---|
| Aug 31, 2026 | Attending ACM AI Leadership Summit 2026 in Atlanta, Georgia from August 31 to September 2, 2026. Looking forward to connecting with fellow researchers and industry leaders in AI! |
| Aug 30, 2026 | Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks has been accepted at ICWSM 2027 (AAAI International Conference on Web and Social Media). See you in Edinburgh, Scotland |
| Jun 15, 2026 | From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents has been accepted at the Mechanistic Interpretability Workshop, ICML 2026 in Seoul, South Korea. |
| May 26, 2026 | Excited to share that our tutorial, “Knowledge-Infused Multimodal Learning,” has been accepted at ICWSM ’26, the 20th International AAAI Conference on Web and Social Media! I will be presenting at the University of Southern California, USC Information Sciences Institute, Los Angeles, alongside Agnik Saha and Professor Ugur Kursuncu. We will discuss some of our lab’s recent work on knowledge graph construction and how to design vision-language, knowledge-guided frameworks for more reliable and interpretable multimodal AI systems. Tutorial website |
| May 15, 2026 | Excited to be spending the summer in Seattle, Washington, for my second stint at Siemens Data & AI Lab. I’m looking forward to collaborating on Physical AI initiatives to help design and develop the next generation of AI systems. If you’re in the Seattle area, please feel free to reach out |
| Apr 01, 2026 | New preprint: From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents, on step-wise conformal probes for early failure detection in LLM agents. |
| Feb 01, 2026 | Playing Devil’s Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models has been accepted at ACM Transactions on Intelligent Systems and Technology (TIST). |
| Nov 27, 2025 | New preprint with collaborators at KAIST and Amazon: Co-Evolving Agents: Learning from Failures as Hard Negatives. |
| Oct 16, 2025 | New preprint: Echoes of Human Malice in Agents, a benchmark for multi-turn online harassment attacks on LLM agents. |
| May 15, 2025 | Started as an Applied Scientist Intern at Siemens GenAI R&D in Seattle, on the Data & AI Research team, working on multimodal industrial foundation models. |
| May 15, 2025 | Just KIDDIN’: Knowledge Infusion and Distillation for Detection of INdecent Memes has been accepted to ACL 2025 Findings (acceptance rate 19.1%). Grateful for the ACL Student Travel Award. |
| Dec 15, 2024 | Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success was presented at IEEE BigData 2024 (acceptance rate 18.7%), supported by the IEEE BigData Student Travel Award. |