Multi-modal LLM Serving
Efficient and predictable serving for mixed multimodal LLM workloads.
I am a Postdoctoral Researcher at KAIST CASYS. My research focuses on systems and architectures for LLM inference and serving, with interests in multimodal workloads, KV-cache management, and processing-in-memory. I received my Ph.D. in Computer Science from KAIST, advised by Prof. Jaehyuk Huh.
Efficient and predictable serving for mixed multimodal LLM workloads.
Memory-efficient LLM serving through hierarchical KV-cache management.
System and architecture support for efficient PIM-based LLM serving.