
Tun Jian Tan
Principal AI Engineer · Embedded LLM
Core contributor of vLLM
ประวัติจะตามมา
- DFlash
- KV-cache compression
- KV offloading
- Prefix caching
- Chunked prefill
- Sleep mode
อาทิตย์ 6 กันยายน · เปิดประตู 13:15 · เซสชัน 14:00–19:00 · ชั้น 3
JAS Green Village Ramkhamhaeng · ชั้น 3 ·ดูเส้นทาง ↗
vLLM's core architecture and the features behind fast, efficient serving — then practical ways to get more from limited GPU resources, including when to reach for each technique and what it costs you.
Take a deep dive into the vLLM ecosystem to see how to optimize, scale, and extend your LLM serving infrastructure. Beyond essential production deployment strategies, we will unpack major platform updates, including the Rust frontend, KV connector, and reinforcement learning rollouts. Finally, learn the practical mechanics of integrating brand-new architectures, with Kimi K3 as an example.
เวลาทั้งหมดเป็นเวลาประเทศไทย (ICT, UTC+7)

Principal AI Engineer · Embedded LLM
Core contributor of vLLM
ประวัติจะตามมา

CTO & Co-Founder · Embedded LLM
A HKUST PhD with 14 years in applied AI, Pin Siang leads the team behind Embedded LLM's contributions to vLLM — working on inference optimisation across GPU platforms, parallelism strategy for large MoE and multimodal models, and private on-premise serving.

APAC General Manager · Tensormesh
Tuan leads APAC for Tensormesh, working on the open-source KV cache orchestration layer that helps teams cut inference cost by reusing context instead of recomputing it. His path into inference infrastructure runs through a decade of strategy and hands-on cloud-native systems work across Europe, APAC and Australia. He also co-founded Vietnam's official OpenInfra user group — 7,000+ members, backed by the OpenStack Foundation and CNCF.

Principal AI Evangelist · KBTG
Two decades in Data and AI, and a Microsoft MVP in Responsible AI. At KBTG, Komes leads AI upskilling programmes and designs innovation processes around high-impact use cases — his "Responsible AI at the Start" approach has become a benchmark for organisations embedding ethics from day one.


Specialist Solutions Architect, AI · APAC · Red Hat
Jing Wen helps enterprises across APAC move AI from experimentation to production, with a focus on model serving and inference optimisation. Her background spans full stack engineering and enterprise pre-sales, with hands-on vLLM and OpenShift AI in production-grade deployments across FSI, telco, healthcare and the public sector.


Specialist Lead of AI Solution and Platform · True Digital Group
ประวัติจะตามมา

Graduate Researcher · Tsinghua University
ประวัติจะตามมา

Senior Solutions Architect · NVIDIA
ประวัติจะตามมา