Agenda

Sunday 6 September · doors open 13:15 · sessions 14:00–19:00 · 3rd floor

JAS Green Village Ramkhamhaeng · 3rd floor ·Get directions ↗

  1. – 14:00
    Registration

    Registration & doors open

  2. – 14:20
    Opening

    Opening — introducing the vLLM roadmap in the Thailand ecosystem

    AGICAFETAGICAFET
  3. – 14:50
    Talk

    Introduction to vLLM: High-Performance LLM Serving & Maximizing Performance on Limited GPUs

    vLLM's core architecture and the features behind fast, efficient serving — then practical ways to get more from limited GPU resources, including when to reach for each technique and what it costs you.

    Tun Jian TanEmbedded LLM
  4. – 15:45
    Panel

    Panel — AI Adoption in Thailand: what actually works

    Dr Pin Siang TanEmbedded LLMTuan LuongTensormeshDr Komes ChandavimolKBTGYimModerator · Moderator
  5. – 16:00
    Talk

    Thai and global perspective: the connected point

    CKFastwork
  6. – 16:30
    Talk

    Inside vLLM: Production Best Practices, Model Integration and Road Map

    Take a deep dive into the vLLM ecosystem to see how to optimize, scale, and extend your LLM serving infrastructure. Beyond essential production deployment strategies, we will unpack major platform updates, including the Rust frontend, KV connector, and reinforcement learning rollouts. Finally, learn the practical mechanics of integrating brand-new architectures, with Kimi K3 as an example.

    Tiezhen WANGInferact
  7. – 17:00
    Talk

    Scaling Distributed Inference with llm-d

    Jing Wen NgRed Hat
  8. – 17:30
    Talk

    Supercharging vLLM with LMCache: KV-Cache Reuse in Practice

    Dai TranTensormesh
  9. – 17:45
    Break

    Break

  10. – 18:20
    Panel

    Panel — vLLM Developers in Thailand

    Sattaya SingkulTrue Digital GroupNatthanan BhukanTsinghua UniversityPongsasit ThongpramoonNVIDIA
  11. – 19:00
    Networking

    Lucky draw & networking

All times are Bangkok time (ICT, UTC+7).

Speakers

Tun Jian Tan

Tun Jian Tan

Principal AI Engineer · Embedded LLM

Core contributor of vLLM

Bio to follow.

  • DFlash
  • KV-cache compression
  • KV offloading
  • Prefix caching
  • Chunked prefill
  • Sleep mode
Dr Pin Siang Tan

Dr Pin Siang Tan

CTO & Co-Founder · Embedded LLM

A HKUST PhD with 14 years in applied AI, Pin Siang leads the team behind Embedded LLM's contributions to vLLM — working on inference optimisation across GPU platforms, parallelism strategy for large MoE and multimodal models, and private on-premise serving.

  • Inference optimisation
  • MoE parallelism
  • On-prem serving
Tuan Luong

Tuan Luong

APAC General Manager · Tensormesh

Tuan leads APAC for Tensormesh, working on the open-source KV cache orchestration layer that helps teams cut inference cost by reusing context instead of recomputing it. His path into inference infrastructure runs through a decade of strategy and hands-on cloud-native systems work across Europe, APAC and Australia. He also co-founded Vietnam's official OpenInfra user group — 7,000+ members, backed by the OpenStack Foundation and CNCF.

  • KV cache orchestration
  • Inference cost
  • OpenInfra community
Dr Komes Chandavimol

Dr Komes Chandavimol

Principal AI Evangelist · KBTG

Two decades in Data and AI, and a Microsoft MVP in Responsible AI. At KBTG, Komes leads AI upskilling programmes and designs innovation processes around high-impact use cases — his "Responsible AI at the Start" approach has become a benchmark for organisations embedding ethics from day one.

  • Responsible AI
  • Enterprise AI adoption
  • AI governance (AIGP)
Tiezhen WANG

Tiezhen WANG

Head of Ecosystem · Inferact

Bio to follow.

  • Rust frontend
  • KV connector
  • RL rollouts
  • Kimi K3
Jing Wen Ng

Jing Wen Ng

Specialist Solutions Architect, AI · APAC · Red Hat

Jing Wen helps enterprises across APAC move AI from experimentation to production, with a focus on model serving and inference optimisation. Her background spans full stack engineering and enterprise pre-sales, with hands-on vLLM and OpenShift AI in production-grade deployments across FSI, telco, healthcare and the public sector.

  • llm-d
  • Distributed inference
  • OpenShift AI
  • Enterprise adoption
Dai Tran

Dai Tran

Solution Architect · Tensormesh

Bio to follow.

  • LMCache
  • KV-cache reuse
  • Inference throughput
Sattaya Singkul

Sattaya Singkul

Specialist Lead of AI Solution and Platform · True Digital Group

Bio to follow.

  • LLM platform
  • AI solutions
  • Production serving
Natthanan Bhukan

Natthanan Bhukan

Graduate Researcher · Tsinghua University

Bio to follow.

  • LLM research
  • Inference systems
  • Open source
Pongsasit Thongpramoon

Pongsasit Thongpramoon

Senior Solutions Architect · NVIDIA

Bio to follow.

  • GPU inference
  • Accelerated computing
  • Deployment at scale

Stay in the loop

Occasional email about vLLM Bangkok Day and what the Thai vLLM community is building. No spam, unsubscribe any time.