กำหนดการ

อาทิตย์ 6 กันยายน · เปิดประตู 13:15 · เซสชัน 14:00–19:00 · ชั้น 3

JAS Green Village Ramkhamhaeng · ชั้น 3 ·ดูเส้นทาง ↗

  1. – 14:00
    ลงทะเบียน

    ลงทะเบียน · เปิดประตู

  2. – 14:20
    เปิดงาน

    เปิดงาน — แนะนำ roadmap ของ vLLM ในระบบนิเวศประเทศไทย

    AGICAFETAGICAFET
  3. – 14:50
    บรรยาย

    Introduction to vLLM: High-Performance LLM Serving & Maximizing Performance on Limited GPUs

    vLLM's core architecture and the features behind fast, efficient serving — then practical ways to get more from limited GPU resources, including when to reach for each technique and what it costs you.

    Tun Jian TanEmbedded LLM
  4. – 15:45
    เสวนา

    เสวนา — AI Adoption in Thailand: อะไรที่ใช้ได้จริง

    Dr Pin Siang TanEmbedded LLMTuan LuongTensormeshDr Komes ChandavimolKBTGYimผู้ดำเนินรายการ · Moderator
  5. – 16:00
    บรรยาย

    มุมมองไทยและระดับโลก: จุดเชื่อมต่อ

    CKFastwork
  6. – 16:30
    บรรยาย

    Inside vLLM: Production Best Practices, Model Integration and Road Map

    Take a deep dive into the vLLM ecosystem to see how to optimize, scale, and extend your LLM serving infrastructure. Beyond essential production deployment strategies, we will unpack major platform updates, including the Rust frontend, KV connector, and reinforcement learning rollouts. Finally, learn the practical mechanics of integrating brand-new architectures, with Kimi K3 as an example.

    Tiezhen WANGInferact
  7. – 17:00
    บรรยาย

    Scaling Distributed Inference with llm-d

    Jing Wen NgRed Hat
  8. – 17:30
    บรรยาย

    Supercharging vLLM with LMCache: KV-Cache Reuse in Practice

    Dai TranTensormesh
  9. – 17:45
    พักเบรก

    พักเบรก

  10. – 18:20
    เสวนา

    เสวนา — นักพัฒนา vLLM ในประเทศไทย

    Sattaya SingkulTrue Digital GroupNatthanan BhukanTsinghua UniversityPongsasit ThongpramoonNVIDIA
  11. – 19:00
    เน็ตเวิร์กกิ้ง

    ลุ้นรางวัล · เน็ตเวิร์กกิ้ง

เวลาทั้งหมดเป็นเวลาประเทศไทย (ICT, UTC+7)

สปีกเกอร์

Tun Jian Tan

Tun Jian Tan

Principal AI Engineer · Embedded LLM

Core contributor of vLLM

ประวัติจะตามมา

  • DFlash
  • KV-cache compression
  • KV offloading
  • Prefix caching
  • Chunked prefill
  • Sleep mode
Dr Pin Siang Tan

Dr Pin Siang Tan

CTO & Co-Founder · Embedded LLM

A HKUST PhD with 14 years in applied AI, Pin Siang leads the team behind Embedded LLM's contributions to vLLM — working on inference optimisation across GPU platforms, parallelism strategy for large MoE and multimodal models, and private on-premise serving.

  • Inference optimisation
  • MoE parallelism
  • On-prem serving
Tuan Luong

Tuan Luong

APAC General Manager · Tensormesh

Tuan leads APAC for Tensormesh, working on the open-source KV cache orchestration layer that helps teams cut inference cost by reusing context instead of recomputing it. His path into inference infrastructure runs through a decade of strategy and hands-on cloud-native systems work across Europe, APAC and Australia. He also co-founded Vietnam's official OpenInfra user group — 7,000+ members, backed by the OpenStack Foundation and CNCF.

  • KV cache orchestration
  • Inference cost
  • OpenInfra community
Dr Komes Chandavimol

Dr Komes Chandavimol

Principal AI Evangelist · KBTG

Two decades in Data and AI, and a Microsoft MVP in Responsible AI. At KBTG, Komes leads AI upskilling programmes and designs innovation processes around high-impact use cases — his "Responsible AI at the Start" approach has become a benchmark for organisations embedding ethics from day one.

  • Responsible AI
  • Enterprise AI adoption
  • AI governance (AIGP)
Tiezhen WANG

Tiezhen WANG

Head of Ecosystem · Inferact

ประวัติจะตามมา

  • Rust frontend
  • KV connector
  • RL rollouts
  • Kimi K3
Jing Wen Ng

Jing Wen Ng

Specialist Solutions Architect, AI · APAC · Red Hat

Jing Wen helps enterprises across APAC move AI from experimentation to production, with a focus on model serving and inference optimisation. Her background spans full stack engineering and enterprise pre-sales, with hands-on vLLM and OpenShift AI in production-grade deployments across FSI, telco, healthcare and the public sector.

  • llm-d
  • Distributed inference
  • OpenShift AI
  • Enterprise adoption
Dai Tran

Dai Tran

Solution Architect · Tensormesh

ประวัติจะตามมา

  • LMCache
  • KV-cache reuse
  • Inference throughput
Sattaya Singkul

Sattaya Singkul

Specialist Lead of AI Solution and Platform · True Digital Group

ประวัติจะตามมา

  • LLM platform
  • AI solutions
  • Production serving
Natthanan Bhukan

Natthanan Bhukan

Graduate Researcher · Tsinghua University

ประวัติจะตามมา

  • LLM research
  • Inference systems
  • Open source
Pongsasit Thongpramoon

Pongsasit Thongpramoon

Senior Solutions Architect · NVIDIA

ประวัติจะตามมา

  • GPU inference
  • Accelerated computing
  • Deployment at scale

ติดตามข่าวสาร

อีเมลเป็นครั้งคราวเกี่ยวกับ vLLM Bangkok Day และสิ่งที่คอมมูนิตี้ vLLM ไทยกำลังสร้าง ไม่มีสแปม ยกเลิกได้ทุกเมื่อ