Get in Touch
 Duration 14 hours

Course Outline

Tencent Hunyuan Production Fundamentals

  • Overview of key scenarios for serving Tencent Hunyuan models.
  • Characteristics of large and MoE models in production settings.
  • Identification of common bottlenecks in latency, throughput, and cost.
  • Defining service-level objectives (SLOs) for inference workloads.

Deployment Architecture and Serving Flow

  • Core components of a production-grade inference stack.
  • Evaluating containerized, on-premise, and cloud deployment models.
  • Basics of model loading, request routing, and GPU allocation.
  • Designing systems for reliability and operational simplicity.

Latency Optimization in Practice

  • Leveraging optimized inference engines such as TensorRT where appropriate.
  • Understanding KV-cache concepts and applying practical cache tuning.
  • Minimizing startup, warmup, and response overhead.
  • Measuring time to first token and token generation speed.

Throughput, Batching, and GPU Efficiency

  • Strategies for continuous and request batching.
  • Managing concurrency and queue behavior effectively.
  • Enhancing GPU utilization while preserving user experience.
  • Handling requests with long contexts and mixed workloads.

Quantization and Cost Control

  • The importance of quantization in production serving.
  • Practical trade-offs between FP16, INT8, and other precision options.
  • Balancing model quality, latency, and infrastructure costs.
  • Developing a simple checklist for cost optimization.

Operations, Monitoring, and Readiness Review

  • Configuring autoscaling triggers for inference services.
  • Monitoring key metrics including latency, throughput, cache usage, and GPU health.
  • Implementing logging, alerting, and incident response protocols.
  • Reviewing a reference deployment and formulating an improvement plan.

Requirements

  • A foundational understanding of large language model deployment and inference workflows.
  • Experience working with containers, cloud or on-premise infrastructure, and API-based services.
  • Practical working knowledge of Python or system engineering tasks.

Audience

  • ML engineers focused on deploying LLMs into production environments.
  • Platform engineers responsible for managing GPU-based inference services.
  • Solution architects designing scalable AI serving platforms.

Number of participants


Price per participant

Upcoming Courses

Related Categories