Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Tencent Hunyuan Production Fundamentals
- Overview of key scenarios for serving Tencent Hunyuan models.
- Characteristics of large and MoE models in production settings.
- Identification of common bottlenecks in latency, throughput, and cost.
- Defining service-level objectives (SLOs) for inference workloads.
Deployment Architecture and Serving Flow
- Core components of a production-grade inference stack.
- Evaluating containerized, on-premise, and cloud deployment models.
- Basics of model loading, request routing, and GPU allocation.
- Designing systems for reliability and operational simplicity.
Latency Optimization in Practice
- Leveraging optimized inference engines such as TensorRT where appropriate.
- Understanding KV-cache concepts and applying practical cache tuning.
- Minimizing startup, warmup, and response overhead.
- Measuring time to first token and token generation speed.
Throughput, Batching, and GPU Efficiency
- Strategies for continuous and request batching.
- Managing concurrency and queue behavior effectively.
- Enhancing GPU utilization while preserving user experience.
- Handling requests with long contexts and mixed workloads.
Quantization and Cost Control
- The importance of quantization in production serving.
- Practical trade-offs between FP16, INT8, and other precision options.
- Balancing model quality, latency, and infrastructure costs.
- Developing a simple checklist for cost optimization.
Operations, Monitoring, and Readiness Review
- Configuring autoscaling triggers for inference services.
- Monitoring key metrics including latency, throughput, cache usage, and GPU health.
- Implementing logging, alerting, and incident response protocols.
- Reviewing a reference deployment and formulating an improvement plan.
Requirements
- A foundational understanding of large language model deployment and inference workflows.
- Experience working with containers, cloud or on-premise infrastructure, and API-based services.
- Practical working knowledge of Python or system engineering tasks.
Audience
- ML engineers focused on deploying LLMs into production environments.
- Platform engineers responsible for managing GPU-based inference services.
- Solution architects designing scalable AI serving platforms.