Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours
Course Outline
Foundations of Scaling Ollama
- Ollama’s architectural design and scaling factors
- Typical bottlenecks in multi-user setups
- Core practices for preparing infrastructure
Resource Management and GPU Enhancement
- Tactics for maximizing CPU and GPU efficiency
- Considerations regarding memory and bandwidth
- Setting resource limits at the container level
Deployment via Containers and Kubernetes
- Packaging Ollama using Docker
- Operating Ollama within Kubernetes clusters
- Managing load balancing and service discovery
Autoscaling and Batching Mechanisms
- Crafting autoscaling policies for Ollama
- Using batch inference to boost throughput
- Balancing latency against throughput
Latency Improvement
- Analyzing inference performance
- Implementing caching and model warm-up procedures
- Minimizing I/O and communication delays
Monitoring and Observability
- Connecting Prometheus for metric collection
- Creating dashboards using Grafana
- Setting up alerts and incident handling for Ollama infrastructure
Cost Control and Scaling Approaches
- Assigning GPUs with cost awareness
- Evaluating cloud versus on-premises deployment
- Methods for sustainable scaling
Recap and Future Actions
Requirements
- Hands-on experience in Linux system administration
- Conceptual understanding of containerization and orchestration
- Knowledge of deploying machine learning models
Target Audience
- DevOps Engineers
- ML Infrastructure Teams
- Site Reliability Engineers