Get in Touch
 Duration 21 hours

Course Outline

Foundations of Scaling Ollama

  • Ollama’s architectural design and scaling factors
  • Typical bottlenecks in multi-user setups
  • Core practices for preparing infrastructure

Resource Management and GPU Enhancement

  • Tactics for maximizing CPU and GPU efficiency
  • Considerations regarding memory and bandwidth
  • Setting resource limits at the container level

Deployment via Containers and Kubernetes

  • Packaging Ollama using Docker
  • Operating Ollama within Kubernetes clusters
  • Managing load balancing and service discovery

Autoscaling and Batching Mechanisms

  • Crafting autoscaling policies for Ollama
  • Using batch inference to boost throughput
  • Balancing latency against throughput

Latency Improvement

  • Analyzing inference performance
  • Implementing caching and model warm-up procedures
  • Minimizing I/O and communication delays

Monitoring and Observability

  • Connecting Prometheus for metric collection
  • Creating dashboards using Grafana
  • Setting up alerts and incident handling for Ollama infrastructure

Cost Control and Scaling Approaches

  • Assigning GPUs with cost awareness
  • Evaluating cloud versus on-premises deployment
  • Methods for sustainable scaling

Recap and Future Actions

Requirements

  • Hands-on experience in Linux system administration
  • Conceptual understanding of containerization and orchestration
  • Knowledge of deploying machine learning models

Target Audience

  • DevOps Engineers
  • ML Infrastructure Teams
  • Site Reliability Engineers

Number of participants


Price per participant

Upcoming Courses

Related Categories