← All open roles
Engineering · Unimatch Lab · AI infrastructure

Lead DevOps Engineer

Lead a small team that builds and runs the bare-metal GPU clusters behind our self-hosted LLMs. Hands-on: Terraform, GPU drivers, the pager, and the cost per GPU-hour.

Up to 12000 USDT totalEurope or UAE
Position summary

You lead a small team that builds and runs the bare-metal GPU clusters behind our self-hosted LLMs. This is a hands-on Lead DevOps role: you write Terraform, debug GPU driver and infrastructure issues, and carry the pager, while also assigning work, reviewing your team's changes, and unblocking them day to day. You report directly to the Head of Engineering. You shape the infra strategy for your area yourself and bring it to the Head of Engineering for sign-off — then you own execution: the build, the incidents, and the cost per GPU-hour.

Our compute runs primarily on our own hardware: racked GPU servers, our own network fabric, our own power and cooling planning. AWS and GCP exist as a hybrid layer for burst and failover, not the default environment. Think in terms of rack space and GPU supply lead times, not just API calls.

Cost is a standing responsibility, not a cleanup task. You own the unit economics of the clusters you run — cost per GPU-hour, cost per token, utilization — and the call between buying more hardware, tightening scheduling, or bursting to cloud.

The load is high. English is the working language, written and spoken.

Key responsibilities

Bare-metal and GPU infrastructure

  • Build, rack, and run on-prem GPU clusters as the primary compute platform: hardware provisioning, network fabric, power and cooling capacity
  • Own GPU infrastructure at a deep level — scheduling, utilization, memory allocation, failure modes, firmware and driver lifecycle
  • Keep hybrid cloud (AWS, GCP) wired in for burst and failover, and know when it's the right call instead of the fallback
  • Plan physical capacity two to three quarters out: rack space, power draw, network throughput, GPU procurement lead times

LLM as a production service

  • Operate and evolve self-hosted open-weight LLMs on the team's clusters: inference and fine-tuning pipelines, day-2 operations
  • Keep RAG and model-serving paths inside SLO under production traffic
  • Track latency, throughput, cost, and memory as one operating picture

Cost optimization and FinOps

  • Own the unit economics of the infrastructure you run: cost per GPU-hour, cost per token/request, cluster utilization
  • Build and maintain cost visibility — tagging, chargeback or showback, dashboards — across on-prem and cloud spend
  • Make and defend the capex-vs-opex call (buy more hardware vs. burst to cloud vs. optimize scheduling) with numbers
  • Report infrastructure budget status and cost trends to the Head of Engineering on a regular cadence; flag overruns before they land
  • Find and reclaim idle or underutilized GPU capacity as a standing practice

Delivery, GitOps, and team leadership

  • Lead a small team (2-4 engineers): assign work, review Terraform and Kubernetes changes, unblock people, and stay in the codebase yourself
  • Own Kubernetes, Terraform, and ArgoCD as the default change path for your clusters
  • Build GitLab CI and custom pipelines so changes are reviewable, repeatable, and reversible
  • Keep Linux, networking, and performance work at the depth the cluster actually needs

Reliability, observability, and security

  • Run Prometheus, Grafana, and OpenTelemetry against SLOs and error budgets for the systems you own
  • Sit in the on-call rotation; lead incident response and postmortems for your team's surface
  • Implement IAM, RBAC, and secrets management (Vault / KMS / SSM) for your clusters, inside the security model the Head of Engineering sets
Required qualifications
  • Hard requirement: direct, hands-on experience building and operating your own data center or owned hardware infrastructure — bare-metal server provisioning, racking, network fabric, and GPU infrastructure, with deep reasoning about compute and memory under load. Not only managed cloud; candidates without real bare-metal/on-prem ownership will not be considered
  • 6+ years in DevOps, Infrastructure, or Platform Engineering
  • Experience owning infrastructure cost directly: capacity planning tied to unit economics, cost tagging/visibility, or FinOps practice in a prior role
  • Production experience operating self-hosted open-weight LLMs on local GPU clusters: inference, fine-tuning, day-2 operations
  • Deep Linux (internals, networking, performance), Docker, and production-scale Kubernetes
  • Terraform as a default; ArgoCD or equivalent GitOps tooling in real production use
  • GitLab CI and custom pipelines you've designed and run in production
  • Working AWS or GCP experience for hybrid and burst scenarios
  • Prometheus, Grafana, OpenTelemetry; SLO / error budgets; incident response and postmortems
  • Experience leading or managing a team of engineers, while staying hands-on — this is a working lead, not a pure people-manager seat
  • English B2, written and spoken; Russian, written and spoken
Preferred qualifications
  • STEM background or equivalent practical experience
  • Hugging Face, LoRA / QLoRA fine-tuning in production
  • Vector databases in production paths: Qdrant, Pinecone, or Weaviate
  • Ansible or Pulumi in real use
  • Enterprise-level AWS or GCP
  • Specialization in AI/GPU-dense infrastructure specifically — high-density GPU racks, power and cooling planning built around AI/ML workloads, beyond generic bare-metal ops
Technology stack
  • Bare-metal servers, GPU infrastructure, on-prem data center operations
  • Self-hosted LLM inference and fine-tuning; Hugging Face, LoRA / QLoRA, RAG pipelines
  • Linux, Docker, Kubernetes (production-scale)
  • Terraform (required), Ansible / Pulumi
  • GitOps: ArgoCD (required)
  • CI/CD: GitLab CI, custom pipelines
  • Hybrid cloud: AWS and GCP (burst and failover, not primary)
  • Vector databases: Qdrant / Pinecone / Weaviate
  • Prometheus, Grafana, OpenTelemetry
  • Vault / KMS / SSM, IAM, RBAC
Success metrics
  • Cost per GPU-hour and cost per token are tracked and hold flat or trend down as workload grows
  • Cluster utilization stays above the agreed threshold; idle capacity gets caught and reclaimed inside an agreed window
  • Infrastructure budget reporting to the Head of Engineering is current and accurate — no cost surprises
  • Self-hosted LLM inference and fine-tuning run as production services with explicit latency, throughput, cost, and memory targets
  • Platform changes go through Terraform and ArgoCD; rollback is a practiced path, not a theory
  • SLOs and error budgets exist for your team's surface; incidents produce postmortems that change the system
  • Your team ships against its roadmap commitments, and the engineers you lead grow in scope