Overview
We are looking for an experienced MLOps Engineer to build and operate the infrastructure that powers Firebird AI’s production machine learning and LLM workloads. In this role, you will own the deployment, reliability, scalability, and observability of GPU-based AI services across Kubernetes environments. You will work at the intersection of machine learning, infrastructure, and platform engineering, building the systems that enable ML and product teams to reliably deploy and operate models at scale. You will work closely with ML Engineering, Platform, and Infrastructure teams to improve GPU utilization, automate deployments, optimize inference infrastructure, and ensure production AI services remain reliable and performant.
- Build, operate, and continuously improve Kubernetes-based infrastructure for production GPU and machine learning workloads.
- Design and maintain scalable deployment patterns for LLM inference and other AI services.
- Own CI/CD pipelines, deployment automation, Helm charts, and infrastructure configuration for ML services.
- Design observability across model serving, gateway, infrastructure, and pipeline layers using Prometheus, Grafana, Alertmanager, and related tooling.
- Define and maintain production metrics, SLOs, dashboards, and alerting strategies.
- Optimize workload scheduling, GPU utilization, autoscaling, traffic routing, batching, and resource allocation.
- Build and operate stateful services supporting the ML platform, including Kafka, Redis, ClickHouse, and PostgreSQL.
- Implement reliable traffic management strategies, including load balancing, canary deployments, shadow traffic, and cache-aware routing.
- Improve container build and deployment performance, including image size, cold-start time, CUDA dependencies, and reproducibility.
- Participate in production on-call, lead incident response, perform root cause analysis, and implement long-term corrective actions.
- Collaborate with ML engineers on productionizing models and translating serving requirements into reliable infrastructure.
- Continuously evaluate emerging technologies across GPU infrastructure, Kubernetes, LLM serving, and ML platform engineering.
- 3+ years MLOps, ML platform, DevOps or SRE, with 2+ running GPU workloads in production.
- Kubernetes with GPU specifics: device plugins, node selectors and taints, pool shapes, resource requests and limits, autoscaling (HPA/KEDA) including on engine queue depth.
- Helm chart authoring and customize as the primary tooling; Terraform or equivalent a plus.
- CI/CD ownership with GitHub Actions on self-hosted runners; workflows that call idempotent scripts, static validation gates on every PR.
- Observability designed, not inherited: Prometheus, Grafana, Alertmanager; defining metrics, SLOs and alert rules for serving, gateway and pipeline tiers.
- Operating stateful services on Kubernetes: Kafka, Redis, ClickHouse, PostgreSQL, from single-node to a production topology with backups and a restore drill.
- Traffic routing and serving optimization — request routing across replicas, prefix/KV-cache-aware routing, load balancing, batching policy, autoscaling on queue depth, canary and shadow traffic splitting
- Docker depth: multi-stage builds, pinned upstream tags, image size, cold-start control, CUDA base images
- Production on-call track record: owned incidents, wrote postmortems, shipped the fixes.
- Strong shell and working Python.
- Bare-metal or on-prem GPU cluster bring-up: GPU operator, driver lifecycle, NVLink and InfiniBand, local registry and object storage.
- Multi-node inference orchestration and its networking.
- Cloud GPU infrastructure (Nebius, AWS, GCP or Azure): instance selection, quota and capacity management.
- Envoy Gateway or Envoy proxy and Gateway API: routes, policies, access-log telemetry.
- NVIDIA Dynamo operator: disaggregated prefill/decode, planner and worker configuration.
- AIPerf or GenAI-Perf for load generation in CI.
- Canary and shadow traffic splitting at the gateway.
- Secrets management, RBAC, NetworkPolicy, image signing and scanning, supply-chain security.
- Experiment and model tracking: MLflow, Kubeflow or Metaflow.
- LLM-aware request routing operated in production: at least one of Gateway API Inference Extension (InferencePool/EPP), llm-d, NVIDIA Dynamo.
- GPU cost attribution and rightsizing (FinOps).
- Argo CD or Flux, if GitOps reconciliation replaces manual deploys.
- MIG or time-slicing.
- Python
- Shell
- Kubernetes
- GPU
- Helm
- Terraform
- GitHub Actions
- Prometheus
- Grafana
- Alertmanager
- Kafka
- Redis
- ClickHouse
- PostgreSQL
- Docker
- CUDA
- Envoy
- MLflow
- Kubeflow
- Metaflow
- Argo CD
- Flux
✨ Our intelligent job search engine discovered this job and republished it for your convenience.
Please be aware that the job information may be incorrect or incomplete. The job announcement remains the property of its original publisher. To view the original job and its full details, please visit the job's URL on the owner’s page.
Please clearly mention that you have heard of this job opportunity on https://ijob.am.
