HCLTech
Dubai, UAEPosted today
Job Description Senior Technical Lead Others, Abu Dhabi Job Summary Job Title: Senior Engineer – Kubernetes Administration, EDGE AI CoE – T&I Cluster Job Purpose Design, build, secure, and operate on‑premises Kubernetes platforms hosting AI/ML workloads across the group. Own end‑to‑end cluster lifecycle (provisioning through decommission), GPU enablement, performance, reliability, and cost efficiency to support training, inference, data pipelines, and MLOps at scale, assuming on-prem deployments without managed cloud services. Key Responsibilities Cluster Architecture & Build (On‑Prem) Design and provision production‑grade clusters on bare‑metal and/or virtualized environments (e.g., kubeadm‑based; OpenShift, Rancher or upstream distributions), including HA control planes, etcd health, and node pools. Implement secure bootstrapping in connected and air‑gapped environments (private registries, image mirrors, OS hardening, CIS benchmarks). Configure networking: CNI (e.g., Cilium/Calico), ingress (NGINX/HAProxy), internal/external load‑balancing (MetalLB/BGP/ECMP), DNS/CoreDNS, IPv4/IPv6, and, where needed, Multus for secondary interfaces. Lifecycle Operations Plan and execute upgrades, patching, and capacity management; manage taints/affinities, quotas, and multi‑tenancy isolation. Implement autoscaling/right‑sizing (HPA/VPA; cluster autoscaling via on‑prem integrations) and node lifecycle automation. Maintain DR strategy, backups, and restores (e.g., Velero/restic), with documented runbooks and regular tests. AI/ML Workload Enablement Enable and manage specific workloads within each Kubernetes deployment that needs access to GPUs. Operate model serving (e.g., KServe/Seldon), batch/streaming pipelines, and ML orchestration (Kubeflow/Argo Workflows); optimize for throughput/latency and GPU utilization. Storage & Data Services Manage CSI drivers and dynamic provisioning; integrate with Ceph/Longhorn/Rook/NetApp/NFS as applicable; support stateful workloads with proper SLAs. Security & Compliance Enforce RBAC, Pod Security, NetworkPolicies, image signing/scanning (e.g., cosign/Trivy), admission controls/policy (OPA Gatekeeper/Kyverno), secrets management (Vault/Sealed Secrets/KMS). Integrate enterprise identity (LDAP/AD/Keycloak OIDC), certificate management /PKI, and audit logging; maintain evidence for audits. Platform Engineering & GitOps Standardize delivery with Helm/Kustomize; implement GitOps (Argo CD/Flux) and IaC (Terraform/Ansible) for repeatable environments. Operate on‑prem CI/CD runners (e.g., GitLab Runners/K8s executors) and artifact registries (Harbor) including air‑gapped workflows. Observability & SRE Implement metrics, logs, tracing (Prometheus/Grafana, EFK/ELK/Loki, OpenTelemetry); define SLIs/SLOs and capacity/health dashboards. Lead incident response, root cause analysis, and continuous improvement; participate in an on‑call rotation. Education & Experience Bachelor’s in Computer Science, Engineering, or related field (or equivalent experience). 7–10 years in platform/SRE/DevOps/systems roles, including 4+ years administering production Kubernetes. Proven experience designing, building, and operating on‑prem Kubernetes clusters (kubeadm/upstream; Rancher, OpenShift acceptable) in data‑center environments—no reliance on managed services. Hands‑on with GPUs for AI workloads, air‑gapped operations, private registries, and enterprise identity/security integrations. Certifications: CKA required (or obtained within 6 months); CKS/CKAD preferred; relevant Linux/networking certifications a plus. Key Skills Kubernetes core: control plane ops, etcd care, scheduling, admission controllers, multi‑tenancy (namespaces, quotas), PDBs, affinities, taints/tolerations. Linux/Containers: containerd/CRI‑O, systemd, kernel/GPU drivers, filesystems, OS hardening. Networking: Cilium/Calico, NetworkPolicies, ingress/egress; optional service mesh (Istio/Linkerd). Storage: CSI, PV/PVC, snapshots, performance tuning, backup/restore, DR patterns; Ceph/Rook/NetApp/NFS familiarity. Key Responsibilities Security: RBAC, Pod Security, image scanning/signing, OPA/Kyverno, Vault/Sealed Secrets, OIDC/LDAP/AD, audit and compliance practices. Observability: Prometheus/Grafana, EFK/ELK/Loki, OpenTelemetry; capacity/performance tuning. Delivery/IaC: Helm/Kustomize, Argo CD/Flux, Terraform/Ansible; GitLab CI/GitHub Actions; scripting in Bash/Python; strong YAML hygiene. Soft skills: ownership, clear documentation, incident communication, stakeholder management, mentoring. Education & Experience Bachelor’s in Computer Science, Engineering, Data/AI, or related field; Master’s preferred. 8+ years in software/ML engineering with 3+ years building GenAI/LLM solutions and production services. Hands-on expertise with on-prem/air-gapped deployments, Kubernetes, service meshes, and secure networking. Demonstrated experience with data engineering at scale: streaming/batch pipelines, metadata/lineage, governance. Practical LLMOps/MLOps exper
Bachelor’s in Computer Science, Engineering, or related field (or equivalent experience). 7–10 years in platform/SRE/DevOps/systems roles, including 4+ years administering production Kubernetes. Certifications: CKA required (or obtained within 6 months); CKS/CKAD preferred; relevant Linux/networking certifications a plus. 8+ years in software/ML engineering with 3+ years building GenAI/LLM solutions and production services. Hands-on expertise with on-prem/air-gapped deployments, Kubernetes, service meshes, and secure networking. Demonstrated experience with data engineering at scale: streaming/batch pipelines, metadata/lineage, governance. Practical LLMOps/MLOps experience.
Design, build, secure, and operate on-prem Kubernetes platforms hosting AI/ML workloads; own end-to-end cluster lifecycle (provisioning to decommission), GPU enablement, performance, reliability, and cost efficiency. Design and provision production-grade clusters (bare-metal/virtualized) with HA control planes; implement secure bootstrapping and private registries. Configure networking (CNI, ingress, load balancing, DNS, IPv4/IPv6, Multus). Plan and execute upgrades, patching, capacity management; manage taints/affinities, quotas, multi-tenancy. Implement autoscaling and node lifecycle automation. Maintain DR, backups, restores with runbooks and tests. Enable and manage AI/ML workloads, model serving, pipelines, and ML orchestration (Kubeflow/Argo). Manage CSI drivers and storage integrations; support stateful workloads with SLAs. Enforce security/compliance controls (RBAC, Pod Security, NetworkPolicies, image signing, OPA/Kyverno, secrets management). Integrate identity and PKI, audit logging. Standardize delivery with Helm/Kustomize; implement GitOps (Argo CD/Flux) and IaC (Terraform/Ansible). Operate on-prem CI/CD runners and artifact registries. Implement observability tooling and define SLIs/SLOs; lead incident response and on-call rotation.
Not sure you fit this role?
Upload your CV and see how you score against HCLTech and every other live job. It's free.
Get my free matchesAED 10k–18k a month· est.