
InnovationTeam
Riyadh, Saudi ArabiaPosted 1 months ago
Looking for an AI Engineer to Develop, deploy, and operate AI/LLM models across Clinets dual environment — GCP for public-cloud workloads, Humain sovereign cloud for classified data. Benefits 5 years ML/AI engineering, in production LLM deployment with knowledge in Python, PyTorch, Hugging Face Kubernetes in production; GPU-served inference GCP Vertex AI or any equivellent cloud
Develop, deploy, and operate AI/LLM models across dual environments (GCP and Humain sovereign cloud). Manage end-to-end model lifecycle including serving stack ownership, versioning, CI/CD, monitoring for latency, token throughput, GPU utilization, and drift. Implement pre-deployment evaluation, safety checks, and accuracy baselines. Optimize inference performance (quantization, batching, context sizing) and manage resource allocation (GPU quotas, partitioning, RBAC). Implement and maintain routing between GCP Vertex AI/GKE and Humain Kubernetes deployments. Ensure compliance with data sovereignty and regulatory guidelines (ZATCA, SDAIA, PDPL) and adhere to AI ethics and GenAI guidelines.
Not sure you fit this role?
Upload your CV and see how you score against InnovationTeam and every other live job. It's free.
Get my free matchesAED 22k–35k a month· est.