Bachelor's degree in Computer Science, Information Technology, Data Science, Mathematics, Statistics, Engineering, or a related quantitative discipline is essential.
A postgraduate qualification is advantageous.
Relevant practical experience may be considered where supported by a strong record of cloud and platform engineering delivery.
Equivalent Terraform, Pulumi, or cloud infrastructure certification (advantageous)
Advantageous Certifications – FinOps
FinOps Certified Practitioner (advantageous)
Equivalent cloud cost management or financial operations certification (advantageous)
Advantageous Certifications – Security
Certified Cloud Security Professional (advantageous)
AWS Certified Security (advantageous)
Microsoft Security, Compliance, and Identity certification (advantageous)
Equivalent cloud or cybersecurity certification (advantageous)
Work Experience
Approximately 4 to 6 years of relevant experience in cloud engineering, platform engineering, DevOps, MLOps, infrastructure engineering, or AI platform engineering.
At least 2 years of practical experience supporting cloud-based data, machine learning, generative AI, or AI platform workloads in a production environment.
Production experience with at least two of the following: AWS Bedrock or Amazon SageMaker, Databricks, Microsoft Azure AI Foundry or Azure Machine Learning, Hugging Face, Kubernetes-based model serving.
Practical infrastructure-as-code experience using Terraform, Pulumi, AWS CDK, or an equivalent technology.
Experience building or supporting CI/CD pipelines for cloud infrastructure, platform components, data services, or machine learning workloads.
Experience with Docker, Kubernetes, Helm, APIs, identity integration, and cloud-native platform services.
Experience implementing monitoring, dashboards, alerts, and operational support processes for production platforms.
Working knowledge of cloud cost management, cost allocation, capacity monitoring, and infrastructure optimisation.
Experience applying cloud security controls, identity and access management, secrets management, and secure API integration.
Experience working within enterprise risk, architecture, security, and change management processes.
Experience in financial services, telecommunications, healthcare, insurance, or another regulated industry is advantageous.
Knowledge and Skills
Multi-Cloud AI Platform Engineering: Practical knowledge of designing, deploying, and supporting AI services across AWS, Microsoft Azure, Databricks, Hugging Face, or Kubernetes-based environments.
Agentic AI Infrastructure: Working knowledge of agent orchestration frameworks, tool-calling API patterns, agent memory, state management, tracing, and multi-agent workflows.
AI FinOps and Cost Management: Knowledge of cloud consumption models, token-based pricing, Databricks DBUs, provisioned throughput, GPU utilisation, chargeback and showback reporting, and spend anomaly detection.
AI Security and Zero Trust: Working knowledge of OAuth 2.0, OIDC, JWT, RBAC, ABAC, API security, managed identities, secrets management, prompt injection controls, data loss prevention, and secure agent tool access.
Infrastructure-as-Code: Strong practical experience with Terraform, Pulumi, AWS CDK, or equivalent infrastructure automation technologies.
Containerisation and Orchestration: Experience with Docker, Kubernetes, Helm, container registries, workload scheduling, resource allocation, and production container operations.
Platform Observability: Experience with Prometheus, Grafana, Datadog, OpenTelemetry, cloud-native monitoring tools, or Databricks Lakehouse Monitoring.
Cloud-Agnostic Model Serving: Working knowledge of containerised model deployment and serving technologies such as ONNX, BentoML, Triton Inference Server, Kubernetes, or equivalent frameworks.
MLOps Tooling: Working knowledge of MLflow, Kubeflow, Airflow, model registries, feature stores, automated testing, and CI/CD for machine learning workloads.
GPU Infrastructure: Understanding of GPU workload deployment, capacity management, right-sizing, spot instance strategies, and cost optimisation for model training and inference.
Enterprise Risk and Governance: Working knowledge of information security, technology risk, architecture governance, responsible AI, privacy, data residency, and change management requirements within a regulated environment.
Agile Delivery: Experience working in agile engineering teams using sprint planning, backlog management, iterative delivery, peer review, testing, and continuous improvement practices.
GK
This is a preview of the role
Sign in to your GoKazini account to see the company name, full job details, salary information, and how to apply.