Postgraduate degree in a quantitative discipline such as Computer Science, Data Science, Mathematics, Statistics, Engineering, or equivalent (Masters - essential; PhD - advantageous).
Cloud Certification: AWS Solutions Architect Professional, AWS Machine Learning Specialty, or Microsoft Azure AI Engineer Associate.
FinOps Certification: FinOps Foundation Certified Practitioner (FOCP) or equivalent AI cost governance credential.
Security Certification: Certified Cloud Security Professional (CCSP) or AWS Security Specialty.
IaC Certification: HashiCorp Terraform Associate or equivalent infrastructure-as-code credential.
Work Experience
5–8 years of progressive leadership experience in Cloud AI Platform Engineering, with production experience managing multi-cloud AI platform stacks across at least two of: AWS Bedrock/SageMaker, Databricks AI, Microsoft Azure AI Foundry, or Hugging Face enterprise deployments.
2–3 years experience in AI FinOps and Cost Governance: Demonstrated ownership of AI compute cost models and FinOps reporting in a multi-BU or multi-cloud environment, with evidence of cost optimisation outcomes.
2–3 years experience in AI Security Architecture: Designing and implementing zero-trust AI security (OAuth/OIDC, JWT), prompt injection controls, data residency compliance in a regulated environment.
2–3 years experience in Agentic AI Infrastructure: Production design of agent orchestration infrastructure (such as LangGraph, AutoGen, Foundry Agent Service, Bedrock Agents), tool-calling APIs, and agent state management.
2–3 years experience in Platform Observability: Operating AI-specific observability tooling for inference latency, drift alerting, and capacity management (such as Prometheus, Grafana, Datadog, or Lakehouse Monitoring).
2–3 years experience in Infrastructure-as-Code: Terraform, Pulumi, or equivalent for multi-cloud, multi-region AI infrastructure deployments; CI/CD pipeline design for platform components.
2–3 years experience in Regulated Industry: AI platform engineering in financial services or a similarly regulated sector with model risk governance and change management obligations.
Advantageous: People leadership — leading or mentoring a team of platform or infrastructure engineers in an agile delivery environment.
Advantageous: Pan-African Deployments — delivering AI platform services across multiple African jurisdictions with awareness of data localisation and cross-border data transfer requirements.
Knowledge and Skills
Multi-Cloud AI Platform Architecture: Expert design and operation of AWS Bedrock, Databricks AI, Azure AI Foundry, and Hugging Face in enterprise production environments across multiple business units and geographies.
Agentic AI Infrastructure: Practical production knowledge of agent orchestration frameworks (LangGraph, AutoGen, Foundry Agent Service, Bedrock Agents), tool-calling API design, agent memory architecture, and multi-agent coordination patterns.
AI FinOps and Cost Management: Chargeback and showback model design; DBU and token cost attribution; provisioned throughput versus on-demand optimisation; GPU cluster cost management; spend anomaly detection and FinOps dashboarding.
AI Security and Zero Trust: OAuth 2.0, OIDC, JWT/JWE/JWS; RBAC and ABAC for AI workloads; prompt injection prevention; data exfiltration controls at the Gateway layer; AI threat modelling and data residency compliance.
Infrastructure-as-Code: Terraform, Pulumi, or AWS CDK for multi-cloud AI infrastructure; CI/CD pipeline design for platform components; container orchestration using Docker, Kubernetes, and Helm.
Platform Observability: Prometheus, Grafana, Datadog, OpenTelemetry, and Databricks Lakehouse Monitoring; custom metric design for AI workload health including inference latency, token throughput, and model drift.
Cloud-Agnostic Model Serving: ONNX, BentoML, Triton Inference Server; containerised model deployment patterns for portability across AWS, Azure, and Databricks environments.
MLOps Tooling: Working knowledge of MLflow, Kubeflow, Airflow, and CI/CD for ML, sufficient to collaborate effectively with AI Solution Engineers on model deployment and lifecycle management.
GPU and HPC Architecture: On-demand GPU cluster management; spot instance strategies; high-performance compute cost optimisation for large-scale model training and fine-tuning workloads.
Enterprise Risk and Governance: Absa Enterprise Wide Risk Management Framework; Group Architecture standards; AI Responsible Use Policy; POPIA; country-specific data localisation requirements across Absa's ten operating countries.
Agile Delivery: Sprint planning, backlog management, and continuous delivery practices in a self-directed squad environment; experience removing delivery barriers in a fast-moving, multi-stakeholder context.
Minimum Education: Bachelor's Degree in Information Technology.
GK
This is a preview of the role
Sign in to your GoKazini account to see the company name, full job details, salary information, and how to apply.