MLOps + DevOps Engineer - Agentic AI & Platform

  • Full-Time
  • On-Site

Job Description:

Role Overview:

We are building an AI-native platform where agentic systems, backend services, and real workflows operate together. This role sits at the intersection of ML Systems, Backend infrastructure and Distributed system operations. You will be responsible for ensuring that models, agents, APIs and workflows run reliably, scale predictably and remain observable end to end. This is not a traditional ML/DevOps role, but it is about operating intelligent systems in production.

What You'll Work On:

1. Design and implement cloud architectures supporting AI/ML workloads and production-grade systems

2. Build and manage ML platforms using AWS services including EC2, ECS/EKS, Lambda, S3, RDS, VPC, and IAM

3. Leverage AWS Bedrock, SageMaker, or similar managed AI services for model training and deployment

4. Use Infrastructure-as-Code tools such as Terraform, CloudFormation, or CDK to automate cloud provisioning

5. Work with vector databases (Milvus, Pinecone, Weaviate) and graph databases (Neo4j) to support retrieval-based and knowledge-driven AI solutions

6. AI/Model infrastructure:

6.1.Deploy and manage Small and Medium Language Models (SLMs)

6.2.Manage external LLM integrations

6.3. Build pipeline for model versioning, evaluation and fine-tuning

6.4. Support RAG systems, embeddings and Vector database infra

7. Agentic System Runtime

7.1.Enable execution of multi-agent workflows

7.2.Enable orchestration layers

7.3.Ensure consistency, fault-tolerance and latency control

8. Design and operate event-driven backend infrastructure – Apache Kafka (or equivalent)

9. Handle async workflows, retries, ordering, idempotency and enable reliable communication between backend and AI

10.Own Kubernetes cluster design, scaling strategies and workload isolation

  1. Support micro-services and model-serving workloads
  2. Design and manage API Gateway between Frontend and Backend layers
  3. Implement routing, auth, rate limiting and service protection at Edge layers
  4. Build end to end CI/CD (beyond basic GitHub Actions) and support multi-service deployment, environment isolation, rollback strategies, secret/vault management, blue/green deployment, feature flag support
  5. Designing, managing and handling data storage infra layer for RDBMS, Vector DB, Document DB, Caching, Object storage.
  6. Instrumentation, Observability and Debugging

    1. You will design and own end to end observability across AI + Backend + Frontend + Infra
    2. Implement structures instrumentation across APIs, async workflows, AI agents and capture request lifecycle, agent decision paths and execution timelines
    3. Design centralised logging – structured logs (JSON) and contextual logging (correlation IDs)
    4. Distributed tracing across tiers and service layers – user journey, agent decisions and backend actions
    5. Build strong monitoring with metrics for system health, API performance, LLM token burns, Queue lag and model latency
  7. Owning and commanding triage & debugging with engineering teams for multiservice failures, AI <> Backend inconsistencies, root cause analysis and replaying failures
  8. You will be managing security and secrets by securing APIs, model configs, infra credentials
  9. You will be implementing RBAC, secret rotation and environment isolation

Must-have:

  1. Expertise in AWS cloud services, EC2, S3, Lambda, ECS/EKS, managed AI platforms, model deployment and distributed systems
  2. Strong hands-on experience with Docker, Kubernetes at Production grade
  3. Experience with Apache Kafka or similar event-driven
  4. Networking protocols, Security concepts, VPC, Load balancers
  5. Strong experience in distributed systems and reliability engineering
  6. Hands-on with CI/CD pipeline - Jenkins, GitLab CI, GitHub Actions, automated testing
  7. Experience with Observability stack (logs, metrices, tracing) – Grafana, Prometheus, ELK, Sentry, Phoenix, Arize, Open Telemetry.
  8. Strong ownership, communication, and code review skills.
  9. You're based in Saudi Arabia and holding transferable Iqama.