MLOps / AI Operations Engineer

PwC · Bucharest, Romania · 2 days ago
4+ yrs mentionedad in EnglishData & AIvia workday
Apply
Job Description & Summary The opportunity Industrialize AI delivery through automated deployment, evaluation operations, observability, reliability engineering and transparent consumption management. What you will be doing ·        Build CI/CD pipelines for AI services, prompts, agent configurations, infrastructure and evaluation assets. ·        Automate environment provisioning, testing, deployment, rollback and release evidence. ·        Implement tracing, logging, model and agent monitoring, alerts and operational dashboards. ·        Operationalize evaluation thresholds, incident handling and continuous-improvement loops. ·        Monitor latency, capacity, token usage, infrastructure consumption and cost drivers. ·        Define runbooks, service ownership and production support handover. What we need from you ·        4+ years in DevOps, platform engineering, ML engineering, SRE or cloud operations. ·        Strong automation, containers, cloud services, observability and Infrastructure as Code capability. ·        Experience deploying or operating ML, generative AI or distributed application workloads. ·        Understanding of release controls, reliability, security and cost optimization. Relevant AI technologies and tooling ·        Hands-on experience with GitHub Actions, Azure DevOps, GitLab CI or equivalent, plus Infrastructure as Code using Terraform, Bicep or comparable tooling. ·        Strong container and orchestration capability using Docker and Kubernetes, together with experience deploying AI or agent services across cloud and hybrid environments. ·        Experience operating model and prompt assets, agent configurations, evaluation datasets and release evidence using MLflow, platform-native registries or equivalent lifecycle tooling. ·        Practical implementation of agent tracing and observability using OpenTelemetry and tools such as LangSmith, MLflow, Langfuse, Azure Monitor, Prometheus or Grafana. ·        Ability to monitor model and agent quality, tool failures, retrieval performance, latency, token usage, cost, capacity and workflow-level service indicators. ·        Experience with progressive delivery, rollback, secrets management, vulnerability scanning, incident response and reliability practices for non-deterministic AI systems. Measures of success ·        Deployment frequency and success rate ·        Mean time to detect and restore ·        Evaluation and monitoring coverage ·        Service reliability and latency ·        Cost and consumption transparency Key interfaces ·        Other members of the AI Transformation & Agentic Systems Practice ·        PwC sector, functional, cloud, cyber, risk, Responsible AI and change specialists ·        Client business owners, product owners, technology teams and operational users ·        Technology alliance and implementation partners where relevant Contribution to the practice ·        Support proposals, client workshops and market development appropriate to seniority. ·        Contribute reusable methods, patterns, code, assets and lessons learned. ·        Coach colleagues and participate in the capability’s continuous learning agenda. ·        Uphold PwC quality, independence, confidentiality and risk-management requirements. #LI-BS1 #LI-Hybrid

Get the passFilters for ad language, visa sponsorship and level, translation of any ad, from €5 — cancel anytime.