Location: 

(

Philadelphia

,

PA

)

Salary: 

$

165k

 - $

235k

Our client, a premier portfolio of globally recognized lifestyle brands, is seeking a Senior DevOps Engineer who will own and evolve the platform layer that keeps their ecommerce sites, order management systems, and integration pipelines running at scale—while helping the organization take its next leap forward with enterprise AI and agentic automation.

This role spans the full platform spectrum, from ensuring high availability of customer-facing commerce systems to building the infrastructure that lets AI agents and internal tools operate safely and reliably. The individual will partner with engineering teams across brands and functions, translating ambiguous problems into stable, scalable systems that power both current operations and future innovation.

The ideal candidate is a seasoned platform engineer with deep expertise in cloud-native infrastructure, Kubernetes, and CI/CD, combined with a growing interest or experience in AI/ML infrastructure. They thrive on bridging the gap between traditional commerce reliability and cutting-edge AI enablement, building systems that are secure, observable, and cost-efficient.

Role Responsibilities

Ecommerce & Commerce Platform Reliability

  1. Support and evolve the infrastructure behind customer-facing ecommerce sites, ensuring high availability, performance, and resilience during peak traffic events including product launches and promotional campaigns.

OMS & Integration Systems

  1. Operate and improve the infrastructure supporting the Order Management System (OMS) and the integration layer connecting commerce, fulfillment, ERP, and third-party services. Build for reliability, observability, and graceful failure.

CI/CD & Deployment Pipelines

  1. Design and maintain deployment pipelines supporting Java, Node, Python, and other languages, with strong standards around testing, rollback procedures, and release safety.

Internal Developer Platform

  1. Build the tooling and abstractions that enable engineering teams across brands to self-serve infrastructure, standardize deployments, and ship faster without sacrificing reliability or security.

Agentic AI Platform Infrastructure

  1. Architect and operate the platform layer for enterprise AI initiatives including LLM orchestration, MCP server deployments, and agent execution runtimes, with careful attention to security, cost efficiency, and performance.

AI Enablement for Engineering Teams

  1. Partner with product and data teams to build infrastructure that makes AI capabilities accessible across the organization, including RAG pipelines, vector search, embedding services, and API gateways for AI tools.

Observability & Incident Response

  1. Own the observability stack across commerce and AI systems, including distributed tracing, metrics, alerting, and on-call support, enabling teams to detect and resolve issues quickly at any layer of the stack.

Security & Compliance Posture

  1. Implement and maintain controls around access governance, secrets management, PII handling, and audit logging across both commerce-critical and AI workloads, ensuring compliance with regulatory requirements.

Cost & Capacity Management

  1. Monitor and optimize cloud spend across ecommerce infrastructure, integration services, and AI inference, leveraging autoscaling, rightsizing, and intelligent traffic routing to maximize efficiency and minimize waste.

Role Qualifications

Must-Have

  1. Experience: 5+ years in DevOps, Platform Engineering, or Site Reliability Engineering roles, with a proven track record of operating production systems at scale.
  2. Container Orchestration: Deep expertise with Kubernetes (EKS, GKE, or AKS) in production environments, including cluster management, application deployment, scaling, and troubleshooting.
  3. Infrastructure as Code: Strong experience with Terraform for managing and provisioning cloud infrastructure in a repeatable, version-controlled manner.
  4. Cloud Architecture: Deep understanding of cloud-native architecture principles and hands-on experience with major cloud providers (AWS, GCP, or Azure).
  5. CI/CD Tooling: Experience designing and maintaining CI/CD pipelines using tools such as GitHub Actions, ArgoCD, Tekton, or equivalent platforms.
  6. Scripting & Automation: Strong scripting and automation skills in Python or Go, with a focus on building maintainable, production-ready tooling.
  7. Observability: Experience with observability stacks including OpenTelemetry, Prometheus, Grafana, New Relic, or similar tools for monitoring, tracing, and alerting.
  8. AI/ML Systems: Demonstrated experience shipping and operating production ML or AI systems, with an understanding of the unique infrastructure challenges they present.

Preferred Skills

  1. Hands-on experience with LLM APIs (Anthropic, OpenAI, or AWS Bedrock) and an understanding of their integration patterns.
  2. Experience developing or deploying MCP (Model Context Protocol) servers.
  3. Vector database operations experience with tools such as Pinecone, Weaviate, Qdrant, or pgvector.
  4. Experience building platforms at enterprise scale (1,000+ internal users), with a focus on developer experience and self-service capabilities.
  5. Background in Developer Experience (DX) engineering, with a passion for creating intuitive, productive workflows.
  6. Security certifications including CKS, CISSP, or CCSP are a plus.

The Perks

Our client offers a comprehensive suite of Perks & Benefits to all eligible employees. Availability and eligibility may vary based on employment status and location, but generally include competitive medical, dental, and vision coverage, generous PTO, substantial employee discounts, and robust retirement savings plans.

Ready to grow your career?

Let's get started