Location: 

(

San Francisco

,

CA

)

Salary: 

$

200k

 - $

400k

Director of Engineering, AI Infrastructure & Platform

Location: Palo Alto, CA (Onsite, 5 days a week)

Team: Platform Engineering

Partner: A mission-driven company building a safety-focused Large Language Model designed specifically for healthcare.

About Our Hiring Partner

Our hiring partner is tackling one of the most ambitious and meaningful challenges at the intersection of technology and society: building a safety-first Large Language Model to revolutionize global healthcare. Their mission is to dramatically improve health accessibility and outcomes by making deep medical expertise available to everyone through AI. This represents a foundational technology with the potential for unprecedented positive impact on human health.

The company was co-founded by a visionary coalition, including a seasoned CEO, alongside leading physicians, hospital administrators, and AI researchers from world-class institutions like Stanford, Johns Hopkins, and technology pioneers from Google and NVIDIA. With a formidable $278 million in funding from elite investors such as Andreessen Horowitz and General Catalyst, they possess the strategic backing, expertise, and resources to execute their long-term vision. For engineering leaders seeking purpose-driven startup jobs in SFC, this represents a pinnacle opportunity to build and scale foundational technology.

Why This is a Foundational Leadership Opportunity

  1. Build the Engine of Innovation: Architect the core infrastructure that powers cutting-edge AI research and global healthcare applications, directly enabling the company's mission.
  2. Lead a World-Class Team: Recruit, mentor, and lead a multidisciplinary team of platform engineers, SREs, and infrastructure specialists.
  3. Solve Unprecedented Scale Challenges: Design and operate a global, multi-cloud GPU fabric to support the most advanced LLM workloads in a security-critical environment.
  4. Execute with Substance: Join a well-capitalized, high-growth company where you can focus on building robust, scalable systems—a key advantage for leadership roles in startup jobs in SFC.

About the Role

We are seeking a Director of Engineering – Platform Engineering to lead the design, implementation, and operation of our core cloud infrastructure, observability systems, and global GPU control plane. This leader will be responsible for scaling a distributed compute fabric to support cutting-edge LLM workloads while maintaining exceptional reliability, security, and cost efficiency.

You will build and lead a high-performing engineering team, working closely with product development, AI research, and compliance to deliver world-class infrastructure for safety-critical healthcare AI. This role is the cornerstone of our technical execution, ensuring that our AI innovations can be reliably delivered to users around the world.

What You'll Do

  1. Team Building & Leadership: Recruit, develop, and inspire a world-class platform engineering team. Foster a culture of operational excellence, innovation, and rigorous accountability.
  2. Infrastructure Strategy: Define and execute the long-term technical roadmap for a multi-cloud, multi-region GPU and compute environment, driving excellence in cost optimization and capacity planning.
  3. Cloud & Control Plane Management: Architect and manage the global GPU control plane, enabling dynamic provisioning and monitoring of inference workloads. Lead the automation of deployments using infrastructure-as-code (e.g., Terraform) and CI/CD best practices across providers (AWS, GCP, Azure).
  4. Security & Compliance: Ensure a robust security posture and strict compliance across all environments, adhering to healthcare standards including HIPAA and SOC 2.
  5. Observability & Reliability: Scale comprehensive observability systems—covering metrics, tracing, and logging—to ensure full visibility into production AI workloads. Establish and uphold SLOs/SLIs and implement rigorous incident management processes.
  6. Technical Collaboration: Partner with AI and product teams to anticipate infrastructure needs and design scalable architectures that support rapid experimentation and global deployment.

Qualifications (Must-Have)

  1. 10+ years of engineering experience, including 5+ years in a leadership role overseeing infrastructure, SRE, or platform teams at scale.
  2. A proven track record of managing large-scale distributed systems and global cloud infrastructure.
  3. Deep, hands-on experience with high-performance computing (HPC) or large-scale AI/ML workloads.
  4. Strong expertise in cloud platforms (AWS, GCP, Azure) and infrastructure-as-code tools (e.g., Terraform, Pulumi).
  5. Mastery of modern observability stacks (e.g., Prometheus, Grafana, Datadog) and a commitment to operational excellence.
  6. Demonstrated experience implementing security and compliance frameworks relevant to healthcare (e.g., HIPAA, SOC 2).
  7. Exceptional communication and cross-functional partnership skills, with the ability to align technical strategy with product and research goals.

Bonus Qualifications (Nice-to-Have)

  1. Experience designing or operating GPU control planes or schedulers (e.g., Kubernetes, Ray, Slurm).
  2. Prior work with ML infrastructure, large-scale data pipelines, or model-serving platforms.
  3. A background in sophisticated cloud cost optimization and sustainability initiatives for large-scale compute operations.
  4. Familiarity with edge or hybrid-cloud deployments for low-latency AI inference systems.

This is a full-time, onsite leadership role based in Palo Alto, CA. The company culture is built on the power of in-person collaboration to accelerate innovation and build a cohesive, high-performing team.

Apply Now

If you are an infrastructure leader passionate about building the foundational platforms that power world-changing AI, we strongly encourage you to apply. This role stands out among startup jobs in SFC as a chance to lead at the absolute frontier of scale, reliability, and technical impact in the service of a historic mission.

Ready to grow your career?

Let's get started