Job Title: Senior Backend Engineer, Infrastructure & Reliability
Location: San Francisco, CA (Onsite)
Team: Engineering
Partner: A pioneering AI company building autonomous, quantitative systems for growth marketing, backed by Quiet Capital.
About Our Hiring Partner
Our hiring partner is redefining the future of digital growth by building an AI Operator that autonomously manages advertising spend with the precision of high-frequency trading. Their platform delivers transformative efficiency gains by replacing manual agency processes with algorithmic decision-making and execution. Already managing critical budgets for leading global brands, they are on a clear trajectory to scale into billions in autonomous volume. As an early-stage team in San Francisco, they are assembling foundational engineering talent to build the robust, reliable systems required to manage financial decisions at scale.
Why This is a Foundational Role
- Build the Bedrock: Architect the core infrastructure that ensures the reliable, uninterrupted operation of a system managing millions in daily spend. Your work is the foundation upon which all autonomous decisions execute.
- Solve Complex Deployment Challenges: Tackle the unique engineering puzzle of deploying and managing software across diverse, secure environments, including private client VPCs.
- Be the Force Multiplier: As the guardian of production quality and developer experience, you will build the paved road that allows the entire engineering team to ship high-frequency systems safely and with velocity.
- SF-Based Technical Leadership: Join a concentrated team in San Francisco in one of the most critical and challenging San Francisco startup jobs, owning the reliability of a cutting-edge financial AI system.
About the Role
We are seeking a Backend Engineer to build the bedrock of our autonomous system. You are the architect of reliability. While our algorithms team designs the trading strategies, you will ensure the engine runs without interruption. Your mission is to operationalize ML pipelines, manage complex deployments across fragmented environments, and build the infrastructure that allows us to ship high-frequency systems safely.
This role is for the seasoned engineer—the adult in the room—who can foresee failure modes before they happen. You will be the guardian of production, ensuring our system can execute with zero downtime.
What You'll Do
- Architect Multi-Environment Deployments: Design and manage the infrastructure to deploy our decision engine into diverse environments, including our cloud and private client VPCs, solving the lifecycle management challenges of secure, isolated instances.
- Operationalize the Quant Engine: Transform experimental ML models into resilient, scalable production services. Build and maintain the robust training, inference, and retraining pipelines that keep our models performant.
- Build the Event-Driven Backbone: Architect the scheduling and event-driven systems that power our high-frequency control loops. Manage the message queues and orchestration layers that coordinate data flow across our platform.
- Guardian of Reliability: Own system health. Implement comprehensive observability (monitoring, logging, tracing) and act as the primary troubleshooter, converting incidents into automated prevention mechanisms.
- Elevate Developer Experience: Treat internal engineers as your customers. Build the CI/CD pipelines, local development environments, and internal tooling that maximize shipping velocity and confidence.
Who You Are
- A Senior Operator: You have 6-8+ years of experience in ML Platform, Infrastructure, or DevOps/SRE roles. You have seen how complex systems break at scale and know how to design them to survive.
- Platform Mindset: You possess a high-leverage mentality, focused on building the paved road that enables other engineers to do their best work efficiently and reliably.
- Distributed Systems Native: You are deeply comfortable with the complexities of data consistency, concurrency, and latency in distributed, high-throughput environments.
- Architect of Reliability: You proactively prioritize system health, designing for observability and automation from the outset to ensure critical financial workloads remain stable.
Bonus Points
- MLOps Expertise: Experience with feature stores, model registries, and tools like Ray, MLflow, or Kubeflow.
- Client Environment Experience: A background deploying and managing software within customer-controlled cloud environments (AWS/GCP/Azure VPCs).
- Security-First Mindset: Experience with SOC2 compliance, IAM policies, and securing sensitive financial data.
Compensation & Benefits
This role offers a competitive package reflective of the seniority and critical nature of the position.
- Salary: $180,000 – $270,000
- Equity: A significant early-stage equity package.
- Relocation: Support for candidates moving to the Bay Area.
- Daily Meals: Provided lunch and optional dinner.
- Comprehensive Benefits: Full health, dental, and vision coverage, alongside an unlimited PTO policy.
Apply Now
If you are a seasoned backend engineer who thrives on building robust, scalable infrastructure and wants to be the foundational pillar for a groundbreaking autonomous system, we encourage you to apply. This is a defining San Francisco startup job for engineers who are passionate about reliability, infrastructure-as-a-product, and enabling others to build with confidence.