Senior Manager, Site Reliability Engineering

Job type: Full Time · Department: Software Development · Work type: Remote · USD 190,000-250,000 / year

San Francisco, California, United States; Los Angeles, California, United States; Denver, Colorado, United States; Austin, Texas, United States; Chicago, Illinois, United States; New York, United States; Seattle, Washington, United States

Sr Manager, Site Reliability Engineering

About Invoca

Invoca is an AI-powered revenue execution platform that brings together marketing, commerce, and contact center teams to turn every customer interaction into measurable, profitable growth. Join our dynamic, fast-growing team, where innovation and collaboration are at the core of our culture.

About the team

Reliability Engineering builds and runs the platform the rest of Invoca ships on. Real-time call routing, audio processing, transcription, and data storage all run on the infrastructure this team owns, across the US and EU, under PCI, SOC 2, and ISO obligations.

Invoca believes in service ownership. Engineering teams are on call 24/7 for what they build, not just SRE. Leadership acts as an escalation point and joins the incident commander rotation roughly two days per month.

About the role

Invoca is building an agentic development environment, one where software is designed, written, tested, deployed, and operated with AI as a first-class collaborator. Infrastructure is where that gets real or doesn't: agents move fast, and the platform underneath them is what makes that safe. The fastest-growing consumer of infrastructure operations isn't a person clicking through a portal, it's software. That’s what we’re building for.

You help shape where it goes next. Getting there starts with shedding weight: finding where we can leverage partners and managed services, building self-service paths engineers and agents can both use safely, and retiring what we shouldn't be running at all.

This is not a caretaking role. You'd inherit a capable team with deep knowledge of our systems, being asked to work differently than it has. You bring outside perspective: you've built and run infrastructure teams elsewhere and you arrive with a specific point of view on what great looks like.

This position reports to the Senior Director, Infrastructure & Security.

What you'll do

Lead the evolution of Invoca's infrastructure

  • Decide what we run and what we don't. We keep a system in-house only when it differentiates us. Apply that test system by system and find the right home.

  • Land the move from self-hosted to managed Kubernetes making sure the operational expertise stays with us.

  • Reduce what the team carries operationally so capacity goes toward building and not maintenance.

  • Own the technical direction for the platform your team runs, partnering with Developer Enablement, Infosec, and architecture where decisions reach beyond it.

Lead the team and own the 3-month delivery plan

  • Lead a team of 8 SREs with deep systems knowledge. Set a clear bar, give people the feedback and context to hit it, and invest in the growth of the team.

  • Own and execute the rolling 3-month delivery plan: translate infrastructure strategy into clear priorities, navigate ambiguity, and hold full accountability for shipping high-impact work.

  • Represent SRE where infrastructure decisions get made, and be in architecture and planning conversations before your team is needed.

  • Work alongside a leadership team including a tech lead, architect, and your engineering management cohort to solve problems bigger than any one team.

Build infrastructure with agents in mind

  • Drive a pipeline, not a bottleneck. Work the team does on another team's behalf becomes a callable API with a policy decision attached — invoked by an agent or a person with the same governance and the same audit trail.

  • Extend policy and scanning controls so agent-driven infrastructure change is safe by default rather than safe by review.

  • Define clean boundaries and contracts between infrastructure and the product teams that build on it.

  • Win on adoption, not mandate. The path with policy and audit attached has to also be the fastest one available for an agent or a person.

Build for the agentic development era

  • Use AI and agentic workflows in how the team builds and operates infrastructure: agent-assisted development, automated incident detection and triage, self-healing deployments, drift and policy detection.

  • Establish safe, observable, and auditable ways to bring AI into how we build and run infrastructure, and help the teams around you adopt them with confidence.

Standardize and multiply engineering leverage

  • Build shared practices and tooling for infrastructure that every team can adopt, replacing one-off, team-specific approaches.

  • Treat infrastructure as a product with internal users: define adoption goals and measure whether teams are actually better off because of your work.

Connect the infrastructure to business outcomes

  • Translate technical decisions in the platform into terms leadership and product teams care about: velocity, cost, risk, customer trust, and revenue impact.

  • Partner with engineering and product leadership to prioritize your work based on measurable business and customer impact, not technical elegance alone.

What we're looking for

We're looking for an operator who moved into leadership: someone who has grown infrastructure teams, holds a high bar, and is deeply invested in the people on them. You're the right shape for this role if you have:

  • 5+ years hands-on in an SRE, DevOps, systems, or infrastructure engineering role, and 3+ years directly managing teams in one of those disciplines.

  • Real depth across the platforms we run on:

    • Cloud infrastructure in AWS/GCP

    • Kubernetes

    • Infrastructure as code (Terraform) and policy-as-code for governing it

    • GitOps and continuous delivery (ArgoCD and Atlantis)

    • Observability (Prometheus, Grafana, ELK)

    • Linux (configured via Chef)

    • MySQL

  • Strong opinions about infrastructure design and automation, held with an open mind, and the judgment to use both what exists today and where the industry is heading to guide decisions.

  • Significant experience as a hiring manager — you know what strong looks like at this level because you've hired it — and a track record of inheriting a team and raising its performance through development, hiring, and clear expectations.

  • You can work alongside highly skilled external specialists, able to hold your own technically, and can hold them to a high bar.

  • A working point of view on AI in infrastructure work: where it genuinely accelerates the work and where it doesn't, what output you can trust and what needs a gate, and how to raise a whole team's fluency rather than leaving it to individuals.

  • Effective in a remote-first, asynchronous culture, and able to report complex technical progress in terms a non-engineering audience can act on.

How you work

  • Direct and kind. You say the hard thing in the room, not afterward, and you receive feedback the same way.

  • High agency. You act without waiting for permission on reversible decisions and own the outcome either way. No work is beneath you.

  • You raise the bar. You think from first principles rather than accepting how things have been done, and you don't reward work that's merely fine.

  • You make it safe to not know. A high bar and a safe team are not in tension — people have to be able to ask, dissent, and be wrong without paying for it, or you never hear the thing you needed to hear. Nobody is made small for the gap between where they are and where the bar is.

  • You bring intensity. Deadlines in days where the work allows, a rough first version to get real feedback, risk flagged early, and you finish what you start.

  • You dig for the real problem. When someone asks for something, you understand what they're trying to accomplish before responding.

  • You disagree productively and commit fully. Decisions you deliver to your team are yours, not something handed down.

How success will be measured

Within 12 months, success in this role will look like:

  • The team's performance is legible without you in the room. Reliability, delivery, and operational load metrics are reviewed regularly and someone outside SRE could tell whether the team and the platform had a good month.

  • When asked, internal customers are delighted when using our infrastructure and working with our team.

  • The team builds for those using the platform, not those running it. What gets prioritized traces back to what other teams actually need, dates are met reliably enough that people plan around them, and the volume of work routed to your team for execution has materially dropped.

  • Agents are a first-class consumer of the platform. Routine infrastructure change happens through interfaces a person or an agent uses identically, with the same identity, policy, and audit trail.

  • Capacity freed by the platform transition went somewhere visible — reinvested in work the team chose, not absorbed into maintenance that grew to fill it.

  • The team’s work has a demonstrable, articulable connection to business or customer outcomes.

📍 Location: This is a remote-first role. We are currently hiring in the following locations: 📍

United States: Greater Los Angeles Area (including Santa Barbara and San Diego) · SF Bay Area · Denver Metro · Austin Metro · Chicago Metro · Greater NYC Area · Seattle Metro

Candidates must be based within ~2 hour drive of these areas. Occasional business travel may be required.

Compensation, Benefits & Perks:

At Invoca, we believe great work starts with taking care of our people. That's why we offer competitive compensation and benefits designed to support your health, your family, your growth, and your life outside of work. For teammates in the U.S., health benefits begin on your first day of employment.

Compensation

  • Base Salary: $190,000–$250,000*

  • Variable Compensation: Bonus and/or commission eligibility, where applicable.

  • Equity: All employees are invited to share in Invoca's success through our stock option program.

*Actual compensation will depend on factors including experience, skills, and location.

Benefits & Perks

Health & Wellbeing

  • Day-One Health Benefits – Medical, dental, and vision coverage begins on your first day of employment for U.S.-based teammates.

  • Mental Wellbeing – Access to mental wellbeing support and an Employee Assistance Program (EAP).

  • Wellness Subsidy – Reimbursement that can be applied toward gym memberships, fitness classes, and more.

Time Away

  • Flexible Time Off – We encourage a healthy work-life balance. Our flexible paid time off policy gives you the flexibility to recharge and take time away when you need it.

  • Paid Holidays – Invoca provides 20 U.S. paid holidays, including a winter break, giving you time to rest, recharge, and spend time with friends and family.

  • Paid Family Leave – Up to 12 weeks of 100% paid leave for baby bonding, adoption, and caring for family members.*

  • Paid Medical Leave – Up to 12 weeks of 100% paid leave for childbirth and medical needs.*

Growth & Career

  • Enterprise-Grade AI Tools – Access to industry-leading enterprise AI tools, including leading LLMs and agentic AI workflow solutions to help you maximize productivity and work more effectively.

  • Professional Development – Annual reimbursement of up to $1,000 per fiscal year to support learning, certifications, conferences, and other professional development opportunities.

Financial Benefits

  • 401(k) – Invoca offers a 401(k) plan through Fidelity with a company match of up to 4%.*

Recognition & Community

  • Recognition Programs – We celebrate great work through company-wide and team recognition programs, including President's Club and quarterly and departmental awards recognizing innovation, execution, and customer impact.

  • InVacation (Sabbatical) – As a thank-you to our long-term team members, we offer a sabbatical after seven years of service.

*Some benefits may vary for teammates outside the U.S. based on local laws, regulations, and country-specific offerings.

Made with

Senior Manager, Site Reliability Engineering | Invoca Careers