CareerFit
/
Sign inGet started
/
CareerFit

Every open role from the boards that matter, in one pool you can search, score against your CV and track.

Explore

  • All roles
  • Remote roles
  • How it works

Account

  • Create a free account
  • Sign in
Back to the pool

Ashby

Engineering Manager, Site Reliability Engineering

Replit·Foster City, CA

RemoteLeadFullTime250k–325kRemote country eligibility unknownSource posted 12d ago
Apply on Ashby

Opens the original posting. You apply there, not here.

Pay
250k–325k
Level
Lead
Type
FullTime
Where
Remote
Source posted
12d ago
First seen by us
Sep 30, 2026
Last seen on its board
Not recorded

Last observation does not confirm this role is still open.

Keep track of this one

Save it, hide it, or log your application and a follow-up date. Free with an account.

Sign in to track

Skills in this posting

  • SRE
  • Distributed Systems
  • Observability
  • OpenTelemetry
  • Kubernetes
  • Mentoring
  • ArgoCD
  • GCP

The posting

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. ABOUT THE ROLE Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows. This Engineering Manager will lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure. You'll lead and grow an existing team that builds and operates production platforms and works hands-on across application and infrastructure boundaries. This is a software-building leadership role, not simply an incident-management function. You'll help teams ship safely, understand production behavior, and remove performance bottlenecks through concrete engineering improvements. You should be comfortable going deep on a rollout failure or performance investigation while developing technical leaders and sustainable ownership across a distributed team. WHAT YOU'LL DO - Observability. Build and operate metrics, logs, traces, and alerting capabilities. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements. - Incident Management. Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures. - Load Testing. Build and maintain load/failure testing capabilities. Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners. - Performance Engineering. Lead deep engagements with internal teams on SLOs and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners—not just recommendations. - Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding tools—including Replit—to prototype and automate. Apply rigorous review and verification to AI-generated changes. - Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate rather than relying on a few permanent escalation points. - Measure outcomes and close the loop. Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements arising from cost/capacity analysis. Agree success measures and continuing ownership with partner teams. WHAT YOU'LL BRING - Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team—not only acted as its strongest individual contributor. - Software-oriented production systems depth. You have built and operated distributed systems or reliability platforms and can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms. - Safe-change and performance judgment. You have led consequential migrations or incidents and used measurement to diagnose reliability or performance problems. You can distinguish symptoms from causes and validate fixes under realistic conditions. - Platform-product and cross-team judgment. You can build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost. NICE TO HAVE - Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo. - Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling. - Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP. - Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match (US Only) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits (In-Office & US Only) 📱 Monthly Wellness Stipend 🧑‍💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement (In-Office Only) 🚀 Quarterly Team Gatherings ☕ In Office Amenities (In-Office Only) Want to learn more about what we are up to? - Self-driving Company https://replit.com/blog/self-driving-company - Replit Agent at Scale https://replit.com/blog/evaluating-and-improving-agent-at-scale - AI Adoption https://replit.com/blog/ai-adoption - Build Open-Source Apps https://replit.com/build/open-source-app-builder   Interviewing + Culture at Replit - Operating Principles https://blog.replit.com/operating-principles - Reasons not to work at Replit https://blog.replit.com/reasons-not-to-join-replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Your apply kit

Apply with a head start

A cover letter written for this role

Sign in, add your CV, and get a first draft written for this exact posting — ready to edit and paste into the application.

Sign in to draft one

Check the company before you apply

Public sources · not verified by us

Background signals to help you judge whether this employer is real, what working there is like, and what has been in the news.

One-click background checks

Works right now, no account needed here

Ready-made searches for Replit on public career sites and in the news.

ReputationGlassdoor Reviews & SalaryCheck employee ratings, culture reviews, and verified salary ranges.VerificationLinkedIn Company ProfileInspect headcount, team growth, and verified employee roster.Due DiligenceNews & Layoffs (Google News)Scan recent news for funding, layoffs, acquisitions, or warnings.AuthenticityOfficial Careers SearchConfirm whether this vacancy is published directly on the employer's portal.

These are search shortcuts built from the employer’s name. Some sites ask for a free account. Never pay an application fee or send money for equipment to get a job.

What the web says

Sign in and we will search public sources and summarise what they say about this employer. The one-click checks above work without an account.

Your research notes

Sign in to save links and notes as you look into this employer. Only you can see them.