Before you read further: This is not a 9-5. There is no runbook for every incident, no team to absorb the blast radius of a missed SLA, and no manager between you and production. If you need someone to tell you what to do when things break at 2am, this role is not for you. If you automate toil with agents, build for reliability before you need it, and take personal ownership of uptime, keep reading.
The 47-second deploy is only a promise if the infrastructure behind it never flinches. As CreateOS scales from thousands to millions of deployments, the reliability bar moves up every week. We are running GPU workloads, multi-chain node infrastructure, and a growing base of AI-native applications that demand near-zero downtime.
The SRE who joins now will not inherit a mature, documented system. They will build it. They will define what reliable means at CreateOS, set the SLO culture, architect the observability stack, and be the person the entire engineering team trusts when things go wrong.
This is the moment when the patterns you establish become permanent. That is rare.
CreateOS is the AI-native deployment OS built for the agentic era. One-click full-stack deploys in 47 seconds, native MCP integration with Claude Code and Cursor, managed databases, GPU compute, and a Skills marketplace. $4.6M in revenue. 80,000+ active builders. Growing fast.
As a Site Reliability Engineer, you will own the reliability, scalability, and performance of CreateOS's core deployment infrastructure. You will work at the intersection of software engineering and infrastructure operations, building the systems that let our 47-second deploy claim hold up under real-world load at scale.
This is a high-ownership role. You will build things that did not exist before, own the on-call rotation, and have direct influence over our infrastructure architecture decisions.