Director of Infrastructure Engineering
Job Description
Job Title: Director of Infrastructure Engineering
Job Type: Full-time
Location: Remote
The Role
We're looking for a Director of Infrastructure Engineering to build and scale the platform that powers production AI systems. You'll own the strategy and execution behind our cloud infrastructure, developer platform, and reliability practices, ensuring our systems remain secure, observable, and resilient as we grow.
This is a hands-on leadership role for someone who enjoys operating at every level—from defining long-term infrastructure strategy to solving production incidents and mentoring high-performing engineers.
What You'll Do
Own the architecture and evolution of our multi-cloud infrastructure across AWS and GCP, optimizing for scalability, reliability, security, and cost.
Lead and grow a high-performing Infrastructure/Platform Engineering team while establishing a strong engineering culture centered on ownership, operational excellence, and continuous improvement.
Build infrastructure as code using Terraform (or equivalent), ensuring reproducible, version-controlled, and auditable environments.
Design and improve CI/CD systems that enable fast, reliable, and secure software delivery.
Build world-class observability across metrics, logs, traces, and alerting to proactively detect and resolve production issues.
Define and drive reliability practices, including incident response, on-call operations, SLOs, error budgets, disaster recovery, and blameless postmortems.
Partner closely with Security and Engineering leadership to embed security by default and maintain compliance with frameworks such as ISO 27001, SOC 2, and CMMC.
Continuously improve the developer experience by reducing operational friction through automation and platform tooling.
What We're Looking For
8+ years building and operating production infrastructure, platform engineering, DevOps, or SRE systems.
3+ years leading engineering teams in high-growth environments.
Deep expertise with AWS, GCP, or multi-cloud production environments.
Strong experience with Terraform, Kubernetes, containers, and modern CI/CD platforms.
Proven experience building highly available, observable, and resilient production systems.
Strong understanding of infrastructure security, compliance, and operational risk management.
Experience scaling engineering organizations and infrastructure in fast-moving startup environments.
Excellent communication skills with the ability to influence technical strategy across engineering and executive stakeholders.
Preferred
Experience supporting AI/ML infrastructure or large-scale data platforms.
Familiarity with model training, inference, evaluation, or data pipeline infrastructure.
Experience operating regulated cloud environments such as FedRAMP, GovCloud, or CMMC Level 2.
Contributions to platform engineering, open-source infrastructure, or developer productivity initiatives.
Compensation & Benefits Notice
The national pay range for this full-time position is base salary of $230,000 –$260,000. All employees are eligible for equity compensation, and employees may also receive performance-based bonuses, dependent on role and subject to company policies. micro1 provides a comprehensive benefits package, including up to 100% reimbursement for health-insurance premiums, paid time off, a 401(K) plan with a company match, and additional benefits designed to support a high-performing, remote-first workforce.
micro1 is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, or gender identity), national origin, age, disability, genetic information, veteran status, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance and/or a reasonable accommodation during the application process, reach out to support@micro1.ai.
Our hiring process utilizes artificial intelligence tools to assist in candidate screening and assessment. Our AI tools are designed to complement, not replace, human decision-making.
Disclaimer
The information contained in this job posting, including but not limited to role responsibilities, qualifications, compensation, and benefits, is provided for informational purposes only and does not constitute a binding offer of employment. micro1 reserves the right to amend, modify, or withdraw any portion of this posting at its sole discretion and without prior notice. All employment decisions are made in accordance with applicable laws and regulations.
About Micro1
Micro1 is building the essential infrastructure for the next generation of artificial intelligence, moving beyond the limitations of static datasets to create a dynamic ecosystem where models truly learn to think and act. At its core, the company understands that the path to frontier intelligence isn't paved with synthetic data alone; it requires the nuance, judgment, and contextual awareness that only expert human guidance can provide. By positioning itself at the intersection of human expertise and advanced reinforcement learning, Micro1 is tackling the hardest problem in AI today: teaching models not just to answer questions, but to reason through complexity and take meaningful actions in unpredictable, real-world environments.
This mission comes to life through Realm, Micro1's flagship training environment. Realm functions as a high-fidelity simulation ground where AI agents are immersed in scenarios that closely mirror the messiness of actual human workflows. Rather than relying on abstract benchmarks, Realm forces models to engage in agentic actions, multi-step decisions, digital navigation, coding, and problem-solving that require genuine comprehension. It is within these realistic sandboxes that world-class human data is generated, creating a continuous feedback loop where expert annotators observe, correct, and guide model behavior. This process does more than just fine-tune outputs; it fundamentally elevates a model's underlying reasoning architecture, ensuring that the intelligence developed in training holds up when deployed into the wild.
But training is only half the equation. Micro1 recognizes that the true test of an AI system lies in its production performance, which is where Cortex enters the picture. Cortex is a contextual evaluation platform designed to move beyond superficial accuracy metrics and provide a granular, real-time view of how agents behave in live settings. It doesn't just flag errors; it surfaces the underlying context behind failures, revealing why an agent struggled with a particular user intent or environmental variable. This insight is invaluable for engineering teams, as it transforms evaluation from a passive checkpoint into an active improvement tool. With Cortex, companies can catch degradation early, understand the nuances of edge cases, and feed those learnings directly back into Realm for retraining, creating a virtuous cycle of continuous advancement.
Ultimately, Micro1 is not merely a data lab or an evaluation tool it is a comprehensive intelligence engine that connects the entire lifecycle of model development. From the expert human feedback that sharpens raw capability, to the realistic RL environments that forge robust agentic behavior, to the contextual oversight that ensures reliability in production, Micro1 provides the integrated foundation that AI labs and enterprises desperately need. As the industry races toward truly autonomous systems, Micro1 stands as the critical bridge between laboratory breakthroughs and dependable, real-world impact, ensuring that frontier models do not just perform well on paper, but deliver genuine value in the complex, dynamic world they are meant to serve.

