LLM Red-Teamer

📍remote
🕒Full-time
💰$40 – $65/hr
📅Posted July 15, 2026

Job Description

Role Title: LLM Red-Teamer

Role Type: Contractor

Location: Remote

micro1 is engaging LLM Red-Teamers to contribute to a high-impact customer project focused on the evaluation and improvement of frontier language models. In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-world input. No prior experience in AI is required — your domain knowledge is what matters.

Scope of Work

  1. Develop complex, adversarial multi-turn conversations and task-based scenarios aligned with detailed project specifications.

  2. Author clear, precise evaluation rubrics to rigorously assess model responses against defined behavioral targets.

  3. Iteratively test conversations and tasks against frontier LLMs, escalating difficulty and nuance until the desired quality threshold is achieved.

  4. Deliver comprehensive task packages, including transcripts, target behaviors, binary rubrics, and supporting rationale or evidence.

  5. Validate LLM outputs, documenting model strengths and failure modes relative to the project specification.

  6. Maintain calibration with team leads and quality control contacts as project requirements evolve.

  7. Contribute independently, producing high-quality deliverables at a steady and consistent pace.

Preferred Qualifications

  1. Exceptional written English skills, with clarity, precision, and strong structural organization.

  2. Prior experience in AI human data environments (RLHF, SFT, evaluations, annotation, or prompt engineering).

  3. Deep familiarity with large language models, including the ability to anticipate and identify common failure patterns.

  4. Demonstrated ability to work autonomously, interpreting and executing complex specifications with minimal oversight.

  5. Proven critical thinking and meticulous attention to detail.

  6. Experience designing evaluation items or rubrics is advantageous.

  7. Background in writing-intensive or analysis-centric fields such as research, editorial, technical writing, or quality assurance is a plus.

Compensation Structure

Compensation is output-based; experts are paid per task that meets the project specifications. The time required to complete work may vary depending on the expert’s experience and workflow. Minimum submission requirements apply. Experts must submit a minimum of tasks per week.

Start Timeline & Availability

We typically fill roles within 48 hours and are looking for experts ready to jump in right away. If selected, we expect you to start your first tasks within 24–48 hours of completing onboarding.

Share this job opportunity:

About Micro1

Micro1 is building the essential infrastructure for the next generation of artificial intelligence, moving beyond the limitations of static datasets to create a dynamic ecosystem where models truly learn to think and act. At its core, the company understands that the path to frontier intelligence isn't paved with synthetic data alone; it requires the nuance, judgment, and contextual awareness that only expert human guidance can provide. By positioning itself at the intersection of human expertise and advanced reinforcement learning, Micro1 is tackling the hardest problem in AI today: teaching models not just to answer questions, but to reason through complexity and take meaningful actions in unpredictable, real-world environments.

This mission comes to life through Realm, Micro1's flagship training environment. Realm functions as a high-fidelity simulation ground where AI agents are immersed in scenarios that closely mirror the messiness of actual human workflows. Rather than relying on abstract benchmarks, Realm forces models to engage in agentic actions, multi-step decisions, digital navigation, coding, and problem-solving that require genuine comprehension. It is within these realistic sandboxes that world-class human data is generated, creating a continuous feedback loop where expert annotators observe, correct, and guide model behavior. This process does more than just fine-tune outputs; it fundamentally elevates a model's underlying reasoning architecture, ensuring that the intelligence developed in training holds up when deployed into the wild.

But training is only half the equation. Micro1 recognizes that the true test of an AI system lies in its production performance, which is where Cortex enters the picture. Cortex is a contextual evaluation platform designed to move beyond superficial accuracy metrics and provide a granular, real-time view of how agents behave in live settings. It doesn't just flag errors; it surfaces the underlying context behind failures, revealing why an agent struggled with a particular user intent or environmental variable. This insight is invaluable for engineering teams, as it transforms evaluation from a passive checkpoint into an active improvement tool. With Cortex, companies can catch degradation early, understand the nuances of edge cases, and feed those learnings directly back into Realm for retraining, creating a virtuous cycle of continuous advancement.

Ultimately, Micro1 is not merely a data lab or an evaluation tool it is a comprehensive intelligence engine that connects the entire lifecycle of model development. From the expert human feedback that sharpens raw capability, to the realistic RL environments that forge robust agentic behavior, to the contextual oversight that ensures reliability in production, Micro1 provides the integrated foundation that AI labs and enterprises desperately need. As the industry races toward truly autonomous systems, Micro1 stands as the critical bridge between laboratory breakthroughs and dependable, real-world impact, ensuring that frontier models do not just perform well on paper, but deliver genuine value in the complex, dynamic world they are meant to serve.