
Turing
Verified CompanyAbout the Company
The Turing story
Based in San Francisco, California, Turing is a leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: accelerating frontier research with high-quality data, advanced training pipelines that push the boundaries of reasoning, multimodality, and STEM, plus top AI researchers with expertise in frontier data for coding, reasoning, STEM, multilinguality, multimodality, and agents; and applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence: systems that perform, deliver impact, and drive lasting results on the P&L.
Turing has received numerous awards, including Forbes’s “One of America’s Best Startup Employers,” #1 on The Information’s annual list of “Most Promising B2B Companies,” and Fast Company’s list of the “Best Workplaces for Innovators." Turing’s leadership team includes AI technologists from industry giants Meta, Google, Microsoft, Apple, Amazon, Twitter, McKinsey, Bain, Stanford, Caltech, and MIT.
Contact Information
No contact details publicly provided.
Active Openings 57
Senior Python Developer
Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:This position is within a project with one of the foundational LLM companies. The goal is to assist these foundational LLM companies in enhancing their Large Language Models.One way we help these companies improve their models is by providing them with high-quality proprietary data. This data serves two main purposes: first, as a basis for fine-tuning their models, and second, as an evaluation set to benchmark the performance of their models or competitor models.For example, for SFT data generation, you might have to put together or be provided a prompt which contains provided code and questions, you will then provide the model responses, and write corresponding Python code to solve the questions.For RLHF data generation, you may need to create a prompt yourself or use one provided by the customer, ask the model questions, and evaluate the outputs generated by two versions of the LLM. You'll compare these outputs and provide feedback, which is then used to fine-tune the models. Please note that this role does not involve building or fine-tuning LLMs.What does day-to-day look like:Design, develop, and maintain efficient, high-quality Python code to train and optimize AI models.Conduct evaluations (Evals) to benchmark model performance and analyze results for continuous improvement.Familiarity with Python frameworks and librariesEvaluate and rank AI model responses to user queries across diverse domains, ensuring alignment with predefined criteria.Develop comprehensive explanations and rationales for evaluations, showcasing excellent reasoning and technical expertise.Lead efforts in Supervised Fine-Tuning (SFT), including creating and maintaining high-quality, task-specific datasets.Collaborate with researchers and annotators to execute Reinforcement Learning with Human Feedback (RLHF) and refine reward models.Design innovative evaluation strategies and processes to improve the model's alignment with user needs and ethical guidelines.Create and refine optimal responses to improve AI performance, emphasizing clarity, relevance, and technical accuracy.Conduct thorough peer reviews of code and documentation, providing constructive feedback and identifying areas for improvement.Collaborate with cross-functional teams to improve model performance and contribute to product enhancements.Continuously explore and integrate new tools, techniques, and methodologies to enhance AI training processes.Requirements:3+ years of strong experience with Python programming language.Industry experience and knowledge of code quality, formatting, and best practices of software developmentExperience with Python’s testing ecosystem, including unit, integration, and property-based testing.Knowledge of multi-threading and asynchronous programming in Python.Ability to work with architectural patterns and refactor code without introducing regressions.Strong debugging skills, including fixing memory and concurrency issues.Fluent in conversational and written English communication skillsPerks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: at least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Engagement type : Contractor assignment (no medical/paid leave)Duration of contract : 1 month; [expected start date is next week]Evaluation Process (approximately 75 mins) :Two rounds of interviews (60 min technical + 15 min cultural &, offer discussion)
Product Manager – Operations & Controls Platform (Fund Controller Accounting & Global Business Finance)
Product Manager – Operations & Controls Platform (Fund Controller Accounting & Global Business Finance)FULLTIMELocation: New York, New YorkEmployment Type: Full-TimeLocation: New York, NYPosition OverviewWe are looking for an experienced Product Manager to lead the strategy, roadmap, and execution for platforms supporting Back Office Operations, especially around Fund controller accounting and Global Business Finance. The ideal candidate brings strong product management experience combined with deep knowledge of financial operations, operational controls, and controller functions.This role will partner closely with business stakeholders, engineering, and operations teams to modernize back-office platforms and improve operational efficiency. While familiarity with AI is a plus, the primary focus is on delivering business value through well-designed operational systems rather than building AI solutions.Key ResponsibilitiesOwn the product vision, roadmap, and prioritization for Operations & Controls platforms.Lead product discovery by understanding business problems, operational workflows, and stakeholder requirements.Partner with Back Office Operations, Controllers, Finance, and Technology teams to improve operational processes.Drive modernization initiatives in fund control accounting workflow automation, and operational reporting.Build and manage the product backlog, define user stories, and prioritize features based on business value.Coordinate delivery across engineering, architecture, and business stakeholders.Track milestones, dependencies, risks, and delivery progress.Promote transparency and effective stakeholder communication throughout the product lifecycle.Identify practical opportunities where AI and automation tools can improve productivity, decision-making, and operational efficiency.Ensure solutions meet governance, audit, and operational control requirements.Required Qualifications5+ years of experience as a Product Owner, Product Manager, or Senior Business Analyst within Financial Services.Strong experience supporting Back Office Operations, Fund Controllers, and Finance Operations.Has fund accounting domain literacy, hands-on exposure to at least one fund accounting/admin platform, exception managementExperience building or enhancing corporate back-office platforms and operational systems.Demonstrated experience translating fund control requirements — NAV integrity, cash/position break resolution, and controller sign-off workflows — into product specs and acceptance criteria that satisfy audit and governance standards.Excellent stakeholder management and cross-functional collaboration skills.Proven ability to manage product roadmaps, planning, prioritization, and execution.Strong analytical, communication, and problem-solving skills.Familiarity with AI tools (e.g., ChatGPT, Copilot, Gemini) and understanding of where AI can improve operational workflows. Hands-on AI solution development is not required.Ability to work effectively with engineering teams and understand modern software delivery practices.Preferred QualificationsExperience with platforms such as Geneva/SS&C, Eagle, Advent, Paxus, or similar fund accounting and operations systems.Experience working with fund administrators or investment operations teams.Experience with Agile product delivery.Exposure to workflow automation, cloud platforms, or modern enterprise applications.
Agentic Coding Annotator - Online / Offline Tasks
About TuringTuring is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Turing helps customers in two ways: working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM, and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role OverviewWe are looking for strong, detail-oriented software practitioners to help evaluate and improve datasets for agentic coding models.This role involves working with realistic coding tasks in an agentic coding harness, reviewing model trajectories, verifying solutions, and producing high-quality annotations.Depending on the assignment, the work may include:Online evaluations: Manually interacting with blinded models on predefined tasks, then ranking and grading resulting trajectoriesOffline evaluations: Designing realistic coding tasks, calibrating them through user simulation, writing task-specific rubrics, and grading generated trajectoriesThis is not a basic annotation role. Candidates are expected to read and debug code, validate behavior, follow detailed process rules, and make consistent judgment calls across model runs.We are specifically looking for candidates with enough engineering maturity to independently work on realistic software tasks, not just toy problems or shallow code-review exercises.What does day-to-day look likeExecute realistic coding tasks within the assigned agentic coding harness while maintaining model blindness and session independenceFollow task instructions, milestones, planned interactions, and evaluation guardrails consistently across runsVerify model outputs by reading code, running commands, checking logs, and inspecting generated artifactsPerform targeted validation of outputs using tests, scripts, and manual checksWrite clear, specific, evidence-based rationales for trajectory rankings and assessmentsDesign multi-step, realistic coding tasks (offline work), including user intent and milestone structureCreate and refine task-specific rubrics and binary evaluation criteriaReview completed work for quality, completeness, consistency, and schema complianceIdentify and escalate broken environments, unclear instructions, or process gaps with clear supporting evidenceRequirementsSoftware Engineering Fluency (Mandatory)5+ years of experience in software engineering, QA, developer tooling, data/ML engineering, or similar code-heavy rolesStrong hands-on experience in at least 1–2 programming languages or ecosystemsRepresentative languages include :Python, JavaScript/TypeScript, Rust, Java, C/C++, Bash/CLI environments, Haskell, Swift, SQL, or other production-relevant ecosystemsAbility to: Read and understand unfamiliar codebases Run and interpret tests, scripts, and CLI tools Debug issues and reason about edge cases or partial fixes Evaluate whether an implementation is functionally correctAdditional Preferred Qualifications (Offline / Senior Candidates)Strong Docker skills and experience building/debugging reproducible environmentsExperience working in large, complex repositories (not just small or greenfield projects)Demonstrated originality and sound engineering judgment in defining technical problemsAbility to design realistic, non-trivial tasks that go beyond tutorials, README flows, or simple bug fixesPerks of Freelancing With TuringWork on cutting-edge AI projects with leading foundation model companiesCollaborate on high-impact work at the frontier of LLM evaluation and reasoningRemote, flexible opportunities with global teamsCompetitive compensation based on experience and project scopeOffer DetailsCommitments Required: 8 hours per day with a 4-hour overlap with PST.Employment Type: Contractor position (Note: this role does not include medical/paid leave).Duration of Contract: 5 weeks; [expected start date is next week].
Software Engineer – Python + Docker
We are staffing a frontier AI data initiative that requires strong software engineers to build the infrastructure and training data used to develop and evaluate AI agents.Key Responsibilities: The Engineer will, depending on assigned track: build Python backend applications that replicate existing SaaS tools (such as Slack, Linear, Jira, Notion, Gmail, and wikis), including implementing integrations and standing up full-service backend functionality; Perform thorough backend testing and integration validation to ensure each connector behaves faithfully like the system it emulates; Adapt and extend previously built connectors as needed, and test existing connectors internally before they are treated as complete; Build new connectors from scratch where required, within agreed delivery timeframes; Conduct data mining and task mining to identify representative workflows suitable for long-horizon task development; Author realistic tasks derived from mined data and workflows; Verify task and data quality, realism, and correctness through structured QA; Write clear evaluation rubrics that define correct, partially correct, and deficient work; Use AI coding agents proficiently throughout all development, QA, and validation work; collaborate across the connectors and tasks tracks, Participate in onboarding, calibration, and quality-review cycles; deliver assigned work to the expected quality bar and on agreed milestones.Must Have Required Skills:Minimum 3+ years of overall experienceStrong proficiency in Python (FastAPI, Flask, or Django)with proven experience in backend software development.Proficiency with Git, Docker, and basic software pipeline setup.Experience building scalable backend applications, REST APIs, and microservices.Proficiency in using AI coding assistants (e.g., Codex, Claude Code, Cursor, GitHub Copilot, or similar) as part of daily development workflows.Strong understanding of software engineering best practices, including version control, testing, and code quality.Strong analytical and problem-solving skills with attention to detail.Ability to work independently while collaborating effectively within distributed engineering teams.Excellent written and verbal communication skills.Nice to have:Experience building connectors or integrations for SaaS platforms.Knowledge of evaluation frameworks, QA methodologies, and rubric creation for AI datasets.Experience developing realistic workflows for long-horizon AI tasks.Offer detailsCommitments required: Must complete and get approved minimum of 1 task per day Payment: $300/TaskEngagement type: Contractor - Pay per TaskDuration of contract: 2-3 weeks and start date ASAP
AI Quality Analyst (Personalization) - Indonesian
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.Key QualificationsIndonesian Proficiency: Ability to read and write in Indonesia with a high degree of comp, as Indonesian is the focus language for this project.Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team.Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities.Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections.Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating.Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.Feedback: Ability to provide constructive feedback and detailed annotations.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Description:In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve:Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences.Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied.Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.Education & ExperienceBS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field).Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.Offer Details:Commitments Required: at least 4 hours per day and minimum 30 hours per week with 4 hours of overlap with PST. (We have 2 options of time commitment: 30 hrs/week or 40 hrs/week)Engagement type: ContractorEngagement Length: 3 monthsOur offered rate for this project is $15 per hour.Evaluation Process -Shortlisted candidates will be sent a Job Interest Form.After the profile review, an assessment will be shared, which must be completed within 24 hours.Based on the assessment outcomes, shortlisted candidates will be contacted to discuss the pre‑onboarding requirements.
Business Analyst (Dutch Language)
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:We are looking for candidates with strong analytical and English comprehension skills. The ideal candidate should have the ability to read, summarize, and break down large content into smaller logical blocks, conduct research online, validate claims made in the content through online research, and work with the LLM (Large Language Models) to solve puzzles!Your role is critical in helping fine-tune and improve large language models (like gpt), and will make you an expert on how to leverage AI to be a better analyst. This is your chance to future-proof your career in an AI-first world!What does day-to-day look like:You would spend time answering a variety of interesting analytical questions and creating scenarios that can train the LLM models to get better. Here are a couple of examples which models might get wrong and you would have to give the correct answer and explanation to enable the models to learn:Based on a given distribution of sales by month across locations, could you analyze which location has grown the most? (Hint: what time period should we look at? should we account for sudden variability at the beginning?)In a small town, there are four distinct neighborhoods: Oak, Pine, Maple, and Elm. A postman is assigned to deliver mail and can only deliver to two neighborhoods in one day, with certain rules (e.g. Oak is always visited before Pine). If he delivers to Oak and Elm on the first day, which neighborhoods does the carrier deliver to on the second day?Note: No other prior specialized domain experience is needed.Requirements:Multilingual capabilities are mandatory. (English and Dutch)Analytical Skills: Good research and analytical skillsFeedback Skills: Ability to provide constructive feedback and detailed annotations.Creative Thinking: Creative and lateral thinking abilities.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Commitment: Ability to commit to 40 hours per week for the contract duration overlapping US hours.Technical Setup: Desktop/Laptop set up with a good internet connection.Preferred Qualifications:Bachelors degree or undergraduate in Engineering, Literature, Journalism, Communications, Arts, Statistics, or a related field. We are open to candidates who do not have a Bachelor's degree but have experience in the area.Experience writing professionally (business analysts, research analyst, copywriter, journalist, technical writer, editor, translator, etc.)Understanding of Excel and Google Suite.Proficiency in Data interpretation, Logical reasoning and Basic arithmetic is highly desirable.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:Engagement type : Contractor assignment/freelancer (no medical/paid leave)Commitments Required : Availability of up to 40 hours/week is preferred, this role will require some overlap with UTC-8:00 (2-5 hrs/day) America/Los_Angeles
AI Quality Analyst (Personalization) -Japanese
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.Key QualificationsJapanese Proficiency: Ability to read and write in Japanese with a high degree of comp, as Japanese is the focus language for this project.Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team.Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities.Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections.Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating.Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.Feedback: Ability to provide constructive feedback and detailed annotations.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Description:In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve:Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences.Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied.Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.Education & ExperienceBS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field).Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.Offer Details:Commitments Required: at least 4 hours per day and up to 20 hours per week with 4 hours of overlap with PST.Engagement type: ContractorEngagement Length: 3 monthsOur offered rate for this project is $15 per hour.Evaluation Process -Shortlisted candidates will be sent a Job Interest Form.After the profile review, an assessment will be shared, which must be completed within 24 hours.Based on the assessment outcomes, shortlisted candidates will be contacted to discuss the pre‑onboarding requirements.
Business Analyst (Chinese Language)
About Turing:Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM, and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:We are looking for candidates with strong analytical and English comprehension skills. The ideal candidate should have the ability to read, summarize, and break down large content into smaller logical blocks, conduct research online, validate claims made in the content through online research, and work with the LLM (Large Language Models) to solve puzzles!Your role is critical in helping fine-tune and improve large language models (like gpt), and will make you an expert on how to leverage AI to be a better analyst. This is your chance to future-proof your career in an AI-first world!What does day-to-day look like:You would spend time answering a variety of interesting analytical questions and creating scenarios that can train the LLM models to get better. Here are a couple of examples which models might get wrong and you would have to give the correct answer and explanation to enable the models to learn:Based on a given distribution of sales by month across locations, could you analyze which location has grown the most? (Hint: what time period should we look at? should we account for sudden variability at the beginning?)In a small town, there are four distinct neighborhoods: Oak, Pine, Maple, and Elm. A postman is assigned to deliver mail and can only deliver to two neighborhoods in one day, with certain rules (e.g. Oak is always visited before Pine). If he delivers to Oak and Elm on the first day, which neighborhoods does the carrier deliver to on the second day?Note: No other prior specialized domain experience is needed.Requirements:Multilingual capabilities are mandatory. (English and Chinese)Analytical Skills: Good research and analytical skillsFeedback Skills: Ability to provide constructive feedback and detailed annotations.Creative Thinking: Creative and lateral thinking abilities.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Commitment: Ability to commit to 40 hours per week for the contract duration overlapping US hours.Technical Setup: Desktop/Laptop set up with a good internet connection.Preferred Qualifications:Bachelors degree or undergraduate in Engineering, Literature, Journalism, Communications, Arts, Statistics, or a related field. We are open to candidates who do not have a Bachelor's degree but have experience in the area.Experience writing professionally (business analysts, research analyst, copywriter, journalist, technical writer, editor, translator, etc.)Understanding of Excel and Google Suite.Proficiency in Data interpretation, Logical reasoning and Basic arithmetic is highly desirable.Benefits:Competitive compensation based on experience and expertise.Flexible working hours and remote work environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:Engagement type : Contractor assignment/freelancer (no medical/paid leave)This role will require some overlap with UTC-8:00 (2-5 hrs/day) America/Los_Angeles
AI Quality Analyst (Personalization) - Spanish
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.Key QualificationsSpanish Proficiency: Ability to read and write in Spanish with a high degree of comp, as Spanish is the focus language for this project.Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team.Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities.Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections.Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating.Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.Feedback: Ability to provide constructive feedback and detailed annotations.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Description:In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve:Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences.Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied.Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.Education & ExperienceBS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field).Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.Offer Details:Commitments Required: at least 4 hours per day and upto 40 hours per week with 4 hours of overlap with PST.Engagement type: ContractorEngagement Length: 3 monthsOur offered rate for this project is $15 per hour.Evaluation Process -Shortlisted candidates will be sent a Job Interest Form.After the profile review, an assessment will be shared, which must be completed within 24 hours.Based on the assessment outcomes, shortlisted candidates will be contacted to discuss the pre‑onboarding requirements.
Ph.D. / Postdoctoral / Master’s Expert
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:In this role, you will be working on projects to help fine-tune large language models (like ChatGPT) using your strong analytical and English comprehension skills.The ideal candidate should have a solid foundation in STEM, particularly at the level expected in engineering entrance exams, as well as graduate or PhD-level programs.You should be able to break down complex STEM concepts into simple, clear explanations and work efficiently.The projects will also help you learn how to leverage AI to be a better analyst. This is your chance to future-proof your career in an AI-first world!What does day-to-day look like:Design and solve challenging STEMs problems to probe the limitations of large language models.Create clear, high-quality, step-by-step solutions with well-articulated reasoning.Collaborate with LLM researchers to align problems with evaluation goals, especially in areas where models typically struggle (e.g., abstraction, multi-step reasoning, symbolic manipulation).Help define new evaluation benchmarks based on Physics curricula spanning early undergraduate to PhD-level topics.Requirements:Analytical Skills: Good research and analytical skillsFeedback Skills: Ability to provide constructive feedback and detailed annotations.Creative Thinking: Creative and lateral thinking abilities.Communication: Excellent structured communication and collaboration skills in a remote setting.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Prefered Qualification:Candidates currently pursuing a Master’s/Ph.D./Postdoctoral degree in STEM, Applied Physics, or a related field are eligible and encouraged to apply.Ability to analyze and solve complex physics problems with a structured and logical approach.Ability to explain STEM concepts clearly using simple language, visuals, and physics reasoning.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:Engagement type : Contractor assignment/freelancer (no medical/paid leave)
Dockerfile Data Validation Engineer
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.About the Role:We are seeking an engineer responsible for designing, implementing, and maintaining data-validation workflows inside Docker-based build pipelines. This role involves creating and managing Dockerfile labels, metadata standards, and validation scripts that ensure datasets, schemas, and model artifacts meet quality and compliance requirements before deployment.You will work closely with data engineering, machine learning, and DevOps teams to build reliable, reproducible, and fully validated containerized data pipelines.What does day-to-day look like:Develop and optimize Dockerfiles with built-in data-validation steps.Implement LABEL metadata for dataset versions, schemas, and lineage.Create validation scripts (Python/Bash) for schema checks, data integrity, and quality control.Integrate validation steps into CI/CD pipelines and enforce fail-on-bad-data checks.Document standards for Dockerfile labeling, validation logic, and data governance.Required SkillsExperienced DevOps engineers with 4+ years of experienceStrong experience with Docker & Dockerfiles.Proficiency in Python or Bash for validation scripting.Knowledge of data formats, schemas, and validation tools.Familiarity with CI/CD systems and container registries.Nice to Have:Previous participation in LLM research or evaluation projects.Experience building or testing developer tools or automation agents.Experience with MLOps workflows, data versioning, or Great Expectations.Knowledge of Kubernetes or container security tools.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Employment type : Contractor assignment (no medical/paid leave)Duration of contract : 2-4 weeks; [expected start date is next week]Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Mexico, BrazilEvaluation Process (approximately 75 mins) :Interviews (30-60 min technical discussion in QODE)
ML/AI/Stats Experts
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Requirements:PhD in Machine Learning, AI, Computer Science, Statistics, or a closely related field2+ publications in top-tier ML/AI conferences such as:NeurIPSICMLICLRStrong research experience in areas such as:Deep learningGenerative AI / LLMsRepresentation learningReinforcement learningOptimizationComputer vision or NLP, depending on the taskExperience with reading, understanding, and critically evaluating research papersAbility to assess whether an ML solution is technically correct, methodologically sound, and aligned with current researchStrong mathematical and statistical foundationsAbility to start immediatelyExcellent structured communication and collaboration skills in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connectionNice to haveExperience as a research scientist, ML researcher, PhD researcher, or university faculty memberStrong Python/PyTorch/TensorFlow skillsRole Overview:Evaluate AI-generated paper reproductionsFor each task, raters will read the paper, inspect the agent’s reproduction artifacts, rank reproduction attempts, and provide written justificationsPerks of Freelancing With Turing:Work in a fully remote environmentOpportunity to work on cutting-edge AI projects with leading LLM companiesPotential for contract extension based on performance and project needsOffer Details:Engagement type: Contractor assignment/freelancer (no medical/paid leave)
Research Analyst - Advanced Math
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:In this role, you will be working on projects to help fine-tune large language models (like ChatGPT) using your strong analytical and English comprehension skills. The ideal candidate should have the ability to read, summarize, and break down large content into smaller logical blocks, conduct research online, validate claims made in content through online research, and work with the LLM (Large Language Models) to solve puzzles! The projects will also help you learn how to leverage AI to be a better analyst. This is your chance to future-proof your career in an AI-first world!Your role is critical in helping fine-tune and improve large language models (like gpt), and will make you an expert on how to leverage AI to be a better analyst. This is your chance to future-proof your career in an AI-first world!What does day-to-day look like:You would spend time answering a variety of interesting analytical questions and creating scenarios that can train the LLM models to get better. Here are a couple of examples that models might get wrong, and You would articulate the correct answer & explanation to enable the models to learn:Based on a given distribution of sales by month across locations, could you analyze which location has grown the most? (Hint: what time period should we look at? should we account for sudden variability at the beginning?)In a small town, there are four distinct neighborhoods: Oak, Pine, Maple, and Elm. A postman is assigned to deliver mail and can only deliver to two neighborhoods in one day, with certain rules (e.g. Oak is always visited before Pine). If he delivers to Oak and Elm on the first day, which neighborhoods does the carrier deliver to on the second day?Note: No other prior specialized domain experience is needed.Requirements:English Proficiency: Ability to read and write in English with a high degree of comprehension skills.Mathematical Proficiency: Strong skills in algebra, geometry, trigonometry, and calculus at the high school/college level.Problem-Solving Skills: Ability to approach complex problems systematically and creatively.Explanation Skills: Capacity to break down solutions into clear, understandable steps.Time Management: Ability to solve problems efficiently and meet deadlines.Communication: Excellent written and verbal communication skills for explaining mathematical concepts.Adaptability: Flexibility to work with various types of math problems and student needs.Commitment: Ability to work full-time, 40 hours per week.Technical Setup: Desktop/Laptop with reliable internet connection and necessary software for mathematical computations and online collaboration.Preferred Qualifications:Bachelor's degree in Mathematics, Engineering, Physics, or a related field. We are open to candidates who have demonstrated exceptional math skills without a formal degree.Experience in tutoring or teaching mathematics at the high school or Engineering Entrance Exam or Engineering College level.Familiarity with standardized test formats and requirements for Engineering Entrance Exam or Engineering College level. Knowledge of mathematical software and tools is a plus (e.g., LaTeX, Google Colab).Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:Commitments Required : at least 4 hours per day and minimum 20 hours per week with 4 hours of overlap with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Engagement type : Contractor assignment/freelancer (no medical/paid leave)Duration of contract : 1 month; [expected start date is next week]This role will require some overlap with UTC-8:00 (2-5 hrs/day) America/Los_AngelesEvaluation Process (approximately 80-90 mins) :Shortlisted candidates will be sent an automated analytical challenge (approximately 30 mins)Once you clear the challenge, you will be redirected to a business writing assessment (English - approximately 30 minutes) followed by a language proficiency assessment (approximately 20 minutes).Once you clear these online assessments, you are ready to go!
Mathematics Expert (Master’s/Ph.D.)
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable on projects that improve and evaluate large language models by applying advanced mathematical reasoning, problem-solving, computational thinking, and clear written communication impact, and lasting results on the P&L.Role Overview.In this role, you will work on projects that improve and evaluate large language models by applying advanced mathematical reasoning, problem-solving, computational thinking, and clear written communication.The ideal candidate should have a solid foundation in mathematics, particularly at the level expected in engineering entrance exams, as well as graduate or PhD-level programs.You should be able to break down complex mathematical concepts into simple, clear explanations and work efficiently.The role also includes computational projects where you design precise, closed-ended prompts, write reliable Python code, verify numerical answers, and provide clear rationales.The projects will also help you learn how to leverage AI to be a better analyst. This is your chance to future-proof your career in an AI-first world!What does day-to-day look like:Design original and challenging mathematics problems that test the reasoning limits of large language models, especially in multi-step, abstract, and proof-based settings.Solve problems independently and write detailed, logically structured solutions with clear justifications.Review model-generated solutions, identify mathematical errors or missing arguments, and provide precise feedback, annotations, and corrections..Contribute to defining new evaluation benchmarks based on Mathematics curricula spanning from early undergraduate to PhD-level topics.Develop and validate Python-based solutions for computational tasks using approved scientific libraries.Work on theorem-prover tasks using Lean, including translating mathematical problems and proofs into formal language, verifying that formal proofs compile correctly.Requirements:Analytical Skills: Good research and analytical skillsFeedback Skills: Ability to provide constructive feedback and detailed annotations.Creative Thinking: Creative and lateral thinking abilities.Communication: Excellent structured communication and collaboration skills in a remote setting.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Preferred Qualifications:Candidates pursuing a Master’s/Ph.D./Postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field are eligible and encouraged to apply.Ability to analyze and solve complex math problems with a structured and logical approach.Ability to explain math concepts clearly using simple language, visuals, and examples.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:Commitments Required: at least 4 hours per day and a minimum of 20 hours per week, with 4 hours of overlap with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week, or 40 hrs/week)Engagement type: Contractor assignment/freelancer (no medical/paid leave)
Senior Software Engineer – Ruby (LLM Evaluation & Repository Validation)
Job Title: Senior Software Engineer – LLM Evaluation & Repository ValidationAbout the projects: we are building LLM evaluation and training datasets to train LLM to work on realistic software engineering problems. One of our approaches, in this project, is to build verifiable SWE tasks based on public repository histories in a synthetic approach with human-in-the-loop; while expanding the dataset coverage to different types of tasks in terms of programming language, difficulty level, and etc.About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and qualityWhy Join Us? Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems. You’ll be at the forefront of evaluating how LLMs interact with real code, influencing the future of AI-assisted software development. This is a unique opportunity to blend practical software engineering with AI research.What does day-to-day look like:Analyze and triage GitHub issues across trending open-source libraries.Set up and configure code repositories, including Dockerization and environment setup.Evaluating unit test coverage and quality.Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs.Opportunities to lead a team of junior engineers to collaborate on projects.Required Skills:Minimum 3+ years of overall experienceStrong experience with at least one of the following languages: RubyProficiency with Git, Docker, and basic software pipeline setup.Ability to understand and navigate complex codebases.Comfortable running, modifying, and testing real-world projects locally.Experience contributing to or evaluating open-source projects is a plus.Nice to Have:Previous participation in LLM research or evaluation projects.Experience building or testing developer tools or automation agents.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Employment type : Contractor assignment (no medical/paid leave)Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, MexicoEvaluation Process (approximately 75 mins) :Two rounds of interviews (60 min technical + 30 min technical & cultural discussion)
Finance Expert (US based)
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole overview:Turing is looking for Experts in finance to work with our researchers to improve performance of AI models. We are looking for experts across a range of topics including capital markets, portfolio management, research, trading, quant, investment banking, private equity, corporate finance, accounting, and others. If you enjoy solving complex problems in finance and are interested in working with AI systems, please apply. No prior AI experience is required.What does day-to-day look like:Evaluate LLM models for areas of finance where models do not perform well.Create rubrics to assess model capabilities on specific areas of your finance expertise (such as deal analysis, M&A assessments, and more).Collaborate with AI researchers and fellow finance experts to shape training methods, evaluation strategies, and benchmarks.Requirements:2+ years experience in Capital Markets, Portfolio Management, Research, Trading, Quant, Investment Banking, Private Equity, Venture Capital, Growth Equity, FP&A, Accounting, or Financial Consulting.Strong grasp of financial concepts (investment analysis, resarch, forecasting, revenue builds, corporate finance, asset management, risk management, etc.) based on your domain of expertise.Excellent English written communication.Bonuses (not at all necessary):CFA (Level I/II/III) or CA/CPA/MBA in Finance.Perks of freelancing with Turing:Work on the cutting edge of AI and finance.Fully remote and flexible work environment.Competitive hourly compensation of ~$100+/hour depending on experience.Offer Details:Commitment: Flexible, 10–30 hrs/week.Duration: ~1 month, with the possibility of extension based on performance and project needs.
MLE Bench – Data Analyst
About TuringBased in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole OverviewWe are looking for experienced Data Analysts (MLE Bench) to contribute to benchmark-driven evaluation projects focused on real-world machine learning systems. This role involves hands-on analytical work with production-like datasets, metrics, and ML outputs to help evaluate, diagnose, and improve the performance of advanced AI systems.The ideal candidate is comfortable working at the intersection of data analysis and machine learning, with strong analytical rigor and the ability to work with real datasets and ML evaluation workflows.What does day-to-day life look like?Analyze structured and unstructured datasets generated from ML training, inference, and evaluation pipelines.Define, compute, and validate metrics used to evaluate model performance and behavior.Investigate data distributions, model outputs, failure modes, and edge cases relevant to benchmark tasks.Write and run Python and SQL code to analyze data, create reports, and support evaluation workflows.Validate data quality, consistency, and correctness across datasets and experiments.Create clear, well-documented analytical artifacts and reproducible analysis workflows.Collaborate with ML engineers and researchers to design challenging, real-world evaluation scenarios for MLE Bench.RequirementsMinimum 3+ years of experience as a Data Analyst or Analytics-focused Engineer.Strong proficiency in Python for data analysis.Solid experience with SQL and relational datasets.Experience analyzing ML outputs and evaluation metrics.Strong understanding of statistics and analytical reasoning.Ability to work with large, complex datasets and draw reliable insights.Experience writing clean, readable, and well-documented analytical code.Excellent spoken and written English communication skills.Perks of Freelancing With TuringWork in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer DetailsCommitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST.Engagement Type: Contractor assignment (no medical/paid leave)Duration of Contract: 3 months (adjustable based on engagement)Location: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, MexicoEvaluation ProcessTechnical Interview with live coding challege (60 mins)
SWE Bench – Data Engineer/Data Scientist
About TuringBased in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole OverviewWe are looking for experienced Software Engineers (SWE Bench – Data Engineer / Data Science) to contribute to benchmark-driven evaluation projects focused on real-world data engineering and data science workflows. This role involves hands-on work with production-like datasets, data pipelines, and data science tasks to help evaluate and improve the performance of advanced AI systems.The ideal candidate has strong foundations in data engineering and data science, with the ability to work across data preparation, analysis, and model-related workflows in real-world codebases.What does day-to-day life look like?Work with structured and unstructured datasets to support SWE Bench-style evaluation tasks.Design, build, and validate data pipelines used in benchmarking and evaluation workflows.Perform data processing, analysis, feature preparation, and validation for data science use cases.Write, run, and modify Python code to process data and support experiments locally.Evaluate data quality, transformations, and outputs for correctness and reproducibility.Create clean, well-documented, and reusable data workflows suitable for benchmarking.Participate in code reviews to ensure high standards of code quality and maintainability.Collaborate with researchers and engineers to design challenging, real-world data engineering and data science tasks for AI systems.RequirementsMinimum 3+ years of overall experience as a Data Engineer, Data Scientist, or Software Engineer (data-focused).Strong proficiency in Python for data engineering and data science workflows.Demonstrable experience with data processing, analysis, and model-related workflows.Solid understanding of machine learning and data science fundamentals.Experience working with structured and unstructured data.Ability to understand, navigate, and modify complex, real-world codebases.Experience writing readable, reusable, maintainable, and well-documented code.Strong problem-solving skills, including experience with algorithmic or data-intensive problems.Excellent spoken and written English communication skills.Perks of Freelancing With TuringWork in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer DetailsCommitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST.Engagement Type: Contractor assignment (no medical/paid leave)Duration of Contract: 3 months (adjustable based on engagement)Location: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, MexicoEvaluation ProcessTechnical Interview with live coding challege (60 mins)
Data Scientist/Analyst
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:We are actively seeking talented Data Scientists & Analysts proficient in Python to join our ambitious team dedicated to pushing the frontiers of AI technology. This opportunity is tailored for professionals who thrive on developing innovative solutions and aspire to be at the forefront of AI advancements. You will work with companies in the US looking to develop cutting-edge commercial and research AI solutions. You will write effective Python code to tackle complex issues but also use your business sense and analytical abilities to glean valuable insights from public databases, fix bugs in the code, and create thorough documentation. The ideal candidate will communicate clearly with researchers and help the organization in realizing its objectives, and clearly express the reasoning and logic when writing code in Jupyter notebooks, or other suitable mediums. It will also utilize extensive data analysis skills to develop and respond to important business queries using available datasets (such as those from Kaggle, the UN, the US government, etc.) What does day-to-day look like:Design, develop, and maintain efficient, high-quality code to train and optimize AI models.Conduct evaluations (Evals) to benchmark model performance and analyze results for continuous improvement.Evaluate and rank AI model responses to user queries across diverse domains, ensuring alignment with predefined criteria.Develop comprehensive explanations and rationales for evaluations, showcasing excellent reasoning and technical expertise.Lead efforts in Supervised Fine-Tuning (SFT), including creating and maintaining high-quality, task-specific datasets.Collaborate with researchers and annotators to execute Reinforcement Learning with Human Feedback (RLHF) and refine reward models.Design innovative evaluation strategies and processes to improve the model's alignment with user needs and ethical guidelines.Create and refine optimal responses to improve AI performance, emphasizing clarity, relevance, and technical accuracy.Conduct thorough peer reviews of code and documentation, providing constructive feedback and identifying areas for improvement.Collaborate with cross-functional teams to improve model performance and contribute to product enhancements.Continuously explore and integrate new tools, techniques, and methodologies to enhance AI training processes.Required SkillsBachelor’s/Master’s degree in Engineering, Computer Science (or equivalent experience) A desire to have a significant impact on the field of artificial intelligence Strong data analytic abilities and business sense are required to draw the appropriate conclusions from the dataset, respond to those conclusions, and clearly convey the key findings Excellent problem-solving and analytical skills Excellent communication abilities to work with stakeholders and researchers successfully Fluent in conversational and written English communication skills Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: at least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Engagement type : Contractor assignment (no medical/paid leave)Duration of contract : 1 month; [expected start date is next week]Evaluation Process (approximately 75 mins) :Two rounds of interviews (60 min technical + 15 min cultural &, offer discussion)
Senior Software Engineer – Go (LLM Evaluation & Repository Validation)
Job Title: Senior Software Engineer – LLM Evaluation & Repository ValidationAbout the projects: we are building LLM evaluation and training datasets to train LLM to work on realistic software engineering problems. One of our approaches, in this project, is to build verifiable SWE tasks based on public repository histories in a synthetic approach with human-in-the-loop; while expanding the dataset coverage to different types of tasks in terms of programming language, difficulty level, and etc.About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and qualityWhy Join Us? Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems. You’ll be at the forefront of evaluating how LLMs interact with real code, influencing the future of AI-assisted software development. This is a unique opportunity to blend practical software engineering with AI research.What does day-to-day look like:Analyze and triage GitHub issues across trending open-source libraries.Set up and configure code repositories, including Dockerization and environment setup.Evaluating unit test coverage and quality.Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs.Opportunities to lead a team of junior engineers to collaborate on projects.Required Skills:Minimum 3+ years of overall experienceStrong experience with at least one of the following languages: GoProficiency with Git, Docker, and basic software pipeline setup.Ability to understand and navigate complex codebases.Comfortable running, modifying, and testing real-world projects locally.Experience contributing to or evaluating open-source projects is a plus.Nice to Have:Previous participation in LLM research or evaluation projects.Experience building or testing developer tools or automation agents.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Employment type : Contractor assignment (no medical/paid leave)Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, MexicoEvaluation Process (approximately 75 mins) :Two rounds of interviews (60 min technical + 30 min technical & cultural discussion)
Senior Software Engineer – Rust (LLM Evaluation & Repository Validation)
Job Title: Senior Software Engineer – LLM Evaluation & Repository ValidationAbout the projects: we are building LLM evaluation and training datasets to train LLM to work on realistic software engineering problems. One of our approaches, in this project, is to build verifiable SWE tasks based on public repository histories in a synthetic approach with human-in-the-loop; while expanding the dataset coverage to different types of tasks in terms of programming language, difficulty level, and etc.About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and qualityWhy Join Us? Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems. You’ll be at the forefront of evaluating how LLMs interact with real code, influencing the future of AI-assisted software development. This is a unique opportunity to blend practical software engineering with AI research.What does day-to-day look like:Analyze and triage GitHub issues across trending open-source libraries.Set up and configure code repositories, including Dockerization and environment setup.Evaluating unit test coverage and quality.Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs.Opportunities to lead a team of junior engineers to collaborate on projects.Required Skills:Minimum 3+ years of overall experienceStrong experience with at least one of the following languages: RustProficiency with Git, Docker, and basic software pipeline setup.Ability to understand and navigate complex codebases.Comfortable running, modifying, and testing real-world projects locally.Experience contributing to or evaluating open-source projects is a plus.Nice to Have:Previous participation in LLM research or evaluation projects.Experience building or testing developer tools or automation agents.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Employment type : Contractor assignment (no medical/paid leave)Duration of contract : 3 month; [expected start date is next week]Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, MexicoEvaluation Process (approximately 75 mins) :Two rounds of interviews (60 min technical + 30 min technical & cultural discussion)
Business Analyst (French Language)
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:We are looking for candidates with strong analytical and English comprehension skills. The ideal candidate should have the ability to read, summarize, and break down large content into smaller logical blocks, conduct research online, validate claims made in the content through online research, and work with the LLM (Large Language Models) to solve puzzles!Your role is critical in helping fine-tune and improve large language models (like gpt), and will make you an expert on how to leverage AI to be a better analyst. This is your chance to future-proof your career in an AI-first world!What does day-to-day look like:You would spend time answering a variety of interesting analytical questions and creating scenarios that can train the LLM models to get better. Here are a couple of examples which models might get wrong and you would have to give the correct answer and explanation to enable the models to learn:Based on a given distribution of sales by month across locations, could you analyze which location has grown the most? (Hint: what time period should we look at? should we account for sudden variability at the beginning?)In a small town, there are four distinct neighborhoods: Oak, Pine, Maple, and Elm. A postman is assigned to deliver mail and can only deliver to two neighborhoods in one day, with certain rules (e.g. Oak is always visited before Pine). If he delivers to Oak and Elm on the first day, which neighborhoods does the carrier deliver to on the second day?Note: No other prior specialized domain experience is needed.Requirements:Multilingual capabilities are mandatory. (English and French)Analytical Skills: Good research and analytical skillsFeedback Skills: Ability to provide constructive feedback and detailed annotations.Creative Thinking: Creative and lateral thinking abilities.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Commitment: Ability to commit to 40 hours per week for the contract duration overlapping US hours.Technical Setup: Desktop/Laptop set up with a good internet connection.Preferred Qualifications:Bachelors degree or undergraduate in Engineering, Literature, Journalism, Communications, Arts, Statistics, or a related field. We are open to candidates who do not have a Bachelor's degree but have experience in the area.Experience writing professionally (business analysts, research analyst, copywriter, journalist, technical writer, editor, translator, etc.)Understanding of Excel and Google Suite.Proficiency in Data interpretation, Logical reasoning and Basic arithmetic is highly desirable.Benefits:Competitive compensation based on experience and expertise.Flexible working hours and remote work environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:Engagement type : Contractor assignment/freelancer (no medical/paid leave)This role will require some overlap with UTC-8:00 (2-5 hrs/day) America/Los_Angeles
Senior Software Engineer – C#(LLM Evaluation & Repository Validation)
About the projects: we are building LLM evaluation and training datasets to train LLM to work on realistic software engineering problems. One of our approaches, in this project, is to build verifiable SWE tasks based on public repository histories in a synthetic approach with human-in-the-loop; while expanding the dataset coverage to different types of tasks in terms of programming language, difficulty level, and etc.About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and qualityWhy Join Us? Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems. You’ll be at the forefront of evaluating how LLMs interact with real code, influencing the future of AI-assisted software development. This is a unique opportunity to blend practical software engineering with AI research.What does day-to-day look like:Analyze and triage GitHub issues across trending open-source libraries.Set up and configure code repositories, including Dockerization and environment setup.Evaluating unit test coverage and quality.Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs.Opportunities to lead a team of junior engineers to collaborate on projects.Required Skills:Minimum 3+ years of overall experienceStrong experience with at least one of the following languages: C#Proficiency with Git, Docker, and basic software pipeline setup.Ability to understand and navigate complex codebases.Comfortable running, modifying, and testing real-world projects locally.Experience contributing to or evaluating open-source projects is a plus.Nice to Have:Previous participation in LLM research or evaluation projects.Experience building or testing developer tools or automation agents.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Employment type : Contractor assignment (no medical/paid leave)Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, MexicoEvaluation Process (approximately 75 mins) :Two rounds of interviews (60 min technical + 30 min technical & cultural discussion)
Senior LLM Engineer
Senior LLM Engineer – GenAI / ML (Python, Langchain)Full TimeLocation: IndiaOverall Experience: 7–12 YearsFocus: Hands-on engineering role focused on designing, building, and deploying Generative AI and LLM-based solutions. The role requires deep technical proficiency in Python and modern LLM frameworks with the ability to contribute to roadmap development and cross-functional collaboration.Key Responsibilities:Design and develop GenAI/LLM-based systems using tools such as Langchain and Retrieval-Augmented Generation (RAG) pipelines.Implement prompt engineering techniques and agent-based frameworks to deliver intelligent, context-aware solutions.Collaborate with the engineering team to shape and drive the technical roadmap for LLM initiatives.Translate business needs into scalable, production-ready AI solutions.Work closely with business SMEs and data teams to ensure alignment of AI models with real-world use cases.Contribute to architecture discussions, code reviews, and performance optimization.Skills Required:Proficient in Python, Langchain, and SQL.Understanding of LLM internals, including prompt tuning, embeddings, vector databases, and agent workflows.Background in machine learning or software engineering with a focus on system-level thinking.Experience working with cloud platforms like AWS, Azure, or GCP.Ability to work independently while collaborating effectively across teams.Excellent communication and stakeholder management skills.Preferred Qualifications:1+ years of hands-on experience in LLMs and Generative AI techniques.Experience contributing to ML/AI product pipelines or end-to-end deployments.Familiarity with MLOps and scalable deployment patterns for AI models.Prior exposure to client-facing projects or cross-functional AI teams.
Scientific Computing SME (Physics)
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole Overview: In this role, you will be working on projects to help fine-tune large language models (like ChatGPT) using your strong analytical and english comprehension skills. The ideal candidate should have a solid foundation in Physics, particularly at the level expected in engineering entrance exams, as well as graduate or PhD-level programs. You should be able to break down complex Physics concepts into simple, clear explanations and work efficiently. The projects will also help you learn how to leverage AI to be a better analyst. This is your chance to future-proof your career in an AI-first world! What does day-to-day look like: Design and solve challenging Physics problems to probe the limitations of large language models. Create clear, high-quality, step-by-step solutions with well-articulated reasoning. Collaborate with LLM researchers to align problems with evaluation goals, especially in areas where models typically struggle (e.g., abstraction, multi-step reasoning, symbolic manipulation).Help define new evaluation benchmarks based on Physics curricula spanning early undergraduate to PhD-level topics. Requirements: Analytical Skills: Good research and analytical skills Feedback Skills: Ability to provide constructive feedback and detailed annotations. Creative Thinking: Creative and lateral thinking abilities. Communication: Excellent structured communication and collaboration skills in a remote setting. Independence: Self-motivated and able to work independently in a remote setting. Technical Setup: Desktop/Laptop set up with a good internet connection. Prefered Qualification: Candidates currently pursuing a Master’s/Ph.D./Postdoctoral degree in Physics, Applied Physics, or a related field are eligible and encouraged to apply. Ability to analyze and solve complex physics problems with a structured and logical approach. Ability to explain physics concepts clearly using simple language, visuals, and physics reasoningTechnical Skills• Strong proficiency in Python for scientific computing• Experience with scientific libraries: NumPy, SciPy, SymPy, Pandas, Matplotlib • Understanding of numerical methods, algorithms, and computational complexity • Ability to design problems with precise, verifiable numerical outcomesPerks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:Commitments Required : At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Engagement type : Contractor assignment/freelancer (no medical/paid leave)Duration of contract : 6weeks; [expected start date is next week]
Medicine Physician (MD/DO/Doctoral study/PhD)
About UsBased in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.What does day-to-day look like:Work with research teams to evaluate and improve how AI systems handle clinical reasoning. You'll design evaluation methods that test AI performance on real medical problems where clinical expertise makes the difference.Design systematic evaluation frameworks for medical AI systemsCreate clinical scenarios that test AI reasoning and decision-making capabilitiesBuild assessment methods that capture the nuance of clinical practiceIdentify gaps in AI medical knowledge and reasoningCollaborate with AI researchers to improve model performanceCandidate Requirements:Licensed physician in active clinical practice (any specialty)Experience with clinical decision-making and evidence-based medicineInterest in how AI can support clinical practiceStrong analytical and communication skillsWhy it Matters:Help shape the next generation of medical AI by ensuring these systems can handle the complexity of real clinical practice. Remote work with flexible scheduling around your clinical commitments.Engagement details:Commitment: flexible engagement, remote working, up to 30 hrs/weekDuration: 1 month, with potential extensions based on performance and fit
Azure DevOps Engineer(LangGraph)
Yoe: 8+Work mode: remote/IndiaAvailability: Immediate-2 weeksJob Title: Mid-Senior Azure DevOps Engineer (LangGraph Experience)Job Description:We are seeking a Mid-Senior Azure DevOps Engineer with strong experience in Azure DevOps and hands-on expertise in LangGraph to support the deployment, automation, and management of AI-powered applications. The ideal candidate will have experience building scalable CI/CD pipelines, managing Azure cloud infrastructure, and deploying LLM-based agentic applications.Key Responsibilities:Design, implement, and optimize CI/CD pipelines using Azure DevOps.Deploy, configure, and maintain applications and AI workflows on Microsoft Azure.Build, deploy, and support LangGraph-based agentic AI applications.Automate infrastructure provisioning using Terraform, Bicep, or ARM templates.Manage containerized applications using Docker and Kubernetes (AKS).Monitor application health, troubleshoot production issues, and optimize cloud resources.Collaborate with development, AI/ML, and platform teams to ensure reliable software delivery.Implement DevSecOps best practices, including security, monitoring, and compliance.Required Skills:8+ years of experience in Azure DevOps or Cloud DevOps engineering.Hands-on experience with LangGraph for building and deploying AI agent workflows.Strong experience with Azure services such as App Service, Azure Functions, AKS, Azure Container Registry, Azure Storage, and Azure Key Vault.Experience designing and maintaining Azure DevOps CI/CD pipelines using YAML.Proficiency with Infrastructure as Code using Terraform, Bicep, or ARM templates.Strong knowledge of Docker, Kubernetes, and container orchestration.Proficiency in Python or PowerShell for automation.Experience with Git, branching strategies, and release management.Familiarity with Azure Monitor, Application Insights, and Log Analytics.Preferred Skills:Experience with LangChain, Azure OpenAI Service, or other LLM frameworks.Knowledge of vector databases, Retrieval-Augmented Generation (RAG), and AI application deployment.Azure certifications such as AZ-104, AZ-400, or AZ-305.Experience implementing DevSecOps and cloud security best practices.Qualifications:Bachelor's degree in Computer Science, Information Technology, or a related field.Excellent problem-solving, communication, and collaboration skills.Experience working in Agile/Scrum environments.
Engineering Manager
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole Overview:As a Delivery Leader for our LLM Training business, you will play a pivotal role in leading our efforts to develop high-quality, foundational LLMs. This position requires managing large teams of software engineers and data scientists dedicated to performing various Supervised Fine Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) tasks. Your leadership will ensure the delivery of superior SFT and RLHF datasets, achieving optimal outcomes in terms of quality, throughput, and cost.What does day-to-day look like:Lead and manage large teams (20+) of Python, JavaScript, Java (and other long tail of languages) developers to execute technical tasks effectively.Collaborate closely with researcher clients, ensuring their satisfaction with the delivered work.Maintain rigorous review processes to ensure the highest quality of datasets.Oversee the transition of projects from initiation to stable states, taking full responsibility for deliveryIdentify and implement measures to enhance the quality of Python/JavaScript/Java etc. code within the team.Ensure the team's work is of high quality, addressing any issues related to the clarity of questions, methodology, result communication, or bugs.Requirements:8+ years of professional software engineering experience, including 3+ years in an engineering management role.Proven experience in managing large technical teams in a delivery-oriented role with at least one of the following, ideally both:Python or JavaJavascriptDemonstrated ability to engage in hands-on technical work, identifying and resolving quality issues.Excellent leadership and people management skills, with a focus on motivating teams and fostering a collaborative work environment.Strong communication and stakeholder management abilities, capable of effectively collaborating with clients and internal partners.Why join Us?:Our company is at the cutting edge of AI and Machine Learning, offering unique opportunities to contribute to the advancement of LLMs. You will lead a team of talented individuals, collaborate with top-tier clients, and be part of a dynamic and innovative culture. We are committed to the professional growth, providing a supportive environment where you can thrive and make a significant impact in the AI industry.Offer Details:Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Employment type : Contractor assignment (no medical/paid leave)Duration of contract : 10 monthsLocation : India & LATAM countriesEvaluation Process (approximately 120 mins) :Two rounds of interviews (60 min technical + 60 min leadership)
AI Quality Analyst - English
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.Key QualificationsEnglish Proficiency: Ability to read and write in English with a high degree of comp, as English is the focus language for this project.Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team.Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities.Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections.Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating.Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.Feedback: Ability to provide constructive feedback and detailed annotations.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Description:In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve:Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences.Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied.Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.Education & ExperienceBS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field).Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.Offer Details:Commitments Required: at least 4 hours per day and upto 40 hours per week with 4 hours of overlap with PST.Engagement type: ContractorEngagement Length: 3 monthsEvaluation Process -Shortlisted candidates will be sent a Job Interest Form.After the profile review, an assessment will be shared, which must be completed within 24 hours.Based on the assessment outcomes, shortlisted candidates will be contacted to discuss the pre‑onboarding requirements.
MLE Bench – ML Engineers
About TuringBased in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole OverviewWe are looking for experienced Machine Learning Engineers (MLE Bench) to contribute to benchmark-driven evaluation projects focused on real-world machine learning systems. This role involves hands-on work with production-grade ML codebases, model training and evaluation pipelines, and deployment-oriented workflows to help assess and improve the capabilities of advanced AI systems.The ideal candidate is comfortable bridging research and engineering, working deeply with models, data, and infrastructure in realistic ML environments.What does day-to-day life look like?Work with real-world ML codebases to support MLE Bench–style evaluation tasks.Build, run, and modify model training, evaluation, and inference pipelines.Prepare datasets, features, and metrics for ML benchmarking and validation.Debug, refactor, and improve production-like ML systems for correctness and performance.Evaluate model behavior, failure modes, and edge cases relevant to benchmark tasks.Write clean, reproducible, and well-documented Python code for ML workflows.Participate in code reviews to ensure high standards of engineering quality.Collaborate with researchers and engineers to design challenging, real-world ML engineering tasks for AI system evaluation.RequirementsMinimum 3+ years of overall experience as a Machine Learning Engineer or Software Engineer (ML-focused).Strong proficiency in Python for machine learning and data workflows.Hands-on experience with model training, evaluation, and inference pipelines.Solid understanding of machine learning fundamentals (supervised/unsupervised learning, evaluation metrics, optimization).Experience working with ML frameworks (e.g., PyTorch, TensorFlow, JAX, or similar).Ability to understand, navigate, and modify complex, real-world ML codebases.Experience writing readable, reusable, and maintainable production-quality code.Strong problem-solving and debugging skills.Excellent spoken and written English communication skills.Perks of Freelancing With TuringWork in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer DetailsCommitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST.Engagement Type: Contractor assignment (no medical/paid leave)Duration of Contract: 3 months (adjustable based on engagement)Location: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, MexicoEvaluation ProcessTechnical Interview with live coding challege (60 mins)
Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)
About TuringTuring is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Turing helps leading AI labs improve the reasoning, problem-solving, and decision-making capabilities of large language models (LLMs) through high-quality human feedback, evaluation, and training data.Role OverviewWe are looking for experienced Software Engineers to help evaluate, benchmark, and improve the coding capabilities of frontier AI models. In this role, you will assess AI-generated code, validate solutions against real-world software engineering tasks, identify correctness and quality issues, and contribute to the development of high-quality evaluation datasets and benchmarks.This position is ideal for engineers who enjoy code review, debugging, problem-solving, and applying strong software engineering judgment to complex technical scenarios. Your work will directly contribute to measuring and improving the performance of advanced AI coding systems.What Does Day-to-Day Look Like?Review and evaluate AI-generated code for correctness, efficiency, maintainability, and adherence to requirements.Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes.Debug code, reproduce issues, and verify fixes across different programming environments.Assess model-generated explanations, reasoning, and implementation approaches for technical accuracy.Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.Identify edge cases, failure modes, and areas where AI systems struggle with software engineering problems.Document findings clearly and provide structured feedback to improve evaluation quality and consistency.Collaborate with project teams to establish quality standards and evaluation methodologies.RequirementsBachelor's or Master's degree in Computer Science, Software Engineering, or a related technical field.3+ years of professional software engineering experience.Strong proficiency in one or more of the following languages: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL.Strong understanding of data structures, algorithms, software design principles, and debugging methodologies.Experience performing code reviews and evaluating code quality in production or large-scale codebases.Ability to analyze complex technical problems and assess solution correctness with minimal supervision.Familiarity with version control systems (e.g., Git) and modern software development workflows.Strong written communication skills and attention to detail.Experience with AI/ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects is a plus.Experience evaluating AI-generated code, benchmark creation, or software quality assessment is highly preferred.Perks of Freelancing With TuringWork in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer DetailsCommitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST.Engagement type : Contractor assignment (no medical/paid leave)Duration of contract : 1 month; [expected start date is next week]Location: US onlyEvaluation ProcessOnline automated coding challenge for Python and Docker test (RHLF)
AI Quality Analyst (Personalization) - Dutch
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.Key QualificationsDutch Proficiency: Ability to read and write in Dutch with a high degree of comp, as Dutch is the focus language for this project.Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team.Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities.Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections.Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating.Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.Feedback: Ability to provide constructive feedback and detailed annotations.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Description:In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve:Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences.Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied.Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.Education & ExperienceBS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field).Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.Offer Details:Commitments Required: at least 4 hours per day and upto 40 hours per week with 4 hours of overlap with PST.Engagement type: ContractorEngagement Length: 1 month Our offered rate for this project is $20 per hour.Evaluation Process -Shortlisted candidates will be sent a Job Interest Form.After the profile review, an assessment will be shared, which must be completed within 24 hours.Based on the assessment outcomes, shortlisted candidates will be contacted to discuss the pre‑onboarding requirements.
AI Quality Analyst (Personalization) - Arabic
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.Key QualificationsArabic Proficiency: Ability to read and write in Arabic with a high degree of comp, as Arabic is the focus language for this project.Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team.Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities.Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections.Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating.Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.Feedback: Ability to provide constructive feedback and detailed annotations.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Description:In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve:Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences.Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied.Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.Education & ExperienceBS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field).Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.Offer Details:Commitments Required: at least 4 hours per day and upto 40 hours per week with 4 hours of overlap with PST.Engagement type: ContractorEngagement Length: 3 monthsOur offered rate for this project is $15 per hour.Evaluation Process -Shortlisted candidates will be sent a Job Interest Form.After the profile review, an assessment will be shared, which must be completed within 24 hours.Based on the assessment outcomes, shortlisted candidates will be contacted to discuss the pre‑onboarding requirements.
Small business owners (AI response evaluation)
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:Evaluate and compare the quality of responses from multiple AI chatbots across real-world small business use cases.Responsibilities:Create realistic business-related prompts based on defined user goalsInteract with multiple AI chatbots (max. 5 turns per conversation)Assess response quality across clarity, usefulness, and accuracyProvide structured feedback and comparative evaluationsSubmit conversation transcripts and evaluation resultsRequirements:Business owner or strong understanding of small business operationsStrong analytical and critical thinking skillsAbility to follow structured evaluation guidelinesComfortable interacting with AI toolsWhat you'll work on:Create engaging visual content for marketingHelp answer and evaluate situations related to day-to-day operations and customer interactionsConduct market research and contribute ideas in your area of expertiseWork with data to support analysis and financial planningReview and evaluate AI-generated responses for small business use casesUse tools and input files such as spreadsheets, PDFs, and images as part of your workflowPerks of Freelancing With Turing:Work at the forefront of AI applications in accounting and finance.Fully remote and flexible work environment.Opportunity to collaborate on high-impact projects with global reach.Offer Details:Project-based with defined number of evaluation tasksEach task includes multi-chatbot comparison and final assessmentDuration: 10 weeks.
LLM Go Developer
A well-established company that is leveraging the advanced power of technology to help realize the science-fiction fantasy of collaborative and open-ended computer dialogues, is looking for Go develoeprs. The engineer will be working together on the definition, design, and delivery of new features with cross-functional teams. The company is developing the next generation of dialog agents, which will have a wide range of uses in areas including education, entertainment, and general question-answering. This is an exciting opportunity for candidates who are keen to learn in a fast-paced setting.Job Responsibilities:Review the code / solutions generated by an AI system, ensuring adherence to quality standards and best practicesOrganize the development cycle, effectively manage project priorities, and specify goals and deadlinesUse your expertise in Go programming to help resolve difficult coding issues that come up during AI validationCreate a collaborative environment within the team that encourages innovation, communication, and continued improvementVerify the accuracy, efficiency, and dependability of AI-generated code by validating itWork with cross-functional teams to identify strategies for enhancing the AI system's capabilities and integrating it with other elementsAnalyze team members' code and provide constructive feedback to promote a high caliber of software developmentJob Requirements:Bachelor’s/Master’s degree in Engineering, Computer Science (or equivalent experience)At least 3+ years of relevant experience as a software engineerDemonstrated ability to lead, ideally overseeing a group of software engineersIn-depth knowledge of the Go programming languages and the best practices for software developmentDesirable to have some experience with AI systems and code-creation technologiesIdentify and add to public GitHub repositories that need improvement in a proactive manner which can involve developing new features and reworking existing code, and it will have a major influence on projects with large codebases (50K+ lines of code). It will be good to have candidates who have used Github in the past.Train LLM models with high-quality, stable, and scalable back-end components using the newest coding best practices in a variety of languages and frameworks.Ability to collaborate closely with developers, inspiring and setting an example for themOutstanding problem-solving abilities as well as the capacity for critical and strategic thoughtExceptional communication skills, with proficiency in English, both written and verbalOffer DetailsThis is a contractual position.Duration of contract & committed hours are flexible.Interview ProcessTwo internal interviews (60 min technical + 15-30 min cultural and offer conditions discussion).
JavaScript / TypeScript Full-Stack Developer
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:We are actively seeking talented developers proficient in Javascript/Typescript to join our ambitious team dedicated to pushing the frontiers of AI technology. This opportunity is tailored for professionals who thrive on developing innovative solutions and aspire to be at the forefront of AI advancements. You will work with different companies in the US who are looking to develop cutting-edge commercial and research AI solutions.What does day-to-day look like:Design, develop, and maintain efficient, high-quality code to train and optimize AI models.Conduct evaluations (Evals) to benchmark model performance and analyze results for continuous improvement.Evaluate and rank AI model responses to user queries across diverse domains, ensuring alignment with predefined criteria.Develop comprehensive explanations and rationales for evaluations, showcasing excellent reasoning and technical expertise.Lead efforts in Supervised Fine-Tuning (SFT), including creating and maintaining high-quality, task-specific datasets.Collaborate with researchers and annotators to execute Reinforcement Learning with Human Feedback (RLHF) and refine reward models.Design innovative evaluation strategies and processes to improve the model's alignment with user needs and ethical guidelines.Create and refine optimal responses to improve AI performance, emphasizing clarity, relevance, and technical accuracy.Conduct thorough peer reviews of code and documentation, providing constructive feedback and identifying areas for improvement.Collaborate with cross-functional teams to improve model performance and contribute to product enhancements.Continuously explore and integrate new tools, techniques, and methodologies to enhance AI training processes.Required Skills:Write readable, reusable, and maintainable codeParticipate in code reviews to ensure that the standards for code quality are metDemonstrate your proficiency with your language of choice, while covering all basesProvide clear, clean, well-organized, correct, and clearly annotated/classifiable code in the responsesBachelor’s/Master’s degree in Engineering, Computer Science (or equivalent experience)Demonstrable experience with developing web apps using modular development and scalable architectures as well as a strong focus on code readability and security/stability (i.e. testing)Proficiency with the language's syntax and conventionsExtensive experience working with JavaScript or TypeScriptSolid understanding of JavaScript ES6 and either Node.js or React or Nest or Angular or VueNice to have some prior software Quality Assurance and Test Planning experienceExcellent spoken and written English communication skillsDocker experience TechnologiesBack-end : Nest.js, Node.jsFront-end : Vue, Angular, React Languages : JavaScript & TypescriptPerks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: at least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Engagement type : Contractor assignment (no medical/paid leave)Duration of contract : 1 month; [expected start date is next week]Evaluation Process (approximately 75 mins) :Two rounds of interviews (60 min technical + 15 min cultural &, offer discussion)
Prompt & Verifier
About TuringBased in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole OverviewWe are seeking an AI Prompt & Policy Specialist to help design, evaluate, and improve how AI systems interact with tools, APIs, and users across multiple locales. In this role, you will work at the intersection of prompt engineering, policy design, and safety, ensuring that AI-driven workflows are accurate, compliant, and user-friendly.You will use your API/JSON literacy, SQL basics, policy creation skills, and strong critical reasoning to review, refine, and design prompts and behaviors for AI systems interacting with products like Slack, PayPal, and other third-party tools. This role also requires strong communication skills and awareness of safety and content filtering, especially across a wide range of languages and locales.What does day-to-day life look like?Design, review, and refine prompts and instructions for AI systems interacting with tools, APIs, and external services.Read and understand API/JSON schemas and arguments to ensure correct parameter usage and tool behavior.Apply SQL basics to validate data, perform simple checks, and support analysis of AI or tool behavior.Create, document, and maintain policies and guidelines governing AI interactions, content handling, and tool workflows.Develop workflow knowledge for specific tools (e.g., Slack, PayPal and similar platforms) and encode that into prompts and policies.Ensure safety and compliance by applying content filtering guidelines and detecting unsafe or non-compliant outputs.Evaluate AI responses for critical reasoning quality, correctness of parameters, and adherence to policies.Support multilingual use cases by validating prompts and responses across multiple locales, with attention to nuance and clarity.Collaborate with product, engineering, and safety teams to iterate on prompts, policies, and workflows based on evaluation findings.Document edge cases, decision rationales, and recommendations with clear, concise written communication.Requirements3–6 years of experience in any combination of: AI operations, content review, QA, product operations, policy, or technical support roles.Strong API / JSON literacy: ability to read schemas, understand arguments, and verify correct parameter usage.SQL basics: ability to run simple queries to validate data, investigate issues, or support analysis.Demonstrated experience with prompt engineering or writing clear, structured instructions for AI or automation systems.Proven policy creation skills: drafting, interpreting, and applying guidelines or policies in an operational setting.Familiarity with tool-specific workflows (e.g., Slack, PayPal, or similar SaaS platforms), including typical user actions and edge cases.Strong communication skills, with the ability to write clear, concise rationales and explanations; comfort working with multilingual contexts across many locales.High safety awareness for content filtering, including sensitivity to harmful, abusive, or policy-violating content.Excellent critical reasoning and judgment, especially when evaluating ambiguous or complex cases.Strong attention to detail, particularly in checking parameters, arguments, and compliance with specified schemas or policies.Perks of Freelancing With TuringWork in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer DetailsCommitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. Engagement Type: Contractor assignment (no medical/paid leave)Duration of Contract: 3 months (adjustable based on engagement)Location: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, MexicoEvaluation ProcessTechnical Interview for 60 mins
Board Game Reasoning Expert (AI Training & Evaluation)
About TuringTuring is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Turing helps leading AI labs improve the reasoning, problem-solving, and decision-making capabilities of large language models (LLMs) through high-quality human feedback, evaluation, and training data.Role OverviewWe are seeking Board Game Reasoning Experts to help train and evaluate next-generation AI models. In this role, you will leverage your expertise in board games, game mechanics, strategic reasoning, probability, and complex rule systems to create, review, and evaluate high-quality datasets that improve AI performance.You will work on challenging tasks involving game logic, decision-making, strategic planning, rule interpretation, and scenario analysis, helping frontier AI systems develop stronger reasoning capabilities.What Does Day-to-Day Look Like?Create and review game-based reasoning tasks designed to evaluate AI systems.Analyze board game scenarios, strategic decision trees, and rule-based systems.Evaluate AI-generated responses for correctness, consistency, and reasoning quality.Develop prompts, rubrics, and evaluation guidelines for board game and strategy-focused tasks.Identify logical errors, rule violations, and flawed reasoning in AI outputs.Contribute to benchmark creation and quality assurance processes for AI evaluation projects.Collaborate with project teams to improve dataset quality and evaluation methodologies.RequirementsRequired QualificationsBachelor's degree in Computer Science, Mathematics, Cognitive Science, Game Design, Philosophy, Economics, or a related analytical field.2+ years of professional or semi-professional experience in board game design, playtesting, tabletop game communities, or related strategy-focused environments.Strong understanding of logic, probabilistic reasoning, game mechanics, and complex rule systems.Background in game theory, behavioral economics, decision science, or formal logic.Familiarity with modern board games, trading card games (TCGs), tabletop RPGs, strategy games, or competitive game systems.Strong analytical and problem-solving skills.Excellent written communication skills.Preferred QualificationsPrior experience in data annotation, AI training, prompt engineering, QA, game design, puzzle design, or rules-based system analysis.Experience creating evaluation rubrics, benchmark datasets, or quality assurance frameworks.Experience with Python, SQL, or data analysis tools.Experience working on RLHF, model evaluation, synthetic data generation, or LLM benchmarking projects.Familiarity with recently released strategy or tabletop games and the ability to quickly learn new rule systems.Perks of Freelancing With TuringWork in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer DetailsCommitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. Engagement type : Contractor assignment (no medical/paid leave)Duration of contract : 2 months; [expected start date is next week]Evaluation ProcessTake home Assessment
Senior Software Engineer – LLM Evaluation
About TuringBased in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole OverviewAs a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections in Python, JavaScript (including ReactJS), C/C++, Java, Rust, and Go; evaluating and refining AI-generated code for efficiency, scalability, and reliability; and working with cross-functional teams to enhance enterprise-level AI-driven coding solutions.What does day-to-day life look like?Working on AI model training initiatives by curating code examples, building solutions, and correcting code in Python, JavaScript (including ReactJS), C/C++, Java, Rust, and Go.Evaluate and refine AI-generated code to ensure that it is efficient, scalable, and reliable.Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks.Build agents that can verify the quality of the code and identify error patterns.Hypothetize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operation maintenance) and evaluate model capabilities on themDesign verification mechanisms that can automatically verify a solution to a software engineering task.RequirementsSeveral years of software engineering experience, including 2+ years of continuous full-time experience at a top-tier product company (e.g., Google, Stripe, Amazon, Apple, Meta, Netflix, Microsoft, Datadog, Dropbox, Shopify, PayPal, IBM Research).Strong expertise in building full-stack applications and deploying scalable, production-grade software using modern languages and tools.Deep understanding of software architecture, design, development, debugging, and code quality/review assessment.Excellent oral and written communication skills for clear, structured evaluation rationales.Offer DetailsCommitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week (partial PST overlap required)Type: Contractor (no medical/paid leave)Duration:1 month (starting next week; potential extensions based on performance and fit)
AI Quality Analyst (Personalization) - Hindi
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.Key QualificationsHindi Proficiency: Ability to read and write in Hindi with a high degree of comp, as Hindi is the focus language for this project.Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team.Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities.Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections.Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating.Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.Feedback: Ability to provide constructive feedback and detailed annotations.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Description:In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve:Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences.Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied.Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.Education & ExperienceBS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field).Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.Offer Details:Commitments Required: at least 4 hours per day and upto 40 hours per week with 4 hours of overlap with PST.Engagement type: ContractorEngagement Length: 3 monthsOur offered rate for this project is $15 per hour.Evaluation Process -Shortlisted candidates will be sent a Job Interest Form.After the profile review, an assessment will be shared, which must be completed within 24 hours.Based on the assessment outcomes, shortlisted candidates will be contacted to discuss the pre‑onboarding requirements.
AI Quality Analyst (Personalization) - Russian
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.Key QualificationsRussian Proficiency: Ability to read and write in Russian with a high degree of comp, as Russian is the focus language for this project.Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team.Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities.Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections.Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating.Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.Feedback: Ability to provide constructive feedback and detailed annotations.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Description:In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve:Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences.Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied.Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.Education & ExperienceBS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field).Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.Offer Details:Commitments Required: at least 4 hours per day and up to 40 hours per week with 4 hours of overlap with PST. Engagement type: ContractorEngagement Length: 3 monthsOur offered rate for this project is $15 per hour.Evaluation Process -Shortlisted candidates will be sent a Job Interest Form.After the profile review, an assessment will be shared, which must be completed within 24 hours.Based on the assessment outcomes, shortlisted candidates will be contacted to discuss the pre‑onboarding requirements.
Audio/Voice/Annotation Trainer - Korean Language
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:We're seeking a skilled Korean Voice Actor to bring characters and stories to life across various domains, using their voice to capture the right emotion, personality, and tone.What does day-to-day look like:Record voice-overs from a professional home studio or in-studio sessionsPerform lines with clarity, emotion, and appropriate pacingInterpret scripts and take creative directionRevise performances based on feedbackDeliver high-quality audio files on time and in the correct formatRequirements:Previous voice acting experienceStrong vocal control and versatilityClear speech and emotional rangeAccess to high-quality recording equipment (for remote roles)Ability to follow direction and meet deadlinesPreferred Qualifications:Experience with dubbing, ADR, or localizationDemo reel showcasing a variety of voice stylesBackground in acting or performing artsFamiliarity with audio editing toolsPerks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:This is a flexible agreement, not a full-time or part-time employment position.Evaluation ProcessShortlisting based on qualifications and relevant professional experience.Shortlisted candidates will undergo a delivery review, after which they will be ready to start!
Audio/Voice/Annotation Trainer - Japanese Language
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:We're seeking a skilled Japanese Voice Actor to bring characters and stories to life across various domains, using their voice to capture the right emotion, personality, and tone.What does day-to-day look like:Record voice-overs from a professional home studio or in-studio sessionsPerform lines with clarity, emotion, and appropriate pacingInterpret scripts and take creative directionRevise performances based on feedbackDeliver high-quality audio files on time and in the correct formatRequirements:Previous voice acting experienceStrong vocal control and versatilityClear speech and emotional rangeAccess to high-quality recording equipment (for remote roles)Ability to follow direction and meet deadlinesPreferred Qualifications:Experience with dubbing, ADR, or localizationDemo reel showcasing a variety of voice stylesBackground in acting or performing artsFamiliarity with audio editing toolsPerks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:This is a flexible agreement, not a full-time or part-time employment position.Evaluation ProcessShortlisting based on qualifications and relevant professional experience.Shortlisted candidates will undergo a delivery review, after which they will be ready to start!
Spanish Voice Actors (Studio-Grade Recording Experience)
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:We are looking for an experienced Spanish Voice Actors with 3+ years of professional recording experience to deliver one hour of recorded audio as per the provided script. The ideal candidate can deliver clear, engaging, and versatile performances for a wide range of projects, including commercials, corporate narrations, e-learning, audiobooks, and digital media. You must be able to adapt tone, pitch, and style to suit different audiences while producing clean, studio-quality audio.One-time Onboarding + Set-Up Fee: USD 20 (Expected to take a max of 1 hour)Fixed Fee: USD 200 per finished hour (“Finished hour” means after editing, mastering, cleanup as per the scripts & directions provided)What does day-to-day look like:Record voice-overs from a professional home studio or in-studio sessions.Perform lines with clarity, emotion, and appropriate pacing.Interpret scripts and take creative direction.Revise performances based on feedback.Deliver high-quality audio files on time and in the correct format.Requirements:Minimum 3 years of experience as an Spanish voice actors.Proven portfolio or samples demonstrating studio-grade work.Excellent pronunciation, diction, and fluency in Spanish.Strong vocal range and adaptability across formats.Proficiency with microphone techniques and recording best practices.Basic knowledge of sound editing (noise reduction, leveling, file formatting).Access to a high-quality microphone and professional recording setup.Ability to follow direction, meet deadlines, and manage quick turnarounds.Preferred Qualifications:Experience with dubbing, ADR, or localizationDemo reel showcasing a variety of voice stylesBackground in acting or performing artsFamiliarity with audio editing toolsOffer Details:This is a flexible agreement, not a full-time or part-time employment position.Evaluation ProcessShortlisting based on qualifications and relevant professional experience.Selected candidates will have to submit a short 30-second demo recording.
Senior Forward Deployed Engineer (Agentic AI, RAG, Enterprise Architecture)
A fast-growing enterprise AI company is looking for a Senior Forward Deployed Engineer to take technical ownership of strategic AI solution deployments for enterprise customers. This role combines hands-on engineering, solution architecture, customer partnership, and delivery leadership, with a focus on building customer-specific applications on an Agentic AI platform.This role is remote (US-based) with up to 25% travel for customer engagements.You will lead the architecture, prototyping, implementation, and post-deployment optimization of large-scale Agentic and Knowledge AI solutions. The role involves working across enterprise data pipelines, multi-agent orchestration, RAG workflows, SLM fine-tuning, platform customization, full-stack delivery, and production-grade AI systems. You will collaborate closely with customers, Product, Platform Engineering, and cross-functional teams to translate ambiguous business needs into scalable engineering plans and successful deployments.Required Skills6–10+ years of engineering experience, including at least 2 years in customer-facing, field engineering, solutions engineering, or forward deployed engineering rolesProven experience building and deploying AI/ML, data-intensive, or enterprise-grade applications in productionStrong full-stack development experience with Python, Node.js or Go, and React or VueDevOps experience with Docker, Kubernetes, CI/CD, and modern cloud-based deployment practicesExperience designing and implementing enterprise data pipelines and system integrationsStrong knowledge of REST APIs, Python, SQL, GraphQL, Webhooks, and enterprise integration patternsSolid understanding of LLMs, prompt engineering, prompt tuning, vector databases, RAG pipelines, and agentic workflowsExperience with vector databases such as AstraDB, Pinecone, or WeaviateFamiliarity with RAG frameworks such as LlamaIndex or HaystackExperience with agent and workflow orchestration tools such as LangChain, LangGraph, or CrewAIAbility to lead technical solution design and implementation for strategic enterprise customersExperience customizing platform components, integrating APIs, building reusable tooling, or extending platform logicStrong understanding of observability, monitoring, versioning, telemetry, and trustworthy AI deployment practicesAbility to translate ambiguous customer needs into clear, actionable engineering plansStrong project ownership, mentoring, communication, and collaboration skills across technical and business stakeholdersUndergraduate degree, master’s degree, or PhD in Computer Science, Data Science, or a related technical fieldBonus SkillsKnowledge of SLM fine-tuning, model distillation, and model optimization techniquesExperience building and delivering enterprise Agentic AI solutionsExperience working with Agentic development platformsFamiliarity with graph databases, multimodal AI systems, evaluation frameworks, security, guardrails, and GPU infrastructure trendsExperience contributing to reusable assets, technical best practices, internal frameworks, and documentationPrior experience supporting post-deployment optimization and production adoption for enterprise customersExperience partnering with Product and Platform Engineering teams to identify feature gaps, customer pain points, and product improvement opportunities
Senior Software Engineer – Python (LLM Evaluation & Repository Validation)
About the projects: we are building LLM evaluation and training datasets to train LLM to work on realistic software engineering problems. One of our approaches, in this project, is to build verifiable SWE tasks based on public repository histories in a synthetic approach with human-in-the-loop; while expanding the dataset coverage to different types of tasks in terms of programming language, difficulty level, and etc.About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and qualityWhy Join Us? Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems. You’ll be at the forefront of evaluating how LLMs interact with real code, influencing the future of AI-assisted software development. This is a unique opportunity to blend practical software engineering with AI research.What does day-to-day look like:Analyze and triage GitHub issues across trending open-source libraries.Set up and configure code repositories, including Dockerization and environment setup.Evaluating unit test coverage and quality.Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs.Opportunities to lead a team of junior engineers to collaborate on projects.Required Skills:Minimum 3+ years of overall experienceStrong experience with at least one of the following languages: PythonProficiency with Git, Docker, and basic software pipeline setup.Ability to understand and navigate complex codebases.Comfortable running, modifying, and testing real-world projects locally.Experience contributing to or evaluating open-source projects is a plus.Nice to Have:Previous participation in LLM research or evaluation projects.Experience building or testing developer tools or automation agents.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Employment type : Contractor assignment (no medical/paid leave)Duration of contract : 3 month; [expected start date is next week]Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, MexicoEvaluation Process (approximately 75 mins) :Two rounds of interviews (60 min technical + 30 min technical & cultural discussion)
Remote Finance & Research Analyst
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole overview:Turing is looking for Experts in finance to work with our researchers to improve performance of AI models. We are looking for experts across a range of topics including capital markets, portfolio management, research, trading, quant, investment banking, private equity, corporate finance, accounting, and others. If you enjoy solving complex problems in finance and are interested in working with AI systems, please apply. No prior AI experience is required.What does day-to-day look like:Evaluate LLM models for areas of finance where models do not perform well.Create rubrics to assess model capabilities on specific areas of your finance expertise (such as deal analysis, M&A assessments, and more).Collaborate with AI researchers and fellow finance experts to shape training methods, evaluation strategies, and benchmarks.Requirements:2+ years experience in Capital Markets, Portfolio Management, Research, Trading, Quant, Investment Banking, Private Equity, Venture Capital, Growth Equity, FP&A, Accounting, or Financial Consulting.Strong grasp of financial concepts (investment analysis, resarch, forecasting, revenue builds, corporate finance, asset management, risk management, etc.) based on your domain of expertise.Excellent English written communication.Bonuses (not at all necessary):CFA (Level I/II/III) or CA/CPA/MBA in Finance.Perks of freelancing with Turing:Work on the cutting edge of AI and finance.Fully remote and flexible work environment.Competitive hourly compensation of ~$100+/hour depending on experience.Offer Details:Commitment: Flexible, 10–30 hrs/week.Duration: ~1 month, with the possibility of extension based on performance and project needs.
Video Content Creator (US based)
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role OverviewWe’re seeking creative and charismatic storytellers to join a cutting-edge initiative designed to train AI systems to better understand and describe the world in a natural, engaging way.As a Video Creator, you’ll act as a local narrator, producing short, dynamic “walk-and-talk” videos that showcase interesting locations near you. You’ll help capture authentic, human perspectives that reflect the character, diversity, and beauty of your surroundings.What You’ll Do Day-to-DayCreate 1–5 minute “walk-and-talk” videos highlighting public places in your city (parks, landmarks, markets, or cultural spots).Speak naturally on camera, providing friendly, conversational commentary as if guiding a viewer in person.Record smooth, steady footage using your smartphone or camera with clear audio.Choose safe, public locations where filming is legally allowed.Follow provided creative guidelines, ensuring each video has a clear introduction, middle section, and conclusion.Submit videos along with accurate location details (name and address or coordinates).Respect privacy and safety best practices when filming in public areas.RequirementsComfortable speaking on camera in English with clear communication.Creative, personable, and confident presence on video.Owns a smartphone or recording device with good video and audio quality.Reliable internet connection for uploading videos.Ability to follow task guidelines and deliver videos on schedule.Nice to HaveBackground or interest in vlogging, photography, tourism, or performing arts.Experience creating videos for YouTube, TikTok, or Instagram Reels.Familiarity with basic video editing or framing techniques..Perks of Freelancing With Turing:Fully remote, flexible work — film in your local area on your own time.Opportunity to contribute to advanced AI research and multimodal data projects.Collaborate on creative, real-world tasks that shape how AI learns from human behavior.Potential for extended collaboration based on performance and project needs.Offer Details:Commitments Required : Minimum 10 hours per week. This contract assignment may require 4 hr overlap with PT Time .Engagement type : Contractor assignment/freelancer (no medical/paid leave)Duration of contract : 3 monthsLocation : United States only (Due to consent regulations, participants must NOT reside in the following states/provinces: Texas, Illinois, California, Connecticut, Massachusetts, New Jersey, Vermont)
Business Analyst
About Turing Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role OverviewWe are looking for detail-oriented professionals to support the development and improvement of AI models through high-quality data review and annotation. In this role, you will evaluate, classify, label, and validate content based on detailed project guidelines, ensuring accuracy and consistency across datasets used to train and assess AI systems.The ideal candidate has excellent attention to detail, strong analytical thinking, and the ability to follow instructions with precision. This is an exciting opportunity to contribute to cutting-edge AI projects while gaining hands-on experience in the rapidly evolving field of artificial intelligence.Key ResponsibilitiesReview, label, classify, and validate data according to project guidelines.Ensure accuracy, consistency, and quality across assigned tasks.Identify edge cases, discrepancies, and annotation issues, escalating them when appropriate.Perform quality checks and provide clear, constructive feedback.Meet productivity and quality targets while working independently in a remote environment.RequirementsStrong English reading, writing, and comprehension skills.Excellent attention to detail and analytical abilities.Ability to interpret guidelines and apply them consistently.Good communication and collaboration skills.Self-motivated with the ability to work independently.Desktop or laptop with a reliable internet connection.Preferred QualificationsPrior experience in data annotation, content review, quality assurance, or similar work is preferred.Familiarity with AI or machine learning data projects is a plus.Basic proficiency with Google Workspace or Microsoft Office.Strong problem-solving skills and the ability to make consistent, objective decisions.Offer DetailsCommitment Required: At least 4 hours per day and a total of 40 hours per week with 4 hours of overlap with PST.Engagement Type: Contractor assignment / Freelancer (no medical or paid leave).Duration of Contract: upto 12 weeksTime Zone Requirement: This role will require some overlap with UTC-8:00 (America/Los_Angeles) for 4 hours per day.Application ProcessShortlisted candidates will be sent automated analytical challenges.Once you clear them, you are ready to go!
Chemistry Specialist
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole OverviewThe Chemical Reasoning & Discovery Engineer is responsible for creating high-quality reasoning datasets that evaluate and improve the scientific reasoning capabilities of Large Language Models (LLMs). This role focuses on designing chemistry-centered discovery tasks that require models to analyze experimental or simulated data, identify chemical relationships, infer reaction rules and molecular behaviors, estimate parameters, and predict outcomes under varying conditions. The ideal candidate combines strong expertise in chemistry, quantitative analysis, and scientific methodology to develop rigorous, reproducible, and logically consistent evaluation tasks.What does day-to-day look likeDesigned chemistry-based reasoning scenarios using experimental results, molecular data, and simulated chemical systems.Authored multi-step tasks involving reaction analysis, chemical property inference, parameter estimation, and predictive reasoning.Developed evaluation problems focused on pattern discovery, reaction mechanisms, model selection, and scientific consistency checks.Created deterministic solutions, reference explanations, and detailed scoring rubrics for model assessment.Collaborated with reviewers and LLM engineers to ensure scientific accuracy, clarity, and reproducibility across datasets.Required QualificationsDemonstrated 3+ years of experience in chemistry, chemical research, scientific computing, laboratory analysis, or related analytical fields.Applied strong knowledge of chemical principles, scientific reasoning, quantitative modeling, and data interpretation.Utilized expertise in experimental design, reaction mechanisms, stoichiometry, thermodynamics, kinetics, or simulation-based analysis.Communicated complex chemical concepts, assumptions, and findings through clear and structured documentation.Evaluated scientific reasoning with exceptional attention to detail, logical consistency, and reproducible outcomes.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Potential for contract extension based on performance and project needs.Offer Details:Commitments Required : 4 hours of overlap with PST.Engagement type : Contractor assignment/freelancer (no medical/paid leave)Duration of contract : 8 weeks.
Android Engineer
About the clientOur mission is to bring community and belonging to everyone in the world. We are a community of communities where people can dive into anything through experiences built around their interests, hobbies, and passions. With more than 50 million people visiting 100,000+ communities daily, it is home to the most open and authentic online conversations.About the RoleDesign, develop, and prototype Android native customer applications for internal and external use. Participate in the full app life cycle: concept, design, build, deploy, test, and release to the app store. Work with product teams on new ideas, designs, prototypes, and estimates. Keep up-to-date on current and upcoming features in relevant products and platforms. Drive a best practices approach to continuously improving our products, processes, and tools. Assist in the creation and maintenance of documentation for all features in development.What you'll doWork cross-functionally with product, design, and other engineering counterparts to execute on product and business strategy and build novel products and features that our users will love.Contribute to the full development cycle: technical design, development, test, experimentation, analysis, and launch. You’ll be reviewing code and design docs, giving feedback on product specs and mocks.Set and define standards that improve developer workflows, recommend best practices, and help coach and mentor engineers on the team to further their professional development.Continuously learn and improve your technical and non-technical abilities.Who You Might BeExpertise in Java or Kotlin, with 8+ years of experience in Android development.Sound software engineering fundamentals.Experience building mobile applications at scale with at least 1MM usersAble to embrace the challenges of building data intensive, highly responsive, and fault tolerant apps in the constrained environment of a smartphone.Passion for developing scalable, well-designed software that improves people’s lives globally.BS degree in Computer Science, a similar technical field of study, or equivalent practical experience.Offer DetailsFull-time contractor (no benefits)Remote only, full-time dedication (40 hours/week)Required 6+ hours overlap with PSTCompetitive compensation package.Opportunities for professional growth and career development.
JavaScript / TypeScript Full-Stack Developer
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:We are actively seeking talented developers proficient in Javascript/Typescript to join our ambitious team dedicated to pushing the frontiers of AI technology. This opportunity is tailored for professionals who thrive on developing innovative solutions and aspire to be at the forefront of AI advancements. You will work with different companies in the US who are looking to develop cutting-edge commercial and research AI solutions.What does day-to-day look like:Design, develop, and maintain efficient, high-quality code to train and optimize AI models.Conduct evaluations (Evals) to benchmark model performance and analyze results for continuous improvement.Evaluate and rank AI model responses to user queries across diverse domains, ensuring alignment with predefined criteria.Develop comprehensive explanations and rationales for evaluations, showcasing excellent reasoning and technical expertise.Lead efforts in Supervised Fine-Tuning (SFT), including creating and maintaining high-quality, task-specific datasets.Collaborate with researchers and annotators to execute Reinforcement Learning with Human Feedback (RLHF) and refine reward models.Design innovative evaluation strategies and processes to improve the model's alignment with user needs and ethical guidelines.Create and refine optimal responses to improve AI performance, emphasizing clarity, relevance, and technical accuracy.Conduct thorough peer reviews of code and documentation, providing constructive feedback and identifying areas for improvement.Collaborate with cross-functional teams to improve model performance and contribute to product enhancements.Continuously explore and integrate new tools, techniques, and methodologies to enhance AI training processes.Required Skills:Write readable, reusable, and maintainable codeParticipate in code reviews to ensure that the standards for code quality are metDemonstrate your proficiency with your language of choice, while covering all basesProvide clear, clean, well-organized, correct, and clearly annotated/classifiable code in the responsesBachelor’s/Master’s degree in Engineering, Computer Science (or equivalent experience)Demonstrable experience with developing web apps using modular development and scalable architectures as well as a strong focus on code readability and security/stability (i.e. testing)Proficiency with the language's syntax and conventionsExtensive experience working with JavaScript or TypeScriptSolid understanding of JavaScript ES6 and either Node.js or React or Nest or Angular or VueNice to have some prior software Quality Assurance and Test Planning experienceExcellent spoken and written English communication skillsDocker experience TechnologiesBack-end : Nest.js, Node.jsFront-end : Vue, Angular, React Languages : JavaScript & TypescriptPerks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: at least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Engagement type : Contractor assignment (no medical/paid leave)Duration of contract : 1 month; [expected start date is next week]Evaluation Process (approximately 75 mins) :Two rounds of interviews (60 min technical + 15 min cultural &, offer discussion)
AI Engineering Lead
AI Engineering LeadLocation: IndiaEmployment Type: Full-TimeRequired skills:12+ years of professional experience as a software engineer and building applications/systems.2+ years of hands-on experience in how LLMs work & Generative AI (LLM) techniques, particularly multi-agent systems.Expert proficiency in programming skills in Python, Langgraph, and SQL is a must.Expert in architecting GenAI applications/systems using various frameworks & cloud services.Expert proficiency in using AI tools like Claude Code, Codex, cursor, windsurf, and the like.Expert proficiency in AI observability & evaluation tools like Langsmith, Langfuse, or similar.Good proficiency in using various cloud services from Azure, GCP, or AWS for building GenAI applications.Experience in driving the engineering team toward a technical roadmap.Excellent communication skills to effectively collaborate with business SMEs.Roles & Responsibilities:Solutioning & LeadBuild the technical roadmap given a business requirement and own the delivery of the same.Lead the engineering team toward a technical roadmap and ensure the timely execution of the roadmap to achieve customer satisfaction.Design robust multi-agent architectures, including supervisor-router patterns with dynamic sub-agent routing and stopping conditions.Mentoring and guidance: Provide technical leadership and knowledge-sharing to the engineering team, fostering best practices in machine learning and large language model development.Hands-on skillsDevelop LLM-based solutions: Lead the design, training, fine-tuning, and deployment of large language models, leveraging techniques like retrieval-augmented generation (RAG) and multi-agent-based architectures.Build and maintain agent evaluation pipelines, including offline eval datasets, LLM-as-judge, and CI-integrated eval runs.Codebase ownership: Build & maintain high-quality, efficient code in Python (using frameworks like LangChain/LangGraph) and SQL, focusing on reusable components, scalability, and performance best practices.Cloud integration: Deployment of GenAI applications on cloud platforms (Azure, GCP, or AWS), optimizing resource usage and ensuring robust CI/CD processes.Communication & Cross-functional collaborationActively follows the frontier and has differentiated, up-to-date views on model releases, agentic architectures, evaluation methods, tool-use and computer-use patterns, multimodal capability, reasoning/test-time compute trends, and the serious open questions in the field.Produce a structured, high-signal answer to an open-ended technical or strategic question — while modulating depth for a non-engineering executive audience.Work closely with product owners, data scientists, and business SMEs to define project requirements, translate technical details, and deliver impactful AI products.
Technical Content Writer
About TuringBased in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole OverviewWe are looking for a Technical Content Writer who can deliver clear, technically accurate content and support data-driven work with hands-on experience using JSON and public datasets. You’ll analyze datasets to extract business insights and present reasoning through well-structured documentation, notebooks, or clearly explained technical artifacts.You will collaborate with researchers and stakeholders, translating complex findings into concise narratives and actionable recommendations. Strong communication, output quality, and attention to detail are essential.What does day-to-day life look like?Produce clear, well-structured technical documentation and written deliverables.Analyze public datasets to derive business insights and respond to key analytical questions.Clearly explain reasoning and logic in notebooks or other suitable formats.Work with JSON-based data and ensure outputs are accurate and well-organized.Use Python scripting (as needed) to support data exploration, validation, and reproducible analysis.Ensure comprehensive documentation and traceability of methods, assumptions, and findings.Collaborate and communicate with researchers and stakeholders to refine insights and deliverables.RequirementsMinimum 3+ years of relevant experience as a technical content writer, analyst, or software developer.Strong technical writing skills with experience producing clear, accurate, and structured content.Comfort working with data formats like JSON and interpreting datasets.Working knowledge of Python for data exploration/validation (not a core engineering role).Familiarity with notebooks and/or scripting workflows to support analysis and documentation.Good understanding of code quality, formatting, and documentation best practices.Strong analytical abilities and business sense to draw appropriate conclusions and communicate them clearly.Excellent spoken and written English communication skills.Bachelor’s/Master’s degree in Engineering, Computer Science, or equivalent experience.Perks of Freelancing With TuringWork in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer DetailsCommitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST.Engagement Type: Contractor assignment (no medical/paid leave)Duration of Contract: 3 months (adjustable based on engagement)Location: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, MexicoEvaluation ProcessTechnical Interview with live coding challege (60 mins)
Python Machine Learning Engineer
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&LRole Overview:We’re looking for a ML Developer to drive the design, development, and delivery of advanced machine learning solutions. The ideal candidate is not just a strong individual contributor but also a technical leader capable of setting direction, mentoring team members, and ensuring that ML initiatives align with business goals. Competitive ML experience (e.g., Kaggle, benchmarks) is a strong plus.What does day-to-day look like:Own end-to-end DS/ML solution development — from data pipelines and model design to deployment and monitoring.Translate business objectives into robust ML architectures that accurately capture business logic and context.Collaborate cross-functionally with Product, Engineering, and Business stakeholders to define problem statements and success metrics.Evaluate and optimize models for performance, scalability, and accuracy using state-of-the-art techniques.Stay current with advancements in AI/ML research and apply relevant innovations to improve outcomesRequired Qualifications:Bachelor’s or Master’s degree in Computer Science, Machine Learning, AI, Statistics, or a related quantitative field.4+ years of hands-on DS/ML development experience,Proficiency in key DS/ML areas and frameworks:Supervised and Unsupervised Learning Time-Series Forecasting Natural Language Processing (NLP) Computer Vision (CV) Statistical Modeling and InferenceExpertise in Python and core libraries (Pandas, NumPy, Scikit-learn, etc.).Ability to understand and apply different models to real-world use casesStrong understanding of data preprocessing, feature engineering, model tuning, and evaluation metrics.Proven ability to design scalable, production-grade ML systems.Preferred Qualifications:Proven expertise in Deep learning (e.g., convolutional neural networks, recurrent neural networks, transformers).Experience with cloud data platforms (Databricks, AWS, etc.)Hands-on experience with PySpark and Databricks PlatformStay up-to-date with the latest advancements in machine learning and artificial intelligence.Bonus:Experience and knowledge in Kaggle competitions and Benchmarks, such as MLEBenchExperimenting with new technologies and frameworks based on Research papers published in top conferences and journalsPerks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Employment type : Contractor assignment (no medical/paid leave)Duration of contract : 3 month; [expected start date is next week]Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, MexicoEvaluation Process (approximately 75 mins) :Two rounds of interviews as follows:Interview 1: Technical, 60 mins Interview 2: Onboarding & cultural discussion, 15 mins
LLM Trainer - Agent Function call
About Turing:Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.Role Overview:This position is within a project with one of the foundational LLM companies. The goal is to assist these foundational LLM companies in enhancing their Large Language Models.One way we help these companies improve their models is by providing them with high-quality proprietary data. This data serves two main purposes: first, as a basis for fine-tuning their models, and second, as an evaluation set to benchmark the performance of their models or competitor models.For example, in the case of Agent Completion (AC) data generation, your task will be to simulate high-quality multi-turn conversations between a user and a smart assistant that utilizes function-calling tools to accomplish user goals. You will craft these dialogues by playing both the assistant and the user, while simulating tool use where necessary to guide the assistant through complex decision-making and real-world reasoning scenarios.What does day-to-day look like:Design multi-turn conversations that simulate real interactions between users and AI assistants using apps like calendar, email, maps, and drive.Emulate both the user and the assistant, including the assistant's tool calls (only when corrections are needed).Carefully select when and how the assistant uses available tools, ensuring logical flow and proper usage of function calls.Craft dialogues that demonstrate natural language, intelligent behavior, and contextual understanding across multiple turns.Generate examples that showcase the assistant’s ability to gracefully complete feasible tasks, recognize infeasible ones, and maintain engaging general chat when tools aren’t required.Ensure all conversations adhere to defined formatting and quality guidelines, using an internal playbook.Iterate on conversation examples based on feedback to continuously improve realism, clarity, and value for training purposes.Collaborate with peers and reviewers to maintain consistency and high standards in deliverables.Requirements:Strong general technical reasoning skills and the ability to model real-world assistant behavior using tool-based APIs.Ability to break down complex tasks and simulate realistic dialogues that reflect user expectations and assistant limitations.Experience in any programming language or tech stack is acceptable; a strong grasp of APIs, data formats (e.g., JSON), and logical thinking is more critical than specific toolsets.Excellent written communication skills in English, with a focus on clarity, tone, and instructional coherence.Creativity and attention to detail in crafting realistic scenarios and responses.Experience working with or around LLMs, virtual assistants, or function-calling frameworks is a plus.Ability to follow detailed guidelines and formatting standards with high consistency.3+ years of overall professional experience in a technical or analytical field.Perks of Freelancing With Turing:Work in a fully remote environment.Opportunity to work on cutting-edge AI projects with leading LLM companies.Offer Details:Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. (We have 3 options of time commitment: 20 hrs/week, 30 hrs/week or 40 hrs/week)Engagement Type: Contractor assignment (no medical/paid leave)Duration of Contract: 6 WeeksLocation: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, Mexico
AI Quality Analyst (Personalization) - Japanese
About Turing:Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.Role Overview:As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.Key QualificationsJapanese Proficiency: Ability to read and write in Japanese with a high degree of comp, as Japanese is the focus language for this project.Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team.Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities.Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections.Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating.Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.Feedback: Ability to provide constructive feedback and detailed annotations.Communication: Excellent communication and collaboration skills.Independence: Self-motivated and able to work independently in a remote setting.Technical Setup: Desktop/Laptop set up with a good internet connection.Description:In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve:Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences.Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied.Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.Education & ExperienceBS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field).Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.Offer Details:Commitments Required: at least 4 hours per day and up to 20 hours per week with 4 hours of overlap with PST.Engagement type: ContractorEngagement Length: 3 monthsOur offered rate for this project is $15 per hour.Evaluation Process -Shortlisted candidates will be sent a Job Interest Form.After the profile review, an assessment will be shared, which must be completed within 24 hours.Based on the assessment outcomes, shortlisted candidates will be contacted to discuss the pre‑onboarding requirements.
Reviews
You must be logged in to leave a review for this company.
