Engineering judgment for capable AI
What this work looks like
Coding evaluator roles use production engineering experience to test whether AI-generated software is correct, maintainable, secure, and useful. Compare current roles by technical focus, compensation, work type, and eligibility before continuing to the external application.Typical work
- Review generated code for correctness, architecture, performance, accessibility, and maintainability.
- Reproduce failures, compare alternative implementations, and explain why one solution is stronger.
- Evaluate frontend interfaces, backend services, data workflows, or specialist languages based on the project.
What teams look for
- Recent professional software development experience in the languages or systems named by the listing.
- Ability to debug unfamiliar code and communicate technical tradeoffs clearly.
- Evidence of shipped production work, careful reviews, or quality ownership.
Apply with stronger evidence
- Match your strongest languages and shipped systems directly to the role requirements.
- Describe measurable engineering outcomes instead of listing tools without context.
- Read platform-specific experience requirements before investing time in an application.
Related searchesAI code reviewer jobssoftware engineer AI training jobsremote frontend AI evaluator work
Current opportunities
Verified listing details with no HumanitApp account or application fee.
Code / RemoteWe are looking for engineers who build and operate LLM agents in production, and who have real visibility into how agents are actually used inside a company. — You have probably: • Shipped an agent that real users depended on, and been on...
$100 - $500 / per-taskVerified / Application completed externally Code / Remote1. Role Overview — Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the...
$300 / per-taskVerified / Application completed externally Code / RemoteWe're looking for highly accomplished Cybersecurity Research Experts to help evaluate cutting-edge AI systems in advanced security reasoning, vulnerability analysis, exploit development, and secure software engineering. This project is...
$200 - $250 / hourVerified / Application completed externally Code / RemoteAbout Mercor’s talent network — Join our Machine Learning Engineer Expert Network to connect with leading AI labs and companies seeking your expertise. This is an open application for future contract opportunities that match your...
$70 - $250 / hourVerified / Application completed externally Code / RemoteMercor is partnering with a frontier AI research lab to engage experienced software engineers with deep expertise in legacy codebase migration and modernisation. We are looking for engineers who have personally executed complex migrations...
$200 / hourVerified / Application completed externally Code / RemoteWe're looking for engineers who have shipped production search systems (especially agentic ones now that we're in the era of LLMs and agents) and are thinking hard about the requirements and optimizations of those systems. — You've...
$80 - $150 / per-taskVerified / Application completed externally Code / RemoteMercor is building a network of experienced data scientists for potential future projects with leading AI research organizations. These projects may focus on evaluating how effectively AI systems perform real-world data science work. —...
$100 - $150 / hourVerified / Application completed externally Code / RemoteOverview — We are hiring Engineering / Platform professionals with technical backgrounds in software engineering, infrastructure, or platform development. In this role, you will review, assess, and provide structured feedback on your...
$80 - $160 / hourVerified / Application completed externally Code / RemoteAbout Mercor’s talent network — Join our Frontend Engineer Expert Network to connect with leading AI labs and companies seeking your expertise. This is an open application for future contract opportunities that match your background and...
$70 - $150 / hourVerified / Application completed externally Code / RemoteAbout Mercor’s talent network — Join our DevOps / Platform Engineer Expert Network to connect with leading AI labs and companies seeking your expertise. This is an open application for future contract opportunities that match your...
$70 - $150 / hourVerified / Application completed externally Code / RemoteAbout Mercor’s talent network — Join our Backend Engineer Expert Network to connect with leading AI labs and companies seeking your expertise. This is an open application for future contract opportunities that match your background and...
$70 - $150 / hourVerified / Application completed externally Code / RemoteAbout Mercor’s talent network — Join our Full-Stack Engineer Expert Network to connect with leading AI labs and companies seeking Full-Stack Engineering expertise. This is an open application for future contract opportunities that match...
$70 - $150 / hourVerified / Application completed externally Code / RemoteMercor is seeking an exceptional open source contributor with deep expertise in Python, Java, C, JavaScript, or TypeScript to collaborate on high-impact projects with global reach. This role is ideal for engineers with a strong command of...
$100 - $150 / hourVerified / Application completed externally Code / United States (Remote)Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team...
$75 - $110 / hourVerified / Application completed externally Code / United States RemoteJoin a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team...
$90 - $110 / hourVerified / Application completed externally Code / RemoteMercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Risk Adjustment and HCC Coding leaders to evaluate AI tools designed to improve risk score accuracy and...
$110 / hourVerified / Application completed externally Code / Bay Area, CAJoin a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — We are hiring a senior engineering and software domain...
$65 - $105 / hourVerified / Application completed externally Code / RemoteWe're looking for experienced machine learning researchers with hands-on experience training and improving deep learning models end-to-end, across vision and language. You'll work on well-scoped empirical open-ended ML research problems. —...
$100 - $120 / hourVerified / Application completed externally Code / RemoteAbout the role — We are hiring expert Evaluators in Document/deck production QA to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep...
$80 - $120 / hourVerified / Application completed externally Code / RemoteAbout the role — We are hiring expert Evaluators in Software / AI / IT / data to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep...
$80 - $120 / hourVerified / Application completed externally Code / RemoteAbout the role — We are hiring expert Evaluators in Spreadsheet QA / workbook maintenance to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply...
$80 - $120 / hourVerified / Application completed externally Code / RemoteAbout the role — We are hiring expert Evaluators in Engineering / manufacturing / technical operations to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You...
$80 - $120 / hourVerified / Application completed externally Code / RemoteAbout Mercor’s talent network — This is an open application for future contract opportunities that match your background and interests. — Once you complete your profile and pass our AI interview, you'll be eligible for relevant projects as...
$70 - $110 / hourVerified / Application completed externally Code / Remote — United StatesEvaluate the quality and correctness of AI-assisted software-development traces used to train and evaluate a frontier AI lab's models. You'll assess end-to-end coding sessions produced with AI-assisted developer tools — judging...
$70 - $90 / hourVerified / Application completed externally Code / Remote — United StatesEvaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading...
$70 - $90 / hourVerified / Application completed externally Code / Remote — United StatesEvaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab's models. You'll assess experiment design, model-selection reasoning, and evaluation methodology —...
$70 - $90 / hourVerified / Application completed externally Code / Remote — United StatesEvaluate the quality, correctness, and cloud-architecture soundness of AWS serverless and infrastructure-as-code tasks used to train and evaluate a frontier AI lab's models. You'll assess multi-service serverless designs, IaC fidelity, and...
$70 - $90 / hourVerified / Application completed externally Code / Remote — United StatesEvaluate the quality, correctness, and production-readiness of Kubernetes tasks used to train and evaluate a frontier AI lab's models. You'll assess cluster-operations scenarios, manifest correctness, and failure-mode troubleshooting — and...
$70 - $90 / hourVerified / Application completed externally Code / Remote — United StatesEvaluate the quality, fidelity, and completeness of vulnerability-reproduction and remediation tasks used to train and evaluate a frontier AI lab's models. You'll assess whether CVE reproductions are faithful, fixes are sound, verification...
$70 - $90 / hourVerified / Application completed externally Code / RemoteRole Overview • Mercor is seeking senior litigation professionals to build evaluation tasks for AI systems operating in civil litigation and dispute resolution contexts. • The workflows are calibrated to the case complexity, evidentiary...
$90 - $100 / hourVerified / Application completed externally Code / RemoteComputational Structural & Mechanical Engineering Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard computational scientific and engineering problems. As a task...
$70 - $85 / hourVerified / Application completed externally Code / Remote — US-basedMercor is seeking computational scientists specializing in atomistic and surface modeling to support a frontier AI research lab building models for materials science and the physical sciences. This is hands-on, expert-level work: you'll...
$84 / hourVerified / Application completed externally Code / RemoteMercor is seeking experimental scientists and engineers across inorganic synthesis, characterization, superconductors, and semiconductors (including advanced packaging) to support a frontier AI research lab building models for materials...
$84 / hourVerified / Application completed externally Code / RemoteMercor is seeking experts in Atomic Layer Deposition (ALD) and thin-film processes to support a frontier AI research lab building models for semiconductors and the physical sciences. This is hands-on, expert-level work: you'll apply deep,...
$84 / hourVerified / Application completed externally Code / RemoteAbout the Role — Mercor is partnering with a leading technology organization to recruit experienced Radio Frequency (RF) and Electromagnetic (EM) Engineering Experts for a short-term project. Experts will apply their hands-on industry...
$80 - $95 / hourVerified / Application completed externally Code / RemoteMercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Coding Managers and HIM Coding leaders to evaluate AI-powered coding solutions and help train...
$80 / hourVerified / Application completed externally Code / RemoteMercor is hiring SOC Investigation Specialist on behalf of high-growth technology and enterprise partners building next-generation SOC automation and AI-driven investigation systems. This role is ideal for experienced SOC analysts who can...
$70 - $95 / hourVerified / Application completed externally Code / RemoteRole Overview — We are seeking expert engineers to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core engineering domains,...
$61 - $77 / hourVerified / Application completed externally Code / RemoteAbout the role — We build materials-science and engineering tasks that test how well an AI model does real expert work. You write a realistic task — a prompt, a data room, and a way to grade it — run it against the model, and tighten it...
$60 - $90 / hourVerified / Application completed externally Code / RemoteAbout the role — A frontier engineering-reasoning evaluation run in collaboration with a leading AI research lab. The work measures whether state-of-the-art models can reason from first principles in your engineering domain rather than...
$60 - $90 / hourVerified / Application completed externally Code / RemoteAbout the work — We're building a high-quality dataset of human preference judgments on AI-generated frontend code. Each task hands you a reference web page — crawled from the real internet, delivered as a full zipped site tree plus...
$90 / hourVerified / Application completed externally Code / RemoteComputational Statistics and Applied Mathematics Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create...
$70 - $90 / hourVerified / Application completed externally Code / RemoteAbout the Role — Mercor is partnering with a leading technology organization to recruit experienced Electrical Test Automation Engineers for a short-term project. Engineers will apply hands-on industry experience in electrical testing,...
$70 - $85 / hourVerified / Application completed externally Code / RemoteAbout the Role — Mercor is partnering with a leading technology organization to recruit experienced Electrical Design Specialists for a short-term project. Specialists will apply their hands-on industry experience to analyze, evaluate, and...
$70 - $85 / hourVerified / Application completed externally Code / RemoteMercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) — Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that...
$70 / hourVerified / Application completed externally Code / RemoteMercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) — Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that...
$70 / hourVerified / Application completed externally Code / RemoteElectrical Engineering & RF/Circuit Design Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create...
$70 - $85 / hourVerified / Application completed externally Code / United States RemoteJoin a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team...
$50 - $65 / hourVerified / Application completed externally Code / RemoteRole Overview • Mercor is seeking senior civil engineering professionals to build evaluation tasks for AI systems operating in large-scale infrastructure and construction contexts. • The workflows are calibrated to the design complexity,...
$70 - $80 / hourVerified / Application completed externally Code / RemoteRole Overview • Mercor is seeking senior electrical engineering professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise electrical systems and product design contexts. • The workflows are calibrated to...
$70 - $80 / hourVerified / Application completed externally Code / RemoteRole Overview • Mercor is seeking senior mechanical engineering professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise product design and manufacturing contexts. • The workflows are calibrated to the...
$70 - $80 / hourVerified / Application completed externally Code / RemoteAbout Mercor’s talent network — Join our Physicist Expert Network to connect with leading AI labs and companies seeking your expertise. This is an open application for future contract opportunities that match your background and interests....
$60 - $80 / hourVerified / Application completed externally Code / US RemoteSupport a high-priority technical-expert panel for a frontier AI lab as an Expert Interviewer. — Overview — We're seeking an experienced interviewer to help vet shortlisted engineers across GPU kernel development, security & vulnerability...
$50 - $60 / hourVerified / Application completed externally Code / RemoteWrite original general-knowledge multiple-choice questions in Norwegian for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote, flexible...
$48.51 - $59.29 / hourVerified / Application completed externally Code / RemoteAre you a Level 3 / Tier 3 network support engineer interested in data science and autonomous infrastructure? Our client is building vertically integrated networking systems and using the data they generate to power the next generation of...
$50 - $70 / hourVerified / Application completed externally Code / RemoteOverview We're looking for audio engineers with strong experience editing and quality-checking voice recordings for speech and text-to-speech (TTS) applications, specifically for French-language audio. This role focuses on preparing...
$50 / hourVerified / Application completed externally Code / RemoteWrite original general-knowledge multiple-choice questions in Chinese (Simplified) for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote,...
$39.69 - $48.51 / hourVerified / Application completed externally Code / RemoteWrite original general-knowledge multiple-choice questions in Chinese (Traditional) for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote,...
$39.69 - $48.51 / hourVerified / Application completed externally Code / RemoteWrite original general-knowledge multiple-choice questions in Czech for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote, flexible hours. —...
$33.96 - $41.5 / hourVerified / Application completed externally Code / India RemoteJoin a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team...
$25 - $30 / hourVerified / Application completed externally Code / RemoteWrite original general-knowledge multiple-choice questions in Arabic for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote, flexible hours. —...
$24.26 - $29.65 / hourVerified / Application completed externally Code / RemoteMercor’s Talent Success team is hiring! Our Talent Success team is responsible for ensuring everyone using our platform has a delightful experience end to end, from applying to job listings, taking our AI-led interviews, accepting offers...
$35,000 - $50,000 / yearVerified / Application completed externally Code / RemoteMercor’s Talent Success team is hiring! Our Talent Success team is responsible for ensuring everyone using our platform has a delightful experience end to end, from applying to job listings, taking our AI-led interviews, accepting offers...
$30,000 - $45,000 / yearVerified / Application completed externally