Amazon Mechanical Turk will permanently close on September 30, 2026, including its public worker type in SageMaker Ground Truth and Amazon Augmented AI. Before the Mechanical Turk shutdown, export your history, finish open tasks, classify every review queue by risk and expertise, and decide which work belongs with automation, a vendor, contractors, or a trained human review team.

That is more than a vendor migration. MTurk made human judgment feel like an API: send a Human Intelligence Task, receive an answer, pay per task. Its closure exposes all the operating decisions hidden behind that interface. You still need to know who may see the data, what qualifies someone to judge it, how you resolve disagreement, and who owns the final answer.

## What is the Mechanical Turk shutdown timeline?

**MTurk closes on September 30, 2026, but several wind-down dates continue after the service stops accepting work.** Put those dates into a migration plan with a named owner now. Do not leave them in a browser tab and hope somebody remembers.

According to [Amazon's closure FAQ](https://www.mturk.com/help):

- Workers can submit HITs until September 30, 2026. Unsubmitted HITs expire when the service closes.
- Requesters have the standard 30-day window to approve or reject work submitted before closure. Unreviewed submissions are automatically approved after that window.
- Requesters can award bonuses until October 30, 2026.
- Prepaid requester balances should be refunded within 30 days.
- Transaction history remains accessible until January 28, 2027.
- The MTurk worker type also stops being available for new or existing human-review workflows in SageMaker Ground Truth and Amazon Augmented AI on September 30.

Amazon does not give a reason for the closure in the FAQ. Do not base a migration decision on speculation about why it happened. Base it on the confirmed deadline and the work your product still requires.

Start with a named owner. That person should inventory every direct MTurk integration, every AWS labeling or review workflow that selects the public workforce, every scheduled job that creates HITs, and every downstream process that expects their results. Include one-off research and operations accounts, not just production code.

Then preserve what you will need after January 28. Export transaction records, task definitions, qualification rules, gold-standard examples, reviewer instructions, rejection reasons, and quality reports. A payment ledger alone will not help you reconstruct why a label was accepted.

## What did MTurk actually provide to AI startups?

**MTurk provided elastic human capacity, task distribution, payment, and a programmable workflow. It did not remove the need to design the judgment itself.** That distinction matters when you choose what replaces it.

Amazon's [MTurk product overview](https://www.mturk.com/) describes machine-learning data collection, annotation, continuous correction, and human-in-the-loop validation as core uses. It also describes the economic appeal clearly: companies could scale a distributed workforce up and down without building a full internal operation.

For a startup, that bundle met several needs at once:

1. **Reach:** you could ask a large public pool to complete small tasks.
2. **Elasticity:** you paid when work entered the queue instead of carrying fixed capacity.
3. **Distribution:** the platform assigned units of work and collected responses.
4. **Administration:** it handled worker accounts, balances, and payments.
5. **Integration:** an API connected human decisions to a software pipeline.

That convenient abstraction also hid decisions your team still owned. Who was qualified? Were instructions clear enough for an outsider? Could the worker see personal or confidential data? Was majority agreement meaningful? What happened when the label affected a customer outcome?

If nobody can answer those questions, your company did not have a human-in-the-loop system. It had a queue with humans somewhere inside it.

This is why the shutdown is useful, despite the disruption. It gives you a hard date to turn an accidental dependency into an explicit operating function.

## Should you replace MTurk with another crowd platform?

**A like-for-like crowd platform is one valid replacement, but only for work that is safe, bounded, and measurable.** Classify the task before comparing vendors. Otherwise you preserve the same hidden risks behind a new API.

Use four questions for every queue:

- **Sensitivity:** What can the reviewer see? Does it include customer data, candidate data, internal documents, personal information, or regulated records?
- **Expertise:** Can a trained generalist make the judgment, or does it require domain knowledge?
- **Ambiguity:** Is there an objectively correct answer, a stable rubric, or a legitimate range of interpretations?
- **Impact:** What happens when the answer is wrong? Can you reverse it cheaply, or does it affect safety, access, employment, money, or a customer relationship?

AWS's [public-workforce documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/sms-workforce-management-public.html) draws a firm privacy boundary. It says not to send confidential information, personal information, or protected health information to the public MTurk workforce. Data sent to that workforce must be marked as free of personally identifiable information.

That restriction should survive MTurk. Moving the same sensitive queue to a different public crowd does not make the data safe.

| Replacement path | Best fit | Main risk |
|---|---|---|
| Automation or model-as-judge | Deterministic, reversible triage with strong tests | One systematic error can spread across the entire queue |
| Managed annotation vendor | Burst capacity with outsourced operations | Reviewer quality and subcontracting can be opaque |
| Private contractor pool | Repeated, bounded work with known reviewers | You now own scheduling, management, and classification obligations |
| In-house review team | Sensitive, expert, or high-impact decisions | Fixed cost and uneven queue volume |
| Hybrid tiered review | Mixed work with clear escalation boundaries | Requires more workflow design and measurement |

For most AI startups, **hybrid review is the defensible default**. Automate rules you can test. Give bounded recurring judgments to trained reviewers. Escalate ambiguity and high-impact cases to experts. Name one person who can change the rubric and adjudicate disputes.

Do not confuse automation with migration. A model can reduce the queue, pre-sort cases, or draft a recommendation. It should not silently become the final judge just because the human supplier disappeared.

## Which human-in-the-loop tasks need experts?

**Expert review is warranted when the task combines ambiguity, specialized knowledge, sensitive data, or costly errors.** Simple labeling can still suit trained generalists. The job is to route work by risk, not to declare one workforce model universally superior.

A practical review ladder has four levels:

1. **Deterministic automation** handles rules with an answer you can verify in code.
2. **Trained generalists** handle repeatable judgments with a clear rubric and safe data.
3. **Domain experts** handle edge cases, sensitive material, and decisions that depend on professional context.
4. **A named adjudicator** resolves disagreement, accepts appeals, and owns changes to the rubric.

The [Data Science with Human-in-the-Loop workshop proceedings](https://aclanthology.org/2024.dash-1.pdf) show the same division of labor in research workflows: nonexperts can perform simple annotation, while complex error analysis and improvement decisions need engineers or domain specialists.

Adding more anonymous votes does not automatically produce truth. AWS notes that assigning more workers can improve complex labeling tasks, while offering little benefit for simple ones. A peer-reviewed [public-health surveillance study](https://www.jmir.org/2022/1/e28749/) used three crowd workers per item and compared them with a gold-standard set. The researchers measured reliability instead of treating agreement as proof.

Complex work also creates a subtler failure mode. A 2026 Harvard Data Science Review study, [Bias in the Loop](https://hdsr.mitpress.mit.edu/pub/nrcn4h7d/release/2), found that nonexpert annotators can struggle with demanding tasks and that AI suggestions can bias human judgments. A reviewer who clicks “agree” after reading a model answer is not an independent check.

Design for independence where it matters. Ask reviewers to score before seeing the model's proposed answer or another person's decision. Show suggestions later, if at all. Preserve the rationale so an adjudicator can distinguish a weak rubric from a careless review.

## How do you hire a human review team?

**Hire reviewers based on representative evidence of judgment, not generic claims of attention to detail.** Define the work, pay candidates for a realistic sample, score it independently, and treat sound escalation judgment as a hiring criterion.

### Write the outcome before the job description

Start with the decision the reviewer must make and the evidence available to them. Describe the normal cases, the edge cases, the data boundary, the expected throughput, and the point at which they must stop and escalate.

“Review AI output” is not a role. “Evaluate whether support answers follow the policy, cite the supplied account facts, avoid prohibited promises, and escalate unclear refund cases” is a role someone can understand and test.

Choose the employment model after defining the queue. A managed vendor can absorb a seasonal backlog. A private contractor pool can handle repeated work with stable reviewers. An employee can own a sensitive queue that requires deep product context and daily collaboration. Check the legal and tax obligations that apply to your location and arrangement.

If this is one of your first specialist roles, [The First Five Hires](/blog/first-five-hires-seed-stage-sequencing) gives you a framework for deciding whether the need is durable enough to justify a permanent seat. If a founder still owns the process, [Hiring Without a Recruiter](/blog/founder-led-hiring-without-recruiter) explains how to run a disciplined search without inventing an HR department.

### Use a paid, representative work sample

Build a small sample from realistic versions of the work, with sensitive details removed. Include straightforward items, ambiguous items, and at least one case that should be escalated rather than forced into a label.

Score candidates on more than agreement with your answer key:

- correct decisions on reference cases;
- false-positive and false-negative costs;
- consistency across similar examples;
- quality of written rationale;
- recognition of missing information;
- correct use of the escalation boundary;
- care with privacy and confidential material.

Pay for the sample. You are asking candidates to perform work that resembles the job, and the exercise may require serious concentration. Keep it short enough that qualified people with existing jobs or caring duties can participate.

Have interviewers score independently before discussion. [Structured interview scorecards](/blog/structured-interview-scorecards-predictive-validity) help prevent the loudest opinion from replacing the evidence. Record a reason for the final decision, especially when the panel overrides the work-sample result.

## How do you measure reviewer quality?

**Measure a review system with calibrated examples, agreement, error costs, escalation behavior, and drift over time.** Raw throughput and majority vote are not enough.

Create a gold set of examples whose expected decisions have been reviewed by your rubric owner or domain expert. Do not make it a museum piece. Add cases when production reveals a new ambiguity, and retire examples when the policy changes.

Track a compact scorecard:

| Measure | What it tells you | Warning sign |
|---|---|---|
| Gold-set accuracy | Whether reviewers apply the current rubric | Accuracy falls after a policy or product change |
| Pairwise agreement | Whether the task produces consistent judgments | Reviewers disagree on supposedly simple cases |
| False-positive cost | Harm from incorrectly accepting or flagging an item | Automation optimizes volume while costly errors rise |
| False-negative cost | Harm from missing a real issue | Reviewers avoid difficult escalations to protect speed |
| Escalation rate | Whether reviewers recognize uncertainty | Near-zero escalation on an ambiguous queue |
| Override rate | How often experts reverse first-line decisions | One reviewer, task type, or rubric section dominates overrides |
| Review latency | Whether capacity matches demand | Backlog age grows even when daily volume looks stable |

Segment the metrics. An overall accuracy number can hide a weak language, customer group, content type, or edge-case category. Break down disagreement by task slice and rubric section.

Pay for calibration, disagreement review, and documentation, not just accepted labels. If compensation rewards speed alone, people will behave rationally: they will avoid uncertainty, write thinner rationales, and rush the cases you most need them to examine.

The [NIST AI Risk Management Framework Core](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/) calls for defined human-AI roles, explicit tasks, oversight responsibilities, and validation in context. You do not need an enterprise governance office to follow that principle. You need named owners, written boundaries, and evidence that the process works on your actual cases.

## What should you do before September 30?

**Use the remaining time to settle financial obligations and preserve the required data, then prove the replacement workflow on tasks that mirror production before MTurk disappears.** A migration is complete only when results reach the downstream system and someone can explain their quality.

Work through this checklist in order:

1. **Name an owner.** Give one person authority over the deadline, inventory, and cutover.
2. **Freeze new dependencies.** Stop adding MTurk to new workflows unless the work will finish before closure.
3. **Inventory every queue.** Include direct API calls, Ground Truth, Augmented AI, research accounts, scheduled jobs, and manual exports.
4. **Close open work.** Plan for submissions, approvals, rejections, bonuses, balances, and worker communication using Amazon's published dates.
5. **Export evidence.** Preserve transactions, task templates, qualifications, instructions, gold examples, decisions, and quality history before January 28, 2027.
6. **Classify the tasks.** Record sensitivity, expertise, ambiguity, impact, volume, latency, and seasonality.
7. **Choose a route per queue.** Automate, outsource, contract, hire, or combine them. Do not force one answer across unrelated work.
8. **Test privacy boundaries.** Remove confidential and personal data from public-workforce paths. Verify vendor access and retention terms for any replacement.
9. **Run both systems briefly.** Compare results on the same safe, representative cases. Investigate disagreement instead of averaging it away.
10. **Test rollback and escalation.** Know what happens when the new system fails, queue volume spikes, or reviewers cannot reach agreement.

Set internal cutover earlier than September 30. You need time to discover that an old cron job still creates HITs, a result parser expects an MTurk-specific field, or a finance account owns the only complete transaction export.

## How can Kit help you build the team?

**Kit can help you recruit and select reviewers with a structured, human-owned hiring process. It does not replace MTurk or run production annotation queues.** Keep that boundary clear when you plan the transition.

Use Kit to define the reviewer role, create hiring stages, schedule interviews, give candidates a representative assignment, and collect independent team reviews before the final decision. [Role templates](/templates) give you a starting structure, which you can adapt to the expertise and risk level of the queue.

Kit does not provide crowd labor, task distribution, inter-rater-agreement dashboards, or production quality management. Once someone joins, their daily review queue and calibration system belong in the specialist tools your team chooses. Kit's job ends where employment selection turns into operations.

That narrow role still matters. The MTurk shutdown gives you a deadline, but it should not pressure you into replacing an anonymous crowd with an improvised hiring decision. Define the work, test real judgment, compare reviewers against reference evidence, and keep a person accountable for the final choice.

If you are building a permanent reviewer team, you can [start a free Kit trial](/users/sign_up) and run that hiring process in one place. The larger migration remains yours: protect the data, measure the decisions, and design a human-in-the-loop function that can survive its next vendor change.