You'll hear most of these terms in your first 90 days. Some have settled meanings across the industry. Others vary platform to platform. Where there's variation, we've given you the most common usage.

This glossary combines the foundational terms from The Intellectual Side Hustle with additional terminology observed across thousands of job listings on the major human data platforms. Treat it as a living reference — we add new terms as the industry's vocabulary continues to settle.

A
AI Safety / Alignment

The field of ensuring AI systems behave in ways that are helpful, harmless, and honest. Much of human data work — particularly RLHF, red teaming, and evaluation — exists specifically to improve AI safety. When a platform says it's working on "alignment," it means training models to follow human values and intentions.

You'll see this in: job titles like "AI Safety Evaluator," project descriptions mentioning "alignment research"

AI Trainer / AI Tutor

The most common job title in the industry. An AI Trainer provides the human input — writing, rating, correcting, evaluating — that helps AI models learn and improve. "AI Tutor" is a newer variation that emphasises the teaching aspect: you're not just labelling data, you're teaching a model how to think. The two titles are largely interchangeable across platforms.

Appears in: 1,300+ job listings across Outlier, Appen, Remotasks, Scale AI, and others

Annotation / Labelling

Marking data with useful information so it can be used to train or evaluate AI systems. This may include tagging sentiment, identifying entities, categorising content, highlighting errors, or assigning labels based on project guidelines.

From The Intellectual Side Hustle by Mucha Murapa

Attempter

A contributor who takes on a task attempt. On some platforms, you "attempt" a task rather than being assigned one — meaning you choose to start it, and your work is then reviewed. If your attempt doesn't meet quality standards, the task may be reassigned. The term reflects the try-and-review model that many platforms use.

You'll see this in: platform dashboards, task queues, contributor guidelines

Average Handling Time (AHT)

The target or average time expected to complete a single task from start to submission. AHT often improves as contributors become more familiar with the workflow, guidelines, and project requirements.

From The Intellectual Side Hustle by Mucha Murapa

C
Calibration

The process of helping multiple contributors interpret guidelines and rubrics in the same way. A team may review the same sample task, compare decisions, and discuss why they scored something differently — so that everyone working on the project is judging to the same standard.

From The Intellectual Side Hustle by Mucha Murapa

Contributor / Tasker

Generic terms for anyone who completes work on a human data platform. "Contributor" is the most common — used by Outlier, Remotasks, and others. "Tasker" appears on some platforms. Both mean the same thing: you're the human providing the data, feedback, or expertise that trains the AI.

You'll see this in: platform onboarding, payment dashboards, contributor agreements

Conversation / Multi-Turn

A task format where you engage in a back-and-forth dialogue with an AI model, playing a specific role or persona. Multi-turn tasks test whether the model can maintain context, follow instructions, and produce coherent responses across several exchanges. These tasks typically pay more than single-turn evaluations because they require sustained attention and consistency.

You'll see this in: Outlier projects, RLHF tasks, chatbot evaluation work

D
Data Collection

Gathering raw data — images, audio recordings, text samples, video clips — that will be used to train AI models. Unlike annotation (which adds labels to existing data), data collection creates the data itself. You might be asked to photograph objects, record yourself speaking, or write sample texts on specific topics.

You'll see this in: Appen, Telus, Prolific task listings

Domain Expert / SME (Subject Matter Expert)

A professional with deep knowledge in a specific field — law, medicine, finance, engineering, science — who applies that expertise to AI training tasks. Domain experts command the highest pay rates in the industry (often £50–£200+ per hour) because their knowledge is rare and difficult to replicate. SME is the corporate abbreviation you'll see in job listings.

Appears in: 600+ job listings, particularly on micro1, Mercor, AfterQuery, and Ethos

E
Evaluation (Eval)

A structured process used to measure the quality, performance, or alignment of an AI output against defined criteria. For example, an eval might ask, "Which of these two answers is better, and why?" You'll hear this term used informally as "eval" in almost every project briefing.

From The Intellectual Side Hustle by Mucha Murapa

G
Generalist

A human data worker who handles a broad range of tasks rather than specialising in one domain. Generalists are the backbone of many platforms — they evaluate search results, rate AI responses, label images, and complete whatever tasks are available. Pay is typically lower than specialist work, but availability is higher and more consistent.

Appears in: 116 job listings, common on Outlier, Remotasks, and Appen

Ground Truth

The expected or accepted correct answer used to compare against human or AI outputs. It may come from an expert answer, an official source, a gold-standard dataset, or a reviewer-approved response. Disagreements between expert answers and ground truth are themselves valuable signal in AI training, so a "wrong" answer in your work isn't always penalised — sometimes it's the data the platform was looking for.

From The Intellectual Side Hustle by Mucha Murapa

Guidelines

The written instructions explaining how a task should be completed. They usually define what to include, what to avoid, how to apply the rubric, and how to handle unclear or unusual cases. Read them slowly. Re-read them. Most rejected work is rejected for guideline drift, not for poor reasoning.

From The Intellectual Side Hustle by Mucha Murapa

H
Human Data Expert (HDE)

A subject-matter expert or skilled professional who provides high-level, nuanced input to train, refine, and stress-test AI models. Unlike traditional data labellers who perform simpler, repetitive work, Human Data Experts apply deep, specialised knowledge to create, review, or evaluate high-quality training data that AI cannot reliably produce on its own. The term was coined by Ali Ansari, CEO of micro1, and has since been adopted across the industry. HDEs are typically paid £50 to £1,000 per hour, depending on the rarity and depth of their domain expertise.

From The Intellectual Side Hustle by Mucha Murapa

Human Data Manager (HDM)

Oversees experts and reviewers during a project. They may support onboarding, monitor project progress, handle first-line escalations, and act as a link between the Human Data Platform's senior team and the Human Data Experts.

From The Intellectual Side Hustle by Mucha Murapa

Human Data Platform (HDP)

A company that supplies the human expertise needed to create, review, evaluate, or improve training data for AI models. Sometimes called a Human Data Provider. Examples include micro1, Outlier, Mercor, Surge AI, and Toloka.

From The Intellectual Side Hustle by Mucha Murapa

Human Data Reviewer (HDR)

Usually a senior or high-performing domain expert who reviews work completed by other experts. They check for content accuracy, reasoning quality, guideline adherence, and overall consistency.

From The Intellectual Side Hustle by Mucha Murapa

J
Judgement

A subjective evaluation or rating made using project guidelines. Unlike some annotation tasks, which may have a clear right or wrong answer, judgement tasks ask contributors to assess quality, usefulness, relevance, or appropriateness. For example: "Rate this search result as Excellent, Good, or Bad."

From The Intellectual Side Hustle by Mucha Murapa

L
Large Language Model (LLM)

The AI systems that human data workers help train. ChatGPT, Claude, Gemini, and Llama are all LLMs. They're called "large" because they're trained on enormous datasets and have billions of parameters. When you do RLHF work, evaluation, or red teaming, you're directly shaping how these models behave.

You'll see this in: nearly every job description in the industry

Localisation / Translation / Transcription

Three related but distinct tasks. Translation converts text from one language to another. Localisation goes further — adapting content for cultural context, not just language. Transcription converts spoken audio into written text. All three are in high demand as AI companies expand their models to work across languages and cultures.

Appears in: 70+ job listings, particularly on Appen, Telus, Welocalize, and RWS

M
Multimodal

AI models that work across multiple types of input — text, images, audio, video, code. When a job listing mentions "multimodal evaluation," it means you'll be assessing AI outputs that combine different media types. For example, evaluating whether an AI correctly describes what's happening in a video, or whether it generates an appropriate image from a text prompt.

You'll see this in: advanced evaluation roles, typically paying higher rates

O
Onboarding

The process of getting set up on a new platform or project. This typically includes creating your profile, verifying your identity, completing qualification assessments, reading project guidelines, and doing practice tasks. Good onboarding can take anywhere from a few hours to several days. Don't rush it — your onboarding performance often determines which projects you're offered and at what pay rate.

You'll see this in: every platform's getting-started flow

P
Persona / Character

A role you're asked to play during conversation tasks. A platform might ask you to "act as a first-year medical student asking questions about pharmacology" or "pretend to be a customer complaining about a delayed delivery." Persona tasks test whether the AI can handle realistic, varied interactions. The better you inhabit the persona, the more useful the training data.

Appears in: 24 job listings, common in RLHF and conversation tasks

Prompt / Prompt Engineering

A prompt is the input you give to an AI model — the question, instruction, or scenario that triggers a response. Prompt engineering is the skill of crafting prompts that produce the best possible outputs. In human data work, you might be asked to write prompts that test the model's limits, or to evaluate how well a model responds to different prompt styles. It's both a task type and a career path.

Appears in: 58+ job listings, with dedicated "Prompt Engineer" roles paying $40–$100/hr

Q
Quality Assurance (QA)

The process of checking completed work for accuracy, consistency, and adherence to project guidelines. QA helps ensure that human data work is reliable enough to be used in AI training, evaluation, or model improvement. QA roles — Peer Reviewer, QA Reviewer, Senior Reviewer — are themselves a tier of human data work, and often pay at higher rates than the underlying expert work.

From The Intellectual Side Hustle by Mucha Murapa

R
Rater / Rating

A person who evaluates AI outputs by assigning scores or rankings. "Search Quality Rater" is one of the oldest roles in the industry — originally created by Google to evaluate search results. Today, raters work across all types of AI output: text, code, images, and more. Rating tasks typically involve choosing between two or more AI responses and explaining which is better and why.

Appears in: 97 job listings, particularly on Appen, Telus, and Scale AI

Red Teaming

Deliberately trying to break, trick, or expose weaknesses in an AI model. Red teamers craft adversarial prompts designed to make the model produce harmful, biased, incorrect, or inappropriate outputs. The goal isn't to be malicious — it's to find the vulnerabilities before real users do. Red teaming is one of the highest-paying task types in human data work because it requires creativity, domain knowledge, and an understanding of how models fail.

Appears in: 17 job listings, typically paying $50–$200/hr on platforms like Outlier, Scale AI, and HumanSignal

Reinforcement Learning From Human Feedback (RLHF)

The technique used to help AI models better align with human preferences. Humans review, rank, or rate AI outputs, and that feedback guides the model toward responses that are more helpful, accurate, safe, or appropriate. RLHF is the technique that turned raw language models into the conversational AI assistants people now use daily — and it's the work many Human Data Experts are paid to do.

From The Intellectual Side Hustle by Mucha Murapa

Rubric

The platform's checklist of criteria you're evaluating against on a given task — what counts as accurate, helpful, safe, complete, or well-reasoned in this specific context. Every serious platform uses rubrics. If a platform asks you to rate or evaluate AI outputs without giving you a rubric, that's a flag.

From The Intellectual Side Hustle by Mucha Murapa

S
Search Quality Rater

A specialised rater who evaluates the relevance and quality of search engine results. This role predates the current AI training boom — Google has employed search quality raters since the mid-2000s. Today, the role has expanded to include evaluating AI-generated search summaries, featured snippets, and conversational search results. It remains one of the most accessible entry points into human data work.

Appears in: 58 job listings, primarily on Appen and Telus

Side-by-Side Comparison (SxS)

A task format where you're shown two AI-generated responses to the same prompt and asked to judge which is better. SxS comparisons are fundamental to RLHF — they're how models learn human preferences. You'll typically rate on multiple dimensions: accuracy, helpfulness, safety, and writing quality. The "side-by-side" format is so common that platforms often abbreviate it to "SxS" in project names.

You'll see this in: RLHF projects, model evaluation tasks

Supervised Fine-Tuning (SFT)

The process of training an AI model on high-quality example responses written by humans. In SFT tasks, you write the "ideal" response to a given prompt — showing the model exactly what a good answer looks like. This is different from RLHF (where you rate existing responses) because in SFT, you're creating the training data from scratch. SFT tasks typically pay well because they require both expertise and strong writing skills.

You'll see this in: expert writing tasks, "response generation" projects, GoAGI and AfterQuery listings

Synthetic Data

Data that's artificially generated rather than collected from real-world sources. In human data work, you might be asked to create synthetic conversations, write fictional but realistic scenarios, or generate example data that mimics real patterns. Synthetic data is increasingly important because it can be produced at scale without privacy concerns — but it still needs human oversight to ensure quality and realism.

You'll see this in: data generation tasks, conversation creation projects

T
Task

The individual unit of work you complete. This might include reviewing one document, labelling one image, rating one AI answer, or answering one question about a piece of content. Performance is often measured by both task quality and task completion rate.

From The Intellectual Side Hustle by Mucha Murapa

The foundational terms in this glossary are adapted from The Intellectual Side Hustle by Mucha Murapa, with additional terms identified from analysis of 3,200+ job listings across 40 human data platforms. This is a living document — new terms are added as the industry's vocabulary continues to settle.

© 2026 Mucha Murapa / Train AI Media™. All rights reserved.