Remote Data Labeling Jobs in Amsterdam

Rex.zone is hiring for remote data labeling jobs in Amsterdam focused on building high-quality training data for AI/ML systems. As a Data Labeling Specialist, you will follow annotation guidelines, perform QA evaluation, and support LLM training pipelines through RLHF, prompt evaluation, and content safety labeling. Your work directly improves training data quality, annotation guidelines compliance, and model performance improvement for NLP and computer vision use cases. Explore full-time remote opportunities with clear feedback loops, measurable quality metrics, and production-grade workflows used by AI labs, tech startups, and annotation vendors.

Job Image

Job Overview

Title: Remote Data Labeling Jobs in Amsterdam | Date: 25-02-2026 | Company: Rex.zone | Country: US | Remote Type: Remote | Employment Type: FULL_TIME | Experience Level: Mid-Senior | Industry: Technology | Job Function: Engineering | Skills: data labeling, data annotation, RLHF, prompt evaluation, QA evaluation, annotation guidelines, named entity recognition, NLP, computer vision annotation, content safety labeling, LLM training pipelines | Salary Currency: USD | Salary Min: 63360 | Salary Max: 126720 | Pay Period: YEAR

About the Role

You will deliver production-quality labeled datasets used in large language model evaluation and multimodal model training. This includes interpreting annotation guidelines, labeling text and images, performing second-pass review, and documenting edge cases to improve consistency. Typical workflows include named entity recognition, classification, summarization quality checks, prompt response evaluation, RLHF preference ranking, and content safety labeling. You will collaborate asynchronously with program managers and QA leads, triage ambiguous examples, and help refine rubric definitions to increase inter-annotator agreement and reduce label noise.

What You Will Do

Perform data labeling and data annotation across NLP and computer vision annotation tasks (classification, spans, bounding boxes, segmentation, ranking). Execute RLHF-style preference judgments and prompt evaluation to support LLM training pipelines and evaluation suites. Run QA evaluation checks: consistency audits, spot checks, rework loops, and error taxonomy reporting. Follow annotation guidelines compliance requirements and propose clarifications for ambiguous cases. Measure and improve training data quality using agreed metrics (accuracy, agreement, precision/recall proxies, defect rates). Label and review content safety labeling categories (policy-based safety, toxicity, self-harm, sexual content, violence, regulated goods). Document decisions, edge cases, and rubric updates so future labeling is repeatable and scalable. Support dataset versioning and change management so model performance improvement can be traced to data changes.

Required Qualifications

Mid-Senior experience in data labeling, data annotation, QA evaluation, or related data operations roles. Proven ability to follow annotation guidelines, maintain consistency, and deliver high throughput without quality regression. Hands-on familiarity with NLP tasks (named entity recognition, intent classification, sentiment, summarization evaluation). Comfort with computer vision annotation concepts (bounding boxes, polygons, segmentation masks) and review workflows. Experience with prompt evaluation and LLM evaluation rubrics (helpfulness, correctness, harmlessness, groundedness). Strong written communication for documenting label decisions and contributing to rubric iterations. Ability to work full-time in a remote environment with reliable connectivity and disciplined self-management.

Preferred Qualifications

Experience with RLHF, preference ranking, or pairwise comparison labeling at scale. Background in content safety labeling, trust & safety operations, or policy-based classification. Exposure to dataset sampling strategies, disagreement analysis, and error analysis for model performance improvement. Familiarity with annotation tooling, QA dashboards, and structured feedback loops.

Work Environment and Collaboration

Remote-first, asynchronous coordination with scheduled quality reviews and calibration sessions. Clear production expectations, review processes, and escalation paths for ambiguous labeling decisions. Work may support AI labs, tech startups, BPOs, and annotation vendors via Rex.zone programs.

How to Apply

Apply through Rex.zone with your resume and a brief summary of relevant data labeling and QA evaluation experience. Highlight examples of annotation guidelines compliance, training data quality improvements, or LLM evaluation work. Selected candidates may complete a short labeling calibration to verify rubric understanding and consistency.

Frequently Asked Questions

  • Q: Are these remote data labeling jobs in Amsterdam fully remote?

    Yes. Remote Type is Remote, and the workflows are designed for fully remote delivery with asynchronous collaboration and defined QA evaluation checkpoints.

  • Q: What kind of data labeling tasks will I work on?

    You may work on NLP labeling (named entity recognition, classification, summarization evaluation), computer vision annotation (bounding boxes, segmentation), RLHF preference ranking, prompt evaluation, and content safety labeling depending on project needs.

  • Q: Is this role full-time or contract/freelance?

    This posting is for FULL_TIME employment. Rex.zone may also host contract or freelance programs, but this role remains full-time as specified.

  • Q: What does QA evaluation mean in this role?

    QA evaluation includes review of labeled outputs for consistency, rubric compliance, error categorization, and rework loops that improve training data quality and downstream model performance improvement.

  • Q: Do I need experience with RLHF and LLM evaluation?

    Mid-Senior candidates are expected to be comfortable with structured evaluation rubrics. Direct RLHF experience is preferred, but strong prompt evaluation and annotation guidelines compliance experience can also be relevant.

  • Q: Which domains does the role support?

    Common domains include NLP, computer vision annotation, content safety labeling, and LLM training pipelines for AI labs, tech startups, and annotation vendors.

  • Q: What skills should I include to match this job intent?

    Include data labeling, data annotation, RLHF, prompt evaluation, QA evaluation, annotation guidelines, named entity recognition, NLP, computer vision annotation, content safety labeling, and LLM training pipelines.

230+Domains Covered
120K+PhD, Specialist, Experts Onboarded
50+Countries Represented

Industry-Leading Compensation

We believe exceptional intelligence deserves exceptional pay. Our platform consistently offers rates above the industry average, rewarding experts for their true value and real impact on frontier AI. Here, your expertise isn't just appreciated - it's properly compensated.

Work Remotely, Work Freely

No office. No commute. No constraints. Our fully remote workflow gives experts complete flexibility to work at their own pace, from any country, any time zone. You focus on meaningful tasks - we handle the rest.

Respect at the Core of Everything

AI trainers are the heart of our company. We treat every expert with trust, humanity, and genuine appreciation. From personalized support to transparent communication, we build long-term relationships rooted in respect and care.

Ready to Shape the Future of AI Data Operations?

Apply Now.