[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-data-labeling-jobs-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a Remote, Full-Time role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation, ranking (RLHF-style), QA evaluation, validation against guidelines, prompt evaluation, and write rationales that support training data quality and model performance improvement.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. We value strong analytical skills, attention to detail, and the ability to follow annotation guidelines compliance; training is provided for project-specific workflows.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in German and English is required, as you will evaluate and label content in both languages.","What languages are required?",{"A":20,"Q":21},"Domains can include general knowledge, customer-support style conversations, reasoning tasks, summarization, translation-style prompts, and content safety labeling, depending on project needs.","What domains are covered?","ai-data-labeling-jobs-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based AI Generalist Trainers (Remote, Full-Time) to support RLHF and large language model evaluation by judging, ranking, and validating model outputs. You will apply annotation guidelines compliance to improve training data quality, perform LLM evaluation and prompt evaluation, and write clear rationales that drive model performance improvement through consistent QA evaluation and training data quality checks.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will evaluate and improve AI systems by assessing model-generated responses across tasks and domains. Your work will focus on RLHF-style ranking, large language model evaluation, and training data quality improvements through structured rubrics, annotation guidelines compliance, and rigorous validation. You will collaborate asynchronously with distributed teams while maintaining high accuracy and consistency in bilingual (German\u002FEnglish) evaluation workflows.",{"h2":31,"desc":32},"Responsibilities","• Perform large language model evaluation by reviewing German and English model outputs for correctness, relevance, and safety.\n• Rank and compare multiple responses (RLHF-style) and select the best output using defined rubrics.\n• Write concise, evidence-based rationales that explain reasoning behind rankings and labels.\n• Execute QA evaluation, including spot checks, disagreement resolution, and error analysis to improve training data quality.\n• Validate labels against annotation guidelines compliance and document edge cases for guideline refinement.\n• Conduct prompt evaluation to identify ambiguous prompts and recommend improvements for more reliable model behavior.\n• Apply content safety labeling where required (toxicity, policy violations, sensitive content) and escalate high-risk items.\n• Track quality metrics, follow workflow instructions precisely, and meet productivity targets without sacrificing accuracy.",{"h2":34,"desc":35},"Basic Qualifications","• Must be based in Germany and authorized to work as a contractor\u002Femployee as applicable.\n• Fluency in German and English (reading, writing, and comprehension) for bilingual evaluation tasks.\n• Strong analytical skills with the ability to compare nuanced outputs and justify decisions.\n• High attention to detail and consistent adherence to annotation guidelines compliance.\n• Comfortable working independently in a remote environment with reliable internet access.\n• Ability to handle repetitive evaluation tasks while maintaining accuracy and training data quality standards.",{"h2":37,"desc":38},"Preferred Qualifications","• Prior experience in data labeling, QA evaluation, or content safety labeling.\n• Familiarity with LLM evaluation concepts, RLHF, and common failure modes of large language models.\n• Experience writing structured rationales, performing ranking tasks, or validating datasets.\n• Self-driven, organized, and proactive in raising guideline gaps and proposing improvements.\n• Background in linguistics, translation, journalism, technical writing, or related fields is a plus.",{"h2":40,"desc":41},"Compensation","USD $35–$40 per hour (Remote, Full-Time).",{"h2":43,"desc":44},"How to Apply","Apply to Rexzone with an updated resume\u002FCV highlighting bilingual (German\u002FEnglish) experience and any work in evaluation, annotation, QA, or AI-related workflows. Qualified applicants may be asked to complete a short skills assessment involving ranking, reasoning, and validation tasks.","AI Data Operations"]