[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-ai-rater-work-from-home-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a remote, full-time role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation tasks such as evaluating outputs, preference ranking (RLHF-style), QA evaluation, validation against guidelines, and writing rationales to support training data quality and model performance improvement.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. Strong bilingual comprehension, analytical reasoning, and consistent annotation guidelines compliance are essential, and you will work within structured evaluation rubrics.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both German and English is required, including the ability to judge nuance, tone, and correctness in each language.","What languages are required?",{"A":20,"Q":21},"Domains vary by project and may include general knowledge, writing quality, instruction following, reasoning, and content safety labeling scenarios, all aimed at improving training data quality and overall model performance.","What domains are covered?","ai-rater-work-from-home-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based English\u002FGerman AI Generalist Trainers to support RLHF and large language model evaluation by assessing, ranking, and validating model outputs to improve training data quality and drive model performance improvement.","Germany-Based English & German AI Generalist Trainer (Remote, Full-Time) 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-Based English & German AI Generalist Trainer at Rexzone, you will contribute to AI\u002FLLM workflows by performing RLHF-style evaluation, prompt evaluation, and QA evaluation across English and German tasks. You will review model-generated responses, rank alternatives, write clear rationales, and validate outputs against annotation guidelines compliance and content safety labeling standards. Your work directly impacts training data quality, large language model evaluation outcomes, and model performance improvement.",{"h2":31,"desc":32},"Responsibilities","Evaluate and rank model-generated outputs in English and German using defined rubrics and annotation guidelines; perform RLHF-style preference ranking and prompt evaluation to identify the best responses; conduct QA evaluation to ensure training data quality, consistency, and annotation guidelines compliance; write concise, evidence-based rationales explaining evaluation decisions and reasoning; validate tasks for completeness, policy adherence, and content safety labeling requirements; flag ambiguous prompts, edge cases, and guideline gaps, and propose clarifications for improved labeling accuracy; track errors, patterns, and failure modes that affect large language model evaluation and downstream model performance improvement.",{"h2":34,"desc":35},"Basic Qualifications","Based in Germany and able to work remotely in a full-time schedule; fluent in German and English (reading, writing, and nuance in both languages); strong analytical skills with the ability to compare outputs and justify rankings with clear reasoning; exceptional attention to detail and consistency when following annotation guidelines; comfortable working with structured workflows, task queues, and quality targets; able to handle sensitive or safety-related content in line with content safety labeling policies.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience with data labeling, LLM evaluation, prompt evaluation, or QA evaluation; familiarity with RLHF concepts, preference ranking, and common LLM failure modes; experience applying annotation guidelines at scale and maintaining high training data quality; self-driven, reliable, and able to manage productivity and accuracy in a remote environment; interest in how evaluation decisions translate into model performance improvement.",{"h2":40,"desc":41},"Compensation","This role pays $35–$40 USD per hour (hourly).",{"h2":43,"desc":44},"How to Apply","Apply to Rexzone with a short summary of your bilingual English\u002FGerman experience, your location in Germany, and any relevant background in evaluation, QA, data labeling, or LLM workflows.","AI Data Operations"]