[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-junior-ai-evaluator-jobs-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a remote, full-time role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation tasks such as ranking model responses, running QA evaluation checks, validating labels for training data quality, applying annotation guidelines compliance, and writing rationales that explain your reasoning.","What tasks will I do?",{"A":14,"Q":15},"AI experience is helpful but not required. Rexzone provides task guidelines; success depends on strong analytical skills, attention to detail, and consistent evaluation quality.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both English and German is required because you will evaluate and write content in both languages.","What languages are required?",{"A":20,"Q":21},"Domains vary by project and may include general knowledge, reasoning, helpfulness, instruction-following, safety, and content quality—supporting RLHF and model performance improvement.","What domains are covered?","junior-ai-evaluator-jobs-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based English\u002FGerman AI Generalist Trainers to support RLHF and large language model evaluation by ranking, QA-checking, and validating model outputs. You will apply annotation guidelines compliance to improve training data quality and drive model performance improvement across real-world AI\u002FLLM workflows.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based, remote AI Generalist Trainer at Rexzone, you will evaluate and improve AI systems by assessing model-generated responses in English and German. Your work supports RLHF, large language model evaluation, and training data quality by performing structured ranking, QA evaluation, and rationale writing that enables measurable model performance improvement.",{"h2":31,"desc":32},"Responsibilities","Perform large language model evaluation by reviewing, comparing, and ranking model outputs; execute QA evaluation for accuracy, consistency, and policy alignment; write clear rationales explaining reasoning behind rankings and corrections; validate labeled data against annotation guidelines compliance and project rubrics; identify edge cases, ambiguity, and failure modes to improve training data quality; apply prompt evaluation and content safety labeling where required; track issues and provide feedback to improve workflows and model performance improvement; maintain productivity and quality targets while working independently in a remote setting.",{"h2":34,"desc":35},"Basic Qualifications","Must be based in Germany; fluent in English and German (reading, writing, and comprehension); strong analytical skills with the ability to evaluate nuanced content and follow rubrics; exceptional attention to detail with consistent annotation guidelines compliance; ability to explain reasoning clearly in written rationales; reliable internet and ability to work full-time remotely.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience in data labeling, QA evaluation, or content review; familiarity with RLHF, prompt evaluation, and LLM evaluation concepts; comfort working with ambiguous tasks and iterating with feedback; self-driven, organized, and able to meet quality standards with minimal supervision; experience with content safety labeling and training data quality best practices.",{"h2":40,"desc":41},"Compensation","USD $35–$40 per hour (hourly). Exact rate is based on skills, performance on assessments, and project requirements.",{"h2":43,"desc":44},"How to Apply","Apply through Rexzone with an up-to-date resume\u002FCV highlighting bilingual English\u002FGerman expertise and any experience in evaluation, QA, data labeling, or AI\u002FLLM workflows. Qualified applicants may be asked to complete a short evaluation task focused on ranking, reasoning, and validation.","AI Data Operations"]