[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"job-entry-level-ai-evaluator-jobs-germany":3},{"Ques":4,"Slug":22,"Header":23,"job_category":45},{"title":5,"content":6},"Frequently Asked Questions",[7,10,13,16,19],{"A":8,"Q":9},"Yes. This is a full-time remote role, and you must be based in Germany.","Is this role remote?",{"A":11,"Q":12},"You will perform large language model evaluation including RLHF-style ranking, prompt evaluation, QA evaluation, validation, and writing bilingual rationales to improve training data quality and model performance improvement.","What tasks will I do?",{"A":14,"Q":15},"AI experience is preferred but not required. Strong analytical skills, attention to detail, and the ability to follow annotation guidelines compliance are essential.","Do I need AI experience?",{"A":17,"Q":18},"Fluency in both English and German is required for reading, writing, and producing evaluation rationales.","What languages are required?",{"A":20,"Q":21},"Domains can include general knowledge, customer-support style prompts, summarization, reasoning, safety-sensitive content, and other tasks related to content safety labeling and training data quality for LLM evaluation.","What domains are covered?","entry-level-ai-evaluator-jobs-germany",{"desc":24,"title":25,"content":26},"Rexzone is hiring Germany-based, bilingual (English\u002FGerman) AI Generalist Trainers to support AI\u002FLLM workflows through RLHF, large language model evaluation, and training data quality improvements. You will evaluate, rank, and QA model outputs, write clear rationales, and follow annotation guidelines compliance to drive model performance improvement. This full-time remote role focuses on LLM evaluation, data labeling, prompt evaluation, QA evaluation, and content safety labeling to ensure reliable training data quality and consistent large language model evaluation.","Germany-Based English & German AI Generalist Trainer 2026 May",[27,30,33,36,39,42],{"h2":28,"desc":29},"About the Role","As a Germany-based English & German AI Generalist Trainer at Rexzone, you will improve AI system behavior by performing large language model evaluation tasks, including RLHF-style ranking, prompt evaluation, and QA evaluation. Your work directly impacts training data quality and model performance improvement by applying annotation guidelines compliance, validating edge cases, and providing high-quality reasoning and rationales in both English and German.",{"h2":31,"desc":32},"Responsibilities","Evaluate and rank model-generated responses for helpfulness, factuality, reasoning quality, and policy adherence; Perform QA evaluation to validate labeling accuracy, consistency, and training data quality; Write clear rationales explaining ranking decisions and reasoning in English and German; Apply annotation guidelines compliance and update notes when ambiguity or edge cases are discovered; Validate prompts and outputs for content safety labeling and escalate policy-relevant issues; Review disagreements, resolve conflicts through evidence-based reasoning, and suggest improvements to evaluation rubrics; Track recurring error patterns to support model performance improvement and reliable large language model evaluation.",{"h2":34,"desc":35},"Basic Qualifications","Based in Germany and able to work remotely full-time; Fluent in English and German (professional reading and writing required); Strong analytical skills with the ability to compare outputs, detect subtle errors, and justify decisions; High attention to detail and consistency when following annotation guidelines compliance; Comfortable working with AI\u002FLLM workflows, including evaluation, ranking, QA, reasoning, and validation tasks.",{"h2":37,"desc":38},"Preferred Qualifications","Prior experience with data labeling, prompt evaluation, QA evaluation, or content safety labeling; Familiarity with RLHF concepts and large language model evaluation; Experience interpreting rubrics, writing structured rationales, and improving training data quality; Self-driven, reliable, and able to manage throughput while maintaining accuracy; Interest in how evaluation feedback supports model performance improvement.",{"h2":40,"desc":41},"Compensation","USD $35–$40 per hour, full-time, remote. Pay is hourly and depends on experience and demonstrated evaluation quality.",{"h2":43,"desc":44},"How to Apply","Apply through Rexzone with a brief summary of your bilingual (English\u002FGerman) experience and any relevant AI evaluation, QA, or annotation work. Selected candidates may be asked to complete a short skills assessment focused on ranking, reasoning, and annotation guidelines compliance.","AI Data Operations"]